arXiv Papers with Code in Computer Vision (January 2026 - June 2026)
Authors:Kartik Narayan, Vishal M. Patel
Abstract:
Low‑resolution face recognition (LR‑FR) remains a challenging task due to poor feature extraction and aggregation, as probe images often contain limited identity information resulting from extreme degradations such as blur, occlusion, and low contrast. Additionally, the domain gap between high‑resolution (HR) gallery images and low‑resolution (LR) probe images poses a significant challenge. A single feature encoder struggles to generalize effectively across both domains when fine‑tuned on an LR dataset, and this issue is further magnified by catastrophic forgetting. To address these challenges, we propose FaceMoE, an effective adaptation of Mixture of Experts (MoE) transfomer architecture for low‑resolution face‑recognition . Specifically, we introduce multiple specialized feed‑forward network (FFN) experts and incorporate a top‑k router, which dynamically assigns tokens to appropriate experts. This design emergently promotes specialization across experts for different semantic regions of the face, which enables FaceMoE to perform resolution‑aware feature extraction. Moreover, the top‑k router facilitates sparse expert activation, enabling the model to preserve pretrained knowledge when finetuned on a LR dataset, while increasing model capacity without proportional computational overhead. FaceMoE is trained with a combined face recognition loss, router z‑loss, and load balancing loss to ensure expert specialization and stable training. To the best of our knowledge, this is the first work leveraging MoE for LR‑FR. Extensive experiments across eleven datasets, spanning HR, mixed‑quality, and LR benchmarks, demonstrate that FaceMoE significantly outperforms state‑of‑the‑art methods. Code: https://github.com/Kartik‑3004/FaceMoE
Authors:Yujie Guo, Yudong Jin, Lingteng Qiu, Zehong Shen, Zhen Xu, Jing Zhang, Xianchao Shen, Hujun Bao, Sida Peng, Xiaowei Zhou
Abstract:
Producing 3D human representations from input views on the fly is essential for immersive live streaming systems, where representation compactness is as critical as high fidelity given limited computational power and transmission bandwidth. Although recent feed‑forward reconstruction methods achieve impressive quality through the view‑centric prediction of 3D representations, they repeatedly encode the same subject content across multiple views, leading to significant inter‑view redundancy. Our key insight is to perform predictions directly in 3D space, enabling the network to learn and produce a highly compact representation. To this end, we propose PointSplat, a novel human‑centric approach that directly infers Gaussian primitives from an input point set. The proposed method first estimates a coarse geometric proxy and performs ray casting to prune redundant points and establish explicit 2D‑‑3D correspondences. Subsequently, it employs a Point‑Image Transformer to fuse appearance and geometry features, predicting Gaussian attributes in a single forward pass. This design restricts predictions to foreground regions of interest, substantially reducing the total number of Gaussians while improving novel‑view rendering quality. Extensive experiments demonstrate that PointSplat achieves higher efficiency and quality while exhibiting strong robustness to variations in view count and image resolution across multiple datasets.
Authors:Or Hirschorn, Aaron Olender, Eli Alshan, Ianir Ideses, Lior Fritz, Sagie Benaim
Abstract:
We present a zero‑shot, training‑free and optimization‑free framework for generating 360 panoramic images and videos by directly injecting spherical priors into pre‑trained diffusion transformers. Existing methods either rely on costly fine‑tuning on scarce panoramic data that limits generalization, or leverage multi‑step optimization that incurs prohibitive inference latency. We observe that contemporary generative models natively exhibit some panoramic priors from large‑scale training. However, these emergent capabilities are insufficient, as the models fundamentally fail to satisfy the rigorous topological constraints imposed by equirectangular projection (ERP). We introduce a zero‑shot and optimization‑free approach that resolves these constraints at inference time. Spherical RoPE replaces standard rotary position embeddings: low‑frequency channels are re‑parameterized as 3D Cartesian coordinates to natively encode the spherical manifold, while high‑frequency channels are harmonically quantized to enforce exact periodicity. Coupled with complementary Semantic Distortion classifier‑free guidance (CFG) that explicitly steers geometry, we avoid retraining and inherit the full creative breadth of state‑of‑the‑art models. Our approach generalizes across diverse backbones and 360 generation modalities. We demonstrate this across text‑to‑panorama using Flux.1, Flux.2, and LTX‑Video backbones, achieving competitive performance against baselines, all while remaining training‑free. Project page: https://orhir.github.io/SpheRoPE
Authors:Sanghyuk Chun, William Yang, Amaya Dharmasiri, Olga Russakovsky
Abstract:
Uncertainty estimation has been a long‑standing challenge in AI models; it amounts to "knowing what you don't know," and metacognition is notoriously difficult even for humans (cf. the Dunning‑Kruger effect). Although it is still far from solved even in simpler classification systems, tackling it in multimodal large language models (MLLMs) is becoming increasingly important. Within MLLMs, uncertainty can stem from any of the diverse sources as well as from their relationships, and further can stem from the unbounded answers in the open‑ended setting. To tackle the issues, we propose CoMet, an MLLM uncertainty estimation method by decomposing uncertainty into a context‑specific term and a multiplicity‑specific term. The former captures ambiguity induced by the given context (e.g., task or prompt), while the latter captures how many plausible answers determined by the context remain compatible with the given input. We train a lightweight post‑hoc uncertainty module to estimate these quantities, which enables efficient uncertainty estimation without autoregressive answer generation or repeated sampling. Experiments on various open‑ended multimodal benchmarks, hallucination detection, and multiple‑choice visual question answering benchmarks show that CoMet consistently improves uncertainty estimation over existing baselines while remaining efficient in practice. Code is available at https://github.com/princetonvisualai/comet_uncertainty
Authors:Lianyu Hu, Shengqian Qin, Zeqin Liao, Qing Guo, Liang Wan, Wei Feng, Yang Liu
Abstract:
Chain‑of‑thought (CoT) reasoning has enabled multi‑modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit intermediate reasoning steps in natural language. However, this text‑based reasoning paradigm is inherently slow at inference time with even thousands of tokens and fundamentally constrained by the expressiveness of natural language. In this paper, we propose CoLT, (Chain of Latent Thoughts), a novel framework that teaches multi‑modal models to reason through a chain of latent thought representations instead of verbose text tokens, which can perform thinking with as few as 3 steps. Naively forcing the model to think with latent states easily produces meaningless semantics and makes training unstable. To effectively regulate the latent reasoning process, we introduce a lightweight external decoder that provides step‑level supervision for each latent reasoning step in two complementary directions: a forward mode that decodes latent thoughts into the textual reasoning of the next step, and a backward mode that aligns decoder hidden states with the model's latent thoughts given preceding textual context. We further incorporate internal supervision that encourages coherent step‑by‑step latent transitions. The decoder and internal supervision are removed during inference to maintain high efficiency of latent reasoning. Extensive experiments on eight benchmarks demonstrate that CoLT not only outperforms existing latent reasoning methods such as CODI and SIM‑CoT, but also surpasses latent visual reasoning approaches that rely on auxiliary images with costly annotation requirements. Compared to text CoT methods, CoLT can notably reduce the inference time by 10.1× and text decoding time by 22.6×. Code is released at https://github.com/hulianyuyy/CoLT.
Authors:Yuhao Wang, Mu Qiao, Haiwen Diao, Yunzhi Zhuge, Pingping Zhang, Xindong Zhang, Lei Zhang, Huchuan Lu
Abstract:
Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token sequences. Training‑free visual token reduction provides an efficient solution. However, existing methods distort attention distributions, giving rise to a phenomenon we term Attention Logit Collapse. To address this issue, we propose ERA, an Entropy‑guided visual token pruning framework with Rectified Attention for efficient MLLMs. Specifically, ERA comprises three crucial components: Dual‑view Entropy Pruning (DEP), Bias‑aware Token Recycling (BTR), and Logit‑preserving Attention Rectification (LAR). First, DEP identifies representative anchor tokens by jointly modeling visual diversity and head‑wise saliency. BTR then recycles pruned tokens into their corresponding anchors while estimating a cluster‑level logit bias. Building upon this, LAR injects the estimated bias into attention logits, effectively rectifying the collapse induced by token reduction. Together, these components preserve visual evidence even under aggressive compression, enabling robust performance across single‑image, multi‑image, and video settings on a wide range of MLLMs. Beyond delivering practical acceleration, ERA establishes logit‑preserving visual token pruning as a principled framework for efficient MLLMs, unifying theoretical foundation, algorithmic design, and practical deployment. The code is at https://github.com/924973292/ERA.
Authors:Peng Li, Rawal Khirodkar, Junxuan Li, Yuan Dong, Chen Cao, Yuan Liu, Wenhan Luo, Yike Guo, Shunsuke Saito
Abstract:
Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressivity and often introduce artifacts due to imperfect fitting. We propose LUNA, an LBS‑free universal neural animation model that directly maps multiple 2D controls like images, keypoints, sketches, and unseen characters into 3D Gaussian deformations, bypassing explicit body fitting. At its core, a transformer‑based motion regressor disentangles global rigid motion from fine‑grained local dynamics to capture both coherent movement and subtle non‑rigid effects. To resolve the inherent ambiguity of 2D‑to‑3D lifting while scaling beyond fitted datasets, we introduce hybrid supervision that distills soft structural priors from an LBS teacher and a loss that supports training on both limited fitted data and large in‑the‑wild unlabeled videos. Extensive experiments show LUNA achieves competitive visual fidelity compared to LBS‑based approaches, while delivering realistic human motion and zero‑shot cross‑identity generalization across diverse driving modalities. To the best of our knowledge, LUNA is the first end‑to‑end 3D animatable model that supports implicit 2D driving.
Authors:Qingyun Liu, Jiwen Zhang, Jingyi Hu, Siyuan Wang, Zhongyu Wei
Abstract:
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBench, a multimodal embodied cooperation benchmark with an evaluation platform spanning diverse real‑world tasks, two cooperation structures, and three collaboration modes. Through extensive experiments across various MLLMs, we summarize three key findings: (i) Collaboration generally improves embodied task completion, but its benefits depend on balancing collaborative gains against coordination complexity. (ii) Communication is essential to collaboration gains, while the best collaboration mode depends on team size and model capability. (iii) Moreover, collaboration improves robustness under noisy priors and exploration conditions. Generally, MECoBench provides a systematic testbed for understanding the mechanisms and limits of multimodal embodied collaboration. Code and dataset are available at https://github.com/q‑i‑n‑g/MECoBench.
Authors:Xinyu Hou, Xiaoming Li, Zongsheng Yue, Chen Change Loy
Abstract:
Depth‑of‑field control is a fundamental tool in photography, yet post‑capture bokeh editing from a single image remains challenging. A practical editor should handle images captured under arbitrary focus and aperture settings. Existing methods typically assume an all‑in‑focus input, or first recover an all‑in‑focus image before rendering new bokeh. Such pipelines can discard useful blur cues from the source image and propagate reconstruction artifacts into the final edit. We introduce AnyBokeh, a physics‑guided framework for any‑to‑any bokeh editing. Instead of treating source blur merely as a degradation to be removed, AnyBokeh estimates the source blur state with a signed circle‑of‑confusion map and a disparity map. By modeling the linear relation between signed circle of confusion and disparity difference, AnyBokeh estimates a source‑specific optical fingerprint and transfers the source optical characteristics to the desired focus and aperture setting. A generative editor conditioned on both source and target circle‑of‑confusion maps then performs relative blur synthesis, enabling spatially adaptive deblurring, preservation, and defocus rendering. To support physically supervised learning, we further construct a high‑fidelity synthetic dataset with accurate depth, focus distance, and full EXIF metadata. Experiments on real‑world benchmarks show that AnyBokeh achieves faithful and controllable editing across any‑to‑any bokeh editing, all‑in‑focus‑to‑bokeh rendering, and defocus deblurring, while avoiding all‑in‑focus reconstruction and test‑time bokeh‑level calibration commonly required by existing approaches. The code and dataset will be available at https://github.com/itsmag11/AnyBokeh.
Authors:Hubert Dymarkowski, Xingjian Fu, Rappy Saha, Jude Haris, José Cano
Abstract:
Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers. This heterogeneity leads to significant variation in tensor shapes, requiring flexible and efficient FPGA‑based acceleration. In this paper, we present FlexViT, a reconfigurable FPGA accelerator for efficient ViT inference on resource‑constrained edge devices. Built on the SECDA‑TFLite framework, FlexViT employs a hardware‑software co‑design approach that maps both fully connected and convolutional layers onto a unified high‑throughput INT8 GEMM engine using a runtime im2col transformation. To efficiently support diverse layer configurations, we propose a dual‑mode dataflow that dynamically switches between input and weight reuse by reconfiguring the compute array at runtime. We further introduce a depth‑first tiling strategy that completes accumulation in a single pass, eliminating off‑chip partial‑sum transfers and reducing memory bandwidth requirements. We implement FlexViT on a PYNQ‑Z2 FPGA and evaluate it across a representative set of ViT models. FlexViT achieves up to 2.74x speedup on accelerator‑executed layers, translating into up to 1.40x end‑to‑end speedup compared to CPU‑only execution. The code is available at: https://github.com/gicLAB/FlexViT
Authors:Haojian Huang, Harold Haodong Chen, Meng Luo, Junjia Du, Shanqing Xu, Ziheng Chen, Yanxiang Huang, Yinchuan Li, Ying-Cong Chen
Abstract:
We introduce VidPair‑Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. Unlike previous benchmarks that primarily rely on text‑based perturbations or adversarial questions while neglecting the consistency of visual backgrounds, VidPair‑Halluc features video pairs with highly similar backgrounds but distinctly different foreground semantics, enabling precise attribution of model errors to genuine hallucination rather than background variation. The benchmark is constructed through PairFlow, a pipeline that leverages recent advances in text‑to‑image and video generation to systematically compose stories, generate coherent video clips, and assemble them into adversarial pairs. Covering both spatial and temporal reasoning across ten semantic aspects, VidPair‑Halluc comprises 1K high‑quality adversarial video pairs and 11K spatio‑temporal QA pairs with control over background and foreground variations. Evaluations on mainstream LVMs show persistent difficulty with robust fine‑grained video understanding in adversarial settings, and code and data are available at the https://jethrojames.github.io/VidPair‑Halluc/.
Authors:Junzhe Jiang, Zipei Ma, Zijie Pan, Li Zhang
Abstract:
A pivotal step in autonomous driving simulation involves inserting foreground vehicles with predefined trajectories into simulated scenes. This process enhances scene diversity and facilitates the creation of various corner cases for testing and improving autonomous driving models. However, existing methods often rely on pre‑reconstructed 3D assets, which frequently lead to lighting inconsistencies between the inserted foreground and the background. Moreover, the reliance on limited, manually‑curated 3D assets hinders large‑scale deployment. To address these challenges, we propose DriveWeaver, a novel framework for controllable vehicle insertion in autonomous driving simulation. Specifically, for a masked target insertion area, DriveWeaver performs video inpainting conditioned on vehicle point clouds to generate high‑quality, temporally consistent vehicles. This video‑inpainting‑based approach ensures seamless blending between the foreground and background, while the readily available point cloud conditions enable superior generalization. To support long‑term generation, we further design a global‑to‑local hierarchical inpainting strategy, ensuring the consistent identity and appearance of the inserted vehicles. Meanwhile, we extract explicit 3D Gaussian representations of the inserted vehicles through an urban reconstruction pipeline to enable real‑time rendering for autonomous driving simulation. Extensive experiments across diverse datasets demonstrate that our method outperforms existing baselines in visual realism and geometric consistency, providing a robust tool for scalable autonomous driving scene augmentation.
Authors:Shaozu Ding, Linan Song, Marco De Vincenzi, Dajiang Suo
Abstract:
LiDAR has increasingly been integrated into traffic cameras to expand coverage and mitigate occlusion in roadside cooperative perception. However, how unimodal and camera‑LiDAR fusion architectures behave under variations in LiDAR point sparsity induced by sensor configurations and scene‑dependent sensing conditions remains underexplored. We introduce RESOLVE, a large‑scale real‑world benchmark dataset featuring multi‑resolution roadside LiDAR and synchronized camera‑LiDAR sensing for systematic evaluation of unimodal and fusion‑based architectures in roadside 3D detection and tracking. RESOLVE contains over 100k images and 26k point cloud frames with 220k manually annotated bounding boxes, captured at a real‑world urban intersection across diverse lighting and weather conditions and spanning 10 classes of traffic participants. In particular, RESOLVE enables controlled evaluation across three LiDAR resolution levels while keeping all other sensing and environmental factors fixed. This allows fair cross‑architecture comparisons under point cloud distribution shifts resulting from resolution variations, sensing distance, and training‑inference resolution mismatches. Results from extensive benchmark experiments reveal insights into how multimodal fusion can compensate for LiDAR point sparsity, offering clues for designing cost‑efficient roadside multimodal perception. The dataset and benchmark codes are available at https://github.com/ASU‑Suo‑Lab/RESOLVE.
Authors:Sairam VCR, Varun Gopal, Poornima Jain, Vineeth N Balasubramanian, Muhammad Haris Khan
Abstract:
Real‑world detectors for autonomous driving, surveillance, and robotics must handle domain‑shifts under strict latency and memory constraints, yet existing source‑free object detection (SFOD) methods rely on heavyweight architectures that prioritize accuracy alone. We show this trade‑off is unnecessary: building on YOLOv10, an NMS‑free dual‑head detector, we achieve state‑of‑the‑art adaptation accuracy while being faster and more compact. We observe that directly applying vanilla mean‑teacher self‑training to dual‑head detectors leads to suboptimal adaptation performance due to two key factors. First, simple pseudo‑label generation strategies, such as using a single head or directly combining high‑confidence predictions from both heads, yield suboptimal supervision under domain‑shift. We propose DHF (Dual‑Head Pseudo‑Label Fusion) which selectively admits one‑to‑one (O2O) and one‑to‑many (O2M) head predictions, preserving precision and recovering missed objects. Second, we observe domain‑shift collapses multi‑scale feature discriminability. We propose the use of our MARD (Multi‑scale Adaptive Representation Diversification) loss which mitigates this by enforcing detection‑aware variance and covariance constraints on multi‑scale feature maps. Both modules are training‑time only, leaving inference unchanged. Across domain‑shift benchmarks, our method, RT‑SFOD yields 1.4 to 3.5% mAP gains, 1.3× higher throughput, with ~2× fewer parameters than prior state‑of‑the‑art SFOD methods, thus advancing the Pareto frontier of the speed‑accuracy‑model size trade‑off. We report main results with YOLOv10, and demonstrate generalizability with additional YOLO‑ and DETR‑based dual‑head detectors. Code is available here: https://github.com/Sairam13001/RT‑SFOD/
Authors:Kyuhwan Yeon, Benjamin Ramtoula, Daniele De Martini
Abstract:
Most end‑to‑end autonomous driving methods rely solely on instantaneous sensor observations, limiting them to reactive behavior without the anticipatory foresight human drivers employ through prior experience. We introduce geospatial visual priors, street‑level visual context anchored to the intended driving route, providing visual‑spatial foresight independent of real‑time sensors. We propose a memory augmentation module featuring a dual‑memory architecture and an adaptive memory gate, which can be easily integrated into existing end‑to‑end approaches. This design pairs a contextual memory for retrieved priors with a persistent fallback memory, and dynamically regulates the influence of memories based on current state compatibility. Evaluated on the NAVSIM‑v2 benchmark, our approach consistently improves performance across diverse end‑to‑end baselines. Furthermore, because these priors are independent of onboard sensors, our method inherently improves robustness against sensor corruption, while the dual‑memory design ensures safe fallback when the retrieved priors themselves become unreliable. Our project page is available at https://ori‑mrg.github.io/PriorEye.
Authors:Junha Jung, Minbyul Jeong, Suhyeon Lim, Sungwook Jung, Jaehoon Yun, Taeyun Roh, Mujeen Sung, Jaewoo Kang
Abstract:
Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post‑training pipelines remain predominantly outcome‑centric, relying on final answer correctness or sequence‑level preferences. This suffers from sparse credit assignment, making it difficult to optimize the reasoning process essential for clinical applications. Our analysis reveals that cascading errors from early‑stage reasoning failures are a leading cause of incorrect predictions in medical visual question answering (VQA) benchmarks. Motivated by this, we propose Medical Reasoning‑aware Policy Optimization (MRPO), an RL algorithm that incorporates step‑wise process rewards. When the final answer is incorrect, MRPO assigns exponentially larger penalties to tokens in earlier invalid reasoning steps, breaking failure cascades without compromising successful paths. Across three multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3‑VL‑8B‑Instruct even surpasses substantially larger medical MLLMs such as HuatuoGPT‑Vision‑34B by 2.79 points. Moreover, MRPO reduces early‑stage reasoning failures from 64.0% to 13.0%, showing that targeted mitigation of cascading failures improves both reasoning quality and final answer accuracy. Our code is available at https://github.com/dmis‑lab/MRPO
Authors:David Montalvo-García, Nicolás Gaggion, María J. Ledesma-Carbayo, Enzo Ferrante
Abstract:
Graph‑based cardiac segmentation with implicit anatomical correspondences provides topological guarantees and population‑level analysis capabilities, but models trained on independent frames of image sequences exhibit temporal discontinuities that affect reliable clinical measurements, particularly in cardiac ultrasound. In this work, we introduce self‑supervised temporal regularization as a post‑training refinement stage that exploits the temporal coherence in image sequences to enforce consistent cardiac segmentation and motion estimation over time, without requiring per‑frame annotations. By penalizing velocity and acceleration discontinuities across consecutive frames, our method achieves temporally consistent segmentations while maintaining the learned anatomical correspondences. We further leverage these correspondences to automatically map landmarks to the AHA 17‑segment clinical standard, enabling standardized regional assessment and detection of pathological myocardial motion patterns. Validation on CAMUS dataset demonstrates the clinical utility of combining temporal consistency with automatic regional mapping. The code is publicly available at https://github.com/david‑montalvoo/MaskHybridGNet‑TempReg
Authors:Ziyuan Liu, Ruifei Zhu, Ouqiao Ma, Yuantao Gu
Abstract:
Remote sensing change detection (CD) traditionally focuses on pixel‑level binary segmentation, which identifies where changes occur but neither what nor why. To bridge this semantic gap, we introduce JL1‑CC&QA, a multi‑task benchmark that extends the JL1‑CD dataset with two complementary annotation layers: change captioning (CC) and change question answering (QA). Built upon 5,000 bi‑temporal image pairs acquired by the Jilin‑1 satellite at 0.5‑0.75m ground sample distance, the benchmark comprises: (i) JL1‑CC, providing 17,021 quality‑verified captions that describe diverse land‑cover transformations; and (ii) JL1‑QA, offering 20,060 question‑answer pairs across eight question types, enabling fine‑grained, interactive interrogation of surface changes. All annotations are produced via a three‑stage pipeline consisting of multi‑modal large language model (LLM) generation, vision‑grounded LLM judging, and human expert verification. We hope that JL1‑CC&QA, as a benchmark unifying binary change masks, change captions, and change‑oriented QA over the same image set, will serve as a valuable resource for the community to advance multi‑task change understanding in remote sensing. The dataset is available at https://github.com/circleLZY/JL1‑CD.
Authors:Ba-Thinh Nguyen, Huu-Dung Nguyen, Thi-Duyen Ngo, Thanh-Ha Le
Abstract:
Remote photoplethysmography (rPPG) estimates physiological signals from facial videos by analyzing subtle pulse induced skin color variations. Despite recent progress, existing self‑supervised rPPG methods mainly reconstruct masked pixels or low‑level visual representations, which can bias the model toward facial appearance rather than latent physiological dy namics. Moreover, most recent Mamba‑based approaches scan facial video tokens only in chronological order, limiting their ability to exploit the cyclic structure of pulse signals. To ad dress these limitations, we propose RhythmJEPA, a rhythm structured joint‑embedding predictive learning framework for rPPG. Instead of reconstructing RGB frames, RhythmJEPA predicts latent teacher representations from masked facial videos, thereby encouraging physiology‑aware representation learning in the embedding space. To explicitly model pulse‑related tem poral structure, we introduce a Cyclic Rhythm‑State Plan ner (CRSP), which estimates frame‑wise latent physiological states and decodes the most plausible cyclic state path via dynamic programming with a constrained transition grammar. Guided by the decoded states, we further design a Dual Order Mamba Encoder (DOM), which combines conventional chronological scanning with state‑ordered scanning to capture both local temporal continuity and long‑range rhythm‑consistent dependencies. Finally, a lightweight Spatial Pulse Mixer (SPM) extracts compact pulse‑sensitive facial tokens with a favorable balance between complexity and performance. Experiments on PURE, UBFC‑rPPG, and MMPD show competitive performance over representative rPPG methods. The codes are available at https://github.com/deconasser/RhythmJEPA.
Authors:Jiwen Yu, Jianxiong Gao, Jianhong Bai, Yiran Qin, Kaiyi Huang, Quande Liu, Xintao Wang, Pengfei Wan, Kun Gai, Xihui Liu
Abstract:
Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world models is the lack of memory, causing inconsistent generated scenes over extended durations. Previous methods explored rule‑based context frame retrieval as memory, but they fail to generalize in scenarios with scene occlusions and dynamic objects. We propose MemLearner, a learning‑based adaptive context query method using query tokens to bridge context and predicted tokens. By leveraging the video generation model itself for context querying, MemLearner exploits pre‑trained visual priors without training additional modules from scratch, and incorporates efficient strategies for training and inference. We collect a dataset of long videos with scene occlusions and dynamic objects, paired with camera pose annotations, and propose a multi‑dataset training strategy leveraging both annotated rendered and unannotated real‑world videos. Extensive experiments demonstrate that MemLearner significantly outperforms prior video world models in terms of scene consistency and memory, particularly under challenging occlusion and dynamic scenarios.
Authors:Jingbo He, Michael Färber, Roberto Calandra
Abstract:
For robots manipulating open‑world objects, tactile representations must generalize to unseen materials. We introduce RCT (Robotic Contact Tactile), a robot‑collected touch‑vision‑language dataset with 29,279 tactile frames from full robot presses on 122 industrial reference materials in 7 categories, recorded with three DIGIT sensors at multiple contact positions. RCT preserves each press as a contact sequence, enabling held‑out evaluation across materials, categories, sensors, contact positions, and contact sequences. Frames from one press are strongly correlated: frame‑random splits can place near‑duplicate observations of the same physical interaction in both training and test. With the encoder held fixed, removing contact‑sequence overlap reduces tactile‑to‑text Recall@1 by 17.7 percentage points. When materials are additionally held out at training time, performance drops sharply, leaving held‑out‑material Recall@1 at 25.1 +/‑ 6.1% averaged over three held‑out draws. The public TVL/HCT split shows the same structure: every test contact sequence appears in training, and raw‑pixel nearest neighbors recover the correct sequence in 98.3% of cases. Uniformly sampling a press improves contrastive training, and RCT‑trained embeddings improve category probes on unseen materials. RCT makes contact‑sequence‑aware, held‑out‑material evaluation reproducible and exposes novel‑material generalization as a central challenge for robotic tactile perception. The RCT dataset is open‑sourced at https://faerber‑lab.github.io/RCT/
Authors:Haoming Liu, Yuanhe Guo, Yijia Cao, Shenji Wan, Hongyi Wen
Abstract:
Diffusion models have emerged as a dominant paradigm in generative modeling, enabling high‑fidelity sampling from complex data distributions. Despite impressive capabilities, controlling diffusion models to produce outputs aligned with user intent remains an open challenge, especially when balancing global coherence with local precision. Existing control mechanisms vary in the granularity of their conditioning signals. For example, textual prompts guide generation globally through high‑level semantics, while ControlNet‑like approaches secure precise local structure via dense conditions. In this work, we introduce Histogram‑constrained Image Generation (HIG), a novel control mechanism that falls into the middle ground of control granularity. Our framework enforces user‑specified distributional constraints (e.g., color histograms or latent token distributions) during the generation process with exact precision. We model such control as an optimal transport (OT) problem and apply explicit guidance transformations during sampling, thereby driving the diffusion trajectory to align with the desired histogram. We demonstrate the versatility of HIG across diverse applications, including constrained generation via color/latent histograms and high‑capacity information embedding through histogram‑level encoding. Our findings underscore the promise of distributional control, a flexible and interpretable control scheme that is fully compatible with existing control mechanisms, diversifying the hybrid strategies for controllable image generation. Our project page is available at: https://maps‑research.github.io/hig/.
Authors:Ruiqi Xu, Daniel Aliaga
Abstract:
Despite advances in indoor scene generation, synthesizing coherent building exteriors consistent with generated interiors remains largely unexplored. Existing methods can generate floor plans and wall layouts but typically stop at a structural shell, lacking stylistically consistent facades and roofs. Completing these exteriors is challenging because the footprint, wall geometry, and opening semantics must remain fixed‑constraints that unconstrained generative models often violate. We introduce ShellMaker, a language‑guided exterior completion framework that operates under these structural constraints. Given a building scaffold and a text style prompt, ShellMaker generates a complete exterior mesh with PBR materials by combining parametric roof generation, LLM‑based part‑aware prompt refinement, joint wall‑roof material retrieval, and geometry‑aware assembly. Operating on a format agnostic scaffold representation, ShellMaker generalizes to indoor generators, CityGML, and CAD inputs, while maintaining structural consistency and improving architectural coherence over retrieval and unconstrained generative baselines. The project page is available at https://ruiqixu37.github.io/ShellMaker_web/
Authors:Ke Wang, Xiaoyi Pan, Zhaoyu Gu, Xiaofeng Ai, Zhiming Xu, Feng Zhao, Shunping Xiao
Abstract:
Synthetic aperture radar automatic target recognition (SAR ATR) is critical for Earth observation and defense, but its practical deployment is constrained by scarce annotated training data. Self‑supervised pre‑training alleviates this label bottleneck, yet prevailing Transformer architectures incur prohibitive quadratic computational complexity, and conventional universal masking neglects the unique electromagnetic scattering properties intrinsic to SAR imagery. To address these limitations, we propose SAMBA (Scattering‑Guided Bidirectional Mamba), an efficient self‑supervised pre‑training foundation model for SAR target interpretation. Our framework features three core innovations: (i) a linear‑complexity Mamba encoder with a mid‑sequence class token to mitigate computational bottlenecks; (ii) a three‑level hierarchical Scattering‑Guided Masked Autoencoder (SG‑MAE) masking strategy guided by SAR physical priors, aligning the pretext task with SAR's intrinsic imaging mechanism; (iii) a lightweight SpatialMix feature interaction module to enhance cross‑region feature fusion. We also design a two‑stage cross‑domain pre‑training pipeline to optimize the overall pre‑training process. Extensive evaluations demonstrate that SAMBA consistently delivers superior performance across all pre‑training configurations, with substantially fewer parameters than both CNN and Transformer baselines. Compared with the default masking strategy in standard MAE, the proposed SG‑MAE strategy further boosts the model's few‑shot transfer capability. Benchmarking on seven downstream datasets covering classification and detection tasks shows SAMBA achieves state‑of‑the‑art (SOTA) performance on most metrics, fully validating its robust generalizability across diverse SAR interpretation tasks. Source code and pre‑trained weights are publicly available at https://github.com/mynswkk/SAMBA.
Authors:Yuxiang Xie, Qi Lv, Jianming Xing, Zijian Hong, Xiang Deng, Weili Guan, Liqiang Nie
Abstract:
Vision‑language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to the RoboSpatial Challenge at the Embodied Reasoning in Action Workshop, CVPR 2026, built on RoboBrain2.5‑8B‑NV. RoboSpatialBrain combines two training‑free, inference‑time mechanisms: a forced <think> prefix activation strategy paired with a task‑specific post‑prompt that elicits deliberate reasoning on context and compatibility tasks, and an explicit reference‑frame redirection pipeline that resolves camera‑centric and object‑centric ambiguity for context tasks. We additionally explore fine‑tuning RoboBrain2.5 on compatibility data and present a detailed analysis of its interaction with prompting. RoboSpatialBrain achieved first place in the RoboSpatial Challenge, with an overall success rate of 80.9% on RoboSpatial‑Home. Code is available at https://github.com/YuxiangXie2003/RoboSpatialBrain.
Authors:Md Raqib Khan, Santosh Kumar Vipparthi, Subrahmanyam Murala
Abstract:
Despite rapid progress in learning‑based stereo matching, high accuracy is often achieved at the cost of heavy backbones and computationally intensive 3D cost volume processing, resulting in substantial memory and runtime overhead. More critically, these methods frequently struggle to generalize across domains, limiting their practical deployment. We present LiteMatch, a lightweight stereo matching framework that achieves strong zero‑shot generalization through cost volume stabilization‑without expensive 3D convolutions. LiteMatch employs two complementary encoders: a Cross‑View Correspondence Encoder (CVCE) to capture global cross‑view interactions, and a High‑Frequency Encoder (HFE) that enhances fine structural details via FFT‑based frequency cues. To stabilize the cost volume, we introduce the Cost Volume Consistency Loss (CVC‑Loss), a voxel‑wise binary cross‑entropy objective applied to softmax‑normalized cost distributions. By encouraging sharp and unimodal disparity probabilities, CVC‑Loss promotes stable cost distributions and enables rapid convergence. A lightweight refinement module further produces sharp full‑resolution disparities with low‑iteration updates, avoiding heavy recurrent refinement. With a flexible design ranging from 3.36M to 9.58M parameters, LiteMatch achieves exceptional zero‑shot generalization, delivering competitive EPE and D1 performance across Scene Flow, KITTI, Middlebury, ETH3D, and DrivingStereo. Our results establish that lightweight architectures can indeed generalize across domains without sacrificing accuracy. \hrefhttps://mdraqibkhan.github.io/Litematch\textcolorblueCode
Authors:Nikolai Röhrich, Julian Gleißner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
Abstract:
Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.g., dense regions or small objects in aerial or autonomous mobility data. While synthetic augmentation is an appealing solution, directly generating new labeled data risks misalignment of labels and generated pixels. Existing solutions to this problem often rely on external models, or employ coarse heuristics such as indiscriminately augmenting all foreground objects or entire backgrounds, which wastes capacity on uninformative pixels. To address this, we propose an uncertainty‑guided synthetic context augmentation strategy that strictly preserves label validity and efficiently maximizes pixel informativeness per synthetic sample ‑ no external guardrails required. Using a baseline segmenter's predictive entropy, we identify uncertain semantic regions and inpaint only the complementary visual context. When fine‑tuning the segmenter on this synthetic data, we compute the loss only over the original pixels, excluding inpainted regions. This focuses learning on the unmodified, uncertain regions while presenting them in novel contexts. We demonstrate substantial mIoU gains on Cityscapes, UAVID, and BDD100K with the largest gains on rare and difficult classes such as buses, trains, or (from the aerial perspective) cars. Our results demonstrate that uncertainty‑guided context augmentation is a highly effective lever to improve segmentation performance on complex datasets, with code provided at https://github.com/XITASO/Preserve‑the‑Hard‑Regenerate‑the‑Rest.
Authors:Clément Fuchs, Tim Bary, Benoît Macq
Abstract:
Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage guarantees. Full conditional guarantee is not an attainable goal, a well known fact in conformal predictions literature. As a result, several approaches have tried to approximate this behavior by adapting the conformal sets of test‑time samples according to their similarity to calibration examples. Although the latter has gained traction and shown impressive performances for regression problems, its application to image classification remains under‑explored. We conduct an extensive benchmarking on natural image classification tasks with vision‑language models (VLMs), using our open source implementation of a recent localized conformal prediction algorithm. We show that straightforward usage of the cosine similarity between test‑time and calibration visual features, an intuitive choice for VLMs, is not sufficient to improve over the non‑local baselines. In response, we propose a simple non‑linear transformation of the cosine similarities, which conserves marginal coverage guarantees and achieves statistically significant mean set sizes reduction. Code is available at https://github.com/cfuchs2023/lcp‑vlm/.
Authors:Zikang Yan, Xiao Wang, Qingquan Yang, Zhendong Yang, Gaoting Chen, Zehua Chen, Bo Jiang, Jin Tang, Guosheng Xu
Abstract:
Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service life of fusion devices. However, conventional numerical methods, such as the Finite Element Method (FEM), are computationally expensive and therefore unsuitable for real‑time applications. Therefore, a fast and generalizable method is required for real‑time reconstruction of the divertor temperature field and subsequent real‑time control. To address the above issue, we propose a Physics‑aware Neural Operator Transformer (PNOT) to characterize the spatiotemporal evolution of the divertor temperature field. It models boundary heat‑flux relations as a structured graph and employs graph attention to explicitly capture spatial physical dependencies. Inspired by physics‑aware attention, we further develop a physics‑aware neural operator module to aggregate query points with similar physical conditions via slicing and model heat diffusion, while a gradient‑constrained Sobolev regularization loss enforces consistency between function values and their derivatives. Experimental results show that these physical constraints improve prediction accuracy while preserving physical consistency. The source code of this paper will be released on https://github.com/Event‑AHU/OpenFusion
Authors:Xu Yan, Huiqun Wang, Chen Wang, Lei Ren, Di Huang
Abstract:
Masked autoencoding has emerged as a prominent paradigm for self‑supervised learning on 3D point clouds, achieving competitive performance across downstream tasks. Unlike its 2D counterpart, 3D masked autoencoding directly reconstructs spatial coordinates, making it inherently susceptible to positional leakage. In this work, we identify that the decoder in existing 3D MAE frameworks tends to over‑rely on positional information, which weakens semantic representation learning and leads to suboptimal feature quality. To address this issue, we propose MPL‑MAE, a masked point learning framework that mitigates positional over‑reliance while enhancing the utilization of encoder features. Specifically, we introduce a recalibrated positional embedding module that suppresses metric‑dominant coordinate signals while preserving geometric topology, together with a gated positional interface module that dynamically regulates positional injection during reconstruction. These designs promote a more balanced interaction between spatial priors and semantic features, yielding robust and informative representations. Extensive experiments across downstream tasks demonstrate that MPL‑MAE consistently achieves competitive performance, validating its effectiveness. Code is available at https://github.com/yanx57/MPL‑MAE.
Authors:Kartik Bali, Roland Aydin
Abstract:
Identifying and grounding precise geometric entities, such as edges, planar regions, and curved surfaces within 3D objects, is foundational to computer‑aided design (CAD), robotic manipulation, and scientific simulation. Although modern Vision Language Models (VLMs) have advanced referring segmentation (RIS) in the image domain, extending such language‑driven localization to structured 3D geometry is substantially harder. The 3D object appearance is highly sensitive to viewpoints; a single perspective may render a target entity clearly observable, while another may suffer from severe occlusion or foreshortening. In this work, we attempt to solve these challenges with MV‑GEL (Multi‑View Geometric Entity Localization), a framework for localizing fine‑grained geometric entities on polygon meshes from natural language queries. Our key insight is that reliable CAD entity (i.e., faces, edges or solids) localization depends on selecting views that make the queried entity maximally interpretable. We introduce GELviews, a prompt‑conditioned ranking module that prioritizes viewpoints based on language prompted observability of geometric CAD entities. Selected views are processed by a VLM‑based reasoning segmentation backbone, and predicted masks are lifted to the corresponding meshes via geometry‑aware ray casting. Our framework is completely CAD agnostic and relies only on 3D meshes. Experiments show up to a 1.7X improvement in face‑level IoU and over 4.5X gains in edge‑level F1 compared to vanilla baselines, substantially outperforming CLIP‑based and random view sampling, particularly for thin and view‑sensitive structures.The dataset, code and trained checkpoints are available at https://github.com/kbali1297/MV‑GEL.
Authors:Jiawei Xu, Qiangqiang Zhou, Zhouping Li, Yanjiao Shi, Yugen Yi, Jiacong Yu
Abstract:
In recent years, most research on multimodal salient object detection (SOD) and camouflaged object detection (COD) typically aims to improve performance through complex cross‑modal feature fusion and decoding structures. However, this approach leads to an excessively large model parameter scale and often fails to deliver satisfactory detection performance due to structural redundancy. In contrast, the human visual process is able to efficiently perform salient and camouflaged object identification without such complex structures. This contrast raises an important question: Can we draw conceptual inspiration from the human visual process to achieve a simpler modeling strategy, and still realize accurate and efficient object detection? To answer this question, we propose HVPNet, a simple yet general bio‑inspired computational architecture. Drawing on the multi‑layered information integration of the retina as a conceptual metaphor, we designed a Retinal Integration Module (RIM), which effectively integrates multimodal features through a level‑specific multi‑stage integration strategy. To fully exploit these features, we further design a cortical decoder (CD) that breaks down the decoding process into low‑ and high‑level visual stages, abstracting the hierarchical processing in the human visual cortex. Benefiting from these designs, HVPNet can readily extend to seven tasks across four modalities. Without bells and whistles, it establishes an excellent accuracy‑efficiency trade‑off across 22 datasets spanning these seven tasks. Our code is available at https://github.com/jiaweiXu1029/HVPNet.
Authors:Deniz Bickici, Michael Pabst, Shohei Mori, Dieter Schmalstieg
Abstract:
Open‑vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision‑language models, leaving the graph unqueryable during exploration. We argue that this sequential coupling is unnecessary and propose an asynchronous architecture in which lightweight online mapping runs concurrently with heavyweight semantic refinement. A probabilistic voxel‑based backbone maintains stable object identities incrementally, while background VLM agents progressively enrich the graph. This framework resolves duplicate object tracks through semantic loop closure, attaches fine‑grained visual attributes and derives spatial relations between objects. A multi‑target frame scheduler amortizes VLM cost by selecting a small set of informative frames that jointly cover multiple targets. The resulting scene graph is queryable during exploration and grows in semantic richness over time. Our method matches or outperforms existing open‑vocabulary 3D scene graph methods on semantic segmentation (ScanNet, Replica) and surpasses the prior state‑of‑the‑art across three visual grounding benchmarks (Sr3D+, Nr3D, ScanRefer) by 15.3 to 18.8 A@0.25. Project page: https://denizbickici.github.io/thinkgraphs/
Authors:Jisung Park, Seohyeon Kang, Daeun Yoo, Eunsu Lee, Seoin Cho, Wooyeop Choi, Ian Choi, James R. Evan, Daesoo Kim, Sonia Gandhi, Minee L. Choi
Abstract:
Artificial intelligence is transforming our capability to solve biological challenges. In dimensionality bottleneck regimes exacerbated by high‑dimensional biological data, neural networks force distinct concepts into the lower dimensions known as superposition. Although this superposition is widely known to hinder interpretability, its impact on corrupting the geometry of latent spaces remains critically overlooked. Here, we utilized sparse autoencoders (SAEs) trained on over 100,000 multiplexed images of patient‑derived Parkinson's disease and healthy neurons to resolve superposition. This approach bypasses the mathematical non‑uniqueness of feature attribution by shifting to interpretable latent representation analysis. We theoretically and empirically demonstrate that superposition contaminates representational metric spaces, and thereby SAEs successfully recover geometric fidelity. By treating these geometrically purified representations as single‑cell state vectors, we adapted single‑cell RNA sequencing (scRNA‑seq) data analysis methodologies directly to the image domain. Finally, we introduce GW‑map, utilizing Gromov‑Wasserstein optimal transport to align these image representations with authentic scRNA‑seq data de novo. This coupling reconstructs hierarchical neuronal pathology pathways such as Calcium‑AIS scaffold, without reference spatial transcriptomics, establishing a scalable foundation for spatial biology. Code is available at https://github.com/jijihihi/Bio\_superposition
Authors:Stefanos-Iordanis Papadopoulos, Zacharias Chrysidis, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis
Abstract:
The proliferation of multimedia content on social platforms has fueled multimodal misinformation, where images are used to reinforce false claims. Consequently, Multimodal Fact‑Checking (MFC) has emerged as an increasingly important research area. However, current progress is hindered by a reliance on synthetic training data and curated benchmarks that fail to capture the complexity of in‑the‑wild data. Furthermore, existing detection models rely on restricted intra‑modality consistency or unconstrained all‑to‑all fusion, failing to capture nuanced relations between posts and external evidence. To address these limitations, we introduce X‑POSE, a benchmark of real‑world, community‑annotated multimodal posts from X (formerly Twitter), augmented with full‑length news articles retrieved via VLM‑optimized search. Additionally, we propose TRENT, a novel MFC model that performs evidence triangulation using three parallel cross‑attention streams alongside a relational fusion mechanism that explicitly models entailment and contradiction. Extensive evaluations demonstrate that TRENT consistently outperforms state‑of‑the‑art specialized models and commercial VLMs. The code, prompt templates, and dataset are available at https://github.com/stevejpapad/evidence‑triangulation
Authors:Hyunsoo Lee, Inwoo Hwang, Young Min Kim
Abstract:
Generating diverse, coherent, and plausible content from partially given inputs remains a fundamental challenge for diffusion models. Existing approaches face clear limitations: training‑based approaches offer strong task‑specific results but require costly computation, and they generalize poorly across tasks. Training‑free approaches offer better efficiency, but they do not explicitly optimize over unobserved variables, leading to globally inconsistent results. To address these limitations, we introduce Accelerated Likelihood Maximization (ALM), a novel training‑free sampling strategy integrated into the reverse diffusion process that significantly extends the applicability of diffusion models beyond simple generation tasks. Unlike previous methods that implicitly influence missing regions through pre‑generated region constraints, we directly optimize the unobserved region during the sampling process, enabling globally coherent and plausible generation. Furthermore, we incorporate an acceleration strategy that significantly improves computational efficiency without sacrificing performance. Experimental results demonstrate that ALM consistently outperforms state‑of‑the‑art methods in various data domains and tasks, establishing a powerful paradigm for versatile content generation.
Authors:Yibing Zhang, Xunpeng Yi, Qinglong Yan, Yeda Wang, Han Xu, Jiayi Ma
Abstract:
With the advancement of imaging technology, ultra‑high‑definition images have become increasingly essential in modern visual applications. However, existing multi‑focus image fusion remains largely confined to low‑resolution images and faces three major barriers in UHD scenarios, namely data availability, model adaptability, and deployment feasibility, which severely hinder its practical application. To shatter these barriers, first, we propose the UHD‑MFF dataset, the first large‑scale ultra‑high‑resolution multi‑focus fusion dataset. Second, we propose a scale‑specialized lookup‑table framework tailored for ultra‑high‑resolution images, termed as UMF‑LUT. It consists of Coarse‑Region Lookup Table (C‑LUT) and Detail‑Edge Lookup Table (D‑LUT). Specifically, C‑LUT performs joint queries of multiple gradient cues and semantic cues at low‑resolution scales to enable region‑level decision‑making. Also, D‑LUT operates at high‑resolution scales, leveraging efficient Laplacian cues to provide complementary edge‑level decision information. Such a design makes the model particularly well‑suited for ultra‑high‑resolution multi‑focus image fusion. Finally, it offers strong deployability with minimal computational overhead, enabling real‑time 4K multi‑focus fusion and showing promising potential for smartphone. Extensive experiments demonstrate that it outperforms SOTA methods in both visual fidelity and quantitative metrics. It effectively advances the development of multi‑focus image fusion toward ultra‑high‑resolution imaging scenarios. The code is available at https://github.com/zyb5/UHD‑MFF.
Authors:Dong Yeong Kim, JunGyu Lee, Jaewon Choi, June Young Seo, Myeongseop Kim, Jinwook Choi, Taek Min Kim, Young-Gon Kim
Abstract:
Real‑time video segmentation of the prostate in Transrectal Ultrasound (TRUS) is essential for image‑guided interventions. While conventional 2D methods suffer from inter‑frame inconsistencies by disregarding temporal context, 3D architectures incur prohibitive latency. To resolve this dilemma, we present a Temporally Consistent Learning Framework that distills temporal coherence into a 2D network during training, preserving single‑frame inference efficiency. Our design is driven by a key clinical observation: the prostate exhibits geometric stability, whereas the surrounding acoustic environment fluctuates due to physiological motion and transducer pressure. Because conventional temporal constraints propagate erroneous gradients from these unstable regions, we introduce a Confidence‑Weighted Temporal Consistency objective derived from optical flow warping residuals, selectively attenuating contributions from unreliable regions. Complementing this pixel‑wise constraint, a Dual‑scale Prototype Alignment Module enforces semantic coherence through contrastive optimization of local boundary and global semantic features. Furthermore, to eliminate the need for dense per‑frame video annotations, we employ geometric equivariance‑based pseudo‑labeling with knowledge distillation from a pretrained teacher. Extensive experiments on SUN‑SEG and our newly introduced TRUS‑V benchmark (2,679 frames) demonstrate state‑of‑the‑art accuracy and temporal consistency at real‑time speed. Code and dataset are available at https://github.com/DYDevelop/DTC‑TRUS.
Authors:Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat
Abstract:
Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero‑shot and few‑shot recognition across standard benchmarks. Yet their ability to deny an action by recognizing when an activity is not happening despite strong contextual cues remains largely unexplored. We introduce UCF101‑AD, a large‑scale benchmark consisting of paired Action‑Presence and Action‑Denial clips, designed to evaluate this capacity for denial. Each negative video in UCF101‑AD preserves the same contextual and motion cues, including persons, objects, and locations, as its positive counterpart, but the defining action itself is explicitly absent. Evaluating 20 state‑of‑the‑art MLLMs reveals a consistent failure: models that exceed 85% accuracy on the positive action classes collapse below 50% on their action‑denial counterparts, indicating a strong inclination to affirm plausible actions rather than verify that they truly occur. This exposes a critical blind spot in modern video understanding: the inability to reason causally about whether a motion actually happens. To probe this issue, we explore a causal graph formulation, CausalAct, which expresses scene structure through natural‑language prompts linking context, interaction, and motion. Incorporating such causal cues substantially reduces false positives, demonstrating that denial is a learnable reasoning skill. UCF101‑AD provides a new lens for diagnosing and improving causal reasoning in multimodal models. Dataset and relevant code: https://github.com/raiyaan‑abdullah/Learn‑to‑Deny.
Authors:Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, Maximilian Rokuss, Mingyu Lu, Timothy Ossowski, Juan Manuel Zambrano Chaves, Cliff Wong, Peniel Argaw, Yashna Hasija, Mu Wei, Wen-wai Yim, Qin Liu, Zilin Jing, Jason Entenmann, Naoto Usuyama, Tristan Naumann, Hoifung Poon
Abstract:
As AI agents become increasingly capable of complex, long‑horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real‑world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end‑to‑end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi‑step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT‑5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT‑5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at https://github.com/microsoft/HealthAgentBench.
Authors:Jiyong Boo, Byeongin Joung, Hyemin Yang, Kuk-Jin Yoon
Abstract:
Monocular 3D lane detection plays a critical role in autonomous driving, yet recovering reliable 3D geometry from a single image remains challenging due to inherent depth ambiguity. Prior methods project image features into Bird's‑Eye‑View (BEV) space under a flat‑ground assumption, causing geometric distortion on real‑world roads. Recent methods instead predict explicit height maps to capture non‑planar surfaces, but still rely on sparse anchor‑based regression and exploit the recovered geometry merely for spatial transformation rather than semantic understanding. To overcome these limitations, we propose HSDF‑Lane, which implicitly models the road surface as a Height‑aligned Signed Distance Field (HSDF) over a densely sampled 3D feature volume. Through differentiable rendering, the HSDF jointly produces an accurate height map and surface‑aligned features. We further introduce Lane‑aware Semantic Positional Encoding (LSPE), which injects a lane‑existence prior derived from the surface‑aligned features into the transformer queries, coupling geometric structure with semantic guidance. Extensive experiments on the OpenLane benchmark show that HSDF‑Lane achieves state‑of‑the‑art performance in both 3D lane detection and height map estimation.
Authors:Oleksii Nasypanyi, Jaemin Cho, Utku Ozbulak, Byungkon Kang, Francois Rameau
Abstract:
Scene Coordinate Regression (SCR) methods are increasingly adopted for visual localization. In these approaches, the scene is implicitly encoded within a neural network that regresses a 3D world coordinate for each image pixel. Because the scene is represented only through the network parameters and not stored explicitly as images or maps, such methods are often assumed to be privacy‑preserving. In this work, we show that this assumption is incorrect in practice. Specifically, we introduce a query‑based attack that reconstructs the 3D geometry of the training environment from an SCR model under different levels of model access. To do so, we repeatedly query the model with batches of proxy images unrelated to the target scene to obtain dense pixel‑wise 3D coordinates. Reliable points are identified through their stability under small input perturbations and can be further refined in a white‑box setting. These stable points are accumulated across independent query batches to recover the scene geometry. From the recovered 3D representation, we also invert the network features to synthesize images from arbitrary viewpoints, revealing additional appearance information. Experiments on indoor and outdoor datasets demonstrate that substantial portions of training environments can be reconstructed with high geometric fidelity. Beyond geometry, we also recover an approximate color appearance, which exposes recognizable layout and potentially sensitive scene elements. This directly contradicts claims in the literature that SCR representations are privacy‑preserving by design, and reveals a real risk when such systems are deployed in private or security‑critical spaces. The project page is available at https://jaeminch0.github.io/seeing‑through‑the‑weights‑privacy‑leakage‑in‑scene‑coordinate‑regression.
Authors:Duc Cao Dinh, Khai Le-Duc, Florent Draye, Chris Ngo, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin
Abstract:
3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions. Existing approaches typically perform reasoning over the entire scene, leading to ambiguous predictions and high computational cost, especially in cluttered environments. We observe that many referential expressions rely on local spatial context and often correspond to restricted spatial regions rather than the full scene. Motivated by this insight, we propose PruneGround, an effective plug‑and‑play framework for 3DVG built upon three key components. First, we introduce Language‑Guided Spatial Pruning (LGSP), which leverages a frozen Vision Language Model (VLM) to identify language‑relevant regions, thereby reducing spatial computation and grounding candidates in the narrower search space. Second, we propose MultiView‑Conditioned Description Reformulation (MCDR), which decomposes complex expressions into simplified target‑anchor relations and augments missing spatial cues through multi‑view reasoning. Finally, we propose LLM‑Grounder, which repurposes a detection‑pretrained spatial LLM into a language‑conditioned grounding model by aligning point cloud and linguistic representations within the pruned region. Extensive experiments on the three most popular point cloud benchmarks demonstrate that our method achieves state‑of‑the‑art results on all three ScanRefer settings and on 9 out of 10 Nr3D/Sr3D settings. Code and models are publicly available: https://github.com/leduckhai/PruneGround
Authors:Junghwan Park
Abstract:
Few‑shot segmentation asks a model to delineate a target class in a query image from only a handful of annotated examples, a setting most acute in remote sensing, where labels are scarce and the imagery departs sharply from the natural images on which vision backbones are pretrained. Prevailing approaches either train a segmenter on labelled episodes, which raises accuracy within the training distribution but binds the model to it, or reduce each class to a lossy summary of frozen features, a single prototype, a few cluster prototypes, or a discrete clustering, none of which preserves the internal structure of a multimodal class. We argue that a class is better described by a distribution than by a point, and that frozen self‑supervised features already carry enough structure to estimate that distribution directly. We introduce FROST, a training‑free few‑shot segmenter that treats the reference foreground and background as two point clouds on the unit sphere of frozen DINOv3 features and labels each query token by a nonparametric density ratio, with a threshold the Bayes rule fixes at zero under equal priors. Because the variance of a density estimate shrinks as its sample grows, the decision sharpens as references accumulate, and every remaining quantity from the kernel bandwidth to the spatial gate is read from the support set rather than tuned. We develop FROST for overhead imagery, where a class is typically a scatter of many small and dissimilar instances that a density tracks but a lossy summary blurs. Across seventeen remote‑sensing benchmarks FROST surpasses both training‑free and learning‑based methods, leading by 5.6 mIoU from a single annotated example and widening its lead as the support set grows, all while remaining among the smallest models compared. Code is available at https://github.com/jhpark‑ai/FROST.
Authors:Björn Braun, Christian Holz
Abstract:
To enable personalized, real‑time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or music, a system must understand not just what a person does, but how well they execute an activity. In an ego‑exo video setting, this requires simultaneously detecting individual skilled actions and classifying each as correct or needing improvement, which Ego‑Exo4D's proficiency demonstration benchmark formalized. We first adapt seven state‑of‑the‑art temporal action detection architectures to this task, extend the evaluation protocol to disentangle detection from grading, and show that existing methods grade near‑randomly. We then introduce SkillSpotter, a pose‑aware multi‑view architecture that jointly detects and grades skilled actions through three task‑specific modules: (1) adaptive temporal suppression to handle the varying density of skilled actions across diverse activities, (2) gated 3D body pose fusion to leverage body kinematics as a complementary signal to visual features, and (3) bidirectional cross‑view attention to combine ego and exo views effectively. SkillSpotter improves class‑specific mAP from 12.40 to 21.82 (+76%) and balanced accuracy from 55.99% to 60.40% over the best baseline. SkillSpotter's modules transfer to other temporal action detection models with consistent gains, and our method generalizes beyond Ego‑Exo4D to HoloAssist. Code: https://github.com/eth‑siplab/SkillSpotter
Authors:Chaeyeon Lee, Khang Nguyen Quoc, Jinsol Song, Yosep Chong, Kwangil Yim, Jin Tae Kwak
Abstract:
Whole slide image (WSI) analysis is central to computational pathology, with multiple instance learning (MIL) emerging as the standard pipeline for slide‑level diagnosis. However, conventional approaches formulate WSI diagnosis as a flat classification task over discrete labels, contradicting the inherently hierarchical, coarse‑to‑fine nature of clinical reasoning. Although recent hierarchical classifiers and vision‑language models (VLMs) have sought to address this structural gap, they either fail to capture semantic continuity between related diagnoses or suffer from unconstrained text generation that produces taxonomic hallucinations and parent‑child label violations. To address these limitations, we propose TaxoMIL, a taxonomy‑constrained framework that reformulates WSI diagnosis as a multi‑granularity text generation task. TaxoMIL utilizes a dual‑head Transformer decoder to generate coarse‑ and fine‑level diagnostic text, and introduces taxonomy‑guided objectives that explicitly structure the label embedding space and strictly ground slide‑level visual representations within the clinical taxonomy. Extensive experiments across three diverse WSI datasets demonstrate that TaxoMIL consistently outperforms state‑of‑the‑art MIL classifiers and VLM‑based generative methods, yielding accurate and hierarchy‑aware diagnostic predictions. The code is released at https://github.com/QuIIL/TaxoMIL
Authors:Geonho Bang, Geunju Baek, Dongyoung Lee, Wonjun Jeong, Jun Won Choi
Abstract:
Long‑range 3D object detection is critical for safe autonomous driving at highway speeds, yet existing radar‑camera fusion methods remain limited at extended ranges. BEV‑based methods capture scene‑level context but incur rapidly growing computation and often lose fine‑grained object detail, while query‑based methods are efficient but provide limited scene‑level context. Temporal fusion further requires both multi‑frame accumulation for sparse distant observations and object‑level motion modeling for fast‑moving objects. We propose Horizon3D, a sparse radar‑camera fusion framework for long‑range 3D object detection that combines Gaussian primitives with sparse BEV features. Horizon3D initializes Gaussian primitives at radar‑ and camera‑estimated object keypoints using Keypoint‑Guided Gaussian Initialization, refines them through Object‑Centric Sparse Fusion, and splats them onto the BEV plane to fuse object‑level detail with sparse radar BEV context. It further introduces Dual‑Path Temporal Fusion, which aggregates temporal cues through a BEV path for scene‑level accumulation and a Gaussian path for object‑level motion propagation. Experiments on TruckScenes show that Horizon3D achieves state‑of‑the‑art radar‑camera 3D detection performance. On the validation set, it outperforms the previous best method by +3.0 NDS and +1.6 mAP while maintaining competitive inference speed.
Authors:Jiaan Wang, Sirui Liu, Yu Li, Kaiyuan Yang, Juan Cao, Sheng Tang
Abstract:
AI‑generated image (AIGI) detection is undergoing a critical transition from laboratory benchmarks to open‑world adversarial defense. The prevalent paradigm focuses on finding static feature spaces, assuming that some invariant artifacts learned from historical data can achieve universal zero‑shot generalization. While achieving saturation on several AIGI benchmarks, this static hypothesis suffers a severe performance drop against rapidly evolving generators (e.g., SD3, Nano Banana Pro). To address these limitations, we propose that the field should expand beyond "static generalization" to a new paradigm of "dynamic adaptation". We introduce Fleet, a framework that pioneers a dynamic paradigm of continuous few‑shot evolution, enabling rapid alignment with emerging generative threats. Fleet improves few‑shot adaptation by replacing unconstrained feature updates with constrained routing correction, where avoidance routing redirects novel AI samples away from Non‑AI‑dominated routes within decoupled subspaces. To validate this, we present Treasure, a benchmark spanning 64 models and 360k images, featuring diverse architectures and 20 closed‑source commercial engines. Experiments reveal that while static SOTA methods fail catastrophically on modern generators, Fleet restores performance from 20.4% to 73.1% with only 10‑shot adaptation on "Doubao Seedream 4.0". Code and data are available at https://github.com/ICTMCG/Fleet .
Authors:Jingwang Ling, Lifan Wu, Feng Xu, Shuang Zhao
Abstract:
Reconstructing physics‑based 3D assets ‑‑ geometry, materials, and illumination ‑‑ from multi‑view images is a core problem in computer graphics and vision, and a prerequisite for realistic relighting and editing. Physics‑based inverse rendering offers an accurate image‑formation model, but is severely underconstrained: without strong priors, illumination is baked into materials, and reconstructions generalize poorly to novel views and lighting. Data‑driven diffusion models, in contrast, predict visually plausible materials, yet their predictions rarely satisfy the rendering equation and are not directly usable for physics‑based rendering. We bridge these two paradigms rather than replacing either. Our key idea is to treat the predictions of a state‑of‑the‑art diffusion model not as target material values but as a similarity kernel for optimization: we introduce a regularization loss that penalizes deviations in the optimized material over surface regions where the diffusion predictions are near‑constant, while leaving the optimization free to match the input images. Built on this regularizer, our end‑to‑end pipeline jointly reconstructs geometry, materials, and illumination, yielding high‑quality assets that drop into standard rendering pipelines and relight faithfully. On the Synthetic4Relight, Stanford‑ORB, and DTC‑Synthetic datasets, our method significantly outperforms state‑of‑the‑art baselines in both reconstruction accuracy and relighting quality.
Authors:Hiroki Takeda, Yuto Miyatake, Daisuke Furihata
Abstract:
Tensor Train (TT) decomposition is a powerful technique for analyzing high‑dimensional data. Existing algorithms for computing TT decompositions can be categorized into two main types: conventional batch‑based approaches and recursive online methods. In the context of streaming data, batch methods typically achieve higher reconstruction accuracy but often suffer from memory exhaustion, while online methods provide greater computational efficiency. In this work, we introduce Online TT‑ALS (Alternating Least Squares), an algorithm that sequentially enforces orthogonality constraints. This approach allows for efficient and exact updates of the core tensor while maintaining high reconstruction accuracy. Theoretically, we prove that enforcing these orthogonal gauge constraints guarantees monotonic decrease of the local objective function and temporal smoothness. Computationally, our deterministic single‑sweep update reduces the rank dependence from quadratic to linear, achieving an overall complexity of \mathcalO(I^n‑1 r). Experimental results demonstrate that the proposed method outperforms existing online techniques not only in terms of mathematical approximation accuracy but also in human perception‑based video quality metrics. Furthermore, compared to recent deep learning‑based paradigms, our algebraic approach achieves speedups of several orders of magnitude. Consequently, our method exhibits high computational efficiency and is suitable for low‑latency real‑time processing applications.
Authors:Zhiyuan Yao, Zheren Fu, Zhixiao Zheng, Jiajun Li, Yi Tu, Zhendong Mao
Abstract:
Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image. In this paper, we identify an internal signature of hallucination: progressive degradation of text‑to‑image cross‑attention during generation, leading to specific failure patterns like unfocused or biased attention. Existing mitigation strategies are largely outcome‑driven and do not explicitly target this failure mode. To address this problem, we propose ADAPT (Attention Dynamics Alignment with Preference Tuning), an attention‑based framework that intervenes directly on text‑to‑image cross‑attention dynamics. We propose ADAPT with three key contributions: a cross‑attention visual anchor refined from early decoding to provide stable spatial grounding, an attention‑supervised inference mechanism that detects and corrects attention drift online, and a Visual Attention Guidance DPO that aligns preferences toward visually grounded responses. Experiments show that each component of ADAPT contributes to hallucination reduction, and the full framework achieves new best results across multiple hallucination benchmarks, reducing hallucination rates by 40%‑60% across mainstream backbones while preserving general multimodal capabilities. Our work provides an attention‑based perspective on mitigating hallucinations by exploring the model's internal text‑to‑image cross‑attention behaviors. Code is available at https://github.com/yao‑ustc/ADAPT
Authors:Brian Wei, Srikumar Sastry, Daniel Cher, Eric Xing, Nathan Jacobs
Abstract:
Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unlike natural imagery, satellite scenes are structured by spatially complex and semantically distinct geometries. Prior work addresses this complexity by adapting natural image frameworks using dense rasters or sparse prompts, trading off annotation cost and fidelity while breaking compatibility with vector primitives commonly used to represent geographic information. We introduce TerraDiT‑Ω, a unified spatial control framework that generates satellite imagery directly from any native geospatial primitive. By jointly leveraging precise annotations (polygons, polylines) and coarser ones (bounding boxes, points), the model supports controllable layouts across varying annotation budgets, broadening applicability to design tasks such as urban planning while remaining naturally compatible with end‑to‑end GeoAI workflows. To effectively leverage these primitives during generation, we propose Geometry‑Aware Local Attention, a conditioning mechanism that injects explicit geometric cues into the attention space. Across all conditioning formats, our approach consistently outperforms both dense‑control and sparse‑control baselines. Furthermore, this flexibility enables controllable synthetic data augmentation using a single generative model, improving downstream performance on land‑cover segmentation, object detection, road graph extraction, and scene classification. Code, data, and weights are available at https://github.com/mvrl/TerraDiT.
Authors:Shen Zheng, Anurag Ghosh, Gaurav Parmar, Srinivasa Narasimhan
Abstract:
Image‑to‑image (I2I) translation has achieved strong results in tasks like human relighting and driving scene translation using latent diffusion models (LDMs). However, compact LDMs often struggle to preserve fine‑grained structures because the encoder compresses high‑resolution inputs into a spatially downsampled latent space. To address this issue, we propose a simple saliency‑guided warp‑unwarp framework that reallocates spatial representation toward salient regions before encoding, enabling better preservation of structural details without increasing latent resolution. The warped image is processed by the original diffusion model and then mapped back via an inverse warp. In addition, we propose a simple and efficient outpainting‑based synthetic data generation pipeline to produce high‑quality paired data for image relighting. Our method is model‑agnostic, requires no architectural modification, and introduces negligible computational overhead. Experiments on human relighting, driving scene relighting, and translation demonstrate improved structural preservation, lighting faithfulness, and image quality, with our framework extending naturally to video via frame‑by‑frame application with good temporal stability. Project Webpage: https://shenzheng2000.github.io/WarpI2I.github.io
Authors:Wencong Wu, Xiuwei Zhang, Hanlin Yin, Hongxi Zhang, Yanning Zhang
Abstract:
Transformer‑based approaches have obtained excellent performance in multispectral object detection tasks due to their ability to model long‑range dependencies and capture complementary information. However, previous transformer‑based multispectral detection methods tend to use all available tokens for similarity calculation, which results in redundant information interaction from irrelevant areas, leading to degraded detection performance. To overcome this challenge, we propose a novel Dual Sparse Aggregation Transformer (DSAFormer) for multispectral object detection, which consists of a Dual Sparse Transformer (DSFormer) and a Learnable Addition Fusion Block (LAFB). Specifically, the DSFormer is designed to exploit and boost cross‑modal complementary information, thereby improving detection performance. It incorporates three key components: A Spatial Sparse Multi‑Head Cross‑Attention (SSMHCA) mechanism selectively captures cross‑modal relationships at the spatial level by reserving only the high query‑key similarity scores, eliminating irrelevant interactions. A Channel Sparse Multi‑Head Cross‑Attention (CSMHCA) mechanism performs similar sparse calculations at the channel level to enhance feature representation and filter out low matching query‑key. A Multi‑Scale Feature Refinement Layer (MSFRL) is developed to aggregate hierarchical features and suppress redundant information. To effectively fuse multimodal features, the LAFB is introduced to aggregate intramodal and intermodal feature information by feature reweighting. Extensive experimental results have demonstrated that our proposed DSAFormer achieves better detection performance against state‑of‑the‑art methods on four public datasets, including the MFAD, FLIR, M^3FD, and LLVIP. The source code of our DSAFormer will be released at https://github.com/WenCongWu/DSAFormer.
Authors:Mert Onur Cakiroglu, Zhihe Lu, Mehmet Dalkilic, Hasan Kurban
Abstract:
AI‑generated video detection benchmarks such as GenVidBench and AIGVDBench are the de facto leaderboards, yet most evaluation protocols leave uncontrolled confounds that can inflate reported generalization. As an existence proof, a three‑feature clip‑length classifier reaches a leave‑one‑generator‑out (LOGO) AUC of 0.998 on GenVidBench under unaudited evaluation, while measuring nothing about motion. A 20‑paper survey finds none applying all six standard controls that would catch this, so we combine them into an audited protocol and apply it to six representative feature sources (three published detectors and three repurposed signal sources), re‑running it cross‑dataset on AIGVDBench. The audit both debunks and certifies: the trivial classifier collapses to near chance (0.529), a CLIP baseline is caught carrying dataset identity, and the 2025 forensic detector WaveRep clears the floor at out‑of‑distribution LOGO AUC 0.996 with chance‑level real‑vs‑real coherence. At a deployable FPR of 0.1%, multiple high‑AUC methods fall to single‑digit recall and the leaderboard order changes, so we recommend an audited tuple (AUC, above‑floor margin, operating‑point recall, and calibration) over a single number. As a white‑box positive control, we add TemporalSpec (codec motion vectors); via cross‑substrate feature fusion (XSFF), a second substrate adds genuine complementarity that survives the audit. We release VidAudit, to our knowledge the largest unified and audited detector collection for this task, providing 14 detectors behind one plugin API, a leaderboard, and Croissant metadata, available at https://github.com/KurbanIntelligenceLab/vidaudit. Together, the protocol and toolkit move evaluation from leaderboard rank toward whether a result measures what it claims.
Authors:Koorosh Roohi, Javad Rajabi, Andrew Fleet, Babak Taati
Abstract:
Photomosaics are large images whose local regions are seen as independent tiles while their overall arrangement forms a coherent scene. Generating them at high resolution, with every tile convincing in its own right, is computationally expensive, since the canvas must hold many detailed tiles at once. We present PhotoQuilt, a training‑free framework that generates photomosaics at arbitrary resolution. Diffusion models struggle to satisfy both scales at once, as direct high‑resolution generation is costly and tends toward one smooth image rather than a mosaic, while patch‑based tiling keeps local detail but loses global structure. PhotoQuilt resolves this with a bootstrapped tiled denoising procedure. We first produce a global composition at low resolution to fix the layout, then upscale it in latent space and re‑inject noise to restore generative capacity. Denoising proceeds within fixed tiles, so each forms its own image while the shared global structure holds them in one layout. Because tile generation is handled separately, PhotoQuilt scales to large canvases without quadratic attention cost. Experiments show that PhotoQuilt outperforms current baselines on both global structure and local realism.
Authors:Mohammad Mahdi Abootorabi, Sina Namazi, Armin Saadat, Lyuyang Wang, Obed Dzikunu, Paul F. R. Wilson, Zhuoxin Guo, Brian Wodlinger, Parvin Mousavi, Purang Abolmaesumi
Abstract:
Micro‑ultrasound (μUS) is a new, emerging, and promising imaging modality for prostate cancer (PCa) detection, but accurate identification of suspicious tissue remains highly dependent on clinical experience, leading to substantial inter‑observer variability. Machine‑learning assistance can reduce this variability; however, training reliable deep models is challenging because supervision is sparse and noisy ‑‑ typically limited to core‑level histopathology outcomes (e.g., cancer grade and its percentage in a biopsy core) without pixel‑level lesion annotations and under severe class imbalance. We introduce Prost‑RL, which reframes μUS PCa detection as a spatially aware, policy‑driven inference problem by learning where to look before decoding. Prost‑RL integrates a lightweight reinforcement‑learning policy into a foundation‑model encoder‑decoder to generate interpretable spatial attention maps that act as soft prompts for both cancer‑likelihood heatmap prediction and image‑level classification. We further propose Adaptive Policy Optimization (APO) to stabilize hybrid supervised‑RL training and a noise‑robust objective combining symmetric cross‑entropy with negative‑entropy regularization to mitigate weak‑label noise and encourage sharp localization. On a cohort of 6,607 biopsy cores from 693 patients across five clinical sites, Prost‑RL achieves 79.0\pm3.5 AUROC with 64.6\pm6.3% sensitivity at 80% specificity for core‑level detection (+2.1 AUROC and +4.5 sensitivity points over the strongest baseline), and 79.3\pm5.8 AUROC for clinically significant cancer classification. The learned policy highlights biopsy‑aligned regions, providing transparent, spatially grounded evidence alongside quantitative risk predictions. Code is available at: https://github.com/DeepRCL/Prost‑RL.
Authors:Rasul Khanbayov, Erchin Serpedin, Hasan Kurban
Abstract:
Prototype‑based medical image classifiers present three clinical limitations: they treat findings as independent, silently amplify unsafe physician feedback, and require full retraining whenever a new finding is needed. We present GRAPE (Graph‑Augmented Prototype Explanations), a unified architecture that addresses all three challenges. First, a Graph Attention Task Head models anatomical concept co‑occurrence, boosting macro‑F1 by +13.8,pp over the prototype baseline on TBX11K. Second, a Concept‑Mismatch Safety Check ‑ the first such mechanism in prototype‑based medical classifiers ‑ warns when the model's dominant finding inside a doctor‑drawn region conflicts with the claimed label, catching 85% of erroneous annotations versus 51% for MC‑Dropout with no extra inference cost. Third, Open‑Vocabulary Prototype Anchoring aligns visual prototypes to clinical text, allowing a new finding to be added from a single labeled image without modifying any other component. On NIH ChestX‑ray14, one Effusion example recovers full‑supervision localization accuracy; on TBX11K, prototype maps achieve 2.6x better lesion localization than end‑to‑end baselines. All three capabilities add only +1~ms latency at interactive batch size. The project page is https://github.com/KurbanIntelligenceLab/GRAPE.
Authors:Brent A. Griffin, Jason J. Corso
Abstract:
Foundation model pseudo‑labeling ‑ labeling data strictly via zero‑shot inference ‑ enables massive scale, but performance is undermined by hallucinations that evade standard thresholds. To eliminate these errors, we introduce the Turing‑inspired Label Imitation Game (LIG), a framework that formalizes pseudo‑label pruning as an adversarial interrogation. Rather than filtering labels via isolated thresholds, we use the LIG to train a Turing Test Network (TTN), a task‑agnostic "judge" that evaluates candidate pseudo‑labels within a dataset‑wide context. Experiments across four diverse datasets demonstrate the TTN's robustness, consistently enhancing label accuracy for three state‑of‑the‑art vision‑language models without costly supervision or retraining. Crucially, we demonstrate that learned semantic‑contextual logic is a robust alternative to spatial‑geometric verification, enabling a unique zero‑shot task transfer capability ‑ a TTN trained strictly on image classification datasets can effectively prune complex object detection pseudo‑labels. This pruning yields F1‑score gains of 28% for the worst‑performing baseline categories and 44% with task‑specific fine‑tuning. Significantly, we also observe Category Revival, where the TTN pruning "detoxifies" the training signal for downstream models and enables them to recover from zero recall on transfer‑vulnerable classes. The pre‑trained TTN models and code are available at https://github.com/voxel51/ttn.
Authors:Yogeswar Reddy Thota
Abstract:
Current operating systems expose interfaces optimized for human users but not for AI agents. Humans benefit from pixels, icons, windows, visual grouping, mouse movement, and keyboard shortcuts; AI agents instead need compact semantic state, grounded actions, and reliable feedback. As a result, many computer‑use agents are forced to interpret screenshots, OCR output, and visual crops, introducing high token costs, visual ambiguity, latency, and coordinate uncertainty. This paper introduces LUMOS (Language Model Unified Machine‑Readable Operating‑System Semantics), a semantic interaction layer between AI agents and operating systems. LUMOS converts native accessibility metadata and browser UI structures into machine readable semantic blueprints with stable identifiers, roles, names, values, bounds, and action affordances. It also supports live semantic pointer grounding by querying the UI element under or near the cursor through operating‑system automation APIs. An LLM then acts through an accessibility grounded observe act loop using constrained visible‑UI primitives rather than application‑specific scripts. LUMOS does not claim to replace visual agents; instead, it reduces dependence on screenshots when operating systems already provide semantic structure. These results suggest a path toward AI‑native operating systems and machine‑readable interaction layers.
Authors:Mohamed el Amine Boudjoghra, Ivan Laptev, Angela Dai
Abstract:
Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and motion from sparse observations remains challenging. Existing approaches remain largely constrained by lack of supervised data or lack the priors needed to reliably recover articulation, hidden geometry, and internal object structure. We present the first debate‑driven agentic approach to articulated 3D object reconstruction from text or image inputs that both grounds articulation reasoning in concrete motion and exposes the occluded geometry revealed under articulation. High‑level agents reason about object semantics and motion using knowledge from vision‑language and video models, while low‑level agents estimate articulation parameters and interaction points; together, they engage in a two‑round structured debate that first exploits global‑‑local disagreement and then grounds the agents in freely generated video. The same video prior, conditioned on the agreed articulation, then drives each part through its motion to expose occluded interiors and geometry that cannot be inferred from a single static view. By combining agentic reasoning with a video generative prior, our approach jointly infers articulation and reconstructs complete 3D articulated objects, producing high‑fidelity geometry, internal structure, and motion‑consistent states beyond directly observed surfaces.
Authors:Sen Liang, Cong Wang, Zhentao Yu, Fengbin Guan, Zhengguang Zhou, Teng Hu, Youliang Zhang, Yuan Zhou, Xin Li, Qinglin Lu, Zhibo Chen
Abstract:
Existing instruction‑based video editing datasets commonly focus on single‑task appearance editing, failing to meet the complex creative demands of real‑world scenarios. To bridge this gap, we present Goku, a large‑scale dataset featuring 2 million high‑quality, instruction‑aligned video editing pairs, which is the first to extend task boundaries from basic appearance editing to multi‑task and structural manipulations(e.g., precise control of subject movement). To tackle the data synthesis challenges inherent in these complex tasks, we design an efficient data synthesis pipeline that decomposes complex edits into controllable sub‑problems and introduce a progressive filtering system for data reliability throughout the whole process. Furthermore, we explore the optimal network structures on Goku, and propose Goku‑Edit. To deeply comprehend complex editing instructions, Goku‑Edit leverages an MLLM as its text encoder and adopts a decoupled dual‑branch design: a dedicated mask branch handles structural control, freeing the main branch for appearance rendering. A comprehensive video editing benchmark, Goku‑Bench, is also proposed with 1,000 human‑verified test cases and 7 novel editing‑specific metrics. Evaluated on Goku‑Bench, Goku‑Edit obtains up to +8% improvement on other open‑source models in terms of instruction following.
Authors:Juntao Jiang, Jinsheng Bai, Linxuan Fan, Yali Bi, Jiangning Zhang, Yong Liu
Abstract:
We present APRIL‑MedSeg, a YAML‑driven modular framework for 2D medical image segmentation. It provides a unified and extensible ecosystem that decomposes segmentation networks into reusable components. Also, the framework integrates a broad spectrum of advanced paradigms, including semi‑supervised learning, domain adaptation, knowledge distillation, weakly supervised learning, and text‑guided segmentation as well as foundation model support. A registry‑based configuration system with inheritance enables flexible and reproducible experiment management, supporting seamless switching across models, datasets, and training strategies. In addition, the framework provides a unified interface for medical datasets, augmentation pipelines, deployment utilities and model ensembling. Overall, APRIL‑MedSeg is designed as a general‑purpose research and development platform that bridges algorithmic innovation and practical deployment, while also serving as a structured ecosystem for systematically organizing and reproducing advances in medical image segmentation. The code is available at https://github.com/juntaoJianggavin/APRIL‑MedSeg under an Apache 2.0 license.
Authors:Haoyang Li, Guanlin Li, Youhe Feng, Chen Zhao, Zhuoran Wang, Yang Li, Qizhe Wei, Shifeng Bao, Haitao Shen, Yihan Zhao, Tong Yang, Jing Zhang
Abstract:
Cross‑embodiment transfer in vision‑language‑action (VLA) models remains challenging because low‑level state and action spaces differ fundamentally across robot platforms. We observe that the high‑level cognitive process underlying manipulation, including scene perception, object identification, task planning, and sub‑task decomposition, is largely shared across embodiments. Based on this observation, we present ZR‑0, a 2.6 billion parameter end‑to‑end VLA model that uses dense Embodied Chain‑of‑Thought (ECoT) supervision to align cross‑embodiment representations within the vision‑language model (VLM). ZR‑0 adopts a dual‑stream architecture: a pre‑trained VLM (System 2) generates structured ECoT reasoning during training, while a Diffusion Transformer‑based action expert (System 1) produces continuous action chunks via flow matching. The two components are coupled through cross‑attention, with an attention mask that restricts the action expert to input prompt features only, enabling ECoT generation to be entirely skipped at inference without any performance loss. ZR‑0 is pre‑trained on ProcCorpus‑60M, a large‑scale dataset comprising approximately 60 million frames (approximately 1,000 hours) from over 400K trajectories, with dense ECoT annotations covering 96.8% of all frames. We evaluate ZR‑0 on three simulation benchmarks spanning single‑arm (LIBERO), bimanual (RoboTwin 2.0), and humanoid (RoboCasa GR‑1 Tabletop) embodiments, as well as real‑world experiments on the xArm platform, demonstrating strong performance across all settings. Code and model checkpoints are available at https://github.com/RUCKBReasoning/ZR‑0.
Authors:Deyin Liu, Jicheng Xu, Lin Yuanbo Wu, Xiaowei Zhao, Xiatian Zhu, Zhe Jin, Anjan Dutta
Abstract:
Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research interest. With the development of diffusion‑based/flow‑based video foundation models, existing animation works have began to upgrade the guidance information from 2D skeleton/pose to 3D modeling conditions. Despite achieving reasonable results, these approaches face challenges in synthesizing trajectory‑controllable human motion within natural scene under changed camera views. In this work, we present a scene‑adaptive human image animation framework that controls both human motion and camera trajectories within a reconstructed 3D environment for video generation. To achieve this, we first develop a ground‑adaptive 3D motion retargeting approach to enable user‑friendly motion trajectory control adapting to the changes of elevations of ground and orientations automatically. Then we design a viewpoint‑adaptive latent fusion mechanism to inject point‑cloud geometric priors through scene‑visibility masking into the generative process, providing precise guidance of viewpoint changes under camera control. Experiments on two standard human image animation benchmark datasets demonstrate remarkable improvements of our method over the state of the arts in related video generation metics. Project page: https://robinhood256100.github.io/web‑disp
Authors:Chengzeng You, Binbin Xu, Soteris Demetriou
Abstract:
The structural vulnerabilities of point cloud‑based 3D object detectors remain poorly understood. Prior work has studied adversarial robustness primarily on isolated 3D object models, while recent LiDAR spoofing attacks target richer and more realistic driving scenes but focus mainly on physical realizability rather than understanding detector behavior or attack efficiency. In this work, we investigate how LiDAR‑based detectors rely on spatial evidence in complex scenes and whether these reliance patterns can be exploited to induce failures more efficiently. To this end, we propose an explainability‑guided adversarial analysis methodology. We introduce the Saliency‑LiDAR (SALL) method, which aggregates Integrated Gradient attributions across scenes to produce universal saliency maps for LiDAR‑based 3D object detectors. Guided by these maps, we design the Explainability‑aware Frustum Attack (EFA), which selectively perturbs only the most influential frustums rather than uniformly attacking entire object regions. Experiments on KITTI and nuScenes, across detectors such as PointPillars and SECOND, show that EFA reduces detection recall by more than 15 percentage points while requiring 25‑50% fewer perturbed frustums than the state‑of‑the‑art non‑saliency‑aware baseline. These findings reveal that modern 3D detectors concentrate discriminative evidence in a small subset of spatial regions, exposing a structural robustness vulnerability in current LiDAR perception systems. Our code is released at https://github.com/SecMindLab/Saliency_LiDAR.
Authors:Yuan Li, Youyuan Lin, Zitang Sun, Yung-Hao Yang, Kiyofumi Miyoshi, Chenhui Chu, Shin'ya Nishida
Abstract:
Blind image quality assessment (BIQA) is commonly built on two basic learning paradigms: regression and ranking. Regression calibrates absolute scores, whereas ranking recovers quality structure from ordinal relations. Although joint regression‑ranking supervision often improves BIQA, the relation between the two paradigms remains largely empirical and underexplored. In this work, we revisit what underlies regression and ranking and identify pairwise relational distance, termed quality margin, as their common bridge. Our derivation shows that, at the objective‑optimization level, both paradigms fit quality margins: regression fits margins induced by score endpoints, while ranking fits transformed or sign‑level margins through preference probabilities. Motivated by this insight, we propose MR‑IQA, a direct quality‑margin optimization framework for reinforcement learning (RL)‑based BIQA. MR‑IQA samples quality scores and optimizes pairwise margin errors as policy rewards, thereby modeling quality structure more explicitly. Experiments on six BIQA benchmarks show competitive general performance, and controlled comparisons demonstrate that MR‑IQA achieves the strongest average PLCC/SRCC over regression‑ or ranking‑based RL methods. Our findings provide a new insight into unifying regression and ranking, offering a theoretical basis for understanding quality‑structure modeling in BIQA and beyond. Code is available at https://github.com/RobinY99/MR‑IQA.
Authors:Guang-Xing Li
Abstract:
Continuous physical fields represent a large fraction of data under scientific investigation. Their multiscale structures are central to discovery, yet useful coordinates are not known in advance. Standard self‑supervised methods define context and targets in fixed image coordinates, posing a predictive task misaligned with fields organized across a continuous scale hierarchy. We introduce ScaleAware‑JEPA, a framework that constructs dense, label‑free latent coordinates for continuous scalar fields. Constrained Diffusion Decomposition (CDD) separates each field into pixel‑registered scale components and provides the scale coordinates that define the masking geometry. The resulting JEPA objective predicts hidden structure with a context footprint tied to the diffusion scale of each component rather than to an arbitrary patch size. Across MHD turbulence, interstellar molecular gas and urban nighttime‑light structure, the learned geometry maps back to coherent morphology, forming dense structural atlases without labels or predefined segmentation rules. By tying latent prediction to the scale hierarchy of a field, ScaleAware‑JEPA constructs latent coordinates through which complex physical patterns can be inspected before their relevant structures have been prescribed. Code is available at https://github.com/gxli/SA‑JEPA.
Authors:Zhongqiang Song, Guanying Chen, Yuqi Zhang, Yin Zou, Chuanyu Fu, Zhiyuan Yuan, Chuan Huang, Shuguang Cui, Xiaochun Cao
Abstract:
This paper addresses the problem of monocular metric depth estimation in aerial UAV imagery. Although recent data‑driven methods have achieved remarkable progress in ground‑level scenarios, models trained primarily on street‑view and indoor datasets exhibit significant domain gaps when applied to aerial viewpoints. To tackle these challenges, we introduce AerialMetric, a benchmark dataset designed to evaluate and facilitate the adaptation of monocular metric depth estimation under UAV aerial viewpoints. The dataset consists of four complementary subsets collected from different sources, jointly covering real‑world photogrammetry data, controlled aerial acquisition settings, photorealistic synthetic scenes, and in‑the‑wild Internet imagery. Totally, AerialMetric provides 52K real‑world and 16K synthetic image‑depth pairs with reliable metric ground truth. Based on this dataset, we conduct systematic evaluations of existing state‑of‑the‑art models under aerial settings and investigate the impact of viewpoint, altitude, and camera parameters on metric depth prediction. In addition, by fine‑tuning representative metric depth model on our dataset, we establish a comprehensive aerial benchmark and achieve state‑of‑the‑art performance across diverse aerial imagery. Our dataset, code, and model weight are publicly available at https://kuieless.github.io/AerialMetric‑ECCV2026‑page/.
Authors:Sunqi Fan, Lingshan Chen, Runqi Yin, Qingle Liu, Yongming Rao, Meng-Hao Guo, Shi-Min Hu
Abstract:
Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to the domain of GUI agents, hoping to build strong GUI agents through a similar paradigm. However, GUI agent data cannot be directly harvested from the internet, making it costly and difficult to collect at scale. As a result, current GUI agents suffer from poor cross‑device generalization and limited visual grounding ability for fine‑grained GUI elements. As an attempt to address data challenge in GUI agents, we propose GUICrafter, a weakly‑supervised GUI agent leveraging massive unannotated screenshots to substantially reduce the reliance on expensive human annotations. GUICrafter explores a curriculum learning framework for training GUI agents through two progressive stages. First, the model learns visual grounding from large‑scale unannotated screenshots and webpages, leveraging the rich contextual signals inherent in GUI interactions without human annotations. Then, in Stage 2, we leverage a small amount of high‑quality data to calibrate the model via reinforcement learning. Experiments show that GUICrafter achieves competitive, or even superior, performance to advanced systems like UI‑TARS while using only 0.1% of its data. Furthermore, under the same amount of annotated data, GUICrafter surpasses all previous methods such as GUI‑R1. Code, data, and models are available at https://github.com/fansunqi/GUICrafter.
Authors:Hairui Chen, Yanwu Yang, Jianfeng Cao, Hanyang Peng, Chenfei Ye, Ting Ma
Abstract:
Brain networks exhibit a modular community structure that varies across individuals and neurological conditions. However, existing self‑supervised learning (SSL) methods often overlook this heterogeneity, relying on generic masking strategies that fail to capture subject‑specific functional organization. We propose BrainPICM, a self‑supervised framework for brain network analysis via progressive individualized community aware masking. BrainPICM formulates ROI‑to‑community mapping as a progressive unbalanced optimal transport process, yielding soft assignments and per‑ROI confidence scores. Guided by these confidence estimates, a curriculum‑style masking strategy gradually incorporates low‑confidence, potentially pathological regions into training, enabling the model to learn both stable modular structures and individual variations. Additionally, a deviation‑aware aggregation module quantifies functional reorganization by measuring mass redistribution relative to a population template, enhancing interpretability and downstream prediction. Experiments on three fMRI datasets (ABIDE‑I, ADHD‑200, ADNI) show that BrainPICM consistently outperforms state‑of‑the‑art supervised and SSL methods in diagnostic accuracy, indicating that explicitly injecting modular community structure into masked modeling yields more functionally consistent and generalizable representations. The source code for this approach will be released at https://github.com/Hrychen7/BrainPICM.
Authors:Zhengyuan Li, Zeyun Deng, Yifan Shen, Liangyan Gui, Miaolan Xie, Joseph Campbell, Xifeng Gao, Kui Wu, Zherong Pan, Aniket Bera
Abstract:
Self‑collision remains a persistent challenge in SMPL‑based human pose estimation and motion generation. Under extreme articulations or stochastic motion synthesis, generated meshes frequently exhibit self‑penetrations, leading to physically implausible results. We propose PoseShield, a neural collision constraint defined directly in SMPL pose space. We formulate collision correction as a constrained optimization problem and connect the learned constraint with the Eikonal equation. Enforcing Eikonal regularization ensures non‑vanishing gradients near the collision boundary, improving numerical stability and robustness of the optimization process. Unlike prior methods that operate in the mesh space or rely on heuristic penalties, our approach operates directly in the low‑dimensional space of human poses and is theoretically grounded. The same learned constraint extends to human motion sequences, providing a generator‑agnostic post‑hoc collision corrector without retraining the underlying motion model. Experiments on a newly constructed SMPL pose benchmark show that our method achieves a 95.8% success rate and outperforms state‑of‑the‑art baselines.
Authors:Piyush Arora, Navlika Singh, Umberto Cappellazzo, Stavros Petridis, Maja Pantic
Abstract:
Audio‑Visual Speech Recognition takes two input modalities, acoustic and visual streams, where visual information from lip movements aids recognition when audio is noisy. Recently, LLM‑based AVSR models have emerged as a promising paradigm by connecting pre‑trained audio‑visual encoders to an LLM, achieving strong results in clean conditions. However, these models are predominantly optimized for clean acoustic conditions, with limited attention to making the LLM backbone robust to noise. No explicit mechanism is employed to produce stable representations under corrupted audio, leading to performance degradation in noisy environments. To address this, we propose VIB‑AVSR, which integrates Variational Information Bottleneck layers at targeted positions within the LLM backbone to regularize representations. VIB‑AVSR reduces degradation under noisy conditions across multiple SNR levels and noise types, without requiring architectural modifications or additional training data.
Authors:Hang Su, Chao Sun, Zhaofan Li, Wei Hu, Juhua Liu, Bo Du
Abstract:
Vision‑language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recently emerged, the domain remains particularly challenging due to severe speckle noise, acquisition variability, and subtle anatomical boundaries, leading to high inter‑observer variability. Existing CLIP‑based models rely primarily on global image‑text alignment, limiting their sensitivity to clinically decisive local structures. We propose SonoCLIP, the first million‑scale region‑controllable fetal ultrasound vision‑language foundation model that integrates segmentation masks as mask‑channel visual prompts within the vision encoder, enabling joint global‑local contrastive representation learning. To support scalable region‑text alignment, we introduce a sigmoid‑based pairwise contrastive loss that improves stability under large‑scale supervision. We further curate a 1.44M‑image multimodal fetal ultrasound dataset spanning 24 standard planes for large‑scale pretraining. Extensive cross‑center evaluations demonstrate that SonoCLIP achieves superior zero‑shot transfer performance under both global and mask‑guided inference, establishing a controllable and clinically oriented foundation model for fetal ultrasound analysis. Our code and data are available at https://github.com/Harrison‑one/SonoCLIP.
Authors:Xiaomeng Fan, Wei Wu, Yuwei Wu, Zhi Gao, Shiyu Luo, Mingyang Gao, Haoyu Zhao, Zhenxin Diao, Yuxuan Ba, Lijia Feng, Yunde Jia, Mehrtash Harandi
Abstract:
Multimodal large language models (MLLMs) are increasingly expected to generate fine‑grained descriptions of visual content. However, we observe and theoretically show that generating fine‑grained responses poses a reliability challenge, i.e., fine‑grained generation is more error‑prone than coarse‑grained generation. This phenomenon suggests that models should generate the finest description that remains reliable rather than simply produce more specific outputs. To investigate this problem, we develop \textscGranFact, a granularity‑aware benchmark consisting of expert‑verified multi‑object images with coarse‑to‑fine category annotations. Then, we design a hierarchy‑aware evaluation algorithm, which assesses both whether model predictions are visually correct and how specific the correct predictions are. We also propose a reliability‑prioritized preference optimization method based on Direct Preference Optimization, which penalizes unreliable fine‑grained claims while rewarding reliable specificity. Experiments on \textscGranFact show that our method improves fine‑grained generation while preserving reliability. Code and data are available \hrefhttps://github.com/WeiWu2025/GranFacthere.
Authors:Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
Abstract:
We propose MotionAtlas, a system for detailed captioning of motion‑centric videos, comprising (1) a dedicated human‑annotated benchmark, (2) a scalable, high‑quality pipeline to construct training samples, and (3) a family of powerful Video‑MLLMs. Unlike conventional global motion captioning datasets, we focus on region‑aware motion captioning: given a video and a spatiotemporal mask, the model generates precise descriptions of motion within the target region, thereby alleviating visual clutter and motion entanglement and enabling reliable, quantifiable evaluation. Concretely, we first build MotionAtlas‑Bench, a comprehensive benchmark comprising 2,073 multiple‑choice questions, meticulously annotated for a curated set of high‑quality, motion‑centric videos, to evaluate fine‑grained motion understanding of the objects in question. Second, we design a rigorous and scalable data pipeline that leverages self‑bootstrap refinement to suppress fine‑grained hallucinations, yielding 159k high‑quality motion captioning data. Third, we design a tailored training data composition strategy, which achieves consistent and substantial performance gains across diverse baseline Video‑MLLMs, including Molmo2 and Qwen3‑VL. For instance, MotionAtlas‑4B surpasses Qwen3‑VL‑4B by an average of 5.2 percentage points across general motion benchmarks. The benchmark, dataset, and code have been released.
Authors:Mijin Yoo, In Cho, Subin Jeon, Jiwoo Lee, Eunbyung Park, Seon Joo Kim
Abstract:
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed‑forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object‑level structure to be recovered after the fact. We propose a feed‑forward framework that decomposes a scene into instance‑structured 3D token groups directly from unposed multi‑view images ‑‑ compact object‑centric units from which reconstruction, segmentation, and manipulation all follow. Each token group pairs an instance token capturing entity‑level identity with anchor tokens that encode local geometry and appearance, which are decoded into a set of 3D Gaussians. This two‑level factorization decouples object identity from local appearance, making object instances a native interface of the representation rather than a derived product. The token groups are learned through differentiable rendering with joint reconstruction and segmentation supervision, requiring no 3D annotations. Our feed‑forward model surpasses per‑scene optimization baselines in class‑agnostic instance segmentation while remaining competitive in novel view synthesis. Beyond these metrics, the same token groups directly unlock instance‑level scene editing ‑‑ removing, translating, or inserting objects by operating on their groups ‑‑ as well as efficient open‑vocabulary 3D instance retrieval, where retrieval complexity scales with the number of instances rather than primitives.
Authors:Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon
Abstract:
Vision‑language dataset distillation (VLDD) compresses a large image‑text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision‑language models under strict data and compute budgets. Most existing methods match expert trajectories or cross‑modal statistics, yet still enforce full‑dimensional alignment in a Euclidean embedding space. This is often overly restrictive due to rank‑deficient image‑‑text correlation, with shared semantics concentrated in a low‑dimensional range and remaining variation spread across a weakly correlated residual subspace. LoRS relaxes alignment at the similarity level by low‑rank factorization, but does not explicitly control dominant alignment capacity and structure in the representation space. We thus propose a rank‑aware hyperbolic alignment (RAHA) that combines hierarchical geometry with explicit alignment‑capacity control. RAHA lifts multimodal representations to hyperbolic space and optimizes distilled pairs with asymmetric objectives that enforce geodesic alignment in the shared range while regularizing the residual subspace to preserve modality‑private diversity and improve transfer robustness. Experiments on benchmarks show that RAHA demonstrates competitive cross‑modal retrieval and improved transfer indicators under fixed budgets.
Authors:Tuo Chen, Minjing Dong, Benlei Cui, Jian Liu, Jie Gui
Abstract:
Self‑supervised learning (SSL) pretrained models have become a dominant paradigm for visual representation learning, but they are vulnerable to backdoor attacks. Existing defenses struggle to defend against such attacks in a fully black‑box setting because they often require access to labels, attack patterns, or training data. To tackle this issue, we propose a new attack‑agnostic, model‑agnostic, and modality‑agnostic black‑box test‑time defense paradigm, called \emphPlatonic Representation Defense. It is inspired by the Platonic Representation Hypothesis, which suggests that large‑scale independently trained encoders converge toward compatible projections of the same underlying reality. We formalize this idea as a conditional energy function defined over source representations and a set of reference representations. The energy function is trained for detection through noise‑contrastive estimation and for representation purification through denoising score matching. Theoretically, the energy gap between matched and mismatched samples is lower bounded by the mutual information between source and reference representations. We demonstrate the effectiveness of our method on multiple self‑supervised encoders and more than 10 attacks. The method can perform both representation detection and purification, and achieves substantial performance gains across multiple attacks. Code is available \hrefhttps://github.com/jsrdcht/Platonic‑Representation‑Defensehere.
Authors:Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo, Shuojin Yang
Abstract:
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks. However, existing benchmarks primarily evaluate whether models can perceive shallow visual cues, while rarely examining whether MLLMs can learn deeper knowledge or procedural skills from video tutorials and generalize them to downstream long‑horizon agentic tasks. To address this gap, we introduce VG‑GUIBench (Video‑Guided GUI Benchmark), a new benchmark designed to evaluate whether MLLM‑based GUI agents can follow video tutorials to complete corresponding GUI interactive tasks. Furthermore, we observe that the performance of models on both VideoQA and video‑guided agentic tasks critically depends on effective keyframe extraction. Based on this observation, we propose TASKER (Task‑driven And Scene‑aware Keyframe searchER), a keyframe extraction algorithm that jointly considers task relevance and scene dynamics to identify informative frames. Experimental results demonstrate that TASKER achieves significant performance improvements on both VideoQA and video‑guided agentic task benchmarks, outperforming the best baseline by 2.0% on the EgoSchema fullset and 1.8% on the NExT‑QA dataset, respectively. These results further highlight the potential of generalized keyframe extraction methods for video understanding tasks. Our code and data are available at https://github.com/VG‑GUI‑TASKER/VG‑GUI‑TASKER.
Authors:Xuanhua Yin, Yuxuan Jia, Chuanzhi Xu, Weidong Cai
Abstract:
High‑resolution Diffusion Transformer (DiT) inference contains substantial spatial redundancy, but many spatially adaptive implementations encode regional computation as attention masks, which can inadvertently move scaled dot‑product attention (SDPA) away from FlashAttention fast paths. We identify this avoidable systems bottleneck as Mask‑Induced Dispatch Tax (MIDT) and show that it grows with latent sequence length. We introduce SAFE‑DiT, a training‑free Semantics‑Aware Fast‑path Execution framework that separates exact mask elision from approximation‑based spatial scheduling. SAFE‑DiT removes only provenance‑certified image self‑attention masks that induce a row‑wise constant shift in attention logits, preserves semantics‑bearing masks such as text‑padding masks, and realizes spatial adaptation through prompt‑conditioned token partitioning, selective state updates with global context, and periodic context refresh. We call this acceleration‑only configuration SAFE‑Core and report sensitivity‑weighted classifier‑free guidance separately as SAFE‑DiT+SW. On the evaluated PyTorch SDPA stack, redundant masks make long‑sequence attention 4.1× to 5.8× slower than the mask‑free path. On Lumina‑Next, SAFE‑DiT achieves 2.69× end‑to‑end acceleration at 1024^2 resolution and 5.09× at 2560^2, reduces peak memory at 2560^2 from 94.1 to 27.9 GB, and enables 3072^2 generation when dense inference runs out of memory. Paired metrics, component ablations, and a blinded human study support visual non‑inferiority of SAFE‑Core to the dense fast‑path baseline, while SAFE‑DiT+SW provides a separate prompt‑alignment operating point without reintroducing spatial self‑attention masks. Code is available at https://github.com/xuanhuayin/SAFE‑DiT.
Authors:Xiao Wang, Liye Jin, Dan Xu, Yuehang Li, Lan Chen, Yaowei Wang, Yonghong Tian, Jin Tang
Abstract:
Vision‑language tracking guided by natural language specifications leverages high‑level semantic cues of target objects to substantially boost tracking accuracy and robustness. Existing studies have verified that adaptively optimizing textual descriptions throughout the tracking process can effectively mitigate the semantic‑visual mismatch induced by dynamic variations in target appearance, position, and other inherent attributes. Nevertheless, mainstream methods that directly generate textual information via sequence models or large language models inevitably suffer from inherent defects, including erroneous target updating, excessive background distraction, and pervasive hallucination artifacts. To address the aforementioned limitations, this paper proposes a novel language dependency parsing mechanism to precisely distill core tracking principal components, encompassing target objects, semantic concepts, and background contextual information. On this basis, we perform component‑aware adaptive textual description updates by exploiting the powerful cross‑modal understanding capability of the pre‑trained vision‑language model Qwen‑VL. By integrating the proposed elaborately designed modules into the baseline framework, our method achieves consistent and superior tracking performance on multiple large‑scale vision‑language tracking benchmarks, including TNL2K, LaSOT, TNLLT, and OTB‑LANG. The source code and pre‑trained models will be released at https://github.com/Event‑AHU/Open_VLTrack.
Authors:Yiming Jiang, Hanzhang Tu, Wenfeng Song, Siyou Lin, Liang An, Shuai Li, Aimin Hao, Yebin Liu
Abstract:
Uncalibrated volumetric video streaming for human reconstruction is essential for holographic communication and AR/VR, yet remains challenging due to the need for temporal consistency and computational efficiency from sparse‑view inputs. Existing methods rely on per‑scene optimization or calibrated cameras, while recent feed‑forward models are limited to low‑resolution (0.5K) single‑frame synthesis. We present HiReFF, a feed‑forward method for 2K‑resolution 360° human video reconstruction from uncalibrated sparse‑view videos. Our framework decomposes the problem into two key tasks: foreground 3D Gaussian reconstruction from sparse‑view videos (four views separated by 90°) and computationally efficient high‑resolution synthesis. To enable the former, we propose Scale‑synchronized Camera Calibration to resolve scale ambiguity for multi‑view supervision, and Gaussian‑wise Foreground Masking to reconstruct clean foregrounds by modulating Gaussian parameters. For efficient high‑resolution synthesis, our High‑resolution Side‑tuning achieves 2K rendering by augmenting the Gaussian head with supplementary features while keeping the backbone at 0.5K, drastically reducing computational overhead. Experiments demonstrate that HiReFF significantly outperforms existing methods in high‑resolution streaming volumetric video reconstruction. https://iridescentjiang.github.io/HiReFF
Authors:Aymen Mir, Riza Alp Guler, Jian Wang, Peter Wonka, Bing Zhou, Gerard Pons-Moll
Abstract:
We study the problem of physically plausible shadow casting when animating 3D Gaussian Splatting (3DGS) avatars, either individually or in multi‑avatar and object‑interaction scenarios, within existing 3DGS scenes. In contrast to prior methods that rely on binary hit tests and mesh‑based shadow casters, our method performs shadow computation entirely in Gaussian space, without requiring any mesh reconstruction. We introduce RAGA, a Ray‑Traced Gaussian Shadow Casting formulation based on exact ray‑Gaussian line integrals. For each occluding Gaussian, we integrate the opacity profile along the shadow ray and normalize by the theoretical maximum integral, producing a weight that captures how the ray traverses the occluder rather than merely whether an intersection occurred. To reduce temporal variance from clothing deformations in animated avatars, we further introduce an avatar proxy representation that stabilizes shadow casting while preserving visual fidelity. We implement RAGA using custom CUDA kernels integrated with the NVIDIA OptiX framework; as such, our shadow tracer runs at rates of about 50 FPS. We evaluate on single‑avatar, multi‑avatar, and avatar‑object interaction scenarios across multiple datasets, demonstrating substantially improved shadow realism, temporal stability, and scene coherence. Our project page is available at https://miraymen.github.io/raga/.
Authors:Zhihong Liu, Zheng Li, Jiachun Jin, Siqi Kou, Yitao Jian, Fengpei Yu, Zhijie Deng
Abstract:
While text‑guided image editing has made remarkable progress, it remains limited in structural portrait retouching. Textual descriptions struggle to convey fine‑grained changes to facial features and body proportions. To address this gap, we introduce Exemplar‑Based Portrait Photo Retouching, where the model is given an exemplar pair and tasked with inferring and applying the same retouching operations to a new query image. Existing exemplar‑based editing methods primarily focus on tasks with pronounced visual transformations. In contrast, structural portrait retouching involves extremely delicate and localized modifications, making accurate extraction and transfer of these edits challenging. To tackle this, we propose MirrorPPR, a novel framework designed to capture and transfer subtle structural retouching operations. Our method uses a Retouching Operation Extractor to capture the subtle differences from the exemplar pair. The extracted representations are then injected into a pre‑trained Diffusion Transformer (DiT) through a connector and Low‑Rank Adaptation (LoRA) modules. Furthermore, constructing perfectly aligned cross‑identity training pairs is severely hindered by operation misalignment. To overcome this, we propose an advanced data self‑augmentation paradigm that ensures strictly aligned retouching operations. To alleviate data scarcity and support this novel task, we introduce MirrorPPR47M, a large‑scale dataset with over 47 million retouched pairs. By structuring the dataset into simulated and professional subsets, we enable progressive curriculum learning to smoothly optimize the network. Extensive experiments demonstrate that MirrorPPR significantly outperforms existing baselines in both retouching quality and identity preservation. The project page is available at https://sjtu‑deng‑lab.github.io/MirrorPPR.
Authors:Dacheng Qi, Chenyu Wang, Jingwei Xu, Yi Ma, Shenghua Gao
Abstract:
Computer‑aided design (CAD) plays a fundamental role in modern manufacturing by providing the high precision required for industrial production. Recent large language model based approaches formulate CAD generation as a sequence prediction problem and have achieved promising results. However, existing methods and evaluation protocols primarily emphasize visual similarity, while overlooking precise geometric parameters and correct metric scale. Small numerical deviations that are negligible at the shape‑level may still violate industrial tolerance requirements, a problem further compounded by current autoregressive paradigms that utilize command sequence representations, aggressively quantize numerical parameters to ease LLM prediction. In this work, we present Pointer‑CAD v2. Compared with v1 (arXiv:2603.04337), this version directly predicts continuous values, bypassing the need for quantized numerical parameters and thereby eliminating quantization errors. Specifically, we propose a unified framework that decouples parameter reasoning from geometric construction through a Plan‑Then‑Construct paradigm. Our method first produces a structured design plan with explicit metric scale parameters. These parameters are organized into a dictionary and directly referenced during sequence generation via a pointer mechanism, eliminating discretization errors and ensuring dimensionally consistent execution. In addition, we construct a new large‑scale dataset with plan‑level annotation and introduce three hierarchical geometry accuracy metrics to evaluate parametric fidelity at the vertex, edge, and face levels. Extensive experiments demonstrate that Pointer‑CAD v2 consistently outperforms existing baselines and achieves substantial improvements in geometric accuracy, enabling reliable CAD generation for precision‑critical engineering applications.
Authors:Dingyi Yao, Xinqi Zhang, Lihui Peng, Jianming Hu, Danya Yao, Yi Zhang
Abstract:
Synthetic data mitigates the data scarcity problem in autonomous driving perception. However, the synthetic‑to‑real gap leads to performance degradation, hindering real‑world model generalization. Although current methods leverage diffusion models for photorealistic style transfer to bridge this gap, they critically ignore a practical asymmetry: while synthetic data possesses perfect pixel‑level annotations, real‑world style reference images generally lack corresponding labels. Consequently, existing methods relying on symmetric semantic guidance suffer from either prohibitive annotation costs or severe semantic misalignment. To address this dilemma, we formally propose a novel task: Asymmetric Style Transfer for Autonomous Driving (ASTAD), which requires semantically consistent transfer using only labeled synthetic content and unlabeled real‑world references. We further introduce the ASTModel, a training‑free two‑stage framework designed to bridge this domain gap under asymmetric constraints. ASTModel first extracts a coarse semantic prior from the unlabeled target, followed by dynamic prior refinement and class‑consistent style injection during the denoising process. Extensive experiments demonstrate that ASTModel significantly outperforms existing methods in downstream perception utility and structural fidelity, while offering a 3.2× inference speedup. This work aligns synthetic‑to‑real adaptation with practical constraints, holding the potential to accelerate the scalable deployment of robust autonomous driving systems. Code: https://github.com/Dingyi‑Yao/ASTAD.
Authors:Cong Wang, Haiyu Wu, Zhiwei Jiang, Zifeng Cheng, Fei Shen, Yafeng Yin, Qing Gu
Abstract:
Concept erasure aims to prevent image generative models from producing unsafe content while preserving their general generative capability. Meanwhile, next‑scale autoregressive (AR) image generation has recently emerged as a new generative paradigm characterized by next‑scale prediction, for which concept erasure remains largely unexplored. In this paradigm, semantic information is highly compressed at early scales, leading to severe entanglement between unsafe and unrelated semantics. In this paper, we propose ScaleErasure, an inference‑time concept erasure method that performs minimal intervention. ScaleErasure precisely selects and guides predicted logits that are most relevant to the unsafe concept, thereby enabling effective erasure under severe semantic entanglement. Specifically, ScaleErasure performs two additional forward passes conditioned on the unsafe concept and the corresponding safe concept, and leverages their outputs to guide the target logits away from unsafe concepts toward safe concepts. To enable precise and minimal intervention, logits selection and guidance are conducted across three dimensions: scales, tokens, and bit channels. Experiments demonstrate that ScaleErasure outperforms adapted baselines in the next‑scale AR paradigm, achieving more precise concept erasure while largely preserving general generative capability. The code is available at https://github.com/coziiizz/ScaleErasure.
Authors:Abdullah Al Shafi, Md Kawsar Mahmud Khan Zunayed, Safin Ahmmed, Sk Imran Hossain, Engelbert Mephu Nguifo
Abstract:
Jointly learning to segment and classify medical images demands cross‑task synergy, yet encoder‑sharing architectures limit decoder reconstruction to task‑private representations, permanently discarding the boundary cues and semantic priors each branch could supply to the other. This work introduces BTI‑Net, which establishes bidirectional communication at every decoder level through two parallel pathways via Task Interaction Modules (TIM). Spatial boundary context is gated into the classification branch, while global semantic priors multiplicatively modulate the decoder, with refined features propagating progressively from coarse semantics to fine boundary detail across all four decoder resolutions. Since cross‑task interaction is not equally reliable for every input, Uncertainty Proxy Attention (UPA) gates each TIM output per instance and per level using three signals that capture cross‑task alignment, scene complexity, and prediction confidence, without external annotations or additional inference passes. Experiments on three medical benchmarks spanning ultrasound, dermoscopy, and brain MRI demonstrate consistent improvements in segmentation IoU and classification accuracy over both encoder‑sharing and decoder‑interaction baselines. Ablation confirms adaptive gating contributes +2.36 IoU over fixed bidirectional interaction, and classification accuracy improves by up to +2.26 points over the strongest multi‑task baseline. UPA's uncertainty proxies serve as reliable single‑pass task‑failure signals without the overhead of stochastic sampling. Code: https://github.com/C‑loud‑Nine/BTI‑Net_MTL
Authors:Darian Fernández-Gutiérrez, Rafael Bello, Marilyn Bello, Natalia Díaz-Rodríguez
Abstract:
Concept‑based Explainable AI (C‑XAI) seeks human‑understandable explanations grounded in semantic concepts, yet validation is limited by the scarcity of fine‑grained concept annotations. We evaluate whether mid‑scale Multimodal Large Language Models (MLLMs) can perform localized concept naming under strict zero‑shot conditions by assigning labels to bounding‑box regions at both object and part levels. We propose a reproducible zero‑shot evaluation protocol for Concept Naming (CoNa) with (i) closed‑set, category‑constrained prompting for moderate vocabularies and (ii) Open‑CoNa, an embedding‑similarity‑based strategy for large label spaces. Experiments with four MLLMs (7B‑32B) show consistent performance trends across datasets, reaching 62%‑88% object‑level exact‑match accuracy, highlighting the potential of training‑free concept annotation from localized regions. We discuss limitations and failure modes and release a reproducible framework to support future low‑cost C‑XAI research.
Authors:Yang Guo, Zihan Yang, Feifei Kou, Yulan Hu, Ran Zhang, Siyuan Yao
Abstract:
Small Object Detection (SOD) is a fundamental yet challenging problem in computer vision due to its limited spatial resolution and weak visual cues. Although recent approaches have achieved remarkable advances, the background distractors in different frequency spectra still degrade the performance. In this paper, we propose a novel small object detection framework termed SFDNet, which is capable of detecting small objects via efficient spectrum‑aware feature disentanglement. Specifically, we propose an Adaptive Spectrum Disentanglement (ASD) module that decomposes backbone features into multiple complementary spectral components, aiming to construct discriminative object‑relevant representations by discarding the background distractors for each component. Afterwards, to strengthen the semantic consistency of the similar objects in the same class, we propose a Class‑Wise Prototype Distillation (CPD) procedure, which establishes class prototypes for the object instances and enforces the compact representation by efficient prototype distillation. Extensive experiments on multiple challenging benchmarks show that SFDNet outperforms existing state‑of‑the‑art methods by a large margin. Our code is available at https://github.com/ManOfStory/SFDNet.
Authors:Chenghao Qian, Nedko Savov, Lingdong Kong, Yeying Jin, Rui Song, Wenjing Li, Zhun Zhong, Jiaqi Ma, Gustav Markkula, Luc Van Gool
Abstract:
Weather synthesis aims to add weather effects to input videos while preserving scene identity, structure, and motion. The key limitation of existing methods is the lack of diversity in weather appearance and effective control over weather dynamics (e.g., temporal evolution and particle motion). Most approaches rely on text prompts, which are inherently underspecified and often fail to produce detailed weather characteristics. Additionally, general‑purpose video editors optimized for clean and aesthetic outputs tend to suppress heavy weather phenomena, making dense particle effects difficult to generate. To address these, we propose a Semantic‑Aware, Physics‑Informed, and Geometry‑Grounded framework that steers an off‑the‑shelf video editor to synthesize diverse global appearances and detailed particle dynamics. We factorize the synthesis into three conditional signals, so that each provides a distinct and stable source of guidance: semantics specifies what the weather should look like, dynamics governs how it evolves over time, and geometry determines where it should appear in the scene. Specifically, we introduce (1) semantic‑aware appearance anchoring to establish the target appearance from scene semantics and user input; (2) physics‑informed dynamic simulation to generate particle effects by simulating a Gaussian‑represented particle field under gravity, wind, and turbulence; and (3) geometry‑grounded video synthesis to align the simulated particles with target scene geometry and synthesize the final video. Experiments demonstrate that our method produces diverse, physically and visually realistic weather effects. Furthermore, we show that our synthesized data significantly improves the robustness of autonomous driving semantic segmentation under adverse weather conditions. Project page: https://jumponthemoon.github.io/w‑crafter/.
Authors:Satyasa Khadka, Sandhya Baral, Sudip Tiwari, Sharad Kumar Ghimire
Abstract:
This paper presents a robust Automatic Number Plate Recognition (ANPR) system tailored for Nepali license plates written in Devanagari script. In this paper, a pipelined model was used that integrates YOLO‑based models for license plate and character detection, followed by a CNN classifier trained on 34 Devanagari characters. Two publicly available data sets were used that incorporate diverse lighting, fonts, and structural variations. Data augmentation and additional training on embossed plates enhanced the generalizability of the model. The system achieved a recognition accuracy of up to 93%, demonstrating strong performance under real‑world conditions and providing a scalable solution for traffic management in Nepal. Code: https://github.com/Satyasakhadka/Nepali‑NumberPlate‑Character‑Recognition
Authors:Deepayan Das, Davide Talon, Yiming Wang, Massimiliano Mancini, Elisa Ricci
Abstract:
Personalizing Multimodal Large Language Models (MLLMs) aims to recognize users' unique concepts from visual data and provide personalized responses. Although prior work has shown the benefit of concept descriptions and reasoning for this task, MLLM descriptions often include information, such as state and context, that does not help and may in fact hinder the unique identification of the target concept among other visually similar items. Effective descriptions of personal concepts should instead be accurate, discriminative, and free of distracting details. To achieve such descriptions, we introduce Reinforced Reference Game (RRG), a learning framework that promotes discriminative descriptions through a novel reinforced multimodal reference game. The MLLM plays both the roles of speaker and listener in a contrastive game setting, whose goal is to effectively communicate discriminative information about a target concept. Our approach formulates a verifiable contrastive reward over hard positives (dissimilar views of the same concept) and hard negatives (visually similar but different concepts). Empirically, RRG achieves state‑of‑the‑art across multiple tasks on three personalization benchmarks. RRG generalizes to unseen domains and outperforms existing methods based on concept descriptions and personalization‑specific RL frameworks. We will release code and models in the project page.
Authors:Zhihui Ke, Yuyang Liu, Xiaobo Zhou, Tie Qiu
Abstract:
3D Gaussian Splatting~(3DGS) has emerged as a promising paradigm for reconstructing streamable free‑viewpoint video~(FVV) from multi‑view videos. However, 3DGS‑based FVVs typically lack user interaction and editing capabilities, which diminishes the immersive experience. Recent research has integrated language features from CLIP into 3DGS via distillation, enabling open‑vocabulary queries and supporting many downstream applications. Nevertheless, the stringent requirements of FVV, low frame size and high FPS, make current language Gaussian representations unsuitable for language‑embedded FVV. In this paper, we propose DLGStream, a novel language‑embedded FVV representation that streams time‑varying language features alongside Gaussian attributes to support 4D environment interaction, scene editing, and spatial intelligence. Specifically, we propose a dual‑opacity dynamic language Gaussian representation, which maintains two opacity attributes for color and language features to deal with performance degradation that occurs when colors and features are jointly optimized. Furthermore, we introduce an interpolation‑based deformation field to reduce temporal redundancy. This deformation field can also be used for 4D frame interpolation, boosting FVV sequences from low to high FPS. Experimental results demonstrate that DLGStream achieves superior performance in both on open‑vocabulary segmentation and reconstruction quality with an average frame size of merely 43 KB. The code is available on \hrefhttps://github.com/kkkzh/DLGStreamhttps://github.com/kkkzh/DLGStream.
Authors:Minh Son Hoang, Dinh Phu Tran, Quyen Nguyen Duc, Dam Hoang Phuong, Daeyoung Kim
Abstract:
Diffusion prior‑based methods have shown impressive results in real‑world image super‑resolution (ISR), yet two key challenges persist: balancing pixel‑level fidelity with semantic quality, and adapting to diverse degradations. Existing dual‑branch approaches freeze the pixel module during semantic training, but the semantic branch can still expand capacity within the pixel subspace, precluding genuine perceptual improvement. Moreover, using a single static adapter cannot generalize across heterogeneous real‑world corruptions. To address both issues, we propose FreqOrtho‑SR, which comprises: Frequency‑guided Mixture of LoRA Experts (FreqMoE), it routes inputs to specialized experts via a non‑parametric FFT‑based degradation‑feature extractor that encodes frequency‑domain signatures, enabling stable and interpretable specialization across corruption types; and Orthogonal Gradient Projection (OGP), which reframes the dual‑objective optimization as a subspace‑constrained problem: by extracting the pixel‑fidelity subspace via SVD on combined expert weight deltas and projecting semantic gradients onto its null space, OGP guarantees orthogonality between the two objectives, enabling genuinely complementary learning without mutual interference. Experiments show that FreqOrtho‑SR achieves competitive overall performance and a strong fidelity‑perception trade‑off across multiple benchmarks with efficient single‑step inference. The source code of our method can be found at \hrefhttps://github.com/sonhm3029/FreqOrtho‑SR\textttsonhm3029/FreqOrtho‑SR.
Authors:Ximiao Zhang, Min Xu, Xiuzhuang Zhou
Abstract:
Current anomaly detection methods primarily focus on structural anomalies, while paying insufficient attention to anomalies that violate logical constraints. Conversely, top‑performing logical anomaly detection approaches address this by modeling global semantic consistency, but perform poorly on subtle structural anomalies due to inadequate detection granularity. In this paper, we propose LogiCo, a unified framework for Logical and structural anomaly detection via Component‑level feature reconstruction. Unlike existing methods that rely on explicit global semantic modeling, LogiCo employs a novel component‑level feature reconstruction technique to capture inter‑component logical constraints. Specifically, LogiCo maps pre‑trained image features into a discrete component‑level feature space and performs collaborative feature reconstruction at both component and patch levels, enabling it to effectively detect both logical and structural anomalies. Furthermore, to address the specific challenge of count‑related logical anomalies, we integrate a segmentation‑map discriminator that extends the model's capability to identify quantitative inconsistencies. LogiCo achieves state‑of‑the‑art performance on both logical and structural anomaly detection across four benchmarks, including MVTec‑LOCO, MVTec‑AD, VisA, and Real‑IAD, demonstrating its superiority and practical feasibility. The code is available at https://github.com/cnulab/LogiCo.
Authors:Ruitao Chen, Mozhang Guo, Jinge Li
Abstract:
Deformable 3D Gaussian Splatting (3DGS) has emerged as an efficient approach for rendering dynamic scenes in a wide range of 3D applications. However, existing deformation field‑based approaches largely lack explicit object‑level modeling, often resulting in inconsistent Gaussian deformations within individual objects and unwanted coupling between different objects. To address this limitation, we introduce a semantics‑guided framework that enforces dynamic regularization at the object level, aiming to achieve spatially consistent object‑wise deformation. Specifically, we first extract segmentation masks using the Segment Anything Model (SAM) and derive semantic features from input images. An object‑ID map is then constructed via feature relevance matching with a predefined object dictionary. Guided by this object‑ID map, we identify the pixel‑wise top‑k contributing Gaussians for each object and impose consistency regularization on their deformation parameters, including position, scale, and rotation. Unlike prior methods that learn deformation fields without explicit object‑level constraints, our approach incorporates semantic cues to guide deformation behavior at the object level. Experimental results demonstrate that our semantics‑aware regularization improves object‑level deformation consistency and outperforms baseline methods in rendering quality, achieving higher PSNR and SSIM and lower LPIPS in dynamic 3DGS rendering. Our project page is available at https://dyn‑reg‑3dgs.github.io/.
Authors:Mohamed Shawky Sabae, Philipp Langsteiner, Jan-Niklas Dihlmann, Hendrik Lensch
Abstract:
Inverse rendering requires separating illumination from surface materials, which is highly ambiguous due to their tight coupling in observed images. While Gaussian Splatting is efficient for novel view synthesis, existing relightable methods approximate scene lighting using discrete point lights, global environment maps, or implicit representations. By ignoring the physical spatial extent of real‑world emitters, these approaches produce incorrect light attenuation and unrealistic shadows. We present AEGIR (Area Emitters for Gaussian Inverse Rendering), a framework that explicitly models local area emitters within a relightable Gaussian Splatting representation. Joint optimization of emitters, materials, and geometry is challenging due to flexible emitter parameterization, which increases both the number of parameters and the ambiguity between illumination and materials. We address this by introducing a differentiable deferred rendering pipeline that integrates multiple importance sampling with targeted regularization. As a result, AEGIR accurately simulates local light transport and achieves more consistent decomposition. Experiments show that explicit area emitters improve illumination reconstruction and enhance downstream tasks, including novel view synthesis, controlled relighting, and virtual object insertion, particularly in scenes with complex local lighting.
Authors:Jun Wang, Peirong Liu
Abstract:
Forecasting longitudinal brain lesion evolution is critical for disease monitoring and treatment planning. Existing approaches typically learn a direct mapping from a baseline image to a future observation, without explicitly modeling the physical mechanisms underlying the lesion progression. Such an entangled modeling of structural deformation and image intensity variation limits physical plausibility, model generalization, and interpretability. To address this, we propose PDF, a Physics‑grounded Disentangled Flow matching framework for longitudinal brain disease forecasting. We explicitly decompose the longitudinal modeling of lesion growth into two processes, each learned by a dedicated flow matching network: morphology evolution, which captures lesion growth and structural deformation; and intensity evolution, which models signal changes driven by variations in lesion concentration. To enforce physics‑grounded constraints, we introduce a PDE‑regularized loss based on lesion growth dynamics, that enforces a diffusion‑reaction‑advection formulation for morphological evolution. Experiments on three public longitudinal datasets spanning diverse brain diseases demonstrate state‑of‑the‑art performance, validating the effectiveness of the disentangled modeling framework and physics‑grounded learning design. Code is publicly available at https://github.com/jhuldr/PDF.
Authors:David Charatan, Daniel Xu, Richard Szeliski, George Kopanas, Vincent Sitzmann
Abstract:
Differentiable rendering has emerged as a powerful approach for 3D reconstruction and novel view synthesis. State‑of‑the‑art differentiable rendering methods combine a variety of custom representations of 3D geometry and appearance with specialized renderers. However, most downstream tasks in computer graphics rely on 3D meshes. While prior work has attempted differentiable rendering with mesh representations, these approaches are limited to object‑centric scenes and fail to reconstruct large‑scale, unbounded scenes. In this work, we introduce Meshtryoshka, a novel mesh differentiable rendering framework that combines an off‑the‑shelf triangle rasterizer with a 3D representation that consists of nested mesh shells which resemble a matryoshka doll. In every forward pass, the mesh shells are extracted anew from a 3D signed distance function via iso‑surface extraction, and the opacities for each vertex are computed as a function of signed distance. Each mesh shell is then rasterized independently, and the final image is created via alpha compositing. Crucially, mesh vertex positions are updated only indirectly via gradients that flow through the opacity values into the signed distance function, and hence, our method is compatible with off‑the‑shelf mesh renderers that need not be differentiable with respect to vertex positions. On object‑centric scenes, our method performs competitively with surface‑based differentiable rendering techniques. Our differentiable mesh rendering method scales to unbounded, real‑world 3D scenes, where it yields high‑quality novel view synthesis results approaching those of state‑of‑the‑art, non‑mesh methods. Our method suggests that it may be possible to solve the differentiable rendering problem without relying on specialized renderers, only using conventional tools from the computer graphics toolbox.
Authors:Anya Ji, Abhijith Varma Mudunuri, David M. Chan, Alane Suhr
Abstract:
While recent vision‑language models (VLMs) have achieved significant improvements on static visual‑to‑code tasks such as generating code for webpages, charts, or SVGs, it remains unclear whether they can recover temporal dynamics when motion is present. To this end, we introduce Animation2Code, a benchmark for evaluating temporal visual reasoning via reconstructing executable web animation code from videos. Animation2Code consists of 1,069 web animation videos with diverse visual appearances and motion patterns, paired with corresponding HTML/CSS/JavaScript implementations. We propose two human‑aligned metrics, appearance similarity and temporal similarity, which allow us to disentangle visual fidelity from temporal alignment when comparing rendered animations against ground‑truth samples. Benchmarking state‑of‑the‑art VLMs on this dataset shows that current VLMs struggle to maintain temporal consistency in reconstruction, even when achieving high appearance similarity, including under finetuning and iterative refinement settings. Code and data are available at https://anya‑ji.github.io/animation2code‑website .
Authors:Shuang Song, Jiyong Kim, Rongjun Qin
Abstract:
High‑resolution satellite imagery demands 3D reconstruction methods that deliver both speed and geometric accuracy. Recent adaptations of 3D Gaussian Splatting (3DGS) to satellite imagery demonstrate strong efficiency, but reconstruction quality often degrades under diverse illumination across multi‑date, high‑altitude acquisitions (with small intersection angles), limiting applicability to remote sensing and vision tasks. We present SatSplat, the first framework to adapt 2D Gaussian Splatting (2DGS) to satellite photogrammetry, with online camera adjustment. We approximate satellite cameras with an affine model and learn a minimal delta parameterization for in‑splat camera refinement from dense observations. The formulation is implemented with a 2DGS scene representation. To handle time‑varying shadows and illumination changes, we integrate geometric shadow mapping and per‑camera color correction during training. Across the evaluated DFC2019 and IARPA2016 benchmark sites, SatSplat achieves strong geometric accuracy while significantly outperforming prior 3DGS‑based baselines. On our processed DFC2019 benchmark, SatSplat reduces mean absolute error by 11.93% and peak video memory by 31% relative to the previous state of the art. Our approach enables large‑scale digital surface modeling with practical computational efficiency. The project page is available at https://gdaosu.github.io/satsplat/.
Authors:Yuexi Du, Leya Barrientos, Laura Sheiman, John Lewin, Hemant D. Tagare, Nicha C. Dvornek
Abstract:
Multiview mammography relies on paired craniocaudal (CC) and mediolateral oblique (MLO) views to provide complementary projections of a 3D breast volume, enabling precise anomaly localization. However, acquiring high‑quality, balanced datasets remains challenging for deep learning applications. We propose a novel method to synthesize multiview mammograms by leveraging the inherent geometric relationship between CC and MLO views. To enforce an implicit 3D consistency prior during generation, we develop an alignment module that searches a 2D affine transformation subspace to establish optimal anatomical correspondence. Leveraging this alignment, we introduce a pixel‑space self‑consistency loss based on the Earth Mover's Distance (EMD) between the 1D anteroposterior (AP) axis tissue distributions of the generated images. Integrated into a pretrained flow matching model, MammoFlow forces synthesized pairs to share physically plausible tissue distributions from the chest wall to the nipple. To our knowledge, this is the first work to guide multiview mammogram generation using implicit geometric tissue correspondence. Our method demonstrates superior image quality, passes expert radiologist evaluation, and generates physically consistent pairs that improve downstream classification AUC by 5%. Code is available at https://github.com/XYPB/MammoFlow
Authors:Xiao Song, Haonan Qin, Zhaoxu Zhang, Jiong Zhang, Yuqi Fang, Caifeng Shan
Abstract:
Large vision‑language models (LVLMs) are increasingly used for clinical image understanding, yet they remain vulnerable to \emphhallucinations‑‑producing textual findings or attributes not supported by the image. We present a vision‑traceable hallucination detection framework that audits arbitrary LVLM responses via visual evidence grounding, requiring neither modification nor internal access to the hidden states of LVLMs. Given an LVLM response, we extract visually verifiable entities and use a medical‑domain‑adapted Qwen‑VL grounding verifier to localize each entity on the input image. To enhance the robustness of our detection method, we introduce a counterfactual entity perturbation method and estimate visual evidence uncertainty by contrasting factual and counterfactual grounding results. Specifically, we compute an entity‑level uncertainty score from the positive confidence, counterfactual confidence, and their grounding overlap for binary hallucination decision‑making. Experiments on multiple medical imaging modalities and LVLM backbones demonstrate that our method consistently improves hallucination detection performance over recent baselines, while providing interpretable localization evidence and strong cross‑model transferability. Code and dataset are available at https://github.com/Agentic‑CliniAI/CounterVHD.
Authors:In Kyu Lee, Sumin Seo, Jaesik Min
Abstract:
Accurate correspondence matching across multiple angiographic views is the prerequisite for 3D coronary reconstruction and interventional guidance. However, the development of robust deep learning models for this task has been stifled by a fundamental data bottleneck. Obtaining ground truth for matching tasks in angiography pairs is prohibitively expensive and hard to scale. To overcome this barrier, we introduce a physically‑grounded data generation framework that synthesizes high‑fidelity Digital Reconstructed Radiographs (DRRs) from 3D Coronary CT Angiography (CCTA) volumes. Our framework generates dense, highly accurate 3D‑to‑2D projection labels by simulating realistic C‑arm acquisition geometry on patient anatomy at zero human cost. Leveraging this dense supervision, we propose a Geometry‑Informed Matching Module (GIMM) that integrates global feature and anatomical structure into correspondence learning. Unlike real angiography where assessment relies on subjective human annotation, our dataset provides 2D correspondence labels with paired images, allowing human‑free evaluation. We comprehensively evaluate our method on the proposed CT‑derived DRR dataset and demonstrate improvements over other matching baseline models.
Authors:Jia-Wei Liao, Li-Xuan Peng, Mei-Heng Yueh, Min Sun, Cheng-Fu Chou, Jun-Cheng Chen
Abstract:
Recently, diffusion models have been widely adopted in generative modeling and have served as foundational models for many image generation tasks. To control the generation without costly re‑training or fine‑tuning, many works seek inference‑time guidance methods to steer the latent via a differentiable objective at inference time. However, these methods cannot effectively preserve the original Gaussian distribution because they introduce distributional drift, thereby degrading the sample quality. To address this gap, we propose DiffRGD, a distribution‑aware guidance framework that explicitly preserves the latent Gaussian structure. DiffRGD formulates each sampling step as a constrained optimization problem on a spherical manifold induced by the latent Gaussian distribution, and solves it efficiently via Riemannian Gradient Descent (RGD). DiffRGD is a plug‑and‑play method that can be seamlessly integrated into any pre‑trained diffusion model. Extensive experiments demonstrate that DiffRGD outperforms previous methods in most image restoration and conditional generation tasks. Our project page is available at https://diffrgd.github.io/.
Authors:Shanwen Wang, Xin Sun, Sirui Wang, Xiao Xiang Zhu
Abstract:
Open‑vocabulary semantic segmentation (OVSS) enables text‑guided segmentation of unseen objects, breaking fixed‑class limitations to achieve open‑world understanding. However, existing OVSS methods primarily focus on modifying the CLIP attention mechanism, which still suffers from unstable local segmentation for remote sensing (RS) domain. To address these limitations, we propose RSGPNet, a training‑free geometric prompting framework for RS OVSS that refines segmentation by leveraging object geometric areas and consistency constraints. Specifically, RSGPNet comprises three core modules: a Text‑guided Coarse Mask module (TCM), a Geometric Re‑prompting Module (GRP), and a Coarse‑to‑fine Consistency Verification Mechanism (CVM). TCM utilizes text prompts and the input image to construct initial coarse segmentation masks. GRP then converts these coarse masks into geometric box prompts, feeding them back into the segmentation model to generate refined masks. Finally, CVM employs consistency computation to prevent prompting from reinforcing erroneous regions. They allow the model to improve segmentation accuracy in complex areas, such as category boundaries. Extensive experiments on RS datasets demonstrate that RSGPNet significantly outperforms state‑of‑the‑art methods across both quantitative and qualitative metrics while exhibiting excellent interpretability. The code is released at \hrefhttps://github.com/wangshanwen001/RSGPNethttps://github.com/wangshanwen001/RSGPNet.
Authors:Jianlong Xiong, ChuanBo Xie, Le Yu, Quansong He, Tao He
Abstract:
Recent advances in network architecture design have introduced layer attention to enhance inter‑layer interactions. In such frameworks, each layer queries all preceding layers to establish cross‑layer connections. However, layer attention results in quadratic computational complexity with respect to network depth. To mitigate this issue, prior works have proposed Recurrent Layer Attention (RLA) and linear attention mechanisms, which suffer from static information updates and limited long‑range cross‑layer dependency modeling. To overcome these limitations, we propose Key‑Correlated Layer Attention (KCLA), inspired by our observation that Key representations in layer attention exhibit high cosine similarity. KCLA achieves linear computational complexity while preserving dynamic information updates, directly derived from the foundational definition of layer attention. Furthermore, KCLA maintains long‑range cross‑layer connections and features a fixed spatial complexity, independent of network depth. Empirical evaluations demonstrate that KCLA delivers good performance across diverse tasks, including image recognition, object detection, and medical image segmentation. The code is publicly available at https://github.com/bgx666/KCLA.
Authors:Yunhun Nam, Jongheon Jeong
Abstract:
Vision‑Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that is not grounded in the image. Preference alignment is a promising approach to improve visual faithfulness, but its success depends heavily on how preference pairs are constructed. Existing methods exhibit two key limitations; (a) intervention‑based methods often introduce significant deviation from the policy distribution, and (b) sampling‑based methods often underuse visual information during the construction. In this paper, we propose ViPSy (Vision‑driven Preference Synthesis), a framework for constructing preference data that are both policy‑aligned and visually grounded. Our framework consists of two stages; in the first stage, ViPSy derives a visual cue from recurring object‑level content across semantically aligned image variants, so preference construction can rely on visual information rather than language priors. In the second stage, ViPSy conditions the policy's own rollouts on this cue, allowing candidates to be guided by visually grounded content while staying close to the policy's response distribution. The resulting candidates remain close to the policy's response distribution while better leveraging visual information from the image. Experiments show that the resulting VLM, preference‑aligned with ViPSy‑constructed preference pairs, achieves a new state‑of‑the‑art in hallucination mitigation. Compared with the previous state‑of‑the‑art method, it reduces hallucination rates on AMBER and Object HalBench by 35.7% and 24.5%, respectively. The resulting model further improves on general visual grounding benchmarks, e.g., MMStar, MMVP, and CV‑Bench, while also yielding gains in semantic segmentation and ImageNet linear probing, underscoring the effectiveness of our framework in enhancing the model's visual capabilities.
Authors:Jiasheng Wang, Tanun Jitwatcharakomol, Piyawadee Jongpradubgiat, Simeng Zhu
Abstract:
Accurate lesion segmentation in PET/CT is critical for oncology, yet remains challenging because physiologic tracer uptake and artifacts can mimic malignant signal. We present RADIANT‑PET, a reasoning‑augmented framework that couples a high‑sensitivity voxel‑level segmentation model with lesion‑level large language model (LLM) adjudication. Candidate uptake regions are generated with a deliberately permissive segmentation stage, then converted into structured textual descriptions that summarize uptake intensity, morphology, and regional and global anatomical context. An LLM classifies each candidate as true lesion vs. false positive, optionally leveraging the radiology report as additional clinical context. To strengthen lesion‑level reasoning, we further optimize a local LLM via reinforcement learning using Group Relative Policy Optimization, rewarding correct lesion classification and anatomically concordant site assignment. Across AutoPET and an OSU test cohort, RADIANT‑PET consistently outperforms strong image‑only baselines, with the largest improvements observed when radiology reports are provided. Overall, these results demonstrate that LLM‑based lesion‑level reasoning adds a novel reasoning layer beyond conventional segmentation, suppressing physiologic false positives and aligning voxel‑level predictions with clinical interpretation. The project repository is available at: https://github.com/jwang‑580/RADIANT‑PET.
Authors:Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min
Abstract:
Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting. We introduce PixelRAG, a new retrieval‑augmented method that represents websites in their native visual form and performs retrieval and reading entirely in pixel space, enabling an end‑to‑end architecture that eliminates text abstraction. PixelRAG is, to our knowledge, the first pipeline to operate over a full Wikipedia corpus in this form, scaling to a datastore of 30 million screenshot images with an efficient visual retrieval index. Built on an existing visual embedding model (i.e., Qwen3‑VL‑Embedding), PixelRAG further fine‑tunes this model on screenshot data with carefully curated contrastive training data. Retrieved screenshots are then fed directly as pixel inputs to a VLM, without intermediate text conversion. PixelRAG consistently outperforms both no‑retrieval and text‑based RAG baselines, most surprisingly on widely studied text‑centric tasks such as NQ and SimpleQA. It also achieves strong gains on multimodal open‑domain QA (e.g., MMSearch), benchmarks over noisy news corpora (e.g., LiveVQA), and agentic benchmarks (e.g., MoNaCo), improving accuracy by up to 18.1% over text‑based baselines. Finally, pixel representations enable a new efficiency lever for RAG through image compression, achieving up to 3x token cost reduction at lower resolutions while maintaining accuracy. Our results challenge the necessity of text representations in web retrieval, suggesting that web RAG can operate directly in the web's native visual form while improving both performance and efficiency.
Authors:Dihong Huang, Zhenyu Wei, Zhuxiu Xu, Yunchao Yao, Sikai Li, Mingyu Ding
Abstract:
Dexterous manipulation policies can solve individual skills, but composing them to perform multiple tasks with a single hand remains challenging. Adding a new task on top of an existing manipulation skill often imposes conflicting demands on overlapping fingers and contact modes, causing destructive interference between preserving an existing manipulation outcome and executing a new one. We propose DexCompose, a role‑aware residual composition framework that reuses pretrained dexterous policies for multi‑task manipulation through explicit finger‑level action ownership. Given two pretrained full‑hand policies, DexCompose first collects successful post‑task states from the first skill and performs release tests over candidate finger masks to identify which fingers are necessary for maintaining the established skill state. It then trains two asymmetric residual modules: a bounded residual stabilizer for task preservation, and a context‑aware residual that adapts the frozen downstream policy only within the action subspace assigned to the new task. We evaluate the framework on 16 composite dexterous manipulation tasks spanning four object‑retention skills and four downstream interactions. DexCompose achieves a 77.4% average composite success rate, demonstrating that structural action ownership with dual residuals offers a promising direction for composing dexterous skills beyond conventional policy chaining.
Authors:Yana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao, Kangheng Lin, En Yu, Keyu Lv, Han Zhou, Yin Tang, Haodong Li, Mitt Huang, Hangyu Guo, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Vishal M. Patel
Abstract:
We introduce PerceptionRubrics, a rubric‑based evaluation framework that addresses the gap between saturated benchmark scores and real‑world brittleness. Shifting evaluation from holistic semantic matching to rigorous atomic auditing, PerceptionRubrics pairs 1,038 information‑dense images with over 10,000 instance‑specific rubrics. These criteria are derived from golden captions constructed via a novel Circular Peer‑Review consensus pipeline and then distilled into a dual‑stream system of Must‑Right (essential facts) and Easy‑Wrong (fine‑grained details) rubrics. Crucially, PerceptionRubrics implements a Gated Scoring mechanism: unlike linear averages, failure on mandatory visual facts triggers sharp binary penalties. Extensive evaluation yields critical insights: (1) The Reliability Gap: models often verify fragmented elements correctly yet fail strict conjunctive constraints, exposing brittleness in dense domains; (2) Open‑Closed Stratification: contrary to reasoning trends, we reveal a persistent 8% perception deficit between open‑source and proprietary frontiers; and (3) Human‑Aligned Rigor: our gated metrics substantially out‑align conventional benchmarks, validating that strict perceptual fidelity is the prerequisite for reliable generation.
Authors:Jia-Chen Zhao, Beiqi Chen, Xinyang Chen, Guangcong Wang, Liqiang Nie
Abstract:
We present StructSplat, a feed‑forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requiring camera parameters. Existing methods either rely on per‑scene optimization or assume known camera poses, and often entangle geometry and appearance within a unified backbone, limiting reconstruction fidelity and generalization. Our key idea is to adopt a structured representation that organizes geometry, semantic, and texture cues with explicit roles in the reconstruction process. Specifically, we introduce a pixel‑aligned feature injection mechanism to enable accurate texture modeling from 2D observations, incorporate semantic‑aware priors to improve global consistency, and design a camera alignment strategy to prevent information leakage and improve generalization. Experiments show that our method significantly outperforms prior approaches on challenging benchmarks. On DL3DV, our method achieves 28.045 PSNR, surpassing AnySplat (22.377) by +5.67 dB. In cross‑dataset evaluation, our method achieves +1.94 dB over AnySplat on ACID and +1.72 dB on RealEstate10K. Project page: https://structsplat.github.io Code: https://github.com/J‑C‑Zhao/StructSplat
Authors:Yelin Wang, Zijia Song, Shuo Ye, Chuanguang Yang, Miaoyu Wang, Yong Xu, Zhulin An, Yongjun Xu, Zitong Yu
Abstract:
Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi‑temporal remote sensing images and holds significant research and application value. However, most existing methods rely on conventional deep learning architectures, and the limited model capacity constrains performance. Although large‑model post‑training techniques have achieved great success in general domains, their direct transfer to RSICC remains challenging due to data scarcity and the need for fine‑grained change understanding. To address this, we propose RSICCLLM, the first post‑training framework for large vision‑language models in RSICC. Specifically, we design a data generation paradigm, release the instruction dataset RSICI, and establish a task‑specific RSICC benchmark. We further introduce Difference‑aware Supervised Fine‑tuning to explicitly extract change representations and guide the model in perceiving and understanding temporal differences. In addition, we propose Dual‑Negative Preference Optimization (DNPO), which employs two complementary negative‑sample construction strategies to construct the preference dataset RSICP and further refine model performance. Extensive experiments validate the superior capability of RSICCLLM, which achieves outstanding results with only 7B parameters, surpassing models of substantially larger scales. The code and dataset will be made publicly available at https://github.com/keaill/RSICCLLM.
Authors:Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
Abstract:
Extracting dynamic 4D object interactions from massive, in‑the‑wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs. However, existing monocular 4D reconstruction methods primarily focus on isolated objects, often failing under the severe occlusions and complex dynamics inherent in multi‑object interactions. To bridge this gap, we propose HAT‑4D, the first agentic framework designed to reconstruct the 3D geometry, temporal dynamics, and physical interactions of multiple objects from a single video. By integrating VLMs with a multi‑level human‑in‑the‑loop feedback mechanism, HAT‑4D efficiently resolves depth ambiguities and interaction‑induced occlusions during 3D generation and 4D propagation, yielding physically plausible assets without relying on expensive multicamera rigs. As a scalable data engine, HAT‑4D facilitates the creation of MVOIK‑4D, an open‑world benchmark for monocular 4D interaction reconstruction, accompanied by a novel multi‑dimensional evaluation protocol focused on physical plausibility and temporal consistency. Extensive experiments demonstrate that HAT‑4D achieves SOTA performance on most evaluation metrics, while maintaining competitive semantic alignment. Ablation studies show that introducing a small amount of human feedback improves interaction reconstruction. Moreover, the data produced by HAT‑4D effectively improves baseline performance when used for fine‑tuning. Our data and code are available at https://lijiaxin0111.github.io/HAT4D/
Authors:Hong Li, Minqi Meng, Yanjun Liang, Chongjie Ye, Houyuan Chen, Weiqing Xiao, Xianda Guo, Guojun Lei, Xuhui Liu, Chaojie Yang, Yanlun Peng, Hao Zhao, Baochang Zhang
Abstract:
Reconstructing high‑fidelity, relightable 3D avatars from a single in‑the‑wild image is a challenging ill‑posed problem, primarily hindered by the scarcity of high‑quality PBR data and the complexity of disentangling illumination from intrinsic materials. In this paper, we present a data‑efficient framework that leverages the robust priors of a unified pre‑trained diffusion backbone to sequentially address texture completion, delighting, and material decomposition. Unlike existing methods that rely on fragmented pipelines or extensive proprietary datasets, we utilize cascaded Low‑Rank Adaptations (LoRAs) to adapt the strong generative prior of the diffusion model for each sub‑task in UV space. Specifically, we first employ an Inpainting LoRA to complete missing UV textures caused by occlusion, leveraging the model's semantic understanding to generate semantically and photometrically coherent details. Subsequently, a Light‑Homogenization LoRA and a novel Cross‑Intrinsic Attention mechanism are introduced to remove baked‑in lighting and collaboratively synthesize pixel‑aligned PBR maps (Albedo, Normal, Roughness, Specular, and Displacement). To ensure physical plausibility, we impose a UV‑space differentiable BRDF shading loss during the decomposition stage, forcing the generative process to adhere to the rendering equation without the artifacts typical of rasterization‑based supervision. Extensive experiments demonstrate that our method, trained on fewer than 100 real 3D scans, generates comprehensive, 4K‑resolution PBR assets with superior realism and generalization compared to state‑of‑the‑art methods, and all training code and model weights will be released upon acceptance.
Authors:Peiwen Zhang, Yufan Deng, Shangkun Sun, Juncheng Ma, Duomin Wang, Jonas Du, Zilin Pan, Ye Huang, Hao Liang, Songyan Huang, Ruihua Zhang, Enze Xie, Ming-Yu Liu, Daquan Zhou
Abstract:
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general‑domain video generators and robot‑specific data fine‑tuned models can still produce physically implausible manipulations, including discontinuous motion trajectories and inconsistent robot‑object interactions, which limits their reliability as world simulators. Through extensive experiments, we find that such physical instability mainly arises from two factors: deformation of moving objects and implausible spatio‑temporal correlations among interacting entities, particularly during contact. Building on this observation, we propose PhysisForcing, a scalable training framework that strengthens physical consistency by focusing supervision on physics‑informative regions through joint optimization of pixel‑level and semantic‑level features. The framework consists of a pixel‑level trajectory alignment loss, which supervises DiT features using reference point trajectories, and a semantic‑level relational alignment loss, which aligns DiT features with inter‑region relations extracted from a frozen video understanding encoder. Extensive experiments on R‑Bench, PAI‑Bench, and EZS‑Bench show that PhysisForcing consistently improves embodied video generation over strong baselines, improving the Wan2.2‑I2V‑A14B and Cosmos3‑Nano base models on R‑Bench by 22.3% and 9.2% (7.1% and 3.7% over vanilla finetuning), with the Cosmos3‑Nano variant attaining the best overall score. Beyond generation, as a world model under the WorldArena action‑planner protocol it raises the closed‑loop success rate from 16.0% to 24.0% and further improves downstream policy success, indicating that physically aligned video models yield stronger representations for robotic manipulation.
Authors:Alex Colagrande, Paul Caillon, Eva Feillet, Alexandre Allauzen
Abstract:
Neural operators provide deep neural networks for learning mappings between function spaces. Among them, the Fourier Neural Operator (FNO) is particularly effective: its spectral convolution relies on low‑dimensional Fourier‑domain representations and can handle inputs at different resolutions. This design aligns well with settings where the Fourier basis diagonalizes the underlying operator, such as linear, constant‑coefficient PDEs on periodic domains, in which Fourier modes evolve independently. However, nonlinear PDEs may benefit from an additional inductive bias, as they exhibit structured interactions between modes, governed by polynomial nonlinearities. To capture this inductive bias, we introduce the Higher‑Order Spectral Convolution, a spectral mixer that extends FNO from diagonal modulation to explicit n‑linear mode mixing, aligned with the dynamics of nonlinear PDEs. Our experiments on standard benchmarks show that the proposed Higher‑Order FNO (HO‑FNO) retains the efficiency of FNO‑based architectures and consistently improves over other spectral neural operators. HO‑FNO also performs on par with or better than state‑of‑the‑art transformers and state‑space models on several datasets, with stronger gains in highly nonlinear regimes, such as the Poisson equation with polynomial forcing, where a single HO‑FNO layer outperforms FNO models with up to 16 layers. We open‑source our code for reproducibility at: https://github.com/AlexColagrande/HO‑FNO.
Authors:Qinming Zhou, Chenxi Sun, Deyang Kong, Junhao He, Xiangheng Tang, Peike Yu, Haotian Wu, Leilei Cao, Linfeng Zhang
Abstract:
Real‑world object removal is challenging due to two key difficulties: the target object's non‑local effects, such as shadows and reflections, which are difficult to model, and the fact that user‑provided masks are often inaccurate or incomplete. With billions of parameters and tens of denoising steps, diffusion‑based models achieve strong removal performance at the expense of substantial computational cost, limiting their use in interactive applications and on edge devices. To address these challenges, we present OSOR (One‑Step Object Removal), which simultaneously achieves efficient, effect‑aware, and mask‑robust object removal. Concretely, OSOR introduces: (1) an occupancy‑guided discriminator for precise boundary supervision, enabling stable single‑step diffusion training; (2) an alpha head that leverages knowledge from pretrained diffusion models to predict appropriate removal regions with minimal overhead, thereby handling imperfect masks; and (3) a semantic‑anchored verification pipeline (SAVP) that filters noisy instruction‑based triplets to produce effect‑aware supervision at scale. Using SAVP, we curate CORNE, which contains 280K verified removal pairs, and further annotate AnimeEraseBench and TextEraseBench to evaluate performance on more complex removal tasks. Experiments show that OSOR surpasses strong multi‑step diffusion baselines in perceptual quality while achieving 4× to 30× faster inference.
Authors:Pragati Shuddhodhan Meshram, Varun Chandrasekaran
Abstract:
Attributing a generated image to its source diffusion model is a fundamental challenge in provenance verification and intellectual property protection. This problem is particularly difficult because diffusion models trained on different datasets can converge to similar score functions and thus similar output distributions, making the generated images themselves unreliable as attribution evidence. Existing non‑invasive methods either fail on architecturally similar variants or rely on signals that vanish when models share the same autoencoder. We propose Spectral Denoising Signatures (SDS), a non‑invasive attribution method that identifies the source model by fingerprinting each candidate model's denoising behavior. Our key insight is that a model's denoising score function exhibits a distinctive spectral geometry, reflected in how it redistributes energy across spatial frequency bands during denoising. By probing this behavior with frequency‑controlled perturbations, SDS extracts a stable signature that is intrinsic to the model, requiring only standard forward passes with no inversion, optimization, or generation‑time enrollment. Our results demonstrate that SDS achieves approximately 99.9% accuracy across eight diverse diffusion models and 96.2% under cross‑domain prompt shift, outperforming non‑invasive baselines across variations in training data, architecture, and training procedure, establishing spectral geometry as a principled and practical basis for diffusion model attribution. Code is available at: https://github.com/Pragati‑Meshram/SGS
Authors:Jiyao Wang, Qingyong Hu, Duoxun Tang, Xiao Yang, Kaishun Wu, Jiangbo Yu
Abstract:
Video‑based remote physiological measurement (RPM) is highly accessible but remains fragile under varying illumination, skin tones, and motion. Radio frequency (RF) radar is largely invariant to illumination and appearance, providing complementary cardio‑respiratory micro‑motion cues; however, requiring radar at inference is often impractical due to its limited ubiquity and deployment overhead. We propose RPM‑Distill, a physiology‑guided cross‑modal distillation framework that leverages synchronized radar only during training while retaining video‑only inference. Our key observation is that although RGB and RF waveforms differ in sensing physics and time‑domain morphology, they share similar latent periodic rhythm in the frequency domain. We thus distill physiology‑structured spectral evidence to improve robustness, via losses that (i) anchor the fundamental peak, (ii) match the off‑peak background distribution, and (iii) preserve spectral morphology and sharpness. To avoid negative transfer under sample‑level teacher quality and alignment uncertainty, a spectral policy network predicts sample‑level distillation gates and component weights from the student‑‑teacher spectral relation map, learned with a meta bilevel objective on a small labeled validation split. Through extensive experiments in challenging conditions and cross‑dataset settings, RPM‑Distill brings 81% MAE and 21% correlation improvement over unimodal baselines. Code is at https://github.com/WJULYW/RPM‑Distill.
Authors:Boyuan Chen, Zichen Dang, Chuang Yang, Lap-Pui Chau, Yi Wang
Abstract:
In real‑world deployments, scene text detectors inevitably face distribution shifts beyond the training distribution. Prior work often depends on large‑scale scene‑text pretraining, yet evaluation under cross‑domain changes and real‑world imaging degradations remains limited. We propose TextDS, an efficient framework for scene text detection under distribution shifts. First, we propose a data‑efficient dual‑encoder design with visual foundation models, eliminating the reliance on large‑scale scene‑text pretraining. Second, we introduce Step‑wise LoRA adaptation (SWLoRA), which performs progressive low‑rank refinement with a dynamic early‑exit mechanism for effective feature adaptation. Third, we propose Common Subspace Fusion (CSF) to align and fuse the two branches in a shared subspace while retaining complementary, shift‑robust information. Finally, we construct adverse‑condition scene text detection datasets to address the gap in evaluating under imaging degradation. Experiments show that TextDS achieves competitive performance in scene text detection, demonstrating robustness across domains and adverse imaging conditions with only 4.9M trainable parameters.
Authors:Dongbin Zhang, Hao Liu, Binquan Dai, Kangjie Chen, Chuming Wang, Chen Li, Jing Lyu, Haoqian Wang
Abstract:
High‑fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods face a dilemma between expressiveness and disentanglement. Mainstream 2D pose‑conditioned approaches suffer from "motion‑shape entanglement", leading to the leakage of the driving subject's body shape. Conversely, methods relying on 3D priors (e.g., SMPL) achieve geometric disentanglement but struggle to capture facial expressions and complex gestures, resulting in rigid animations. To this end, we propose EMOSH, a novel framework for high‑fidelity controllable human video generation. First, an Expressive Human Model (EHM) is introduced as the core control representation. By explicitly disentangling shape and pose parameters, we fundamentally resolve the body shape leakage issue. Alongside this, a robust motion tracker is designed to accurately estimate EHM parameters from video. Second, we propose a Coarse‑to‑Fine Hybrid Motion Injection strategy, enabling more fine‑grained control over expressions and gestures. Furthermore, we introduce a Spatially‑Aligned Conditioning mechanism to bridge the domain gap between training and inference, improving identity consistency. Extensive experiments demonstrate that EMOSH outperforms previous methods in both self‑driven and cross‑driven scenarios, producing high‑fidelity videos with vivid expressions while maintaining shape disentanglement.
Authors:Xirui Teng, Nan Xi, Junsong Yuan
Abstract:
Analyzing fine‑grained skill activities (e.g., sports, surgery) requires not only recognizing visual patterns but also performing step‑by‑step visual reasoning that leads to the final judgment. While recent advances in action quality assessment have achieved remarkable progress in evaluating performance, existing models remain black boxes, where they lack the ability to explicitly reveal the reasoning processes underlying their judgments. To address this limitation, we propose Latent Visual Diffusion Reasoning (LVDR), a novel framework that integrates keypoint‑guided Monte Carlo Tree Search (MCTS) to model and visualize the latent visual reasoning process. LVDR not only produces more accurate skill assessments but also uncovers the critical visual reasoning sequences that contribute to the final evaluation. Extensive experiments across four datasets spanning diverse sports and surgical domains demonstrate that LVDR achieves competitive quantitative performance while providing interpretable visual reasoning trajectories leading to the final predictions. Source codes and models can be found through the following link: https://github.com/XiruiTeng/LVDR_Official.git.
Authors:ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, Haoqian Wang
Abstract:
Knowledge‑based Visual Question Answering (KB‑VQA) requires models to combine image understanding with external knowledge. Most prior methods use a fixed retrieve‑then‑generate pipeline with a pre‑selected retriever and a static top‑k setting, which is not adaptive during reasoning. We propose ProMSA, a progressive multimodal search agent for KB‑VQA. Given an image‑question pair, the agent iteratively chooses image search, text search, or stop, under explicit tool‑call budgets and with deduplication to avoid redundant retrieval. For training, we first use rejection‑sampling SFT to learn valid tool‑use formats, then optimize the agent with TN‑GSPO, a sequence‑level RL objective that normalizes updates by both generation length and tool‑interaction depth. Experiments on E‑VQA and InfoSeek show consistent gains over strong RAG and agent baselines, and improved retrieval and end‑to‑end accuracy. The code is available at https://github.com/DingWu1021/Promsa.
Authors:Haoyuan Wang, Yabo Chen, Haibin Huang, Chi Zhang, Xuelong Li
Abstract:
Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accumulation and temporal degradation during extended rollouts. This issue is further amplified under heterogeneous controls such as human motion and camera trajectories, which may interfere and destabilize a pretrained video prior, while existing methods often trade off controllability and visual quality. We propose "Directing the World", a fast autoregressive framework for controllable world‑model video generation with compositional human‑motion and camera‑trajectory control. Our key idea is to decouple control learning while preserving a unified autoregressive video prior. We introduce a Fast‑Slow Memory training strategy to stabilize long‑horizon rollout learning and improve convergence. For human motion control, we design a t‑guided Dynamic Projection mechanism and a refined Motion‑CFG strategy, enabling temporally smooth and accurate motion alignment without degrading visual fidelity, and supporting multi‑person control.After learning a robust motion prior, we introduce a second‑stage camera‑trajectory control module to compose human dynamics with viewpoint changes for coherent world exploration. We further construct a large‑scale dataset with synchronized video, text, human‑motion, and camera‑trajectory annotations, organized into motion‑centric and camera‑centric subsets for decoupled training. Extensive experiments show stable long‑horizon generation with precise controllability and high visual quality. See more at https://whydahuzi.github.io/Directing‑the‑World.github.io/.
Authors:Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi, Eva Cetinic, Gennaro Vessio, Giovanna Castellano
Abstract:
Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identifies a subject, or recognizes an iconographic symbol, does it ground each claim in the relevant region of the canvas, draw on an undifferentiated visual signal, or rely primarily on textual priors? We study this using the Token Activation Map (TAM), which produces, for each generated token, a heatmap isolating the visual evidence specific to that token from prior‑context interference. Applying TAM to a curated set of paintings spanning multiple periods and genres, we analyze grounding patterns across five semantically distinct token categories: common visual objects, style descriptors, metadata, iconographic tokens, and affective expressions. We find that visual grounding varies substantially with token semantics. We further show that MLLMs attempt to identify artworks and artists, achieving higher accuracy in artist attribution than in title prediction, where hallucinations are more frequent. Finally, we compare TAM with SAM~3 open‑vocabulary segmentation. To ensure reproducibility, we release our code, experimental configurations, prompts, and qualitative results on the project page at https://nicolafan.github.io/tamart/.
Authors:Yuheng Qiu, Jingyi Luo, Chenfei Ye, Ting Ma, Jianfeng Cao
Abstract:
Deep learning has demonstrated remarkable success in high‑throughput histopathology image analysis. However, the performance of learning‑based models critically depends on the quality and size of annotations by expert pathologists, which is a resource‑intensive and time‑consuming process. To address the limitations of data scarcity and annotation burden, several methods have been proposed to synthesize paired histopathology data. Nevertheless, these frameworks typically still require annotation data, albeit in reduced quantities, to impose structural constraints during training. In this work, we present CHIS, a plug‑in framework that guides the sampling trajectory of a pretrained diffusion model through two key stages: structural initialization at the start and textural modulation during generation. The initial noise state is refined by fusing the phase information from a prior mask with the amplitude of Gaussian noise in the frequency domain, yielding a structurally informed starting point. During the reverse diffusion process, we adaptively modulate both coarse‑grained and fine‑grained textures at different wavelet decomposition levels. This enables a diffusion model pretrained solely on unlabeled images to generate outputs that align with prior structural masks while preserving the reference tissue style. We conducted extensive experiments demonstrating the superiority of CHIS in generation fidelity and its substantial benefits for downstream segmentation tasks. Code is available at https://github.com/IBIL‑Code/CHIS.
Authors:Qiaoyue Yang, Sven Heutger, Christopher Niemann, Magnus Jung, Ayoub Al-Hamadi, Sven Wachsmuth
Abstract:
Human motion describes the three‑dimensional full‑body movement of a person. Anticipating such motion holds significant relevance across a wide range of application domains such as human‑robot interaction, autonomous driving, animation, and healthcare. In recent research, spatial and temporal dependencies are modeled by bidirectional attention mechanisms. These typically anticipate human motion in an autoregressive manner which could cause an accumulation of errors over time. As a consequence, they solely focus on local pose forecasting. To address these limitations, we propose a non‑autoregressive transformer based on spatio‑temporal attention, and train it not only for local pose anticipation, but also for global motion prediction in space. Furthermore, to enhance its applicability in real‑world scenarios, our model is also trained to recover missing joints due to occlusions, and is capable of processing varying lengths of history observations. Our code is publicly available at https://github.com/Q‑Y‑Yang/Prediction‑of‑Local‑and‑Global‑Human‑Motion.
Authors:Zhaotong Yang, Ying Tai, Jiahui Zhan, Yu Zheng, Jianjun Qian, Jian Yang
Abstract:
Unified fashion generation integrates tasks like virtual try‑on and garment reconstruction into a single model to reduce task‑specific adaptation costs. However, naive parameter sharing across semantically distinct tasks induces negative transfer through severe inter‑task gradient conflict. We propose OrthoTryOn, a unified framework mitigating this interference within a shared Low‑Rank Adaptation (LoRA) module. Its Orthogonal Subspace Projection (OSP) applies task‑specific orthogonal rotations to bottleneck features, mapping them into decorrelated coordinate frames. To address residual semantic coupling at inference time, we further propose Fisher‑guided Negative Guidance (FNG), a parameter‑free strategy that utilizes diagonal Fisher information to quantify inter‑task sensitivity overlap and explicitly repels generation trajectories from the most confusable task via Classifier‑Free Guidance. Extensive experiments demonstrate that OrthoTryOn avoids the severe performance degradation typical of naive unified training and even surpasses independently trained task‑specific models, achieving state‑of‑the‑art results across multiple benchmarks while generalizing robustly across diverse diffusion backbones. Code is available at https://github.com/NJU‑PCALab/OrthoTryOn.
Authors:Haoyu Zhang, Meng Liu, Qianlong Xiang, Kun Wang, Yaowei Wang, Liqiang Nie
Abstract:
Spatial intelligence is essential for low‑altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image‑level recognition, single‑view understanding, or narrow answer formats, leaving 3D spatial inference, multi‑view collaboration, scene dynamics, and diverse task formulations insufficiently evaluated. To address these gaps, we introduce SpatialUAV, a real low‑altitude UAV benchmark comprising 4,331 curated instances across 14 fine‑grained task types, covering semantic discrimination, spatial relation, aerial‑‑aerial collaboration, aerial‑‑ground collaboration, and motion understanding. SpatialUAV organizes all samples into a unified visual‑input‑‑question‑‑answer schema, while supporting seven input configurations and nine answer formats, including option labels, region identifiers, geometric values, cross‑view correspondences, and free‑form motion descriptions. To ensure reliable and grounded evaluation, our data construction pipeline integrates detector‑assisted regions, depth supervision, metadata‑derived rules, extensive manual annotation, blind filtering, and multi‑turn human validation, together with task‑specific metrics for heterogeneous outputs. Evaluating representative vision‑language models across three categories, we show that current models remain far from human‑level performance, with pronounced bottlenecks in cross‑view association, structured grounding, geometric reasoning, and temporal viewpoint understanding. These results offer empirical guidance for advancing low‑altitude UAV spatial intelligence. Code and data are available at https://github.com/Hyu‑Zhang/SpatialUAV.
Authors:Zhaoning Shi, Bo Ma, Hao Xu, Zepeng Yang, Bo Liang
Abstract:
This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus‑DETR, a novel detection framework based on biological hippocampal memory modeling. This framework integrates a hippocampal memory network module, HipNet, into the DETR architecture and systematically simulates the anatomical structure and functional organization of hippocampal subregions, including the entorhinal cortex, dentate gyrus, CA3, CA1, and subiculum. Through this design, Hippocampus‑DETR realizes pattern separation, pattern completion, importance filtering, and information integration of visual encoding features. During training, different memory submodules are optimized using a layer‑wise training strategy, ultimately forming a memory system with memory retrieval and completion capabilities. Experimental results demonstrate that Hippocampus‑DETR achieves higher detection accuracy than current mainstream models. More importantly, models equipped with this framework also exhibit excellent generalization ability and data efficiency in tasks such as few‑shot image classification, multimodal feature construction, and image restoration. Subsequent experiments further validate the functional necessity and internal interpretability of each memory submodule. This study not only provides a novel object detection framework, but also offers a feasible technical pathway for integrating neurocognitive mechanisms with deep learning models, highlighting its significant value in improving model learning efficiency and task robustness. The project is available at https://github.com/2186cloud/hipnet.
Authors:Mingcheng Wang, Junbo Qiao, Yunchen Li, Lingfu Jiang, Wei Li, Jie Hu, Jiao Xie, Zhou Yu, Xinghao Chen, Guixu Zhang, Shaohui Lin
Abstract:
Speculative decoding (SD) has emerged as a key solution to accelerate the inference of autoregressive models. However, in the field of image generation, it faces the challenge of low acceptance rates, and directly relaxing its criteria leads to degradation in image quality. In this paper, we propose a novel content‑aware speculative decoding algorithm, termed CSD, which integrates an entropy‑based probability relaxation mechanism with an optimal resampling strategy to enhance the inference efficiency for autoregressive image generation. By leveraging the informational uncertainty inherent in different regions of an image, CSD dynamically adjusts the acceptance probability of candidate tokens, increasing the acceptance rate in low‑detail areas to accelerate generation. Moreover, a distribution alignment filter is introduced to ensure the output distribution to be aligned with the target model, which significantly improves the generative quality. Experiments conducted on Lumina‑mGPT and Janus‑Pro demonstrate that the superiority of the proposed CSD. Our source code is available at https://github.com/aderfebr/CSD.
Authors:Jian Shi, Cheng Zhen, Pingping Zhang, Rui Xu, Yanan Lv, Yili Ma, Huan Bi, Haojie Li, Huchuan Lu
Abstract:
Language‑guided Medical Image Segmentation (LMIS) has shown great potential to improve the delineation of anatomical structures and lesions by integrating clinical textual information. Existing methods generally rely on either implicit interaction between textual and visual features or auxiliary coarse‑grained supervision for cross‑modal alignment. However, these methods lack explicit and fine‑grained constraints to ensure semantic consistency, causing a mismatch between language and the segmentation outputs. To address this issue, we propose Text‑as‑Illumination Retinex Network (TIRNet), a novel Retinex‑inspired framework that treats text embeddings as semantic illumination for feature modulation, thereby improving semantic consistency in LMIS. TIRNet introduces two key blocks integrated at each decoder stage: (1) the Retinex‑inspired Text Modulation Block (RTMB), which employs positive and negative illumination maps to enhance text‑relevant foreground features and suppress background interference; and (2) the Consistent Detail Compensation Block (CDCB), which selectively recovers high‑frequency details via a consistency‑gated mechanism conditioned on illumination reliability. Furthermore, we propose a Multi‑Scale Illumination Supervision Loss (MSIS‑Loss), comprising a Region‑Grounded Contrastive Loss (RGC‑Loss) that enforces cross‑modal similarity to be concentrated in text‑relevant foreground regions and suppressed in background regions, and a Background Suppression Loss (BS‑Loss) that provides pixel‑level supervision for negative illumination maps, jointly ensuring a precise cross‑modal alignment at each decoder stage. Extensive experiments on the MosMedData+ and QaTa‑COV19 datasets demonstrate that TIRNet achieves state‑of‑the‑art performance in LMIS. The code is available at: https://github.com/anaanaa/TIRNet.
Authors:Taïga Gonçalves, Yongsong Huang, Tomo Miyazaki, Shinichiro Omachi
Abstract:
The existence of adversarial attacks is often attributed to the presence of non‑robust features in neural networks. While prior defenses reduce their impact via pruning, masking, or feature recalibration, we instead propose to jointly learn to amplify and attenuate these signals through a simple activation scaling mechanism. To this end, we introduce Activation Amplification and Attenuation (A3), a lightweight plug‑in module that enhances adversarial robustness with minimal modifications of the activations. A3 dynamically rescales the activations using a learnable mask and a scaling factor derived from the original activation magnitudes. The influence of adversarial perturbations can be amplified or attenuated using the same learnable parameters by simply flipping the sign of the scaling operation. The amplified signals serve as negative references to construct novel contrastive and ranking loss functions. Experimental analysis shows that learning to degrade the predictions in amplification mode simultaneously improves adversarial robustness in attenuation mode. Moreover, A3 relies on only a small number of learnable parameters, with most of its behavior being determined by the scaling mechanism rather than additional network capacity. Extensive experiments demonstrate that integrating A3 into different backbones, datasets, and training methods consistently improves adversarial robustness while introducing negligible computational and memory overhead compared to existing plug‑in modules. Code is available at: https://github.com/tgoncalv/A3.
Authors:Anuki Pasqual, Dulan Lokugeegana, Manimohan Thiriloganathan, Nuthya Rathnayake, Kithsiri Samarasinghe, Udaya S. K. P. Miriya Thanthrige
Abstract:
Vehicle license plate recognition is an integral component of intelligent transportation systems. In this work, we present an embedded real‑time license plate recognition system customized for developing countries. We address the challenge of handling complex, unstructured traffic scenes with diverse vehicle types while implementing the system on an embedded platform for low‑cost deployment. Our method consists of license plate detection on a multi‑vehicle image, followed by character recognition on the detected license plates. Both steps use lightweight convolutional neural networks to balance accuracy and efficiency. We also introduce the SL‑LPR dataset of Sri Lankan road images, which contains a variety of vehicle types and traffic conditions typically seen in developing countries. On this dataset, the license plate detection and character recognition models achieved 93.6% mAP and 87.88% accuracy, respectively, and were competitive against larger models on several public datasets. To achieve real‑time performance in a resource‑constrained embedded environment, we applied low‑bitwidth quantization using the Brevitas library and implemented FPGA acceleration for the models using the FINN framework. The end‑to‑end system can operate at 11.5~FPS when implemented on the Xilinx Kria KV260 platform. These results demonstrate that our system is effective for real‑time license plate recognition on an embedded device, even in complex traffic scenarios. The SL‑LPR dataset is available for research use at: https://github.com/sl‑lpr‑uom/SL‑LPR.git.
Authors:Qinfeng Zhu, Lei Fan
Abstract:
Panoramic images capture the complete visual sphere in a single frame, providing spatial context unattainable by conventional cameras. Yet this completeness comes at a geometric cost: the 2‑sphere cannot be faithfully mapped to the plane, and every planar representation introduces distortions that violate the assumptions underlying standard vision architectures. This survey traces the evolution of panoramic scene analysis along a methodological trajectory, from projection‑based adaptation, through distortion‑aware engineering, to sphere‑native modeling and geometry‑aware tokenization for foundation models, and argues that this evolution reflects a progressive deepening of geometric commitment rather than a simple accumulation of techniques. We organize the literature along two orthogonal dimensions: architectural design (how operators interact with spherical geometry) and training paradigm (how knowledge is transferred across domains). Covering dense prediction (semantic segmentation, depth estimation, and room layout estimation), unified multi‑task understanding, open‑world perception, vision‑language reasoning, and dynamic video analysis, we identify a central unresolved tension: among the methods surveyed, none simultaneously delivers strict spherical equivariance and full reuse of perspective‑pretrained foundation‑model weights, and we argue that this is a structural rather than incidental gap. We further expose five systematic gaps in current evaluation protocols, namely the absence of spherical‑area‑weighted metrics, seam‑consistency testing, polar‑robustness stratification, cross‑projection generalization, and open‑world protocol standardization, and propose a six‑point research roadmap toward general‑purpose panoramic intelligence. The corresponding repository is publicly available at: https://github.com/zhuqinfeng1999/Awesome‑Panoramic‑Scene‑Analysis.
Authors:Kai Wang, Zhaopeng Gu, Yixiang Chen, Yuan Xu, Qisen Ma, Peng Su, Zhaowen Li, Yan Huang, Liang Wang
Abstract:
World‑action models have shown promising robot‑manipulation performance by jointly predicting future visual states and actions. However, existing methods mainly rely on short‑term history and short‑horizon future prediction, which is insufficient for long‑horizon tasks whose correct execution depends on earlier observations and task progress. Such temporally dependent tasks require effective use of complementary temporal information, including recent local context, cross‑stage historical events, immediate future dynamics, and global task progress. To address long‑term forgetting and poor awareness of the global task state, we introduce DiM‑WAM, a memory‑augmented world‑action model that integrates multi‑scale historical context, local future dynamics, and global task progress. The memory extracts compact visual event information from real observations, updates multiple memory banks through independent similarity‑based merging, and then reads the bank‑identity‑ and time‑embedded long‑term context to condition video and action denoising. A progress‑supervision objective further encourages memory tokens to encode not only completed historical events but also the current task stage and its implications for the remaining task. On RMBench, DiM‑WAM raises average success from 28.4% with LingBot‑VA to 69.8%, exceeding the explicit‑memory Mem‑0 baseline at 42.0%. On four real‑world Franka tasks, it improves average stage success from 70.7% to 91.5% and full‑task success from 52.5% to 80.0%. Project page: https://wangkai‑casia.github.io/dim‑wam/\texttthttps://wangkai‑casia.github.io/dim‑wam/.
Authors:Liu Yu, Can Chen, Ping Kuang, Zhikun Feng, Fan Zhou, Gillian Dobbie
Abstract:
Large Vision‑Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination. Deviating from the prevailing attention intensity assumption, we reveal a deeper dynamic structural misalignment: hallucination is triggered at decision‑critical steps where specific attention heads, acting as risky mediators, decouple from visual evidence to lock onto language priors. This establishes a pathological shortcut that bypasses visual grounding. To dismantle this, we propose Fox (Faithfulness and Observational‑flow via eXpression‑rectification), a training‑free inference‑time framework. Fox diagnoses structural misalignment using a visual attention entropy probe to localize risky mediators unsupervisedly. We then execute a targeted causal intervention via numerical logit saturation to physically sever the shortcut path. Finally, a conflict‑gated cooperative decoding strategy reconciles interventional faithfulness with observational fluency. Extensive experiments demonstrate that Fox achieves SOTA performance, outperforming SID by 29.1% while preserving linguistic richness. Code is available at https://github.com/Cc2021start/Fox.
Authors:Tim Alexander Bader, Tim Dieter Eberhardt, Maximilian Dillitzer, Wilhelm Stork
Abstract:
Camera‑based perception systems for autonomous driving are typically developed and evaluated using fixed sensor rigs, while real‑world vehicle fleets exhibit substantial variation in camera placement, orientation, field of view, and camera count. This mismatch introduces a cross‑rig domain gap in which only the geometric observation process changes. To study this effect under controlled conditions, we introduce Plentiful CARLA Camera Rigs, a benchmark that renders identical driving scenes under 14 systematically designed camera rigs. This setup enables direct analysis of cross‑rig generalization without confounding changes in scene content or appearance. Using the benchmark, we analyze cross‑rig transfer behavior of representative multi‑view perception architectures and observe substantial performance shifts induced by geometric rig variation. To facilitate structured analysis, we further introduce two calibration‑based descriptors derived from rig metadata: Rig Variance, capturing internal rig diversity, and Rig Contrastive Distance, measuring geometric discrepancy between rigs. Our experiments show that geometric rig differences strongly correlate with relative cross‑rig performance shifts and that Rig Contrastive Distance provides a reliable proxy for ranking transfer difficulty between sensor rigs.
Authors:Thomas Shih-Chao Liang, Zhuoran Yu, Yong Jae Lee
Abstract:
Large Language Models (LLMs) possess broad conceptual knowledge acquired through large‑scale text pretraining, yet their potential to supervise models in other modalities remains underexplored. In this work, we propose LaViD‑‑Language‑to‑Visual Knowledge Distillation‑‑a simple and effective framework for transferring high‑level semantic knowledge from a language‑only teacher to a vision‑only student model. Instead of relying on paired multimodal data, LaViD elicits conceptual signals from an LLM by prompting it to generate multiple‑choice questions (MCQs) that probe semantic distinctions between visual classes. Each class is mapped to a soft label distribution over these MCQs, forming a rich conceptual signature that guides the student through an auxiliary distillation loss. Notably, despite using a language‑only teacher without access to image data, LaViD consistently outperforms recent methods like MaKD that distill from vision‑language models across multiple fine‑grained benchmarks. It also achieves competitive or superior performance compared to state‑of‑the‑art visual distillation methods such as DKD and MLKD, with further gains when combined with logit standardization. On the Waterbirds dataset, LaViD substantially improves worst‑group accuracy, demonstrating enhanced robustness to spurious correlations with distillation. Code is available at https://github.com/lliangthomas/lavid.
Authors:Daniel Cher, Hamza Iqbal, Eric Xing, Brian Wei, Nathan Jacobs
Abstract:
Geolocation encoders, which map geographic coordinates to learned representations, are emerging as an effective means of capturing visual and non‑visual characteristics from a latitude‑longitude pair alone. However, existing approaches project coordinates onto fixed bases (e.g., spherical harmonics), allocating representational capacity uniformly and devoting equal resources to the open ocean and to a developing city. We introduce Tessellating the Earth (TTE), a location encoder built from learnable Spherical Voronoi partitions that concentrates representational capacity where it is needed in a fully differentiable, end‑to‑end manner. Each Voronoi site carries its own embedding and migrates during training toward discriminative areas. To bridge the gap between local spatial structure and global semantic understanding, we introduce \emphglobal semantic tokens: a set of shared learnable concept tokens that distill semantic knowledge from the satellite imagery into a compact vocabulary the location encoder can reference at inference, enabling geographically distant sites covering similar environments to share semantics. TTE sets a new state of the art for location encoders across a suite of geospatial classification and regression tasks, and achieves the strongest results when used as a geographic prior for fine‑grained species classification on iNaturalist‑2018. Code, and weights are available at https://github.com/mvrl/TTE.
Authors:Yujin Tang, Chenming Shang, Ruize Xu, Nikhil Singh
Abstract:
Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write down. We introduce DMV‑Bench (Code: https://github.com/yyyujintang/DMV‑Bench), the first interactive benchmark for multimodal‑agent visual memory. DMV‑Bench is built on a controlled home‑furnishing e‑commerce catalogue of 1,000 product variants in which a text‑leakage contract keeps the discriminative signal of each task in the pixels alone. Across a chain of autonomous shopping sessions, every visited product image carries a unique, pre‑rendered incidental cue, and the agent is later asked to recall a particular cued product and navigate to its URL. Inspired by dual‑coding theory, we propose DualMem, a memory architecture that maintains a visual and a verbal code in parallel. On DMV‑Bench, DualMem outperforms a caption baseline and three recent multimodal agent‑memory systems at every chain length J in 5, 10, 15, 50 on both Gemini 2.5 Flash and Qwen2.5‑VL‑7B, with the lead surviving controls for memory‑bank size and encoding‑position bias, and an asymmetric dual‑coding regime in which vision carries the cue end‑to‑end while the verbal channel plays a smaller query‑grounding role.
Authors:Trung Thanh Nguyen, Daniel Lusk, Kilian Gerberding, Janusch Vajna-Jehle, Tuan-Anh Vu, Duc Viet Le, Tu Vo, Phi Le Nguyen, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide, Julian Frey, Teja Kattenborn
Abstract:
Automated instance segmentation of forest LiDAR point clouds is increasingly critical as forest monitoring moves toward scalable, detailed, 3D measurement. Yet, progress is constrained by label scarcity for tree instances; a single hectare can hold millions of points and hundreds of overlapping, complex crowns, making manual annotation from scratch with raw data laborious and error‑prone. Annotations are often corrected from automatic pre‑segmentations, but remain costly as these provide no interactive or AI‑assisted refinement. Inspired by the promptable paradigm of foundation segmentation models, we propose SelectAnyTree, a promptable instance segmentation model that delineates any individual tree in a 3D forest point cloud from a few clicks. It introduces two key components: Click‑to‑query prompt encoder and Canopy Height Model (CHM)‑guided first prompt. The former turns each click into a single content query, encoding its 3D position and positive/negative polarity together with a pooled local backbone feature. The latter provides treetops as a geometry‑ and ecologically guided first prompt without any user input. The resulting prompt query is then decoded into one tree mask by a state‑space query decoder to efficiently capture long‑range context in large‑scale forest scenes with linear‑time complexity. We evaluate SelectAnyTree in interactive and instance‑level settings across seven diverse forest regions and an independent held‑out test dataset, demonstrating strong generalization beyond the training domains. It segments a target tree to 78.2 Intersection over Union (IoU) from a single click, 24.8 points above the strongest promptable baseline, and reaches every accuracy target with the fewest clicks, while using far fewer parameters and less inference time than prior promptable models. The source code is available at https://github.com/thanhhff/SelectAnyTree.
Authors:Jingfeng Mao, Xuyang Chen, Qilin Zhang, Oussema Dhaouadi, Guangming Wang, Brian Sheil, Daniel Cremers, Yan Xia, Olaf Wysocki
Abstract:
Aerial 6DoF localization typically relies on precise GNSS signals or radiometrically rich 3D reconstructions, limiting scalability and on‑board deployment. We propose SemCityLoc, a semantic‑geometric alignment system that reframes aerial pose estimation as structured surface registration between foundation‑model‑derived visual priors and standardized LoD‑compliant 3D city models. Instead of matching sparse contours or dense texture, our method aligns semantic surfaces and monocular depth with lightweight semantic 3D building models, increasing pose discriminability in repetitive and occluded urban environments. To enable accurate evaluation, we introduce SemCityLockeD, the first real‑world benchmark combining centimeter‑accurate UAV poses with standardized LoD1‑‑LoD3 semantic city models and challenging low‑altitude imagery. Experiments demonstrate substantial improvements over existing map‑based approaches, improving recall by up to 36% and reducing mean positional error from 9.89m to 2.62m in challenging urban canyons. Our results indicate that semantically structured geometry provides sufficient and scalable constraints for high‑precision aerial localization without radiometric scene reconstructions. The code and data are available at https://albertchen98.github.io/SemCityLoc.
Authors:Jingjun Sun, Chaowei Wang, Zhirui Liu, Jiaxu Tian, Ming Yang, Yaoxing Wang, Shan Gao
Abstract:
3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object‑relation‑object graphs, providing a compact relational abstraction for spatial understanding. In embodied intelligence settings, the same 3D scene may be observed by agents from viewpoints that differ by yaw rotations. However, current 3DSGG models often fail to produce relation predictions that follow the expected transformation behavior under such viewpoint shifts. This behavior reveals an empirical mismatch related to predicate‑level transformation heterogeneity: directional predicates such as left, front, right, and behind should transform with the observation frame, whereas most contact, support, and semantic predicates such as standing on and attached to should remain stable. To reduce this mismatch, we propose Transformation‑Aware Decoupling (TAD), a viewpoint‑robust 3DSGG framework that decouples relation reasoning according to predicate transformation behavior and is supported by viewpoint‑stable object representations. TAD decomposes relation reasoning into two parts: one learns cues that should stay stable across viewpoints, while the other learns directional cues that should change with the observation frame. The two parts are merged for standard multi‑label predicate prediction. Transformation‑specific descriptors and group‑aware auxiliary supervision encourage the two branches to capture complementary relation cues. Extensive experiments on 3DSSG show that TAD achieves state‑of‑the‑art robustness under yaw viewpoint changes without training‑time rotation augmentation, while maintaining competitive performance under the standard benchmark. The project page is available at https://tad‑predicate.github.io/.
Authors:Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Khan
Abstract:
Recently, self‑evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi‑role self‑play and self‑consistency reward schemes in existing self‑evolving LMMs optimize answer agreement without ensuring the decoder attends to visual content, relying instead on statistical language priors to produce self consistent outputs. This leads to a persistent failure mode we term visual under‑conditioning, where the decoder relies on language priors rather than the image during generation, manifesting as insufficient attention to visual tokens. As a result, current self‑evolving LMMs struggle on vision‑‑language understanding tasks such as image captioning and visual question answering. To address this, we propose VISE (Visual Invariance Self‑Evolution), a purely unsupervised self‑evolving framework that directly regularizes the model's visual conditioning policy through two complementary invariance‑based rewards: a geometric invariance reward that enforces spatial consistency under known transformations, and a semantic invariance reward that penalizes evidence‑agnostic generation by requiring the model to recognize the absence of evidence when predicted regions are perturbed. VISE operates within a single model without specialist roles, external reward models, or annotations, and is trained on raw unlabeled images. Experiments on 18 benchmarks demonstrate the efficacy of our approach. Using Qwen3‑VL‑2B as the base model, VISE achieves gains of +16.85 CIDEr on COCO and +19.66 CIDEr on TextCaps, reduces object hallucination by 5.0 Chair‑I points, and generalizes across four model families and scales. Our code and models are available at https://mbzuai‑oryx.github.io/VISE
Authors:Yiming Chen, Yushi Lan, Andrea Vedaldi
Abstract:
We present PhysiFormer, a diffusion transformer for physically‑plausible 3D object motion. Unlike video world models that operate in view‑dependent pixel space, PhysiFormer represents objects as 3D meshes expressed in world coordinates. Given the initial vertex positions and velocities, as well as object material type, rigid or elastic, the model samples future vertex trajectories. While related neural physics approaches build on ad‑hoc latent spaces or explicitly enforce rigidity and causality, PhysiFormer shows that excellent results can be obtained without any such inductive biases, by casting vertex trajectory prediction as a single denoising diffusion process directly in world coordinates. The probabilistic formulation captures uncertainty in the learned dynamics, enabling diverse plausible futures from initial conditions, making this framework potentially useful for applications with unobserved uncertainty. The model features attention factorised over time, space, and objects for efficiency, enabling permutation‑invariant multi‑object reasoning without needing explicit object encoding. Trained on over 100k simulated trajectories, PhysiFormer generates rigid and elastic mechanics, and generalises to mixed‑material settings, unseen real‑world geometries, and larger object counts. It substantially outperforms autoregressive baselines in trajectory accuracy, rigidity preservation, and momentum‑based physical consistency. Our results position coordinate‑space diffusion as a promising step toward view‑invariant, geometry‑aware world modelling for robotics, graphics, and physical design. Visualisations, code, and models are available at https://yimingc9.github.io/physiformer.
Authors:Haina Jiang, Liam Wang, Peng-Chen Chen, Min Seop Kwak, Seungryong Kim, Brian Bell, Jeong Joon Park
Abstract:
Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical task: once trained, they struggle to correct their own constraint violations and extrapolate beyond the training distribution. Recent hybrid methods promote physical correctness by targeting the PDE residual via gradient descent or Gauss‑‑Newton steps, but inherit the compute cost and instability of the underlying classical optimizers. We show, theoretically and empirically, that numerically minimizing the PDE residual can be an unreliable proxy for reconstruction accuracy in ill‑conditioned systems, explaining why these methods often do not make accurate predictions despite achieving low residuals. We propose error‑conditioned Neural Solvers (ENS), built on a different principle: rather than an optimization target, the PDE residual field is passed as a direct input to the network at each iteration, enabling it to read the spatial structure of its own errors and learn an update policy to iteratively correct its predictions. Across four PDE families, ENS attains the highest prediction accuracy in the large majority of settings, with gains reaching 10× on turbulent Kolmogorov flow, while avoiding the expensive compute cost of hybrid methods. ENS's learned correction policy generalizes under distribution shift, including zero‑shot parameter changes and cross‑equation transfer, where its relative advantage is largest in the ill‑conditioned regimes where residual minimization is least reliable. Project website: https://neuralsolver.github.io/.
Authors:Junwei Luo, Shuai Yuan, Zhenya Yang, Yansheng Li, Zhe Liu, Hengshuang Zhao
Abstract:
Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather‑driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land‑surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion‑based methods typically treat weather variables as undifferentiated conditioning signals, and existing benchmarks focus mainly on reconstruction accuracy rather than whether forecasts respond correctly to changed weather forcing.We introduce EO‑WM, a video diffusion transformer for multispectral EO forecasting. EO‑WM incorporates a physically informed conditioning framework that represents meteorological forcing through a climatological baseline, weather anomalies, and cumulative physical stress signals. Specifically, it separates baseline and anomaly through distinct conditioning pathways, and accumulates anomalous forcing over time to capture sustained heat and drought stress. To evaluate weather‑response behavior beyond standard metrics, we introduce two diagnostic benchmarks: an Extreme Summer Benchmark for severity‑aware prediction of vegetation degradation under extreme weather, and a Seasonal Matched‑Pair Benchmark for testing response fidelity under changed weather forcing. Experiments show that EO‑WM reduces the error in predicted Normalized Difference Vegetation Index (NDVI) decline amplitude by a relative 5.63% and improves directional hit rate by a relative 7.80%, while remaining competitive on standard pixel‑level metrics. The benchmarks and model will be made open‑source at https://github.com/Luo‑Z13/EO‑WM.
Authors:Jiyong Kim, Shuang Song, Ronjgun Qin
Abstract:
Gaussian Splatting has been recently explored for satellite 3D reconstruction, demonstrating flexibility and efficiency in representing radiometrically diverse satellite scenes. However, the limited top viewpoint of satellite imagery results in insufficient supervision on building facades, leaving surface holes and degraded visual fidelity. Generative refinement, which leverages pretrained generative priors to iteratively refine and update the rendered images used as supervision targets, has recently been investigated to improve the visual fidelity of Gaussian‑rendered images. However, since these models refine each view independently, the resulting images can generate hallucinations and break photo‑consistency, leading to geometric degradation. To address these limitations, we propose SatSplatDiff, which aims to minimize geometric degradation prevalent in generative refinement. Building on photogrammetric DSM initialization and 2DGS‑based shadow casting established in our prior work SatSplat, we first introduce monocular depth supervision and multi‑scale geometric refinement to establish a geometrically accurate and well‑regularized surface representation. We then apply shadow‑guided generative refinement, where geometrically calculated shadow maps guide the Gaussians to maintain consistency with the underlying geometry, improving visual fidelity while reducing geometric degradation. Extensive evaluations on the IARPA2016 and DFC2019 datasets demonstrate state‑of‑the‑art performance, reducing geometric MAE by up to 18% and improving visual fidelity (FID‑CLIP) by 28‑45% over existing baselines. Our method delivers up to 5x resolution enhancement with minimal hallucination and sensor‑consistent appearance, demonstrating seamless cross‑tile consistency and strong scalability for large‑scale reconstruction. Source code is available at https://github.com/GDAOSU/SatSplatDiff
Authors:Jinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang Jiang
Abstract:
Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically assess understanding and generation capabilities in isolation, overlooking the synergy between comprehension and generation. To bridge this gap, we introduce Unison, a comprehensive benchmark comprising 2,169 high‑quality unified task samples, designed to evaluate joint understanding and generation in unified multimodal models. Unison offers three key strengths: 1) Comprehensive Dimensions: Unison encompasses internal consistency, understanding‑guided generation, generation‑guided understanding, and mutual enhancement to enable holistic evaluation. 2) Diagnostic Evaluation: it provides both unified and decoupled tracks for understanding and generation, allowing fine‑grained attribution of failure modes and quantitative analysis of the gains from unified modeling. 3) Human Alignment: we also introduce Unison‑Judge, an evaluation model well aligned with human judgments to ensure reliable assessment. Based on systematic evaluations of state‑of‑the‑art models on Unison, we uncover critical limitations in current unified multimodal systems and highlight promising directions for future research. Codes, Unison and Unison‑Judge are publicly available at https://github.com/FudanCVL/Unison.
Authors:Jiahe Chen, Qian Shao, Qiyuan Chen, Jiaying He, Jintai Chen, Jian Wu, Hongxia Xu
Abstract:
Open‑set semi‑supervised learning aims to leverage unlabeled data that may contain out‑of‑distribution outliers while maintaining performance on in‑distribution classes. Existing methods mainly follow two paradigms: filtering suspicious samples or incorporating unlabeled objectives with soft weighting. We argue that both face a common trade‑off: aggressive filtering can discard informative but hard ID samples, whereas utilization can introduce auxiliary gradients that conflict with supervised learning when pseudo labels are wrong. We therefore shift the focus from sample selection to gradient‑level control. We propose Geometric Gradient Rectification (GGR), a plug‑in framework that uses the supervised gradient as an anchor and projects conflicting auxiliary gradients onto an admissible region in gradient space. This makes the applied auxiliary update first‑order non‑opposing within the rectified coordinate block while preserving orthogonal components that may still carry useful representation signals. We further extend GGR with subspace‑aware rectification to stabilize the anchor under noisy mini‑batch gradients. Experiments on CIFAR and ImageNet benchmarks show that GGR improves representative OSSL baselines in most settings and yields gains in both closed‑set generalization and open‑set robustness. Code will be available at https://github.com/JiaheChen2002/GGR.
Authors:Ricardo da Rocha Carvalho, Eloísa Oliveira, Luiz Bernardo Martins Kummer, Emerson Cabrera Paraiso, Rayson Laroca
Abstract:
Introduction: Most Multiplayer Online Battle Arena (MOBA) analytics studies rely on structured data, which does not directly capture what each team could actually see during a match. Objective: This work introduces Dota2‑Vis, a video‑based dataset, and a baseline pipeline for visibility analysis in professional Dota 2 matches. Methodology: The dataset comprises all 144 matches from The International 2025, recorded from both team perspectives, totaling 288 Full HD videos, together with 2,477 manually annotated minimap images. We evaluate multiple variants of a modern object detector for player‑icon detection and use the best‑performing model to estimate opponent‑visible player presence over time. Results: YOLO11l (large) achieved the best overall performance, reliably identifying player icons even in dense and visually cluttered minimap scenes. The resulting visibility curves reveal player, hero, role, and team‑level patterns that complement conventional MOBA analytics, highlighting behavioral differences that are difficult to obtain from structured data alone. The dataset and code are publicly available at https://github.com/RicardoRCarvalho/dota2‑vis/.
Authors:Shuchao Duan, Alan Whone, Hossein Rahmani, Jun Liu, Majid Mirmehdi
Abstract:
Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable facial motion evidence that supports the prediction. This limits interpretability and makes it difficult to inspect the basis of model outputs in Parkinson's disease assessment. To address this gap, we propose TraMP‑LLaMA, a unified multimodal framework that jointly predicts severity scores and generates structured textual reports from facial motion cues. The framework integrates RGB appearance and landmark trajectory cues, and adopts a decoupled instruction‑tuning strategy to reduce task interference between severity prediction and language generation. To support this task, we further extend the PFED5 dataset with expert‑guided textual motion descriptions and construct PFED5‑plus. Experiments on PFED5‑plus show that TraMP‑LLaMA outperforms competitive video‑language baselines in report generation and achieves the best severity prediction performance among the compared methods under joint multi‑expression training, improving Spearman's rank correlation by at least 4.39 percent over all competing methods. The text annotations and code are available at https://github.com/shuchaoduan/TraMP‑LLaMA.
Authors:Kexu Cheng, Zicheng Liu, Mingju Gao, Chunhe Song, Hao Tang
Abstract:
Developing physically aware video generation models remains a significant challenge due to the difficulty in capturing diverse physical phenomena, such as thermal dynamics, mechanics, and optics. In this work, we introduce PhysRAG, a novel pipeline that enhances physical awareness in video generation through Retrieval‑Augmented Generation (RAG). To address the issue of limited high‑quality data, we design a two‑stage data filtering pipeline based on the WISA‑80K dataset, resulting in a curated set of 7K high‑quality videos for training. Furthermore, we construct a physical video database and develop a mechanism to inject physical knowledge into a video diffusion model using learnable queries. Our method achieves state‑of‑the‑art performance in both visual quality and physical rule compliance, surpassing existing models in benchmarks such as PhyGenBench and VBench. We conduct extensive ablation studies to validate the effectiveness of our key components, including the data filtering pipeline, RAG mechanism, and method for physical information extraction. To facilitate future research, our code, data, and models are prepared for release at https://github.com/sediment1024/PhysRAG.
Authors:Daniel Barath
Abstract:
Rolling shutter (RS) cameras equip virtually all consumer devices, yet RS‑aware relative pose estimation has remained impractical: the state‑of‑the‑art solver requires a minimum of 20 point correspondences, making RANSAC‑based robust estimation prohibitively expensive due to the exponential dependence of the iteration count on the sample size. We make RS relative pose estimation practical by introducing affine correspondences (ACs) into the RS two‑view geometry. We derive novel \emphRS‑corrected affine constraints that account for the coupling between point perturbations and the row‑dependent essential matrix, providing two equations per correspondence beyond the standard epipolar constraint. Building on these constraints, we develop a linearized algebraic solver that estimates pose and RS motion from only 7 ACs. The solver exploits the physical smallness of RS parameters to linearize the constraints, eliminates the 12 RS unknowns via null‑space projection, and solves the remaining degree‑20 system via action matrices in 1.2\,ms. On the TUM RS benchmark, our method achieves the best pose and RS parameter accuracy among all tested methods and, uniquely among RS solvers, provides accurate translational velocity estimates ‑‑ which are poorly conditioned from point correspondences alone due to a \vecv‑\vect coupling. On the global‑shutter EuRoC MAV dataset, the solver achieves comparable accuracy to the standard 5‑point algorithm, demonstrating that it generalizes well to the GS setting. Code is at https://github.com/danini/rolling_shutter_made_practical.
Authors:Ke Chen, Ling Zhou, Guangqi Jiang, Gengshen Wu, Yi Liu, Shoukun Xu
Abstract:
General Salient Object Detection (SOD) aims to identify and segment visually interesting objects from uni‑modality or multi‑modality scenes, recently advanced by cutting‑edge State Space Models (SSMs). However, a critical limitation of current approaches is their neglect of the inherent spectral biases exhibited by different neural network paradigms. By digging to the dataset‑level spectral analysis of Convolutional Neural Networks (CNNs) and SSMs, their semantic representations are inherently complementary based on their complementary frequency preferences. Inspired by this, we harmonize heterogeneous representations from SSMs and CNNs to bridge their spectral biases for general salient object detection. To this end, inspired by the dynamic information propagation of Liquid Neural Networks (LNNs), we introduce a liquid fusion to dynamically integrates features from two backbones, including VMamba and ConvNeXt, referred to Liquid Fusion Network (LFNet). Concretely, by treating the continuous VMamba features and ConvNeXt features as evolving states and exogenous stimulus, respectively, LFNet employs a dynamic gating mechanism for content‑aware feature aggregation. Crucially, this state‑stimulus paradigm enables to scale to multi‑modal cues, resulting in flexibility in general SOD. Besides, a Saliency‑Guided Upsampling (SGU) operator to propagate the features to the shallow layer, which leverages a spectral‑spatial co‑design to suppress upsampling artifacts while preserving semantics. Extensive experiments across five diverse tasks (RGB, RGB‑D, RGB‑T, VSOD, and VDT) demonstrate that LFNet achieves state‑of‑the‑art performance, offering a superior trade‑off between detection accuracy and model efficiency. Code has been released at https://github.com/cke520/LFNet.
Authors:Xilai Li, Xiaosong Li, Haishu Tan, Tao Ye, Huafeng Li, Hongbin Wang
Abstract:
Multi‑modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, causes significant image degradation, disrupting feature representation and requiring simultaneous feature restoration and cross‑modal complementarity. Existing methods often struggle with effective representation learning under such conditions, limiting their practical performance. To address these challenges, we propose a mask‑guided MMIF method that integrates feature restoration and interaction. We first introduce "Pseudo Ground Truth" to simplify training, promoting faster and more effective feature learning. Then, we design a mask generation mechanism based on the mapping relationship between the fused result and the source images, quantifying the relative contribution of each modality during the fusion process. By incorporating the proposed mask‑guided cross‑modal cross‑attention mechanism, the network is encouraged to selectively attend to informative features during modality interaction, mitigating the risk of overfitting to the static distribution of the "Pseudo Ground Truth". Additionally, we propose a mask‑guided learning strategy and a task‑coupled degradation‑aware learning strategy to balance feature restoration and interaction. Extensive experiments on synthetic and real‑world datasets demonstrate that our method surpasses state‑of‑the‑art approaches in visual quality, quantitative metrics, and downstream tasks. The source code is available at https://github.com/ixilai/AMG‑Fuse.
Authors:Yuan Xu, Yixiang Chen, Kai Wang, Jiabing Yang, Peiyan Li, Qisen Ma, Yan Huang, Liang Wang
Abstract:
Vision‑Language‑Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine‑tuning, however, action supervision applies equally across all timesteps, without structured supervision on which manipulation stage the robot is in or what the next gripper‑event target should be. This causes failures to concentrate around challenging gripper‑event transitions. To address this, we propose StaKe, a plug‑in auxiliary supervision framework that automatically derives two complementary signals from demonstration gripper states without manual annotation: a stage classifier that identifies the current manipulation stage, and a keyframe predictor that estimates the target joint action at the next gripper transition. Both are modeled as lightweight auxiliary heads that enrich the learned representations during training, while leaving the base VLA policy architecture and inference loop unchanged. Experiments on bimanual simulation and single‑arm Franka real‑robot tasks show that StaKe consistently improves success rates (relative gains of 14% and 56%, respectively), with larger improvements on longer‑horizon tasks that involve more gripper‑event transitions. Ablation studies validate each design choice, and qualitative analysis confirms that the learned representations faithfully track manipulation stages. These results indicate that structured supervision is an effective and general strategy for enhancing VLA fine‑tuning in long‑horizon manipulation. Project website: https://hi‑yuanxu.github.io/StaKe‑Web/
Authors:Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto, Shi Qiu, Jamal Bentahar, Naveed Akhtar, Mubarak Shah
Abstract:
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image‑text alignment. As downstream applications increasingly demand visually grounded commonsense inference and compositional reasoning, it remains unclear whether CLIP‑style encoders can support such reasoning without architectural changes. To address this, we present ReasonCLIP‑58M, a continual pretraining framework that integrates large‑scale reasoning supervision into CLIP‑style models through our two‑stage strategy, which progressively integrates reasoning signals while preserving descriptive alignment, followed by category‑structured reasoning supervision. To support this framework, we construct two complementary datasets and a benchmark: ReasonLite‑42M, with open‑form, visually verifiable reasoning captions; ReasonPro‑16M, with category‑specific reasoning supervision; and RCLIP‑Bench for diagnostic evaluation of visually grounded reasoning. We train a family of ReasonCLIP that improves visually grounded commonsense and compositional reasoning while also enhancing zero‑shot retrieval performance. As a drop‑in visual encoder for multimodal large language models such as LLaVA‑NeXT, ReasonCLIP delivers consistent gains without additional inference cost, demonstrating that structured reasoning supervision enhances the expressive capacity of CLIP‑style visual representations. All datasets, models, and training code are available at https://github.com/RISys‑Lab/ReasonCLIP.
Authors:Xuyue Huang, Zhe Chen, Wang Shen, Xiao-Ping Zhang
Abstract:
Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs. Feature caching accelerates inference by reusing intermediate representations. Existing methods rely on historical features for implementation simplicity, yet suffer from severe error accumulation at high acceleration ratios. To address this limitation, we investigate the nature of the requisite feature correction. We demonstrate that the optimal calibration update is characterized by a shared low‑rank subspace across diverse prompts. Guided by this structural insight, we propose LearniBridge, a learnable calibration mechanism for feature caching that bridges multiple timesteps through lightweight LoRA updates. This mechanism enables effective calibration requiring only 3‑5 training samples. Extensive experiments on image and video generation show that LearniBridge achieves up to 5.87×, 5.75×, and 4.10× acceleration on FLUX, HunyuanVideo, and WAN2.1, respectively. On WAN2.1, it improves VBench by 1.28% over the previous SOTA at 4.10× acceleration. Our code is available at https://github.com/Iiiiiiirene/LearniBridge.
Authors:Yiheng Cao, Gustavo Andrade-Miranda, Jiatian Zhang, Lingxiao Zhao, Xin Gao
Abstract:
Developing robust artificial intelligence models for 4D (3D + time) medical imaging is constrained by limited annotated data, inter‑device domain shifts, and privacy restrictions. To address this, we propose a 4D controllable generative framework for anatomically consistent data augmentation. A semi‑supervised variational autoencoder learns a compact latent representation of anatomical volumes while jointly predicting aligned segmentation masks in a unified framework. Anatomical structure is then disentangled from temporal dynamics through a cascaded latent diffusion model (LDM). A static LDM generates subject‑specific anatomy conditioned on clinical priors (diagnosis and volumes measures) and a subsequent motion LDM estimates residual latent motions, ensuring strict temporal coherence across the 4D sequence. The proposed approach was evaluated on cine cardiac MRI as a representative 4D imaging application. Experiments across multiple datasets demonstrate high controllability of static anatomy (Pearson r > 0.8) and strong temporal coherence (FVD = 288.08). In cross‑vendor generalization experiments, augmenting training sets with synthetic 4D sequences significantly improves downstream segmentation performance. Using nnU‑Net, the proposed augmentation strategy improves the average Dice score by 1.4% and reduces the Hausdorff Distance by 3.0mm compared to training on real data alone, for the left ventricle, Dice improves by 2.8% with a 5.4mm reduction in boundary error. Overall, this framework provides a scalable and controllable solution for 4D medical image synthesis, supporting the development of more robust models with limited annotations and cross‑vendor variability. Code available on https://github.com/cyiheng/4DCardiacMRISynthesis.
Authors:Honghang Chen, Xiujun Zhang, Xiaoli Sun, Mingqing Xiao
Abstract:
Implicit neural representation (INR) has emerged as a powerful prior for multi‑dimensional data (e.g., multispectral images and videos). However, most INR methods employing periodic activation functions (e.g., Sine) predominantly rely on function composition. This mechanism introduces optimization instability as network depth increases, thereby limiting their performance. Meanwhile, these methods fail to incorporate proper physical priors to effectively alleviate spectrum bias. To address these issues, inspired by the commonalities between deep periodic networks and generalized Fourier series, we propose a novel Calibrated Harmonic Overlaid Implicit Neural Representation (CHOIR). Specifically, we utilize Coordinated Harmonic Superposition (CHS) to replace the conventional function composition used in most INRs, thereby ensuring optimization stability when scaling network depth. Furthermore, we introduce a Perceptual Spectrum Calibration (PSC) to mitigate spectrum bias. This calibration embeds the ubiquitous power‑law spectrum prior of natural images and adjusts the globally fixed spectrum towards a physically plausible log‑uniform distribution. Extensive experiments on various multidimensional data recovery problems demonstrate that our method achieves superior performance over state‑of‑the‑art approaches. Code is available at https://github.com/chorl0229/CHOIR.
Authors:Zhihao Wen, Yixin Yang, Bojian Wu, Yang Zhou, Dani Lischinski, Daniel Cohen-Or, Hui Huang
Abstract:
While 3D Gaussian Splatting (3DGS) provides an efficient and explicit representation for novel view synthesis, enforcing stylistic coherence across viewpoints remains challenging. Existing 3D stylization methods typically apply 2D feature‑matching losses independently per rendered view, which leads to unstable style allocation, many‑to‑one feature reuse, and limited cross‑view consistency. We propose a capacity‑controlled framework for multi‑view stylization of 3DGS, grounded in optimal transport. Specifically, we reformulate local style matching as a semi‑balanced optimal transport problem. By introducing explicit column‑capacity constraints with tunable strength, our formulation mitigates many‑to‑one matching and enables controllable allocation of style features. This transport‑based objective provides a principled mechanism for balancing feature coverage and stylistic diversity while maintaining stable correspondences across viewpoints. To further enhance cross‑view coherence, we incorporate a novel cross‑view matching guidance to constrain correspondences between scene content and style patterns. In addition, we introduce several geometric regularizations to enhance the vanilla 3DGS, thereby enabling optimized Gaussian primitives to represent finer‑grained textures during stylization. Extensive experiments demonstrate that our approach significantly improves multi‑view stylistic consistency and produces stable, expressive 3D stylizations while preserving the core semantic structure of the scene.
Authors:Haofei Song, Siyuan Xu, Xintian Mao, Shaojie Guo, Qingli Li, Yan Wang
Abstract:
Arbitrary slice super‑resolution reconstructs isotropic volumes from anisotropic clinical acquisitions by synthesizing intermediate slices at arbitrary scales. However, treating this ill‑posed inverse problem as unconstrained residual‑based regression risks hallucinating anatomically implausible structures or altering the originally observed data. To address both concerns, this paper presents the Dual‑Prior Null‑space Learning (DP‑NSL) framework, which reformulates the task as a constrained recovery process guided by two complementary priors. A Measurement‑Consistent Projection (MCP) enforces a Deterministic Observation Prior: the reconstruction undergoes an exact orthogonal projection that reproduces every acquired slice with zero error, confining all learned details to the unobservable null space. Within this null space, a Mixture‑of‑Splines (MoS) module imposes a Geometric Continuity Prior by dynamically mixing B‑spline experts of different analytic orders, allowing each anatomical region to be modeled with a content‑aware level of continuity. To promote spatial coherence, a Local Spatial Consistency Decoder (LSCD) further injects local inductive bias. Experiments on three CT and one MRI benchmark show that DP‑NSL outperforms existing approaches while strictly preserving measurement consistency. Code is available at https://github.com/DeepMed‑Lab‑ECNU/Medical‑Image‑Reconstruction.
Authors:Kim Youwang, Jon Hasselgren, Peter Kocsis, Andrea Weidlich, Tae-Hyun Oh, Jacob Munkberg
Abstract:
Neural materials can represent complex specular reflections and scattering effects in a compact, universal basis. However, acquiring and authoring such materials remains challenging. We present NeuMatEx, a differentiable inverse rendering method for extracting spatially varying neural materials from images. The nonlinear structure of neural material latent spaces makes optimization with naive inverse rendering infeasible. To address this, we train a Large Material Reconstruction Model (LMRM) that directly predicts initialbase color, neural material latents, and aleatoric uncertainty guides from images. This material prior provides a good initialization and better constrains our subsequent optimization using inverse path tracing. The predicted uncertainty further helps by anchoring high‑confidence regions more tightly to the LMRM prediction, preventing lighting and complex specular effects from being baked into materials. Experiments on synthetic and real assets show that NeuMatEx extracts complex materials with better visual quality and material decomposition than PBR‑based methods.
Authors:Jingjun Gu, Chaojie Shen, Yifeng Cao, Wei Zhang, Yiliu Li, Aobo Fan
Abstract:
Skin lesion segmentation is a key task in computer‑aided dermatological diagnosis, where accuracy directly impacts downstream analysis and disease classification. However, dermoscopic images are challenging due to blurred boundaries, low contrast, large shape variations, and artifacts such as hair and shadows. Recently, diffusion models have shown strong performance in medical image segmentation thanks to their progressive denoising and distribution modeling capabilities. Nevertheless, existing diffusion‑based methods still suffer from limited cross‑level feature interaction and insufficient boundary detail recovery. To address these issues, we propose MLFFM‑SegDiff, a multi‑level feature fusion diffusion model for skin lesion segmentation. Built on a diffusion framework, the method introduces a dual‑path U‑Net encoder, a Multi‑Level Feature Fusion Module (MLFFM), and a boundary‑sensitive loss function. The dual‑path encoder enhances interaction between noisy mask features and dermoscopic image features. MLFFM improves skip connections via attention, scale alignment, and adaptive cross‑level fusion. These designs enable the decoder to jointly leverage shallow boundary cues and deep semantic representations, improving mask reconstruction quality. Experiments on ISIC2018, PH2, and HAM10000 demonstrate that MLFFM‑SegDiff outperforms representative methods including DermoSegDiff, U‑Net, and SwinUNETR across Accuracy, F1‑score, Jaccard index, Recall, and Dice. In particular, it achieves an average Jaccard index of 0.8546 and Dice coefficient of 0.9207. These results validate the effectiveness of the proposed multi‑level feature fusion strategy for improving lesion segmentation performance. The code will be released at https://github.com/Qacket/MLFFM‑SegDiff.git after publication.
Authors:Quan Zhou, Shaoqing Zhai, Qiang Hu, Jia Chen, Qiang Li, Zhiwei Wang
Abstract:
Transforming foundation segmentation models from human‑prompted tools into auto‑promptable annotators is critical for scalable medical data annotation. Current methods commonly depend on external feature matchers or auxiliary networks to automate geometric prompting, but introducing architectural overhead and limiting performance scalability. Although SAM3 natively supports concept segmentation via reusable text prompts, its direct use in medical imaging is hindered by a lack of fine‑grained clinical knowledge and the ambiguity of human‑written descriptions. In this work, we propose Mask to Concept (M2C), an efficient framework that adapts SAM3 for medical few‑shot annotation without external modules, parameter retraining, or manual text engineering. Using only a few labeled images, M2C enables SAM3 to automatically search for transferable visual concepts entirely within its frozen architecture: it initializes a learnable concept embedding, uses it to prompt segmentation, and updates the embedding by gradients of minimizing the concept segmentation error. We further introduce a Hybrid Uncertainty Estimation (HUE) module that calculates the prediction entropy and maps concept predictions back to the box prompts, measuring concept‑geometry prompting inconsistency. Highly uncertain samples are flagged actively for human correction, and the corrected masks are then fed back to M2C to continuously search for more precise concept embeddings, forming a self‑enhancing annotation loop with minimal expert effort. Experiments on medical segmentation benchmarks show that our method achieves SOTA few‑shot segmentation performance and outstanding annotation efficiency, offering a practical and efficient pathway toward scalable medical image labeling. Codes are at https://github.com/Huster‑Hq/M2C.
Authors:Bin Hu, Yanwen Ma, Jiehui Huang, Ziliang Zhang, Haoning Wu, Ruicheng Zhang, Yaokun Li, Zijun Wang, Yuechen Zhang, Chun-Mei Tseng, Hanhui Li, Shengju Qian, Jun Zhou, Kaipeng Zhang, Xiaodan Liang, Jiaya Jia, Xiu Li
Abstract:
Recent game world models can synthesize visually plausible, action‑conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering trajectories, and physical dynamics are typically learned as implicit correlations from data rather than as controllable variables. This limitation hinders their applicability to authored game environments, where physical rules are deliberately designed and require explicit manipulation. We introduce PhysEditWorld, a multimodal dataset with physical parameters, with a primary focus on gravity in this initial version. At its core, PhysEditWorld is built upon a replay paradigm implemented with a UE5 replay‑and‑rendering pipeline. Each scenario records a normalized action trace and replays the same initial state, character controller, action sequence, and camera policy under multiple gravity configurations, enabling controlled and attributable physical variation. PhysEditWorld contains 12 cinematic UE5 scenes, over 100 hours of gameplay interactions, and more than 60 million rendered rollout frames. Each sample provides synchronized multimodal signals, including RGB, depth, normals, audio, action traces, camera trajectory, engine states, semantic annotations, and explicit gravity labels. We further conduct initial utility studies on both generative video models and world understanding models, demonstrating that PhysEditWorld enables improved gravity‑faithful dynamics modeling, enhances consistency under physical edits, and provides a scalable foundation for controllable world modeling research.
Authors:Hongjae Lee, Sojung Kang, Jaeseong Yu, Seung-Won Jung
Abstract:
While traditional image restoration focuses on perceptual quality, Task‑Driven Image Restoration (TDIR) aims to maximize the performance of downstream high‑level vision tasks. Recent approaches leveraging generative priors have shown promise for TDIR; however, they typically suffer from computational inefficiency and potential semantic alteration by indiscriminately updating all latent tokens. In this paper, we posit that not all visual information is equally important for machine perception. Through an analysis of the latent token space, we observe that task‑relevant cues are unevenly distributed across the token sequence, exhibiting index‑wise specialization. This suggests that selectively refining a subset of tokens can be sufficient for task‑driven objectives. Leveraging this insight, we propose TaskTok, a novel framework that selectively restores only task‑relevant tokens via a learnable token switch and a lightweight token refinement module. Extensive experiments across image classification, semantic segmentation, and object detection demonstrate that TaskTok significantly enhances task performance with high computational efficiency. The source code is available at https://github.com/jimmy9704/TaskTok
Authors:Hongjae Lee, Myungjun Son, Jaeseong Yu, Seung-Won Jung
Abstract:
Image restoration aims to reconstruct high‑quality images from degraded low‑quality inputs. As the computational demands of image restoration models continue to rise, there is growing interest in lightweight architectures optimized for fast and efficient inference. Logic gate networks (LGNs), which operate using fundamental logic operations such as NAND and XOR, have recently emerged as a promising direction for achieving highly efficient computation. However, their potential remains largely untapped in the domain of image restoration. In this work, we introduce LogicIR, the first LGN specifically designed for image restoration tasks. LogicIR incorporates a UNet‑inspired architecture composed entirely of logic gates. In addition, we propose a differentiable bit decoding layer and an index shuffling mechanism that improves information propagation across logic gates. Experimental results across multiple image restoration benchmarks demonstrate that LogicIR achieves strong performance with significantly reduced computational cost, establishing LogicIR as a viable and efficient alternative for image restoration. The source code is available at https://github.com/jimmy9704/LogicIR
Authors:Geng Li, Yuxin Peng
Abstract:
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine‑grained perception capabilities. However, existing benchmarks predominantly rely on explicit textual cues or low‑resolution inputs, failing to evaluate a model's ability to autonomously perceive implicit visual cues in high‑resolution. To bridge this gap, we introduce DiCoBench, a comprehensive, multi‑image high‑resolution benchmark designed for cross‑image fine‑grained perception. DiCoBench consists of 765 meticulously curated samples categorized into two progressive tracks: Differential Visual Cues and Commonality Visual Cues, covering 8 distinct perception tasks. By formulating the benchmark as a multiple‑choice question task and utilizing high‑resolution imagery (approaching 2K), we eliminate evaluation metric bias and pose a substantial challenge to current state‑of‑the‑art MLLMs. Our extensive evaluation of 18 diverse MLLMs reveals a striking performance gap compared to human accuracy (98.3%), with top‑performing models struggling significantly with micro‑scale detail capture. We believe DiCoBench will serve as a challenging testbed to drive future research in autonomous, high‑resolution multi‑image perception.
Authors:Feifan Luo, Ting Li, Zhao Li, Hongyang Chen
Abstract:
Non‑rigid 3D shape matching is a fundamental task in computer vision and graphics. In this paper, we propose a hybrid self‑supervised method based on a coarse‑to‑fine strategy, which ensures consistency between the coarse mapping and the refined correspondence produced by our refinement module. The architecture features a dual‑branch design, consisting of two symmetric functional map learning streams: one based on the Laplacian basis and the other utilizing the elastic basis. Extensive experiments show that our approach not only maintains computational efficiency, but also achieves state‑of‑the‑art performance across a variety of challenging scenarios, including non‑isometric deformations and topological noise. Finally, we rigorously demonstrate that contrastive energies promote feature discrimination. Furthermore, integrating these energies with existing methods yields consistent improvements, validating the overall efficacy of our approach. Our code is available at https://github.com/LuoFeifan77/Coarse‑to‑Fine‑Hybrid‑Self‑Supervised‑Matching.
Authors:Shengbin Guo, Shaokang He, Chaoyue Meng, Shengpeng Xiao, Xunzhi Xiang, Shaofeng Zhang, Qi Fan
Abstract:
While instruction‑based image editing, enabled by multi‑modal generative models, has advanced significantly, existing benchmarks lack a comprehensive evaluation of physics‑based reasoning, a critical capability for handling real‑world scenarios. To address this, we introduce PhyEditBench, a benchmark designed to assess the physical understanding of editing models. Guided by a hierarchical taxonomy, we establish 4 primary classes and 12 subclasses. It comprises 238 high‑quality, high‑resolution, real‑world instances meticulously extracted from videos to capture authentic physical dynamics, alongside 35 synthetic Anti‑Physics instances. Our empirical analysis of current SOTA editing methods exposes substantial limitations in their physics‑based reasoning. We further propose a training‑free baseline named PhyWorld that uses test‑time scaling and a latent reduction strategy. PhyWorld outperforms comparable models and suggests that the video generation process can effectively serve as a reasoning mechanism for image editing. The project page is available at https://github.com/Previsior/PhyEditBench.
Authors:Zhixing Li, Yinan Yu
Abstract:
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural‑diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. Unlike traditional black‑box QA, CRISP utilizes metric 3D Scene Graphs and an oracle intervention protocol to decouple latent reasoning capabilities from perceptual bottlenecks. This granular diagnosis uncovers a systematic perception‑reasoning disconnect. Crucially, we reveal that while proprietary models possess robust latent reasoning engines, they suffer from inaccurate metric estimation and a critical failure to leverage their implicit structural representations. Conversely, open‑source models remain fundamentally bottlenecked by their lack of multi‑hop compositional reasoning. By shifting the focus from merely ``guessing correctly'' via language priors to genuinely ``perceiving, verifying, and reasoning,'' CRISP offers a rigorous roadmap for multimodal alignment beyond end‑to‑end post‑training. The code and dataset are available at https://github.com/iiyamayuki/CRISP‑Bench.
Authors:Xiao Wang, Xufeng Lou, Zikang Yan, Lan Chen, Sibao Chen, Yaowei Wang, Yonghong Tian, Jin Tang
Abstract:
RGB‑Event tracking improves localization robustness by fusing RGB appearance textures and dense temporal motion cues from event sensors. While this multi‑modal scheme broadens tracking applicability, real‑world scenes suffer diverse structured signal degradations that hinder traditional multi‑modal fusion. In harsh environments, either modality can lose reliability drastically, and targets frequently appear incomplete due to occlusion, edge truncation and foreground clutter.To tackle the above challenges, we present a hierarchical perturbation and retrieval framework tailored for RGB‑Event tracking with robustness against partial target missing and modal degradation, termed APRTrack. To mimic real‑world signal corruption, APRTrack constructs structured degradation via two adversarial perturbation branches at the modality and spatial levels, which separately simulate full‑modal failure and localized target region absence. A hierarchical routing mechanism is designed to disentangle the training pipelines of the two perturbation types, effectively eliminating feature collapse induced by superimposed degradation constraints. Furthermore, we devise Footprint‑guided Channel‑calibrated Hopfield Retrieval (FCHR) for reliable historical information compensation. This module evaluates retrieval confidence based on association footprints between queries and memory banks, and calibrates the retrieval metric space prior to Hopfield matching, realizing controllable historical feature compensation bounded to target regions. Extensive experiments on FE108, COESOT, VisEvent, and FELT datasets demonstrate the effectiveness of our proposed strategies for the RGB‑Event visual object tracking. The source code and pre‑trained models will be released on https://github.com/Event‑AHU/OpenEvTracking
Authors:Baiqi Li, Ce Zhang, Yu Fang, Yue Yang, Shangzhe Li, Mingyu Ding, Gedas Bertasius
Abstract:
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real‑world human‑action video and a language instruction with an aligned simulator scene and an executable LIBERO task, enabling scalable and reproducible evaluation. WatchAct comprises 3,000 long‑horizon instances across 14 tasks in four capability domains drawn from the cognitive demands of watching another agent: parsing events (Event Grounding), recovering procedural structure (Procedural Reasoning), inferring unstated intent (Implicit Intent Inference), and tracking how the scene was changed (Episodic Reasoning). We further propose a disentangled evaluation protocol that separately measures (i)~video‑to‑plan reasoning by vision‑language models, (ii)~policy execution under oracle plans, and (iii)~full task completion by integrated planner‑‑policy pipelines. In both simulation and on a Franka Research 3 robot, current systems remain far from solving WatchAct. The best pipeline, Gemini‑3.1‑Pro with π_0.5, reaches only 16.3% Success Rate (SR) in simulation and 14.0% on the real robot. Gemini‑3.1‑Pro attains just 36.8% Plan SR (vs. 97.1% for humans), while π_0.5 reaches only 21.5% Task SR under oracle plans and drops to 10.6% on out‑of‑domain scenarios. Dataset and code are available at https://baiqi‑li.github.io/watchact_page/.
Authors:Tianle Zhu, Haohua Que, Handong Yao, Hongyi Xu, Zhipeng Bao
Abstract:
High‑precision remote perception is often hindered by the severe bandwidth constraints of Vehicle‑to‑Everything (V2X) networks. We propose DinoLink, a token‑centric compression framework that replaces raw pixel streaming with discrete semantic communication for vehicle‑cloud collaborative inference. DinoLink employs a dual‑sparsity architecture: a saliency‑aware selector prunes redundant background tokens, while a Residual Vector Quantization (RVQ) module collapses features into compact codebook indices. By transmitting only lightweight indices and positional priors, DinoLink achieves a 139× bitrate reduction compared to uncompressed transmission while maintaining a competitive 32.8% mAP on the nuScenes dataset. Deployment simulations further demonstrate a 34.5× acceleration in narrow‑band environments, such as LoRa. Our results substantiate DinoLink as a robust, bandwidth‑efficient frontend for high‑fidelity remote perception in constrained V2X scenarios. The code is publicly available at https://github.com/UGA‑MOBILITY‑LAB/dino_link.
Authors:Hao Sun, Hao Yan, Mengting Chen, Quanjian Song, Yu Li, Juan Cao, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Sheng Tang
Abstract:
While Video Virtual Try‑on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms remains fundamentally constrained by a passive dependency on source camera trajectories, failing to accommodate the requisite interactive freedom for omnidirectional viewpoint exploration. To address this limitation, we define a pioneering research frontier: Camera‑controllable Video Virtual Try‑on (CaM‑VVT). Unlike conventional VVT, CaM‑VVT not only necessitates viewpoint‑agnostic texture hallucination but also strict structural synchronization between non‑rigid human dynamics and background contexts under arbitrary, unconstrained camera movements. To tackle these challenges, we present TryOnCrafter, the first unified DiT‑based framework specifically architected for the CaM‑VVT task. Departing from implicit pixel‑space manipulation, we introduce a Renderable 4D Try‑on Proxy that explicitly decouples the human subject from the environment. This is achieved by distilling high‑fidelity 2D try‑on priors into a clothed 3DGS‑based avatar, which is subsequently animated via SMPL‑X sequences and metric‑aligned into a reconstructed background point cloud. This proxy establishes a robust structural foundation with superior texture density and motion integrity. Our Proxy‑Anchored Video DiT leverages this robust structural foundation as a primary geometric anchor, ensuring that the synthesized photorealistic videos are strictly constrained by prescribed trajectories and physically plausible deformations. Benefiting from the inherent editability of the 4D proxy, TryOnCrafter facilitates diverse downstream applications, including human relocalization, ``bullet time'' effects, and 360‑degree orbital viewing.
Authors:JoungBin Lee, Jaewoo Jung, Jongmin Lee, Tongmin Kim, Hyunsung Kim, Takuya Narihira, Kazumi Fukuda, Jahyeok Koo, Jisang Han, Yuki Mitsufuji, Seungryong Kim
Abstract:
Synthesizing a novel‑view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with respect to the reference video. Existing methods based on explicit 3D representations are limited by the accuracy of off‑the‑shelf reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos. In contrast, camera‑conditioning‑only methods can achieve high visual quality but often struggle to preserve geometric and motion consistency. In this work, we introduce MVTrack4Gen (Multi‑View point Tracking for Novel‑View Generation), a motion‑aware training framework that leverages multi‑view point tracking as an additional geometric and motion supervision signal for camera‑conditioning‑only novel‑view video diffusion models. Our key finding is that specific attention layers encode strong correspondence cues, where query features attend to key features at geometrically corresponding locations across views and over time, and the misalignment of these correspondences causes motion inconsistency. Based on this observation, we route these features into an auxiliary multi‑view tracking head and jointly train the diffusion model with a point‑tracking objective. By explicitly strengthening these motion‑aware correspondences, MVTrack4Gen improves existing models to better follow the motion in the reference view and maintain cross‑view geometric consistency. Across diverse benchmarks, our method achieves state‑of‑the‑art geometric consistency and competitive camera accuracy.
Authors:Yang Chen, Xiaowei Xu, Shuai Wang, Xinwen Zhang, Qiushi Guo, Tiezheng Ge, Limin Wang
Abstract:
Normalizing Flows (NFs) are powerful generative models capable of exact density estimation and sampling. However, their strict invertibility often forces the model to exhaust its capacity on low‑level pixel details, hindering the capture of high‑level semantic structures. While Masked Image Modeling (MIM) has excelled in representation learning, its integration into generative pipelines has remained largely modular and disjointed. In this paper, we propose MIMFlow, a unified end‑to‑end framework that jointly optimizes latent semantics, pixel reconstruction, and generative flow. By employing a VAE encoder to infer semantic latent from masked images, MIMFlow achieves a principled decoupling of the generative task: the Normalizing Flow focuses on modeling a simplified, low‑frequency semantic manifold, while a specialized decoder handles high‑frequency synthesis. This design effectively resolves the inherent capacity bottleneck of NFs, allowing the model to prioritize global structural coherence over redundant noise. Empirical results on ImageNet 256×256 show that MIMFlow‑L reaches 71.3% linear probing accuracy and an FID of 2.50. Despite using only 128 tokens (50% fewer than standard models), it yields a 32.8% performance gain over similar‑scale NF baselines. Our code is available at https://github.com/MCG‑NJU/MIMFlow.
Authors:Long Cao, Zhongquan Wang, Jie Li, Yuhan Chen, Kefei Qian, Xiangfei Huang, Guofa Li
Abstract:
Image priors can synthesize target conditions for 3D Gaussian street scenes, but independently edited views do not define a coherent 3D target. Direct fitting can propagate view‑specific noise, while existing pipelines do not jointly handle imperfect sparse anchors and standard‑rasterizer deployment. To address this gap, teacher‑relative appearance residual distillation is introduced for appearance baking. A structured space for frequency decomposition, confidence estimation, and primitive‑level lifting is formed by residuals between teacher anchors and original renders. The direct optimization signal is supplied by renderer‑space matching, while primitive assignment is regularized by support‑aware Gaussian‑space aggregation. Supported detail is admitted and unsupported noise is suppressed through confidence‑gated coarse‑to‑fine optimization, after which all residuals are baked into fixed‑geometry spherical‑harmonic coefficients. The teacher and auxiliary training modules are discarded at inference. Evaluation across Waymo street assets, Tanks and Temples scenes, and multiple target conditions shows a favorable overall balance of target alignment, content preservation, artifact suppression, and cross‑view consistency over editing‑based baselines. Ablations confirm the effectiveness of the main components. Code will be released at https://github.com/Cagares/Baking‑for‑3D‑Gaussian.
Authors:Nathan Painchaud, Tristan Habémont, Morgane des Ligneris, Allan Serva, Pierre Croisille, Laurent Bertoletti, Thomas Lampert, Johannes F. Lutzeyer, Odyssée Merveille
Abstract:
Risk stratification for pulmonary embolism (PE) is critical for clinical decision‑making. Stratification guidelines are based on patient medical records, parameters measured from computed tomography pulmonary angiography (CTPA), and blood tests. However, blood tests are often missing in routine practice. This work studies whether state‑of‑the‑art models can accurately classify risk stratification from only medical records and biomarkers extracted from CTPA images. We benchmark different approaches to combine medical records and cardiac biomarkers with rich pulmonary vascular information; we add vascular biomarkers to tabular models and apply graph neural networks (GNNs) on the vascular tree's intrinsic graph representation. We use a private dataset (n=353) with uniquely complete data for PE risk stratification. Our results show that, among global features, medical records and cardiac biomarkers are the most significant predictors, while vascular biomarkers do not further improve stratification. Even more surprising, even GNNs on vascular graphs fail to outperform strong tabular baseline on global features. We consider hypotheses, on both models and data, that could explain this suboptimal performance. Our investigation suggests that, counter‑intuitively, vascular graphs might hold no discriminative information for PE risk stratification. Code is available from https://github.com/creatis‑myriad/GENESIS.
Authors:Pengwei Wang, José Morano, Virginia Mares, Hrvoje Bogunović
Abstract:
Color fundus photography (CFP) is the most common ophthalmic imaging modality for large‑scale screening. However, it is highly susceptible to degradations, making robust fundus image quality assessment (FIQA) crucial. The criteria for what constitutes high‑quality at the image level vary across clinical tasks, making FIQA dependent on expert knowledge. This motivated the development of automated methods and datasets. While existing datasets aim to standardize image‑level quality, their criteria often differ. Furthermore, image‑level labels preclude the quantitative evaluation of localized degradations, which is essential for trustworthy FIQA. We argue that pixel‑level FIQA based on anatomical visibility represents a more task‑agnostic, explainable approach. In this work, we introduce FunPiQ, the first FIQA benchmark to provide pixel‑level quality annotations. In addition, we propose EFIQA‑CP, an explainable‑by‑design (EBD) method that uses quality pseudo‑labels based on anatomical visibility to train a CNN via Non‑Negative Positive‑Unlabeled learning. Extensive evaluations of classification methods with post‑hoc explanations, anomaly detection methods, and EBD methods demonstrate the superior performance of the last and, particularly, of EFIQA‑CP.
Authors:Jiacheng Sui, Tianyu Hao, Bingjie Gao, Li Niu, Guangtao Zhai
Abstract:
Diffusion models have shown promise in drag‑style editing. Previous works mainly focus on point‑based drag, which is inherently ambiguous. This paper focuses on region‑based drag and introduces a novel In‑Context Region‑based Drag (ICRDrag) method. Under the in‑context learning framework, ICRDrag consumes a source image, a source region mask, and a target region mask, producing the target dragged image. Built upon the basic in‑context learning model, we introduce two novel attention regularization: 1) image‑mask attention consistency to ensure that a target region attends to similar source regions for image and mask modalities; 2) source‑target attention correspondence to ensure the mutual correspondence between source and target regions. To facilitate region‑based drag, we also construct Paired Region Dataset (PRD), a large‑scale dataset with paired masks and images. Extensive experiments show that ICRDrag significantly outperforms existing methods in both quantitative metrics and user studies, achieving superior editing accuracy and visual fidelity. The dataset, code, and model are available at https://github.com/bcmi/ICRDrag‑Region‑Drag‑Editing.
Authors:Yuchen Xie, Xinyu Zhou, Kuangji Zuo, Yanshuo Lu, Fengrui Huang, Boyu Ma, Jianfei Yang
Abstract:
Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments. However, prevailing EVT paradigms predominantly rely on language‑based target indication. While language is expressive and convenient, cluttered scenes often contain multiple objects that satisfy the same semantic description, leading to ambiguous target grounding. We therefore propose a paradigm shift, reframing target indication in EVT from text‑only specification to unified spatial‑semantic prompting. Based on this paradigm, we introduce Unified Spatial‑Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning, USS, an end‑to‑end embodied tracking framework that supports text, point, bounding box, and mask prompts within a unified architecture. USS encodes heterogeneous prompts with modality‑specific encoders, fuses prompt tokens with visual features through hybrid attention, and decodes compact prompt‑conditioned representations into egocentric waypoints. To further improve temporal robustness, USS incorporates a latent world model that predicts future representations through self‑supervised alignment. Real‑robot experiments demonstrate that explicit spatial target cues yield higher success rates than text‑only prompts, particularly in scenarios involving similar distractors and longer‑horizon tracking where maintaining instance‑level target identity is critical. In the simulation benchmark, USS also achieves state‑of‑the‑art performance among non‑MLLM‑based methods and competitive results against recent MLLM‑based approaches with faster inference speed. Our findings reveal that spatial‑semantic prompting provides a more precise and flexible target indication interface for embodied visual tracking. Project site: https://arescheah.github.io/uss‑project‑page/.
Authors:Khawar Islam, Arif Mahmood, Xin Jin, Naveed Akhtar
Abstract:
Data augmentation is known to improve generalization of deep visual models. Recent methods favor mixup strategies that generate interpolated samples to improve model performance. However, these techniques not only incur significant computational overhead, they also lead to semantic disruption of augmentation data due to cross‑sample mixing. We first propose Self‑Saliency (S^2) Mixup, which constructs challenging yet label‑consistent samples by extracting multi‑scale salient patches and reinserting them into non‑salient regions of the same image. This promotes scale‑invariant feature learning while avoiding cross‑sample interference. To further enhance model robustness, we introduce FracMix, a mixing scheme that injects self‑similarity patterns into salient regions using adaptive ratios. Collectively, our unified framework, S^2‑FracMix, enables simultaneous learning from fractal and non‑fractal structures within a single image, yielding a targeted and structurally coherent augmentation strategy. We theoretically analyze the advantage of our technique, and empirically establish its superiority over the existing methods by achieving state‑of‑the‑art performance in extensive evaluation with seven benchmarks across classification (coarse and fine‑grained), robustness, calibration, object detection, and transfer learning tasks. Project page is available at \hrefhttps://fracmix‑data‑augmentation.github.io/fracmix‑data‑augmentation.github.io
Authors:Muhammed Furkan Dasdelen, Fatih Ozlugedik, Anastasia Litinetskaya, Nassir Navab, Carsten Marr, Ario Sadafi
Abstract:
Data scarcity is a major bottleneck in medical Multiple Instance Learning (MIL), especially for rare diseases or expensive modalities. We introduce a statistically grounded patient augmentation approach that generates realistic patients directly in embedding space. Using Gaussian Mixture Models as a probabilistic clustering approach on pooled instance embeddings from all patients, our method learns disease‑specific "recipes"‑statistical distributions of instances across unsupervised clusters. New patients are then generated by sampling embeddings from clusters based on learned recipes. Unlike existing methods that require examples from all categories, our method can generate patients offline by re‑mixing pooled embeddings. Generated patients are further selected based on uncertainty quantification to improve MIL performance. We evaluate our method across three clinically relevant scarcity scenarios: (i) cross‑dataset transfer, where an entirely missing "healthy" class is generated using statistics from an external cohort; (ii) low‑data regimes, where class sizes are extremely limited; and (iii) small‑cohort non‑image tasks, including single‑cell RNA‑seq and flow cytometry. Across all experiments, our method improves performance over baseline, often outperforming other bag‑mixing strategies. Notably, in the missing‑class scenario, a performance comparable to full‑dataset training is achieved, demonstrating its potential for rare disease diagnostic and privacy‑preserving patient augmentation. The code is available at https://github.com/marrlab/RECIPE
Authors:Jiayu Li, Yixiao Fang, Tianyu Hu, Wei Cheng, Ping Huang, Zheheng Fan, Gang Yu, Xingjun Ma
Abstract:
Real‑world photography requires capture‑time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evaluate post‑hoc crop prediction and overlook subject‑side recommendations, leaving the capture‑time guidance capabilities of multimodal large language models (MLLMs) underexplored. To address this gap, we introduce CaptureGuide‑Bench, a benchmark with two complementary tasks: photographer‑side composition decision and refinement, and subject‑side scene‑conditioned pose recommendation. Our evaluation reveals limitations: general‑purpose MLLMs can make composition decisions but lack precise refinement localization, while specialized aesthetic cropping models localize crops effectively but are limited to refinement; neither provides actionable pose guidance. To support model development, we further construct CaptureGuide‑Dataset, comprising 130K samples with textual rationales and structured visual annotations, and develop ShutterMuse, a unified MLLM trained with supervised and reinforcement fine‑tuning. Experiments on CaptureGuide‑Bench show that ShutterMuse achieves the best overall photographer‑side performance among evaluated baselines and competitive subject‑side pose recommendation with substantially lower inference cost, demonstrating the potential of MLLMs as interactive assistants for photography during image capture.
Authors:Wenjie Zhu, Yabin Zhang, Liang Xu, Xin Jin, Wenjun Zeng, Lei Zhang
Abstract:
While test‑time adaptation (TTA) empowers vision‑language models to adapt without costly retraining, it remains highly vulnerable to out‑of‑distribution (OOD) outliers prevalent in real‑world applications. This discrepancy motivates Noisy TTA (NTTA), an online task to filter noisy OOD samples on the fly while maximizing in‑distribution (ID) classification accuracy. Existing zero‑shot NTTA approaches typically rely on test‑time discriminative training, leading to overconfident misclassifications and significantly degraded inference efficiency. To address these limitations, we propose a novel framework named Dual Distribution Estimation (DDE), shifting the zero‑shot NTTA paradigm from instance‑level learning to training‑free Gaussian distribution modeling. DDE incorporates two novel modules: Positive Feature Distribution Estimation (PFDE) and Negative Label Distribution Estimation (NLDE). PFDE explicitly models class‑wise inclusion and exclusion Gaussian distributions to formulate a calibrated contrastive score, robustly enhancing ID accuracy. In parallel, NLDE improves OOD identification by explicitly modeling the negative label distribution to mine highly discriminative labels, effectively mitigating spurious correlations. Extensive experiments show that on the large‑scale ImageNet benchmark, DDE achieves an improvement of 3.70% in harmonic mean accuracy and reduces the FPR95 for OOD detection by 6.20%, while ensuring highly scalable and efficient online inference. Furthermore, DDE is zero‑shot and training‑free, demonstrating remarkable robustness in data‑scarce scenarios. Codes are available at https://github.com/ZhuWenjie98/DDE.
Authors:Yifei Qu, Ru Li, Junjie Chen, Jinyuan Wu
Abstract:
Real‑world single image dehazing is highly ill‑posed due to spatially and spectrally varying scattering, while practical deployment demands lightweight and low‑latency models. Existing approaches either rely on fragile physical inversion under simplified assumptions or adopt heavy blind architectures unsuitable for edge deployment. To overcome these limitations, we propose PGL‑Net (Physics‑Inspired Global‑Local Decoupling Network), a lightweight framework that incorporates physical inductive biases via operator‑level emulation, avoiding explicit parameter estimation. It decouples dehazing into global distribution rectification and local structural refinement. A Physics‑Inspired Affine Fusion (PAF) module performs globally conditioned alignment across hierarchical skip connections to compensate for haze‑induced bias, while a compact Degradation‑Aware Modulation (DAM) block adaptively restores spatially and spectrally variant details through dynamic feature modulation. Extensive experiments on multiple real‑world benchmarks demonstrate that PGL‑Net achieves state‑of‑the‑art restoration quality with significantly reduced complexity. Compared with the recent SOTA SGDN, the Tiny variant (PGL‑Net‑T) improves PSNR by up to 2.6dB and consistently enhances downstream object detection accuracy, while achieving over a 10x reduction in inference latency. Code is publicly available at: https://github.com/sc‑30‑bit/PGL‑Net.
Authors:Baiyang Song, Yuli Lin, Qiong Wu, Tao Chen, Jun Peng, Xiao Chen, Yiyi Zhou, Rongrong Ji
Abstract:
Currently, streaming video understanding is still a daunting task for existing \emphmultimodal large language models (MLLMs). Its difficulties not only lie in handling the ever‑increasing video frames, but also in the unpredictability of future video content and input instructions. In this paper, we study this task from the perspective of constructing a dynamic but fixed‑budget memory bank, and propose a novel and training‑free approach termed \emphCausalMem. CausalMem is dedicated to constructing a dynamic visual memory update mechanism, thereby maximizing the amount of information in streaming video within a limited memory space, much like the human brain. In practice, CausalMem estimates the redundancy of visual tokens and updates the memory bank via an online semantic basis, which models the principal semantics of the observed video stream. To validate CausalMem, we apply it to two representative MLLMs, namely LLaVA‑OneVision and Qwen2.5‑VL respectively, and conduct extensive experiments on both streaming and offline video understanding benchmarks. The experimental results not only show the great advantages than existing methods under both streaming and offline settings, \emphe.g., +3.2% and +3.0% average accuracy gains respectively, but also witness the superior semantic preservation for streaming videos, \emphe.g., using 12k token budgets to memorize hour‑long streaming videos, which achieves more than 20× visual token compression ratio and only occupies about 82 MB storage. Our code is given in \hrefhttps://github.com/hktk07/CausalMemCausalMem.
Authors:Tianchen Guo, Chen Liu, Ling Chen, Xin Yu
Abstract:
Multimodal Large Language Models (MLLMs) have shown remarkable progress in single‑image perception, yet their ability to reason about complex cross‑view human‑centric scenes remains largely unverified. Current multi‑view benchmarks evaluate models using a fixed "bag of frames" and thus conflate a model's robustness to visual distraction with its genuine ability to fuse fragmented cross‑view evidence. To address this issue, we introduce SSMNBench, a diagnostic benchmark comprising 3,300 curated QA pairs for cross‑view human and human‑object understanding. SSMNBench uniquely categorizes tasks into Single‑View Sufficiency (SVS) and Multi‑View Necessity (MVN). By systematically perturbing view availability across 17 state‑of‑the‑art MLLMs, critical limitations are revealed: models suffer from severe "distraction degradation" when presented with redundant views (SVS), and fail to integrate fragmented geometric evidence across cameras (MVN). Our evaluations demonstrate that modern MLLMs rely on multiple single‑image semantic averaging and view preference rather than genuine cross‑view synthesis. By exposing these fundamental vulnerabilities, SSMNBench provides a rigorous diagnostic framework to drive the advancement of future cross‑view‑aware multimodal architectures. The code is available at: \hrefhttps://github.com/gtc‑gh/SSMNBench\textSSMNBench
Authors:Felipe Moreno, Sharifa Alghowinem, Hae Won Park, Cynthia Breazeal
Abstract:
Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for early diagnosis and intervention. As a multidisciplinary topic, these objective measures should be interpretable and accessible to health care professionals, ensuring effective collaboration and treatment planning in the realm of mental health care. Even though current automated depression diagnosis approaches improved over the last decade, a critical gap exists as they often lack affect‑specificity and interpretability, limiting their practical application and potential impact on mental health care. In particular, interpretability from temporal activities from videos when deep models are used is not fully explored. In this study, we present a novel framework for analyzing Deep Neural Networks' decisions when trained on facial videos, specifically focusing on automatic depression severity diagnosis. By fine‑tuning Deep Convolutional Neural Networks (DCNN) pre‑trained on Action Recognition datasets on depression severity facial videos from AVEC depression dataset, our framework is able to interpret the model's saliency maps by examining face regions and temporal expression semantics. Our approach generates both visual and quantitative explanations for the model's decisions, providing greater insight into its reasoning. In addition to this interpretability, our video‑based modeling has improved upon previous single‑face benchmarks for visual depression diagnosis, resulting in enhanced predictive performance. Overall, our work demonstrates the successful development of a framework capable of generating hypotheses from a facial model's decisions while simultaneously improving depression's predictive capabilities.
Authors:Seulgi Jeong, Yunseong Cho, Sanghun Park
Abstract:
Hairstyle transfer has practical applications such as virtual try‑on, yet remains challenging when the source and reference exhibit large head‑pose discrepancies. We propose H‑Adapter, which improves pose robustness by training with a region‑specific loss that disentangles hair and non‑hair objectives and thereby induces spatially disentangled cross‑attention, from which a source‑aligned hair edit mask is derived to guide diffusion‑based inpainting. Experiments on pose‑agnostic and pose‑different subsets demonstrate strong quantitative results, including the best FID, \mathrmFID_\mathrmCLIP, and CLIP‑I under pose differences, while maintaining competitive non‑hair preservation and improving qualitative fidelity to fine‑grained reference hairstyle details. Beyond source‑conditioned transfer, H‑Adapter supports practical extensions including text‑to‑image generation, auxiliary prompt‑based hair color control, and compatibility with an identity‑preserving IP‑Adapter variant. We also introduce a VLM‑as‑a‑judge protocol and observe consistent gains in hairstyle faithfulness, non‑hair preservation, and artifact quality.
Authors:Sining Ang, Yuan Chen, Liu Haiyan, Xuanyao Mao, Jason Bao, Xuliang, Bingchuan Sun, Yan Wang
Abstract:
Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast‑slow planners often rely on hand‑designed triggering rules that either over‑call the slow system or call it at the wrong times. We formulate slow‑system invocation as a resource‑aware sequential decision problem and propose the Adaptive Slow‑System Control Gate (ASSCG), which makes frame‑level Query/Cache/Drop decisions to refresh, reuse, or suppress slow guidance. ASSCG uses an RWKV backbone for efficient long‑horizon gating and is trained with supervised fine‑tuning followed by GRPO‑style compute‑aware reinforcement fine‑tuning. We apply ASSCG to two different fast‑slow architectures: (i) AsyncDriver on nuPlan Hard20 closed‑loop evaluation, where ASSCG improves score to 67.28 (+2.28) while reducing average end‑to‑end inference latency by 60%; and (ii) a RecogDrive‑based dual system that we build by replacing its original VLM‑2B module with a lightweight ViT‑based fast planner and adding an LLM slow planner, evaluated on NAVSIM, where ASSCG achieves 91.4 PDMS (+0.6) and increases average speed by 25%. The project page, including video visualizations and additional results, is available at https://williamxuanyu.github.io/asscg/.
Authors:Hualong Zhang, Siyang Feng, Zihan Huan, Yi Qian, Zhenbing Liu, Rushi Lan, Xipeng Pan
Abstract:
Histopathological tissue segmentation is essential for computer‑aided diagnosis, yet weakly supervised methods often suffer from noisy pseudo‑labels generated by Class Activation Mapping (CAM). Existing CAM approaches tend to focus on staining‑driven appearance cues rather than true causal tissue morphology, resulting in spurious localization and poor structural consistency. To address this issue, we propose C^2RM‑Seg, a two‑stage framework that integrates causal pseudo‑label refinement with structure‑aware semantic enhancement. For classification, we introduce a Causal Counterfactual Reasoning Module (C^2RM) that decomposes features into latent factors and performs counterfactual intervention via a learned causal structure matrix, suppressing confounding context and producing morphology‑aligned CAMs. For segmentation, we design a Dual‑Path Structural‑Semantic Architecture that combines fine‑grained structural features from ResNeSt with global semantic priors from a frozen DINOV3 foundation model. A cross‑path gating mechanism adaptively regulates semantic injection using local structural cues to preserve boundary fidelity. To further mitigate residual pseudo‑label noise, we propose an Uncertainty‑Gated Margin (UGM) loss, which dynamically balances margin enforcement and confidence learning based on prediction uncertainty. Extensive experiments on two public histopathological tissue datasets show that C^2RM‑Seg achieves state‑of‑the‑art performance.
Authors:Salman Shaik, Truong Thanh Hung Nguyen, Hung Cao
Abstract:
Accurate classification of diffuse gliomas is often hindered by domain shifts across centers and a lack of large, annotated datasets. We propose the Anatomically‑conditioned Latent Diffusion Model (ALDM), a novel framework for data‑efficient, few‑shot 3D volumetric MRI synthesis. ALDM utilizes a two‑stage approach: a 3D variational autoencoder learns anatomical priors from a data‑rich source domain, while a conditional latent diffusion model, guided by tumor masks via a ControlNet, generates structurally coherent volumes for a data‑scarce target domain. Evaluated in an extreme few‑shot setting with only 16 target images, ALDM outperformed GAN and hybrid baselines, achieving a superior Frechet Inception Distance (FID) of 85.40 and a downstream classification AUC of 0.987. Qualitative results confirm that the model preserves sharp pathology boundaries and cross‑modal consistency, with visual fidelity improving progressively during training. By capturing essential diagnostic features, ALDM provides a robust tool for clinical data augmentation in low‑resource settings. Our implementation is available at https://github.com/Analytics‑Everywhere‑Lab/anatomically‑conditioned‑LDM.
Authors:Hongye Xu, Bartosz Krawczyk
Abstract:
Exemplar‑free class‑incremental learning (EFCIL) requires stable decision boundaries within a shifting feature space. While maintaining class‑conditional Gaussian statistics provides a principled classification strategy, these parametric summaries remain sensitive to anisotropic representation drift. Existing methods often transport these statistics across tasks using a decoupled, post‑hoc paradigm: optimizing a backbone without explicit geometric constraints can distort the legacy manifold, limiting the precision of retroactive alignment. In this paper, we formulate feature transport as an endogenous training constraint rather than a separate post‑task step, presenting the Geometry‑Anchored Transport Framework. First, we derive an Analytic Geometric Anchor via Mahalanobis‑aligned regression to mitigate macroscopic anisotropic drift. Second, we introduce a Topology‑Aware Evolution objective that regularizes localized manifold degradation while calibrating a residual network against the analytic prior. By coupling manifold evolution with transport constraints during the primary training phase, our framework mitigates evaluation errors without requiring decoupled fine‑tuning. Experiments across CIFAR‑100, TinyImageNet, and ImageNet‑100 demonstrate that the proposed framework consistently improves upon existing post‑hoc alternatives under strict exemplar‑free constraints.
Authors:Qinzhe Yang, Chenyang Liu, Jia Xu, Zhenwei Shi, Zhengxia Zou
Abstract:
State Space Models (SSMs), designed for long‑range modeling, offer linear computational complexity and strong capabilities in capturing long‑range dependencies. In the field of remote sensing, SSMs have gained popularity due to their effectiveness in addressing unique challenges such as dense visual predictions, multi‑modal remote sensing data, and temporal remote sensing data, which have also yielded significant advancements in customized architectures. This paper presents a comprehensive review of SSM‑based approaches in remote sensing, covering most of the relevant studies since SSMs were first introduced to the field. We offer a multi‑dimensional analysis examining SSM applications in remote sensing tasks and discussing advancements in architecture design. This paper not only synthesizes the rapid progress in SSM‑based research but also identifies key challenges and future opportunities. By providing a detailed perspective, this paper aims to serve as a foundational resource for remote sensing researchers, offering actionable insights to foster further advancements in this evolving domain. We will keep tracing related works at https://github.com/QinzheYang/Awesome‑RS‑State‑Space‑Model.
Authors:Qinzhe Yang, Keyan Chen, Jia Xu, Zhenwei Shi, Zhengxia Zou
Abstract:
The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision models, particularly recent ViT‑based foundation models in dense prediction tasks. Instance segmentation, a typical dense visual prediction task in the remote sensing field, faces similar challenges. In this paper, inspired by the recent advances of knowledge distillation in large language models, we introduce RS4D ‑ a new remote sensing instance segmentation method with linear computational complexity, which addresses the inefficiency of long sequence modeling through distilled state space modeling (SSM). We propose an adaptive noise and masking knowledge distillation training method for pre‑training lightweight SSM backbones, which effectively compresses knowledge from the vast self‑attention space into a compact, dense linear state space. We also design a remote sensing image instance segmentation architecture based on this lightweight visual encoder, where we explore variants of three different backbones and two segmentation heads. Extensive experiments are conducted on multiple benchmark datasets, including SSDD, WHU, and NWPU. Compared to ViT‑based approaches, our proposed SSM backbone achieves an 8x reduction in parameters and a 9x reduction in FLOPs while maintaining comparable or superior accuracy to both ViT‑ and CNN‑based instance segmentation methods. The implementation codes have been publicly available at https://github.com/QinzheYang/RS4D.
Authors:Haoxiang Sun, Zhihang Yi, Langxuan Deng, Yuhao Zhou, Peiqi Jia, Jian Zhao, Li Yuan, Jiancheng Lv, Tao Wang
Abstract:
Fine‑grained visual reasoning requires multimodal large language models (MLLMs) to identify task‑relevant visual evidence and ground their reasoning in local image regions. Existing agentic methods typically rely on reinforcement learning with verifiable rewards or supervised fine‑tuning on large‑scale annotated reasoning traces, leading to costly exploration, hand‑designed verification rules, or heavy dependence on textual supervision. A natural way to avoid such external answer labels is to learn from trajectories sampled by the student itself, which points to On‑Policy Distillation (OPD). To understand what OPD can and cannot provide for visual reasoning, we revisit it as negative‑free stop‑gradient alignment. This perspective shows that, although OPD provides effective token‑level correction, its ceiling is constrained by the absence of trajectory‑level discrimination. Motivated by these observations, we propose V‑Zero, an answer‑label‑free framework for visual reasoning with contrastive evidence gating. V‑Zero uses no annotated textual answer labels; instead, during training it pairs a question‑relevant regional crop with a negative visual view to evaluate student‑sampled trajectories and gate dense token‑level distillation. Experiments on multiple visual reasoning benchmarks show that V‑Zero consistently improves fine‑grained visual reasoning while preserving strong generalization. Notably, V‑Zero is more than 5× faster than previous supervised fine‑tuning methods and more than 10× faster than reinforcement learning baselines. Code and dataset will be released at https://github.com/eVI‑group‑SCU/V‑Zero
Authors:Qinzhe Yang, Dongyu Wang, Haohan Niu, Jia Xu, Zhenwei Shi, Zhengxia Zou
Abstract:
Remote sensing object detection has advanced rapidly with the development of large‑scale benchmarks and modern detection architectures. However, existing datasets and detectors remain fragmented. Most benchmarks focus on limited categories, fixed spatial resolutions, or a single sensor, while detectors still struggle to work across different sensors and categorical systems. In this paper, we introduce LEVIRDet‑159, the largest and most comprehensive remote sensing object detection dataset to date, with 159 categories, 2.56 million bounding boxes, and 700k fine‑grained annotations under a multi‑level taxonomy. In each key scale dimension, LEVIRDet‑159 exceeds the corresponding largest existing remote sensing object detection dataset, containing approximately (7x) more images, (6x) more object instances, and (4x) more categories. Based on this dataset, we design LEVIRDetNet, a scale‑hierarchy‑aware detection foundation model for universal remote sensing object detection. LEVIRDetNet couples online visual Ground Sampling Distance (GSD) prediction, GSD‑conditioned query modulation and allocation, and a hierarchy‑aware detection head for mixed‑granularity remote sensing supervision. Under stringent evaluation settings, LEVIRDetNet demonstrates strong cross‑domain generalization. Even without target‑domain training or fine‑tuning, it achieves state‑of‑the‑art detection performance on 9 external benchmarks, improving the strongest fully supervised competing methods by 5.02 mAP on average under each benchmark's primary metric. We hope this study will facilitate the development of strongly generalizable remote sensing object detection across diverse category systems, spatial resolutions, and sensor platforms. The dataset and trained models will be released at https://qinzheyang.github.io/LEVIRDet/, accompanying the final paper.
Authors:Atin Pothiraj, Jaemin Cho, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal
Abstract:
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question‑based evaluation pipeline. PQSG evaluates generated videos by checking their faithfulness to a prompt across objects, actions, and adherence to physical laws using a graph‑based hierarchy of questions generated by a vision‑language model (VLM), guided by high‑quality in‑context examples. By representing questions as a graph, PQSG introduces logical dependencies within questions, ensuring that each query is contextually valid. Moreover, PQSG provides granular assessments of which qualities of the video violate physical plausibility constraints. We validate PQSG by creating FinePhyEval, a dataset with physics‑based prompts and corresponding generated videos from diverse state‑of‑the‑art video generation models (Sora 2, Veo 3, and Wan 2.1), with each video annotated across multiple categories by humans. Using FinePhyEval, we measure the correlation between PQSG's fine‑grained scores and human judgments, showing higher overall correlations than prior work. We also find that PQSG ranks closed‑source models higher than Wan 2.1 on physical realism. Lastly, we show that the annotations we provide in FinePhyEval can also be used for subtask evaluation: we benchmark two strong VLMs on generating and answering questions, finding that while models can create human‑like questions, they still fall short of human performance in answering them.
Authors:Hoai-Danh Vo, Trung-Nghia Le
Abstract:
Generative models have significantly advanced image generation, resulting in synthesized images that are increasingly indistinguishable from authentic ones. However, the creation of fake images with malicious intent is a growing concern. Low‑configured smart devices have become highly popular, making it easier for deceptive images to reach users. Consequently, the demand for effective detection methods is increasingly urgent. In this paper, we introduce a simple yet efficient method that captures pixel fluctuations between neighboring pixels by calculating the gradient, which highlights variations in grayscale intensity. This approach functions as a high‑pass filter, emphasizing key features for accurate image distinction while minimizing color influence. Our experiments on multiple datasets demonstrate that our method achieves accuracy levels comparable to state‑of‑the‑art techniques while requiring minimal computational resources. Therefore, it is suitable for deployment on low‑end devices such as smartphones. The code is available at https://github.com/vohoaidanh/adof.
Authors:Ke Xu, Xinle Wang, Yanning Hou, Xueliang Ma, Juan Xie, Jianfeng Qiu
Abstract:
Zero‑shot 3D anomaly detection is essential for industrial quality inspection, where labeled anomaly samples are scarce. Meanwhile, existing methods lack an effective mechanism to fuse complementary 2D color images with 3D geometric structures, limiting their ability to detect both surface and structural defects in a unified framework. To address these issues, we propose CoGeoAD, a unified CLIP‑based framework that fuses color and geometric features by constructing pixel‑aligned paired multi‑view images. The framework introduces a Data‑Driven Multi‑View Attention (MVA) mechanism to adaptively aggregate 3D features and a Multi‑Stage Color‑Geometric Fusion (MS‑CGF) module to hierarchically integrate multi‑level features from both modalities. Extensive experiments on the MVTec3D‑AD and Eyecandies benchmarks demonstrate that CoGeoAD achieves state‑of‑the‑art performance, effectively capturing both structural and textural anomalies in complex industrial scenarios. our source code is available at https://github.com/kingdomShu/CoGeoAD.
Authors:Qing Lian, Kent Yu, Lei Zhang
Abstract:
Most vision‑language‑action (VLA) models are reactive: they predict the next action from the current instruction and observation, implicitly assuming that the current observation fully specifies the action‑relevant state. In embodied control, however, embodiment‑specific factors such as camera‑to‑robot geometry, robot calibration, or systematic actuation bias are often hard to identify from a single observation. As a result, reactive policies cannot reliably disambiguate these factors in general, overfitting to training environments and generalizing poorly at deployment. We propose Reflective VLA, which conditions each decision on a context of observation‑action‑consequence triplets. Each triplet records not only what the robot observed and executed, but also how the scene changed afterward, exposing the deployment‑specific mapping from actions to observed effects. Architecturally, Reflective VLA routes all observation modalities through the VLM under shared attention, so the action expert reasons directly over past triplets and the current observation. A block‑causal mask enables parallel multi‑frame training without leakage and supports KV‑cached real‑time inference. On standard LIBERO and SimplerEnv‑Bridge, Reflective VLA preserves strong in‑distribution performance. Under distribution shift on LIBERO‑Plus and the harder LIBERO‑Plus‑Hard, it improves average success rate by 5.4 and 4.2 percentage points over a matched reactive baseline. Ablations with a matched history‑only baseline further show that action consequences ‑‑ rather than additional context length alone ‑‑ are the key to cross‑environment generalization. Project page: https://lianqing11.github.io/reflective‑vla‑page/
Authors:John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
Abstract:
Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning. Taking one of the longest continuously used writing systems, ancient Greek, as a case study, we introduce three datasets for diachronic representation learning: Hell‑Char, a curated training set spanning the 3rd‑1st centuries BCE, and two evaluation sets, PaLit‑Char (2nd‑5th c. CE) and Med‑Char (9th‑14th c. CE). To address the challenges of symbolic variation, scarce data, and systematic degradation, we propose: a similarity‑weighted supervised contrastive loss that biases embeddings using dynamically estimated inter‑class similarities, and a lacuna‑driven augmentation scheme that simulates realistic manuscript corruptions. Trained with these strategies, both a lightweight CNN and a pretrained ResNet achieve strong recognition performance and produce embeddings that more coherently separate character classes than PCA or generic pretrained models. These embeddings enable clustering, identification of stylistic subgroups, and construction of prototype images that visualize diachronic evolution and transitional letterforms. Our results demonstrate that respecting intrinsic inter‑letter relationships and augmenting with domain‑informed corruptions yield robust, interpretable representations, offering a transferable paradigm for representation learning under scarce, temporally evolving, and noisy conditions. Code and data available at: https://github.com/ipavlopoulos/diachronic‑greek‑letterforms.
Authors:Leshu Li, Jie Peng, Yang Zhao
Abstract:
3D Gaussian Splatting (3DGS) has garnered significant attention in Simultaneous Localization and Mapping (SLAM) due to its advances in capturing fine‑grained geometry features and synthesizing novel views. For SLAM in large‑scale scenes, such as autonomous driving, 3DGS‑SLAM faces a critical limitation: memory consumption increases continuously over time as Gaussian points accumulate, leading to poor memory efficiency and limiting its applicability. In this work, we propose a rendering‑area‑aware pruning strategy that selectively removes Gaussians based on their contribution to the effective rendering area, rather than solely relying on Gaussian‑level heuristics such as opacity or gradient magnitude. This perspective directly targets the sources of memory redundancy, effectively reducing the peak memory footprint of 3DGS‑SLAM during runtime. Evaluations on the EuRoC and KITTI datasets demonstrate that our method consistently outperforms existing pruning approaches in large‑scale outdoor scenes, achieving over 60% memory reduction and more than 2 times FPS improvement while preserving localization and mapping accuracy. These results highlight rendering‑area‑aware pruning as a promising direction for scaling 3DGS‑SLAM to real‑world autonomous driving scenarios. Our code is publicly available at https://github.com/UMN‑ZhaoLab/Pocket‑SLAM.git.
Authors:Dimitri Gominski, Maurice Mugabowindekwe, Qiue Xu, Xiaowei Tong, Martin Brandt, Hieu Le, Rasmus Fensholt, Dimitris Samaras, Loic Landrieu
Abstract:
Counting individual trees is a fundamental task for environmental monitoring, yet remains largely unexplored with satellite imagery. At these resolutions, isolated trees may still be identifiable, but crown boundaries become ambiguous in dense forests, making the notion of an individual tree inherently ill‑defined. Moreover, large‑scale manual annotations of individual trees are prohibitively expensive. While scalable supervision can be derived from airborne LiDAR, the resulting annotations are noisy and difficult to exploit effectively. We address these challenges by formulating tree counting as a spatial density matching problem supervised through Unbalanced Optimal Transport. This formulation naturally accommodates both precise localization of isolate trees and robust density estimation in dense forests. We further introduce a self‑correction mechanism that leverages transport residuals to progressively refine noisy supervision during training. We evaluate our approach on TinyTrees, a new benchmark spanning three continents and three satellite sensors, comprising over 216 million tree annotations (including 639k manually verified instances) across 25\,890 km^2. Our method consistently outperforms detection‑based, regression‑based, and transport‑based distribution‑matching baselines, demonstrating the effectiveness of unbalanced transport and reliability‑aware supervision for large‑scale tree counting from satellite imagery. Code, data and models are available at https://github.com/dgominski/treematch.
Authors:Daniel Lengerer, Mathias Pechinger, Klaus Bogenberger, Carsten Markgraf
Abstract:
High‑resolution aerial imagery has recently emerged as a complementary modality for automated driving perception and has shown potential to improve birds‑eye‑view (BEV) scene understanding when fused with onboard sensors. Prior work demonstrated performance gains for online high‑definition (HD) map construction through aerial‑onboard fusion; however, conventional end‑to‑end fusion does not fully exploit the structural information contained in aerial representations. In this work, we introduce AerialFusionMapNet, a fusion‑based mapping framework with a structured two‑stage training strategy that explicitly enhances the contribution of aerial features within a unified pipeline. The proposed training scheme enables more effective integration of structural aerial priors. On the nuScenes geographic split, AerialFusionMapNet achieves up to 54.7 mAP, improving over prior aerial‑onboard fusion baselines from 48.8 mAP by +5.9 absolute and +12.1% relative. The results suggest that structured training design, rather than increased architectural complexity, plays a more decisive role in unlocking the full potential of aerial imagery for online HD map construction. Code and trained models are available at https://github.com/DriverlessMobility/AerialFusionMapNet.
Authors:Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
Abstract:
Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade‑off between temporal reasoning and spatial precision. Models that rely on single‑frame or low‑resolution inputs often miss small, distant, or partially occluded hazards, while language‑centric driving models frequently provide limited grounded evidence for their explanations. To address this gap, we propose UniDrive, a unified visual‑language and grounding framework for interpretable risk understanding in autonomous driving. UniDrive combines a temporal reasoning branch that models scene dynamics from multi‑frame visual input with a high‑resolution perception branch that preserves fine‑grained spatial details from the latest frame. The two branches are integrated through a gated cross‑attention fusion module, enabling dynamic context to be aligned with precise spatial evidence. Based on the fused representation, UniDrive jointly generates natural‑language risk descriptions and grounded bounding‑box outputs for risk objects. Experiments on the DRAMA‑Reasoning benchmark show that UniDrive outperforms representative image‑based and video‑based baselines in both captioning and risk‑object grounding. In particular, UniDrive achieves the best overall performance on the validation split and demonstrates clear advantages in small‑object localization, zero‑shot generalization to NuScenes and BDD100K, and human‑rated interpretability and trustworthiness. These results suggest that explicitly combining temporal semantics and high‑resolution perception provides a stronger foundation for interpretable and safety‑oriented autonomous driving systems. The code is available at https://github.com/pixeli99/unidrive‑dev.
Authors:Jiaxiang Liu, Tianxiang Hu, Juwei Guan, Yujie Wu, Yusong Wang, Yao Mu, Zuozhu Liu, Mingkun Xu
Abstract:
Recent advances in vision‑language models (VLMs) such as CLIP have demonstrated strong generalization across natural‑image domains. However, adapting these models to biomedical imaging is non‑trivial: full‑model fine‑tuning is computationally expensive, while medical data are often scarce and exhibit subtle, fine‑grained inter‑class differences, making parameter‑efficient adaptation particularly critical. Visual Reprogramming (VR) offers a parameter‑efficient alternative by injecting learnable perturbations into the input space, but existing VR approaches for VLMs mainly focus on positive class prompts and overlook confusing negatives, leading to miscalibrated predictions in fine‑grained medical scenarios. We present BioMedVR, the first VR‑based framework for biomedical imaging, enabling few‑shot adaptation of pretrained VLMs through compact learnable VR modules. To mitigate class confusion, we introduce a Confusion Minimization Mechanism that leverages LLM‑generated confusion‑aware attributes together with a Confusion‑Suppression Loss to explicitly reduce false‑positive alignment. Moreover, the designed Mixture‑of‑Prompt Experts combines a positive expert for main‑class discrimination and a negative expert for confusion suppression, balanced via adaptive gating. Extensive experiments on 18 datasets, including 11 biomedical datasets and 7 natural image benchmarks, demonstrate that BioMedVR achieves superior accuracy and generalization, effectively bridging VR and VLMs in biomedical domains.
Authors:Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Begüm Demir
Abstract:
Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence. We present a human‑grounded evaluation framework that quantifies alignment between SAE latents and human‑annotated concepts, without requiring user studies, and validate this matching through targeted attribute perturbations. To enable this intervention‑style evaluation in vision, we construct synCUB and synCOCO, synthetic benchmarks of paired images that differ in exactly one attribute. We introduce Fully‑Binary Matching Pursuit (FBMP), a coalition‑based matching procedure that supports many‑to‑one mappings between SAE latents and annotated concepts, and consistently outperforms one‑to‑one baselines. For functional validation, we propose a Targeted Attribute Perturbation Alignment Score (TAPAScore), which tests whether matched concepts respond selectively and in the expected direction under targeted image‑level attribute perturbations. Under sanity checks, our matching and TAPAScore are the only evaluated metrics that reliably distinguish trained SAEs from untrained ones. Across SAEs trained on CLIP and DINOv2 embeddings, we find that increased overcompleteness can reduce perturbation alignment, indicating a reduction in interpretability. Our evaluation framework suggests that moderate dictionary sizes provide the best trade‑off, yielding the most interpretable SAEs. Code and datasets are available at https://github.com/JonasKlotz/sae‑concept‑eval.
Authors:Wenxin Wang, Bo Zhang, Feng Chen, Zixuan Wang, Wen Li, Changsheng Li, Yinjie Lei
Abstract:
Recent advancements have explored agentic zero‑shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language Models (MLLMs). However, existing methods face an intrinsic bottleneck due to the finite observation perspectives inherent in videos and the implicit perception of 3D scenes. In this paper, we propose a collaborative multi‑agent framework that assigns a Planning Agent to handle high‑level viewpoint planning and supplement novel perspectives, and a Perception Agent to explicitly summarize the 3D scene into a structured holistic cognitive map. Specifically, Planning Agent first analyzes this cognitive map to determine query‑relevant viewpoints and supplements missing critical perspectives to ensure comprehensive observation. Subsequently, Perception Agent documents object‑level attributes from these views by assigning consistent instance identifiers across viewpoints, thereby integrating fragmented observations into the holistic cognitive map. In parallel, it provides feedback to filter out mismatched candidate objects and guide subsequent viewpoint planning. Through this closed‑loop iterative process, two agents collaboratively figure out candidates until Perception Agent determines that sufficient information has been captured to complete the task. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance on 6 benchmarks, with improvements of 11.1% Acc@0.5 on ScanRefer, 14.6 BLEU‑1 on 3D‑assisted dialog, and 2.1 EM on SQA3D.
Authors:Zhenyang Li, Lutao Jiang, Yizhou Zhao, Ying-Cong Chen, Xin Wang, Weikai Chen, Yifan Peng
Abstract:
Reconstructing realistic, physically plausible garments from a single image remains a fundamental challenge. Template‑free methods capture surface geometry but lack explicit sewing structure for simulation; while programmatic systems are simulation‑ready but constrained by predefined templates. This reveals a fundamental representation gap between geometric reconstruction and structured garment construction. We present PatternGSL, a structured garment representation in the form of a template‑free and learnable specification language that encodes complete sewing patterns, including panel boundaries, parameterized seams, and explicit stitch topology, in a compact and standardized form. PatternGSL preserves the physical rigor of pattern‑based models while removing template dependence, elevating sewing structure as a first‑class target for generative modeling. We further propose a vision‑language framework that predicts PatternGSL specifications directly from a single image and decodes them into garments using lightweight deterministic validity handling, without optimization‑based refinement or manual cleanup. In addition, we introduce PatternGSLData, the first large‑scale image‑to‑GSL paired dataset comprising 300K samples with complete sewing pattern annotations, enabling supervised VLM training for structured garment reconstruction. Experiments demonstrate improved pattern accuracy over prior baselines, explicit sewing‑structure recovery, reliable cloth simulation, and pattern‑level editing through the same deterministic decoding pipeline. Code and data‑processing scripts will be released at https://lagrangeli.github.io/PatternGSL/.
Authors:Jiayi Lei, Yuandong Pu, Xingyu Han, Rongpeng Zhu, Jing Xu, Jinyao Wang, Zijian Zhou, Bin Fu, Yuewen Cao, Yihao Liu, Hongsheng Li
Abstract:
Text‑to‑image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their success reflects genuine causal understanding or sophisticated pattern matching over visual‑textual correlations. Inspired by Russell's inductivist turkey, we introduce Counterfactual‑World (CF‑World), a counterfactual benchmark designed to investigate whether text‑to‑image models can generate images under rules that systematically contradict real‑world priors. CF‑World organizes each scenario into three progressive levels: factual generation under ordinary world knowledge, explicit counterfactual generation with direct visual instructions, and implicit counterfactual generation requiring causal deduction from altered rules. We evaluate both open‑source and closed‑source T2I models using a Vision Language Model (VLM)‑based evaluator (CF‑Eval). Furthermore, we introduce two metrics: Prior Resistance Rate (PRR), which measures a models' ability to overcome entrenched real‑world priors, and Reasoning Retention Rate (RRR), which assesses whether models can maintain reasoning‑dependent counterfactual generation without explicit visual cues. Experiments show that all models exhibit sharp degradation from factual to counterfactual settings. Further analyses suggest that these failures arise because current T2I models encode world knowledge and visual appearances as tightly coupled patterns. Consequently, their heavy reliance on frequent visual co‑occurrences within the training data forces them to default to familiar commonsense priors when tasked with rendering counterfactual worlds.
Authors:Ling Li, Bowen Liu, Zinuo Zhan, Jianhui Zhong, Ziyu Zhu, Bingcai Wei, Kenglun Chang, Zhidong Deng
Abstract:
Pointing‑based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Traditional methods typically encode input images into static feature representations and perform reasoning primarily within the linguistic domain, often overlooking the rich perceptual cues and explicit spatial geometry inherent in images. In this study, we aim to mitigate the cognitive vulnerability of models in interpreting gestural spatial relations by proposing PointVG‑R, a reasoning‑guided Multi‑modal Large Language Model (MLLM). PointVG‑R introduces geometric‑aware reasoning for pointing‑based grounding, enabling the model to think with images through the strategic integration of Reinforcement Learning (RL) and cold‑start data. Specifically, we design a novel geometric reasoning pipeline that simulates the iterative cognitive process humans employ when interpreting pointing gestures. Furthermore, we construct EgoPoint‑CoT, a high‑quality visual Chain‑of‑Thought (CoT) dataset featuring detailed reasoning trajectories to guide the model via Supervised Fine‑Tuning (SFT) and RL. To address the varying quality of learning signals encountered during training, we further propose an Adaptive Importance Weighting strategy based on Group Variance, which dynamically adjusts reward signals to optimize the learning process. Experimental results demonstrate that PointVG‑R achieves SOTA performance, outperforming the baseline by 15.86 points in mIoU. Extensive ablation studies further validate the efficacy of our proposed modules. Code: https://github.com/lingli1724/PointVG‑R.
Authors:Ling Li, Zhizhen Cai, Xinkun Wu, Ziyu Zhu, Jiaqing Lyu, Bowen Liu, Zhidong Deng
Abstract:
Grounding deictic gestures in natural images is fundamental to AR and human‑robot collaboration, providing a basis for seamless spatial interaction. While Transformer‑based visual models have achieved significant progress in general object detection, their global attention mechanisms often neglect micro‑geometric relationships, degrading orientation accuracy. In pointing tasks, this deficiency manifests as an inability to accurately capture the pointing ray implied by finger poses, which results in pointing drift and localization ambiguity when dealing with distant or densely packed objects. To address this, we propose VistaRef, a framework designed to explicitly enhance spatial orientation awareness. First, we develop the Local Hand Entity Modeling (LHEM) module, which incorporates hand‑pose embeddings to strengthen the model's capability to capture subtle finger deviations. Second, drawing inspiration from multi‑view geometry, we construct the Geometric Ray Modeling (GRM) module to transform implicit orientation information into explicit spatial geometric features, guiding feature aggregation and deep fusion via attention mechanisms. Furthermore, we introduce a novel Orientation‑Consistent Alignment Loss (OCAL) to synergistically supervise hand presence and pointing consistency, ensuring that all architectural improvements collectively serve the core objective of spatial localization. Experimental results demonstrate that VistaRef significantly outperforms the baseline, achieving a 14‑point absolute gain in grounding accuracy. Qualitative analysis further confirms that VistaRef effectively models the geometric correlation from hand to target, bridging the spatial perception gap inherent in traditional Transformers for complex scenarios. Code: https://github.com/lingli1724/VistaRef.
Authors:Inam Ullah, Imran Razzak, Shoaib Jameel
Abstract:
Learning causal models from fragmented biomedical data is challenging because clinical, molecular, and imaging variables are often incomplete or not jointly observed. We propose RetiSEM, a domain‑constrained structural equation modelling (SEM) framework for causal graph recovery and mediation analysis under limited multimodal resources. This proposed work organises variables into biologically informed blocks, applies forbidden‑edge constraints, and decomposes pathway‑level effects into TE, NDE, and NIE components. We evaluate RetiSEM across ten synthetic benchmark scenarios that vary in dimensionality, nonlinearity, causal depth, and pathway structure, together with a fragmented real‑world setting that combines NHANES clinical variables with externally derived retinal representations. This approach achieves lower structural error and higher causal accuracy than unconstrained baselines across the synthetic benchmarks. In the real‑data analysis, retinal variables behave mainly as downstream biomarker‑like indicators, with smaller but detectable indirect effects. These findings support our strategy as an interpretable framework for testing structured causal hypotheses in limited‑resource biomedical AI. The code and resources for this work are publicly available at: https://github.com/Inamullah‑Colab/ReitSEM.
Authors:Xingsong Ye, Yongkun Du, Jiaxin Zhang, Haojie Zhang, Chong Sun, Chen Li, Jing Lyu, Zhineng Chen
Abstract:
WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt‑oriented scene TExt Recognition (WATER) substantially more challenging than general Scene Text Recognition (STR). Existing STR datasets and methods, typically built around regular scene text and fixed‑template inputs, struggle to scale to WATER. Thus, we aim to advance this task from both data and model perspectives. On the data side, we construct a 2M synthetic dataset, WATER‑S, with the scale improved by hundreds of times compared to existing artistic text data. WATER‑S consists of two complementary subsets. One rendered by an upgraded rendering pipeline (SynthWordArt), which provides highly accurate and controllable synthetic WordArt data. The other is generated by combining Qwen3‑VL for prompt mining and Z‑Image for image synthesis, which improves the coverage of realistic and diverse data. On the model side, we propose WATERec. It adopts an visual encoder supporting arbitrary‑shaped inputs and an autoregressive decoder to model complex layouts, structurally breaking the bottleneck of fixed‑template STR on WordArt. Experiments show that this architecture outperforms prior STR methods, achieving state‑of‑the‑art performance on irregular texts such as WordArt. Together with WATER‑R, carefully reorganized from existing real STR data, our strong baseline with the new synthetic data and model design reaches 90.40% accuracy on WordArt‑Bench, surpassing both general‑purpose and OCR‑specialized vision‑language models by a large margin. Code and data are available at https://github.com/YesianRohn/WATER.
Authors:Peize Li, Fanhu Zeng, Tongda Xu, Xingguo Xu, Xinjie Zhang, Xingtong Ge, Haotian Zhang, Yan Wang
Abstract:
In‑camera JPEG previews are ubiquitous in raw image formats and provide an sRGB reference at negligible storage cost. Although existing metadata‑based reconstruction frameworks can exploit this side information when recovering raw images, their context models often become computationally expensive especially at high resolution, eg, 4K raw image, given that attention mechanisms scale quadratically with feature maps, hindering its practical application. To address these limitations, we propose MambaRaw, a JPEG‑conditioned metadata‑based raw image reconstruction framework that uses State Space Models (SSMs) to estimate entropy parameters efficiently. Our key contribution comprises a Spatial‑Energy Coupled Context Modeling mechanism with two lightweight modules: (1) TileMambaBlock, which performs Mamba‑style selective scanning only on information‑dense tiles to improve the efficiency; and (2) Energy‑Aware Refinement (EAR), an identity‑initialized residual module that enhance feature representation to match the long‑tail energy distribution of raw signals. Extensive experiments on three camera datasets (Sony, Olympus, Samsung) show consistent improvements over strong metadata‑based baselines and set a new state of the art for JPEG‑guided raw reconstruction with great efficiency. Notably, at low metadata bitrates, MambaRaw increases PSNR by 1.2‑‑1.4 dB and reduces end‑to‑end coding latency by about 9%. Code is released at https://github.com/Peizeli1/MambaRaw.
Authors:Tianyu Zhu, Yingping Liang, Hesong Li, Ying Fu
Abstract:
Text‑driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive segmentation losses, which overlooks the geometric consistency across frames and leads to weak spatial understanding. In this paper, we propose Geometry‑enhanced Language‑guided Video segmentation (GeoLaV), a two‑stage framework that distills 3D geometric knowledge from images to enhance text‑driven video segmentation. In the first stage, we perform monocular geometry pretraining with monocular novel‑view synthesis, enabling the model to acquire geometry‑consistent visual representations via spatial alignment on large‑scale single‑image datasets. In the second stage, we introduce geometry‑aware distillation and fine‑tune the model on video segmentation datasets, transferring 3D structural knowledge from a general 3D prior model. This process reinforces 3D awareness and improves both spatiotemporal coherence and language grounding in segmentation. Extensive experiments show that our method using only image segmentation data already provides notable zero‑shot generalization in RVOS. When combined with geometry‑aware distillation for fine‑tuning on videos, our method achieves state‑of‑the‑art performance across multiple RVOS benchmarks. The code is available at https://github.com/Tony1882880/GeoLaV.
Authors:Junpeng Jing, Ronglai Zuo, Zhelun Shen, Shangchen Zhou, Rolandos Alexandros Potamias, Stefanos Zafeiriou, Krystian Mikolajczyk, Jiankang Deng
Abstract:
Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computation, or additional foundation‑model priors, making them difficult to deploy on resource‑constrained platforms. In contrast, efficient stereo models offer faster inference but are commonly considered less capable of strong zero‑shot generalization. In this paper, we challenge this assumption by introducing Lite Any Stereo V2 (LAS2), an ultra‑fast model series designed for efficient zero‑shot stereo matching. LAS2 is developed from both architecture and training perspectives. Architecturally, we revisit efficient stereo design under practical deployment settings and propose a 2D‑only cost aggregation framework, optimized for real inference latency rather than theoretical MACs alone. For training, we develop a three‑stage strategy that combines synthetic supervision, self‑distillation, and real‑world knowledge distillation. To improve the reliability of real‑world pseudo supervision, we further introduce pseudo‑label filtering and an error‑clamping operation, enabling smoother synthetic‑to‑real transfer. We instantiate LAS2 as a family of models, including feed‑forward variants for different efficiency budgets and an iterative variant for higher accuracy. Extensive experiments show that LAS2 achieves state‑of‑the‑art accuracy among efficient stereo methods while maintaining significantly lower latency. Specifically, LAS2‑H achieves stronger overall zero‑shot performance than the iterative method Fast‑FoundationStereo, with 1.8x and 2.7x faster inference on H200 and Orin, respectively. The project page, demos, and code are available at https://tomtomtommi.github.io/LiteAnyStereoV2/.
Authors:Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi, Muzammal Naseer, Naoufel Werghi, Sajid Javed
Abstract:
We revisit the memory update mechanism in SAM2‑based visual object tracking and identify confidence‑only mask selection as the dominant cause of drift under occlusion, rapid motion, and distractors. We introduce SENTRY, a training‑free, plug‑and‑play, refine‑before‑write module that validates each memory update for short‑horizon temporal consistency before committing it. SENTRY aggregates diverse segmentation hypotheses per frame, backtracks them into short tracklets, and uses neighbor‑aware cycle‑consistent matching against recent trajectories to favor temporally and geometrically consistent masks. It leaves the base architecture untouched, replacing confidence‑driven writes with consistency‑validated ones. For fair evaluation, we re‑evaluate major open‑source SAM2‑based trackers across all available scales and datasets, filling gaps in prior reports. Integrated into five strong baselines, SENTRY delivers consistent gains across nine benchmarks, achieving new zero‑shot SOTA on LaSOT, LaSOT_ext, GOT‑10k, VOT20, VOT22, and DiDi. Despite these checks, the SAM2‑L version runs at 32.8 FPS on an A100, and across compatible hosts adds only about 0.4‑‑0.6 GB VRAM. Our results provide the first unified all‑scale evaluation of SAM2‑based trackers and show that enforcing temporal validity at write time stabilizes memory‑augmented tracking without retraining. Project page: https://hamadya.github.io/SENTRY/page/
Authors:Yijia Lei, Jinzhao Li, Yichi Zhang, Jiacheng Hua, Yin Li, Miao Liu
Abstract:
We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision‑language models (VLMs). The benchmark targets streaming interaction understanding, where video frames arrive sequentially and models must continuously interpret evolving visual context. EgoSAT unifies several previously distinct tasks within a single streaming framework. In this formulation, queries about completed events correspond to retrospective reasoning, queries about ongoing activities require online understanding, and queries about future actions involve prospective anticipation. This unified setting requires models to reason about the past, present, and future while operating under the constraint that only previously observed frames are available. EgoSAT contains 1,997 unique videos spanning 165 hours of egocentric footage and around 4,800 high‑quality question‑answer pairs, carefully designed to probe reasoning across varying temporal contexts. Using this benchmark, we evaluate a diverse set of both open‑weight and closed‑weight VLMs, providing a systematic assessment of their ability for streaming interaction understanding. By distinguishing answerability and conducting diagnostics on confidence of models, we find existing models not only struggle with prospective and retrospective modeling, but also exhibit severe mis‑calibration: confidence often fails to track inherent answerability, leading to dangerous "confidently wrong" behaviors. Project page: https://leiyj23.github.io/EgoSAT/
Authors:Hojun Choi, Seulbin Hwang, Dae Jung Kim, Kisung Kim, Hyunjung Shim, Jinhan Lee
Abstract:
Bird's‑eye view (BEV) perception fuses multi‑camera images into a unified top‑down representation for autonomous driving. Despite recent progress, state‑of‑the‑art methods remain confined to closed‑set scenarios, making them vulnerable to unpredictable real‑world environments. In this work, we introduce open‑vocabulary BEV segmentation (OVBS), which leverages vision‑language models (VLMs) to recognize categories beyond the training set while maintaining precise BEV perception and real‑time efficiency. A key challenge in OVBS lies in the 3D geometric inconsistency inherent in the ill‑posed lifting of 2D VLM semantics into BEV. To address this, we propose OVBEVSeg, a geometry‑aware OVBS framework that enhances efficient Gaussian splatting (GS)‑based unprojection by leveraging robust 3D geometric constraints across three progressive stages: (1) 2D‑to‑BEV pseudo‑labeling via reliable 3D projection for OV generalization; (2) joint 2D‑BEV per‑scene optimization with BEV structural constraints for 3D geometric consistency; and (3) 3D geometric distillation for online efficiency. On the nuScenes dataset, OVBEVSeg achieves state‑of‑the‑art performance, outperforming closed‑set methods by 15.3 mIoU on unseen categories. Remarkably, even with no novel‑class ground‑truth labels, it remains competitive with self‑ and semi‑supervised baselines trained with up to 40% of ground‑truth annotations. Furthermore, it achieves 2.5x faster inference with only 0.22x the memory consumption of projection‑based methods. Project page: https://hchoi256.github.io/projects/ovbevseg/.
Authors:Yang Zhou, Wenxue Li, Peng Zhang, Yifei Chen, Fei Wang, Daiguo Zhou
Abstract:
Face Video Restoration (FVR) aims to recover high‑fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods often struggle to simultaneously address three key challenges: identity shift, viewpoint‑entangled guidance, and perceptual realism. To tackle these issues, we propose TIGER, a structured tri‑prior fusion framework that Tames Identity, Geometry, and gEnerative pRiors for high‑quality FVR. Specifically, an Identity Prior is first established by injecting subject‑discriminative embeddings into the latent space, effectively anchoring the subject's identity against severe degradations. Then, to provide temporally consistent structural guidance for dynamic videos, TIGER constructs a Geometry Prior by lifting 2D reference cues into a disentangled 3D parameter space, creating a geometric anchor through cross‑source parameter fusion. Moreover, to achieve maximum efficiency without compromising realism, we harness the video generation model's Generative Prior through a one‑step rectified flow. We further design a progressive three‑stage training optimization strategy that refines structural fidelity, textural reconstruction, and distribution‑level realism to ensure robust optimization. We also construct a large‑scale FVR dataset to facilitate robust training and standardized evaluation. Extensive experiments demonstrate that TIGER achieves state‑of‑the‑art performance in both identity fidelity and temporal stability, delivering a high‑quality, efficient and identity‑consistent FVR. Project page: https://yzhoulv.github.io/Tiger/.
Authors:Jiahao Lyu, Pei Fu, Zhenhang Li, Shaojie Zhang, Jiahui Yang, Yu Zhou, Can Ma, Zhenbo Luo, Jian Luan
Abstract:
In‑Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appearance. Recent unified multimodal models provide a promising solution by combining visual‑text understanding and image generation within a single framework. However, directly adapting such models to IIMT remains challenging. In particular, they often suffer from understanding‑generation conflicts, where the translation inferred during understanding is inconsistent with the text supervision used in generation, and spatial position misalignment, where the rendered text does not accurately match the target text regions. To address these issues, we present UniTranslator, a unified multimodal framework for IIMT that tightly couples translation understanding and text editing. Specifically, we introduce an Understand‑Generation Alignment Module (UGAM) to bridge the representation gap between understanding and generation, encouraging semantic consistency between translated content prediction and text rendering. We further propose a Spatial Mask Decoder (SMD) with pixel‑level supervision over text regions to improve spatial grounding, geometric alignment, and layout controllability during generation. Extensive experiments on multiple benchmarks demonstrate that UniTranslator achieves state‑of‑the‑art performance across diverse language directions and complex real‑world layouts. Moreover, our results reveal a strong mutual reinforcement effect between translation understanding and image generation, highlighting the advantage of unified translation multimodal learning. Code is available at https://github.com/SeerRay‑Lab/Unitranslator.
Authors:Yinji Ge, Guixu Zheng, Wulong Guo, Qian Feng, Xu Wu, Kai Zhou, Xinyuan Liu, Fei Xing
Abstract:
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in‑plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data‑driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs. Consequently, current frameworks typically freeze VFMs and shift the entire burden of rotation generalization to the downstream decoder. To break this architectural bottleneck, we propose REDI‑Match, an efficient framework driven by a novel Rotation‑Equivariant Distillation (REDI) paradigm. Instead of relying on rotation data augmentation to establish rotational correspondences, REDI distills the non‑equivariant semantic representations of a VFM into a lightweight, strictly rotation‑equivariant encoder, leveraging an equivariant geometric architecture to constrain robust high‑dimensional semantics. To fully exploit these features, we equip the decoder with an entropy‑driven spatial alignment module. By evaluating discrete rotation hypotheses, this mechanism explicitly locks onto the canonical coordinate system, eliminating global ambiguity before continuous refinement. Extensive experiments demonstrate that REDI‑Match establishes a new state‑of‑the‑art (SOTA) across multiple benchmarks. Notably, it achieves a 13.89% absolute pose accuracy improvement on the highly challenging SatAst dataset while operating 1.9x faster than the current SOTA (RoMa v2), enabling real‑time inference (~41 FPS) on a single RTX 4090 GPU. Code: https://github.com/YinjiGe/REDI‑Match.
Authors:Sachin Sharma, Michele Flammini, Federico Simonetta
Abstract:
Fine‑tuning transformer‑based handwritten text recognition (HTR) models on medieval manuscripts is challenging because these models are pre‑trained on modern text and must adapt to a very different visual domain. This paper studies how three controllable fine‑tuning choices (contrast normalization, data augmentation, and layer freezing) affect recognition accuracy when adapting TrOCR to small historical datasets. We run controlled experiments on a 13th‑century Italian manuscript (I‑CT 91 "Cortonese") and replicate the same experimental grid on the public READ‑16 benchmark as robustness evidence. On Cortonese, our best configuration achieves 8.03% character error rate (CER). Statistical comparisons across 13 configurations show that freezing up to three encoder layers or six decoder layers does not significantly harm accuracy, while deeper freezing becomes progressively detrimental. Removing contrast normalization (CLAHE) yields 7.84% CER, comparable to a domain‑specialized baseline, suggesting strong optimization can reduce reliance on image preprocessing. Cross‑dataset validation on READ‑16 shows that decoder freezing thresholds transfer more robustly than encoder thresholds, and combined freezing strategies require dataset‑specific re‑validation. Finally, we use Grad‑CAM gradient attributions and decoder cross‑attention maps to diagnose error patterns and failure modes revealed by the ablations. Source code is available at https://github.com/LaudareProject/TrOCR‑analysis
Authors:Hongli Xiao, Youjian Zhang, Yucai Bai, Chaoyue Wang, Yaohui Jin, Xiaoguang Ren, Wenjing Yang, Long Lan
Abstract:
Recovering realistic 3D vehicle models from autonomous driving scenes is crucial for synthesizing training data and building simulation environment. However, most existing vehicle generation methods fail to fully exploit multimodal sensors i.e. multi‑view images and LiDAR point clouds) and rely on neural rendering based reconstruction, leading to low‑quality mesh. Recently, native 3D generative models have made significant progress, yet they are not built for arbitrary multi‑view inputs and often struggle with in‑the‑wild driving images. In this work, we present MM‑TRELLIS, a multi‑modal version of TRELLIS for in‑the‑wild 3D vehicle generation that integrates LiDAR and image sensors from autonomous driving datasets into native 3D generative models. Specifically, multi‑view images are cycled as conditioning inputs, while LiDAR point clouds provide test‑time guidance to ensure geometric accuracy and cross‑view consistency. During denoising, we first align the guidance point cloud with the model priors, then enforce consistency between the generated geometry and the guidance point cloud. Finally, we introduce a voxel filtering strategy based on the opacity of 3D Gaussian Splatting to suppress floaters and produce clean meshes. Comprehensive experiments on Waymo dataset demonstrate our method outperforms existing methods in high‑fidelity 3D vehicle generation. Code is available at https://github.com/HongliXiao/MM‑TRELLIS.
Authors:Yajing Wang, Chao Bi, Junshu Sun, Shufan Shen, Zhaobo Qi, Shuhui Wang, Qingming Huang
Abstract:
Multimodal Large Language Models (MLLMs) have demonstrated impressive vision‑language understanding, yet still struggle with fine‑grained perception in high‑resolution images. While existing training‑free methods typically rely on attention‑based localization or coarse‑to‑fine search, they are often misled by distractors and fail to locate multiple targets. Our investigation attributes these failures to Contextual Dominance, where salient distractors overwhelm target attention and cause inaccurate localization, and Semantic Bias, where global semantics cause the model to fixate on the most salient concept, resulting in incomplete localization in multi‑object scenarios. Built on these insights, we propose ActiveScope, a training‑free framework that enhances MLLMs by actively seeking and correcting perception. ActiveScope features two modules. The Semantic Anchor Localization (SAL) utilizes fine‑grained semantic anchors to independently localize key targets, thereby mitigating semantic bias. The Interference‑Suppressed Refinement (ISR) refines localization by suppressing attention on salient distractions to overcome contextual dominance. Extensive experiments on high‑resolution image understanding benchmarks demonstrate that ActiveScope outperforms existing training‑free methods (e.g., 96.34 percent accuracy on V^ Bench), validating the superiority of the active search and self‑correction paradigm. Our code is available at https://github.com/jasmine‑ww/ActiveScope.
Authors:Kim Youwang, Zhengyu Yang, Liuhao Ge, Yu Rong, Timur Bagautdinov, Su Zhaoen, Nir Sopher, Jovan Popović, Teng Deng, Tae-Hyun Oh, Chen Cao
Abstract:
We introduce FiCA, a Feed‑forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Generating a photorealistic and drivable avatar from just a single image is significantly challenging due to the limited visual information available to accurately infer the 3D appearance and geometry of human heads. To address this, we develop a novel system that combines human‑centric vision foundation models with a diffusion model. This system is designed to fully exploit partial visual observations to generate lifelike human avatars. Our proposed diffusion model learns a generative mapping from these partial observations to complete and authentic 3D mesh reconstruction. Additionally, we introduce a feed‑forward mesh refinement network that enhances the fidelity and identity preservation of the generated avatars, eliminating the need for person‑specific test‑time optimization. By leveraging a universal prior model that decodes a generated mesh into a set of 3D Gaussians, we generate a photorealistic 3D Gaussian avatar, capable of being driven with novel expressions in real‑time. Our experiments demonstrate that the avatars generated by our feed‑forward approach faithfully represent diverse identities and surpass the visual quality of avatars produced by recent competing methods.
Authors:Chirui Chang, Xiaoyang Lyu, Yi-Hua Huang, Haoru Tan, Shizhen Zhao, Yikang Ding, Jianmin Bao, Xin Tao, Pengfei Wan, Xiaojuan Qi
Abstract:
Object‑level geometric edits, including translating, rotating, scaling, duplicating, or removing an object, are routine operations in digital content creation (DCC) workflows, yet they remain unreliable in generative video editing. The key challenge lies in specifying the target object's 3D state change unambiguously across viewpoint and time, while consistently updating geometry‑dependent secondary effects such as shadows and reflections. We introduce GIVE, a geometry‑instructed video editing framework that represents edits through a unified object‑state formulation. Two video‑aligned geometry streams describe the target object before and after editing: a depth‑box encoding coarse 3D placement and extent, and an orientation‑box providing an appearance‑agnostic orientation cue. Together, these streams provide a compact pre/post geometric specification for object‑state transitions. To provide paired supervision for learning these edits, we build a scalable graphics‑engine pipeline that executes object‑level edit programs and renders controlled before/after pairs, isolating the intended geometric edit while keeping secondary effects consistent with the transformation. Experimental results demonstrate that GIVE produces faithful geometric edits with temporal coherence and consistent secondary effects across operators in a unified framework, and shows promising transfer to in‑the‑wild videos. Project page: https://geometry‑instructed‑video‑editing.github.io/give/
Authors:Fuyou Mao, Yifei Chen, Beining Wu, Lixin Lin, Jinnan Dai, Zhiling Li, Yilei Chen, Yaqi Wang, Hao Zhang, Yan Tang, Huiyu Zhou, Feiwei Qin
Abstract:
Accurate pulmonary vessel segmentation remains challenging due to the sparse, tortuous, and multi‑scale nature of vascular structures, where small branches are easily lost and topology integrity is difficult to preserve under voxel‑wise supervision. Existing deep segmentation models primarily optimize binary masks, lacking explicit geometric constraints, thus struggling to recover continuous tubular morphology and fine vascular connectivity. In this study, we introduce MorVess, a morphology‑aware segmentation framework that integrates differentiable geometric priors with large‑scale foundation model adaptation to achieve fine‑grained vascular parsing. MorVess jointly predicts vessel masks, distance maps, and thickness maps, providing explicit supervision for vascular boundaries, centerline consistency, and smooth diameter transitions. A lightweight 2.5D adapter bridges 3D spatial context and 2D SAM representations, while a global‑local fusion block aggregates multi‑level semantics and geometric cues for high‑fidelity topology reconstruction. Across two challenging pulmonary CT benchmarks, MorVess delivers superior Dice, clDice, and HD95 scores, substantially improving small‑vessel recovery and global connectivity. These results demonstrate that embedding geometric intelligence into pretrained vision models offers a principled and scalable pathway toward precise vessel analysis and clinically reliable structural quantification. Our source code is available at https://github.com/MaoFuyou/MorVess.
Authors:Miso Kim, Georu Lee, Yunji Kim, Hoki Kim, Jinseong Park, Woojin Lee
Abstract:
Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co‑occurring concepts. As illustrated in Fig.1, unlearning nudity can unintentionally suppress the concept of person, preventing a model from generating images with person. We define these undesirably suppressed co‑occurring concepts that must be preserved CARE (Co‑occurring Associated REtained concepts). Then, we introduce the CARE score, a general metric that directly quantifies their preservation across unlearning tasks. With this foundation, we propose ReCARE (Robust erasure for CARE), a framework that explicitly safeguards CARE while erasing only the target concept. ReCARE automatically constructs the CARE‑set, a curated vocabulary of benign co‑occurring tokens extracted from target images, and leverages this vocabulary during training for stable unlearning. Extensive experiments across various target concepts (Nudity, Van Gogh style, and Tench object) demonstrate that ReCARE achieves overall state‑of‑the‑art performance in balancing robust concept erasure, overall utility, and CARE preservation.
Authors:Inam Ullah, Imran Razzak, Shoaib Jameel
Abstract:
Automated diabetic retinopathy (DR) grading from colour fundus photographs can achieve strong predictive performance, but clinical interpretation requires more than an image‑level label. It requires understanding how lesion evidence is distributed around retinal vessels and how this evidence relates to quantitative vascular biomarkers. We present a dual‑edge spatial‑Jacobian image graph for interpretable DR grading. Each fundus image is represented as a graph node with four aligned evidence streams: AutoMorph vessel information (X_1), DR‑XAI‑style lesion evidence maps (X_2), a 128‑dimensional lesion‑based contrastive image embedding (X_3), and AutoMorph morphometric biomarkers (X_4). The spatial edge branch (X_12) encodes vessel‑lesion geometry, while the Jacobian branch (X_34) models embedding‑biomarker sensitivity. Lightweight two‑token attention fuses both edge families into a final image graph. On 2,910 matched non‑augmented APTOS images, the full graph achieves 0.8076 accuracy, 0.8312 quadratic weighted kappa, 0.5915 macro‑F1, and 0.9330 adjacent‑grade accuracy; referable DR reaches 0.9055 accuracy and 0.9711 AUROC. The framework is positioned as an explainable representation‑learning tool for lesion‑biomarker hypothesis generation, rather than as a deployment‑ready clinical classifier. The code is available at https://github.com/Inamullah‑Colab/dual‑edge‑dr‑graph‑xai.
Authors:Muyuan Zhang, Jiancheng Zhang, Haijin Zeng, Yin-ping Zhao
Abstract:
While Deep Unfolding Networks (DUNs) dominate video Snapshot Compressive Imaging (SCI), they remain constrained by a uniform design philosophy. Existing methods repeatedly stack high‑complexity priors with identical structures, ignoring the fact that optimization trajectories converge toward static states. This results in representation stagnation, where high‑cost computations are wasted on minimal feature updates. To address this inefficiency, we present Differential Unfolding (DU), a heterogeneous framework that replaces uniform repetition with dynamic evolution. Central to DU is the Differential Evolutionary Framework (DEF), which partitions the unfolding process into two complementary roles: structural anchoring and differential evolution. In this scheme, high‑parameter general stages are sparsely deployed to generate high‑fidelity feature foundations. Complementing these, lightweight differential stages employ a Differential Representation Prior (DRP) to propagate and refine these foundational features through a differential mechanism. By integrating Differential Representation Attention (DRA) for evolving attention maps and a Differential Modulated FFN (DM‑FFN) for feature rectification, DRP effectively models cross‑stage variations with minimal overhead. By focusing computational resources on dynamic evolution rather than static redundancy, DU achieves a superior trade‑off between accuracy and efficiency. Extensive experiments verify that our method establishes new state‑of‑the‑art results while significantly slashing computational overhead. https://github.com/Muyuan‑Zhang/DU
Authors:Min Hyeok Bang, Jun Hyeong Kim, Seung-Wook Kim, Se-Ho Lee
Abstract:
In this paper, we present a novel geometry‑aware style transfer framework for 3D Gaussian splatting (3DGS) that simultaneously transfers appearance attributes and geometric structures. Unlike prior works that primarily focus on color‑based stylization and often overlook structural adaptation, our method explicitly incorporates geometry adaptation through a decoupled optimization scheme that alternately updates color and geometry parameters. This strategy alleviates potential interference between color and geometry updates, leading to stable and consistent scene‑level geometry transformation. The decoupled optimization is enabled by the proposed geometry‑aware contrastive feature matching (GCFM). GCFM integrates RGB, depth, and edge cues into a contrastive objective and is employed in both optimization phases to effectively transfer structural characteristics from style images to Gaussian primitives. Extensive experiments show that our approach achieves superior performance in both qualitative fidelity and quantitative metrics, significantly outperforming existing 3DGS‑based stylization methods. Our code is available at \hrefhttps://github.com/oweixx/gasthttps://github.com/oweixx/gast.
Authors:Tongyan Hua, Dongli Wu, Jinjing Zhu, Yinrui Ren, Zhongcheng Hong, Ying-Cong Chen, Hui Xiong, Wufan Zhao
Abstract:
Generating explicit 3D city assets from a single satellite image is important for digital twins, urban simulation, and geospatial intelligence. Unlike satellite‑to‑street‑view synthesis, the task requires a reusable textured mesh with plausible geometry and controllable appearance rather than a 3D proxy optimized only for rendering a small set of images or videos. The ICCV Sat2City framework made a first step by conditioning cascaded sparse‑voxel latent diffusion on satellite‑derived height maps, but its appearance was random, its training data were synthetic, and its task‑specific VAE did not scale well to noisy real‑world reconstructions. We present Sat2City v2, a journal extension that adapts a pretrained native structured‑latent 3D foundation model to weakly aligned satellite images and textured meshes. We build a real‑world dataset with 16,241 satellite‑mesh pairs across 24 regions in 9 cities. Instead of learning a 3D representation from noisy city meshes, Sat2City v2 encodes each mesh into a pretrained native 3D latent space, fine‑tunes a satellite‑conditioned geometry flow, and uses the decoded shape to anchor satellite‑conditioned texturing. This retains Sat2City's geometry‑to‑appearance cascade while enabling appearance‑controllable generation from the satellite input. Experiments on metric‑scale DSM reconstruction and generative city‑asset benchmarks for geometry and appearance show that Sat2City v2 achieves the best overall performance among evaluated baselines. Overall, Sat2City v2 advances satellite‑to‑city generation from rendering‑oriented 3D proxies to explicit textured mesh assets, supported by, to the best of our knowledge, the first documented satellite‑mesh paired dataset collected from matched geographic crops for this asset‑level task. Project page: https://ai4city‑hkust.github.io/Sat2City‑v2/
Authors:Hengji Zhou, Sijie Liu, Jianrun Chen, Xingchen Zou, Lianghao Xia, Liqiang Nie
Abstract:
Short dramas, with their rapid shot rhythms, dialogue‑driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt‑level or text‑only video generation pipelines struggle to meet. We study plot‑to‑short‑drama generation, where a global plot and local context are transformed into visually grounded multi‑shot videos. We propose DramaDirector, a geometry‑grounded framework that lets the planner borrow cinematographic geometry from a gallery of real short‑drama shots indexed by depth and pose. DramaDirector decouples each shot into static visual and dynamic narrative conditions, trains the planner with schema‑constrained SFT and GRPO under a learned text‑visual alignment reward, and retrieves depth‑pose references to guide first‑frame generation and image‑to‑video synthesis. We also introduce DramaBoard, a benchmark built from 35 live‑action dramas, 2.8K episodes, and 81K shots, with structured storyboards and multi‑dimensional evaluation protocols. Experiments show that DramaDirector improves over representative multi‑agent and video generation baselines on faithfulness, consistency, and controllability. Our code is released at: https://github.com/iLearn‑Lab/DramaDirector
Authors:Tian Qiu, Jifeng Shen, Xin Zuo
Abstract:
Effective cross‑modal feature alignment and interaction are central challenges in multispectral object detection. Although global cross‑attention provides strong long‑range modeling ability, its quadratic complexity with respect to feature size limits deployment on resource‑constrained platforms. We therefore propose Progressive Pixel‑Neighborhood Deformable Cross‑Attention for multispectral feature fusion, termed PNAFusion. The proposed framework is motivated by two observations: weak misalignment between visible and thermal images is usually concentrated around local neighborhoods, and semantic correspondence across modalities often follows non‑linear spatial mappings that fixed receptive fields cannot model well. To address these issues, PNAFusion incorporates local spatial priors into its architectural design to concentrate feature interaction and alignment on the most relevant neighborhoods. Specifically, a Pixel‑Neighborhood Cross‑Attention (PNCA) module is introduced to avoid redundant global feature matching and suppress background noise. Meanwhile, an Adaptive Deformable Alignment (ADA) module captures non‑linear spatial correspondences through learned pixel‑wise offsets. These components are further integrated through an iterative feedback mechanism to progressively refine cross‑modal feature alignment. Experiments on FLIR, M3FD, and DroneVehicle show that PNAFusion achieves 84.2, 90.5, and 85.5 mAP@0.5, respectively, under the YOLOv5 detector, and further reaches 86.8 mAP@0.5 on FLIR and 90.8 mAP@0.5 on M3FD when transferred to Co‑DETR. Efficiency analysis indicates that PNAFusion reduces allocated GPU memory by 33.0% compared with ICAFusion and reduces theoretical FLOPs from 194.8 G to 156.4 G, although the deformable sampling and iterative refinement introduce additional latency. Our code will be available at https://github.com/DanielQiuTian/PNAFusion.
Authors:Amirhossein Kardoost, Lion Gleiter, Tingying Peng, Carsten Marr
Abstract:
Self‑supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three‑dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE‑2D vs. MAE‑3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE‑3D consistently outperforms 2D max‑projection and slice‑based variants on downstream single‑cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross‑modal supervision yields larger gains for volumetric models. Channel cross‑attention and frequency‑domain regularization are critical for leveraging 3D spatial context. On a protein‑‑protein interaction task, MAE‑3D achieves a ROC‑‑AUC of 0.865, outperforming prior methods by up to +0.025. For protein localization, our best 3D model attains state‑of‑the‑art AUC_\textmicro (0.952) and F1_\textmicro (0.742), improving over previous approaches by +0.003 and +0.010 absolute, respectively. Overall, these results demonstrate the advantages of native 3D modeling and multimodal alignment for representation learning in single‑cell microscopy.
Authors:Qian Wang, Zhenyu Li, Abdelrahman Eldesokey, Peter Wonka
Abstract:
Subject‑driven image generation faces an "Identity‑Diversity Paradox", where strong identity preservation often leads to rigid and low‑diversity outputs. We propose a post‑training framework called DivRL that jointly optimizes identity consistency and structural diversity simultaneously by leveraging disentangled visual features from a robust similarity model. Specifically, we introduce a Negative Self‑Similarity Measure (nSSM) to quantify structural diversity, and Visual Semantic Matching (VSM) to evaluate identity consistency. We propose an "Explore‑and‑Suppress" strategy that treats VSM as a gated constraint: the model freely explores structurally diverse configurations, and only samples that violate the identity threshold are penalized via a quadratic hinge loss. This converts identity preservation from a competing objective into a feasibility constraint, allowing nSSM and VSM to improve jointly. Experiments demonstrate that our method effectively pushes the model to generate both consistent and diverse images and improves structural diversity while maintaining comparable identity consistency through a gated optimization formulation.
Authors:Yifei Zhao, Qian Lou, Mengxin Zheng
Abstract:
Vision‑language models (VLMs) are increasingly used as perception‑reasoning backbones for embodied intelligence in safety‑critical physical systems, where perception or reasoning errors can lead to unsafe decisions or actions. Although many red‑teaming methods have been developed to probe VLM vulnerabilities, their evaluation remains fragmented across datasets, metrics, and threat models, making direct comparison difficult and obscuring whether observed differences arise from stronger attacks, more vulnerable models, or incompatible evaluation settings. Existing chatbot‑centric red‑teaming benchmarks mainly standardize jailbreak and content‑safety evaluation, but they do not systematically capture physically grounded functional failures or cover red‑teaming methods that target physical‑world VLMs. This raises the key challenge of comparing diverse attack methods under a unified protocol while targeting the same scenario‑specific failures. We introduce REALM, to our knowledge the first unified red‑teaming benchmark for physical‑world VLMs. REALM integrates 12 red‑teaming methods, 3 model‑agnostic defenses, and 13 VLMs under a practical black‑box threat model with shared datasets and metrics. To align adversarial objectives across attack families, REALM introduces an agentic target‑generation pipeline that constructs shared, scenario‑specific, and physically grounded attack objectives for each scene, enabling fair comparison of diverse red‑teaming methods under aligned adversarial goals. Our evaluation shows that text and typographic injection attacks induce the most failures, multimodal co‑optimization yields the strongest visual‑perturbation transfer, single‑pass attacks approach iterative methods at much lower cost, and model scale alone does not confer adversarial robustness. Code is available at https://github.com/UCF‑ML‑Research/REALM.
Authors:Qian Ma, Qiong Wu, Zhengyi Zhou, Yao Ma
Abstract:
Knowledge‑Based Visual Question Answering (KB‑VQA) requires grounding visual queries to external knowledge beyond directly observable content in images. While recent multi modal large language models (MLLMs) show strong perceptual abilities, they struggle on KB‑VQA tasks requiring groundings from both fine‑grained entity and evidence levels. Most existing multi‑modal retrieval augmented generation (MM‑RAG) methods tightly couple entity discrimination and section‑level evidence ranking into a single re‑ranking stage, leading to high cost and limited generalization. In this work, we revisit existing MM‑RAG solutions from a workflow perspective and argue both entity‑level and fact‑level groundings are key bottlenecks. We observe that although MLLMs often fail under open‑ended entity naming, they can better identify the correct entity when selecting from a small set of candidate names. Based on this insight, we propose a simple and training‑free identify‑before‑answer IBA framework that decouples entity identification from section‑level re‑ranking. Our approach prompts an MLLM to select high‑confidence entities using only candidate names, followed by an off‑the‑shelf textual re‑ranker for evidence selection. Experiments on Encyclopedic‑VQA and InfoSeek show that our method consistently outperforms fine‑tuned multi‑modal re‑ranking baselines while reducing training and inference complexity. Additional analyses reveal that the improvements arise not only from better entity identification, but also from selecting more informative evidence once correct entity is fixed. Our implementation is made public to ease reproducibility.
Authors:Anindya Mondal, Sauradip Nag, Anjan Dutta
Abstract:
ABACUS is a unified vision‑language model that handles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation without any benchmark‑specific training required. Our model is built on existing 3B‑parameter unified foundation model and is adapted for object localization tasks using three key innovations: density‑aware adaptive zooming with objectness maps for spatial grounding; a boundary‑aware count policy via GRPO to eliminate crop‑boundary errors; and a cycle‑consistent GRPO strategy where the understanding branch self‑critiques generated outputs, closing the understanding‑generation gap without any external annotations. ABACUS achieves state‑of‑the‑art results across seven benchmarks, outperforming both task‑specific specialists and larger generalist models.
Authors:Yashkumar R Lukhi, Harsh Rameshbhai Moradiya, Radu Timofte, Dmitry Ignatov
Abstract:
We present an automated large‑scale search pipeline for heterogeneous 4‑Expert Mixture‑of‑Experts (MoE4) architectures within the LEMUR neural network dataset ecosystem. Building on a hand‑crafted heterogeneous MoE reference model, we replace manual design with a deterministic code‑assembly generator that systematically combines base architecture families drawn from the LEMUR database into MoE4 ensembles, each governed by a convolutional gating network with temperature scaling, mixup augmentation, and cosine‑annealed learning rate scheduling. Over a 28‑day campaign on an NVIDIA RTX 4090, the pipeline generated 4,463 candidate models across 197 batches, of which 1,021 were evaluated successfully. A critical finding emerged from the campaign: due to alphabetical enumeration via itertools.combinations, the entire explored search space (4.8% of the theoretical 23,751 possible 4‑family combinations) is anchored to a single family, AirNet. We characterise this coverage bias precisely, identify the root cause in the generator, and propose a stratified random sampling fix. Within the AirNet anchored scope, ShuffleNet and MobileNetV3 consistently co‑produce the highest‑accuracy ensembles (mean accuracy up to 0.632), while FractalNet and MNASNet are identified as low‑yield families warranting exclusion in future campaigns. The pipeline, analysis artefacts, and corrected generator are released as part of the open‑source NNGPT project at https://github.com/ABrain‑One/nn‑gpt
Authors:Yubo Zhou, Jianghao Wu, Ping Ye, Shaoting Zhang, Guotai Wang
Abstract:
Concept segmentation models like Segment Anything Model 3 (SAM3) show strong generalization on natural images, yet their performance degrades in medical imaging due to the domain gap caused by different imaging principles and styles. Test‑Time Adaptation (TTA) is essential for improving the testing performance by updating the model on the fly without annotations. However, existing vision‑language TTA methods are mainly driven by image‑level uncertainty minimization, which does not necessarily reflect region‑level semantic correctness in medical segmentation. Moreover, they often lack mechanisms to maintain stability in continual one‑pass adaptation, leading to limited performance when reliable dense supervision is missing for segmentation. To address these issues, we propose Concept Alignment Contrast and LongShort Prompt Memory for Test‑Time Adaptation (CM‑TTA) of SAM3 for medical images. First, for a test sample with multiple augmentations, we introduce a novel Concept Alignment Contrast (CAC) metric, which leverages textual‑visual semantic consistency to robustly evaluate prediction quality to select the best augmented view as the supervision. Second, to balance rapid and stable adaptation, we design a Long‑Short Prompt Memory (LSPM) module. The short memory dynamically fuses recent prompts based on CAC scores for agile local adaptation, while the long memory maintains a stable global prompt to generate enhanced pseudo‑labels. Finally, a Densely Supervised Prompt Update (DSPU) strategy is proposed to optimize the prompt embeddings with enhanced pseudo labels as dense supervision. Extensive experiments on prostate and skin lesion segmentation demonstrate that our CM‑TTA framework significantly outperforms existing methods for TTA of SAM3. The code is available at https://github.com/SherlockZYB/CM‑TTA.
Authors:Junrong Huang, Zhiyuan Zhang, Rui Tang, Hongbo Fu, Jnig Liao
Abstract:
Realistic integration of user‑specified textures into scene images is a fundamental task in computer graphics and image editing. While existing material transfer and reference‑guided inpainting methods can edit surface appearances, they often fail to address the specific requirements of texture tiling. This task necessitates precisely repeating a reference pattern according to user‑defined parameters such as frequency, orientation, and scale. Furthermore, current generative approaches often struggle to maintain the structural fidelity of the reference texture, limited by either destructive pixel‑level resampling or the lack of fine‑grained spatial information in semantic image encoders, and they frequently fail to preserve the coherent lighting and geometry of the original scene. In this paper, we propose a novel framework for controllable and high‑fidelity texture tiling based on Diffusion Transformers. Our approach introduces two key technical innovations to decouple spatial manipulation from content generation. First, we propose a Coordinate‑Transformed Rotary Embedding mechanism. By applying 2D affine transformations directly to the relative positional embeddings between the target latent and the image condition, we achieve precise control over tiling patterns without explicit pixel warping, thereby utilizing the full information of the reference condition without degradation. Second, a Disjoint Attention Mask is employed to shield reference features from semantic leakage. This preserves structural integrity while seamlessly blending the synthesized texture with the scene's original lighting and geometry. Extensive experiments demonstrate that our method outperforms state‑of‑the‑art baselines in both control accuracy and texture fidelity.
Authors:SingGuard Team
Abstract:
Vision‑language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can arise from multimodal question answering, assistant responses, and cross‑modal composition, while moderation policies may vary across products, regions, and deployment stages. Most existing guardrails either rely on fixed taxonomies or target only a narrow set of interaction settings, which limits their adaptability when safety rules change at deployment time. We present SingGuard, a policy‑adaptive multimodal guardrail model family for safety assessment in multimodal conversations. SingGuard treats the active policy as a runtime input: given natural‑language rules, it checks the target content against the active policy rule by rule and predicts both the safety label and the triggered rule. To balance efficiency and interpretability, SingGuard supports fast, hybrid, and slow inference regimes along a fast‑to‑slow reasoning spectrum, ranging from direct safety judgments to policy‑grounded deliberation. We further optimize this behavior with fast‑‑slow decoupled reinforcement learning. We also introduce SingGuard‑Bench, a multimodal guardrail benchmark with 56,340 examples spanning 80+ fine‑grained risk types across multimodal QA, adversarial attack, and dynamic‑rule evaluation settings, including cross‑modal joint‑risk cases where each modality is harmless in isolation but their composition implies unsafe intent. Across six benchmark families (35 datasets), SingGuard achieves state‑of‑the‑art average F1 in every family. Dynamic‑rule evaluation further shows improved policy‑following accuracy from 0.6465 to 0.7415 under runtime policy shifts. Our code is available at https://github.com/inclusionAI/Sing‑Guard.
Authors:Francesco Di Salvo, Sebastian Doerrich, Christian Ledig
Abstract:
Foundation models provide highly descriptive representations for medical images, yet their reliability degrades under distribution shifts arising from changes in patients, devices, or acquisition conditions. Reliable out‑of‑distribution (OOD) detection is therefore essential for safe deployment. Recent post‑hoc detectors efficiently exploit frozen embeddings (e.g., kNN), whereas reconstruction‑based OOD detection in latent feature space has seen limited adoption due to inconsistent performance. In this work, we show that the limitation of reconstruction‑based methods in latent space does not stem from poor reconstruction quality, but from how reconstruction errors are scored. Standard L2 residual norms collapse the anisotropic residual structure, thereby suppressing informative deviations. To address this limitation, we introduce MaRS (Mahalanobis Residual Scoring), a label‑free OOD detector that learns an in‑distribution manifold using a lightweight autoencoder and measures deviation via a Mahalanobis distance on reconstruction residuals, yielding variance‑aware OOD scores. Across three imaging modalities, multiple types of distribution shift, and different model families and scales, MaRS outperforms established confidence‑, distance‑, and reconstruction‑based baselines, while remaining fully post‑hoc and lightweight. The code is available at https://github.com/francescodisalvo05/mars.
Authors:Srinivas Venkatanarayanan, Clement Pakkam Isaac
Abstract:
Vision‑language models (VLMs) are increasingly used to read maps for logistics, delivery, and accessible navigation, where the output is an actionable decision (a route, a pin, a parking choice) that must respect the road network. Yet most map benchmarks grade free text or multiple‑choice answers that cannot be verified against the underlying graph. We present MapReason‑OSM, a benchmark and evaluation harness for graph‑verifiable mobility decisions on self‑rendered OpenStreetMap panels. We render fixed‑style maps for ten U.S. downtowns at two aligned zoom scales, overlay a consistent marker grammar, and pair each panel with a hidden street graph and exact oracles, yielding 6,000 instances (12,000 panels across the two zooms) over 12 routing, facility‑location, and visual disambiguation tasks. Models return structured decisions that we snap back to the graph and score for validity, legality, optimality, and constraint satisfaction, plus cross‑zoom consistency. Across seven VLMs, models read maps and route simply but fail at graph cost reasoning (single‑facility pin placement is near chance even for frontier reasoning models), and are frequently scale‑inconsistent. We release the benchmark, harness, and deterministic generator. Code and data: https://github.com/Vi‑Sri/mapreason‑osm
Authors:Nicolò Savioli
Abstract:
Modern Referring Image Segmentation (RIS) systems generate multiple candidate masks per expression but rely on a simple heuristic‑‑typically the argmax detection score‑‑to select the final output. We identify query selection as a failure‑case bottleneck: although heuristic selection succeeds on 82‑93% of samples, the residual 7‑18% of failures dominate the error budget, leaving a best‑query selection gap of 3‑11% mIoU. We introduce Venice‑H1, a lightweight, backbone‑decoupled post‑hoc re‑ranking module that encodes each candidate through multi‑scale grid signatures‑‑compact spatial descriptors pooled onto 4x4, 8x8, and 16x16 grids‑‑and feeds them to a Transformer‑based re‑ranker with a Failure Gate (ROCAUC 0.78‑0.82) that intervenes only when the default choice is likely suboptimal. Instantiated on DeRIS‑L and DeRIS‑B, Venice‑H1 achieves delta_fail of +1.40 and +0.89 mIoU with strictly positive 95% CIs on all 16/16 (split, backbone) pairs and harmful‑switch rates below 0.53%. Zero‑shot transfer to medical referring segmentation (MS‑CXR, M3D‑RefSeg‑2D) yields +1.16 and +0.51 mIoU without RIS‑backbone fine‑tuning. The module adds approximately 11.3M parameters and under 1 ms latency.
Authors:Xianghui Wang, Feng Chen, Wenbo Zhang, Hua Yan, Zixuan Wang, Changsheng Li, Yinjie Lei
Abstract:
Vision‑Language‑Action (VLA) models provide a unified paradigm for robotic manipulation, yet their real‑world deployment is often bottlenecked by execution efficiency. While existing efforts predominantly focus on compute‑centric efficiency to reduce per‑step inference latency, the intrinsic policy efficiency of these models remains largely unexplored. Policy efficiency is fundamentally affected by two factors, namely the effective executable length of predicted action chunks and the total physical steps required to complete a task. These two factors jointly determine the total number of forward inference calls during execution. We observe that current VLA policies struggle with planning unreliability and action redundancy, suffering from severe prediction degradation at the tail of action chunks and tending to generate unnecessarily redundant physical steps. To address this, we propose PolicyTrim, a reinforcement learning‑based post‑training framework that extends the reliable action chunk length and reduces redundant physical steps. For reliable chunk extension, we employ a dynamic exploration strategy that explicitly rewards the successful completion of longer executable lengths, progressively pushing the trustworthy prediction horizon to its empirical limit. For step efficiency, we design a redundancy‑aware reward that directly favors successful task completions with fewer steps while penalizing unreproducible shortcuts, effectively eliminating redundant physical actions. Extensive experiments across three benchmarks and three VLA models demonstrate that PolicyTrim improves action chunk utilization by 3× and reduces physical execution steps by 51.4%. Ultimately, our framework delivers up to a 5.83× end‑to‑end deployment speedup without compromising task success rates.
Authors:Xinlong Chen, Jiafu Tang, Yue Ding, Yizhuo Jia, Bozhou Li, Bohan Zeng, Yang Shi, Shihao Li, Yiyan Ji, Qiang Liu, Weihong Lin, Yuanxing Zhang, Pengfei Wan, Liang Wang, Tieniu Tan
Abstract:
Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can objectively and comprehensively evaluate these properties across diverse durations and scenarios, thereby hindering the advancement of video captioning models. To bridge this gap, we propose CapRiCorn‑1K, a comprehensive benchmark designed to evaluate both video captioning quality and subject referential consistency across long temporal horizons and diverse video domains. To accommodate varied evaluation needs, our benchmark supports both audiovisual and visual‑only settings. Extensive experiments on CapRiCorn‑1K reveal that current models generally struggle to generate accurate and comprehensive captions while maintaining consistent subject references. Moreover, as video duration increases, both the overall caption quality and subject referential consistency decline. Notably, our evaluation metrics exhibit strong correlations with the performance of downstream understanding and generation tasks conditioned on the generated captions, further validating their effectiveness. The project is available at https://github.com/xlchen0205/CapRiCorn‑1K .
Authors:Abdirashid Omar, Jonghyuk Park
Abstract:
Remote sensing change detection (CD) from bi‑temporal imagery is critical for applications such as urban monitoring, disaster assessment, and environmental management, yet robust localization remains challenging under sparse changes, noisy labels, and appearance variations. In this paper, we propose Context Sampling Attention (CoSA), a lightweight decoder‑side refinement module that explicitly leverages bi‑temporal feature correlation as a control signal for adaptive change‑aware feature enhancement. This differs from conventional attention mechanisms that rely on implicit feature weighting without explicit temporal control. In the implemented FC‑Siam setting, CoSA computes normalized same‑location cross‑correlation between paired decoder features, converts low correlation into a change gate, and injects the resulting gated residual at native 1/8 and 1/16 feature scales through learnable residual scaling. This design enables effective discrimination between stable and ambiguous regions without relying on computationally expensive global attention. Extensive experiments on four benchmark datasets (LEVIR‑CD, S2Looking, DSIFN, and CLCD) demonstrate consistent improvements over strong baselines, achieving 1.5‑2.6% gains in changed‑class F1 while introducing negligible parameter overhead. Ablation studies confirm that multiscale placement and learnable residual gating are both important for peak performance. These results indicate that CoSA establishes a practical and effective refinement paradigm for enhancing temporal discriminability in Siamese change detection frameworks.
Authors:Qing Xu, Xiangjian He, Wenting Duan, Jiebo Luo, Zhen Chen
Abstract:
Cell segmentation is critical for computational pathology and biomedical discovery. While recent Vision Foundation Models (VFMs) have demonstrated remarkable universal feature representations, unlocking their full potential for cellular imaging is currently bottlenecked by resource‑intensive adaptation paradigms. Existing methods typically rely on fine‑tuning heavy visual encoders, leading to extensive computational overhead and a dependency on large‑scale annotations. To address this, we propose the EffiCell‑Seg framework for highly efficient cell segmentation without re‑training the visual encoder. Our core insight is that pretrained VFMs intrinsically encode complementary structural priors: global saliency for localizing potential cells, and local morphological patterns for delineating cellular structures. To harness these priors, we devise a Cell Structure Prompt Encoder (CSP‑Encoder) that synthesizes semantic‑aware saliency and principal morphological features from frozen VFM representations into explicit structural prior maps. Moreover, we propose a Synergistic Mask Decoder (SM‑Decoder) that enforces contextual consistency by jointly predicting geometric distance fields and semantic maps via mutual cross‑guidance. Extensive experiments demonstrate that EffiCell‑Seg outperforms state‑of‑the‑art methods across diverse cell imaging modalities while requiring only ~5M trainable parameters, over 130x fewer than fully fine‑tuned VFM counterparts. The code is available at https://github.com/xq141839/EffiCell‑Seg.
Authors:Yu-Syuan Xu, Hao-Lun Sun, Hao-Wei Chen, Hsien-Kai Kuo, Chun-Yi Lee
Abstract:
Arbitrary‑scale image super‑resolution (ASISR) aims to reconstruct high‑resolution images from low‑resolution inputs over a continuous range of upscaling factors. While traditional pixel‑regression approaches often produce overly smooth results that lack realistic details, recent diffusion methods can produce sharper and more realistic textures. However, these diffusion techniques frequently introduce the risk of structural hallucinations. To address these issues, we propose Fidelity‑ and Perception‑Aware Local Implicit Attention (FPLIA), a framework that effectively integrates fidelity‑oriented features into a diffusion pipeline to produce realistic and faithful reconstructions for ASISR. We introduce a Fidelity and Perception Attention Module (FPAM), which applies both self‑attention and cross‑attention to fidelity‑oriented and perceptual features to enhance representational capacity. To further exploit their complements, we design a Fidelity and Perception Select Module (FPSM) that adaptively selects the most representative features for RGB values prediction. We conduct extensive experiments to validate the effectiveness of these components. Both qualitative and quantitative results show that FPLIA delivers superior perceptual realism while maintaining reconstruction accuracy on standard ASISR benchmarks. The source code is accessible at the following repository: https://github.com/XUSean0118/FPLIA.
Authors:Yanghui Song, Nanqing Liu, Haonan Yin, Yingjie Gao, Chengfu Yang, Qi Ming
Abstract:
Open‑vocabulary semantic segmentation (OVSS) in remote sensing images aims to segment categories beyond a fixed label space. Recent SAM 3‑based methods provide a promising training‑free foundation, yet three key issues remain: (1) a single class‑name prompt lacks sufficient semantic coverage for complex remote sensing categories; (2) expanding each category into multiple prompts introduces redundant online text encoding; and (3) directly aggregating multiple prompt responses propagates noisy activations into the final prediction. To address these issues, we propose ProC‑SAM3, which calibrates SAM 3's prompt interface for remote sensing OVSS from three complementary aspects. First, we construct an offline prompt pool where a Category Matcher groups MLLM‑generated candidates into per‑category sets, and Expansion Constraints further refine each set using category‑specific prior knowledge. Second, the resulting text embeddings are cached and reused across all test images, eliminating repeated text encoding. Third, we introduce Presence‑Guided Residual Fusion to gate unreliable decoder outputs by prompt presence and confidence, followed by peak‑preserving class aggregation that retains fine‑grained activations for small and sparse objects. Experiments on eight benchmarks show that ProC‑SAM3 achieves an average mIoU of 56.1%, outperforming the previous best training‑free method by 3.9 percentage points. Code will be available at https://github.com/YanghuiSong/ProC‑SAM3.
Authors:Prithvi Raj Singh, Satyendra Singh
Abstract:
We present MARLNet (Motion‑Aware Reinforcement Learning Network), a PPO‑based bounding‑box refinement agent that incorporates a constant‑velocity motion prior into the observation state and an action smoothness penalty into the reward function. The agent operates on 268‑dimensional observations encoding the current proposal, a kinematic prediction, the previous action, and a 256‑dimensional EfficientNet‑B0 crop feature, and learns a five‑dimensional policy controlling coordinate adjustments and a binary termination trigger. Evaluated on Pascal VOC 2012 and VisDrone 2019, MARLNet trains stably across all regularization strengths tested and achieves consistent gains in detection success rate at \textIoU \geq 0.5: up to +0.011 on VOC (λ_\textphys=0.10), where the motion prior prevents the overshooting that causes plain PPO to regress on this metric, and +0.007 on VisDrone (λ_\textphys=0.70), where unconstrained PPO achieves a larger gain (+0.025) owing to the weaker base detector. Through reward design ablations and training dynamics analysis, we identify a reward interference in which combining a constant‑velocity deviation penalty with an absolute IoU term causes trigger collapse, and show that replacing it with the action smoothness penalty resolves this failure. We further characterize a representational ceiling facing crop‑feature refinement agents that share a backbone with their base detector, confirmed through a global‑plus‑local observation ablation. Project page: https://prithviraj97.github.io/marl‑net
Authors:Chu-Hsuan Lin, Alberto Mario Ceballos-Arroyo, Jisoo Kim, Shrikanth M. Yadav, Huaizu Jiang, Lei Qin, Geoffrey S. Young
Abstract:
In this work, we present SemanticVessel, a dataset for fine‑grained brain vessel segmentation in computed tomography angiography scans. Based on the detailed contrast provided by dynamic 4D‑CTA scans, we generate segmentation traces for arteries and veins. We then use intensity‑guided region growing to obtain segmentations of the majority of vascular territories in the human brain, which are refined and annotated with 20 unique arterial classes by an expert radiologist. Unlike existing datasets, where minor arteries are discarded as background content, we merge these minor arteries into a generic arterial class. Due to the multiple‑phase acquisition of dynamic 4D‑CTA, labels for a single phase can be re‑used for other phases in the same series, greatly increasing the size of our dataset with no additional annotation cost. The results show that models trained with the additional generic artery class produce better fine‑grained segmentations across the board. We will make our code, annotation GUI, and model weights available to the scientific community. Code, weights, and data will be made available on https://github.com/alceballosa/robust‑vessel‑segmentation
Authors:Jiehui Huang, Yuechen Zhang, Bin Xia, Jiahao Wang, Xu He, Zhenchao Tang, Meng Chu, Xin Tao, Pengfei Wan, Jiaya Jia
Abstract:
Generating a coherent multi‑shot video requires structured cross‑shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches either train end‑to‑end over fixed‑length sequences and cannot scale, generate shot‑by‑shot with memory banks that grow linearly, or orchestrate pretrained generators under an LLM planner without a multi‑shot‑aware backbone. We present UnityShots, a memory‑driven multi‑shot audio‑video generation system built on LTX‑2.3, trained on annotated cinematic and music‑video shots. The video stream maintains two fixed‑size slots, a long‑term memory (LTM) slot anchored to the opening shot and a short‑term memory (STM) slot holding the immediately preceding tail, both updated at every cut by a boundary‑conditioned gate that fuses visual cut probability and beat‑tracker signals. The audio stream injects a reference speaker token at every shot to preserve vocal timbre without a sliding audio bank. A discrete cut‑type prior, learned through AdaLN, becomes an inference‑time control knob over transition strength. We release a benchmark of 200 multi‑cultural multi‑shot sequences spanning six ethnic regions and ten or more languages, with per‑shot reference identities, reference audio, and per‑boundary transition labels. Evaluated across I2V, T2V, and R2V conditioning modes, UnityShots leads open‑source baselines on every cross‑shot coherence metric and matches the strongest closed‑source system on the multi‑shot axes.
Authors:Zijian Fu, Xiangyang Chu, Mengshi Qi, Huadong Ma, Guanghao Zhang, Wei Li
Abstract:
Instruction‑conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and historical motion, but also from a natural‑language maneuver instruction. This paper presents our submission to the doScenes Instructed Driving Challenge, built upon OmniDrive, a vision‑language‑action driving agent with 3D perception, reasoning, and planning capabilities. We adapt OmniDrive to the doScenes setting by training it on instruction‑annotated nuScenes scenes and generating a 6‑second ego trajectory represented by 12 future waypoints. To improve multi‑view visual grounding, we further introduce a DVPE‑style divided‑view perception module into the OmniDrive perception head. Instead of attending globally to all camera features, the proposed module groups query features and image tokens into divided local view spaces and performs visibility‑aware cross‑attention within each view. This design reduces irrelevant cross‑view interference and helps the model better align language instructions with local driving‑relevant visual evidence. The code is publicly available at: https://github.com/feel12348/doscenes‑omnidrive.
Authors:Efe Ilıcak, Baris Imre, Chloé Najac, Ruben van den Broek, Beatrice Lena, Andrew Webb, Marius Staring
Abstract:
Deep unrolled networks (DUNs) integrate physical forward models with learned regularization in cascaded network architectures, achieving exceptional performance in inverse problems while maintaining interpretability. While most DUNs operate in the object domain (e.g., image space), recent variants explored representation spaces for improved information flow. However, these methods rely on heuristic methods for data consistency (DC), sacrificing fidelity with measurements. In this work, we introduce DUNE (Deep Unrolled Networks in rEpresentation space), a framework that maintains exact adherence to physical measurements while operating in learned representation spaces. By deriving the DC gradient via the chain rule and implementing it through the Vector‑Jacobian Product (VJP), we enable exact backpropagation of measurement residuals into the representation space. This formulation supports diverse architectural backbones, including pre‑trained encoders to guide the iterative process. We assess DUNE against state‑of‑the‑art baselines on accelerated MRI reconstruction tasks, demonstrating that exact VJP‑based gradients yield superior reconstruction quality and structural fidelity across both single‑channel portable low‑field and multi‑channel clinical high‑field MRI acquisitions. The code will be available upon publication at https://github.com/EfeIlicak/DUNE.
Authors:David Pascual-Hernández, Roberto Calvo-Palomino, Inmaculada Mora-Jiménez, Jose María Cañas-Plaza
Abstract:
In this report, we present our submission to the GOOSE 2D Fine‑Grained Semantic Segmentation Challenge, organized as part of the Workshop on Field Robotics at ICRA 2026. The challenge combines data from the GOOSE and GOOSE‑Ex datasets, which comprise more than 13k images captured from 4 distinct camera setups, annotated using a hierarchical taxonomy of 56 fine‑grained classes and 11 broader categories. Starting from SegFormer as a baseline, we progressively improve segmentation performance through increased training crop sizes, a transition to the query‑based Mask2Former architecture, and test‑time augmentation. Our experiments show that query‑based segmentation significantly outperforms the baseline model. Furthermore, increasing the crop size used during training yields substantial gains, highlighting the relevance of preserving scene context for fine‑grained semantic disambiguation. Our final submission, using test‑time augmentation, achieves an mIoU of 69.6% on the challenge test set, providing a strong baseline for fine‑grained semantic segmentation in outdoor environments. To facilitate reproducibility and future research, code and weights will be made publicly available at https://github.com/RoboticsLabURJC/outdoor‑fine‑grained‑segmentation .
Authors:Hanzhi Chen, Anran Zhang, Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers, Oier Mees, Stefan Leutenegger
Abstract:
A central question in robot learning is how to acquire skills from the kinds of data that humans learn from: passive observation, embodied practice, and the experience of failure. Human videos provide the first of these in abundance, and prior work has shown they can initialize useful policies. Far less clear is whether they can support the second and third: whether priors extracted from human videos can ground a robot's own attempts well enough to evaluate them, correct them, and improve from them. In this work, we show that human videos can be used to learn embodiment‑agnostic action, dynamics, and value representations that transfer across robot embodiments, providing the predictive foundation required for robots to autonomously improve from their own rollouts and failures. We introduce Dynamics‑Guided Action Correction (DGAC), a training‑free approach that leverages these adapted models to repair failed states: each failure becomes a query for which the learned models propose and rank corrective actions, turning failures into supervision for the next policy update. Across seven real‑world manipulation tasks spanning both a mobile manipulator and a static manipulator arm, our approach improves success rates from 40% to 81% across multiple policy backbones, demonstrating cross‑embodiment robot self‑improvement from human‑video priors. These results show that human priors and robot failures can be combined to enable scalable autonomous policy improvement. Project page: https://ethz‑mrl.github.io/robot‑self‑improvement‑website/.
Authors:Dwarikanath Mahapatra, Abhijit Das, Behzad Bozorgtabar, Zongyuan Ge, Sudipta Roy, Deepak Nayak, Mauricio Reyes, Imran Razzak
Abstract:
Multimodal medical imaging fuses complementary anatomical and functional information, yet modalities frequently disagree in pathologically heterogeneous regions. Current segmentation models handle this in one of two inadequate ways: deterministic fusion that averages away disagreement, or post‑hoc uncertainty estimation decoupled from the fusion process that produces it. Both obscure the clinically critical question: why is this prediction unreliable? We present EnTrust, a framework that treats inter‑modal conflict as the primary source of predictive uncertainty. Our EnFuse module decomposes multimodal features into three disentangled components: shared anatomical consensus (F_c), modality‑specific cues (F_u,m), and spatially localized conflict signals (F_cf), with independence enforced via a cross‑covariance objective. This structured decomposition conditions SegDiff, a diffusion‑based generative segmentation model whose sampled hypotheses diverge specifically in regions of modal disagreement. TrustMap then translates this hypothesis divergence into calibrated, pixel‑wise uncertainty using ensemble entropy, conflict‑guided perturbation probing, and a learned calibration head, enabling clinicians to understand not only where predictions are uncertain, but why. Across four benchmarks spanning brain, cardiac, lesion, and oncology domains, EnTrust achieves state‑of‑the‑art segmentation accuracy while reducing calibration error by 40% compared to the strongest baseline. Notably, it outperforms 5x deep ensembles using a single model at roughly half the memory footprint. Code and checkpoints are available at https://github.com/GenMI‑Lab/EnTrust.git.
Authors:Nichula Wasalathilaka, Abhijit Das, Imran Razzak, Dwarikanath Mahapatra
Abstract:
Medical image re‑identification (MedReID) enables longitudinal patient linkage but remains vulnerable to shortcut learning and often produces decisions that clinicians cannot audit against named anatomy. We propose Graph‑of‑Differences (GoD), which grounds identity comparisons in explicit anatomical structure. Each image is represented as an anatomy graph whose nodes correspond to named anatomical regions; given an image pair, soft node correspondence is established, and differences are computed over matched anatomy. A graph‑level difference alignment objective ties these anatomy‑matched differences to the global backbone difference, ensuring the retrieval signal is anchored in homologous structures rather than arbitrary spatial tokens. Explanations are defined over named graph nodes and quantitatively audited via node insertion/deletion tests, replacing unstable pixel heatmaps with verifiable structure‑level evidence. On internal benchmarks, GoD improves Rank‑1 by +7.1 pp on fundus and +3.1 pp on CXR over a strong frozen‑backbone baseline, with further gains on zero‑shot external transfers confirming that anatomy grounding improves both accuracy and generalization. Code is available at https://github.com/GenMI‑Lab/GoD.git.
Authors:Zheng Zhang, Lihe Yang, Tianyu Yang, Chaohui Yu, Yixing Lao, Xiaoyang Guo, Biao Gong, Fan Wang, Hengshuang Zhao
Abstract:
We present SCOPE (Scale‑Consistent One‑Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequences, where existing methods struggle to maintain both geometric accuracy and temporal consistency across hundreds of frames. Our approach generates affine‑invariant 3D point maps with shared parameters across entire sequences, enabling consistent scale‑invariant representations. We introduce three key innovations: viewpoint‑invariant geometry aligning multi‑perspective points in a unified reference frame; appearance‑invariant learning enforcing consistency across exponential timescales; and frequency‑modulated positioning enabling extrapolation to sequences vastly exceeding training length. Experiments across diverse datasets demonstrate significant improvements, reducing relative point map error by 24.2% and temporal alignment error by 34.9% on ScanNet compared to state‑of‑the‑art methods. Our approach handles challenging scenarios with complex camera trajectories and lighting variations while efficiently processing extended sequences in a single pass. Project page: https://scope3d.github.io/.
Authors:Adnan Mustafic, Halim Benhabiles, Adnane Cabani, Kristhian André Oliveira Aguilar, Romain Amigon, Clément Bardin, Chiara Bentifece, Marin Boehm, Kévin Bouchard, Laura Burattini, Diedre Carmo, Fahima Idiri, Matthis Lahargoue, Ilaria Marcantoni, Hicham Messaoudi, Cyril Meyer, Farid Meziane, Léon Morales, Letícia Rittner, Agnese Sbrollini, Léonard Zipper, Karim Hammoudi
Abstract:
We propose NoduLoCC2026, a challenge on lung nodule detection and localization in chest X‑ray images. We have provided a dataset for both tasks and received submissions from 5 international teams. The participating teams' solutions are presented in this work along with results on an external dataset used for testing. Proposed methods show good performance on the classification task. The best method shows a balanced accuracy score of 0.72 and AUC‑ROC of 0.79. We highlight the limitations of current approaches for the localization task, with the best approach having predicted the correct number of nodules on 53% of the test images with a median distance of 12.83mm, showing that it is a more challenging task than the first one. The challenge website is available via https://gt‑i2mdp.github.io/website/nodule_challenge.html.
Authors:Talha Ilyas, Deval Mehta, Zongyuan Ge
Abstract:
Video‑based seizure detection is essential for the management of epilepsy patients, offering a non‑invasive complement to electroencephalography. While several deep learning approaches have been developed for video‑based seizure detection, none are inherently interpretable, limiting their adoption and translation into clinical practice. We present, to our knowledge, the first exploration of a neurosymbolic framework for video‑based seizure detection that directly addresses this gap. Our approach (1) extracts patient‑centric skeleton sequences from epilepsy monitoring units via a prompt‑guided foundation model, (2) predicts binary spatio‑temporal concept activations grounded in clinical motor semiology guidelines, and (3) composes them via differentiable logic into interpretable Boolean rules with auditable contributions. Furthermore, to mitigate false positives arising from the traditional binary formulation (seizure vs.\ non‑seizure), we sub‑classify non‑seizure segments into clinically relevant normal activities, providing the model with fine‑grained discriminative supervision. Evaluated on two public seizure video benchmarks, our framework achieves 89.78% sensitivity with 0.06 false detections per hour on SAHZU and 85.27%,0.09 on IEEE, while producing complete three‑level interpretability: every prediction decomposes into which motor primitives were detected, how they were logically composed, and how much each rule contributed to the clinical decision. We publicly release all annotations, extracted pose sequences, our data pipeline and code, https://github.com/Mr‑TalhaIlyas/CDSD/.
Authors:Sergio Lanza, Jae Hee Lee, Stefan Wermter
Abstract:
Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioning and Visual Question Answering (VQA), but our understanding of their internal processes remains limited. Recently, Sparse Autoencoders (SAEs) have emerged as a promising tool to support the interpretation of concepts encoded in VLMs. However, most SAE‑based approaches focus only on textual or visual concepts separately, ignoring multimodal concepts. This limitation hinders a comprehensive understanding of VLMs, since concepts that integrate both modalities can be misclassified. Moreover, previous visual approaches often produce low‑quality visual concept descriptions that are vague or incomplete, limiting their usefulness for understanding model reasoning. We propose a framework based on SAEs to extract and analyze visual, textual, and multimodal concepts from VLMs. For each neuron, we propose a candidate human‑interpretable concept and compute the alignment between the concept and the dataset samples using cosine similarity scores. Experiments on a VQA dataset (LLaVA‑NeXT) demonstrate that our framework improves visual concept quality by up to 45% compared to existing SAE‑based methods, while maintaining high textual concept quality and enabling systematic identification of multimodal concepts. This work contributes new insights into the conceptual space of VLMs, providing a structured approach to distinguish between visual, textual, and multimodal concepts. The code is available at https://github.com/PHDLanza/Multidata_SAE
Authors:Joohyeok Kim, Taejin Jeong, Jinyeong Kim, Seong Jae Hwang
Abstract:
The high cost of spatial transcriptomics (ST) has driven extensive studies into predicting gene expression directly from H&E histology images. However, this prediction task faces an inherent limitation, as tissue morphology alone provides insufficient information to fully resolve underlying gene expression. To address this limitation, a recent study leverages partial gene expression to guide the prediction process alongside histology images. Building on this paradigm, we approach the prediction task as a spatial imputation problem, employing a Masked Autoencoder (MAE) to utilize a small fraction of gene expression as genetic anchors for inferring whole‑slide gene expression profiles. Specifically, we propose a bio‑saliency score and a learning‑to‑rank strategy to adaptively identify the most informative spots within the tissue. Based on these identified spots, our framework selects contiguous regions as genetic anchors to ensure suitability for real‑world ST profiling hardware. To effectively leverage these anchors, we design a cross‑modal joint encoder that integrates visual and genetic modalities. By aligning the selected anchors with their corresponding visual features via contrastive learning, the encoder generates robust joint representations to accurately predict gene expression across the whole slide. Notably, our framework consistently surpasses existing methods in both histology‑only prediction and spatial imputation, achieving superior accuracy even without genetic anchors and further excelling with as little as 10% transcriptomic coverage. Our code is available at https://github.com/Kyyle2114/CAMMST.
Authors:Kahim Wong, Kemou Li, Yiming Chen, Haiwei Wu, Jiantao Zhou
Abstract:
AI‑assisted image editing threatens trust in financial, legal, and identity records. The GenText‑Forensics Challenge at ACM MM 2026 addresses this by requiring structured forensic reports, in which integrating detection, pixel‑level localization, and natural language explanation for multilingual text‑centric forgery images. We present SEED, a modular system with three components. First, a similarity‑guided pipeline augments training with diverse synthetic forgeries. Second, a single ViT, built on DINOv3 with LoRA adaptation, jointly performs detection and pixel‑level localization while preserving pre‑trained priors with minimal trainable parameters. Third, an evolving harness takes the detector's predictions and generates a complete forensic report via an MLLM, iteratively improved through a proposer‑evaluator loop optimizing report quality. SEED ranked 3rd in the GenText‑Forensics Challenge. Code and data are available at https://github.com/KahimWong/GenText‑Forensics‑3rd‑Place.
Authors:Jeff Brown, Tim Farkas, Gleb Razgar, Edward S. Boyden
Abstract:
Proofreading‑‑correcting segmentation errors in 3D brain reconstructions‑‑is the rate‑limiting step in synapse‑resolution connectomics. We release ConnectomeBench2, a unified multi‑species dataset of over 716,485 expert‑labeled proofreading decisions with >4,500,000 associated images spanning four major open connectomes (mouse, human, zebrafish, fly), spanning both split and merge error correction. Trained on this dataset, a single Vision Transformer with shared encoders for mesh geometry and electron microscopy reaches human‑level accuracy across species for split error correction and merge error identification, with performance scaling with data size and modality. Beyond accuracy, we show that the model is well‑calibrated within distribution, that measures of distribution distance predict where calibration and accuracy will degrade on unseen data, and that connectomics‑specific pretraining and active learning‑based sample selection show potential to substantially reduce the labeling effort needed to extend to new species and brain regions. The benchmark provides the infrastructure to train and evaluate increasingly capable vision models for connectomic proofreading. Data and code availability. The ConnectomeBench2 dataset is released on Hugging Face at https://huggingface.co/datasets/jeffbbrown2/ConnectomeBench2. The accompanying codebase is available on GitHub at https://github.com/timfarkas/ConnectomeBench2.
Authors:Vasile Marian, Yong-Bin Kang, Alexander Buddery
Abstract:
Object‑centric image generation is important in settings with few labeled examples, including pedestrian analysis in smart‑city scenes, traffic‑sign inspection, and domain‑specific object detection. Synthetic images are most useful for training and evaluation when datasets preserve object structure, bounding boxes, visual diversity, and realistic context. Existing image datasets usually target classification, detection, or scene understanding rather than controlled object‑centric generation and augmentation with limited class‑specific data. We present a shareable collection of three object‑centric dataset resources: Cityscapes‑Pedestrian, TrafficSigns, and COCO PottedPlant. The collection standardizes 256‑by‑256 object‑centric crops and bounding‑box annotations across three regimes: dense pedestrian scenes with privacy blur and occlusion, cleaner high‑contrast traffic signs, and context‑diverse potted‑plant scenes. The release contains 3,009 TrafficSigns samples, 2,156 Cityscapes‑Pedestrian manifest records, and 7,679 COCO PottedPlant manifest records. The larger COCO‑derived manifest preserves contextual and multi‑instance diversity, while equal‑size subsets can be drawn with a fixed random seed for controlled comparisons. The release provides direct TrafficSigns data where redistribution is permitted, together with scripts, manifests, box‑level annotation tables, checksums, and reconstruction documentation for the Cityscapes‑ and COCO‑derived subsets. It is available through the Latzi/object‑centric‑low‑data‑datasets GitHub repository and Zenodo DOI 10.5281/zenodo.20573001. The collection supports label and split inspection, subset creation, reconstruction from upstream data, and evaluation of object‑centric image generation or synthetic‑data augmentation methods on shared records.
Authors:Dong-Hyun Moon, Ju-Hyeon Nam, Sang-Chul Lee
Abstract:
Image forgery localization remains challenging due to diverse manipulation techniques and distribution shifts. Existing forgery localization models achieve high accuracy on benchmarks but often struggle with cross‑domain generalization and robustness. In this paper, we propose SARIF (Segment Anything for Robust Image Forensics), a framework that leverages the Segment Anything Model (SAM), which has a promptable architecture and strong generalization ability. SARIF introduces a feedback‑guided mask decoder and a dual‑encoder design that extracts forgery‑specific information to capture forensic traces while exploiting the SAM architecture. To localize manipulated regions, we design a block‑wise prompting mechanism that derives forgery‑specific cues from residual features between an adapted encoder and its frozen counterpart. These features are fused with the previous mask prompt to drive a feedback‑based mask refinement process, enabling automatic forgery segmentation without manual input. Extensive experiments on standard forgery‑localization benchmarks show that SARIF achieves strong average cross‑dataset performance and robustness to common image corruptions.
Authors:Aryan Das, Koushik Biswas, Moloud Abdar, Vinay Kumar Verma
Abstract:
We introduce UNITY, a Universal‑to‑Specialized adapter for efficient and scalable composite conditioning in diffusion based image generation. Unlike prior methods that train separate adapters for each conditioning modality, UNITY jointly learns shared semantics across multiple conditioning types and subsequently specializes without modifying the underlying architecture. The proposed two stage training paradigm consists of a Universal Stage that captures cross modal representations across all conditioning modalities using half of the total training steps, followed by a Specialization Stage that refines modality specific features using the remaining training budget. At the core of UNITY are the Morphable Attention Flow (MAF) Network and Morph Wrapper modules, which enable channel aware and spatially adaptive feature alignment through learnable flow fields and attention based fusion. This constant complexity formulation supports flexible operation under both single and composite conditioning settings while significantly reducing inference latency and memory consumption. Extensive experiments across multiple datasets demonstrate that UNITY achieves state of the art image fidelity while maintaining superior memory efficiency. Code: https://github.com/arya‑domain/UNITY
Authors:Austin T. Wang, Dongchen Yang, Angel X. Chang
Abstract:
Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for enabling agents to correspond spatial language with objects in the physical world. However, the lack of diverse descriptions at scale prevents models from generalizing beyond simple linguistic patterns. Recent such attempts lack diversity in the constraint types and language used to ground objects. Captioning methods cannot precisely contrast objects, which is important for visual grounding. We therefore propose ViGiL3D++, a scalable, scene‑agnostic method that generates diverse visual grounding queries by combining constraint sampling in scene graphs with the language generation of LLMs. We show that it has greater diversity over existing scaled datasets and improves model performance over several 3DVG benchmarks but also illuminates outstanding limitations of VLMs.
Authors:Qingtao Pan, Kai Ye, Zhihao Dou, Bing Ji, Shuo Li
Abstract:
In multi‑object text‑to‑image (T2I) diffusion, ensuring semantic consistency between textual prompts and generated visual content is crucial for image synthesis. However, such consistency constraint is often underemphasized in the denoising process of diffusion models. Although token supervised diffusion models can mitigate this issue by learning object‑wise consistency between the image content and object segmentation maps, it tends to suffer from the problems of segmentation map bias and semantic overlap conflict, especially when involving multiple objects. In this paper, we propose ELDiff, a new evidential learning‑supervised T2I diffusion model, which leverages the advantages of uncertainty metric and conflict detection to enhance the fault tolerance of unreliable segmentation maps and suppress semantic conflicts, strengthening object‑wise consistency learning. Specifically, a pixel evidence loss is proposed to restrain overconfidence in unreliable labels through evidential regularization, and a token conflict loss is designed to weaken the contradiction between semantics through optimizing a measured conflict factor. Extensive experiments show that our ELDiff outperforms existing training based and train‑free based T2I diffusion models on SD v1.4, SD v2.1, SDXL, SD v3.5, and Qwen‑Image, without requiring additional inference‑time manipulations. Notably, ELDiff can be seamlessly extended to the existing training pipeline of T2I diffusion models. Code can be found at https://github.com/QingtaoPan/ELDiff.
Authors:Abhijit Das, Nichula Wasalathilaka, Yifan Lu, Adinath Dukre, Dwarikanath Mahapatra, Shadab Khan, Imran Razzak
Abstract:
Medical vision‑language models (VLMs) enable zero‑shot clinical image classification, yet reliably detecting out‑of‑distribution (OOD) inputs at deployment remains an open problem. No static scoring method works across all shift types: Maximum Concept Matching (MCM) on FLAIR achieves 76.4% AUROC for far‑OOD but only 42.4% for covariate shifts such as ultra‑wide‑field fundus images, effectively random. We trace this to a structural mismatch: covariate‑shifted inputs are indistinguishable from in‑distribution samples in softmax space, yet occupy distinct regions in the VLM embedding space. To exploit this untapped signal, we propose PROTON (PROtotype‑based Test‑time ONline OOD detection), a lightweight post‑hoc module that maintains an online prototype bank from high‑confidence test predictions and adaptively fuses prototype distance with MCM scoring via stream‑level variance statistics, requiring no model modification, training data, or prompt engineering. On the ophthalmology benchmark FLAIR + FIVES, PROTON improves MCM by +23.9 AUROC on covariate shift, +8.8 on semantic shift, and +8.1 on far‑OOD, making it the only zero‑shot method to improve all three without hierarchical prompts or labeled data. Code is available at https://github.com/GenMI‑Lab/PROTON, and the project page is available at https://genmi‑lab.github.io/PROTON.
Authors:Koichi Namekata, Yash Kant, Zhizheng Liu, Ryan D Burgert, Yuancheng Xu, Kuan Heng Lin, Emmett Steven, Julien Philip, Li Ma, Andrea Vedaldi, Paul Debevec, Ning Yu
Abstract:
Filmmaking demands precise motion control and reference image compositing ‑‑ capabilities that existing methods treat separately. Point‑track‑conditioned image‑to‑video models restrict content insertion to the first frame, while reference‑to‑video models lack fine‑grained spatial‑temporal control over how reference content integrates across frames. We present Go‑with‑the‑Track, which unifies both capabilities by jointly conditioning on multiple reference images and reference‑anchored point‑tracks ‑‑ extending conventional point‑tracks to explicitly establish correspondences between generated frames and reference images, thus enabling precise compositing and motion control throughout the video. To achieve this, we introduce spatially‑aware point‑track embeddings that encode the full sequence of point‑track coordinates using a coordinate‑wise MLP followed by temporal pooling. This representation captures the spatial characteristics of each point‑track (serving as a unique identifier), while the embedding similarity correlates directly with spatial proximity, enhancing the model's ability to distinguish and associate point‑tracks. We inject these point‑track embeddings into a video diffusion transformer via a lightweight adapter, resolving the pixel‑to‑patch resolution mismatch while avoiding the substantial motion detail loss inherent in naive point‑track subsampling. We use a hybrid training strategy to train jointly on dynamic, static, and synthetic scene video datasets to boost motion controllability. Experiments demonstrate that Go‑with‑the‑Track achieves superior motion and reference control in a single model and enables new capabilities: multi‑reference conditioned video generation with point‑track driven compositing, as well as camera control for both static and dynamic scenes. Project Page: https://eyeline‑labs.github.io/Go‑with‑the‑Track/
Authors:Luan Marko Kujavski, Rayson Laroca, Paulo Lisboa de Almeida
Abstract:
As urban areas expand, automatic monitoring of parking lots becomes essential for efficient and sustainable cities. This work proposes a self‑supervised approach for parking spot occupancy recognition that requires no labeled samples from the target parking lot. Building upon a self‑supervised transfer learning fine‑tuning protocol, the proposed training strategy consists of two self‑supervised stages: first on unlabeled generic data and then on unlabeled target‑specific data, followed by supervised fine‑tuning using only generic parking lot labels. We adopt SimCLR with a ResNet‑50 encoder and evaluate the method under a leave‑one‑out cross‑environment protocol on three public datasets: PKLot, CNRPark‑EXT, and PLds. We also introduce a two‑stage deployment strategy in which a Strong General Model is initially deployed, followed by a Specialized Model that incorporates unlabeled images collected during the first N days of deployment in a self‑supervised manner. Experimental results show that the Strong General Model alone outperforms supervised and self‑supervised baselines, achieving an average accuracy of 97.2%, which further improves to 97.8% with the proposed two‑stage strategy. These results demonstrate that self‑supervised learning enables a scalable and labelefficient solution for real‑world parking occupancy monitoring. Our trained models and source code are publicly available at https://github.com/LoanMaikon/Parking‑Spot‑Occupancy‑Recognition.
Authors:Hiroki Sakuma, Masatoshi Okutomi
Abstract:
Multi‑view surface reconstruction is a core problem in computer vision. One prominent line of work represents the surface implicitly as a signed distance field (SDF), optimizing it based on the photometric loss between rendered and observed pixel colors. These approaches typically employ SDF‑based volume rendering to obtain a differentiable relaxation of discontinuous visibility along rays, thereby reducing reliance on silhouette supervision. In this paper, we reformulate SDF‑based volume rendering as probabilistic surface rendering, where each pixel color is modeled as a mixture distribution induced by the random first ray‑surface intersection. To this end, we introduce Stochastic Signed Distance Processes (SSDP), which model the SDF along each ray as a stochastic process, inducing a first‑passage‑time distribution for each ray. We then derive the first‑passage probability for each sampling interval based on Bayesian filtering, together with its practical approximation for parallel rendering. We further show that NeuS, an existing SDF‑based volume rendering method, arises as a special case of our formulation. Experiments on the DTU and MobileBrick datasets demonstrate that our method outperforms baselines in both surface reconstruction and uncertainty quantification, supporting the effectiveness of our first‑passage formulation. Our code is available at https://github.com/skmhrk1209/SSDP.
Authors:Qiuhong Shen, Shihua Zhang, Yue Liao, Qi Li, Zhenxiong Tan, Shizun Wang, Shuicheng Yan, Xinchao Wang
Abstract:
World Action Models (WAMs) are embodied predictive‑action models that make a forecast of the future available to action. Recent WAMs repurpose large video generation models, and a parallel line relies on language or vision‑language backbones without a video‑generation core. This rapid expansion has blurred the boundary among broad world models, video generation models, action‑grounded video world models, Vision‑Language‑Action policies, and WAMs. This survey gives the field a common account. It first clarifies these boundaries, then organizes existing works through two complementary views. The first view asks what each method is required to generate, spanning rendered futures, latent futures, and video‑generation‑free action reasoning. The second view decomposes each method by predictive substrate, backbone, action coupling, and deployment regime. This anatomy supports a unified discussion of interactability, causality, persistence, physical plausibility, and generalization, followed by data, evaluation, and open challenges. Across these axes, a consistent design pattern emerges: WAMs are not simply video generators with action heads, but predictive‑action methods whose design choices trade representational richness against compute, memory, latency, and action‑label cost. The field is moving toward methods that generate less of the future while preserving what control requires. The survey homepage is available at https://world‑action‑models.github.io/.
Authors:Francesca Morandi, Omayma Moussadek, Federico Venturini, Mauro Suardi, Alessandro Banzatti, Francesco Cannarile, Angelo Porrello, Simone Calderara
Abstract:
Open Vocabulary Action Recognition (OVAR) enables the recognition of novel actions by leveraging vision‑language representations, overcoming the limitations of traditional closed‑set approaches. However, achieving robust performance in real‑world scenarios typically requires domain‑specific fine‑tuning, which is often costly and raises privacy and regulatory concerns. In this work, we propose an alternative paradigm that bypasses target‑domain training and recombines knowledge from existing datasets and models. Leveraging model merging and task arithmetic, we extract and combine task vectors from models fine‑tuned on diverse public OVAR datasets. We show that, in out‑of‑distribution settings, the resulting merged model achieves superior zero‑shot generalization to the pre‑trained base model. Code is available at https://github.com/omaymaMoussadek/robust‑ovar
Authors:Shiwen Zhang, Yifan Xu, Haibin Huang, Chi Zhang, Xuelong Li
Abstract:
Given a content reference and a style reference, content‑preserving style transfer requires the model to generate stylized outputs with content and style consistency. We introduced TeleStyle V1 to tackle this problem. However, TeleStyle V1 is trained with photorealistic content reference and artistic style reference, which makes it incapable to cope with artistic content reference and realistic style reference in most cases. In this paper, we designed a Self‑Distillation data synthesis strategy to construct such triplets from TeleStyle V1. Trained with such self‑distilled triplets, our TeleStyle V2 supports Content‑Style references in the forms of Realistic‑and‑Realistic (RnR), Realistic‑and‑Stylized (RnS), Stylized‑and‑Realistic (SnR), Stylized‑and‑Stylized (SnS). In addition, we found Distribution Matching Distillation could preserve the general text‑guided image editing capability of the foundation model and fix the content consistency degradation caused by SFT process. Through quantitative evaluations, our TeleStyleV2‑QIE‑2509‑DMD performs at least on par with Qwen‑Image‑Edit‑2509‑DMD, demonstrating strong general image editing skills beyond content‑preserving style transfer. We observed the content/style reference order confusion problem in TeleStyle V1 and further introduced prompt enhancer to solve it. TeleStyle V2 uses Qwen‑Image‑Edit's VLM encoder, Qwen2.5‑VL‑7B, to generate content prompt and style prompt for free. TeleStyle V2 could achieve comparable style transfer performance with state‑of‑the‑art commercial model, gemini‑3‑pro‑image‑preview.
Authors:Siang-Ling Zhang, Huai-Hsun Cheng, Tsung-Ju Yang, Yu-Lun Liu
Abstract:
Creating 3D visual illusions, a single 3D mesh that reveals entirely different semantics from various viewing angles, is a fascinating but tough challenge. Existing optimization‑based methods are slow and can produce oversaturated colors. In contrast, naive stitching approaches fail to produce geometrically coherent objects. This results in visible unnatural seams and semantic leaks. In this paper, we present a fast and training‑free framework for generating text‑driven 3D visual illusions. Our approach decouples the generation into two stages. First, we propose a cross‑space dual‑branch denoising process. This process dynamically decodes 3D latents into voxel space for CLIP‑guided orientation alignment and Signed Distance Field (SDF) blending, which ensures seamless geometric fusion. Second, we introduce a view‑conditioned texture synthesis module that projects and aggregates view‑specific 2D diffusion priors onto the fused geometry. Extensive experiments demonstrate that our method generates highly realistic, dual‑semantic 3D illusions in just 3‑5 minutes. It significantly outperforms existing methods in geometric integrity, semantic recognizability, and efficiency. Project page: https://siang1105.github.io/JanusMesh.github.io/
Authors:Shaghayegh Kolli, Timo Cavelius, Nafiseh Nikeghbal, Samantha Dalal, Jana Diesner
Abstract:
Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential settings, yet the visual cues that shape how these models judge people remain poorly understood. Prior work often compares different (groups of) individuals, making it difficult to separate appearance effects from identity differences. We introduce StylisticBias, a controlled benchmark for evaluating attribute‑level social bias in MLLMs. We generate 500 photorealistic base faces and create about 50 single‑attribute variations per face, producing about 25K images. This design keeps identity fixed and changes one visual attribute at a time. It lets us measure how specific cues shift model judgments. We evaluate six MLLMs across 25 binary social judgment scenarios. We find that age and body type dominate identity‑level effects, while fashion style and other visual cues drive the largest attribute‑level shifts. We further find that about 15 attributes account for nearly 80% of the total variation, showing that bias is concentrated in a small set of visual cues. Sensitivity is strongest in judgments that are semantically aligned with appearance, especially socioeconomic and style‑related judgments. We release StylisticBias as a benchmark for fine‑grained bias evaluation in multimodal models. Code and dataset: https://github.com/timo‑cavelius/StylisticBias and https://hf.co/datasets/shaghayegh/stylistic‑bias‑dataset.
Authors:Juncheng Ma, Jianxin Bi, Yufan Deng, Xuanran Zhai, Kewei Zhang, Ye Huang, Bo Liang, Shukai Gong, Jiankai Tu, Xiaotian Tang, Jiaxin Li, Kaiqi Chen, Duomin Wang, Yuqi Wang, Bingyi Kang, Eric Huang, Zhiyang Dou, Zhen Dong, Enze Xie, Wojciech Matusik, Tat-Seng Chua, Daquan Zhou
Abstract:
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real‑robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower‑cost, and more diverse alternative for embodied model pretraining. However, its effectiveness compared to teleoperated real‑robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real‑robot trajectories as pretraining data sources for embodied foundation models, under fixed post‑training and validation protocols. Surprisingly, we find that egocentric data, when processed through a carefully designed filtering and labeling pipeline, is not merely a viable substitute for model pretraining but can lead to superior performance. With the same amount of pretraining data, models pretrained on egocentric data achieve a 24% lower validation loss on real‑robot action prediction, as well as 52.5% and 90% higher success rates on in‑distribution and out‑of‑distribution real‑robot task execution, respectively. This finding verifies a scalable paradigm for embodied foundation models: pretrain on egocentric human video to learn diverse world representations, then adapt with a small amount of labeled real‑robot data for action‑space alignment. We hope this study encourages broader exploration of egocentric data and offers guidance for data quality assessment before costly robot data collection.
Authors:Yalun Dai, Hao Li, Shulin Tian, Runmao Yao, Yuhao Dong, Fangzhou Hong, Zhaoxi Chen, Fangfu Liu, Baoliang Tian, Dingwen Zhang, Tao Wang, Kim-Hui Yap, Ziwei Liu
Abstract:
Real‑world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool‑augmented agents largely remain tied to static, stateless inference from isolated visual observations. We introduce \textscS‑Agent, a spatial tool‑use agentic paradigm for understanding and reasoning over continuous multi‑view images and videos. By formulating spatial reasoning as spatio‑temporal evidence accumulation rather than isolated frame‑level prediction, \textscS‑Agent reshapes spatial perception into scene‑centric understanding beyond frame‑centric recognition. Specifically, \textscS‑Agent casts the VLM as a semantic planner that decides what evidence is needed, while a hierarchy of spatial tools and experts grounds objects in 2D, lifts them into 3D geometric evidence, and aggregates this evidence into high‑level spatial knowledge (e.g., counting, measurement, orientation, and relative position). Additionally, a temporal memory mechanism, including Scene Memory for maintaining the evolving scene state and Agent Memory for accumulating reasoning context, enables evidence integration across frames and reasoning steps. Comprehensive experiments on multi‑view and video spatial reasoning benchmarks show that \textscS‑Agent consistently improves both open‑source and closed‑source VLMs in a training‑free manner. Beyond inference‑time augmentation, supervised fine‑tuning (SFT) on \textscS‑Agent‑generated spatial trajectories \textscS‑300K yields \textscS‑Agent‑8B, a compact spatial agent that significantly surpasses similar‑scale baselines (e.g., Qwen3‑VL‑8B) and performs comparably to advanced closed‑source models (e.g., GPT‑5.4 and Gemini 3).
Authors:Jinghong Lan, Wei Cheng, Yunuo Chen, Ziqi Ye, Peng Xing, Yixiao Fang, Rui Wang, Yufeng Yang, Xuanyang Zhang, Xianfang Zeng, Difan Zou, Gang Yu, Chi Zhang
Abstract:
Style‑content dual‑reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style reference.Despite recent progress, this setting remains challenging because models must balance content fidelity, style alignment, and instruction following avoiding semantic leakage from the style reference.A key bottleneck is the lack of large‑scale triplet data with clean content‑style separation and broad long‑tail style coverage.In this work, we propose FreeStyle, a scalable dual‑reference generation framework based on community LoRA mining.We treat community LoRAs as compositional anchors for style and content, and design a rigorous generation and filtering pipeline to construct large‑scale Style‑Reference and Content‑Reference triplets across multiple base models.To address content leakage, we adopt a two‑stage curriculum with stage‑specific disentanglement mechanisms: an attention‑level enrichment constraint that suppresses style‑reference leakage in the style‑transfer stage, and a frequency‑aware RoPE modulation strategy that targets positional‑correspondence‑based leakage in the harder dual‑reference stage.We also introduce a benchmark covering both style‑reference and dual‑reference generation, with evaluations on style similarity, content preservation, aesthetics, instruction following, and leakage rejection. The benchmark incorporates a style‑invariant Content Alignment Score (CAS) and introduces a calibrated VLM‑based Rejection Score for evaluating generation reliability and leakage suppression.Extensive experiments show that our model achieves a strong balance among style alignment, content preservation, and leakage suppression.
Authors:Ning Dong, Yingna Su, Xin Dong, Ziyun Jiao, Xinnian Guo, Zhuangzhuang Pan
Abstract:
Pose‑flow video anomaly detectors are attractive for one‑class surveillance because they provide likelihood‑based rankings for tracked skeleton windows. However, a single likelihood score may hide multimodal normal behavior and be sensitive to pose‑observation noise. We study a frozen‑detector setting in which the pose‑flow backbone, cached skeleton tracks, and evaluation pipeline are fixed. Reliability‑Aware Prototype Calibration (RPC) is a post‑hoc score calibration method for this setting. It adds a standardized nearest‑prototype deviation in the frozen latent space to the standardized flow score, and uses keypoint confidence only to gate this added geometric evidence. Thus, RPC preserves the original density signal while correcting the ranking with empirical normal‑mode structure under pose reliability. Across two frozen pose‑flow backbones and four datasets, RPC improves frame‑level AUROC in all eight backbone‑dataset pairs, with gains ranging from 0.34 to 4.49 percentage points and averaging 2.03 points. Ablation and reliability analyses show that prototype deviation is the main corrective signal, while reliability gating is most useful when pose observations are less trustworthy. These results suggest that lightweight post‑hoc calibration can strengthen cached pose‑flow systems when retraining or reproducing the full pose pipeline is impractical.
Authors:Giovanni Affatato, Sara Mandelli, Edoardo Daniele Cannas, Paolo Bestagini, Stefano Tubaro
Abstract:
Deepfakes targeting a high‑profile individual, known as Person‑of‑Interest (POI), are a threat to modern democracies and societies. Current POI deepfake detection methods still struggle to combine robustness to post‑processing, efficiency and interpretability, focal aspects of modern deepfake detectors. In this paper we propose CUPID, a POI video deepfake detector that combines UV texture maps, a facial appearance representation derived from 3D face reconstructions, with the representation learning capabilities of the Masked Autoencoder (MAE). Our method does not require any deepfake videos in its training phase. Moreover, it does not even require to include a specific POI in the training set: the combination of UV texture maps extracted from real video frames and the MAE context‑guided reconstruction yields a latent space that captures rich and discriminative facial features also for identities unseen during training. In the testing phase, the embeddings extracted from a query video depicting the POI can be matched against pristine reference videos to assess the video authenticity. Furthermore, operating in the UV space naturally provides an additional layer of interpretability. Specifically, we can extract decoded residual maps that highlight which facial regions of a test video deviate most from the identity representation of the corresponding POI. Experiments on four deepfake datasets show that CUPID outperforms current state of the art on most datasets and achieves the best overall robustness against strong downscaling and compression, providing also substantially faster inference. Our experimental code will be released at https://github.com/polimi‑ispl/CUPID.
Authors:Junhui Li, Jialu Li, Youshan Zhang
Abstract:
Mamba‑based models have emerged as a promising alternative for salient object detection (SOD), offering significant advantages in modeling long sequences. However, existing models often fail to explore contextual information and the depth of the entire architecture. This paper introduces U^2Mamba, a powerful and innovative U‑structured network for salient object detection. We propose multiscale Mamba U‑blocks (MMUBs) that enhance the model depth to improve local feature extraction capabilities. Our newly developed nested U‑structure, incorporating MMUBs, enables the network to integrate various receptive fields from shallow and deep layers, thereby collecting richer contextual information and longer‑range data without being constrained by resolution. Instead of using the traditional deep supervision scheme and top‑level supervised training, we propose a hierarchical training supervision method where the loss is computed at each level during the training process. Extensive experiments demonstrate that U^2Mamba achieves highly competitive performance against state‑of‑the‑art methods. The source code is available at \urlhttps://github.com/JL021/U2Mamba.
Authors:Duc T. Nguyen, Hoang-Long Nguyen, Thanh-Ha DO, Huy-Hieu Pham
Abstract:
Existing weakly supervised semantic segmentation (WSSS) methods in computational pathology rely on a multi‑stage paradigm: class activation map (CAM) generation, offline pseudo‑mask refinement, and fully supervised retraining. While established, this decoupled approach presents fundamental limitations. The multi‑stage process not only incurs high computational training costs but also suffers from error propagation: local texture biases in shallow CNN layers generate false‑positive artifacts that subsequent refinement steps often fail to correct. To address these persistent challenges through a simple yet highly effective approach, we propose the Single‑Stage Hierarchical Rectification (SSHR) framework. Rather than passively refining CAMs post‑hoc, our method proactively purifies intermediate feature representations during the forward pass. We introduce a Hierarchical Feature Rectification Module (HFRM) that utilizes deep global semantic context to filter out local anomalies in shallow layers. This mechanism generates high‑fidelity activation maps directly within a single training loop. Experiments on the LUAD‑HistoSeg and BCSS datasets demonstrate that SSHR outperforms state‑of‑the‑art multi‑stage methods. Furthermore, SSHR reduces training duration by 2 to 5 times. This efficiency minimizes computational overhead and accelerates clinical translation for large‑scale histopathology workflows. The code is available at: https://github.com/trongduc‑nguyen/SSHR
Authors:Bo Yin, Xiaobin Hu, Chengming Xu, Ruolin Shen, Mo Yang, Jiangning Zhang, Peng-Tao Jiang, Cheng Tan, Shuicheng YAN
Abstract:
Vision‑language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evidence readout even when high‑level reasoning is intact. Prior inference‑time visual interventions can improve grounding without retraining, but they are largely open‑loop and lack a mechanism to verify whether highlighted evidence is actually used. We study answer‑span prediction entropy as a model‑internal feedback signal and show that naive entropy minimization is ambiguous, since low entropy may arise from evidence‑grounded confidence or shortcut collapse. To resolve this ambiguity, we introduce low‑entropy anchors and an entropy‑shaping objective that reduces answer uncertainty while preserving baseline high‑confidence tokens. We instantiate this principle in SPOT‑E, a plug‑and‑play test‑time method that produces question‑conditioned spotlights, optimized per instance via light‑weight tuning based on Group Relative Policy Optimization (GRPO). Across all benchmarks and different VLM families, SPOT‑E yields consistent gains and improved robustness under visual corruptions. Code is publicly available at: \urlhttps://github.com/YinBo0927/SPOT‑E
Authors:Hyun-Kurl Jang, Jihun Kim, Hyeokjun Kweon, Kuk-Jin Yoon
Abstract:
Continual Test‑Time Adaptation (CTTA) aims to maintain model performance under evolving target domains by adapting online without labeled data. However, practical deployments often cannot retain the source dataset due to privacy or licensing constraints, and purely source‑free CTTA methods tend to become unstable under long‑term distribution shift, suffering from compounding self‑training errors and catastrophic forgetting. We introduce DO‑ALL (Distill Once, Adapt Life‑Long), a plug‑and‑play framework that revisits source information in a compact and privacy‑conscious form via Dataset Distillation (DD). Before deployment, DO‑ALL performs DD to produce a small set of synthetic distilled anchors that summarize the source distribution. During adaptation, each target sample is matched with its most semantically aligned anchor, which provides a stable reference for various CTTA via source replay, representation alignment, and manifold‑smoothing regularization. DO‑ALL can be seamlessly integrated into existing CTTA algorithms, consistently improving long‑term robustness across CIFAR100‑C, ImageNet‑C, and the CCC benchmark. This demonstrates the potential of leveraging DD to enable stable and continuous adaptation without retaining raw source data. The code is available at https://github.com/blue‑531/DOALL.
Authors:Maciej Wozniak, Jesper Ericsson, Hariprasath Govindarajan, Truls Nyberg, Thomas Gustafsson, Patric Jensfelt, Olov Andersson
Abstract:
Leveraging Vision Foundation Models (VFMs) for camera‑to‑LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real‑world autonomous driving (AD). However, current approaches typically treat VFMs as black‑box teachers, relying exclusively on frame‑wise feature similarity. Consequently, they do not fully exploit the teacher's layer‑wise semantic structure and global context, as well as the rich spatiotemporal information inherent in LiDAR sequences. We propose HilDA, a self‑supervised pretraining framework for LiDAR backbones that better captures the semantic what and geometric where needed for driving tasks. HilDA combines hierarchical distillation comprising multi‑layer distillation for progressive semantic alignment and global context distillation for scene‑level semantics, with a temporal occupancy diffusion objective promoting spatiotemporal consistency. Models pre‑trained with HilDA achieve state‑of‑the‑art results on cross‑modal distillation benchmarks and outperform models trained via prior distillation approaches on 3D object detection, scene flow, and semantic occupancy prediction. Code available at: https://maxiuw.github.io/hilda.
Authors:Tong Wang, Siwen Wang, Yaolei Qi, Jinxing Zhou, Yuting He, Guanyu Yang, Yutong Xie
Abstract:
Imperfectly supervised video polyp segmentation (VPS) aims to learn dense, temporally consistent masks from inexpensive supervision, including weak annotations (points, scribbles) and semi‑supervision with few densely labeled frames. This setting is clinically valuable but challenging due to weak contrast, ambiguous boundaries, motion blur, and specular highlights, compounded by sparse pixel‑level guidance. While SAM2 can generate dense masks from sparse inputs, direct pseudo‑labeling often yields geometry‑degraded masks with boundary leakage, underutilizes temporal consistency, and ignores reliability. To address these issues, we propose ARTEMIS, a unified framework for imperfectly supervised VPS driven by agent‑guided reliability‑aware temporal mask evolution. ARTEMIS initializes coarse masks from available supervision: SAM2 converts points/scribbles, while dense labels serve as reliable anchors. A debate‑and‑judge vision‑language agent selects reliable temporal anchors under weak supervision, which are propagated bidirectionally with SAM2 to refine unreliable or unlabeled frames. Finally, ARTEMIS trains the segmenter using temporal reliability‑aware robust learning, incorporating reliability‑guided reference selection, a Reference Prototype Transport Module, and reliability‑aware robust loss. These components assess mask reliability, evolve anchors over time, transport target identity across frames, and down‑weight noisy supervision instead of discarding difficult samples. Experiments on SUN‑SEG and CVC‑ClinicDB‑612 under scribble, point, and limited‑label settings demonstrate that ARTEMIS achieves state‑of‑the‑art performance. Code will be released at https://github.com/wangtong627/ARTEMIS.
Authors:Zhenkai Zhang, Markus Hiller, Krista A. Ehinger, Tom Drummond
Abstract:
Generating high‑resolution 3D CT volumes with fine details remains challenging due to substantial computational demands and optimization difficulties inherent to existing generative models. In this paper, we propose the Pixel‑Level Residual Diffusion Transformer (PRDiT), a scalable generative framework that synthesizes high‑quality 3D medical volumes directly at voxel‑level. PRDiT introduces a two‑stage training architecture comprising 1) a local denoiser in the form of an MLP‑based blind estimator operating on overlapping 3D patches to separate low‑frequency structures efficiently, and 2) a global residual diffusion transformer employing memory‑efficient attention to model and refine high‑frequency residuals across entire volumes. This coarse‑to‑fine modeling strategy simplifies optimization, enhances training stability, and effectively preserves subtle structures without the limitations of an autoencoder bottleneck. Extensive experiments conducted on the LIDC‑IDRI and RAD‑ChestCT datasets demonstrate that PRDiT consistently outperforms state‑of‑the‑art models, such as HA‑GAN, 3D LDM and WDM‑3D, achieving significantly lower 3D FID, MMD and Wasserstein distance scores.
Authors:Pengwei Wang, José Morano, Qian Wan, Hrvoje Bogunović
Abstract:
Image quality control is vital for a wide range of downstream applications. Deep learning‑based image quality assessment methods typically train classifiers on dataset‑specific quality labels, inheriting two limitations: (1) generalization is tied to the labeling criteria of the training set and (2) these methods cannot provide spatial feedback on where the quality is degraded, lacking explainability. In this work, we propose EFIQA, a framework that requires no quality‑related supervision and produces spatial quality maps by design. Rather than learning ``what is degradation" from human‑annotated labels, EFIQA learns ``what should be there" by leveraging anatomical priors. For fundus photography, we instantiate this as a two‑stage approach, by first training an unsupervised anomaly detector via masked anatomical inpainting to identify regions of missing vasculature, and then distilling this prior knowledge into a shallow adapter mapping features of a frozen foundation model to precise quality maps. External‑dataset evaluation demonstrates that this label‑free approach with minimal adaptation achieves better performance and explainability compared with supervised methods across benchmarks with different quality criteria, highlighting its potential for real‑world applications.
Authors:Xiangchen Yin, Wenzhang Sun, Jiahui Yuan, Zijie Liu, Yinda Chen, Wei Li, Dachun Kai, Chunfeng Wang, Xiaoyan Sun
Abstract:
Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls remain isolated, and weather generation typically relies on a source video or reconstructed scene that already specifies future structure. We study a first‑frame‑anchored source‑to‑state setting, where the model starts from a single image and follows explicit camera and object controls and an optional weather instruction, then generates a video that either preserves the source world or transfers it to a target weather state. To address these challenges, we first build HoloStateData, a state video dataset that turns diverse videos into unified control samples for camera, object, and weather supervision. Second, we introduce Holo‑World, a unified controllable video world model that jointly controls scene from a single image. Its Unified Scene Adapter factorizes world preservation and weather transfer into distinct parameter subspaces, using rendered background, geometry buffers, and object controls to maintain controlled scene structure while modeling weather‑dependent appearance and particle effects. Additionally, Scene‑Weather Decomposed CFG guides scene and weather residuals separately, strengthening target weather effects without over‑amplifying the full condition. Quantitative and qualitative experiments demonstrate that Holo‑World maintains precise camera and object control with consistent scene structure while transferring scenes into diverse target weather state, outperforming video‑to‑video weather editing baselines on weather‑state generation. Our project page is available at \urlhttps://xiangchenyin.github.io/Holo‑World/.
Authors:Dong Hoon Lee, Seunghoon Hong
Abstract:
Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality‑compute trade‑off is largely constrained by the tokenizer's fixed compression ratio. Variable‑length tokenizers (VLTs) promise adaptive compression by varying token counts, allowing diffusion models to flexibly balance quality and compute. However, conventional VLTs modulate length by truncating ordered token sequences, which makes token semantics depend on token position and breaks representational alignment across lengths. This leads to a cross‑length shift in the latent distribution that hinders a single variable‑length diffusion model from operating effectively. To address this, we propose a novel variable‑length tokenizer that modulates length by merging tokens. We show that encouraging similar tokens to merge enables direct cross‑length representation alignment when the diffusion transformer operates according to the merging pattern. Since conventional merging methods are data‑dependent, making the merging pattern inaccessible during generation, we introduce learnable global merging, which is data‑independent, to ensure compatibility with diffusion transformers. On ImageNet 256×256 generation, our merging‑based variable‑length tokenizer integrated with a diffusion transformer achieves a superior gFID‑compute trade‑off compared to prior VLT methods. Code is available at [this https URL](https://github.com/movinghoon/lgm)
Authors:Fanfu Xue, En Yu, Yantian Shen, Zhikun Hu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun
Abstract:
UAV Vision‑Language Navigation (UAV‑VLN) is typically formulated as a holistic search‑and‑reach problem, where long‑range target discovery and final target approach are optimized and evaluated jointly. This formulation makes it difficult to assess a critical capability of aerial embodied agents, namely whether a UAV can accurately ground a visible target and translate vision‑language evidence into precise 3D motion once the target enters its field of view. To address this limitation, we introduce UAV‑VLN‑FOV, a target‑visible navigation task that isolates the see‑and‑reach stage and enables a more diagnostic evaluation of terminal reaching ability. We further propose 3DG‑VLN, a vision‑language waypoint prediction framework guided by dynamic 3D direction cues to enhance fine‑grained visual grounding and spatial direction alignment for precise target reaching. Specifically, 3DG‑VLN adaptively processes high‑resolution front‑view and downward‑view observations to preserve fine‑grained visual and geometric details for target grounding. It also updates the target‑relative direction online during closed‑loop navigation, allowing the agent to maintain spatial alignment with the target and reduce accumulated direction drift. To support this task, we construct a dedicated high‑resolution benchmark which contains 2,717 trajectories with target‑oriented high‑level instructions, high‑resolution front‑view and downward‑view egocentric observations, and continuous 3D waypoint annotations. Experiments show that 3DG‑VLN outperforms competitive UAV‑VLN baselines, achieving a 13.82% improvement in success rate. Real‑world trials further demonstrate the potential of 3DG‑VLN for practical see‑and‑reach navigation. The source code and benchmark are available at https://github.com/xuefanfu/3DG‑VLN.
Authors:Hongming Zhu, Huaji Chen, Bowen Du, Sicong Liu, Qin Liu
Abstract:
Unlike traditional remote sensing change detection that relies on predefined categories, Open‑Vocabulary Change Detection (OVCD) identifies land cover changes flexibly using arbitrary text prompts. However, existing methods suffer from an inherent trade‑off when modeling changes: instance‑level comparison overlooks fine‑grained semantic variations (e.g., partial building extensions), while direct pixel comparison proves unreliable, yielding unstable responses and boundary artifacts due to semantic ambiguity and spatial inconsistency. To this end, we propose an efficient training‑free Reliability‑Aware Open‑Vocabulary Change Detection (ReA‑OVCD) framework. It first derives candidate change regions from pixel‑wise semantic discrepancies to ensure flexible and detailed localization. To ensure reliability, it subsequently introduces a collaborative refinement strategy to explicitly model change validity from both semantic and spatial perspectives. Specifically, we develop a Semantic Change Reasoning (SCR) module that reassesses changes by jointly analyzing distributional divergence and response variation, enabling the suppression of incidental inconsistencies while preserving reliable semantic shifts. In addition, a Boundary‑aware Change Refinement (BCR) module is designed to mitigate artifacts stemming from boundary misalignment and uncertainty through validating whether candidate regions are supported by reliable interior pixels. Extensive experiments across multiple datasets (LEVIR‑CD, WHU‑CD, DSIFN, and SECOND) demonstrate that our method consistently outperforms state‑of‑the‑art approaches, achieving \mathrmF_1^C improvements of 2.13% to 9.75% with higher computational efficiency. The code is publicly available at \https://github.com/Funny0101/ReA‑OVCD
Authors:Luca Zedda, Davide Antonio Mura, Cecilia Di Ruberto, Maurizio Atzori, Muhammed Furkan Dasdelen, Carsten Marr, Andrea Loddo
Abstract:
Attention‑based Multiple Instance Learning aggregators in medical imaging are prone to attention concentration, producing overconfident and unstable predictions. We introduce QG‑MIL, a gated transformer aggregator that addresses this through four synergistic architectural components: RMSNorm‑based pre‑normalization, per‑head QK normalization, fine‑grained attention output gating, and SwiGLU‑style feed‑forward modules. Together, these design choices stabilize training and distribute attention more uniformly across instances without auxiliary losses, masking, or multi‑stage regularization. We evaluate QG‑MIL across six benchmarks spanning whole‑slide pathology and cell‑level hematology, covering two fundamentally different MIL scales. The best‑performing QG‑MIL variants outperform leading baselines on all six benchmarks, with an average improvement of +6.1 mean macro F1 points. Attention overlays and attention mass analysis confirm more distributed instance weighting. Ablation studies show that while individual components can match the full model on specific datasets, the QG‑MIL design provides the most consistent cross‑domain performance and tightest variance when compared to selected baselines. We release a configurable implementation to support reproducibility at: https://github.com/unica‑visual‑intelligence‑lab/QG‑MIL
Authors:Chengwen Liu, Hao Peng, Jisheng Dang, Hong Peng, Bin Hu, Tat-Seng Chua
Abstract:
In multimodal video reasoning, reinforcement learning‑based methods typically rely on simplistic and inflexible reasoning‑length control strategies that fail to adapt to the model's evolving competence. This mismatch may suppress necessary exploration at early stages, while encouraging redundant reasoning and inefficient decoding once the model becomes more competent. In this paper, we propose CARE, a competence‑aware reward shaping framework for adaptive reasoning length optimization in multimodal reasoning. Specifically, CARE maintains a smoothed competence estimate via an exponential moving average of pass rates, and uses it to route training into progressive stages that shift the reward preference from exploration‑oriented long‑form reasoning to efficiency‑oriented concise reasoning. To avoid conflating verbosity with intrinsic task complexity, CARE further normalizes reasoning effort with batch‑level statistics, and introduces a posterior amplifier to strengthen reward signals for unexpectedly strong performance on historically difficult samples. The proposed mechanism is seamlessly integrated into the GRPO training pipeline and incurs no additional inference‑time overhead. Extensive experiments on multiple video reasoning and general video understanding benchmarks demonstrate that CARE consistently improves reasoning accuracy, stabilizes reinforcement learning, and significantly enhances token efficiency. Moreover, CARE exhibits a characteristic inverted‑U trajectory of reasoning length during training, and yields shorter yet more informative reasoning traces at convergence, indicating effective adaptive allocation of reasoning budget. We provide the source code for our proposed CARE framework and experiments at https://github.com/1Pansy/Video‑CARE.
Authors:Mingyu Choi, Woo Kyoung Han, Sunghoon Im, Kyong Hwan Jin
Abstract:
Linear recurrent unit (LRU), designed with a principled formulation for stable linear recurrence, has demonstrated promising accuracy and robustness on long‑range dependency tasks. However, its static parameterization and single‑scan method limits its applicability to 2D vision tasks. In this study, we propose a LRU‑based restoration network with a semantic modulating unit (SMU) to achieve a harmonious balance between performance and efficiency in single‑image super‑resolution. The SMU plays three key roles: LRU modulation, spatial categorization, and feature enhancement through learned prototype. Extensive experiments demonstrate that our method quantitatively and qualitatively surpasses recent state‑of‑the‑art methods. Notably, our approach achieves superior performance with computational complexity on par with existing methods. The source code and models are available at https://github.com/MingyuChoi‑run/LSM
Authors:Jiwoong Yang, Haejun Chung, Ikbeom Jang
Abstract:
Multi‑view imaging, such as mammography and chest radiography, is a standard component of clinical practice. However, medical images are often unregistered and contain view‑specific artifacts or irrelevant background cues that can obscure diagnostically relevant findings. Many existing methods directly fuse per‑view representations, allowing such irrelevant content to contaminate the fused embedding and reducing robustness under varying view configurations. We propose OTCHA, a confidence‑aware latent hub token alignment module based on optimal transport (OT) that refines patch tokens before fusion for multi‑view classification. OTCHA introduces a set of learnable latent hub tokens shared across views. For each view, we compute an OT plan between patch tokens and hub tokens that jointly considers feature similarity and geometry, and augment the OT formulation with token‑conditional dustbins to enable partial matching and discard irrelevant tokens. The resulting transport plan provides token‑wise matching confidence, which gates hub‑mediated message passing and weights a novel optimal‑transport‑based representation alignment loss to stabilize refinement. Experiments on three multi‑view medical image datasets demonstrate consistent improvements over competing baselines across diverse anatomies and view configurations. Our code is available at https://github.com/labhai/OTCHA.
Authors:Junho Moon, Haejun Chung, Ikbeom Jang
Abstract:
Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains challenging due to low contrast, frequent discontinuities, and severe class imbalance. Although recent convolutional and Transformer‑based models have improved performance, they often yield fragmented predictions and fail to recover fine branches. We propose CSWinUNETR, a general‑purpose backbone for 2D and 3D thin‑structure segmentation. It employs cross‑shaped stripe self‑attention to model long‑range principal‑axis context and incorporates cyclic shifts to enhance information exchange across stripes. To better preserve fine‑grained details, we further introduce a detail‑enhanced multi‑scale self‑attention module that aggregates contextual features from multi‑resolution representations. In addition, we propose sparse‑control dynamic snake convolution, which reconstructs reliable dense curvilinear kernels from sparsely predicted control points to better follow tortuous geometry. Extensive experiments on four benchmarks across ophthalmology, neurovascular imaging, and dermatology demonstrate that CSWinUNETR consistently outperforms state‑of‑the‑art methods without task‑specific post‑processing or topology‑aware losses. The code is available at https://github.com/labhai/CSWinUNETR.
Authors:Tai Hyoung Rhee, Dong-Guw Lee, Ayoung Kim
Abstract:
Thermal infrared (TIR) imaging has been a popular choice for field robotics due to its robust perception capability under low light visual degradation, but it suffers from severe stochastic and fixed‑pattern noise that breaks downstream estimation. This noise is intensified indoors due to low thermal contrast and uniform temperature distributions, contributing to the relative lack of indoor TIR deployments. Existing TIR denoising methods exhibit a poor accuracy‑efficiency tradeoff, either too slow for online deployment required in robotics or insufficiently robust to severe degradation, while typically being trained on synthetic noise. Addressing these problems, we propose TIDY, a lightweight wavelet‑domain denoiser trained on real clean‑noisy TIR data. By reformulating TIR denoising in the wavelet domain, TIDY explicitly disentangles noise from structural content, enabling targeted suppression with reduced spatial complexity, significantly improving inference speed over prior methods (~34Hz). TIDY introduces two new metrics, Wavelet Entropy and Wavelet Directional Stripe Index, as complementary loss terms to explicitly suppress stochastic noise and stripe artifacts. Across severe indoor corruption and zero‑shot settings, TIDY improves robustness and yields consistent gains in downstream robotics tasks including thermal inertial odometry and monocular depth estimation. Code and dataset is available at: https://github.com/williamrheeth/TIDY
Authors:Victoria Wu, Nima Hashemi, Hooman Vaseli, Christina Luong, Purang Abolmaesumi, Teresa S. M. Tsang
Abstract:
Echocardiography (echo) is a widely used imaging modality for assessing cardiac function, with Left Ventricular Filling Pressure (LVFP) serving as a critical physiological marker for conditions such as heart failure. Standard LVFP classification into normal \emphvs elevated categories relies on the Doppler‑derived E/e' ratio, which is operator‑dependent and often unavailable in resource‑limited settings, motivating methods that infer LVFP directly from B‑mode echo. Existing deep learning approaches achieve high performance but remain largely black‑box, limiting clinical interpretability. We propose HypOProto, a hyperbolic, ordinal prototype‑based framework for interpretable LVFP classification using a frozen, explainable foundation model backbone. HypOProto arranges prototypes along the physiological E/e' scale, placing borderline cases near the hyperboloid root where small angular differences separate similar cases, while normal and elevated cases occupy outward positions reflecting increasing diagnostic certainty. This hyperbolic geometry encodes clinically meaningful ordinal relationships and improves interpretability. We also introduce a novel Hyperbolic Prototype Angular Separation (HyperPAS) loss, enforcing inter‑class prototype separation in hyperbolic space. HypOProto achieves SOTA performance while maintaining transparency, and highlights clinically relevant regions in visualizations. This work represents the first prototype‑based framework for LVFP classification in echo. Our code can be found at https://github.com/DeepRCL/HypOProto.
Authors:Shenjian Gong, Kangkan Wang, Shanshan Zhang, Jian Yang
Abstract:
This paper addresses the challenge of one‑shot novel view and pose human image synthesis. The existing methods transfer the reference human image to a target pose using a set of 2D pose keypoints or synthesize human images based on generalizable human NeRF which uses human model priors to extract point‑wise features. However, pose transfer based methods can not handle complex human pose using ambiguous 2D pose as the condition, while generalizable human NeRFs may be inaccurate to recover occluded/invisiable human parts without extracted reliable features. To solve these problems, we propose a novel approach for novel view and pose synthesis from a singe human image via conditional denoising diffusion model. Our diffusion model divides the novel view and pose synthesis problem into a sequence of conditional denoising steps. Specifically, to generate humans with complex and arbitrary poses, we introduce 3D human priors, i.e., 3D normal map and color prompt, as geometry and color conditions into the generation process. By transferring the reference human into the target human with a series of diffusion steps, our diffusion model enables high‑quality synthesis including the occluded/invisible parts. Further, we propose a self‑reconstruction based customized refinement to enhance fine details when tested on novel persons.Experimental results on different public datasets demonstrate that our approach significantly outperforms previous methods and also shows better generalization ability across datasets. The code will be made publicly available at https://github.com/Yankeegsj/3DPGDM.
Authors:Bingshuo Qian, Xiang Cheng
Abstract:
Multi‑representation diffusion models can improve visual synthesis by denoising complementary views of an image, but their performance depends critically on the asynchronous schedule that determines when each representation is denoised. We propose to learn this schedule. Our method formulates asynchronous flow matching over multiple representation spaces and uses a schedule‑corrected objective that keeps each representation's local noising‑time weights fixed as the schedule changes. We instantiate the schedule with a flexible parametric class that is convex and monotone by construction, and learn it using a fast joint probe with less than 1% additional training compute. On ImageNet 256x256, the learned schedule substantially improves both convergence speed and final quality under a matched 675M‑parameter XL backbone. With AutoGuidance, our 200‑epoch model reaches FID 1.05, matching the 800‑epoch SFD‑XL baseline with 4x less training. Training to 600 epochs further improves to FID 1.02, outperforming the 1B‑parameter SFD‑XXL result of FID 1.04 while using a smaller model. In the unguided setting, our 200‑epoch model reaches FID 2.37, already below the best 800‑epoch SFD‑XL result (2.54) at 4x less training, and improves to FID 2.14 at 600 epochs. Code is available at https://github.com/bsq532087/LWD
Authors:Yueyi Sun, Yuhao Wang, Jason Li, Ye Tian, Tao Zhang, Jacky Mai, Yihan Wang, Haochen Wang, Jinbin Bai, Ling Yang, Yunhai Tong
Abstract:
Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that require captioning multiple regions. In this work, we propose PerceptionDLM, a multimodal diffusion language model optimized for efficient parallel region perception. Built upon PerceptionDLM‑Base, a strong foundational baseline that achieves state‑of‑the‑art performance among open‑source diffusion MLLMs, our architecture fully leverages the parallel decoding nature of DLMs. Specifically, we introduce efficient prompting and structured attention masking to enable simultaneous perception of multiple masked regions, allowing the model to generate region descriptions in parallel at both the sequence and token levels. This design significantly improves inference efficiency compared with existing approaches that process regions sequentially. To systematically evaluate the parallelism property of visual perception capability for DLMs, we construct a new Parallel Detailed Localized Captioning Benchmark (ParaDLC‑Bench) by scaling the DLC‑Bench to include multiple region masks per image, enabling joint evaluation of both caption quality and inference efficiency. Experiments demonstrate that PerceptionDLM maintains competitive performance in region captioning while achieving substantial speed improvements for multi‑region perception tasks. Our results highlight the potential of multimodal diffusion language models for efficient, parallel visual perception. To the best of our knowledge, we are the first to achieve parallel region caption and perception by leveraging the advantages of diffusion language models. Code, models, and datasets are released.
Authors:Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang, Haitao Lin, Jingbo Zhang, Yao Mu, Xiaokang Yang, Wenjun Zeng, Xin Jin
Abstract:
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video‑based WAMs face three coupled limitations: dense multi‑frame future tokens make inference costly, full video prediction spends capacity on action‑irrelevant temporal and appearance details, and long‑horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better‑matched prior: it only needs to model a target‑frame transformation, focuses on action‑relevant current‑to‑target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow‑matching action expert on the KV caches produced by image‑editing denoising, using them as a compact world‑action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real‑world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video‑based WAMs. Attention analysis further shows that editing caches focus on task‑relevant change regions, supporting image editing as an effective alternative to video‑based world‑action modeling.
Authors:Shariq Farooq Bhat, Niloy J. Mitra, Kalyan Sunkavalli
Abstract:
Precise 3D spatial orchestration in text‑to‑video generation remains a significant challenge, particularly for multi‑object scenes where semantic layout and temporal dynamics are often entangled. While existing depth‑conditioned models achieve good structural fidelity, they necessitate dense, frame‑accurate guidance that is labor‑intensive to author for dynamic events involving deformable objects. We present LooseControlVideo, a framework that enables intuitive and expressive control by using sparse, oriented 3D boxes as a "blocking" proxy. This allows users to author high‑level layout and trajectory while leveraging a video generative model to generate realistic occlusions, dynamics and interactions. We achieve this by fine‑tuning a Wan 2.2 backbone on a video dataset annotated with DNOCS, a novel encoding for 3D size, orientation and depth‑ordered occlusions. Furthermore, our method allows for localized refinement, such as adjusting a jump trajectory or adding an interaction, with minimal disruption to the global scene context. Extensive evaluations on the nuScenes, HO‑3D, and BEHAVE benchmarks demonstrate that LooseControlVideo significantly outperforms existing 2D‑box and flow‑based baselines. Our findings indicate a 1.2x to 3x improvement in Trajectory Error; 2x improvement in Rigid Motion Consistency; and a 1.5x to 2x increase in Occlusion Accuracy over current state‑of‑the‑art layout‑conditioned models, demonstrating that oriented 3D primitives provide good geometric prior for complex, multi‑agent video authoring.
Authors:Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero
Abstract:
Vision Foundation Models (VFMs) with Vision Transformer (ViT) backbones, such as DINOv2, have become essential for downstream tasks like object recognition and semantic segmentation. The immense computational requirements of backbones often necessitate distillation into smaller architectures for edge deployment. Feature‑based knowledge distillation (KD) often suffers from the teacher‑student gap; the student struggles to imitate teacher's complex feature map due to its limited capacity. To mitigate this bottleneck, we propose LEAP: Layer‑skipping Efficiency via Adaptive Progression, a training curriculum for ViT feature‑based knowledge distillation. By utilizing the teacher's intermediate feature maps as a sequence of progressively more difficult targets, our curriculum allows the student to build a foundational representation before tackling higher‑level abstractions. Our results demonstrate that this paradigm significantly accelerates convergence through adaptive difficulty selection across various student model sizes and dataset scales. With our curriculum, the LEAP‑distilled ViT‑S achieves 90.1% accuracy on ImageNet‑100, a +12.24% improvement compared with baseline. On ImageNet‑1K, LEAP achieves +3.84% and +7.75% improvement for the instance retrieval task on the Oxford and Paris datasets, respectively. Furthermore, the curriculum enables 25.1% savings in training FLOPs and 21% savings in training time on ImageNet‑100 by implementing early‑stopping for teacher inference during the initial stages of training. Code is available at https://github.com/KevinZ0217/LEAP
Authors:Ellina Zhang, Madhaven Iyengar, Amir Zadeh, Chuan Li, Deepak Pathak, David Held, Tal Daniel
Abstract:
We introduce 3D‑DLP, a self‑supervised object‑centric representation learning model that decomposes scene‑level RGB‑D or voxel observations into a set of 3D latent particles. Building on the Deep Latent Particles (DLP) framework, each particle encodes disentangled attributes, including 3D keypoint position, bounding box dimensions, and appearance features, and represents a distinct entity in the scene. The model learns interpretable per‑particle segmentation maps through an end‑to‑end self‑supervised reconstruction objective. We demonstrate on both simulated and real‑world datasets that the learned latent space is interpretable and controllable: by manipulating particle positions and decoding, we can generate novel scene configurations. Furthermore, we show that leveraging these compact 3D latent particles for downstream robotic manipulation improves performance over baselines that either lack explicit 3D information or rely on memory‑intensive dense 3D inputs without object‑centric structure. Code and videos are available at https://eubooks3003.github.io/3d‑dlp.
Authors:Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He, Ziyang Ma, Qize Yang, Yunfei Chu, Jin Xu, Junyang Lin, Chi-Wing Fu, Pheng-Ann Heng
Abstract:
Passive models for long video understanding typically rely on a "watch‑it‑all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although interactive frameworks have emerged, they often rely on global pre‑scanning, and their context cost still scales with video length. We propose OmniAgent, the first native omni‑modal agent that formulates video understanding as a POMDP‑based iterative Observation‑Thought‑Action cycle. OmniAgent executes on‑demand actions to selectively distill audio‑visual cues into a persistent textual memory, effectively decoupling reasoning complexity from raw video duration. To operationalize this, we introduce (1) Agentic Supervised Fine‑Tuning to bootstrap native active perception via best‑of‑N trajectory synthesis with dual‑stage quality control, and (2) Agentic Reinforcement Learning with TAURA (Turn‑aware Adaptive Uncertainty Rescaled Advantage), which leverages turn‑level entropy to steer credit assignment toward pivotal discovery turns. Crucially, OmniAgent exhibits positive test‑time scaling, where performance improves as the number of reasoning turns increases, validating the efficacy of active perception. Empirical results across ten benchmarks (e.g., VideoMME, LVBench) demonstrate that OmniAgent achieves state‑of‑the‑art performance among open‑source models. Notably, on LVBench, our 7B agent outperforms the 10× larger Qwen2.5‑VL‑72B (50.5% vs. 47.3%).
Authors:Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
Abstract:
Existing multi‑speaker dialogue systems bind speakers to utterances through structured supervision: per‑turn tags, multi‑stream transcriptions, or learnable speaker embeddings. These systems operate within speech‑only pipelines that produce clean vocal sequences without the ambient texture of real conversations. We take a different approach. Our method, ScenA, conditions a text‑to‑audio flow‑matching foundation model, pretrained on large‑scale in‑the‑wild data, directly on multiple reference voices and a free‑form natural language prompt that describes an entire multi‑speaker audio scene. Leveraging such a foundational model allows us to inherit its capacity for natural, non‑studio audio: background noise, room acoustics, overlapping dialogue, and spontaneous paralinguistic events, while adding multi‑speaker control without any per‑turn structure. Concretely, reference latents are concatenated into the model's token sequence and distinguished by lightweight identity‑aware positional encodings. However, we identify a critical obstacle to this approach: the Reference Shortcut. During training under standard noise schedules, the model can identify the matching reference by acoustic similarity to the noisy target, bypassing the text prompt entirely. We address this with a high‑noise‑biased timestep distribution that forces the model to rely on the text prompt for speaker assignment. We evaluate ScenA on the CoVoMix2‑Dialogue benchmark, showing that it outperforms existing multi‑speaker systems on speaker‑binding metrics while generating rich conversational audio with overlapping speech, emotional vocalizations, and ambient sound. Our results demonstrate the advantage of using a general‑purpose audio model conditioned on a free‑form scene description, rather than passing structured dialog scripts through a speech‑only pipeline.
Authors:Chong Bao, Yuan Li, Bangbang Yang, Yujun Shen, Hujun Bao, Zhaopeng Cui, Yinda Zhang, Guofeng Zhang
Abstract:
Recently neural implicit rendering techniques have evolved rapidly and demonstrated significant advantages in novel view synthesis and 3D scene reconstruction. However, existing neural rendering methods for editing purposes offer limited functionalities, e.g., rigid transformation and category‑specific editing. In this paper, we present a novel mesh‑based representation by encoding the neural radiance field with disentangled geometry, texture, and semantic codes on mesh vertices, which empowers a set of efficient and comprehensive editing functionalities, including mesh‑guided geometry editing, designated texture editing with texture swapping, filling and painting operations, and semantic‑guided editing. To this end, we develop several techniques including a novel local space parameterization to enhance rendering quality and training stability, a learnable modification color on vertex to improve the fidelity of texture editing, a spatial‑aware optimization strategy to realize precise texture editing, and a semantic‑aided region selection to ease the laborious annotation of implicit field editing. Extensive experiments and editing examples on both real and synthetic datasets demonstrate the superiority of our method on representation quality and editing ability. Project page: https://zju3dv.github.io/neumeshplusplus/
Authors:Bartłomiej Baranowski, Dave Zhenyu Chen, Matthias Nießner
Abstract:
Existing approaches to 3D scene understanding in Vision‑Language Models (VLMs) either rely on complex, model‑specific geometry encoders or large training budgets in pursuit of spatial reasoning. Instead, OneCanvas aggregates patch features from all views onto a single equirectangular panoramic canvas. Namely, each patch is unprojected to a 3D world coordinate using its depth and camera pose, then placed on the canvas at the continuous longitude and latitude of that point as seen from the canvas origin, with no rasterization or aggregation across overlapping views. A 3D position embedding of the patch's metric coordinates is added to its feature, restoring the depth lost when collapsing the world position to an angular canvas coordinate. Patches from all frames thus share one spatial coordinate system with no fusion or major architectural modifications of the backbone. The pretrained VLM consumes this representation as if it were an ordinary image. Because the canvas can be centered on any pose of interest, the same representation directly supports situated reasoning from a specific viewpoint, a common requirement in robotics and embodied AI. Thanks to this representation, we can also introduce a spatial pretraining curriculum: by procedurally placing patch features of objects, drawn from real images, at chosen 3D world positions on an otherwise empty canvas, we generate on‑the‑fly supervision spanning a broad range of spatial reasoning tasks, with answer distributions controlled to reduce spatial reasoning shortcuts. OneCanvas achieves state‑of‑the‑art accuracy on SQA3D and VSI‑Bench, and generalizes to out‑of‑distribution data on SPBench, using an order of magnitude less training compute than the strongest competing methods.
Authors:Kangsheng Duan, Ziyang Xu, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, Xinggang Wang
Abstract:
While 10B‑level industrial foundation models have pushed the boundaries of image inpainting, their prohibitive computational costs severely hinder practical deployment. Constructing a highly optimized task‑specific specialist offers a promising solution; however, extreme structural compression inevitably triggers a severe representation bottleneck. To conquer this, we propose Moebius, a highly efficient lightweight inpainting framework. We systematically reconstruct the diffusion backbone by introducing the Local‑λ Mix Interaction (LλMI) block. Comprising Local‑λ and Interactive‑λ modules, it elegantly summarizes spatial contexts and global semantic priors into fixed‑size linear matrices, preserving complex latent interactions while drastically shedding parameters. Furthermore, to unlock the full representational capacity of this highly compact architecture, we synergistically pair it with an adaptive multi‑granularity distillation strategy. Operating strictly within the latent space to avoid expensive pixel‑space decoding, this strategy dynamically balances multiple gradient‑based losses to achieve high‑fidelity alignment. Extensive experiments across natural and portrait benchmarks demonstrate that this optimal synergy enables Moebius to rival or even surpass the generation quality of the 10B‑level industrial generalist FLUX.1‑Fill‑Dev. Remarkably, Moebius achieves this using less than 2% of the parameters (0.22B vs. 11.9B) while delivering a >15× acceleration in total inference time, setting a new efficiency standard for high‑fidelity inpainting. Project page at https://hustvl.github.io/Moebius.
Authors:Jeongmin Bae, Seoha Kim, Marc Pollefeys, Mahdi Rad, Youngjung Uh, Taein Kwon
Abstract:
Dynamic 3D hand reconstruction from egocentric videos is essential for next‑generation computing platforms such as AR/VR and AI glasses. Despite its importance, most prior works focus either on multi‑view 3D hand reconstruction or on 4D human body reconstruction. Egocentric 4D hand reconstruction remains challenging due to fast head motion, rapid hand dynamics, severe occlusions, and inherent ambiguity from single‑view observations. To address these challenges, we introduce Hand‑4DGS, the first feed‑forward framework for reconstructing dynamic 4D hands directly from egocentric videos, enabling both fast (~60 FPS) inference and strong generalization. Our approach incorporates a mesh‑guided representation for structural priors and temporal convolutions to model dynamic motion. We evaluate our framework on two challenging egocentric datasets, H2O and ARCTIC, and demonstrate significant improvements over baselines. Our method benefits from the generalization capability of feed‑forward networks and effective 2D image supervision through Gaussian splatting, without requiring expensive 3D hand pose ground‑truth annotations.
Authors:Jiayi Gao, Qingchao Chen, Yuxin Peng, Yang Liu
Abstract:
Current image editing methods excel at static attributes but fail at complex Human‑Object Interactions (HOI), a critical challenge unaddressed by existing benchmarks that conflate HOI with static attributes, relying on global metrics incapable of simultaneously assessing dynamic interaction validity and entangled human‑object pair preservation. Thus, we first introduce HOI‑Edit, a comprehensive benchmark with three progressive cognitive levels, which features an automated metric HOI‑Eval that reliably evaluates instance‑level interaction by letting VLM Q&A after thinking with images containing grounded Human‑Object pairs. Considering the task's essence of remodeling dynamic relationships, we benchmark Image‑to‑Video (I2V) models, finding them inherently suited for dynamic editing due to their temporal generation capabilities. Crucially, beyond superior performance, this capability provides a "replay of the failure process," offering unique diagnosability into why errors occur. We thus propose SCPE (Self‑Correcting Process Editing), a novel, agentic self‑correcting framework that constrains the generation of I2V models through iteratively refined prompts, enabling the generated videos to more accurately present the target HOI. Extracted frames from these videos are the final editing results. On HOI‑Edit, SCPE achieves performance competitive with state‑of‑the‑art (SOTA) editing models like Nano Banana on interaction. Code is available at https://github.com/oceanflowlab/HOI‑Edit.
Authors:Hong-Tao Yu, Chen-Wei Xie, Yuxin Peng, Serge Belongie, Xiu-Shen Wei
Abstract:
Recent advancements in Large Vision‑Language Models (LVLMs) have demonstrated remarkable multimodal perception and reasoning capabilities. While numerous benchmarks have evaluated LVLMs from holistic or task‑specific perspectives, their capabilities on fine‑grained image tasks‑fundamental to computer vision‑remain insufficiently understood. To address this gap, we introduce FG‑BMK, a comprehensive fine‑grained evaluation benchmark containing 1.01 million questions and 0.28 million images, covering diverse scenarios from common object‑centric domains to specialized domains. FG‑BMK jointly evaluates dialogue‑level fine‑grained semantic recognition and feature‑level visual discriminability through human‑oriented and machine‑oriented paradigms, enabling diagnostic analysis of whether LVLM failures arise from insufficient visual representations, weak visual‑to‑semantic grounding, or limited fine‑grained knowledge. Through extensive experiments on a diverse set of representative LVLMs/VLMs, we find that current LVLMs remain inadequate fine‑grained recognizers, with failures arising from intertwined bottlenecks in visual representations, semantic grounding, modality alignment, and category‑level knowledge. We further analyze training design factors for improving fine‑grained capabilities and examine how visual and linguistic perturbations affect LVLM predictions. These findings provide diagnostic insights into the limitations of current LVLMs and offer guidance for future data construction and model design in developing more reliable LVLMs for fine‑grained visual tasks. Our code is open‑source and available at https://fg‑bmk.github.io/.
Authors:Yuchen Rao, Xuqian Ren, Yinyu Nie, Sayan Deb Sarkar, Biao Zhang, Vincent Lepetit, Friedrich Fraundorfer
Abstract:
Recovering complete 3D representations of objects from few casual image captures remains a significant challenge. Recent 3D generative models, particularly those based on Flow‑Matching (FM), can synthesize high‑quality textured assets; however, they often suffer from ''synthetic bias'' where learned priors override observational evidence, alongside a lack of alignment with the observed instance. Conversely, optimization‑based methods like 3D Gaussian Splatting (3DGS) provide high fidelity on visible surfaces but fail to reason about unobserved geometry. In this paper, we present FlowObject, a framework that reformulates sparse‑view 3D reconstruction as a training‑free, guided inverse problem. Our approach applies a dual‑space guidance strategy to steer the Ordinary Differential Equation (ODE) trajectory of a flow‑matching model, enabling the completion of unseen regions through learned generative priors while enforcing strict consistency with real‑world observations. By integrating a 3DGS refinement stage, FlowObject further bridges the gap between ''synthetic‑looking'' generative outputs and photorealistic reconstructions. Comprehensive benchmarks on synthetic and real‑world datasets demonstrate that current state‑of‑the‑art methods often struggle to achieve geometric completeness and observational consistency simultaneously, especially under severe occlusions. In contrast, our method significantly outperforms state‑of‑the‑art generative models and optimization‑based frameworks in both geometric completeness and view‑dependent appearance fidelity.
Authors:Tim Rädsch, Yuki M Asano, Hilde Kuehne, Stefan Bauer, Priyank Jaini, Robert Geirhos, Carsten T. Lüth
Abstract:
Video generative models ( VGMs) have become a new frontier that can be used not just for video generation but for a multitude of downstream tasks, including world modeling. To advance these tasks, a good video model must understand the physical reality of the world. Evaluating this understanding is an emerging field and has led to the Physics‑IQ benchmark, which quantifies this explicitly by comparing model‑generated videos to real‑world videos of physical experiments. In this work, we present a systematic audit of the Physics‑IQ benchmark, expose shortcomings and propose three solutions that sharpen how we can measure physical understanding of VGMs. Specifically, we improve prompt and ground‑truth quality to reduce the influence of confounding factors and further introduce a sample‑level scoring system that weights each sample and metric equally. Our resulting benchmark, Physics‑IQ Verified, refines 57.6% of all samples and improves over 34.8% of prompts. In a comparison study using six image‑to‑video generative models, we observe moderate but meaningful ranking changes (Kendall's τ= 0.46). We hope Physics‑IQ Verified advances the community by providing a more reliable signal toward physically accurate VGMs. The code for the benchmark can be accessed at https://github.com/google‑deepmind/physics‑iq‑benchmark
Authors:Abdulmalik Alquwayfili, Faisal Almeshal, Jumanah Almajnouni, Leena Alotaibi, Faisal Alhajari, Mohammed Alkhrashi, Alreem Almuhrij, Abdullah Aldwyish, Raied Aljadaany, Huda Alamri, Muhammad Kamran J. Khan
Abstract:
Image retrieval in crowded scenes is particularly challenging due to the salience bias of conventional visual encoders, which tend to focus on dominant objects while neglecting low‑attention regions that are often crucial for fine‑grained retrieval. We propose LARE (Low‑Attention Region Encoding), a framework that explicitly models these overlooked regions. LARE adopts a dual‑encoding strategy that encodes low‑attention regions of an image and the full image in parallel, leading to more diverse and informative image embeddings. To evaluate image retrieval performance in challenging crowded scenes, we introduce Dense‑Set, a challenging subset derived from COCO and Flickr30K. In this subset, images are re‑captioned to provide richer descriptions of low‑attention or previously overlooked regions. This dataset highlights the limitations of existing retrieval models and enables a more rigorous evaluation under densely crowded scene conditions. Experimental results demonstrate that the proposed framework improves retrieval performance by preserving subtle, non‑dominant visual cues within the shared latent space.
Authors:Veit Hucke, Thomas Pinetz, Gregor Reiter, Ursula Schmidt-Erfurth, Hrvoje Bogunović
Abstract:
Optical coherence tomography (OCT) is essential in ophthalmology, but inconsistent image quality especially in low‑cost devices hinders automated analysis. To address this, we introduce a flow‑matching‑based test‑time adaptation method that generates high‑quality surrogate images from noisy inputs. Typically, domain gaps between test and training data cause pixel distribution mismatches during the denoising process. We overcome this by matching the test image's histogram to synthetic reference trajectories, successfully aligning the input with expected distributions. Additionally, we remove the network's time conditioning to account for slight deviations in real‑world noise distributions. Our approach achieves state‑of‑the‑art performance in segmenting critical biomarkers for two stages of Age‑related Macular Degeneration (AMD). Code is available: https://github.com/Veit21/tta‑flow.
Authors:Hana Jebril, Thomas Pinetz, Günter Klambauer, Hrvoje Bogunović
Abstract:
Reliable pixel‑level uncertainty quantification holds the potential to transform clinical workflows by enabling high‑fidelity longitudinal monitoring and distinguishing true pathological changes from artifacts. Ideally, these models provide the stability required for critical treatment planning and surgical intervention. However, standard deep learning models often suffer from miscalibration, yielding overconfident predictions that mask underlying vulnerabilities at subtle pathological boundaries. To address this, we propose QUAM‑SM, a post‑hoc framework using targeted adversarial search to identify "adversarially fragile" pixels. By actively seeking perturbations that expose predictive instability, our method highlights regions where decisions are most vulnerable to being flipped. Importantly, the framework disentangles epistemic uncertainty from aleatoric uncertainty. Experiments on two public datasets with multiple expert annotations demonstrate that QUAM‑SM outperforms both standard and recent uncertainty estimation approaches in terms of reliability and boundary sensitivity. Code is available at https://github.com/HanaJebril/quam_sm
Authors:Like Zhang, Runliang Niu, Shiqi Wang, Xiyu Hu, Qianli Xing, Pan Wang, Qingzu He, Qi Wang
Abstract:
Vision‑language models (VLMs) are rapidly advancing toward sophisticated grounded structured visual reasoning. Training models for such advanced capabilities demands a new genre of data that seamlessly unifies spatial coordinates, open‑vocabulary descriptions, structured attributes, and topological relationships into a singular representation. However, existing data annotation tools fundamentally fail to meet these intricate demands, suffering from three systematic bottlenecks: limited expressiveness, severe annotation‑training decoupling, and poor data reusability. To bridge this infrastructure gap, we introduce an open‑source annotation tool, ScreenAnnotator. First, we define a unified annotation atom schema that binds spatial, semantic, and structural primitives into a single unit. Second, we implement an on‑policy annotation loop embedded with a Bayesian Annotation Verifier (BAV). Finally, we design a template‑driven multi‑task data synthesis process dynamically transforms static atoms into diverse multi‑dimensional reasoning tasks, eliminating redundant re‑annotation. The on‑policy loop drives the annotation accept rate to nearly 100% on flowcharts and 77% on GUI screenshots, while steadily reducing per‑image annotation time as labeled data accumulate. In the flowchart scenario, fine‑tuning a VLM yields 76.1% average accuracy, which is a 35.1% point absolute gain. Our code is available at: https://github.com/WnQinm/Annotator.
Authors:Zhoupeng Guo, Yunqi Zhu, Zhihe Fan, Xinjie Yao, Ruipu Zhao, Boan Tao, Yiming Sun, Zhen Wang, Pengfei Zhu
Abstract:
Air‑ground collaborative perception is crucial for robust visual understanding in real‑world dynamic environments. However, existing studies typically formulate collaboration as single‑task cross‑view fusion, overlooking the functional dependencies among localization, target association, and fine‑grained parsing. In addition, the heterogeneous nature of aerial and ground views introduces substantial geometric, scale, and occlusion discrepancies, making uniform feature sharing vulnerable to negative transfer. To tackle these issues, we model air‑ground perception as a progressive cross‑task collaboration task and construct the Air‑Ground Progressive Collaboration (AGPC) benchmark, a spatio‑temporally aligned benchmark comprising more than 745K raw video frames. Built upon this benchmark, we propose Socialized Co‑Perception (SCP), a coarse‑to‑fine framework that organizes collaboration progressively from aerial global localization to ground target association and identity‑aware parsing. Its core module, the Dual‑Layer Router (DLR), decouples input‑side multi‑scale expert selection from output‑side task‑conditioned modulation, enabling selective cross‑view and cross‑task interaction while suppressing harmful interference. Extensive experiments demonstrate the effectiveness of SCP. It achieves a 3.73% coevolutionary gain and a 7.86% improvement in average downstream performance. These results show that task‑conditioned collaboration is more effective than uniform fusion for heterogeneous air‑ground perception. The code is available at https://github.com/g1136639260‑spec/AGSCP.
Authors:Hong-Jun Choi, Jongho Lee, Jaeyoung Kim
Abstract:
Table Structure Recognition (TSR) using a pointer network achieves impressive results by predicting HTML sequences while aligning tags to detected text (or cell) regions. However, our analysis reveals that when pointer networks fail, 79.6% of errors occur between spatially adjacent cells (Manhattan distance <= 2). Despite this, standard cross‑entropy loss weights all negative candidates equally. In this work, we propose Geometry‑Aware Pointer (GAP) Loss, which reweights the cross‑entropy objective based on spatial proximity to ground truth. By applying inverse distance weighting, GAP focuses gradient flow where the model struggles most: immediate neighbors receive stronger gradients than distant cells. Our approach requires only a straightforward modification to the loss computation, maintaining the same model architecture with zero additional inference cost. Extensive experiments on PubTabNet and SynthTabNet demonstrate that GAP consistently reduces adjacent‑cell errors, achieving new state‑of‑the‑art performance. Our findings suggest that incorporating geometric inductive biases at the loss level provides a simple yet effective approach to robust TSR. Our code is available at https://github.com/teamreboott/GAP
Authors:Lin Zhang, Sicheng Mo, Zefan Cai, Jinhong Lin, Zihao Lin, Jiuxiang Gu, Krishna Kumar Singh, Yuheng Li, Yin Li
Abstract:
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexible generation order, e.g., conditioning on future context to extend backward, or on both past and future context for inbetween generation. We bridge this gap by training an autoregressive model that supports generation in arbitrary temporal directions. A key technical challenge arises from the Causal 3D VAE widely used in video diffusion models, which encodes latents strictly conditioned on past context. While suited for forward generation, this causal structure causes inter‑block discontinuities when generation proceeds backward. To address this, we introduce blockwise anchor latents, a set of auxiliary latents that restore the missing past context at block boundaries during backward generation. Built on this design, we propose UniTemp, a bidirectional distillation framework that trains a single autoregressive student model for any‑direction video generation. At inference time, UniTemp conditions on arbitrary past and/or future frames, improving controllability for both bidirectional and inbetween generation. Experiments show that UniTemp maintains competitive performance on short and long video generation compared to forward‑only methods, while enabling diverse workflows such as bidirectional video extension, inbetween generation, looping video generation, scene transition, and visual story generation. Project website: https://lzhangbj.github.io/projects/unitemp/
Authors:Minh-Loi Nguyen, Xuan-Vu Le, Long-Bao Nguyen, Hoang-Bach Ngo, Trung-Nghia Le
Abstract:
Traditional image captioning methods often struggle to generate comprehensive, context‑rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval‑augmented image captioning framework that generates captions with deeper insights, such as object attributes, event context, and underlying significance, by leveraging external knowledge. Our approach features a hierarchical multi‑modal article retrieval mechanism that moves beyond monolithic text entities. This retrieval considers article structure‑aware features, including weighted textual components (e.g., headlines, body sections) and visual placement patterns, alongside multi‑faceted similarity computations (content‑‑visual, visual‑‑visual, and discourse positioning). A subsequent contextual relevance refinement stage further enhances the retrieved information. The retrieved articles then serve as the knowledge base for caption generation: first, a VLM generates a concise image description; second, we segment relevant information from the retrieved articles based on this description; and finally, an LLM utilizes both the description and extracted knowledge to generate a comprehensive, contextually detailed caption. We participated in the ACM Multimedia EVENTA 2025 Challenge and achieved 5th place with an overall score of 0.2824 on the private test set of the OpenEvent‑V1 dataset. Source code is publicly released at https://github.com/mf0212/EVENTA‑Challange.
Authors:Kecia G. de Moura, Robert Sabourin, Rafael M. O. Cruz
Abstract:
Offline handwritten signature verification aims to distinguish genuine from forged signatures using static images. Since real forgeries are rarely available, negative samples are usually randomly drawn from genuine signatures of other users to create training data. However, this random selection often lacks diversity, increases redundancy, and escalates computational cost, leading to inefficient training. We propose a data‑driven strategy to generate diverse, informative negative samples using prototypical signatures, which are compact, non‑identifiable summaries of genuine signature features. Based on the experiments results, we conclude that (i) prototypical signatures yield more informative negative samples, improving the detection of skilled forgeries; (ii) the proposed approach is backbone‑agnostic, showing robustness across architectures; and (iii) when combined with a primal‑form linear SVM, it serves as an alternative to RBF‑based models while significantly improving scalability and computational efficiency. Implementation of the method is available at https://github.com/kdmoura/proto_hsv.
Authors:Chengwen Liu, Zhe Huang, Jisheng Dang, Hong Peng, Qi Tian, Tat-Seng Chua
Abstract:
Reinforcement learning has improved the reasoning ability of large language models, but applying outcome‑only rewards to video multimodal large language models (Video‑MLLMs) provides limited guidance on which visual evidence should support the answer. Inspired by multisensory integration, where consistent cues can enhance the salience and reliability of perceptual estimates, we introduce Consensus Frame GRPO (CF‑GRPO), a temporal‑annotation‑free process‑level reward framework for evidence‑aware video reasoning. CF‑GRPO constructs a consensus frame prior from intrinsic video cues, including temporal coverage, scene‑transition cues, and query‑conditioned visual relevance. It then computes a model‑side frame‑use score from visual and response representations and optimizes their agreement through the Consensus Frame Reward (CFR). With salience‑aware sparse aggregation and distribution sharpening, CFR provides a high‑contrast reward signal without requiring human temporal annotations. Experiments show that VideoCFR achieves competitive performance across complex video reasoning benchmarks and improves several metrics over representative Video‑MLLM and RL baselines, while the consensus prior provides an interpretable view of the evidence frames emphasized during training. The implementation is available at https://github.com/1Pansy/VideoCFR.
Authors:Hiranya Garbha Kumar, Minhas Kamal, Balakrishnan Prabhakaran
Abstract:
Accurately aligning CAD models to their corresponding objects in indoor RGB‑D scans is a central challenge in 3D semantic reconstruction. The task requires estimating a 9‑Degree‑of‑Freedom (DoF) pose‑position, rotation, and scale along three axes‑but is hindered by noisy and incomplete scans, as well as segmentation errors that cause geometric distortions. We present Completion‑Assisted Object‑CAD Alignment (CAOA), a method that integrates a semantically and contextually aware point cloud completion module with a symmetry‑aware relative pose estimation algorithm, enabling precise alignment of CAD models to scanned objects. Existing completion methods are typically trained and evaluated on synthetic datasets, which often fail to generalize to real‑world scans. To bridge this gap, we introduce a synthetic data generation strategy tailored to indoor scenes, significantly reducing the synthetic‑to‑real domain gap‑validated through quantitative comparisons with widely used completion datasets. In addition, we release S2C‑Completion, an expert‑annotated dataset of over 8,500 object‑CAD pairs from Scan2CAD, created for real‑world indoor single‑object completion and intended as a new benchmark for this task. For object‑CAD alignment, we incorporate symmetry information via a symmetry‑aware loss, improving robustness to symmetric ambiguities. On the Scan2CAD benchmark, CAOA achieves a 17% accuracy improvement over state‑of‑the‑art methods.
Authors:Wujian Peng, Lingchen Meng, Yuxuan Cai, Xianwei Zhuang, Yuhuan Yang, Rongyao Fang, Chenfei Wu, Junyang Lin, Zuxuan Wu, Shuai Bai
Abstract:
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinders truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re‑encoding. UniAR adapts a pretrained vision encoder with multi‑level feature fusion and a lookup‑free bitwise quantization scheme, preserving both high‑level semantics and low‑level details while scaling the effective visual vocabulary at minimal cost. Building on this, the unified autoregressive model adopts parallel‑bitwise‑prediction to jointly predict spatially grouped, multi‑level visual codes, substantially reducing visual sequence length and accelerating generation. Finally, a diffusion‑based visual decoder operates on discrete visual tokens to decode high‑fidelity images. Through large‑scale pre‑training, followed by supervised fine‑tuning and reinforcement learning, UniAR achieves state‑of‑the‑art performance on image generation and image editing while remaining competitive on multimodal understanding benchmarks. The project page is available at https://sharelab‑sii.github.io/uniar‑web.
Authors:Jiye Lee, Yonghun Choi, Jungdam Won
Abstract:
Collaborative human‑object interaction shows dynamic and complex movements that require mutual anticipation and continuous adjustment between participants and the shared object. Modeling such collaborative multi‑human object interaction (MHOI) scenarios requires high‑quality data acquisition as a foundational step; however, this is challenging due to the inherent complexity of MHOI where human‑human and human‑object interactions occur simultaneously. Such complexity leads to noisy MHOI captures characterized by several artifacts: contact misalignment between hands and objects, motion jitter and temporal inconsistencies in the captured sequences, and missing or incomplete finger‑level articulation details. To address these challenges, we present MOCHI (MOtion Enhancement of Collaborative Human‑object Interactions), a two‑stage framework for enhancing noisy MHOI data. Our approach first generates physically plausible hand grasps through optimization from noisy body input, producing grasps that are both physically plausible and semantically consistent with the body pose, where these optimized grasps are extended into complete hand‑object interaction sequences. Consequently, the full‑body motion for all participants are refined through a diffusion‑based noise optimization framework that uses single‑person motion priors. During the optimization process, we introduce optimization objectives to encode human‑object and human‑human interaction information within these single‑person priors. Experimental results demonstrate the effectiveness of our pipeline across diverse MHOI data, either acquired by existing capture methods or synthesized by generative models. We further show robustness of our system across varying numbers of participants and types of interactions, and demonstrate various applications including keyframe‑based MHOI creation and data augmentation through varying object geometries.
Authors:Dongyue Lu, Rong Li, Ao Liang, Lingdong Kong, Wei Yin, Lai Xing Ng, Benoit R. Cottereau, Camille Simon Chane, Wei Tsang Ooi
Abstract:
Event cameras sense the world through asynchronous brightness changes with microsecond latency and high dynamic range, offering motion fidelity far beyond frame‑based sensors and capturing temporal structure that conventional exposures often miss. These properties make events a powerful complement to RGB in autonomous driving, especially under blur, glare, and rapid motion, where frame‑based perception can become unreliable. However, existing event‑aware vision‑language models remain limited to generic perception and do not reveal how event sensing contributes to reasoning and decision‑making across the full driving loop. We present EventDrive, a large‑scale benchmark and model suite that unifies event streams, RGB frames, and language supervision across four core dimensions: Perception, Understanding, Prediction, and Planning, covering captions, structured QA, grounding, motion‑state recognition, trajectory forecasting, and planning tasks. Building on this foundation, EventDrive‑VLM introduces a multi‑horizon event pyramid and a temporal‑horizon mixture‑of‑experts module to adaptively encode and fuse asynchronous and frame‑based information for downstream reasoning. Comprehensive evaluation across diverse tasks shows that event streams provide substantial gains in temporal precision, motion awareness, and robustness, bringing event sensing into the center of driving intelligence.
Authors:Zihan Gu, Ruoyu Chen, Junchi Zhang, Li Liu, Xiaochun Cao, Hua Zhang
Abstract:
Visual attribution is a fundamental tool for interpreting modern vision and vision‑language models, particularly when their decisions must be inspected, diagnosed, or audited. Its goal is to explain how a model's decision depends on local regions of the visual input, typically by assigning an importance ordering over candidate image regions. Given an image partitioned into n regions, faithful attribution can be cast as an ordered subset‑search problem, in which progressively inserting the selected regions should recover the target model response as early as possible. Exhaustive search over region subsets incurs exponential cost, while the widely used greedy search still requires a quadratic number of model evaluations, because every selection step rescores all remaining candidates. We propose PhaseWin, an efficient subset‑search algorithm for faithful visual attribution. PhaseWin reorganizes greedy region selection into a phased window‑search procedure: rather than re‑evaluating the full candidate set at every step, it alternates between global candidate screening, adaptive pruning, and localized window refinement, while preserving the essential region‑ranking behavior of greedy search. We analyze PhaseWin under monotone evidence‑accumulation conditions and show that, under feature‑level structural assumptions, it attains controllable linear evaluation complexity together with near‑greedy faithfulness guarantees. Extensive experiments on image classification, object detection, visual grounding, and image captioning show that, among all compared attribution methods, PhaseWin reaches high faithfulness with the fewest forward passes, empirically realizing the predicted reduction from O(n^2) to O(n). The code is available at https://github.com/Qihuai27/phasewin‑va.
Authors:Yonghao Chen, Sicheng Yang, Rui Tang, Lei Zhu
Abstract:
Multi‑contrast magnetic resonance imaging (MRI) provides complementary information for clinical diagnosis. However, acquiring all MRI sequences is often time‑consuming and costly. Recent generative models perform cross‑contrast synthesis to address this issue by inferring absent contrasts from the available ones. Nevertheless, synthesizing 3D MRI presents significant challenges. Due to the massive volume sizes, operating directly in the pixel space is computationally prohibitive; therefore, a common approach is to first compress the 3D volumes into a latent space and subsequently train generative models in that space. We observe that existing compression architectures face several critical issues: they under‑preserve long‑range anatomical coherence, discard clinically meaningful semantics, and rely on optimization objectives that lead to over‑smoothed reconstructions. Ultimately, these shortcomings compromise the performance of subsequent generative models. In this work, we propose a semantics‑first latent modeling framework for 3D MRI reconstruction and cross‑contrast synthesis. Specifically, we introduce a Latent Harmonization Encoder (LHE) to capture global anatomical dependencies, ensuring coherent volumetric representations. To mitigate semantic degradation during latent compression, we further design a Semantic Recovery Block (SRB) that injects high‑level priors from a self‑supervised semantic teacher, enhancing contrast‑aware separability in the latent space. Additionally, we propose an Anatomy‑aware Frequency Loss (AFL) to adaptively preserve diagnostically relevant high‑frequency structures. Extensive experiments on two public multi‑contrast MRI datasets demonstrate consistent improvements in reconstruction fidelity and cross‑contrast synthesis quality. Our code is available at https://github.com/script‑Yang/RSF.
Authors:Sicheng Yang, Hongqiu Wang, Zhaohu Xing, Sixiang Chen, Qiuxia Yang, Yize Mao, Guang Yang, Lei Zhu
Abstract:
Self‑supervised DINO models provide strong transferable visual representations, yet applying them directly to image segmentation remains challenging. Existing approaches commonly rely on heavy decoders with complex upsampling, introducing substantial parameter and computational overhead. We observe that introducing scale into DINO features is far more critical than increasing decoder capacity. In this work, we present SegDINO, an efficient segmentation framework that integrates a DINOv3 backbone with lightweight scale modeling. SegDINO introduces Token Pyramid Adaptation (TPA) to reorganize intermediate DINO features into a pseudo multi‑scale hierarchy, and Scale‑Aware Decoding (SAD) for efficient intra‑scale refinement and top‑down multi‑scale propagation. We further curate PanCT, a new CT dataset containing 284 patients with expert‑annotated pancreatic tumors, to assess SegDINO's ability to handle difficult small‑lesion cases. Extensive experiments on PanCT and three public benchmarks demonstrate that SegDINO achieves state‑of‑the‑art results with high efficiency. The code is available at https://github.com/script‑Yang/segdino_v2.
Authors:Yuming Chen, Yuxin Xie, Tao Zhou, Yi Zhou
Abstract:
Semi‑supervised medical image segmentation has emerged as a dominant research problem in medical image analysis, mitigating annotation scarcity by leveraging consistency regularization on unlabeled data. However, existing approaches operate predominantly via visual pattern matching, relying heavily on pixel‑level similarities. This visual‑centric dependency often falters in clinical scenarios characterized by the visual‑semantic mismatch, where visually similar lesions warrant distinct diagnostic conclusions, thus failing to capture the underlying diagnostic logic used by experts. To address this, we move beyond visual cues and propose CERS (CoT‑Enhanced Reasoning Segmentation), a framework that integrates Chain‑of‑Thought (CoT) reasoning to distinguish pathologically distinct cases. Specifically, we construct a knowledge pool enriched with linguistic reasoning descriptions generated by large language models (LLMs). A semantic‑aware reference selection strategy is introduced to identify historical evidence, filtering candidates first by morphology, and then refining them via CoT consistency to eliminate hard negatives. Furthermore, a multi‑scale coordinate attention module (MCAM) is designed to effectively fuse this reasoning‑derived context into the decoding process. Extensive experiments demonstrate the superiority of CERS against state‑of‑the‑art approaches, particularly in resolving boundary ambiguities and semantic inconsistencies. The code is available at https://github.com/cymasuna/CERS.
Authors:Guo Pu, Yixuan Han, Haofeng Li, Yao Zhang, Hui Zhou, Zhouhui Lian
Abstract:
Online 3D reconstruction from monocular image sequences is a challenging and ongoing research topic. 3D Gaussian Splatting (3DGS), leveraging its high‑quality real‑time rendering capability, empowers online 3D reconstruction to represent dense scenes with enhanced expressiveness, and thus holds great promise for a wide range of applications such as robotics and AR/VR. However, existing online 3DGS methods still suffer from some key challenges: fragile camera pose estimation due to the lack of global optimization, and low optimization efficiency in large‑scale or long‑sequence scenarios. To address these issues, we propose a robust and efficient online voxelized 3DGS reconstruction framework integrated with global \textSim(3) optimization, which enables reliable camera tracking and efficient global loop closure for both camera poses and voxelized 3DGS. To accelerate the convergence of the voxelized 3DGS, we further introduce a color residual learning strategy, which not only boosts optimization speed but also enhances rendering quality. Extensive experiments on diverse indoor and outdoor datasets demonstrate that our method achieves state‑of‑the‑art performance in both camera pose estimation accuracy and rendering quality, while retaining real‑time efficiency. Additionally, we develop and deploy a real‑world UAV‑based active reconstruction system grounded on our proposed method, validating its robustness and generalizability for practical online 3D reconstruction tasks. Our code and data are available at https://github.com/TrickyGo/MoonSplat.
Authors:Antonio Scardace, Daniele Ravì
Abstract:
Despite increasing adoption of multimodal approaches in Alzheimer's Disease (AD) research ‑‑ aimed at integrating molecular, structural, clinical, and genetic biomarkers to enhance disease characterization ‑‑ the relationships among these modalities remain poorly understood. A systematic analysis of their dynamic interaction is essential for improving disease modeling, identifying redundant assessments, and reducing patient burden and acquisition costs. In this paper, we present a quantitative analysis of multimodal AD biomarkers by integrating tau‑PET, structural MRI, cognitive scores (MMSE and CDR), and APOE4 data from 789 subjects drawn from the ADNI dataset. In our analyses, we (A) quantify cross‑modal mutual information and explained variance to assess redundancy and predictive dependencies; (B) examine associations between tau topologies and structural atrophy across brain regions to select informative ROIs; (C) perform a statistical decomposition of the tau‑cognition association into atrophy‑related and atrophy‑independent components; (D) and identify a dominant neurodegenerative trajectory that aligns with cognitive decline. This study provides a systematic characterization of cross‑modal relationships, improving the interpretability and selection of biomarkers in AD. Code is publicly available at: https://github.com/antonioscardace/Multimodal‑AD.
Authors:Zhenyu Yang, Kairui Zhang, Bing Wang, Shengsheng Qian, Changsheng Xu
Abstract:
Despite the remarkable progress of Video Large Language Models (Video‑LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long‑horizon contextual memory. These obstacles undermine real‑time responsiveness and cause severe forgetting throughout prolonged interactions. In this work, we introduce LiveStarPro, a live streaming assistant that is designed for proactive video understanding over long‑horizon streams. The design of LiveStarPro rests on three complementary components. The first component is Streaming Verification Decoding (SVeD), an inference framework that identifies the appropriate response timing through single‑pass perplexity verification, thereby eliminating the dependency on explicit silence tokens. The second component is Streaming Causal Attention Masks (SCAM), a training strategy that enforces incremental video‑language alignment over variable‑length streams. The third component is Tree‑Structured Hierarchical Memory (TSHM), a recursive memory architecture that organizes evicted historical information into event chains and consequently enables efficient retrieval from effectively unbounded video streams. To facilitate a comprehensive evaluation under realistic online conditions, we further present OmniStarPro, a large‑scale benchmark that spans 15 diverse real‑world scenarios and that extends to hour‑scale streams for the assessment of long‑term recall. Extensive experiments demonstrate that LiveStarPro consistently surpasses existing methods, attaining a 28.9% improvement in semantic correctness and an 18.2% reduction in timing error, while its streaming key‑value cache further yields a 1.58x inference speedup over the same model without caching. The model and the code are publicly available at https://github.com/sotayang/LiveStarPro.
Authors:Zhexiao Xiong, Yizhi Song, Hao Kang, Qing Yan, Liming Jiang, Jenson Yang, Zhoujie Fu, Stathi Fotiadis, Angtian Wang, Zichuan Liu, Bo Liu, Yiding Yang, Xin Lu, Nathan Jacobs
Abstract:
Interactive world models aim to simulate environment dynamics under real‑time user actions. However, their action vocabulary is largely confined to navigation: most actions correspond to motion (e.g., walk, turn, look around), while interaction with objects in the scene (e.g., pick up plates, open doors, or trigger physical responses) is either absent, restricted to game domains, or relegated to prompt‑to‑full‑video scenarios. The resulting worlds are visually explorable but not truly actionable. In this work, we present ActWorld, an interactive world model that extends prior navigation‑centric generators to support mid‑rollout object interaction within a chunk‑autoregressive framework. We argue that the navigation‑interaction gap stems from two bottlenecks. First, a data bottleneck: the lack of human‑object interaction data with accurate, dense labels. Second, a memory bottleneck: recency‑biased history compression in existing world models discards the event‑transition frames that causally determine subsequent object states, leading to an action‑forgetting pathology. On the data side, we construct a 100K interaction video dataset, each annotated with per‑chunk captions via chain‑of‑thought reasoning. On the model side, we introduce a hierarchical action‑aware memory design that routes history compression by interaction importance, complemented by a persistent memory bank that maintains event‑update and object‑identity tokens across long rollouts. Experiments show that ActWorld supports both flexible navigation and rich object interaction within a single model, substantially improving interaction fidelity over navigation‑only baselines without sacrificing viewpoint control. Project page is available at https://interactwm.github.io/ActWorld.
Authors:Jiangong Xu, Weibao Xue, Xiaoyu Yu, Jun Pan, Xinlian Lianga, Mi Wang
Abstract:
Optical remote sensing imagery is frequently degraded by cloud and cloud‑shadow contamination, which limits its reliability for near‑real‑time land use and land cover (LULC) mapping. Although synthetic aperture radar (SAR) can provide cloud‑penetrating structural information, existing SAR‑optical fusion methods often assume reliable optical observations and insufficiently address the semantic uncertainty introduced by cloud contamination. To address this issue, we propose CloudLULC‑Net, an end‑to‑end heterogeneous SAR‑optical fusion framework that directly predicts LULC maps from cloud‑contaminated Sentinel‑2 imagery and temporally adjacent Sentinel‑1 SAR observations. The proposed network incorporates optical reliability modulation to suppress unreliable optical responses, heterogeneous information adaptive aggregation to model high‑order spatial‑channel interactions between optical and SAR representations, and a unified semantic mapping transformer to organize fused features in a LULC‑oriented latent space. A semantic anchor‑guided optimization strategy is further introduced to improve the consistency of intermediate semantic representations. To support this task, we construct CloudLULC‑Set, a large‑scale benchmark dataset containing 40,223 curated SAR‑optical‑label triplets with pixel‑level LULC annotations across diverse geographic regions and cloud conditions. Experimental results show that CloudLULC‑Net achieves an OA of 86.60%, an F1‑score of 83.29%, and an mIoU of 73.51%, outperforming representative heterogeneous reconstruction‑first and end‑to‑end SAR‑optical mapping methods. Comparisons with existing global LULC products and analyses under different cloud‑cover levels further demonstrate the robustness and practical value of CloudLULC‑Net for target‑date LULC mapping in cloud‑prone regions.The project is publicly available at: https://github.com/RSIIPAC/CloudLULC
Authors:Jens Bayer, Stefan Becker, David Münch, Michael Arens, Jürgen Beyerer
Abstract:
Pixel‑wise adversarial patches are computationally heavy and often visually detectable, limiting utility in security‑critical systems. We present adversarial Voronoi camouflage that optimizes only seed‑point locations under fixed, printable palettes using a soft assignment, producing structured, splinter camouflage‑like patterns without additional regularization. Evaluated on person detection with COCO‑style AP@[.5:.95], naive placement (Inria ‑> COCO) performs comparably bad, while garment‑level application via segmentation mask (3DPeople) results in a significant AP drop. The attack transfers to out‑of‑domain backgrounds and across detector families (YOLOv9/10/11/12), indicating robustness in black‑box settings. Repainting with different palettes largely nullifies the effect, and single‑color tweaks show limited tolerance (<=0.17), highlighting a structure‑palette coupling. The parameter‑efficient, palette‑constrained design improves visual plausibility while degrading real‑time detector performance. Physical validation and color calibration are left for future work. Code: https://github.com/JensBayer/Voronoi This paper was originally presented at the International Conference on Military Communication and Information Systems (ICMCIS), organized by the Information Systems Technology (IST) Scientific and Technical Committee, IST‑224‑RSY ‑ the ICMCIS, held in Bath, United Kingdom, 12‑13 May 2026.
Authors:Hong Yang, Basura Fernando
Abstract:
Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visual observations. Yet existing visual and embodied question answering benchmarks often provide limited control over the reasoning dependencies being tested, making it difficult to distinguish grounded embodied reasoning from shortcut‑driven visual or linguistic pattern matching. We present ERQA‑Plus, a diagnostic benchmark for reasoning in embodied AI. ERQA‑Plus contains 1,766 question‑answer instances grounded in 711 robot‑centric images and organized according to a structured taxonomy spanning perceptual, action‑centric, social‑interaction, navigation‑environmental, and contextual commonsense reasoning. The dataset is constructed using a multi‑stage generation and validation pipeline that combines taxonomy‑guided question generation, automatic quality judging, iterative revision, and human assessment to improve visual grounding, answer validity, and reasoning quality. We benchmark representative general‑purpose vision‑language models and embodied models, including LLaVA‑NeXT‑8B, Prismatic‑7B, MiniCPM‑V‑4.5‑8B, Qwen3‑VL, RoboRefer‑8B, and RoboBrain2.5‑8B. Although the strongest model, Qwen3‑VL‑32B, achieves 83.4% overall accuracy and 61.4 SBERT score, category‑level results reveal persistent weaknesses in spatial reasoning, procedural reasoning, event prediction, and intention inference. ERQA‑Plus therefore provides a fine‑grained evaluation framework for measuring not only whether embodied agents answer correctly, but also which forms of embodied reasoning they can and cannot perform reliably. The dataset is available https://huggingface.co/datasets/huggingdas/erqa‑plus and the project page at https://github.com/LUNAProject22/erqa‑plus.
Authors:Jie Wang, Tao Wang, Ru Zhang, Jianyi Liu
Abstract:
The widespread deployment of face recognition (FR) systems exposes personal images shared on social media and public platforms to identity linkage and privacy risks. Existing adversarial privacy protection methods can degrade unauthorized FR performance but are not compatible with generative face editing. Artificial intelligence‑driven face editing tools are gaining popularity, which has significantly increased user demand for personalized portrait generation and social sharing. However, current editing methods often preserve identity features, making the edited images still susceptible to tracking by malicious FR systems. Thus, this paper proposes Flux‑Guard, a privacy‑preserving face editing framework based on adversarial attacks, which integrates face editing and privacy protection within a unified generative process. Specifically, we design a flow trajectory control method to align semantic manipulations with the generative process and introduce latent‑space adversarial optimization with an adaptive perceptual‑loss‑driven weighting strategy, dynamically adjusting adversarial strength to maximize attack effectiveness while preserving visual quality. Extensive experiments demonstrate that Flux‑Guard supports face editing while significantly improving attack success rates against cross‑domain face recognition models on the CelebA‑HQ and LADN datasets. Furthermore, evaluation results for commercial APIs have confirmed its effectiveness in real‑world applications. The code is released at https://github.com/JLMWang/Flux‑Guard.
Authors:Semin Kim, Jihwan Yoon, Seunghoon Hong
Abstract:
Finding the initial noise that generates a given data sample, known as inversion, is a key component for downstream applications such as training‑free image editing. Existing fixed‑point inversion methods improve inversion accuracy by formulating each inversion step as a fixed‑point problem, but they lack a principled mechanism for selecting among multiple fixed‑point solutions that can arise in practice. We observe that different selections induce different inversion trajectories, leading to substantial variation in reconstruction and editing quality. For rectified flows, we further find that this variation is closely associated with trajectory straightness, motivating straightness as a principled selection criterion. We propose SelFix, a fixed‑point inversion method that selects fixed‑point solutions inducing straighter inverse trajectories while retaining convergence to an exact inverse root under standard local assumptions. Experiments on FLUX.1‑dev and PIE‑Bench show that SelFix improves fixed‑point inversion, achieving stronger real‑image reconstruction and better source‑preserving prompt‑based editing than prior inversion baselines. The code is available at https://github.com/seminkim/selfix.
Authors:Hao-Yuan Ma, Li Zhang, Zhiwei Zhu, Jie Gao
Abstract:
Text‑guided open‑vocabulary object counting (TOOC) aims to count objects belonging to the categories specified by natural language descriptions. Although vision‑language pre‑trained models have been successful applied to TOOC tasks, they still struggle with fine‑grained spatial understanding and real‑time inference requirements in counting scenarios. To address these limitations, this paper proposes a real‑time TOOC framework, called the Real‑Time Counter (RT‑Counter), that achieves not only good counting accuracy but also high computational efficiency. RT‑Counter designs a novel Visual Prototype Textualization (VPT) module that can project learned visual features into a text feature space and then generate features containing the abstract information that is hard to capture with visual prototypes and the detailed prototype information that is difficult to describe in text, enhancing the object‑level visual‑language model's counting capabilities. Additionally, RT‑Counter incorporates our Weaving Transformer (Weaformer) layers, maintaining high descriptive power at a fraction of the computational cost. The Weaformer layer adopts a novel hybrid attention mechanism that can efficiently weave together local and global visual features. Extensive experiments on three public datasets show that RT‑Counter successfully breaks the accuracy‑speed trade‑off in TOOC. While achieving a competitive MAE of 13.30 on FSC147, RT‑Counter operates at 112.48 FPS, making it 7.4x faster and over 4× more parameter‑efficient than the existing leading methods in TOOC. Our work aims at balancing high accuracy and real‑time performance in TOOC. Code is available at: https://github.com/Jason‑Mar1/RT‑Counter.
Authors:Yu Guo, Zhengru Fang, Shengfeng He, Senkang Hu, Yihang Tao, Phone Lin, Yuguang Fang
Abstract:
Image restoration seeks to recover high‑quality images from degraded inputs but becomes highly ill‑posed under complex, mixed degradations. While unified all‑in‑one models are common, their performance declines as degradation complexity increases. Recent works adopt Chain‑of‑Thought (CoT) reasoning for multi‑round restoration using specialized modules. However, this approach faces two key limitations: (i) increased computational cost due to multi‑step processing, and (ii) weak modeling of interactions between degradations during stepwise inference. We introduce CoTIR, a universal image restoration framework that internalizes CoT reasoning within a single model. Concretely, we view image restoration as a specialized subtask of image editing, which implies that a large‑scale pre‑trained editing model provides a more favorable optimization starting point. Building on this, we fine‑tune the model for restoration and further encode structured CoT‑style reasoning into the learning objective via a differentiable formulation inspired by Lagrangian optimization, enabling holistic restoration without chaining specialized restorers. To facilitate training and evaluation, we further present CoTIR‑Bench, a large‑scale benchmark comprising 5.2 million samples with CoT‑style reasoning traces. Extensive experiments on CoTIR‑Bench and broad real composite degradation scenes show that CoTIR achieves stronger perceptual quality and more competitive fidelity than both all‑in‑one models and multi‑round restoration methods. The source code is available at https://github.com/gy65896/CoTIR.
Authors:Aryan Bhagat
Abstract:
Melanoma is the most dangerous form of skin cancer with five‑year survival rates exceeding 99% when detected early but falling sharply once the disease spreads. This paper proposes and evaluates a two‑stage fine‑tuning approach for ResNet50 applied to binary melanoma classification on dermoscopic images. The core challenges addressed are class imbalance and suboptimal transfer learning from single‑stage fine‑tuning. After stratified train/validation/test splitting, random oversampling was applied exclusively to the training set to achieve a 1:1 class balance. Stage 1 trained only the classification head with the ResNet50 base frozen, while Stage 2 fine‑tuned all layers jointly at a low learning rate of 1e‑5 to prevent catastrophic forgetting of learned visual features. On an independent test set of 3,826 images, the model achieved an AUC‑ROC of 0.9559, accuracy of 88.34%, sensitivity of 87.56%, specificity of 89.13%, and F1‑score of 88.29%. An ablation study confirms the two‑stage protocol significantly outperforms single‑stage fine‑tuning, with sensitivity gains of over 4%. Grad‑CAM visualizations demonstrate correct lesion localization. A fully deployable Streamlit detection application is provided alongside all training code.
Authors:Haoyu Wang, Guoqing Ma, Zeyu Zhang, Yandong Guo, Boxin Shi, Hao Tang
Abstract:
Generalist vision‑language‑action systems need object‑centric 3D evidence and reusable manipulation experience to plan reliable robot trajectories. GeneralVLA provides a hierarchical interface for converting language and RGB‑D observations into 3D end‑effector paths, but two bottlenecks remain. First, monocular SAM3D‑style object reconstruction can hallucinate pose and unseen geometry, while manipulation benefits from stable object shape when calibrated multi‑view observations are available. Second, the original KnowledgeBank mainly retrieves semantically similar snippets and appends new knowledge, which makes it difficult to control memory quality, conflicts, confidence, and geometric relevance. To address the first challenge, we introduce GeoFuse‑MV3D, a geometry‑prior‑guided MV‑SAM3D reconstruction branch that verifies external geometry cues with input‑view masks, applies soft visual‑hull support, performs axis‑wise refinement, and fuses only geometry while preserving appearance. To address the second challenge, we upgrade KnowledgeBank into a governed long‑term memory system with explicit quality, confidence, lifecycle, verifier, and conflict metadata, together with precision‑oriented retrieval. Finally, we evaluate the reconstruction branch on GSO‑30 and the memory module on Terminal‑Bench 2.0 and SWE‑Bench Verified; GeoFuse‑MV3D improves over the MV‑SAM3D baseline by reducing CD and LPIPS by 2.20% and 2.02% while increasing PSNR and SSIM by 2.36% and 1.03%, and KnowledgeBank improves over ReasoningBank by 4.53% on Terminal‑Bench SR and 3.73% on SWE‑Bench resolve rate, while reducing AS by 4.95% and 5.65%, respectively. Code: https://github.com/AIGeeksGroup/GeneralVLA‑2. Website: https://aigeeksgroup.github.io/GeneralVLA‑2.
Authors:Xianda Guo, Pinhan Fu, Ruilin Wang, Wenke Huang, Mang Ye, Qin Zou
Abstract:
Stereo matching has advanced through foundation models trained on large‑scale datasets, yet this paradigm suffers from a scalability bottleneck: incorporating new data requires costly joint retraining. Model merging offers a scalable post‑hoc alternative by integrating knowledge from specialized models after source checkpoints are available. However, existing merging methods typically retain all available models or rely on greedy inclusion, which can preserve harmful task‑vector interference. We propose StereoFactory, a coarse‑to‑fine evolutionary framework for adaptive model merging. Stage~1 employs a genetic algorithm to search the combinatorial space of model subsets, determining which models should participate. Stage~2 addresses module‑level knowledge specialization (different functional modules exhibit distinct preferences for knowledge sources) through CMA‑ES optimization of architecture‑adaptive routing over the selected task vectors, with optional module‑level scaling. Experiments across two architectures and four benchmarks demonstrate that StereoFactory consistently achieves the best four‑benchmark average under the same checkpoint pool, reducing the average error from 3.80 to 3.30 on NMRF and from 2.88 to 2.19 on FoundationStereo relative to the strongest controlled baseline. The post‑hoc search requires only 2.7‑‑3.7% of the corresponding joint‑retraining wall‑clock time. Analysis reveals that knowledge contributions are inherently module‑specific, and selected subsets can transfer across architectures with minimal degradation. Code will be publicly released upon acceptance at: https://github.com/XiandaGuo/StereoFactory.
Authors:Bo Gou, Jicheng Zhang, Jianlong Xiong, Tao He, Bentian Liu, Hai Wu, Yijiao Wang, Yu Zhang, Yujia Yang, Yun Dai, Jian Liu, Jie Wang
Abstract:
Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scarce and limited in scale and view coverage. Second, the performance of some modern video‑level architectures for echocardiographic view classification remains underexplored. Third, some view categories exhibit highly similar spatial appearances, making single‑frame features insufficient for discrimination, while heterogeneous frame quality complicates robust temporal information fusion. To address these challenges, we release the Echocardiographic Videos of Nine Views (EV9V) dataset, comprising 5,138 videos, 910,579 frames, and 9 standard views, which is, to the best of our knowledge, the largest publicly available echocardiography video dataset. Using EV9V, we systematically benchmark representative video classification architectures, including Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformers. Furthermore, we propose a Spatio‑Temporal Fusion Model (STFM), an efficient dual‑stream CNN‑LSTM (Long Short‑Term Memory) framework that jointly captures spatial anatomical structures and temporal cardiac dynamics. The proposed framework leverages uncertainty‑aware learning to preferentially sample representative video segments during training and evidence‑based fusion during inference, improving robustness to variations in frame quality across echocardiographic videos. Extensive experiments demonstrate that our method achieves competitive performance across diverse video classification models, validating the effectiveness of uncertainty‑aware spatio‑temporal learning for echocardiographic view classification. The code is available at https://github.com/bgx666/stfm.
Authors:Xiongjun Guan, Jianjiang Feng, Jie Zhou
Abstract:
Fingerprint recognition is still dominated by task‑specific pipelines, where enhancement, structural parsing, alignment, and matching are optimized in isolation. Although effective in narrow settings, this design limits representation reuse across sensors, qualities, and downstream applications. We therefore present UoU, short for ``a Universal fingerprint foundation model based on large‑scale Unsupervised learning,'' which reframes fingerprint feature extraction as a domain‑specific foundation‑model problem. UoU is organized around a multi‑level representation hierarchy spanning image restoration, structural fields, semantic tokens, point‑level biometric entities, and compact global descriptors. Its training recipe combines a supervised cold start on precise annotations, large‑scale weakly supervised refinement, and large‑scale unsupervised consolidation, with the latter two stages iterated during large‑scale training so that weak supervision broadens semantic coverage while unsupervised learning stabilizes correspondences, invariances, and representation geometry. Rather than treating fingerprint imagery as generic texture, UoU exploits domain‑specific symmetries and intermediate structure, including orientation flow, periodic ridge patterns, sparse biometric entities, and spatial equivariance. The framework is intentionally architecture‑agnostic: while the present study includes an initial transformer‑based structured‑prediction instantiation, the broader design supports multi‑task learning, scalable model configurations, and downstream specialization for matching, alignment, enhancement, registration, and related fingerprint applications. This paper presents the technical motivation, system design, and validation protocol of UoU, and part of the baseline implementation is publicly available at https://github.com/XiongjunGuan/UoU.
Authors:Logan Mann, Yi Xia, Ajit Saravanan, Ishan Dave, Saadullah Ismail, Shikhar Shiromani, Emily Huang, Ruizhe Li, Kevin Zhu
Abstract:
Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical. A common intuition, which we call the Attention‑Confidence Assumption, holds that reliability follows from "structural" visual perception: tight attention on relevant regions should signal a trustworthy answer, while scattered attention signals confusion. We challenge this through the VLM Reliability Probe (VRP), a systematic cross‑family study of reliability signals in contemporary Vision‑Language Models (VLMs). We introduce structural‑attention metrics, cluster counts (C_k) and spatial entropy (H_s), to quantify the visual encoder's gaze, and track its evolution (Delta H_s) across layers. This reveals a "Symbolic Detachment": models often "Early Lock" visual features only to diffuse attention later, severing early perception from final generation. Contrary to the grounding hypothesis, we find a "Cluster Failure": spatial attention has near‑zero correlation (R approx 0.001) with accuracy. Instead, reliability is a phenomenon of generation dynamics and internal‑state distributions. Self‑Consistency, the agreement rate across sampled reasoning paths, is the dominant predictor of truth (R = 0.429). Scaling causal interventions exposes a sharp architectural divergence: LLaVA locks its prediction in a fragile late‑stage bottleneck, whereas PaliGemma and Qwen2‑VL distribute reliability globally, staying resilient even when ~50% or more of their most predictive layer is destroyed. For current VLMs, reliability signals are detached from visual grounding maps and are best inferred from generation‑time dynamics and hidden‑state probes.
Authors:Ahmad Darkhalil, Dima Damen, David Fouhey
Abstract:
Understanding hands and the objects they interact with, both directly and through tools, is a key step for tasks ranging from action perception to 3D reconstruction and robotics. Our paper provides several contributions to the Hand‑Object Interaction (HOI) understanding literature: (1) HOI‑DETR, a new framework that introduces hand‑object and object‑object interactions to the Co‑DETR architecture to produce a state‑of‑the‑art method; (2) a comprehensive HOI evaluation suite of 4 diverse datasets, including a video benchmark derived from the HD‑EPIC dataset and fresh annotations that improve the Hands23 benchmark and (3) a trained checkpoint that significantly improves the state of the art across Hands23, HOIST, FineBio, and HD‑EPIC, including mAP gains of over 20 percentage points on Hands23 and FineBio. Our ablations confirm the contributions of each model component.
Authors:Suttisak Wizadwongsa, Hyelin Nam, Supasorn Suwajanakorn, Jeong Joon Park
Abstract:
Generating novel renderings of a scene along user‑defined camera trajectories from a single monocular video, dubbed video retaking, is a compelling but difficult problem in content creation and visual effects. Existing geometry‑guided approaches reconstruct a 4D representation from the source video and render it along the target trajectory to condition video diffusion models. However, this guidance degrades as the target camera departs from the source trajectory, leaving newly revealed regions sparse or entirely missing. We propose SierpinskiCam, which addresses this limitation by augmenting geometry‑based guidance with Sierpinski dome texture cues that contains rich trackable features even under large viewpoint changes. We further introduce a reference video conditioning mechanism that appends source‑video tokens to the target‑token sequence and separates the two streams with negative RoPE indices, enabling appearance grounding without architectural modification or per‑video adaptation. Extensive experiments show that SierpinskiCam achieves significant gains in camera controllability, geometric consistency, and video quality across diverse and challenging retaking scenarios. Project page: https://hyelinnam.github.io/SierpinskiCam/.
Authors:Prabhjot Singh, Bhushan Pawar, Madhu Reddiboina, Rajvee Sheth
Abstract:
Current multilingual evaluations for Vision‑Language Models (VLMs) assume a one‑to‑one mapping between language and orthography, overlooking billions of users of multi‑script languages. We introduce PuMVR (Punjabi Multimodal Visual Reasoning), a benchmark of 1,000 strictly parallel image‑text instances across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman. Evaluating 10 state‑of‑the‑art VLMs, we expose a substantial and systematic Script Gap. Models frequently solve visual tasks in one script while failing identical tasks in another, with accuracy deltas reaching 16%. Crucially, visual input boosts absolute performance uniformly yet does not close the orthographic gap. Furthermore, cross‑script in‑context transfer is highly brittle, exposing script‑locked knowledge representation. Supported by McNemar tests across all script pairs, our findings demonstrate that current "multilingual" VLMs are not truly multi‑script. We propose the Script Consistency Rate (SCR), which falls as low as 24.8% on our benchmark, as a mandatory metric for script‑agnostic evaluation to ensure equitable AI access. Data and code are available at: https://github.com/prabhjotschugh/Not‑Truly‑Multilingual‑PuMVR.
Authors:Sahith Reddy Chada, Isht Dwivedi, Nirav Savaliya
Abstract:
Reliable autonomous driving requires vectorized HD maps that are geometrically accurate, semantically rich, and scalable to long‑horizon driving. However, existing public HD map datasets are limited in scale, provide sparse semantic attributes, and lack modalities such as aerial imagery that could enable new research directions. We present HRDX, a large‑scale dataset for vector HD‑map construction, spanning about 40 hours (1,400 km) of minimally overlapping drives, which is several times larger than prior public HD map datasets. Data is captured using six synchronized surround cameras, a 128‑beam LiDAR, and centimeter‑level RTK GNSS/IMU, and is further complemented by precisely aligned aerial orthoimagery. Annotations cover 10 vector map classes, complemented with over 20 semantic and topological attributes. To evaluate this richer ontology, we introduce the Composite Score (CS) to jointly assess geometric fidelity and attribute correctness. Benchmark experiments show that HRDX's scale improves online vector‑map construction, and that aligned aerial imagery provides a useful structural prior: using aerial imagery at training and/or inference improves geometric map quality, while aerial‑augmented teachers can transfer part of this benefit to camera‑only students without increasing inference‑time sensor requirements. HRDX is intended to support reproducible research on large‑scale HD‑map learning, multimodal BEV fusion, and training‑time privileged information. HRDX dataset and benchmarks are available at https://github.com/honda‑research‑institute/HRDX
Authors:Jisang Han, Seonghu Jeon, Jaewoo Jung, René Zurbrügg, Honggyu An, Tifanny Portela, Marco Hutter, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
Abstract:
Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision‑language‑action models (VLAs) and video world‑action models (WAMs) inherit strong semantic or temporal priors from large‑scale foundation models, but they still operate primarily on 2D image frames or 2D‑derived latent spaces, leaving implicit the 3D geometry required for contact‑rich manipulation. We propose the Geometric Action Model (GAM), a language‑conditioned manipulation policy that directly repurposes a pretrained geometric foundation model (GFM) as a shared substrate for perception, temporal prediction, and action decoding. GAM splits the GFM at an intermediate layer: the shallow layers serve as an observation encoder, and a causal future predictor inserted at the split layer forecasts future latent tokens conditioned on language, proprioception, and action history. The predicted future tokens are then routed through the remaining GFM blocks for feature propagation and decoding, allowing a single backbone to produce both future geometry and actions. This design equips the GFM with language‑conditioned temporal world modeling through minimal architectural modification while preserving its rich geometric priors. Across a broad suite of simulation and real‑robot manipulation benchmarks, GAM is more accurate, more robust, faster, and lighter than current foundation‑model‑scale baselines.
Authors:Shengyu Gong, Weiming Zeng, Yueyang Li, Zijian Kang, Hongjie Yan, Wai Ting Siok, Nizhuan Wang
Abstract:
Non‑invasive brain‑computer interfaces exhibit significant performance degradation when moving from controlled laboratory stimuli to real‑world natural images. This degradation occurs because conventional multimodal contrastive representation learning models focus exclusively on optimizing geometric distance alignment, thereby failing to account for semantic consistency and inter‑subject variability in neural representation and selective attention. As a result, these models are prone to producing spurious zero‑shot matches. To address these limitations, we propose SUP‑MCRL, a unified framework integrating three collaborative mechanisms: (1) a Semantic‑entity Aware Visual Encoder (SAVE) that learns spatial attention to extract semantic content without relying on pre‑trained saliency models; (2) a Unified EEG Enhancer (UEE) that employs multi‑scale atrous convolutions and inter‑band attention for adaptive cross‑subject robustness; and (3) a Prototype‑based Progressive Augmenter (PPA) that maintains an EMA‑updated pseudo‑feature pool to prevent representation collapse. Zero‑shot experiments on the THINGS‑EEG achieve 66.0%/91.9% (Top‑1/Top‑5) intra‑subject and 24.0%/52.9% LOSO accuracy, significantly surpassing state‑of‑the‑art methods and demonstrating that structured alignment supervision is key to overcoming the limitations of cross‑modal decoding. Code is available at https://github.com/NZWANG/SUP‑MCRL.
Authors:Shuai Yang, Bingjie Gao, Ziwei Liu, Jiaqi Wang, Dahua Lin, Tong Wu
Abstract:
Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs struggle to maintain long‑term consistency after such modifications, as stored contexts may become outdated or invalid. To address this, we propose PermaVid, a novel framework built upon a multi‑modal context memory that disentangles spatial context into semantic appearance and geometric structure, together with an edit‑aware memory update and retrieval strategy that keeps memory evolution aligned with subsequent observations. Specifically, we develop two complementary memory banks: an RGB context memory that captures appearance‑aware observations while implicitly encoding geometry, and a depth context memory that preserves geometry‑only structure disentangled from semantics. Building on this design, we introduce a memory‑guided video generation model that performs multi‑modal feature fusion under reference conditions drawn from mixed‑modality memory contexts. Experiments demonstrate that our method maintains strong long‑term semantic and structural consistency after edits, significantly outperforming state‑of‑the‑art methods.
Authors:Pengyu Zhu, Xiaojing Zhang, Kunbo Zhang, Chunyan Zhang, Zhenyu Wang
Abstract:
Medical image segmentation plays a critical role in clinical diagnostics, treatment planning, disease monitoring, and neurological disorder identification. This article presents a comprehensive review of its systematic development, covering widely used public datasets, representative methods built on the U‑Net, Transformer, and SAM architectures, and key evaluation metrics with their differences, followed by an analysis of major challenges from multiple perspectives. Unlike surveys that focus on a single model family or a specific clinical application, this review organizes U‑Net‑, Transformer‑, and SAM‑based methods within a unified analytical framework, with a particular focus on their effectiveness in improving segmentation accuracy and efficiency. This work aims to guide future research and support clinical translation of medical image segmentation, with all related resources publicly available in our GitHub repository: https://github.com/andrew‑pengyu/Awsome_MedSeg/tree/main.
Authors:Adi Ahituv, Anat Ilivitzki, Moti Freiman
Abstract:
Anisotropic volumetric acquisitions are common in clinical MRI and volume electron microscopy (vEM), where sparse through‑plane sampling creates thick slices or sections that degrade orthogonal reformats and downstream analysis. We present CRIS, a cross‑plane self‑supervised framework for isotropic restoration without paired isotropic ground truth. CRIS casts 3D restoration as 2D stripe completion on orthogonal reformats of an isotropic grid: high‑resolution in‑plane slices are synthetically degraded and periodically masked for training, while at inference blank slices define the isotropic grid, two orthogonal reformats are restored, and predictions are fused by multi‑view averaging. We evaluate CRIS on two MRI cohorts and two microscopy benchmarks up to 8x anisotropy. On brain MRI, CRIS achieves 32.921 +/‑ 0.436 dB PSNR and 0.9631 +/‑ 0.0027 SSIM, outperforming interpolation, SMORE4, SIMPLE, SA‑INR, and ATME, and gives the best segmentation consistency (Dice 0.940 +/‑ 0.004, ASSD 0.245 +/‑ 0.014 mm, HD99 1.275 +/‑ 0.061 mm). On reference‑free abdominal MRI, CRIS reduces FID/KID to 48.714/0.023. On vEM, CRIS outperforms interpolation, NIIV, and vEMINR, reaching 29.133 dB/0.834 3D PSNR/SSIM at 4x, 27.123 dB/0.734 on EPFL at 8x, and 21.915 dB/0.699 on noisy hemibrain data. In a robustness experiment, one variable‑gap CRIS model evaluated across gap factors 3‑‑7 and coronal, axial, and sagittal degradations maintained higher PSNR/SSIM than interpolation (36.36‑‑31.14 dB and 0.977‑‑0.932 vs. 33.07‑‑27.85 dB and 0.951‑‑0.853). These results support CRIS as a modality‑flexible route to isotropic restoration without paired isotropic targets or configuration‑specific retraining. Code is available at https://github.com/adi‑hatav/CRIS.
Authors:Zhengyang Shen, Kai-Hung Chang, Erroll Wood, Deying Kong, Bo Peng, Timo Bolkart, Jinlong Yang, Bowen Zhao, Danhang Tang, Sasa Petrovic, Emre Aksan, Jérémy Riviere, Vassilis Choutas, Delio Vicini, Jay Busch, Shichen Liu, Zhe Cao, Hugh Liu, JingJing Shen, Jonathan Taylor, Mingsong Dou
Abstract:
Robust, high‑fidelity 3D hand capture, while fundamental to digital human creation, remains challenging with practical multi‑view systems that balance rich photometry with the geometric ambiguities of reconstruction arising from limited viewpoint density. This paper presents an end‑to‑end pipeline for dynamic hand performance capture and registration, specifically designed for view‑efficient setups (~20 views). We address key challenges with two primary innovations. First, to overcome reconstruction difficulties like limited view overlap and background clutter, our mask‑free neural method robustly extracts detailed hand geometry and appearance from unmasked images using scene parameterization and scenario‑specific density regularization. Second, addressing registration challenges such as accurately capturing non‑linear skin deformations and ensuring plausible results during severe self‑contact, we propose a physics‑inspired framework. It aligns reconstructions to a personalized hand model by optimizing intrinsic volumetric offsets within its canonical tetrahedral mesh, alongside pose parameters. This approach, supported by robust losses and optimization, captures fine surface deformations, ensures plausible results under severe articulation and self‑contact, and demonstrates strong tolerance to input noise. We demonstrate the scalability and robustness of our automated pipeline on an extensive dataset of over 12,000 sequences, from which we also derive a large‑scale, high‑quality synthetic 2D/3D hand dataset for training downstream tasks. This showcases its effectiveness for single hands, intricate two‑hand interactions, and natural hand‑object manipulations. Our method achieves state‑of‑the‑art reconstruction fidelity in view‑efficient, unmasked scenarios and highly accurate registration. Our project page are available at https://vephand.github.io/.
Authors:Zhangfeng Hu, Zefan Yang, Ge Wang, Tanveer Syeda-Mahmood, Anushree Burade, Mannudeep Kalra, Pingkun Yan
Abstract:
Chest X‑ray (CXR) interpretation often requires longitudinal comparison to assess disease progression. Existing approaches typically rely on temporal feature fusion or inter‑study discrepancy modeling, yet remain limited in capturing subtle progression semantics and overlook the inherently directional nature of disease trajectories. In this paper, we propose ProTrans, a novel vision‑language pretraining framework that formulates disease progression as a directional semantic transition between paired CXR studies. ProTrans leverages radiology reports to anchor individual CXR representations within interpretable disease states, and introduces a learnable progression feature map to explicitly encode semantic shifts between states, aligned with report‑derived progression descriptions. To enforce direction‑aware perception, ProTrans incorporates a reversed temporal modeling process and imposes bidirectional reconstruction consistency across states and transitions, thereby disentangling directional semantics and promoting coherent trajectory modeling. Extensive experiments on longitudinal downstream tasks, including disease progression classification and progression captioning, demonstrate that ProTrans consistently outperforms existing methods, establishing a unified pretraining framework for longitudinal CXR understanding. https://github.com/RPIDIAL/ProTrans
Authors:Jyothiraditya Lingam, Nikhileswara Rao Sulake, Sai Manikanta Eswar Machara
Abstract:
We present GOOSE‑M2F, a task‑specific adaptation of Mask2Former for the GOOSE 2D Fine‑Grained Semantic Segmentation (FGSS) Challenge at ICRA 2026. The GOOSE benchmark spans 64 fine‑grained classes across unstructured outdoor terrain with a severely long‑tailed distribution, where rare classes occupy fewer than 50 pixels per image. We extend the Swin‑Large Mask2Former baseline with three targeted contributions: (1) 200 object queries to eliminate representational saturation; (2) a Feature Refinement Module (FRM) combining ASPP‑lite and CBAM dual‑attention; and (3) an Auxiliary Supervision Head that delivers direct per‑pixel gradients for rare classes. A multi‑stage training strategy pairs Distribution‑Balanced loss, Rare‑Class Copy‑Paste augmentation, dynamic IoU‑aware re‑weighting, and EMA. At inference, a dense sliding‑window engine with 2D Gaussian kernel blending and 4‑scale TTA adds +10.57%. GOOSE‑M2F achieves 70.08% Official Composite mIoU (63.55% fine, 76.61% coarse), placing 3rd on the GOOSE 2D FGSS leaderboard. Code and trained models are publicly available at GitHub: https://github.com/Aditya‑Lingam‑9000/GOOSE‑M2F and Hugging Face: https://huggingface.co/XYZ9843/GOOSE‑M2F.
Authors:Bo Peng, Xu Chen, Yi Gu, Hidenobu Matsuki, Mingsong Dou, Jingjing Shen, Deying Kong, Juyong Zhang, Zhengyang Shen
Abstract:
The growing demand for high‑fidelity 4D hand‑object interaction (HOI) data in embodied AI and spatial computing is currently bottlenecked by the reliance on pre‑scanned object templates and physical markers. While recent methods have demonstrated promising results in reconstructing 4D hand‑object interaction from videos, they are highly sensitive to initial estimates of hand and object poses. Yet, estimating these poses from images is challenging, in particular under severe occlusion which is inherent in hand‑object interaction scenarios. We propose a novel system for the robust and accurate reconstruction of hands and objects from synchronized and calibrated multi‑view videos without requiring any templates or markers. Our system consists of two main components with key innovations: (1) a multi‑view feed‑forward transformer model that aggregates cross‑view geometry and temporal cues to provide a reliable, metric‑consistent initialization for both poses and dense object geometry, and (2) a hand‑object physics‑aware Gaussian‑based optimization framework to refine the initial estimates, integrating tetrahedral constraints, collision refinement, and appearance decomposition to produce physically plausible and visually accurate reconstruction. Validated on public benchmarks and an extensive internal dataset, our pipeline achieves highly robust, artifact‑free reconstruction, providing an efficient foundation for automated 4D asset generation. Our project page are available at https://zyshen021.github.io/HOSTPG/.
Authors:Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen, Ke-Yue Zhang, Yue Zhou, Caiyong Piao, Bin Li, Taiping Yao, Bo Wang, Youchang Xiao, Shouhong Ding
Abstract:
Multimodal large language models (MLLMs) have been increasingly adopted in forensics for their robust semantic understanding. As AI‑generated images become realistic, semantic‑level inconsistencies alone are often insufficient for reliable detection. This motivates a critical question: whether MLLMs can achieve full‑spectrum forensic signal perception, i.e., capturing low‑level generator artifacts without sacrificing pre‑trained semantic knowledge. We further perform a layer‑wise analysis of forensic signal perception in MLLMs, showing that semantic information is primarily formed in the early‑to‑middle layers, whereas direct fine‑tuning for artifact learning disrupts these semantic representations. Based on this insight, we propose Deep Visual Residual MLLM (Deep‑VRM) to preserve early semantic processing while injecting artifact‑specific visual signals as a residual path into an intermediate layer, where they are fused with semantic token representations and propagated through subsequent trainable layers. This enables later layers to jointly model semantic reasoning and signal‑level forensic cues, and surprisingly, the model learns to adaptively leverage different levels of forensic signals depending on the input, achieving robust and generalizable detection performance. Extensive experiments show that our method achieves state‑of‑the‑art across most benchmarks. The code and data are available at https://github.com/KQL11/Deep‑VRM.
Authors:Siya Yang, Nanxiang Jiang, Zhaoxin Fan, Yunfeng Diao
Abstract:
The rapid progress of visual autoregressive (VAR) models has unlocked a transformative frontier for high‑fidelity text‑to‑image synthesis, while heightening concerns over the safety alignment of generated content. Naive application of existing erasure techniques to VAR models causes catastrophic semantic collapse and visual artifacts, since they are predominantly designed for the homogeneous denoising steps of diffusion models. To address this foundational challenge, we first propose the Semantic Singularity Axiom, which posits that any target semantic concept embedded within a prompt is definitively locked at Scale‑0. Then rigorously validate this axiom through our proposed Incremental Semantic Saliency Analysis (ISSA),which also enable the community to transparently inspect the coarse‑to‑fine semantic injection process. Guided by this insight, we introduce the first scale‑aware concept erasure framework (SACE) for VAR models. By strictly confining interventions to the first scale, our approach couples an Entropy‑Regularized Erasure Objective to prevent high‑entropy sampling degeneration, alongside a restorative preservation loss to safely anchor the integrity of entangled benign priors. Extensive experiments demonstrate that our method achieves surgical concept erasure performance across various domains with minimal training overhead, timely and elegently resolute the critical safety vulnerabilities inherent in emerging VAR architectures. Code is available at: https://github.com/limerenceysy/SACEhttps://github.com/limerenceysy/SACE.
Authors:Artyom Mazur, Nina Konovalova, Aibek Alanov
Abstract:
Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits. While transcoder‑based circuit tracing has recently enabled detailed causal analyses of large language models, multimodal diffusion transformers for image generation remain comparatively opaque. We still lack tools for understanding how semantic information propagates across denoising steps and how text and image representations interact within double‑stream MM‑DiT architectures. Existing methods provide only partial insight: attention maps expose a limited view of token interactions, while sparse autoencoders can discover interpretable features but do not directly reveal how these features are transformed and composed through nonlinear MLP layers. In this work, we extend transcoder‑based circuit tracing to multimodal diffusion transformers. We train timestep‑conditioned transcoders that faithfully approximate the input‑output behavior of MLP sublayers in FLUX.1[schnell]. By replacing MLPs with transcoders and linearizing the remaining computation, we obtain exact feature‑to‑feature attribution and recover compact, interpretable circuits. Empirically, our transcoders match or slightly outperform sparse autoencoders on the sparsity‑faithfulness tradeoff. The resulting circuits reveal mechanisms underlying attribute binding and cross‑stream semantic propagation, and provide causal explanations for systematic generation errors. Moreover, circuit‑guided interventions are substantially more precise and effective than standard SAE‑based steering. Our results demonstrate that transcoder‑based circuit analysis is feasible for state‑of‑the‑art diffusion transformers and provides a powerful framework for understanding and controlling multimodal generative models. The code is available at https://github.com/Artalmaz31/DifFRACT
Authors:Shuaike Zhang, Shaokun Wang, Haoyu Tang, Jianlong Wu, Liqiang Nie
Abstract:
Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed‑loop control. Compared with conventional continual learning, ECL suffers from more severe catastrophic forgetting. Feature drift accumulated under closed‑loop control progressively propagates through sequential decision‑making, leading to degradation of previously learned behaviors. A key challenge in ECL lies in structured skill reuse across continually evolving tasks, since existing methods primarily focus on skill learning without explicitly organizing them for coherent task execution. To address this issue, we propose SCE, a Skill‑Compositional Experts framework for ECL. SCE builds a skill base via Compositional Skill Grounding (CSG), which decomposes task demonstrations into reusable skills. Based on this, Dual Execution‑and‑Transition Experts (DETE) enable new task learning through skill composition, where one branch ensures skill execution and the other supports transitions between skills for coherent behavior. Experiments on LIBERO benchmarks and real‑world manipulation tasks demonstrate that SCE consistently improves retention and overall task performance. Further feature drift analyses and ablation studies verify the effectiveness of our method. Project website: https://eqcy.github.io/sce/.
Authors:Yiran Wang, Zeyu Zhang, Yuanming Li, Ziming Wang, Yang Zhao
Abstract:
High‑quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital‑human interaction. 3D Gaussian Splatting (3DGS) has emerged as the dominant representation, with two complementary regimes (generalizable feed‑forward predictors and per‑subject refiners) maturing in parallel. However, existing feed‑forward predictors are trained on a single dataset family with a hard‑coded source count, inheriting the corresponding domain bias. Per‑subject refiners require 300K‑‑600K iterations and rely on adaptive densification that destroys upstream Gaussian layouts, preventing the two regimes from sharing a representation end‑to‑end. To bridge both regimes we propose SpatialAvatar‑0 on a shared FLAME‑mesh‑bound Gaussian representation: a feed‑forward generator with a parameter‑free K‑source mean‑pool and a monocular‑temporal to multi‑view‑spatial two‑phase schedule that anchors against identity‑prior collapse onto the smaller multi‑view set. We further introduce a 10K‑iter layout‑preserving per‑subject refinement loop that freezes the FLAME‑binding and Gaussian count and replaces densification with a three‑component anti‑spike regularization. On VFHQ/HDTF cross‑domain zero‑shot we surpass the in‑domain leader GAGAvatar by +1.5 dB PSNR despite never training on either test domain, and on the SplattingAvatar monocular benchmark we lead every reported metric, surpassing the 300K‑iter GeoAvatar by +1.3 dB PSNR at up to 60x shorter per‑subject schedule than common SOTA baselines. Website: https://spatialwalk.github.io/SpatialAvatar‑0.
Authors:Haochen Hu, Yanrui Bin, Zhengyan Zhang, Minchen Wei, Chih-yung Wen, Bing Wang
Abstract:
The underwater images are captured within diverse water‑medium conditions, leading to complex degradation, including color bias, low contrast, and blur effect. Recently, learning‑based methods have demonstrated their potential for underwater image enhancement (UIE). However, most of the previous work focus on the training strategy or network design to make the enhanced result aligned well with the labels in datasets, ignoring that the labels are selected from the enhanced results of previous UIE methods and these pseudo‑labels are noisy. Consequently, the performance of their models is not satisfactory to a certain extent. However, collecting the true labels of the underwater images is challenging. In this work, we propose a transfer learning‑based UIE that does not require underwater images to have paired noisy or true labels for learning. Instead, the UIE task is first divided into global color correction, haze removal, and background noise suppression following the underwater physics. Then multiple types of prior from other vision tasks are leveraged as cross‑domain supervision in each step. In this way, a novel UIE is available via transfer learning, and the physics‑aligned UIE decomposition provides theoretical soundness. Qualitative and quantitative experiments demonstrate that our proposal based on physics and priors fusion achieves SOTA performance in the UIE task and effectively boosts downstream vision tasks, significantly outperforming benchmark methods. Project repo: https://github.com/Haru2022/P2‑UIE.
Authors:Cheng Zhang, Qing Cai, Xingzheng Wu, Xun Yang, Xiaojun Chang, Bingkun Bao, Liqiang Nie, Xinwang Liu, Yi Yang
Abstract:
Foundation models have demonstrated impressive performance in enhancing healthcare efficiency across a wide range of medical applications. Nevertheless, their limited ability to perceive, understand, and interact with the physical world significantly constrains their effectiveness in real‑world clinical workflows, where safety‑critical decision‑making and physical execution are tightly coupled. Recently, embodied artificial intelligence (AI) has emerged as a promising physical‑interactive paradigm for intelligent healthcare, enabling agents to operate in complex medical environments. As research in this area rapidly expands, understanding how intelligent agents function as integrated, end‑to‑end systems in clinical environments becomes increasingly critical. However, existing surveys on medical embodied AI largely emphasize individual aspects or functional components, lacking a unified system‑level organization of the field. To support and consolidate recent advances, we systematically survey the core components of medical embodied AI, with a particular emphasis on the coordinated integration of perception, decision‑making, and action. We further review representative medical applications and relevant datasets, and we analyze the major challenges encountered in real‑world clinical practice. Finally, we discuss key directions for future research in this rapidly evolving field. The associated project can be found at https://github.com/VMVLab/Medical_Embodied_AI_Paper_List.
Authors:Hyunsoo Lee, Farrin Marouf Sofian, Kushagra Pandey, Stephan Mandt
Abstract:
Collaborative generation, which coordinates multiple diffusion trajectories to extend the capabilities of pretrained priors, has emerged as a powerful paradigm for extending the applicability of diffusion models. Among existing approaches, diffusion synchronization provides a scenario‑agnostic solution by introducing general guidance mechanisms. However, current synchronization approaches rely heavily on heuristics and still require task‑specific tailoring, which limits their generalizability and performance. In this work, we mathematically derive a synchronization framework based on optimal control, providing a principled explanation of diffusion synchronization. During sampling, we optimize control variables to guide multiple trajectories toward coherent solutions while remaining close to the underlying diffusion prior. Our method operates entirely at test‑time without additional training, thereby enabling broad applicability across diverse generation scenarios when combined with strong pretrained priors. We demonstrate consistent improvements over baselines on three representative collaborative generation tasks, covering a wide range of modalities and applications. Beyond performance gains, our work establishes a novel foundation for collaborative generation, opening a principled path toward extending pretrained generative models to new collaborative generation settings.
Authors:Fuyou Mao, Beining Wu, Yanfeng Jiang, Bohan Xu, Lixin Lin, Naye Ji, Hao Zhang, Yan Tang
Abstract:
Organ segmentation from PET/CT is critical for quantitative analysis and radiotherapy planning in oncology. To ease the high annotation cost of PET/CT segmentation, semi‑supervised learning (SSL) provides a practical and effective solution for developing deep models with limited labeled data. Recent developments in visual foundation models have demonstrated remarkable adaptability with improved efficiency. In this work, we propose a mutual distillation framework that seamlessly exploits both structural and functional foundation models, which act as modality‑specific generalists for distilling knowledge from structural CT and metabolic PET imaging. By bridging the gap between the task‑specific precision of student models and the segmentation priors of generalist foundation models, we propose MuDuo, a mutual distillation framework that synergistically leverages SAM‑Med3D for CT and SegAnyPET for PET to distill their knowledge into a lightweight student network. Our approach eliminates the need for manual prompts while maximizing the utility of unlabeled data for automatic segmentation, achieving state‑of‑the‑art performance on the AutoPET dataset with only 5 labeled cases. Our source code is available at https://github.com/Wu‑beining/MuDuo.
Authors:Zihan Wang, Guansong Pang, Zelin Liu, Wenjun Miao, Jin Zheng, Xiao Bai
Abstract:
Multimodal Large Language Models (MLLMs) are increasingly used as automated judges, e.g., for image quality and safety assessment. However, their adversarial robustness remains largely unexplored, threatening the fairness and reliability of automated judging. To bridge this gap, we introduce RobustMLLMJudge, the first general framework for evaluating the adversarial robustness of general‑purpose MLLMs when functioning as judges. It covers diverse attacks against popular judge approaches across quality and safety evaluation scenarios. Using RobustMLLMJudge, we reveal that i) different MLLM judges are highly vulnerable to score‑inflating adversarial attacks; and ii) although effective, these attack methods face a critical challenge due to unique constraints in the evaluation protocols of MLLM judges. We further propose MGSIA, namely Manifold‑Guided Semantic Induction Attack, a novel method that bypasses these constraints to enable more effective and transferable attacks on MLLM judges. The core idea of MGSIA is to combine affirmative semantic induction with high‑score manifold alignment: it maximizes the probability that judges yield affirmative responses (e.g., "Yes") to binary semantic queries, while regularizing adversarial representations toward high‑score centers estimated from proxy protocols. Together, these objectives yield transferable score‑inflating perturbations. Extensive experiments demonstrate the superiority and generalizability of MGSIA in deceiving advanced MLLM judges under different evaluation scenarios, highlighting the need for robust MLLM judges. Code and data will be made available at https://github.com/mala‑lab/RobustMLLMJudge.
Authors:Xiongjun Guan, Jianjiang Feng, Jie Zhou
Abstract:
Small‑area fingerprint sensing on mobile devices creates a fundamental mismatch between acquisition and recognition: each touch captures only a tiny, pose‑varying local patch, while reliable biometric matching ultimately requires a stable and sufficiently complete fingerprint representation. Existing pipelines largely cope with this mismatch by treating repeated touches as independent partial templates, which leads to repeated registration, repeated matching, and no guarantee of adequate global coverage. In this paper, we advocate a different formulation, namely \emphaccumulative fingerprint mapping and reconstruction for small‑area mobile sensing. Rather than matching every partial patch separately, the proposed perspective converts a sequence of local observations into a unified fingerprint state that is progressively refined as new touches arrive and can be matched only once after consolidation. As a concrete baseline, we present a classical pipeline that performs patch‑wise structural feature extraction, feature‑level registration and fusion, fingerprint map construction, and phase‑based ridge reconstruction. More importantly, we position this baseline within a broader mobile fingerprint framework that integrates structured token learning, two‑stage pose reasoning, and diffusion‑based generative reconstruction. This viewpoint reframes mobile fingerprint recognition from multi‑capture multi‑match processing to accumulative map building, state refinement, and one‑shot matching, offering a principled route toward efficient, pose‑robust, and deployment‑friendly biometrics for small‑area mobile platforms. The baseline implementation has been publicly released at https://github.com/XiongjunGuan/FpReconstruction.
Authors:Yiwei Ma, Ke Ye, Weihuang Lin, Jiayi Ji, Xiaoshuai Sun, Tat-Seng Chua, Rongrong Ji
Abstract:
In recent years, there have been notable advancements in the area of instruction‑based image editing (IIE), which focuses on the automatic alteration of input images using a model. Nevertheless, assessing the effectiveness of these editing models poses a considerable challenge due to the intricate nature of instructions and the wide variety of edits. To tackle this problem, one urgent task in this domain is the development of a robust evaluation framework that can precisely gauge the quality of editing outcomes and offer valuable benchmarks to guide future improvements. To address this challenge, we present a comprehensive evaluation benchmark named I2EBench2.0, designed for single‑round and multi‑round assessment of IIE models. I2EBench2.0 has four key features: 1) Evaluation Across Single and Multi‑rounds: I2EBench2.0 simultaneously evaluates both single‑round and multi‑round instruction‑based edits, assessing the precision and consistency of the edits. 2) Extensive Evaluation Criteria: I2EBench2.0 encompasses a broad range of criteria, evaluating both high‑level and low‑level aspects of each IIE model. Specifically, it incorporates 16 dimensions for single‑round evaluations and 7 for multi‑round evaluations. 3) Alignment with Human Judgment: To ensure our benchmark aligns with human evaluation, we conducted a comprehensive user study for each criterion. 4) Research‑driven Insights: By analyzing the strengths and weaknesses of current IIE models across all 16 single‑round and 7 multi‑round dimensions, we provide critical insights aimed at directing future research in this area. We tested eight recently developed IIE models using I2EBench2.0 and derived academic insights through meticulous comparison and analysis. The related code, dataset, and images generated by all IIE models are available on GitHub: https://github.com/cocoshe/I2EBench.
Authors:Feng Qiao, Zhaochong An, Zhexiao Xiong, Serge Belongie, Nathan Jacobs
Abstract:
Re‑rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the original scene across every frame. Existing methods rely on per‑frame pose embeddings, noisy point‑cloud renderings, or implicit learned correspondences, none of which provides an explicit, temporally continuous link between source and target pixels. We propose Track2View, which conditions a video diffusion transformer on paired 3D point tracks: sparse trajectories of scene points projected into both the source and target camera views. These tracks provide explicit spatiotemporal correspondences that are temporally continuous by construction, encoding what content should appear where and when. At the core of Track2View is a dual‑view track conditioner that transfers visual context from source to target view through parameter‑free geometric operations and learned temporal aggregation, ensuring generalization to arbitrary camera trajectories without memorizing specific motions. We further introduce a data curation pipeline that extracts one‑to‑one track correspondences by running a 3D point tracker on temporally concatenated multi‑camera view pairs. On a 400‑video benchmark spanning static and dynamic scenes, Track2View achieves state‑of‑the‑art results across visual quality, view synchronization, and camera accuracy, reducing rotation error by 30‑65% and translation error by 61‑72% relative to leading baselines. Project page is available at this https URL: https://qjizhi.github.io/track2view
Authors:Sivaperuman Muniyasamy, Surendar Devasundaram
Abstract:
Vision‑based perception is fundamental to Space Situational Awareness and autonomous on‑orbit operations such as rendezvous, docking, servicing, and navigation. However, progress in this area is limited by the scarcity of annotated space imagery and by challenging visual‑domain characteristics including severe illumination changes, low signal‑to‑noise ratio, and high contrast. We address Stream 1 of the SPARK 2026 Challenge, which requires a single model for spacecraft classification, detection, and fine‑grained component segmentation across multiple target types. We propose a compact architecture that integrates a MobileNetV3 encoder with a U‑Net‑style decoder, combining computational efficiency with accurate dense prediction. Detection is derived analytically from the union of predicted component masks, avoiding a separate bounding‑box regression head in the single‑spacecraft setting. Our method achieved an overall leaderboard score of 0.9482, with task‑specific scores of 1.0000 in classification, 0.9788 in detection, and 0.8917 in segmentation. The proposed approach ranked second overall in the SPARK 2026 Challenge, demonstrating that lightweight encoder‑decoder architectures can deliver strong multi‑task performance for practical onboard space vision systems.
Authors:Mengshi Qi, Changsheng Lv, Zijian Fu, Xianlin Zhang, Huadong Ma
Abstract:
In this paper, we propose SGFormer++, a novel Semantic Graph Transformer for 3D scene graph generation (SGG), which aims to parse point cloud scenes into semantic structural graphs, where nodes denote detected object instances and edges encode their pairwise relationships, with the core challenge lying in modeling complex global scene structure. While existing graph convolutional network (GCN)‑based methods suffer from over‑smoothing and limited receptive fields, SGFormer++ leverages Transformer layers as its backbone to enable global message passing. Specifically, we introduce two key components tailored for 3D SGG: (1) a Graph Embedding Layer++ that efficiently integrates edge‑aware global context with linear computational complexity, and (2) a Semantic Injection Layer++ that enriches visual features with linguistic priors from large language models (LLMs) and vision‑language models (VLMs), boosting semantic representation without introducing extra trainable parameters. To further address the practical challenge of incremental SGG (I‑SGG), where new relationship categories arrive sequentially, we equip SGFormer++ with a novel Spatial‑guided Feature Adapter, which calibrates predicate features using subject‑object spatial geometry to counter scale variation, and a Cascaded Binary Prediction Head that mitigates catastrophic forgetting via task‑incremental classifier expansion and logit distillation. Extensive experiments on the 3DSSG benchmark demonstrate that SGFormer++ achieves state‑of‑the‑art performance in both standard and incremental settings: it yields a significant 4.49% absolute improvement in Predicate A@1 under the incremental setting. Code and data are available at: https://github.com/Andy20178/SGFormer.
Authors:Mustafa Bora Çelik
Abstract:
Fine‑grained action recognition demands temporal reasoning that general‑purpose architectures address through different cost‑accuracy tradeoffs: 3D dense operators couple computation to the input volume, while difference‑based methods approximate motion through rigid, hand‑crafted subtraction of uncontextualized features ‑ each reflecting a deliberate design choice with corresponding limitations in expressiveness or flexibility. We present MamBOA, a backbone‑agnostic temporal framework built upon a novel interleaved scan structure that recasts the selective state‑space recurrence (S6) as a native motion synthesizer. By interleaving consecutive feature representations extracted from a pretrained backbone into a single alternating sequence, the proposed scan structurally drives the recurrence to encode both temporal observations of each position within a shared hidden state, separated by only a single decay step ‑ rendering the inter‑frame transition an intrinsic component of the state dynamics rather than an externally computed quantity. A cascade of dedicated alignment and decoding operations then distills this joint encoding into an explicit motion representation, which a dual‑path pooling mechanism adaptively aggregates by balancing attention‑driven selection with uniform temporal coverage. The framework interfaces seamlessly with CNN, Transformer, and Mamba backbone families, adding only ~2.1 GFLOPs per feature pair. On Diving48, MamBOA achieves 85.02% Top‑1 accuracy with an image‑pretrained backbone and 86.24% with a video‑pretrained backbone processing the entire video in a single forward pass ‑ demonstrating that structurally induced state‑space dynamics constitute a principled and general foundation for motion modeling.
Authors:Weichen Fan, Haiwen Diao, Penghao Wu, Ziwei Liu
Abstract:
Pixel‑space diffusion models are trained on full‑bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified‑flow diffusion and natural‑image power‑law spectra, the per‑band data‑to‑noise contour k^(t) = (1‑t)^‑2/α separates a signal‑bearing low‑frequency region from a noise‑dominated high‑frequency region at each time t. We show that this implicit coarse‑to‑fine structure is not merely descriptive: it induces a capacity‑allocation problem. A standard pixel‑space denoiser must discover the moving bandwidth boundary internally and can spend computation on frequency‑time regions where the optimal prediction collapses to deterministic baselines rather than data‑distribution modeling. To make this boundary explicit, we introduce Spectral Forcing, a parameter‑free, time‑conditional 2D‑DCT low‑pass operator applied to the noisy input before the patch embedder. Its cutoff expands monotonically with the diffusion time and becomes the identity at the data endpoint. Through controlled synthetic experiments, we identify the regime in which the operator is beneficial: coarse patch tokenization and data whose high‑frequency content is predominantly noise rather than essential signal. On ImageNet‑256 with JiT‑700M/32, Spectral Forcing consistently improves both FID and Inception Score across different training epochs, demonstrating robust gains throughout training; at finer tokenization, the spectral forcing is still competitive. We further insert the unchanged operator into SenseNova‑U1, a unified text‑to‑image model, where it improves DPG‑Bench and GenEval, showing that the input‑side spectral prior transfers beyond class‑conditional generation. These results suggest a route to capacity‑efficient pixel‑space diffusion by showing the signal and hiding the noise.
Authors:Yun Wang, Junbin Xiao, Han Lyu, Yifan Wang, Jing Zuo, Zhanjie Zhang, Hong Huang, Dapeng Wu, Angela Yao
Abstract:
We introduce UCS‑Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User‑Centric Continual Spatial intelligence in egocentric video streams. UCS‑Bench targets a new problem that emphasizes dynamic spatial reasoning, long‑term memory, and their alignment with users' real‑time locations. We propose DirectMe, a framework that incrementally constructs and maintains a structured spatial memory from streaming egocentric observations. DirectMe enables robust tracking and recall of object locations, all relative to the user's movement over time. By tightly coupling visual perception with memory updates and spatial reasoning, our approach supports long‑horizon queries that require recalling interactions, resolving viewpoint‑induced ambiguities, and adapting to dynamic scenes. Our experiments show that DirectMe significantly improves the spatial reasoning of leading multimodal LLMs; it also surpasses many spatially aware and long‑form streaming video models. We hope our benchmark and solution will advance spatial intelligence research for egocentric AI assistants. Data and code are available at https://github.com/cocowy1/UCS‑Bench.
Authors:Jeahun Sung, Dahyeon Kye, Soo Ye Kim, Jihyong Oh
Abstract:
Reference‑guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation: the object‑centric high‑resolution reference image (HRRI) provided by users is downsampled to a fixed low‑resolution (LR) before being fed into the model, so the fine‑grained details are discarded before the output is even produced. In addition, the generation step then introduces its own artifacts (e.g., identity distortion) on top of this loss. Existing reference‑guided generated content refinement (RefGCR) methods can correct some of these artifacts but still operate in the LR domain; reference‑guided super‑resolution (RefSR) methods recover resolution but assume natural‑image degradations and ignore the artifact distribution of generative pipelines. To address both gaps in a single formulation, we introduce a new task: reference‑guided generated content super‑resolution‑refinement (RefGC‑SR^2), where the original HRRI is reused at the post‑processing stage to recover lost details, refine generative artifacts, and upscale the output simultaneously. We construct the first real‑world triplet data generation pipeline for this RefGC‑SR^2 task, training a diptych‑conditioned generator to synthesize paired low‑quality anchors that public pretrained models cannot provide. We further present a frequency‑aware diffusion transformer model for RefGC‑SR^2 that selectively injects fine details from the HRRI while removing generative artifacts. Extensive experiments demonstrate that our RefGC‑SR^2 model successfully (i) refines the object identity faithfully with respect to the reference, and (ii) recovers high‑resolution details, so that the final result is significantly higher quality and practically more usable compared to existing RefGCR and RefSR baselines.
Authors:Nonghai Zhang, Siyu Zhai, Yanjun Li, Zeyu Zhang, Zhihan Yin, Yandong Guo, Boxin Shi, Hao Tang
Abstract:
Generating realistic humanoid motion from scene images and text involves both low‑frequency pose semantics and high‑frequency physical dynamics. However, many existing methods tokenize motion with a single shared codebook, forcing heterogeneous motion signals into the same quantization space. Our frequency‑domain analysis of human motion data reveals a clear mismatch between single‑codebook quantization and motion statistics: five DCT coefficients capture 93% of joint‑position energy but only 37% of joint‑velocity energy, which can bias quantization toward pose statistics and under‑represent high‑frequency velocity components. A second challenge lies in adapting a standard autoregressive model to effectively model high‑frequency physical signals in motion sequences. Therefore, we propose DSFT, a dual‑stream frequency tokenizer that separates motion into Base and physical streams and compresses them independently with DCT truncation and BPE. Furthermore, we present MotionVLA, a Qwen3.5‑based model that arranges Base and physical tokens in a unified sequence, where Phys tokens are predicted after Base tokens. Experiments on HumanML3D and MBench show that, despite using a lightweight 2B backbone, MotionVLA reduces the Diversity gap to real data by over 50% on HumanML3D and improves Motion‑Condition Consistency by 3.8% on MBench, supporting frequency‑aware dual‑stream decoupling as an effective formulation for autoregressive motion generation. Code: https://github.com/AIGeeksGroup/MotionVLA. Website: https://aigeeksgroup.github.io/MotionVLA.
Authors:Tianshan Zhang, Yijia Duan, Yanjun Li, Zeyu Zhang, Hao Tang
Abstract:
Dexterous interaction with articulated objects is important for household, assistive, and humanoid manipulation, where multi‑finger hands can provide compliant contact patterns beyond parallel‑jaw grasping. However, articulated‑object manipulation differs from static‑object manipulation: the target part cannot be directly actuated, and its motion must emerge through sustained physical hand‑‑handle contact. This makes the transition from object‑centric articulated generation to hand‑driven dexterous hand‑‑object interaction non‑trivial, since geometric trajectory replay or open‑loop execution does not model the contact dynamics required to move the articulated part. Moreover, policies trained only for task completion under fixed dynamics can overfit nominal contact loads, especially without tactile or force feedback, and may degrade when the contact load changes. To address these challenges, we present DragMesh‑2, a contact‑driven framework for dexterous interaction with articulated objects that extends articulated interaction from object‑centric generation to hand‑driven dexterous hand‑‑object interaction, where articulated motion must arise through physical contact. We further propose PICA, a physically informed contact‑aware training mechanism that injects physical signals into policy learning without tactile or force feedback, improving robustness and task success under changing contact loads. Finally, we conduct systematic evaluation across multiple damping conditions and articulated‑object categories to study robustness under contact‑load variation, and provide a pure‑geometry dexterous interaction resource to support future loco‑manipulation and humanoid hand‑‑object interaction research. Across seven GAPartNet objects, DragMesh‑2 achieves stronger robustness under contact‑load variation than the compared methods while maintaining high task success across damping conditions.
Authors:Weilong Guo, Yuhan Sun, Shengyang Li
Abstract:
Weak objects are common in images and videos of space applications. However, it is hard to learn proper representations from their limited appearance information. Inspired by multi‑view learning, we develop simple multi‑view attentions, treating their outputs as multi‑view features. We also propose a multi‑view feature high‑order fusion method (MHF) to aggregate more accurate and richer features of weak objects. Our MHF extends the commonly used low‑order feature fusion method to higher orders. It enhances the model's capacity to capture relevant and complementary information about weak objects. This is achieved by introducing high‑order multi‑view features perception and a recursive task‑contribution gated selection of multi‑view features. The new operation is highly flexible and customizable. It is compatible with various variants of multi‑view feature representations. We conduct extensive experiments on two newly constructed space science datasets and an open, large‑scale satellite video dataset. Our MHF serves as a plug‑and‑play module and significantly improves various vision transformers and convolution‑based detection and segmentation models. We achieve all state‑of‑the‑art accuracies on both tasks across three datasets. Our MHF can be a new basic module for visual modeling that effectively represents weak objects in terms of multi‑view learning. The code will be available at https://github.com/Kingdroper/MHF.
Authors:Lingtong Zhang, Wenlei Li, Mu He, Li Xiao, Yang Ji
Abstract:
Zero‑Shot Self‑Supervised Learning (ZS‑SSL) has emerged as a promising paradigm for accelerated Magnetic Resonance Imaging (MRI) reconstruction, eliminating the reliance on fully‑sampled external datasets. However, learning solely from a single under‑sampled scan suffers from supervision scarcity and optimization instability, often leading to overfitting or artifacts. To address these challenges, we propose a robust physics‑driven ZS‑SSL framework that synergizes physical consistency with image‑domain non‑local priors. Our method introduces three core innovations: (1) a Coil Sensitivity Map (CSM)‑Guided Dynamic Repository, which stabilizes the training trajectory by filtering physically inconsistent artifacts based on coil sensitivity constraints; (2) a SPIRiT‑based regularization, which enforces k‑space self‑consistency via a learned correlation kernel and stochastic masking; (3) a Non‑Local Self‑Similarity (NSS) Pixel Bank, which leverages the high‑fidelity reference established by the former modules to explicitly mine non‑local anatomical similarities, thereby augmenting supervision in the image domain. Extensive experiments on the FastMRI dataset demonstrate that our approach achieves state‑of‑the‑art performance, particularly under high acceleration factors, effectively bridging the gap between zero‑shot learning and supervised methods. The code is available at https://github.com/Zolento/NS‑SSL.
Authors:Huan Kang, Hui Li, Tianyang Xu, Tao Zhou, Xiao-Jun Wu, Josef Kittler
Abstract:
Infrared and visible image fusion aims to integrate complementary modalities, while existing Euclidean methods impose rigid distance metrics that distort multi‑modal interactions and parent‑to‑child semantic hierarchies. To overcome these limitations, we introduce a text‑driven fusion framework empowered by hyperbolic manifold learning. During training, BLIP‑extracted text prompts serve as topological anchors within the hyperbolic space, guiding vision‑attribute alignment through hyperbolic embeddings that naturally accommodate varying semantic granularities. By exploiting the exponential volume growth dictated by the Poincaré ball's negative curvature, this approach seamlessly embeds hierarchical trees to encode coarse‑to‑fine semantics without metric saturation, while the vast peripheral space prevents texture distortion during cross‑modal fusion. At inference, the fusion process autonomously adapts to input content using the learned text‑attribute priors, completely eliminating the need for textual input. Experimental results show our method outperforms state‑of‑the‑art approaches on benchmark datasets, with code available at https://github.com/Shaoyun2023/TEDFusion.
Authors:Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao, Deep Shah, Manan Dey, Shriya Sudhakar, Zacharie Bugaud, Wissam Siblini, Ayush Sunil Munot, Yashwanth Devavarapu, Rakshitha Ireddi, Michelle Yang, Márton Kardos, Niklas Muennighoff, Kenneth Enevoldsen
Abstract:
We introduce the Massive Video Embedding Benchmark (MVEB), a 23‑task benchmark for video embeddings spanning classification, zero‑shot classification, clustering, pair classification, retrieval, and video‑centric question answering. We evaluate 33 models and find that no single model dominates: MLLM‑based embeddings lead on classification, clustering, pair classification, and QA; multimodal binding leads on retrieval and zero‑shot classification; generative MLLMs without contrastive adaptation collapse on cross‑modal tasks. Paired video‑only vs. audio+video evaluations show that audio's contribution depends on dataset annotation provenance: audio helps when labels were produced from both modalities and hurts when they were produced from visuals alone, a six‑point gap consistent across model families. MVEB is derived from MVEB+, a 184‑task pool, and is designed to maintain task diversity while reducing evaluation cost. It integrates into the MTEB ecosystem for unified evaluation across text, image, audio, and video. We release MVEB and all 184 tasks along with code and a leaderboard at https://github.com/embeddings‑benchmark/mteb.
Authors:Yizhao Huang, Haoyang Chen, Shiqin Wang, Pohsun Huang, Jiayuan Li, Haoyuan Du, Yandong Shi, Zheng Wang, Zhixiang Wang
Abstract:
This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision‑Language Models (VLMs). Implicit reasoning refers to the ability to autonomously discover and utilize hidden visual evidence to bridge information gaps, rather than merely relying on explicitly specified targets. This capacity underlies human visual understanding and everyday reasoning. We argue that this limitation arises from a tendency to approach visual reasoning primarily as passive semantic retrieval, rather than as active, situated reasoning that depends on autonomous visual exploration. As a result, most existing benchmarks primarily assess Passive Capacity, leaving this aspect of reasoning largely unmeasured. To address this gap, we introduce the Visual Implicit Reasoning Diagnosing Benchmark (V‑IRD), which targets this missing quadrant by requiring models to derive answers strictly through autonomous visual analysis. Our results show that, despite strong retrieval abilities, prominent VLMs struggle to utilize reference objects and to attend to visual evidence that requires self‑directed inquiry. Simply put, strong semantic recognition does not equate to active visual exploration, revealing a critical gap in current VLMs. More information can be found at https://haoychen.github.io/Implicit‑Reasoning/
Authors:Tianhao Chen, Yuheng Wu, Kelu Yao, Xiaogang Xu, Xiaobin Hu, Dongman Lee
Abstract:
Multimodal Large Language Models (MLLMs) achieve strong vision‑language reasoning, but long visual contexts enlarge the KV cache and increase decoding latency. Existing compression methods rely on observation window attention for stable token‑importance estimation, yet this aggregation can dilute sparse visual evidence and discard answer‑critical tokens under aggressive compression. Therefore, we identify last‑query attention as a complementary source for recovering such evidence, but its answer‑irrelevant signals can mislead retention. We propose BACON, a plug‑and‑play method that calibrates observation window attention with last‑query evidence and suppresses isolated noise via intra‑layer coherence and inter‑layer persistence. Across diverse benchmarks, models, budgets, and compression methods, BACON improves multimodal KV compression by 7.5% on average under the most aggressive budget, with gains up to 30.9%. Our project page is available at https://ryu1ion.github.io/official_BACON/
Authors:Daniel Torres, Julia Navarro, Catalina Sbert, Joan Duran
Abstract:
Underwater imaging plays a crucial role in ocean engineering, although captured data often suffer from poor visibility and color distortion. To address these challenges, we propose a model‑based deep unfolding network for underwater image enhancement that integrates variational modeling into a learnable architecture. The framework is guided by a variational formulation based on a dehazing decomposition, incorporating a multiplicative residual component to absorb remaining artifacts and a nonlocal gradient‑type constraint to preserve structural details and enhance edge sharpness. We provide a theoretical analysis establishing the existence of solution for the associated minimization problem. The proposed unfolding method incorporates Mamba layers to efficiently capture self‑similarities in the scene. In addition, we introduce a proximal trajectory loss that enforces consistency between the unfolding stages and the iterations of an ideal restoration regularizer. Experimental results demonstrate that the proposed unfolding approach achieves improved visual quality and competitive quantitative performance compared with recent state‑of‑the‑art methods. The source code will be available at https://github.com/MIA‑UIB/Variational‑Unfolding‑Mamba‑Underwater‑Enhancement .
Authors:Jinwen Wen
Abstract:
We present Double‑Helix Vision (DH), a geometry‑based visual sampler that compresses 2D images into compact 1D signals using paired golden‑ratio‑inspired spiral trajectories. Rather than processing every pixel uniformly, DH employs two phase‑shifted helices (Alpha and Beta, offset by 180 degrees) to sample the image with biologically‑inspired foveation: high density at the center, sparse coverage at the periphery. At 4K resolution, DH achieves a 1,433x compression ratio (99.93% reduction) while preserving the geometric structure of the scene. The full perception pipeline ‑‑ including spatial mapping, temporal collision detection, and intra‑frame structural disparity estimation ‑‑ runs in 0.52 ms at 1080p on CPU‑only hardware, with no neural network dependencies. On CIFAR‑10 at extreme sampling budgets (K=128 points per helix), DH achieves a +6.03% accuracy gain over uniform random sampling. A JSON‑serializable Robotics API is provided, delivering sub‑millisecond spatial perception reports in 2.7 KB packets. Code and benchmarks are available under the MIT License.
Authors:Emirhan Bilgiç, Baptiste Caramiaux, Zhi Yan, Gianni Franchi
Abstract:
As Vision‑Language Models are increasingly deployed in safety‑critical applications, the trustworthiness of their explanations becomes crucial. Explainable AI (XAI) methods for Vision‑Language Models often suffer from semantic hallucination, where attribution maps highlight prominent image regions even when prompted with incorrect text descriptions (e.g., highlighting a dog when prompted ``cat''). Although this problem is widespread, a formal mathematical analysis of XAI methods and CLIP embeddings is largely missing in the literature. We demonstrate that this phenomenon is not specific to a single architecture but is a fundamental consequence of Linear Semantic Leakage in high‑dimensional embedding spaces. We propose a unified theoretical framework, Linear Semantic Attribution (LSA), which generalizes across discriminative methods. We introduce OSP, a geometric intervention that utilizes the residual property of OMP to disentangle unique semantic signals from shared concepts. We prove theoretically and demonstrate empirically that OSP minimizes hallucination by orthogonalizing the query vector against distractor concepts, rendering the attribution model blind to shared features while preserving fidelity for correct prompts. Our code is available at: https://github.com/emirhanbilgic/Orthogonal‑Semantic‑Projection
Authors:Nadav Orenstein, Aviad Cohen Zada, Shai Avidan, Gal Oren
Abstract:
Texture segmentation stresses foundation segmentation because meaningful regions are defined by material or repeated appearance rather than object identity. Segment Anything Models (SAMs) often fail by default on such texture‑defined partitions, but this failure is ambiguous: the texture evidence may be absent, missing from the proposal bank, or present but selected or assembled incorrectly by an object‑centric readout. We ask what texture‑relevant evidence is already preserved in frozen SAM before adaptation. We study two frozen evidence spaces: multiscale features, probed with a minimal clustering readout, and the automatic proposal bank, treated as evidence for a supervised consolidation readout. SAM is frozen throughout; we do not fine‑tune the backbone or retrain the proposal generator. Across RWTD, STLD, an ADE20K‑selected refined‑crop complement, and a ControlNet‑stitched PTD bridge archive, frozen SAM is not a texture segmenter by default, but its failures are not simple texture blindness. Coarse frozen features preserve texture organization, and proposal banks often contain texture‑aligned masks or fragments. Natural scenes more often require assembly and commitment over fragments, while cleaner synthetic cases more often reduce to selecting an already coherent proposal. Default mask failure should therefore be decomposed into representation evidence, proposal‑bank support, readout mismatch, and commitment failure.
Authors:Aviad Cohen Zada, Nadav Orenstein, Shai Avidan, Gal Oren
Abstract:
Images can be segmented based on visual cues (i.e., texture segmentation) or into objects (i.e., semantic segmentation). We propose a new category of sub‑semantic image segmentation that blurs the line between the two. In sub‑semantic image segmentation, language is not used to name whole objects. Instead, it is used to partition an image into stable appearance patterns that can be described by language. To do that, we couple a general‑purpose vision‑language model to SAM 3, a promptable segmentation backbone whose native text pathway can ground rich descriptions into masks. Simple coupling fails for a number of reasons that we identify in the paper, and we overcome them by introducing DETECTURE that resolves three concrete failure modes ‑‑ language leakage between texture regions, prompt competition inside the segmentation backbone, and semantic distortion at the language‑to‑mask interface. Since there is no dataset of sub‑semantic image segmentation, we introduce one, termed TextureADE. The new dataset is derived from the ADE20K dataset using a system we designed. We compare DETECTURE to a number of baselines and find that it achieves the strongest performance on several datasets using different metrics. Code is available at https://github.com/Scientific‑Computing‑Lab/TextureDetecture.
Authors:Xirui Kang, Yanpei Shi, Lucy Liang, Roy Gan, Dongxiu Liu, Pushi Zhang, Danpeng Chen, Xiaoyi Qin, Yinan Zheng, Jinliang Zheng, Hao Wang, Xianyuan Zhan, Hang Su
Abstract:
Modern Vision‑Language‑Action (VLA) models must bridge pretrained vision‑language reasoning and precise continuous robot control. Existing action tokenizers discretize actions primarily for reconstruction, producing codes that preserve motion geometry but provide only weak semantic supervision to the backbone. We therefore formulate action tokenization not as mere compression, but as semantic interface learning between multimodal reasoning and executable control. To this end, we introduce X‑Tokenizer, a lightweight encoder‑Semantic Residual Quantization (SRQ)‑decoder architecture that provides a shared action interface across diverse robotic arm embodiments. Its key component, SRQ, imposes an asymmetric structure on residual vector quantization: the first level is trained with Masked Action Modeling (MAM) to form a discrete action language that captures coarse motion intent, while deeper levels remain reconstruction‑oriented residuals that preserve fine‑grained details. To further align action tokens with multimodal semantics, X‑Tokenizer is pretrained with contrastive alignment to the representation space of a pretrained foundation model and with next‑frame vision‑language feature prediction. Pretrained on 2.4M trajectories (2.0B action frames), a single frozen X‑Tokenizer plugs into a mixed discrete‑continuous VLA as a representation‑shaping supervision signal. X‑Tokenizer achieves top real‑world aggregate and strong RoboTwin 2.0 simulation results. Outperforming FAST in multimodal grounding (+13.5%) and long‑horizon tasks (+8.25), it shows that action tokenizers serve as semantic interfaces for VLA pretraining beyond mere action compression.
Authors:Shiwen Zhang, Haoyuan Wang, Xianghao Zang, Haibin Huang, Chi Zhang, Xuelong Li
Abstract:
Content‑Preserving Style transfer, given content and style references, remains challenging for Diffusion Transformers (DiTs) due to entangled content and style features. With a reverse triplet synthesis pipeline to build a million‑scale training set and a dual‑branch Style‑Content DiT (SC‑DiT) that decouples style and content via separate ROPE embeddings and causal masking, we observe that such a one‑stage training paradigm on mixed style categories causes semantic styles to dominate, hindering texture style learning, and harming content preservation. To address these issues, we propose Style‑CCL, a Multi‑Stage Curriculum Continual Learning framework that trains SC‑DiT from semantic (easy) to texture (hard) styles, and from clean to synthetic data, with Random Memory Rehearsal across stages to avoid catastrophic forgetting. Extensive experiments demonstrate that our Style‑CCL achieves state‑of‑the‑art performance in three core metrics: style similarity, content consistency, and aesthetic quality.
Authors:Romiyal George, Sathiyamohan Nishankar, Selvarajah Thuseethan, Roshan G. Ragel
Abstract:
Vision Transformers (ViTs) have demonstrated strong representation capability in image classification. However, their quadratic self‑attention complexity and large parameter counts limit deployment on resource‑constrained mobile and edge devices. This paper introduces UtVAA, an ultra‑tiny Vision Transformer architecture designed for efficient visual recognition under strict computational budgets. It incorporates a novel Affix Attention block that combines depthwise‑pointwise local feature extraction, linear self‑attention, coordinate attention for spatial dependency modelling, and a lightweight ternary fusion strategy to integrate local and global representations. In addition, Dilated Bottleneck blocks expand the receptive field using dilated depthwise separable convolutions while maintaining low FLOPs and stable optimisation through residual connections. UtVAA is implemented in scalable Tiny, Medium, and Large variants, with the smallest model containing 204.67K parameters and 53.95M FLOPs. Experimental results on CIFAR‑10, CIFAR‑100, PlantVillage‑Tomato and SLIF‑Tomato datasets show that UtVAA achieves competitive accuracy within a sub‑million‑parameter regime. Overall, the results demonstrate that transformer‑based vision models can be redesigned into ultra‑tiny architectures without significant loss in discriminative performance, making UtVAA suitable for mobile and edge deployment. Code is available at https://github.com/romiyal/UtVAA
Authors:Matiur Rahman Minar, Seunghun Oh, GangHyeon Jeong, Unsang Park
Abstract:
Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene layouts drift, while mechanisms that improve spatial stability tend to suppress motion, causing natural flows such as water, fire, or smoke to stagnate. We study this stability‑motion trade‑off in fixed‑camera long‑horizon nature video generation, where the two failure modes can be more clearly separated than in moving‑camera settings. We propose Steady‑Forcing, a memory and training framework combining a persistent visual anchor (V‑Sink), an exponential moving‑average motion memory (EMA‑Sink), block‑relative temporal encoding, periodic cache purification, and distillation from a Wan2.1‑14B teacher with motion‑rewarded priors under task‑focused configurations. Together, these components are designed to preserve background identity while sustaining visually plausible fluid dynamics over multi‑minute autoregressive rollouts. Evaluations across seven baselines show that Steady‑Forcing improves long horizon background consistency and imaging quality, while a blind user study indicates stronger perceived stability and motion continuity. The benchmark evaluation further suggest that generic VBench aggregate scores under‑penalize fixed‑camera artifacts as well as rewarding drift‑induced optical flow as Dynamic Degree while not directly penalizing texture hardening or flow stagnation ‑ motivating future task‑specific benchmarks for static‑camera nature‑flow evaluation. Project page: https://minar09.github.io/steadyforcing/
Authors:Xinyue Cai, Chaoyou Fu, Yi-Fan Zhang, Ran He, Caifeng Shan
Abstract:
Current automated pipelines for audio‑visual Question Answering (QA) generally adopt a ``video‑caption‑QA'' paradigm. However, these methods typically segment videos into short clips and generate separate descriptions for audio and visual modalities. This decoupled processing severs inherent associations between sounds and their visual sources, while independent clip processing often causes inconsistent descriptions of the same entity across segments. Furthermore, coupling long‑text comprehension and QA synthesis into a single step often restricts models to localized events, yielding questions lacking long‑term temporal connections and deep cross‑modal reasoning. To address these issues, we propose an automated data engine featuring two mechanisms: (1) Entity‑Anchored Video Scripting transforms videos into structured scripts, comprising summaries, main entity lists, and segment‑wise audio‑visual descriptions. The entity list serves as a global prior to ensure cross‑segment referential consistency and reconstruct audio‑visual associations. (2) Clue‑Guided QA Generation prompts models to first mine cross‑segment, multimodal clues from the script, and subsequently generate QA pairs based on these high‑value clues. Leveraging this pipeline, we construct the instruction‑tuning dataset OmniVideo‑100K and a human‑verified test set, OmniVideo‑Test. Fine‑tuning VITA‑1.5, Qwen2.5‑Omni‑7B and Qwen3‑Omni‑30B on OmniVideo‑100K yields performance gains of up to 20.59% on OmniVideo‑Test, demonstrating strong generalization (up to 12.64% improvements) across established benchmarks like Daily‑Omni and JointAVBench.
Authors:Sicheng Yang, Hangjie Yuan, Wenjun Zhang, Jinwang Wang, Yichen Qian, Weihua Chen, Fan Wang, Lei Zhu
Abstract:
Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision support. Existing medical hallucination benchmarks mainly focus on data collection, but often ignore where hallucinations originate within the reasoning process. We find that hallucination sources vary across samples: errors may arise from visual misrecognition, incorrect medical knowledge recall, or flawed reasoning integration. To enable source‑level hallucination diagnosis, we introduce ClinHallu, a benchmark for stage‑wise hallucination diagnosis in medical MLLM reasoning. ClinHallu contains 7,031 validated instances, where each instance is augmented with a structured reasoning trace decomposed into Visual Recognition, Knowledge Recall, and Reasoning Integration. We also use stage‑replacement interventions to measure how correcting specific stages affects the final answer. Beyond evaluation, we show that trace‑supervised fine‑tuning reduces stage‑wise hallucinations. ClinHallu provides a fine‑grained hallucination testbed for diagnosing and mitigating reasoning failures in medical MLLMs. The benchmark is publicly available at https://github.com/alibaba‑damo‑academy/ClinHallu.
Authors:Xuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Qingqi Hong
Abstract:
Long‑form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing plausible next‑shot continuations without verifying whether the historical memory preserves identity‑critical subject evidence. Consequently, as generation proceeds, recurring subjects may be diluted, overwritten, or forgotten. In this paper, we propose Memento, a subject‑reconstruction‑guided framework that treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone. Specifically, Memento jointly trains autoregressive next‑shot generation with memory‑based subject reconstruction, recovering target appearances using historical memory and global story captions. To disentangle long‑range subject evidence from short‑range cues, Memento introduces a dual‑query memory mechanism, where one query retrieves identity‑relevant memory and the other selects short‑context keyframes for coherent continuation. Additionally, a subject‑aware cinematic data pipeline provides precise reconstruction supervision via consistent, pronoun‑free subject descriptions. Experiments demonstrate that Memento achieves state‑of‑the‑art performance in long‑term subject consistency, cross‑shot coherence, and visual quality.
Authors:Yijun Liu, Jie Huang, Zeyue Xue, Yuming Li, Ruizhe He, Haoran Li, Shijia Ge, Siming Fu
Abstract:
Reward models guide text‑to‑image (T2I) systems toward outputs aligned with human preferences. However, typical reward models such as HPSv3 are trained on pre‑annotated data from earlier T2I models, without accounting for quality discriminative shifts arising from evolving model capabilities and reinforcement learning (RL) iterations, limiting their broader applicability. In this work, we propose HPSv3++, a reward model framework that elevates the HPSv3 model for varying T2I model capabilities and their RL iteration changes across the full capability‑iteration spectrum. Specifically, we first introduce HPDv3++, a 212K dual‑dimension preference dataset annotated for text fidelity and aesthetic quality using a recent high‑capability (Qwen‑Image) model with human supervision. We then propose a two‑stage training framework. Stage 1 employs data‑aware orthogonal gradient projection to incorporate diverse aesthetic perception from HPDv3++ while preserving the original effective human preference knowledge in HPSv3. Stage 2 further leverages unlabeled data from T2I models spanning different capability levels and RL iterations, and introduces a joint capability‑iterations conditioned signal for the reward model together with a standard deviation‑driven unsupervised guidance mechanism, strengthening reward model across the capability‑iteration spectrum. HPSv3++ achieves state‑of‑the‑art preference prediction, outperforming HPSv3 9.8% on HPDv3, 5.5% on GenAI‑Bench, while achieving 79.1%/88.1% on our proposed HPDv3++. When used for T2I RL training, it consistently improves GenEval scores across diverse T2I models, demonstrating its wide‑range capabilities. The code is available at https://github.com/PlantPotatoOnMoon/HPSv3‑PlusPlus.
Authors:Imane Meddour, Andréa Macario Barros, Cédric Gouy-Pailler
Abstract:
In this work, we propose StereoGeo, an end‑to‑end network‑based approach for stereo camera calibration. Our method estimates the focal lengths and gravity directions of the left and right cameras, as well as the relative extrinsic transformation relating them. Existing methods often rely on calibration patterns in structured environments or address only a single camera configuration, being limited to either intrinsic or extrinsic estimation, and depending on a multi‑view setups. StereoGeo extends the GeoCalib algorithm, integrating deep neural network feature extraction with a differentiable optimizer. Extensive experiments on real‑world benchmarks demonstrate that StereoGeo achieves competitive performance for intrinsic calibration and provides accurate stereo extrinsic estimation, outperforming existing methods that are limited to monocular settings. The dataset used in this work is partially publicly available at https://github.com/meddourimane/StereoGeo‑dataset.
Authors:Sihan Zhuang, Xinyuan Chen, Tianfan Xue, Yaohui Wang
Abstract:
Recent advances in diffusion‑based video generation have significantly improved visual quality and short‑term temporal coherence. However, existing methods still struggle to produce videos with physically consistent and causally plausible dynamics, especially in scenarios involving long‑horizon interactions. This limitation arises from the fact that video diffusion models primarily learn physical consistency implicitly, while vision‑language models can directly model physical laws. Based on this idea, in this work, we propose CausalMotion, a training‑free framework that injects explicit physical reasoning into video generation through structured intermediate representations. Our key idea is to decouple reasoning from generation by leveraging a vision‑language model to decompose a text prompt into a sequence of causally consistent keyframes and object‑centric motion trajectories. These representations are then aligned and integrated as soft constraints to guide a pretrained video diffusion model during inference. This design enables explicit modeling of object dynamics and causal transitions without requiring additional training or supervision. Extensive experiments show that our method consistently improves physical plausibility and temporal coherence, particularly in dynamics‑intensive scenarios, while maintaining high perceptual video quality.
Authors:Victor Barberteguy, Ahmet Iscen, Mathilde Caron, Alireza Fathi, Gül Varol, Cordelia Schmid
Abstract:
Recent advances in 3D feedforward reconstruction neural networks have achieved remarkable success in dense reconstruction from images without any camera parameters. Yet, equipping these models with robust semantic understanding remains an open problem. Here we introduce an approach that performs 3D reconstruction and 3D panoptic segmentation in a unified framework. We build on existing 3D reconstruction models and augment them with a set‑based mask decoder. The approach is jointly trained with a geometric and semantic loss, which are shown to be mutually beneficial. More precisely, the features are initialized from the geometric information and then finetuned to capture jointly geometry and semantics. We demonstrate the generality of our approach by successfully applying our framework both to online and all‑to‑all attention reconstruction backbones. Our method achieves state‑of‑the‑art performance in 3D panoptic segmentation across ScanNet, ScanNet200, and ScanNet++ datasets. Ablation studies show that such joint training of a unified model equips 3D feedforward reconstruction neural networks with panoptic segmentation and yields mutually beneficial improvements.
Authors:Hyejin Oh, Woo-Shik Kim, Sangyoon Lee, YungKyung Park, Je-Won Kang
Abstract:
Multispectral (MS) imaging extends beyond conventional RGB imaging by capturing more spectral bands, thereby improving illuminant spectrum estimation (ISE). However, existing methods often fail to fully exploit spectral information, resulting in suboptimal performance under diverse lighting conditions and across different sensor domains. Hence, we propose a deep learning framework with a spatio‑spectral feature extraction block, which incorporates spectral attention mechanisms to enhance spectral correlation and preserve illuminant‑relevant spatial features. Through the inclusion of an illuminant prior (IP), our approach prioritizes specific channels that provide more meaningful information in an MS image. We also propose a spectral‑domain transform across different MS sensor spaces. The results demonstrate that illuminant spectra learned in high‑dimensional sensor spaces can be effectively transformed to various lower‑dimensional camera sensor spaces without any additional training. To facilitate evaluation, we introduce a real‑world MS dataset containing high‑dimensional ground‑truth illumination spectra captured under diverse lighting conditions. Through extensive experiments, we demonstrate that our method achieves superior accuracy compared to existing models, thus providing a practical solution for real‑world ISE. The code and dataset are available at https://github.com/hyejin5/Spectrum‑Aware‑Illumination‑Estimation‑Using‑Multispectral‑Image.
Authors:Zheyuan Zhan, Hongchen Li, Can Wang, Yinfei Ma, Mingzhen Huang, Ruoshi Bai, Jiawei Chen, Siwei Lyu, Defang Chen
Abstract:
Inversion‑based image editing offers flexible and training‑free control but still struggles with inversion accuracy and the trade‑off between editing fidelity and background preservation. While recent methods improve inversion formulations or attention interactions, the role of textual conditioning in shaping diffusion dynamics and editing behavior remains underexplored. We show both empirically and theoretically that the precision of textual conditioning influences inversion stability by modulating the geometry of the diffusion velocity field, while also affecting the consistency of cross‑branch attention during editing. These effects directly impact background preservation and semantic fidelity. Building on this analysis, we propose SimEdit, a conditioning‑aware framework with two complementary components: (a) conditioning refinement, which constructs conditioning signals with improved semantic precision and structural alignment to facilitate stable inversion and consistent attention manipulation, and (b) token‑wise cross‑branch attention control, which separates edit‑relevant and structure‑preserving components and modulates them asymmetrically during attention manipulation. Extensive experiments on PIE‑Bench demonstrate that SimEdit consistently improves both inversion reconstruction quality and editing performance over previous attention‑manipulation approaches. Our code is available at https://github.com/zju‑pi/SimEdit.
Authors:Yanbin Hao, Pengyu Liu, Xing Wei, Xun Yang, Dan Guo, Meng Wang
Abstract:
Micro‑actions are short‑duration, low‑amplitude subtle body movements at the whole‑body level that can reveal latent intentions, involuntary reactions, and fine‑grained affective changes. Our previous MA‑52 benchmark has provided an important foundation for micro‑action recognition, but it remains limited in scale, scene diversity, task coverage, and evaluation protocols. To advance micro‑action analysis toward more realistic and comprehensive settings, we introduce MMA‑82, a large‑scale multi‑domain extension of MA‑52. MMA‑82 expands the label space from 52 to 82 fine‑grained micro‑action categories and covers four distinct domains, including laboratory interviews, street interviews, psychiatric patient interviews, and emotion‑rich television videos, resulting in 77,856 annotated instances from 454 subjects. Built upon MMA‑82, we establish two core tasks: Micro‑Action Recognition and Multi‑label Micro‑Action Detection. For recognition, we further define in‑domain and cross‑domain protocols, including few‑shot and zero‑shot settings, to evaluate model robustness, transferability, and generalization. Extensive experiments show that current methods still struggle with realistic micro‑action understanding, especially under domain shift, long‑tailed category distributions, and complex temporal localization. Beyond benchmarking, we investigate the relationship between micro‑actions and emotion, showing that micro‑actions are strongly associated with emotional states and provide complementary cues to facial micro‑expressions for improved emotion recognition. These results demonstrate that MMA‑82 serves as a comprehensive and challenging benchmark for realistic micro‑action analysis and a valuable resource for human‑centered AI. MMA‑82 is available at https://lpynow.github.io/MMA‑82‑AIM/.
Authors:Shiao Wang, Xiao Wang, Chao Wang, Yitao Li, Menghao Liu, Bo Jiang, Yaowei Wang, Yonghong Tian, Jin Tang
Abstract:
Conventional RGB cameras have been widely used in multi‑object tracking due to their ability to capture rich appearance and semantic information. However, their performance is often degraded under complex real‑world challenges, such as motion blur, low illumination, and overexposure. Bio‑inspired event cameras offer high temporal resolution and high dynamic range, providing complementary cues under extreme scenarios. Nevertheless, RGB‑event multi‑object tracking remains underexplored due to the lack of large‑scale and well‑annotated datasets. To address this issue, we propose FEMOT, a large‑scale RGB‑event multi‑object tracking dataset that covers diverse real‑world scenarios and 14 challenging attributes. With both RGB and event data as well as high‑quality annotations, FEMOT provides a reliable platform for systematically evaluating RGB‑event multi‑object tracking methods. Based on FEMOT, we retrain and evaluate over ten strong trackers, thereby establishing a comprehensive benchmark for future research. Furthermore, we propose FEMOTR, a multimodal tracking framework that decouples RGB and event features and fuses them in the frequency domain, thereby effectively exploiting their complementary characteristics for robust object localization and identity association. Extensive experiments on FEMOT and DSEC‑MOT datasets demonstrate the effectiveness of the proposed method. The source code and benchmark dataset have been released on https://github.com/Event‑AHU/FEMOT.
Authors:Shiyao Wang, Xijuan Zeng, Hui Wang, Shiwan Zhao, Feng Deng, Chen Zhang, Yong Qin
Abstract:
We present FoleyGenEx, a unified video‑to‑audio (VTA) framework integrating multi‑modal control, frame‑level temporal alignment, and fine‑grained semantics, enabling synchronized, versatile audio synthesis for diverse tasks. Existing VTA methods either have multi‑modal control but weak temporal alignment or strong alignment but lack reference audio conditioning and semantic precision. FoleyGenEx fills this gap via three core innovations: a conditional injection mechanism for audio‑controlled VTA and Foley extension, a multi‑modal dynamic masking strategy preserving training synchronization, and an adverb‑based data augmentation algorithm leveraging signal processing and large language models to enhance textual supervision with nuanced semantics. Experiments on AudioCaps, VGGSound, and Greatest Hits demonstrate its competitive controllable VTA performance against existing methods. Demo samples are available at https://foleygenex.github.io/FoleyGenEx.
Authors:Minghan Li, Jeremy Moebel, Mengyu Wang
Abstract:
One‑step image editing is important for making text‑guided editing fast, practical, and easy to deploy, but its underlying mechanism is still not fully understood. We revisit ChordEdit through reproduction, ablation, and simplification. Our analysis shows that a) the chord window δ largely acts as an effective timestep shift from t to t ‑ δ; b) chord transport acts on high‑noise images and mainly performs low‑frequency semantic editing; and c) proximal alignment acts on low‑noise images and complements it by adding high‑frequency target details. In this view, ChordEdit naturally decomposes editing into a coarse low‑frequency transport stage and a fine high‑frequency alignment stage. These findings suggest a path toward prompt‑conditioned dynamic timestep selection for adaptive image editing. All code and results can be found at \hrefhttps://github.com/Harvard‑AI‑and‑Robotics‑Lab/ChordEdit‑Reproductionlink.
Authors:Dinh-Khoi Vo, Nhut-Thanh Le-Hinh, Viet-Tham Huynh, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
Abstract:
Zero‑shot text‑guided diffusion has significantly advanced image editing; however, its practical usability remains constrained by three persistent challenges: prompt brittleness that requires meticulous prompt engineering, spillover edits that unintentionally affect non‑target regions, and failures on small or cluttered objects caused by limited fine‑grained supervision in training data. We propose FocusDiff (Target‑Aware Refocusing for Tuning‑Free Diffusion Editing), a tuning‑free framework for precise and region‑specific image manipulation based on refocusing cross‑attention. Given a target region obtained through automated segmentation or manual selection, FocusDiff applies selective blurring to non‑edit areas to guide attention toward the masked region while accurately transferring the object's identity, structure, and appearance to the edited output. Integrated context‑preserving modules further ensure background fidelity and global coherence, enabling accurate edits from simple text prompts in a single pass. We also extend FocusDiff to 360‑degree indoor panorama editing and demonstrate its effectiveness within virtual reality environments. Extensive experiments on our localized editing benchmark LIMB, comprising 30 multi‑object images and 100 annotated examples including challenging small‑object cases, show that FocusDiff outperforms existing zero‑shot editors in text‑image alignment and background preservation, achieving superior precision, photorealism, and usability. The project page is available at https://vdkhoi20.github.io/FocusDiff.
Authors:Duong-Duy-Khang Bui, Minh-Tan Pham, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
Abstract:
Fashion sketching is a cornerstone of design workflows, allowing rapid visualization of creative concepts prior to physical prototyping. Yet, progress in sketch‑based fashion image synthesis has been hindered by the absence of large‑scale, high‑quality paired resources. To bridge this gap, we present GarmentSketch, a novel dataset comprising 26,249 fashion sketches across 21 garment categories, each paired with detailed textual descriptions. Captions were produced through a multi‑stage pipeline that integrates multiple multimodal large language models (MLLMs) with human‑in‑the‑loop refinement, ensuring both semantic accuracy and descriptive richness. We benchmark GarmentSketch on state‑of‑the‑art generative models, providing baseline performance for sketch‑guided text‑to‑image generation. Our experiments reveal both the promise and the current limitations of existing methods. By offering a comprehensive and richly annotated resource, GarmentSketch establishes a foundation for advancing sketch understanding, fine‑grained fashion image generation, and creative human‑AI collaboration in design. The dataset will be available at: https://khangbdd.github.io/garmentsketch.
Authors:Krispin Wandel, Jingchuan Wang, Hesheng Wang
Abstract:
Vision Transformers (ViTs) have become a dominant architecture for visual representation learning, providing exceptionally strong and broadly reusable backbone features. However, ViTs are commonly operated on relatively small patch‑token grids due to the quadratic cost of global self‑attention, which creates a persistent bottleneck for dense prediction tasks such as semantic segmentation and depth estimation. This has motivated the development of task‑agnostic feature upsamplers. While recent state‑of‑the‑art methods produce visually sharp dense representations, their reliance on shallow image encoders for guided upsampling can introduce feature leakage, fragmentation, and blur. We introduce ViT‑Up, an implicit feature upsampling framework that replaces external image guidance with layer‑wise query construction from intermediate ViT hidden states. This enables feature prediction at arbitrary continuous image coordinates while preserving alignment with the backbone feature space. Experiments demonstrate that ViT‑Up consistently outperforms state‑of‑the‑art image‑guided upsamplers across dense prediction and semantic correspondence. On DINOv3‑S+, ViT‑Up improves over prior methods by up to +2.07 mIoU on Cityscapes and +4.17 PCK@0.10 on SPair‑71k. With the larger DINOv3‑B backbone, these gains increase to +3.36 mIoU and +8.09 PCK@0.10, demonstrating that ViT‑Up scales favorably with backbone capacity.
Authors:Yijun Liang, Hengguang Zhou, Ming Li, Lichen Li, Cho-Jui Hsieh, Tianyi Zhou
Abstract:
Vision‑language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non‑trivial, visual‑centric and grounded questions remains underexplored. Existing visual questioners' performance is bottlenecked by the availability of high‑quality training data or the cost of curating them. We show that a VLM can continuously improve itself as a visual questioner without any external supervision. We propose a self‑evolving framework that uses a VLM itself as both a proposer and a filter to produce harder, more informative, and visual‑centric questions, while maintaining their exploration diversity to avoid training collapse. These questions are then used to train the VLM in both questioner and answerer modes. To evaluate the questioner, we introduce an agentic protocol that assesses questions along perception, reasoning, and diversity dimensions. Experiments across various backbone VLMs show that our method substantially enhances the quality and substantially expands the difficulty boundary of autonomous question generation. Under the same budget, our self‑supervision is more effective than training on the static source data. Moreover, the self‑evolving questioner remains a competitive or even better answerer.
Authors:Isai Daniel Chacón, Zhongqi Miao, Bruno Demuro, Caleb Robinson, Rahul Dodhia, Lasha Otarashvili, Jason Holmberg, Kirk Larsen, Howard Frederick, Nathan J. Pamperin, Pablo Arbeláez, Juan M. Lavista Ferres
Abstract:
Automated aerial wildlife surveys increasingly rely on deep learning, yet standard object detectors require bounding‑box annotations, reported to be up to seven times slower and three times more expensive to produce than point‑level labels. To address this bottleneck, we introduce the Overhead Wildlife Locator (OWL), a weakly supervised density‑estimation framework with three variants: OWL‑C, a fully convolutional model for high‑throughput screening; OWL‑T, a Swin‑augmented hybrid for heterogeneous, cluttered scenes; and OWL‑D, built on a frozen DINOv3 ViT‑H+/16 encoder with a DPT‑style fusion decoder. We benchmark all three against POLO, YOLOv11n, and YOLOv11l across five public aerial datasets, from sparse fixed‑wing savanna surveys to dense UAV paddock imagery, and against the published HerdNet baseline on its native Delplanque split. OWL‑D sets a new state of the art on Delplanque (0.934 AP vs. HerdNet's 0.840) and records the highest AP on four of the five datasets. Performance is regime‑dependent: on the extreme‑density SheepCounter UAV dataset the hybrid OWL‑T leads (0.978 AP) and the convolutional variants attain the lowest counting error, whereas the foundation‑based OWL‑D degrades, indicating which variant suits which survey type. We further validate operational readiness on the Alaska Department of Fish and Game's 2022 Central Arctic Caribou census: under cross‑herd and cross‑temporal transfer, OWL‑C fine‑tuned on the 2017 Porcupine Caribou Herd split attains F1 = 0.965 on a held‑out patch test set, with a signed count error of +3.1% aggregated across the released test patches. We release the OWL code, model weights, and the annotated Porcupine Caribou Herd 2017 (PCH) and Central Arctic Herd 2022 (CAH) patches, the first open patch‑level datasets for large‑scale caribou aerial surveys, at https://github.com/microsoft/MegaDetector‑Overhead.
Authors:Stella Katharina Wermuth, Qazi Arbab Ahmed, Klaus Neumann, Thorsten Jungeblut
Abstract:
Autonomous staff‑free public transport requires reliable in‑vehicle passenger monitoring. However, perception inside moving vehicles is challenged by confined spaces, variable illumination, motion‑induced background variation, occlusion, and limited viewpoints. To mitigate these spatial constraints, ceiling‑mounted fisheye cameras provide full‑scene coverage from a single viewpoint. Yet existing public overhead fisheye datasets are recorded in static environments and do not capture the domain shift introduced by vehicle motion. To fill this gap, we introduce PMOF, Passenger Monitoring using Overhead Fisheye cameras, the first public dataset of top‑view fisheye imagery captured inside a moving vehicle, comprising over 19k manually annotated frames. PMOF provides rotated bounding boxes, tracking identifiers, and action labels, supporting object detection, tracking, and action recognition. We benchmark PMOF using YOLO26m‑obb models fine‑tuned under multiple dataset configurations that combine PMOF with existing overhead fisheye datasets. Cross‑domain fine‑tuning with custom rotation‑aware augmentation achieves 94.8% AP50 on PMOF and 96.5% AP50 on an unseen overhead fisheye dataset from a different domain. Our results highlight the domain gap between static and moving environments and show that incorporating PMOF improves detection performance and advances generalization beyond passenger monitoring to broader fisheye‑based person detection tasks. The dataset and code are available at https://swermuth.github.io/pmof/.
Authors:Nadav Benedek, Tomer Koren, Ohad Fried
Abstract:
AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter‑sized buffers to training memory. We propose Gefen, a memory‑efficient optimizer that automatically shares second‑moment estimates across parameter blocks and quantizes the first moment using a learned codebook, thereby reducing AdamW's memory footprint by ~8x while maintaining the same performance, corresponding to a reduction of 6.5 GiB per billion parameters. The method is motivated by a theoretical result showing that large mixed Hessian entries constrain the ratio of squared gradients toward one, suggesting that Hessian‑aligned parameters are natural candidates for sharing second‑moment statistics. Since computing Hessians is impractical at scale, Gefen infers block structure from the initial squared gradients, requiring no architecture‑specific metadata or hyperparameters beyond AdamW defaults. Gefen learns an exact histogram‑based dynamic‑programming quantization codebook and reuses the same blocks for first‑moment scaling. Across diverse experiments, Gefen achieves the lowest peak optimizer memory among the compared AdamW‑like methods while maintaining AdamW‑level performance. In FSDP and DDP training, the reduced memory footprint enables larger microbatches and improves throughput significantly over AdamW, providing a practical drop‑in replacement with lower memory usage that can increase throughput and enable training larger models or using larger batch sizes. We provide the complete Python implementation, including fused CUDA kernels at https://github.com/ndvbd/Gefen
Authors:Sharath Girish, Tsai-Shien Chen, Zhikang Dong, Mukesh Singhal, Hao Chen, Sergey Tulyakov, Aliaksandr Siarohin
Abstract:
Cinematic video depicts multiple subjects acting or interacting at specific moments, captured with deliberate camera movement, and stitched together by shot transitions. Together, these elements demand a level of fine‑grained control beyond current text‑to‑video models. Existing work addresses each axis in isolation: multi‑subject personalization, temporal control, multi‑shot synthesis, or camera control; no prior framework jointly integrates all four. We present CineOrchestra, a unified video diffusion model that controls subjects, events, cameras, and shot transitions simultaneously. Our key insight is that these heterogeneous cinematic elements share a fundamental structure: each is an entity acting over a specific temporal interval, which can therefore all be expressed through one shared structure of entity‑centric conditioning primitives, augmented with reference images for visual entities. This formulation reduces the architectural challenge to a single positional encoding problem, which we solve with two parameter‑free coordinated rotary embeddings: (a) an interval‑sampled temporal RoPE that yields consistent attention behavior across events of dramatically varying duration, and (b) a 2D entity‑temporal cross‑attention RoPE that disambiguates per‑entity conditions and routes each to its corresponding spatiotemporal region. On two new benchmarks, CineOrchestra outperforms six per‑axis specialists on dense caption following and shot‑transition timing, with consistent gains in a pairwise user study and component ablations. Project page: https://snap‑research.github.io/CineOrchestra
Authors:Phuc Nguyen H
Abstract:
Human pose estimation (HPE) utilizing wireless WiFi signals has emerged as a promising technology owing to its device‑free nature, privacy preservation, and robustness against occlusion and poor lighting. However, existing methods often overlook the physical complex phase information of WiFi signals and fail to generalize across diverse environments due to severe domain shifts. In this paper, we present C‑MambaPose, a physics‑informed complex‑valued Mamba‑GraFormer hybrid framework for robust cross‑environment WiFi‑based 3D HPE. Our framework first sanitizes raw WiFi Channel State Information (CSI) phase errors and constructs a phase‑preserving complex‑valued representation. We then employ a Spatiotemporal Complex Mamba encoder with a dynamic selective receptive field to capture fine‑grained phase dynamics. A cross‑attention joint‑query mapper maps the unstructured sequence tokens to human joints, which are decoded by a Graph Convolutional Network (GCN) to predict anatomically coherent 3D coordinates. Extensive evaluations on the MM‑Fi dataset show that C‑MambaPose achieves competitive or superior performance to state‑of‑the‑art baselines across all settings, setting a new state‑of‑the‑art specifically on the challenging cross‑environment split, requiring only 3.78 M parameters‑an 83.1% reduction compared to GraphPose‑Fi~\citechen2026graph and an 85.7% reduction compared to MetaFi++~\citezhou2023metafi++, while maintaining a comparable size to DT‑Pose~\citechen2025towards (which is only 18% smaller) but achieving significantly superior performance without requiring any pretraining. Our code is publicly available at https://github.com/phucngvinuni/cmampose.git.
Authors:Dian Zheng, Harry Lee, Manyuan Zhang, Kaituo Feng, Zoey Guo, Ray Zhang, Hongsheng Li
Abstract:
Recent image generators have demonstrated impressive photorealism and instruction‑following capabilities in single‑image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text‑image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation. Even the latest open‑source Unified Multimodal Models (UMMs) exhibit limited performance in this regard. In this paper, we introduce InterleaveThinker, the first multi‑agent pipeline designed to endow any existing image generator with interleaved generation capabilities. Specifically, we employ a planner agent to organize the image‑text input sequence, instructing the image generator on the required execution at each step. Subsequently, we introduce a critic agent to evaluate the generator's outputs, identify samples that deviate from the planned instructions, and refine the instructions for regeneration. To implement this pipeline, we construct the Interleave‑Planner‑SFT‑80k and Interleave‑Critic‑SFT‑112k to perform a format cold‑start. Then we develop Interleave‑Critic‑RL‑13k to reinforce the step‑wise instruction correction capability within a generation trajectory using GRPO. Since a single interleaved generation trajectory may involve over 25 generator calls, optimizing the entire trajectory is computationally impractical. Therefore, we propose accuracy reward and step‑wise reward, allowing single‑step RL to effectively guide the entire generation trajectory. The results show that InterleaveThinker improves performance across various image generators. On interleaved generation benchmarks, it achieves performance comparable to Nano Banana and GPT‑5. Surprisingly, it also significantly enhances the base model on reasoning‑based benchmarks; for example, on 4‑step FLUX.2‑klein, we observe substantial gains on WISE and RISE.
Authors:Zhao-Heng Yin, Guanya Shi, Pieter Abbeel, C. Karen Liu
Abstract:
Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal degrees of freedom and contact‑rich interactions. While prior work has largely focused on rigid objects, articulated tool use remains underexplored because of its physical complexity and the difficulty of learning functional grasping and manipulation policies. We present Mana (Manipulation Animator), a general sim‑to‑real framework that reinterprets dexterous manipulation as an animation problem. Inspired by computer animation, Mana employs a coarse‑to‑fine pipeline that transforms procedurally‑generated grasp keyframes into manipulation trajectories through motion planning and reinforcement learning. The data generation process is largely automatic, requiring only a few mouse clicks to specify functional affordances (<1 minute per tool). Across four articulated tools spanning different scales and joint types, Mana achieves zero‑shot sim‑to‑real transfer for both grasping and in‑hand manipulation, demonstrating a scalable approach to dexterous articulated tool use.
Authors:Junke Wang, Qihang Zhang, Shuai Yang, Yiming Luo, Yujun Shen, Zuxuan Wu, Yu-Gang Jiang, Yinghao Xu
Abstract:
This work presents RepWAM, a representation‑centric world action model (WAM) built on representation visual‑action tokenizers. Existing WAMs typically inherit reconstruction‑oriented video tokenizers from pretrained video generation models. Although these tokenizers preserve visual fidelity, pixel reconstruction alone provides limited guidance for learning instruction‑following dynamics that connect future prediction with robot control. To address this, we explore a semantic visual‑action latent space for representation‑centric world action modeling. Specifically, we train a representation visual‑action tokenizer that maps visual inputs into aligned visual and latent action tokens. We then pretrain our WAM to jointly model future visual states and the latent actions that connect them under language instructions, followed by adaptation to real robot trajectories for closed‑loop manipulation. Experiments on real‑world manipulation tasks and simulation benchmarks show that RepWAM delivers strong performance across diverse manipulation settings, while ablations highlight the value of semantic visual‑action tokenization over reconstruction‑oriented alternatives. These results establish representation visual‑action tokenization as a promising foundation for world action models and a step toward generalist robot policies. Code and weights will be available at https://github.com/wdrink/RepWAM.
Authors:Hao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang, Yi Hua, Ben Mildenhall, Christoph Lassner, Narendra Ahuja, Gengshan Yang
Abstract:
Image‑to‑3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but stop at the visible surface, while image‑to‑3D models generate complete shapes that are often misaligned with the input. We introduce World Tracing, a generative pixel‑aligned geometry representation that predicts 3D points aligned with observed pixels while completing geometry beyond the visible surface. For each input pixel, World Tracing predicts an ordered stack of camera‑space 3D points, where the first layer represents the visible surface and subsequent layers represent front‑to‑back intersections with occluded surfaces. We instantiate this representation with a world‑tracing diffusion transformer, WT‑DiT, which treats multiple geometry layers as separate denoising tokens coupled through factorized and global attention. WT‑DiT is trained with pixel‑space flow matching and a mixed noise schedule that balances visible‑surface reconstruction with occluded‑geometry generation. World Tracing achieves strong performance on visible‑surface reconstruction and complete geometry generation across object, scene, and dynamic benchmarks, outperforming both depth predictors and image‑to‑3D generators. It also preserves 2D‑to‑3D correspondence, enabling text‑driven 3D scene editing, geometry‑conditioned novel‑view video synthesis, and training‑free integration with textured‑mesh generators.
Authors:Antoine Guédon, Shu Nakamura, Nicolas Dufour, Jiahui Lei, Ko Nishino, Angjoo Kanazawa
Abstract:
Geometry is invariant to viewpoint, which makes any collection of images a redundant encoding of a single 3D state. Existing feed‑forward reconstruction models fail to exploit this: per‑view methods emit overlapping, unaligned pointmaps that grow linearly with input count, while global‑latent methods commit to a fixed, low‑resolution output. We introduce Surflo, which compresses a variable number of unposed RGB views into K latent tokens‑one global state‑and decodes oriented 3D surface points by independently transporting them from noise onto the surface via flow matching. This frees the output from any fixed grid or token budget: the same latent yields from a few thousand to a million points in a single forward pass. To suppress the local inconsistencies inherent to independent per‑point decoding, an inference‑time guidance term correlates nearby points by injecting a photometric gradient during ODE integration. Surflo matches or surpasses feed‑forward baselines on surface metrics, runs an order of magnitude faster than optimization‑based methods that require hundreds of views, and is the only feed‑forward approach to combine a global latent with arbitrary‑resolution decoding.
Authors:Vinícius Orrú, Bruno H. Foggiatto, Gabriel E. Lima, David Menotti, Rayson Laroca
Abstract:
Vehicle color recognition is an important cue for vehicle identification in surveillance systems, especially when license plates are illegible due to low resolution, occlusion, motion blur, or poor illumination. However, real‑world vehicle color distributions are highly imbalanced, making overall accuracy insufficient to assess performance on rare but operationally relevant colors. This paper presents a comprehensive study of vehicle color recognition under severe class imbalance using UFPR‑VeSV, a challenging real‑world surveillance dataset. We investigate synthetic minority‑class augmentation through two off‑the‑shelf generative strategies: text‑conditioned image generation with RunDiffusion/JuggernautXL and image‑conditioned color editing with Gemini 2.0 Flash. The curated synthetic data are combined with modern visual representations, loss reweighting, learning‑rate scheduling, color‑safe augmentation, foreground‑aware preprocessing, and ensemble fusion. The bestperforming approach achieves 94.6% micro accuracy and 79.7% macro accuracy, improving macro accuracy by 8.2 percentage points over recent literature. A manual error analysis further shows that many remaining failures are visually ambiguous even for human annotators, highlighting the practical limits of color‑based vehicle identification in unconstrained surveillance imagery. The generated images and source code are publicly available at https://github.com/viniciusorru/vcr‑synthetic
Authors:Dachun Kai, Jiayao Lu, Yueyi Zhang, Xiaoyan Sun
Abstract:
Event‑based vision has drawn increasing attention owing to its distinctive properties, including ultra‑high temporal resolution and extreme dynamic range. Recent works have introduced it to video super‑resolution (VSR) to enhance flow estimation and temporal alignment. In contrast, this paper shifts the focus of event signals from motion refinement to texture enhancement in VSR. We propose EvTexture++, the first event‑driven framework dedicated to texture enhancement in VSR. It leverages high‑frequency spatiotemporal details from events to improve texture recovery. EvTexture++ incorporates a customized texture enhancement branch, along with an iterative texture enhancement module that progressively exploits high‑temporal‑resolution event information for texture restoration. This enables gradual refinement of texture regions across iterations, yielding more accurate and detailed high‑resolution outputs. Besides intra‑frame texture recovery, large motions could degrade inter‑frame temporal consistency, particularly in texture regions, leading to texture flickering. To mitigate this, we further exploit the continuous‑time motion cues of events to enhance temporal consistency, introducing a temporal texture alignment module that estimates event‑guided texture‑aware flow for precise inter‑frame texture alignment. Moreover, EvTexture++ is designed as a plug‑and‑play tool to flexibly boost the performance of existing VSR models. Experiments on five datasets demonstrate that EvTexture++ achieves state‑of‑the‑art performance. When integrated into recent VSR models, it yields significant improvements, with gains of up to 1.55 dB in PSNR on the texture‑rich Vid4 dataset. Code: https://github.com/DachunKai/EvTexture.
Authors:Mingkun Lei, Tong Zhao, Liangyu Yuan, Chi Zhang
Abstract:
Step‑level caching accelerates diffusion models by exploiting temporal redundancy across denoising steps. Existing methods make per‑step cache decisions using threshold‑based heuristics, without directly optimizing for final output quality. As a result, their inference latency varies across inputs and is difficult to control at deployment. In this work, we propose BudCache, which inverts this formulation: rather than letting per‑step error thresholds dictate the runtime cost, we fix the compute budget in advance and search for the cache policy that best preserves the final output. To tackle the combinatorial complexity of step selection, we combine Simulated Annealing with deterministic Hill Climbing. This offline search identifies high‑quality cache policies within minutes and introduces no online search or thresholding overhead during inference. When the compute budget is very tight, we further introduce cache‑aware schedule alignment, which adapts the time discretization to the selected cache policy to reduce cache‑induced trajectory mismatch. Experiments on FLUX.1‑dev and Wan2.1 show that BudCache achieves better generation quality than heuristic caching baselines under the same inference budgets. Code is available at https://github.com/Westlake‑AGI‑Lab/BudCache
Authors:Daichi Azuma, Taiki Miyanishi, Koya Sakamoto, Shuhei Kurita, Yaonan Zhu, Petr Khrapchenkov, Motoaki Kawanabe, Yusuke Iwasawa, Yutaka Matsuo
Abstract:
Goal‑conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed‑loop control. We propose Navigation World Action Model (NavWAM), a diffusion‑transformer policy that turns navigation world‑model prediction into executable action by representing future observations, goal‑progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed‑loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real‑robot adaptation, and evaluate it on image‑goal navigation against planning‑based world models and a representative direct navigation policy. Across offline benchmarks and closed‑loop real‑robot deployment, NavWAM improves over planning‑based world‑model baselines in our evaluations while using the default policy mode without CEM‑style action search. Project page: https://dachii‑azm.github.io/navwam/
Authors:Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, Zhimin Zhang, Yawen Luo, Guoxin Zhang, Yu-Shen Liu, Pengfei Wan
Abstract:
Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametric representations that fail to handle multi‑shot generation or synthesize cross‑paired data, which suffer from data scarcity, resulting in poor performance in complicated camera motion cloning. To address these issues, we introduce a general camera motion representation that encodes cameras as grid motion videos. This camera grid represents the camera parameters visually and supports the integration of diverse trajectories for multi‑shot video generation. Building upon this, we propose OmniDirector, a unified framework trained on a million‑scale camera grid‑video pairs that coordinates characters, actions, and cameras to provide director‑level control for multimodal diffusion transformers. Furthermore, we design a novel hierarchical prompt expansion agent that harmoniously integrates different control signals by systematically describing camera motion and visual content through understanding signal relationships. Extensive experiments demonstrate the superior performance and outstanding controllability of our framework. Project page: https://ymlinfeng.github.io/OmniDirector.github.io/
Authors:Hoang-Nguyen Cao, Le-Hoang Bui, Dinh-Khoi Vo, Minh-Triet Tran, Trung-Nghia Le
Abstract:
Cultural garments pose a unique challenge for visual retrieval systems, as their identity often depends on subtle structural and symbolic details that are poorly captured by standard AI models. We introduce VietFashion, a new benchmark for sketch‑text composed image retrieval centered on the Ao Dai, a traditional Vietnamese garment. VietFashion enables designers and researchers to retrieve culturally meaningful outfits using a combination of hand‑drawn sketches, which convey garment structure, and textual descriptions, which encode cultural semantics. The dataset is initialized with 650 sketches and expanded using generative models to produce over 21,000 photorealistic images with aligned captions. Textual prompts that describe detailed outfit attributes, which are extracted from fashion magazines to ensure authenticity and diversity. To better reflect the inherent ambiguity of design intent, VietFashion adopts a multi‑target retrieval setting, where a single query may correspond to multiple valid results. We establish standardized evaluation protocols and benchmark state‑of‑the‑art composed image retrieval methods. Experimental results reveal significant performance gaps in modeling fine‑grained cultural semantics and multi‑modal composition, positioning VietFashion as a challenging benchmark for fine‑grained fashion retrieval. The dataset is publicly available at: https://hng0303.github.io/VietFashion.
Authors:Xinnan Zhu, Ruijie Xu, Jiayu Ying, Daoguo Dong, Jiachen Xu, Yuan Xie, Xin Tan
Abstract:
Existing 3D scene editing methods typically rely on per‑scene optimization over explicit 3D representations or cascaded edit‑and‑reconstruct pipelines, resulting in high test‑time cost, limited 3D awareness, and structural inconsistencies. To couple appearance synthesis and geometry prediction during editing, we build on a unified RGB‑geometry reconstruction‑generation latent space and adapt it to feed‑forward 3D scene editing. The resulting framework, JointEdit3D, performs asymmetric latent inpainting by observing only a single edited RGB reference latent and generating the remaining RGB views and edited geometry latent under source‑scene anchoring. JointEdit3D introduces a dedicated SceneAnchor Branch to inject source‑scene structure without forcing direct copying, and adopts edit/background‑aware losses to balance edited‑region fidelity with unedited‑content preservation. To address the lack of paired resources for standardized 3D scene editing evaluation, we introduce SceneEdit3D‑15K, a dataset with 15K paired editing samples and renderer‑provided 3D annotations, together with SceneEdit3D‑Bench, a curated 100‑sample benchmark. Experiments show that JointEdit3D improves edited‑region quality and 3D structural completeness over prior baselines while maintaining competitive background preservation.
Authors:Wei Li, Zhen Huang, Xinmei Tian
Abstract:
Contrastively trained vision‑language models like CLIP, have made remarkable progress in learning joint image‑text representations, but still face challenges in compositional understanding. They often exhibit a "bag‑of‑words" behavior‑‑struggling to capture the object relations, attribute‑object bindings, and word order dependencies. This limitation arises not only from the reliance on global, single‑vector representations for optimization, but also from the insufficient exploitation and modeling of the rich compositional information inherently present in paired image text data. In this work, we propose MACCO (MAsked Compositional Concept MOdeling), a framework that masks compositional concepts in one modality and reconstructs them conditioned on the full contextual information from the other, enabling the model to capture and align cross‑modal compositional structures more effectively. To facilitate this process, we introduce two auxiliary objectives that jointly align and regularize masked features both inter‑modally and intra‑modally. Extensive experiments on five compositional benchmarks, along with in‑depth analyses, demonstrate that our approach not only significantly enhances compositionality in VLMs but also improves their ability to capture syntactic structure and linguistic information. Additionally, the improved compositionality also benefits text‑to‑image generation and multimodal large language model. Code is available at https://github.com/hiker‑lw/MACCO.
Authors:Anugrah Aidin Yotolembah, Novanto Yudistira, Gembong Edhi Setyawan
Abstract:
This paper presents Custom ZeroCLIP, a retrieval‑augmented vision‑language framework for zero‑shot captioning of Indonesian traditional garments. The dataset contains 3,800 expert‑annotated images from all 38 Indonesian provinces. Using a province‑level inductive zero‑shot protocol, the model is trained on 24 seen provinces, validated on 6 seen provinces, and evaluated on 8 unseen provinces. The framework combines a frozen CLIP ViT‑B/32 image encoder, a CLIP text encoder, a BERT text encoder, and an LSTM caption decoder. During inference, unseen‑province labels and captions are unavailable, and retrieval uses only captions from training provinces. No unseen‑province image, label, or caption is used during training, validation, or retrieval‑bank construction. Custom ZeroCLIP achieves a CLIPScore of 0.8536, BLEU‑4 of 0.3342, and METEOR of 0.4859, outperforming existing baselines. Ablation results show that retrieval improves cultural vocabulary recovery with a 19.3% METEOR gain, while human evaluation confirms stronger cultural accuracy and fluency. The results demonstrate the effectiveness of retrieval‑augmented domain adaptation for culturally grounded caption generation in low‑resource heritage settings. The dataset is publicly available at https://github.com/AnugrahAidinYotolembah/Traditional‑Indonesian‑Clothing‑Captioning‑Dataset.
Authors:Sathira Silva, Abrham Kahsay Gebreselasie, Muhammad Umer Sheikh, Kartik Kuckreja, Daniel Harari, Muhammad Haris Khan
Abstract:
Learning grounded word meaning from natural experience requires resolving two ambiguities in infant‑view recordings: when the named referent appears and where it is in a cluttered frame. In SAYCam‑style data, caregiver speech is sparse and weakly synchronized with egocentric video, so single‑frame contrastive pairing yields noisy positives in which the intended object is absent or entangled with distractors. We propose BabyMind, an object‑first bias for child‑view contrastive learning under sparse, noisy supervision. BabyMind extracts candidate object embeddings using an offline mask‑based region interface, links candidates across a short utterance‑centered window into lightweight object files via tracking, and aligns utterances to bags of object files with a prototype‑space multiple‑instance contrastive objective. Track‑coherence and global‑object agreement regularizers stabilize learning and transfer object‑file structure into the global frame embedding used at evaluation. On SAYCam‑S, BabyMind improves Labeled‑S 15 forced‑choice accuracy by +2.6 points over CVCL and yields consistent gains on in‑vocabulary out‑of‑distribution benchmarks. Code is available at https://github.com/sathiiii/BabyMind.
Authors:Ching-Yu Tsai, Chia-Min Lin, Chih-Hsiang Yang, Yung-Che Wang, Jen-Shiun Chiang
Abstract:
Crack detection plays an important role in infrastructure inspection and Structural Health Monitoring (SHM). However, cracks typically appear as thin, low‑contrast structures and are easily affected by background noise, posing challenges for existing object detection models. This study proposes an improved YOLO‑based architecture with integrated attention mechanisms, termed YOLO‑AMC (YOLO with Attention Mechanisms for Crack Detection), to enhance automated crack detection performance. Based on YOLOv11, the original C2PSA module is removed, and multiple attention mechanisms, including Global Attention Mechanism (GAM), Residual Convolutional Block Attention Module (Res‑CBAM), and Shuffle Attention (SA), are introduced into the multi‑scale feature fusion layers of the Neck to strengthen cross‑scale feature integration. Experimental results demonstrate that YOLO‑AMC consistently outperforms baseline models YOLOv11n and YOLOv8n across multiple evaluation metrics. Among the evaluated attention modules, GAM achieves the best detection performance, obtaining mAP@0.5 = 0.9917 and mAP@0.5:0.95 = 0.9506 on the test dataset, which are higher than those of YOLOv11 (0.9833 / 0.9112) and YOLOv8 (0.9707 / 0.8921). Furthermore, while maintaining a computational complexity of 7.6 GFLOPs, the proposed model achieves 110.95 FPS on an NVIDIA RTX 4090 platform and approximately 5 FPS on a Raspberry Pi 5 edge device, demonstrating a favorable trade‑off between accuracy and deployment efficiency. The implementation code for this study is available on GitHub at https://github.com/CY‑Tsai24/YOLO‑AMC.
Authors:Inseok Kong, Geunyoung Jung, Jiyoung Jung
Abstract:
3D point cloud models suffer significant performance degradation under distribution shifts caused by sensor noise, occlusions, and environmental changes. Test‑time adaptation (TTA) has emerged as a practical paradigm for mitigating this issue during inference. Recently, leveraging multi‑view augmentation has shown promise in improving 3D TTA performance. However, existing multi‑view approaches are often constrained by sequential optimization that treats each view independently. This sequential optimization leads to substantial inference latency due to repetitive optimization steps, making real‑time adaptation impractical. To address this, we propose Masked Multi‑View Test‑Time Adaptation (MAMVI), which replaces sequential optimization with a unified single‑step adaptation. Specifically, MAMVI utilizes a hybrid masking strategy that combines fixed ratios for stability with Beta‑distributed sampling for diversity. By aggregating losses across multiple views, MAMVI performs adaptation through a single backward pass based on multi‑view consensus. Additionally, a confidence‑based adaptive learning rate is used to dynamically adjust the adaptation intensity for each sample. Extensive experiments on ModelNet‑40C, ShapeNet‑C, and ScanObjectNN‑C demonstrate that MAMVI achieves state‑of‑the‑art accuracy on ShapeNet‑C and ScanObjectNN‑C. Moreover, it remains competitive on ModelNet‑40C while delivering 4.9‑8.9 times faster inference, making it highly suitable for real‑time applications. Our code is available at https://github.com/Inseok‑kong/MAMVI
Authors:Allison Andreyev, Landon Eum, Nestor Tiglao, Romel Gomez
Abstract:
For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural‑language prompts in real time. Although Vision‑Language Models (VLMs) have enabled zero‑shot generalization in robot task and motion planning (TAMP), current state‑of‑the‑art approaches often remain computationally "heavyweight" or require extensive training on thousands of demonstrations. We present GRASP (Grounded Reasoning and Symbolic Planning), a framework designed as a step toward open‑vocabulary tabletop manipulation. Our approach leverages a pretrained VLM to translate natural‑language queries into neuro‑symbolic goal states, grounded in the physical world via a bounding‑box detection pipeline. Unlike methods that rely on fixed color lists or hard‑coded coordinates, GRASP enables robots to interpret abstract spatial concepts such as "top shelf" and execute tasks without additional fine‑tuning. We achieve 73.3% overall success across 90 real‑robot trials at three difficulty levels, requiring no task‑specific training.
Authors:Xu-Jing Ye, Yuan-Gen Wang, Ruping Wang
Abstract:
The Abstraction and Reasoning Corpus (ARC) is viewed as a critical avenue to Artificial General Intelligence (AGI), as it enables models to learn abstract transformation rules from few‑shot examples and then generalize to new tasks. However, prevalent ARC methodology is either pure language or vision‑only (i.e., VARC). The former depends heavily on LLMs, consuming billions of parameters. The latter often struggles to capture high‑level semantics, leading to overfitting on pixel‑level patterns. To bridge this gap, we propose L‑VARC, a novel framework that enhances visual reasoning via a language‑guided Learning Using Privileged Information (LUPI) branch. Specifically, we design a Semantic Compression Module by feeding a unified, task‑agnostic prompt into DeepSeek‑V3. In this way, the raw LARC (a crowd‑sourced language description dataset) can be substantially refined and structured, fitting with the context length constraint of standard text encoders (e.g., CLIP). Moreover, we design a Cross‑Attention Projector to align visual features with semantic embeddings, aiming to guide the training of the ARC model. Notably, the LUPI branch is taken in the training process and will be discarded during inference, thereby yielding a lightweight model with a mere 18 million parameters. Extensive experiments demonstrate that our L‑VARC effectively leverages linguistic priors to boost visual reasoning and outperforms state‑of‑the‑art. Ablation studies further confirm the contribution of the two new designs towards the L‑VARC framework. The code is available at https://github.com/GZHU‑DVL/L‑VARC.
Authors:Jiangtao Kong, Peijun Zhao, Chun-Fu Chen, Youngwook Do, Shaohan Hu, Tianyi Zhou, Huajie Shao
Abstract:
Incremental Learning (IL) for Open‑ended Image‑to‑Text Generation (OpenITG) enables models to continuously generate accurate, contextually relevant text for new images while preserving previously acquired knowledge. Unlike prior studies, this paper addresses a more practical scenario in which the predominant category of visual data shifts over time as environments evolve. In this context, we introduce a new notion of continual alignment, which incrementally adapts the alignment module within pre‑trained VLMs to preserve high‑quality cross‑modal representations. Based on this idea, we propose Efficient Continual Alignment (ECA), a novel exemplar‑free IL approach for OpenITG. The key challenge is enabling the model to acquire new, task‑specific features while minimizing interference with the established alignment without accessing raw data from previous tasks. To address this, ECA employs three core mechanisms: a Mixture of Query (MoQ) module that adapts task‑specific query tokens, a Fisher Dynamic Expansion (FeDEx) that dynamically expands model structure based on a Fisher Information Matrix (FIM)‑based metric, and an embedding dictionary with Dictionary Replay (DR) to retain past knowledge. To evaluate ECA's performance, we construct four new IL OpenITG benchmarks that better reflect real‑world scenarios. Experimental results demonstrate that ECA significantly mitigates catastrophic forgetting and improves IL performance compared to baseline methods. Code and benchmarks are available at https://github.com/Snowball0823/ECA.
Authors:Binay Kumar Singh, Niels Da Vitoria Lobo
Abstract:
Object detection in autonomous driving requires precise localization and an inherent understanding of the relational context between co‑occurring objects. In extremely complex heterogeneous environments rare classes, small‑scale objects, and frequently appearing objects are difficult for standard object detection frameworks to handle. In this paper, we propose a novel framework called Context‑Centric Feature Fusion (CCFF), which utilizes two attention‑based modules, Local Context Fusion Module (LCFM) uses the RoI‑to‑RoI self‑attention mechanism to resolve spatial interactions, mainly considering small and partially obscured objects, while Global Context Attention Module (GCAM) converts the co‑occurrence of objects priors by pooling top‑K RoI features into a global context attention token, avoiding the computational overhead of pixel‑level global pooling. This fusion of local and object‑centric global features yields contextualized embeddings that enhance classification results and co‑occurring objects detection. Our method is evaluated on two datasets, Cityscapes and BDD100K which demonstrate significant improvement on relational consistency, achieving a Category‑level Consistency Strategy (CCS) of 0.973 and 0.969, respectively. Furthermore, our approach produces substantial gains in small object detection (AP_S: 14.1%) and successfully recovers rare classes such as "Train" that are typically lost in large distributions. Our efficiency report shows that the framework processes images in real time with a 0.2 FPS overhead. The code is available at https://github.com/BinayKSingh/CCFF.
Authors:Alireza Heidari, Amirhossein Alimohammadi, Wallace Michel Pinto Lira, Adi Bar-Lev, Ali Mahdavi-Amiri
Abstract:
Transferring hairstyles between images is an important but challenging task in computer graphics, computer vision, and visual effects. It enables users to explore new looks without physically altering their hair, with applications in virtual try‑on systems, augmented reality, and entertainment. Most prior works operate best under small pose gaps, and they fall short under large viewpoint and scale differences, where missing hair content must be synthesized rather than transferred. We propose HairPort, a 3D‑aware hairstyle transfer framework that attempts to solve these issues by explicitly separating hair removal from transfer and enforcing geometric consistency before synthesis. We introduce a Bald Converter, which produces realistic bald versions of faces through LoRA‑based in‑context adaptation of FLUX.1 Kontext. To train our Bald Converter, we introduce a new dataset, Baldy, containing 6,000 paired bald and original images across diverse identities and conditions. We also use a 3D‑Aware Transfer Pipeline that reconstructs and re‑renders the reference hairstyle from the target viewpoint before compositing it onto the source image. Being 3D aware, our method supports large pose and scale discrepancies between the source and target. Finally, a conditional flow‑matching generator synthesizes the transferred result from the bald source and geometry‑aligned reference guidance. Together, our method enables accurate, pose‑consistent, and identity‑preserving hairstyle transfer, outperforming existing methods both qualitatively and quantitatively.
Authors:Zeyue Tian, Lei Ke, Zhaoyang Liu, Ruibin Yuan, Liumeng Xue, Yujiu Yang, Weijia Chen, Xu Tan, Qifeng Chen, Wei Xue, Yike Guo
Abstract:
Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework, 2) large‑scale, high‑quality training data, and 3) the prohibitive inference cost of multi‑step diffusion sampling. As such, we propose AudioX‑Turbo, a unified and efficient framework for anything‑to‑audio generation that integrates varied multimodal conditions (i.e., text, video, and audio signals) in this work. AudioX‑Turbo follows a teacher‑student paradigm. The teacher AudioX‑Base is built on a Multimodal Diffusion Transformer with a Multimodal Adaptive Fusion module that aligns diverse multimodal inputs for high‑fidelity synthesis, and is then distilled into the few‑step student AudioX‑Turbo via Distribution Matching Distillation adapted to flow matching, complemented by a diffusion‑based discriminator for high‑quality few‑step generation. To support the training of AudioX‑Turbo, we construct a large‑scale, high‑quality dataset, IF‑caps‑Pro, comprising approximately 9.2M samples curated through a two‑stage data collection and annotation pipeline. We benchmark AudioX‑Turbo across a wide range of tasks, finding that our model achieves superior performance, especially on text‑to‑audio and text‑to‑music generation, while operating at only 4 sampling steps and requiring approximately 25x fewer function evaluations (NFE) than multi‑step baselines. These results demonstrate that our method is capable of audio generation under flexible multimodal control, showing efficient and powerful instruction‑following capabilities. The code and datasets will be available at https://zeyuet.github.io/AudioX‑Turbo/.
Authors:Cheng-Yu Yang, Shao-Yuan Lo, Yu-Lun Liu
Abstract:
Vision‑language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computation and KV‑cache memory. Existing visual‑token reduction methods largely follow a rank‑and‑remove paradigm: they score visual tokens, keep a compact subset, and permanently discard the rest. We show that this irreversible action is fragile because visual‑token importance changes across decoder depth; tokens ranked low at one stage may become relevant in later layers, especially for grounding‑sensitive queries. We propose Reroute, a training‑free plug‑in that replaces removal with recoverable routing. At each routing stage, selected vision tokens pass through decoder blocks, while deferred tokens bypass the stage and re‑enter the candidate pool at the next routing decision. Reroute reuses existing attention‑score ranking rules and stage‑wise schedules, preserving the theoretical TFLOPs and KV‑cache budget class of the pruning method it augments. Across FastV, PDrop, and Nüwa variants on LLaVA‑1.5 and Qwen backbones, reroute improves grounding under aggressive token reduction while maintaining general VQA performance. These results suggest that VLM token reduction should not be viewed only as irreversible pruning, but also as recoverable routing. The code can be found here: https://github.com/elmma/mllm‑reroute/
Authors:Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, Burhan Yaman
Abstract:
Vision‑language‑action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision‑language‑action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per‑pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open‑loop and closed‑loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open‑loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50\,m average) and 3‑second collision rate (0.18%). On closed‑loop Bench2Drive, VLGA attains the state‑of‑the‑art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.
Authors:Zhen Zhao, Gang Zhang, Xiaolin Hu, Liang Tang
Abstract:
Object detection and instance segmentation tasks are closely related. Existing top‑down instance segmentation methods usually follow a detect‑then‑segment paradigm, where an initial detector is used to recognize and localize objects with bounding boxes, followed by the segmentation of an instance mask within each bounding box. In such methods, the detection accuracy directly influences the subsequent segmentation performance. However, previous research has seldom explored the impact of the instance segmentation task on object detection. In this paper, we present a turbo‑inference strategy for the top‑down methods that leverages the complementary information between detection and segmentation tasks iteratively. Specifically we design two modules: turbo‑detection head and turbo‑segmentation head, which facilitate communication between the tasks. The two modules form a closed loop that interlaces the detection and segmentation results without retraining the model. Comprehensive experiments on the COCO, iFLYTEK, and Cityscapes datasets demonstrate that our method substantially enhances both detection and segmentation accuracies with a certain increase in computational cost. The proposed method represents a tradeoff between prediction accuracy and inference speed. Codes are available at https://github.com/zhaozhen2333/Turbo‑Learning.git.
Authors:Gege Gao, Bernhard Schölkopf, Andreas Geiger
Abstract:
Memory is not merely the storage of data; it is the scaffolding of reality. When biological memory fades, the world does not simply turn black; it regresses into an unrecognizable chaos. Echoes of the Prior is an interactive installation that attempts to visualize this subjective phenomenology of forgetting. By inducing controlled synaptic decay within a Feed‑Forward 3D Reconstruction model, we create an artistic analogy for the erosion of the brain's predictive priors. We position the Neural Network not as a tool for engineering, but as a cognitive proxy ‑ a silicon brain whose structural degeneration evokes the disorienting, poetic, and terrifying experience of losing one's grip on the world. Ultimately, we offer this framework as a catalyst, inviting the wider community to explore the uncharted potential of neuromorphic aesthetics in visualizing the fragility of intelligence. Interactive demo see https://decart‑4d.github.io/.
Authors:Yuchen Xian, Yunqiu Xu, Yang He, Yi Yang
Abstract:
Multimodal image fusion aims to integrate complementary information from different modalities into a fused image that preserves rich local details while maintaining globally consistent appearance. Existing approaches build shared representations on 2D feature grids, which excel at modeling local structures but offer limited leverage over image‑level global appearance factors. To balance these objectives, we introduce a compact 1D token interface based on a frozen pretrained image tokenizer for modeling non‑local appearance/base factors. Rather than using the tokenizer as a reconstruction backbone, our design uses the 1D token space as a global carrier while retaining the 2D spatial pathway for local structure restoration. Specifically, we introduce Selective Token Editing (STE), which sparsely updates/replaces a small set of critical tokens, providing a lightweight mechanism to steer global appearance coherence while keeping the fusion backbone unchanged and avoiding extra losses. Experiments on four commonly used benchmarks show that our method achieves the best overall performance, with consistent, multi‑metric improvements in both global coherence and local fidelity. Project page: https://zju‑xyc.github.io/1D‑Fusion‑Project‑Page/
Authors:Sukmin Seo, Geewook Kim
Abstract:
Temporal grounding‑‑returning the interval [t_s, t_e] for a natural‑language query over a video‑‑is the language interface to long‑form video, yet has been studied on short videos; the dynamics of hour‑scale natural‑language grounding remain underexplored. We take the position that at hour‑scale, the binding constraint is search, not recognition: Video‑LLMs are bottlenecked not by localizing a nearby event, but‑‑given a natural‑language query‑‑by searching for the relevant region of a long video. To test this, we release ExtremeWhenBench, the first open hour‑scale grounding benchmark (2,273 queries over 194 videos, mean 75.7 min, max 9 hr) with an open‑form query distribution. Every open Video‑LLM collapses while a frame‑level retrieval baseline outperforms them; a failure taxonomy attributes 85% of failures to search; and a retrieve‑then‑ground hybrid recovers 6.7x over the monolithic Video‑LLM‑‑mirroring retrieve‑then‑read in open‑domain QA.
Authors:Alexander Martin, Dengjia Zhang, Joel Brogan, Francis Ferraro, Jeremy Gwinnup, Reno Kriz, Teng Long, Kenton Murray, Andrew Yates, Xiang Xiang
Abstract:
This overview paper presents the results of the shared task for the second workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR). In this shared task participants submitted systems focused on either (i) video retrieval or (ii) grounded generation of articles given retrieved videos. Teams could submit to either task. For the retrieval task, we had 2 participating teams that submitted a total of 17 systems ‑‑ all of which beat a baseline derived from the winner of last year's shared task. On the generation side, we had 4 teams submit 16 systems. All teams had at least one generated report that was labeled the best by a human annotator.
Authors:Benjamin Eckhardt, Dmytro Fishman, Stuart Fawke, Andrew Curtis, Bo Fussing, Constantin Pape
Abstract:
Counting living cells is an important step in many biological research workflows. Our collaborators at the Wellcome Sanger Institute study vital genes in humans via large scale saturation genome editing screening, which requires repeatedly counting cells a great number of times. Computer Vision based automation is crucial for high throughput and resource efficiency. In this work, we develop a regression‑based deep learning computer vision algorithm to detect and count cells in phase‑contrast microscopy images. To reduce annotation effort, which in practice often becomes a bottleneck, we focus on counting cells only using sparse point annotations, which are fast and easy to acquire. By comparison to state‑of‑the‑art 0‑shot methods, we show that regression‑based counting is a promising alternative in low data regimes. Through developing methods to automatically count living cells in microscopy images, we contribute to valuable research on the human genome. The code is available at https://github.com/beijn/cellnet.
Authors:Weirong Chen, Keisuke Tateno, Hidenobu Matsuki, Michael Niemeyer, Daniel Cremers, Federico Tombari
Abstract:
We address 4D reconstruction from partial point cloud sequences, where depth‑sensor observations are incomplete, unordered, and lack explicit temporal correspondences. This geometry‑only setting is challenging due to missing observations and ambiguous dynamics. While recent progress has largely relied on image‑based methods, existing point‑based approaches typically focus on single objects, assume relatively complete inputs, or require explicit correspondences. To address these limitations, we propose DynaTok, a point‑based framework for correspondence‑free 4D reconstruction from partial point cloud sequences without images. DynaTok encodes frames into compact latent tokens, aggregates incomplete observations over time with a Transformer‑based spatiotemporal encoder, and decouples geometry and motion through residual tokens in a unified model. A flow‑matching decoder then reconstructs complete, temporally consistent 4D point‑cloud sequences conditioned on the latent tokens. Experiments on object‑ and scene‑level benchmarks demonstrate improved reconstruction quality and temporal coherence from partial point cloud observations. Project page: https://wrchen530.github.io/dynatok/.
Authors:Jiawei Niu, Jian Chen, Di Zhang, Junbo Lu, Zhangcheng Liao, Xuhao Liu, Honglin Zhong, Mireia Crispin-Ortuzar, Chen Li, Zeyu Gao, Yi Cai
Abstract:
Existing computational pathology methods predominantly operate within whole‑slide image (WSI)‑level multiple instance learning (MIL) paradigms, while patient‑level modeling remains underexplored. In routine pathological practice, however, pathologists derive diagnostic and prognostic conclusions by integrating evidence across multiple WSIs rather than relying on any single slide. This discrepancy creates a fundamental misalignment when patient‑level supervision is directly imposed on conventional MIL frameworks, often leading to unstable optimization and degraded predictive reliability. To address this issue, we propose Anchor‑Guided Evidence MIL (AGE‑MIL), a weakly supervised framework for patient‑level prediction. AGE‑MIL constructs a patient‑level anchor from slide representations to capture global pathological context and guide the retrieval and integration of diagnostically relevant local patches, enabling robust patient‑level modeling. Patient‑level risk is further modeled as an evidence accumulation process, promoting stable optimization under weak supervision. AGE‑MIL is evaluated on six clinically relevant patient‑level prediction tasks from two independent cohorts. Experimental results show that the proposed framework consistently outperforms eight state‑of‑the‑art MIL methods. Code is available at https://github.com/wodeniua/AGE‑MIL.
Authors:Pankhuri Vanjani, Zhuoyue Li, Jakub Suliga, Moritz Reuss, Gianluca Geraci, Xinkai Jiang, Rudolf Lioutikov
Abstract:
Vision‑language‑action (VLA) models inherit a shared synchronous clock from vision‑language pretraining, processing every input at one rate. This is misaligned with physical interaction, where a high‑frequency modality changes at hundreds of hertz, vision evolves more slowly, and language stays constant across an episode. A synchronous VLA oversamples slow modalities, undersamples fast ones, and caps action generation at the lowest effective frequency. We hypothesize that decoupling temporal processing per modality, letting each update and retain information at its own sensor rate, yields stronger representations and more robust control. We present DAM‑VLA, which maintains per‑modality latent buffers refreshed at sensor rates and read continuously by the action head, integrating new high‑frequency modalities through gated cross‑attention that leaves the pretrained backbone intact. Across seven contact‑rich real‑world manipulation tasks, DAM‑VLA more than doubles the average success rate of the strongest synchronous baseline (95.2% vs.\ 40.95%) while sustaining smooth, reactive 100\,Hz control. Project website: \hrefhttps://intuitive‑robots.github.io/DAM‑VLA/intuitive‑robots.github.io/DAM‑VLA/
Authors:Tahar Chettaoui, Guray Ozgur, Eduarda Caldeira, Naser Damer, Fadi Boutros
Abstract:
Recent advances in Vision Transformers (ViTs) for face recognition (FR) have moved beyond the standard CLS‑token paradigm. In this paradigm, a special classification token (CLS) is prepended to the patch embeddings and used as a representation of the input for downstream tasks. An alternative approach, Concatenated Patch Embeddings (CPE), instead leverages all patch tokens by concatenating them into a single vector, which is then projected into a compact face representation. CPE has been shown to improve recognition performance in comparison to CLS‑based ones, but our qualitative analysis of attention maps showed the presence of artifacts that limit their interpretability. To address this issue, we incorporate register tokens, learnable tokens concatenated to the initial patch embeddings, and processed jointly through the ViT encoder blocks. This mechanism has been shown to produce more structured and interpretable attention maps compared to baseline ViT. We empirically demonstrate that these artifacts consistently appear across various ViT backbones, including small and large models, and that introducing register tokens effectively mitigates them. Adding four or eight registers significantly enhances interpretability, with eight registers providing the highest verification accuracies and smoothest attention structures. Our resulting model, ViT‑8R, corresponds to a CPE‑based ViT‑B architecture augmented with eight register tokens achieves state‑of‑the‑art performance among ViT‑based FR models on large‑scale IJB‑B and IJB‑C benchmarks. Also, ViT‑8R produces substantially clearer attention maps compared with the baseline model, which offer deeper insight into the model's attention behavior (https://github.com/TaharChettaoui/ViT‑FR‑Registers)
Authors:Min Yang, Mi Zhou, Limin Wang
Abstract:
Video understanding is a crucial part of computer vision, with numerous application scenarios. With the increasing popularity of mobile devices, an increasing number of efforts are trying to deploy video understanding models on them. However, existing video understanding models are difficult to deploy due to their large size and prohibitive power consumption. Spiking Neural Networks (SNNs) have shown bioplausibility and low power advantages over Artificial Neural Networks (ANNs), especially on neuromorphic chips which are regarded as essential components of future mobile devices. However, excessively long conversion time‑steps and severe performance degradation problems limit their application. To solve the problems above, we explore the application of SNNs on temporal action detection (TAD), which is an important task in video understanding, and propose the first SNN‑based end‑to‑end TAD architecture coined as SpikeTAD. While maintaining extremely low power consumption, SpikeTAD achieves an average mAP of 67.2% in THUMOS14 and 37.42% in ActivityNet‑1.3, demonstrating the feasibility of a low‑power TAD model. Our code is available at https://github.com/MCG‑NJU/SpikeTAD.
Authors:Yiqun Ning, Ao Shen, Chenhang He, Lei Zhang
Abstract:
While diffusion‑based virtual try‑on has achieved impressive visual realism, most methods treat the task as 2D inpainting, prioritizing texture preservation over physical plausibility. Consequently, they often produce plausible‑looking images that fail to reflect authentic garment fit across diverse body shapes. We present FitVTON, a Fit‑aware virtual try‑on model on different bodies in the wild. FitVTON encodes garment‑body size through structured text prompts, and learn from simulated try‑on triplets from parameterized garment model. To improve the fitting effects over garment silhouettes, we introduce two auxiliary head to predict the masks for both the garment and the exposed body. We further introduce a texture rectification stage to improve realistic appearance from simulated data. To evaluate the fitting fidelity, we curate a real‑world dataset, FittingEffect3K, combining VLM‑based scoring protocol. Both subjective and quantitive experiments show that FitVTON demonstrate authentic fitting fidelity, with significant sizing accuracy and shape preservation over state‑of‑the‑art methods while maintaining competitive image quality. Project Page: https://zenoning.github.io/FitVTON/.
Authors:LeKai Yu, Hao Liu, Kun Wang, Zhiran Li, Ruping Cao, Fan Liu, Yupeng Hu
Abstract:
In this report, we present our third‑place solution for the DataMFM Challenge Track 1: Document Parsing. This track requires models to recover structured Markdown documents from document page images while preserving textual content and document structure. To address the complementary requirements of accurate content recovery and faithful structure reconstruction, we propose ParseFixer, an agentic framework for backbone parsing and selective correction. ParseFixer consists of two key modules: Full‑Page Backbone Parsing (FBP) and Agentic Selective Correction (ASC). FBP produces stable initial Markdown outputs with MinerU2.5 Pro, while ASC detects high‑value parsing failures and repairs them through a verify‑and‑rollback correction process. By placing selective multimodal correction after open‑source backbone parsing, ParseFixer improves the recovery of key document elements without rewriting reliable backbone predictions. On the test set, our final system achieves an overall score of 61.78 and ranks third in Track 1, demonstrating its effectiveness for accurate document parsing. Our code will be released at: https://github.com/iLearn‑Lab/CVPRW26‑ParseFixer.
Authors:Zsolt Robotka, Ádám Rák, Jalal Al-Afandi, András Horváth, György Cserey
Abstract:
Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non‑signing communities. While large weakly‑aligned datasets have enabled pre‑training at scale and gloss‑free methods have reduced reliance on expert annotation, high‑quality parallel sign video‑text pairs for fine‑tuning remain scarce, limiting generalisation on long‑tail vocabulary and unseen constructions. We propose a corpus augmentation approach that requires no additional human annotation, external sign‑language video corpora, or generative video models, relying only on the existing gloss‑annotated training corpus and an LLM for sentence generation: per‑gloss clips are extracted from training videos via CTC forced‑alignment, novel gloss‑sentence pairs are generated by a corpus‑anchored LLM, and synthetic sequences are assembled through random sentence sampling and clip assignment. The resulting synthetic RGB video‑text pairs are architecture‑agnostic at the downstream training stage and can be consumed directly by RGB‑based SLT models, or converted into pose or feature representations by pipelines that derive such inputs from video. Sincan et al. re‑evaluated five recent gloss‑free methods under strictly identical conditions; the largest verified gain over the GFSLT‑VLP baseline was only 0.98 BLEU‑4. Our augmentation, applied within the same framework, achieves +2.92 BLEU‑4 without any change to architecture or training protocol. We further identify that synthetic data harms vision‑language pretraining despite improving its objectives, and that optimising clip transitions for visual smoothness is counter‑productive under L2‑based criteria; we propose that abrupt boundaries may act as a form of implicit regularisation. Code is available at https://github.com/robizso/slt‑datagen.
Authors:Yuto Furutani, Takashi Otonari, Kaede Shiohara, Toshihiko Yamasaki
Abstract:
Feed‑forward 3D Gaussian Splatting (3DGS) removes the need for time‑consuming per‑scene optimization required by traditional 3DGS. However, existing feed‑forward approaches struggle with real‑world photo collections that include diverse lighting conditions and transient objects. In this paper, we present Wild3R, a feed‑forward approach for unconstrained sparse photo collections. The main bottleneck is the lack of training data that provides multiple viewpoints, a variety of illuminations, and transient variations necessary for learning robust scene representations. To address this, we introduce the WildCity dataset, which comprises 200 scenes, 170 lighting conditions, and transient objects, resulting in 337,500 images in total. By leveraging the dataset, our model learns appearance consistency across viewpoints conditioned on reference views, while removing transient content. Extensive experiments demonstrate that our method outperforms existing feed‑forward approaches and achieves results competitive with prior per‑scene optimization‑based methods.
Authors:Nicole Damblon, Olga Vysotska, Federico Tombari, Marc Pollefeys, Daniel Barath
Abstract:
Visual localization in complex indoor environments remains a critical challenge for robotics and AR applications. Sequential localization, where pose estimates are refined over time, is important for autonomous agents. However, traditional methods often require storing extensive image databases or point clouds, leading to significant overhead. This paper introduces a novel, lightweight approach to sequential visual localization using 3D scene graphs. Our method represents the environment with a compact scene graph, where nodes represent objects (with coarse meshes) and edges encode spatial relationships. For each image in the localization phase, we extract per‑patch semantic features, predicting object identities. Localization is performed within a particle filter framework. Each particle, representing a camera pose, projects the coarse object meshes from the scene graph into the image, assigning object identities to patches based on visibility. The similarity of the per‑patch features, in the input image, and object features from the scene graph determines the weight of a particle. Subsequent images are incorporated sequentially, refining the pose estimate. By leveraging a compact scene graph and efficient semantic matching, our method significantly reduces storage while maintaining performance on real‑world datasets. The code will be available at https://github.com/DmblnNicole/sg2loc.
Authors:Mingzhe Lyu, Jinqiang Cui, Hong Zhang
Abstract:
Low‑light novel view synthesis is challenging because dark multi‑view images contain noise, weak structural detail, and compressed dynamic range. Recent 3D Gaussian Splatting (3DGS) methods address these challenges by generating pseudo ground‑truth (pseudo‑GT) images as supervision targets when paired normal‑light references are unavailable. Existing pseudo‑GT methods apply a uniform linear gain to all pixels, which clips bright regions while providing insufficient enhancement in dark regions, limiting reconstruction quality. We observe that nonlinear tone mappings, long established in 2D low‑light enhancement, have not been explored for pseudo‑GT generation in 3D reconstruction. Accordingly, we propose a scene‑adaptive nonlinear tone‑curve framework that replaces linear pseudo‑GT with nonlinear alternatives. The framework introduces percentile‑based normalisation for scene‑agnostic curve application, a scene‑adaptive offset for automatic black‑level adjustment, and two complementary curves: Adaptive SoftExp (ASE), a bounded exponential curve, and Adaptive Poly3 (AP3), a data‑driven cubic polynomial. The module changes only the pseudo‑GT computation and leaves the 3DGS backbone unchanged. Experiments on three benchmarks covering 21 scenes show that both curves consistently outperform the linear baseline with PSNR improvements up to +4.34 dB on LOM and +3.25 dB on RealX3D. Both curves achieve similar performance despite their different mathematical forms, suggesting the improvement is curve‑agnostic. Code is available at https://github.com/lvmingzhe/adaptiveToneCurve
Authors:Hang Xu, Xiaoxiao Ma, Guohui Zhang, Yu Hu, Siming Fu, Jie Huang, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao
Abstract:
Multi‑turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps. While existing research leverages video priors for consistency, their reliance on bidirectional attention is fundamentally misaligned with the causal, sequential nature of interactive editing. In this paper, we propose AnchorEdit, the first autoregressive (AR) diffusion‑based framework designed specifically for high‑resolution, long‑term multi‑turn editing. AnchorEdit bridges the gap between video priors and causal inference through a three‑stage training curriculum: identity‑preserving sing‑turn pretraining, causal AR forcing fine‑tuning with a novel self‑rollout strategy to mitigate exposure bias, and consistency distillation for efficient 4‑step generation. During inference, we introduce a memory mechanism to anchor the initial subject identity and ensure stable extrapolation across extended editing trajectories. To evaluate performance, we provide a new high‑resolution multi‑turn editing benchmark designed to stress‑test long‑horizon stability. Extensive experiments demonstrate that AnchorEdit achieves state‑of‑the‑art results, maintaining exceptional subject fidelity and instruction following even over 10+ interaction rounds.
Authors:Mengzhuo Chen, Yan Shu, Chi Liu, Hongming Piao, Xidong Wang, Derek Li, Bryan Dai
Abstract:
We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a common reasoning interface. We introduce UniReason‑Med, a single‑checkpoint framework that processes either a 2D image or a slice‑serialized 3D volume at inference time, generating interleaved textual reasoning and localized visual evidence through shared box syntax, region‑token injection, and a common grounded reasoning policy. To train this interface, we construct UniMed‑CoT, a 220K instruction‑tuning dataset with interleaved textual reasoning and grounded visual evidence, including 170K 2D and 50K 3D samples. Through supervised fine‑tuning followed by outcome‑level reinforcement learning, UniReason‑Med learns to generate grounded reasoning traces without IoU/Dice‑based localization rewards during RL. Data‑mixture and component ablations show that joint 2D+3D grounded supervision substantially improves 3D reasoning over 3D‑only training, while grounding and region‑token injection consistently benefit both 2D and 3D tasks. These results suggest that a shared grounded reasoning interface can transfer reasoning structure from 2D images to slice‑serialized volumetric medical understanding. The code and data are publicly available at https://github.com/IQuestLab/unireason‑med.
Authors:Evgeny Gorelik, Kenny Dean Karrow, Fikret Sivrikaya, Sahin Albayrak, Christian Baumann
Abstract:
We introduce a multi‑view in‑cabin monitoring dataset for public transportation with synchronized RGB and depth images from four inward‑facing cameras and a rotating LiDAR covering the vehicle interior of a digitalized and partly automated German city bus. The dataset contains 9.136 synchronized samples with annotations and is accompanied by a calibration and pseudo‑labeling pipeline that generates 3D human pose estimates and oriented 3D bounding boxes for occupants. We further provide a nuScenes‑format conversion and benchmark representative multi‑view 3D detection models (e.g., Lift‑Splat‑Shoot and BEVFusion), supporting comparative evaluation and small‑scale training of multi‑view in‑cabin perception models. The dataset and tools are available at https://github.com/EvgenyGorelik/multiview_incabin_dataset.
Authors:Tajamul Ashraf, Hyewon Jeong, Fida Mohammad Thoker, Bernard Ghanem
Abstract:
To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration. Existing benchmarks largely evaluate isolated perception or single‑turn question answering, and therefore provide limited visibility into failures of planning, tool recruitment, and rollout reliability. We introduce MedCTA, a benchmark for evaluating medical tool agents on clinician‑validated, step‑implicit tasks grounded in realistic multimodal clinical inputs, including radiology images, pathology slides, and reports. MedCTA comprises 107 real‑world clinical tasks with clinician‑verified executable trajectories over 5 deployed tools, and supports process‑aware evaluation of tool selection, argument validity, execution stability, trajectory fidelity, and outcome quality. We benchmark 18 open‑ and closed‑source multimodal models and find that even frontier systems remain brittle in multi‑step clinical tool use: autonomous rollouts are dominated by protocol failures, premature stopping, and incorrect tool recruitment, while gold‑standard tool routing yields large but still incomplete gains. These results show that strong backbone perception does not translate into reliable agentic behavior in clinical settings. MedCTA provides a rigorous testbed for auditing, diagnosing, and advancing trustworthy medical AI agents. The dataset and evaluation suite are available at https://ivul‑kaust.github.io/MedCTA/
Authors:Marius Bayizere
Abstract:
Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This paper presents DroneShield‑AI, a unified open framework integrating six processing layers: RF signal classification, acoustic motor‑signature detection, YOLOv8‑based visual detection, evidence‑weighted sensor fusion, a Behavioral Intent Classification Engine (BICE), and a Graph Neural Network Swarm Intelligence Module (GNN‑SIM). BICE introduces the first systematic six‑class threat taxonomy for drone flight patterns, enabling predictive operator alerts with a 30‑second advance‑warning horizon. GNN‑SIM is the first open framework for adversarial multi‑drone formation analysis using Graph Attention Networks. Evaluated on three publicly available real‑world datasets, the fused pipeline achieves 96.1% detection accuracy, 3.2% false alarm rate, AUC‑ROC: 0.981, and 142ms end‑to‑end latency on commodity CPU‑class hardware at approximately 500‑780 USD total system cost. All code, model weights, and simulation datasets are publicly released at submission.
Authors:Chaofan Ma, Zhenjie Mao, Yuhuan Yang, Fanqin Zeng, Yue Shi, Yingjie Zhou, Xiaofeng Cao, Jiangchao Yao
Abstract:
Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods rely on single‑turn inference, forcing models to resolve geometric ambiguity through semantic priors rather than verifiable evidence. We argue that spatial reasoning should be revisitable: conclusions formed under limited evidence should remain open to revision when complementary viewpoints become available. Building on this insight, we propose Reason, then Re‑reason (ReRe), a training‑free, inference‑time framework with two phases: in the Reason Phase, an MLLM forms a spatial hypothesis from the original video; in the Re‑reason Phase, it verifies or revises the hypothesis by observing a synthesized novel‑view video. To enable effective cross‑view revisiting, we design a Geometry‑to‑Video pipeline that renders strategically complementary novel views from predicted 3D geometry. These views feature an elevated, oblique perspective with scene‑spanning coverage, while preserving the MLLM's native video interface without architectural modifications. Extensive evaluations on VSI‑Bench and STI‑Bench demonstrate that ReRe substantially boosts open‑source MLLMs to rival proprietary state‑of‑the‑art performance. Project page: https://zhenjiemao.github.io/ReRe/
Authors:Cheng Chen, Jingyu Zhou, Yifan Zhao, Jia Li
Abstract:
Understanding multi‑label images remains a challenging task in computer vision. With the rapid progress of vision‑language multimodal learning, vision‑language models (VLMs) enable zero‑shot recognition without labeled data. However, due to their intrinsic design, these models often prioritize the most iconic object and omit other contextual positives. This intrinsic bias conflicts with the nature of multi‑label learning, thereby limiting their applicability. In this work, we propose an unsupervised framework that adapts VLMs from iconic recognition toward inclusive understanding, enabling label‑free multi‑label image recognition. Our approach consists of two key stages, ``cutting'' and ``sewing'': In the cutting stage, we present the multi‑sampling response estimator to prevent the model from concentrating only on one single object. In the second sewing stage, the multi‑object blend adaptation is introduced to adjust the labels to better conform to the multi‑label distribution while preserving the intrinsic characteristics of the original model within only one epoch. Extensive experiments show that our framework significantly outperforms existing unsupervised approaches on four public datasets, even surpassing several representative weakly supervised baselines. These results demonstrate the potential of adapting pre‑trained VLMs for more comprehensive visual understanding without manual annotations. Our code is publicly available at https://github.com/iCVTEAM/TailorCLIP.
Authors:Zequn Yang, Yake Wei, Haotian Ni, Zhihao Xu, Di Hu
Abstract:
Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the first systematic, information‑theoretic analysis highlighting why learning these dynamic, sample‑specific interactions is critical for effective multimodal learning. Our analysis further reveals deficits in conventional paradigms at learning these distinct interaction types: modality ensemble approaches struggle to capture synergy, while joint learning paradigms often under‑utilize redundant information. This highlights the need for an approach that can adaptively learn from different interaction types on a per‑sample basis. To this end, we propose Decomposition‑based Multimodal Interaction Learning (DMIL), a novel paradigm that explicitly models and learns from sample‑specific interactions. First, we design a variational decomposition architecture to isolate the constituent interaction components. Second, we employ a new learning strategy that leverages these explicit interaction components in a fine‑tuning process to achieve comprehensive interaction learning. Extensive experiments across diverse tasks and architectures demonstrate that DMIL consistently achieves superior performance by adapting to holistic sample‑specific interactions. Our framework is flexible and broadly applicable, establishing an interaction‑centric paradigm for multimodal learning. The code is available at https://github.com/GeWu‑Lab/DMIL.
Authors:David Hall, Joshua Knights, Mark Cox, Peyman Moghadam
Abstract:
Natural environments present a complex challenge to robotics perception systems. Current models, particularly vision foundation models, are largely trained on structured, urban environments leading to weaknesses in their perception for field robotics tasks. We showcase the limitations of current models using our recently released WildCross benchmark, a new cross‑modal benchmark for place recognition and metric depth estimation in large‑scale natural environments. WildCross comprises over 476K sequential RGB frames with semi‑dense depth and surface normal annotations, each aligned with accurate 6DoF pose and synchronized dense lidar submaps. In this work, we provide an expanded analysis of the benchmark results from the recent WildCross benchmark, with particular emphasis on expanded metric depth estimation experiments. Access to the code repository and dataset for this work can be found at https://csiro‑robotics.github.io/WildCross.
Authors:Shengkai Sun, Zhiyong Cheng, Zefan Zhang, Jianfeng Dong, Zhihui Li, Meng Wang
Abstract:
Recently, masked skeleton reconstruction models have emerged as strong action representation learners, driving significant progress in self‑supervised skeleton‑based action recognition. However, existing state‑of‑the‑art methods must predict an exceedingly large number of spatiotemporal patches, significantly prolonging training time. Besides, by treating all spatiotemporal regions equally during reconstruction, these models are distracted from learning the critical motion patterns that underlie action semantics. To address these challenges, we propose Adaptive Masked Reconstruction (AMR), a faster and stronger pre‑training framework. We first decouple the decoder from the encoder, enabling flexible prediction of larger spatiotemporal patches and dramatically reducing reconstruction complexity. Given that larger patches contain more complex information, which is challenging to predict and consequently degrades performance, we accordingly introduce an adaptive guidance module. This module identifies regions of high motion informativeness, guiding the model to focus on the most discriminative parts of each patch and alleviating reconstruction difficulty. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU‑MMD datasets demonstrate that AMR not only accelerates pre‑training substantially but also improves downstream recognition accuracy, surpassing current state‑of‑the‑art approaches.
Authors:Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Abstract:
Diffusion models have consistently driven progress in text‑to‑image generation. However, it is challenging to attribute recent progress to specific modeling and data choices: state‑of‑the‑art open‑weight models provide limited ablations, and do not disclose their training data and full training details. The research community needs fully open (weights, data, and code) models as a foundation for further research; yet existing fully open models still fall significantly short of leading models in performance. In this project, we conduct a systematic investigation of the modeling and data design choices in text‑to‑image diffusion training and inference with 300+ controlled experiments totaling 700K+ TPU v6e hours. Our experiments highlight several empirical findings (e.g., equal weighting is a strong default for mixing curated datasets) and simple design decisions (e.g., larger text encoder adapters improve performance with minimal added parameters) for training strong models. Guided by these insights, we train i1, a 3B‑parameter text‑to‑image diffusion model using only publicly available datasets. i1 is competitive with leading models on five representative benchmarks (GenEval, DPG, PRISM, CVTG‑2K, and LongText), and outperforms the best existing fully open model by 29.5 absolute percentage points on average. We provide the i1 checkpoints, training and inference code, and the data processing pipeline. Together, our findings and the i1 recipe establish a practical foundation for future open research in text‑to‑image diffusion models. Our code is available at https://github.com/zlab‑princeton/i1.
Authors:Jia Li, Qian Chen, Wei Wang, Xinyu Li, Zhenzhen Hu, Dongsheng Shao, Richang Hong, Meng Wang
Abstract:
Personality assessment aims to infer stable personality traits from dynamic behaviors across language, voice, and facial cues. Since different personality dimensions are revealed through distinct behavioral perspectives, modeling trait‑specific evidence is challenging. However, most existing approaches adopt a uniform multimodal fusion strategy across all dimensions, assuming identical modality contributions. This overlooks trait‑specific modality preferences and introduces cross‑modal interference. To address this issue, we propose a novel personality assessment framework called Traits Run Deeper, which consists of three components. Specifically, the Multimodal Foundation Representation (MFR) module constructs personality‑oriented multimodal inputs and leverages psychology‑informed semantic templates as anchors, enabling foundation models to capture trait‑relevant information. Building upon MFR, the Trait‑Specific Modality Fusion (TSMF) module acts as an asymmetric fusion mechanism, allowing each dimension to selectively exploit different modality pathways from modality‑specific modeling to complementary fusion. Thus, TSMF captures heterogeneous modality preferences while reducing cross‑modal contamination. Furthermore, the Distribution‑Calibrated Personality Regression (DCPR) module mitigates label imbalance and central tendency bias through target distribution calibration, improving robustness and stability. Experimental results on the AVI Challenge 2026 validation set demonstrate the effectiveness of the proposed framework, reducing mean squared error (MSE) by approximately 25% compared with the baseline. Consistent improvements are observed on the official test set, where our method achieves the best performance and ranks first in the Personality Assessment Track. The source code will be made available at https://github.com/MSA‑LMC/AVI2026.
Authors:Yechan Kang, Yongjin Kweon, Mingyeong Seo, Sohee Park, Yeonguk Jeon, Jongkil Park, Hyun Jae Jang, Jaewook Kim, YeonJoo Jeong, Suyoun Lee, Seongsik Park
Abstract:
Training deep spiking neural networks (SNNs) remains challenging due to sharp loss landscapes and temporal inconsistency caused by surrogate gradients. To address these challenges, we propose a unified framework: adaptive and asymmetric surrogate gradients A2SG. The adaptive gradients adjust an effective window for spatio‑temporal adaptation, reducing spatial gradient variation and maintaining directional consistency of gradients over time. The asymmetric gradients reflect neuronal dynamics by assigning larger gradients to neurons with higher membrane potentials, and we prove that they yield lower variation than symmetric surrogates. Our analysis further establishes a direct connection between local gradient variation and the curvature of the loss landscape, providing a principled explanation for how A2SG promotes convergence to flatter minima and improves generalization. We conduct extensive experiments on diverse models, including CNN‑based and Transformer‑based SNNs, across various tasks such as image classification using both static and neuromorphic datasets, as well as segmentation. The results demonstrate that A2SG consistently improves accuracy and energy efficiency, establishing it as a general and reliable solution for training deep SNNs. Our code is available at https://github.com/KIST‑NCL/A2SG.git.
Authors:Suhang Li, Osamu Yoshie, Yuya Ieiri
Abstract:
Vision‑language reinforcement learning has recently shown strong target‑present localization for camouflaged object detection (COD). Yet localization is only one side of the decision: when the agent faces an ordinary image with no camouflaged target, will it still claim that a camouflaged object exists? Standard COD training and evaluation data are positive‑only, so agents optimized under this setting can acquire an over‑detect bias, a task‑specific form of object hallucination that standard COD evaluation leaves unmeasured. To quantify this target‑absent behavior, we construct Counterfactual COD (CF‑COD), a paired benchmark that removes the camouflaged target from each held‑out COD evaluation image while preserving a plausible background. CF‑COD evaluates whether a model detects the target on the original image and abstains on the target‑absent counterfactual, summarized by Pair Accuracy (PA). We further introduce CFCamo, a paired counterfactual framework for COD with abstention. For training, CFCamo optimizes a Qwen3‑VL‑4B‑Instruct agent with Counterfactual Sequence Policy Optimization (CSPO), which samples paired original‑counterfactual rollouts and uses a Counterfactual Paired Reward (CPR) to couple original‑image detection with counterfactual abstention. On CAMO‑test, CFCamo improves S_alpha by +3.7 pp over the prior RL‑based COD baseline; across CF‑COD, it reaches 80.0‑90.8% PA. Ablations show that removing counterfactual coupling reduces PA to 1.4‑5.2% despite strong target‑present COD scores, showing that target‑present evaluation alone does not characterize detect‑or‑abstain behavior. Overall, these results indicate that CFCamo improves COD agents by coupling target‑present detection with target‑absent abstention, rather than merely strengthening target‑present localization. Code and data are available at https://github.com/suhang2000/CFCamo.
Authors:Junke Wang, Xiao Wang, Jiacheng Pan, Xuefeng Hu, Feng Li, Jingxiang Sun, Chaorui Deng, Zilong Chen, Yunpeng Chen, Kaibin Tian, Matthew Gwilliam, Hao Chen, Danhui Guan, Kun Xu, Weilin Huang, Zuxuan Wu, Haoqi Fan, Yu-Gang Jiang, Zhenheng Yang
Abstract:
This paper introduces ARM, a discrete representation‑based AutoRegressive Model that unifies image understanding, generation, and editing within a next‑token prediction framework. ARM is built on three efforts: first, we train a discrete semantic visual tokenizer that maps images into compact token sequences. Our tokenizer is supervised with multiple objectives that jointly promote semantic discriminability, language alignment and faithful reconstruction, thereby supporting diverse tasks in a shared latent space. With this, we train a 7B autoregressive model over large‑scale text and image token sequences, seamlessly developing vision‑language perception and generation capabilities. Finally, to further improve preference‑aligned behavior for text‑to‑image generation and instruction‑guided editing, ARM applies reinforcement learning (RL) to optimize task‑level objectives such as visual quality, instruction adherence, and edit consistency. Surprisingly, the results show that RL not only substantially improves performance on the target tasks (e.g., raising WISE overall from 0.50 to 0.56, GEdit‑Bench‑EN G_O from 5.75 to 6.68), but also induces cross‑task synergy between text‑to‑image generation and editing. Collectively, these findings highlight autoregressive modeling, when paired with strong representations and preference optimization, as a scalable foundation for multimodal intelligence. Code: https://github.com/wdrink/ARM.
Authors:Gangwei Xu, Qihang Zhang, Jiaming Zhou, Xing Zhu, Yujun Shen, Xin Yang, Yinghao Xu
Abstract:
Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence and limited converged accuracy, particularly at high frame rates, as the training supervision is confined to the current chunk without explicit signals about future dynamics; they also suffer from slow inference due to iterative video denoising. In this paper, we present Next Forcing, a multi‑chunk prediction (MCP) framework for causal world modeling that enables faster training, higher accuracy, and accelerated inference. Inspired by multi‑token prediction in large language models, Next Forcing introduces an MCP training objective that augments the main model with lightweight auxiliary MCP modules to simultaneously denoise video chunks at multiple future temporal horizons (next^1, next^2, next^3 chunks). These MCP modules form a causal chain across prediction depths, where intermediate features fused from multiple layers of the main model are leveraged to predict future dynamics, allowing near‑future predictions to inform farther‑future ones and providing dense multi‑scale temporal supervision back to the main model. During training, the MCP modules significantly accelerate convergence and improve converged accuracy, especially at high frame rates: at 50 fps, Next Forcing achieves a 93.1% relative improvement over LingBot‑VA at 5k training steps and 2.3x faster convergence, and establishes new state‑of‑the‑art results on the RoboTwin benchmark (94.1/93.5% on Clean/Random). At inference, the MCP modules can be retained to predict the next video chunk in parallel with the current one, achieving 2x inference acceleration. Next Forcing also demonstrates significant improvements on PhyWorld, a benchmark evaluating adherence to physical laws in video generation, and over 50% FVD reduction on general video pretraining.
Authors:Hangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu, Wenqi Shao, Ying Fu
Abstract:
Low‑light video enhancement (LLVE) remains a challenging task due to severe information degradation under low‑illumination conditions. Recent multimodal approaches have significantly improved enhancement performance by incorporating auxiliary modalities, such as event streams and infrared images. However, these methods typically assume the availability of these modalities at inference, which is often not feasible in real‑world scenarios. To solve this problem, in this work, we propose AMNet, a unified multimodal framework for LLVE, to support flexible modality‑agnostic inference, where auxiliary modalities may be unavailable. To address the issue of modality absence, we introduce a Spatial‑Spectral Dual‑Gated Translator that learns the correspondence between auxiliary modalities and RGB inputs, producing implicit auxiliary representations to support the robust enhancement. Additionally, to fully facilitate the learning of cross‑modal correspondence, we conduct large‑scale multimodal pretraining based on the RGB‑only dataset with synthetic auxiliary modalities. Extensive experiments demonstrate that AMNet could handle arbitrary inference‑time modality combinations and exhibits superior performance for LLVE under modality absence conditions. Code and models are available on the project page.
Authors:Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee, Joungbin Lee, Siyoon Jin, Heeseong Shin, Jung Yi, Yunjin Park, Chulmin Park, Seungryong Kim
Abstract:
Diffusion‑based lip synchronization models achieve strong visual quality and audio‑visual alignment, but full‑sequence bidirectional attention and many denoising steps make them impractical for real‑time inference. We present Lip Forcing, to our knowledge the first autoregressive diffusion method for video‑to‑video (V2V) lip synchronization, which distills a 14B audio‑conditioned bidirectional video diffusion teacher into causal students. At inference, the students generate each chunk in only two denoising steps without inference‑time CFG, enabling real‑time lip synchronization. A lip‑sync‑specific teacher‑trajectory analysis reveals a CFG fidelity‑sync tradeoff: no‑CFG predictions favor reference fidelity, whereas CFG‑guided predictions favor synchronization within a mid‑trajectory band. Lip Forcing translates this finding into three analysis‑derived components: Sync‑Window DMD, a two‑step inference schedule, and a SyncNet‑based reward. We validate Lip Forcing at two student scales, both distilled from the 14B teacher. The 1.3B student crosses into real‑time streaming at 31 FPS, 17.6× faster than its same‑scale bidirectional model. The 14B student, the largest diffusion model reported for V2V lip synchronization, runs 39.8× faster than its teacher at comparable reference fidelity. Time‑to‑first‑frame is sub‑millisecond at both scales, far below every diffusion baseline.
Authors:Kevin Qinghong Lin, Batu EI, Yuhong Shi, Pan Lu, Philip Torr, James Zou
Abstract:
Data tells stories that shape society; the data journalist's job is to turn raw information into stories non‑experts can trust. A high‑quality news feature takes a newsroom team weeks: hunting for context, running statistics, choosing an angle, and designing visuals. Recent agents handle individual steps well: data‑science agents close the analysis loop, while design agents synthesize beautiful websites. But can an agent serve as a data journalist end to end? We introduce Data Journalist Agent (Data2Story), a multi‑agent framework that orchestrates specialized roles into a single virtual newsroom. Data2Story contributes two innovations. (i) Claims are evidence‑grounded: an Inspector links every number, angle, and asset back to data, code, or an external reference. (ii) Articles are multimodally generative: rather than defaulting to plain text and static charts, Data2Story reasons about what readers will want to see, then deploys multimodal tools, such as interactive maps for geography and audio for music. We evaluate Data2Story on 18 articles, each paired with the originally published expert piece, along four axes: (a) human‑agent angle coverage; (b) rubric evaluation with 53 participants across five dimensions; (c) computer‑use agents as judges, a cost‑saving proxy for how readers navigate interactive articles; and (d) verifiability, where a coding verifier re‑executes statements against the data and checks claims against references. Data2Story produces competitive, evidence‑traceable multimedia stories, with particular strength in transparency and auditability. Human articles retain an edge in editorial angle, creative design, and presentation. We position Data2Story as a collaborator for journalists, enabling more evidence‑based, transparent, and verifiable reporting. Code and demos are available at https://data2story.github.io.
Authors:Yikang Yang, Zhanpeng Hu, Youtian Lin, Mengqi Zhou, Jingxi Xu, Feihu Zhang, Jiaheng Liu, Yao Yao
Abstract:
Multimodal large language models can write code to produce complex programs as well as use programs to do 3D modeling, which opens up a new avenue for 3D generation powered by their priors, world knowledge and reasoning. Yet existing benchmarks rarely evaluate 3D modeling through code. Such modeling demands more than runnable code: from a text or visual specification, a model must generate a parametric 3D program that is geometrically precise, semantically aligned and assembly‑consistent. We introduce P3D‑Bench, a benchmark for parametric 3D generation. Unlike a 3D mesh, a parametric 3D program exposes explicit dimensions, construction operations and part relations, revealing whether a model recovers a design's structure, not just its appearance. Under a unified protocol, P3D‑Bench covers three task families (Text‑to‑3D, Image‑to‑3D and Assembly‑3D) and scores each output for executability, geometric fidelity, topology, text‑grounded constraints, multiview semantic alignment and part‑level structure. We evaluate frontier MLLMs and text‑only LLMs on 400 text cases, 400 image cases and 203 annotated assemblies, with domain‑specific models as reference points. Our extensive evaluation yields three findings. First, assemblies are the hardest setting, where models still fail to compose multiple parts into a coherent structure. Second, models can often recover the global shape and semantic identity of the target object, yet fail to reproduce the precise parametric geometry specified by the input. Third, part‑level modeling remains weak on assemblies, where models recover neither the geometry of each part nor the right number of parts. These results position P3D‑Bench as a benchmark for evaluating precise parametric geometry and part‑level structure in parametric 3D generation.
Authors:Yuke Zhao, Wangbo Zhao, Weijie Wang, Zeyu Zhang, Dakai An, Akide Liu, Yinghao Yu, Jiasheng Tang, Fan Wang, Wei Wang, Bohan Zhuang
Abstract:
We introduce WorldOlympiad, a benchmark for diagnosing video‑based world models across physical faithfulness, geometric consistency, and interaction fidelity. While existing benchmarks often focus on visual quality, semantic alignment, or short‑term temporal coherence, they provide limited insight into whether generated videos obey physical rules, preserve coherent 3D structure, and sustain controllable interactions over long horizons. To address this gap, WorldOlympiad decomposes world‑model evaluation into three complementary dimensions. The physical track uses object segmentation and MLLM‑as‑judge to assess whether generated videos follow interpretable rules in mechanics, thermal phenomena, and material properties. The geometry track reconstructs generated videos with Gaussian splatting and evaluates structural consistency, cross‑view coherence, and camera‑trajectory alignment. The interaction track assesses whether generated rollouts follow complex action prompts and maintain smooth, coherent transitions across consecutive video chunks. WorldOlympiad further covers three major downstream scenarios, including gaming, robotics, and general real‑world videos, capturing diverse challenges from interactive control and embodied manipulation to open‑domain motion and camera dynamics. Together, these tracks and scenarios form a scalable and interpretable evaluation suite that exposes failure modes beyond generic video quality. Experiments on state‑of‑the‑art models reveal substantial gaps in physical reasoning, 3D consistency, and long‑horizon interaction, underscoring the need for more structured evaluation protocols for generative world models.
Authors:Mahmood Alzubaidi, Uzair Shah, Raden Muaz, Ines Abbes, Nader Mohammed, Abdullatif Magram, Khalid Alyafei, Mowafa Househ, Marco Agus
Abstract:
A global shortage of trained sonographers limits prenatal ultrasound screening in low‑ and middle‑income countries, where over half of pregnant women receive no skilled sonography. Current deep learning approaches address detection, segmentation, or classification in isolation, each demanding a separate model and expert‑specified labels at inference. We present FADA, a unified vision‑language model built on Qwen3.5‑VL that performs clinical interpretation, classification, detection, and segmentation through a single interpretation‑first pipeline without external labels. FADA distills knowledge from four domain‑specific foundation models (FetalCLIP, UltraSAM, USF‑MAE, UltraFedFM) via offline pre‑computed feature caching. Selective distillation, which applies feature alignment only to annotation tasks while interpretation relies on standard fine‑tuning, consistently outperforms full distillation across most evaluation axes. The recommended variant, FADA‑SKD, achieves 0.8820 mean Dice for segmentation, 0.7671 mAP@0.50 for detection, and 100% structured interpretation compliance. Expert sonographer validation across 237 images confirms clinically acceptable outputs in both autonomous and human‑in‑the‑loop modes, with 73.5% of interpretations scoring perfectly under clinician guidance. The system is trainable on a single consumer GPU and deployable without cloud connectivity. We validate edge deployment by running the compressed 0.8B model on a commodity smartphone (Qualcomm Snapdragon 7 Gen 1, 12 GB RAM) using llama.cpp with GGUF quantization, completing the full 5‑phase pipeline in approximately 60 seconds entirely offline. This establishes a practical pathway for integrating AI‑assisted fetal assessment with portable ultrasound devices, directly addressing diagnostic access gaps in resource‑constrained settings. Code, models, and data are available at https://github.com/mahmoodphd/FADA.
Authors:Yitong Chen, Zijie Diao, Junke Wang, Lingyu Kong, Yixuan Ren, Bo He, Yu-Gang Jiang, Zuxuan Wu
Abstract:
Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spaces for image generation. However, their reconstruction quality often remains suboptimal, largely because deep VFM representations do not preserve sufficient fine‑grained visual detail. This limitation becomes even more severe after discretization, where missing low‑level information is difficult to recover. In fact, we observe that shallow VFM features retain considerably richer local appearance and structural detail, which complements the high‑level semantics carried by deep features used in existing RAEs. Motivated by this complementary property, we propose Ideal, an In‑depth Alignment framework for discrete representation autoencoding. By jointly aligning quantized tokens with both shallow and deep VFM features, Ideal enables the resulting discrete visual tokens to preserve both visual fidelity and rich semantics. Extensive experiments demonstrate that Ideal yields superior reconstruction performance, achieving 0.61 rFID on ImageNet and outperforming the previous best method by 0.28. When used for autoregressive image generation, Ideal further produces a gFID of 1.89, establishing a new state of the art for autoregressive image generation.
Authors:Jaewoo Lee, Zaid Khan, Archiki Prasad, Justin Chih-Yao Chen, Supriyo Chakraborty, Kartik Balasubramaniam, Sambit Sahu, Elias Stengel-Eskin, Hyunji Lee, Mohit Bansal
Abstract:
Various test‑time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre‑execution action evaluation in complex Graphical User Interface (GUI) environments. However, existing critics suffer from two key limitations: they (1) focus primarily on short‑sighted decision loops (e.g., forgetting earlier actions) and (2) lack the visual grounding needed to detect flawed actions (e.g., clicking wrong UI elements). To address these, we introduce HiViG, a History‑aware Visually Grounded test‑time framework, built around a multimodal critic trained on real GUI trajectories to abstract past interactions into a compact record and to evaluate actions with visual grounding. At test time, HiViG integrates the critic into the policy decision loop to provide macro‑action history, which summarizes the policy's completed achievements, and visually grounded critique, which verifies raw execution coordinates against the current screenshot to intercept errors before execution. Across web, mobile, and desktop benchmarks, HiViG consistently outperforms existing scalar and verbal critics, improving average success rates over the strongest baseline by 5.8% for Qwen3‑VL‑32B and 9.0% for Gemini‑3‑Flash, and demonstrates strong cross‑platform generalization. Ablations show that macro‑action history mitigates short‑sighted planning and visually grounded critique reduces execution errors, with both components being critical for test‑time scaling in long‑horizon GUI tasks.
Authors:Zhiwen Yang, Jiayin Li, Hao Lu, Hui Zhang, Zihua Wang, Bingzheng Wei, Yan Xu
Abstract:
Existing deep learning models for Positron Emission Tomography (PET) image denoising often suffer from severe performance degradation under distribution shifts, fundamentally restricting their robust clinical deployment. This lack of generalization stems from the conventional paradigm of fixed‑parameter models that cannot adapt to variations in test data (e.g., dose levels or scanner types) after training. To overcome this limitation and achieve robust generalization, we introduce U‑TTT, a novel U‑shaped model that integrates Test‑Time Training (TTT) layers to dynamically adjust model parameters during inference through self‑supervision, thereby adapting to the specific characteristics of each test instance. Furthermore, to comprehensively capture the complex degradations of 3D PET data, U‑TTT features a dual‑domain adaptation mechanism comprising a Spatial Test‑Time Training (S‑TTT) layer and a Frequency Test‑Time Training (F‑TTT) layer. The S‑TTT layer captures and corrects spatial structural degradations, while the F‑TTT layer suppresses global noise spectra and restores delicate high‑frequency details. Extensive experiments demonstrate that U‑TTT achieves state‑of‑the‑art PET denoising performance and exhibits superior generalization under challenging distribution shifts, including both unseen dose levels and unseen scanners. Our code will be available at https://github.com/Yaziwel/U‑TTT.
Authors:Taishan Li, Jiwen Zhang, Siyuan Wang, Xuanjing Huang, Zhongyu Wei
Abstract:
Vision‑Language‑Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task‑relevant objects are fully visible. This assumption often fails in realistic settings, where occlusion makes manipulation partially observable. In this paper, we study scene‑induced occlusion as a fundamental challenge for VLA models and introduce LIBERO‑Occ, an occlusion‑oriented extension of LIBERO. Experiments show that state‑of‑the‑art VLAs suffer substantial performance degradation under occlusion. To address this issue, we propose Viewpoint Imagination (VIM), which generates a complementary view from an occluded primary observation and conditions action prediction on both observed and imagined evidence. VIM improves robustness across task suites, occlusion types, and severity levels without requiring additional cameras at deployment time, suggesting that viewpoint imagination is an promising mechanism for perception completion in partially observable manipulation. Our benchmark and corresponding code are available at: \hrefhttps://github.com/litsh/Libero‑Occhttps://github.com/litsh/Libero‑Occ.
Authors:Cong Wang, Zhentao Yu, Hongmei Wang, Weicong Liang, Zixiang Zhou, Zilin Yang, Jiarong Ou, Rui Chen, Yuan Zhou, Qinglin Lu
Abstract:
Current identity‑consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi‑view reference input offers a natural solution, progress remains constrained by the lack of effective frameworks for multi‑view inputs and the scarcity of multi‑view data. We address these challenges by proposing HarmoView, a robust framework for identity‑consistent video generation that effectively integrates multi‑view cues through three architectural refinements complemented by a staged training curriculum. Specifically, we first introduce Multi‑level Feature Injection to anchor identity fidelity; by injecting raw ViT features from frontal references alongside text tokens via cross‑attention, MFI provides persistent low‑level appearance anchors that complement the high‑level identity features within DiT blocks, leading to enhanced identity preservation. Then, we employ learnable proxy tokens to unify heterogeneous reference layouts across single‑/multi‑view settings while simultaneously resolving the reference‑view mismatch problem. Jump‑RoPE is further developed for identity‑wise feature isolation to reduce identity crosstalk. To activate these structural capabilities while preserving the original generative priors, we propose the Progressive View Curriculum. This four‑stage training strategy employs view dropout to facilitate a stable transition from vanilla T2V generation to high‑fidelity, identity‑persistent spatial reasoning. Furthermore, we construct a large‑scale multi‑view dataset to address the issue of data scarcity. Extensive evaluation on our multi‑view benchmark, comprising 100 manually‑curated cases spanning 52 unique identities, demonstrates that HarmoView significantly outperforms open‑source baselines and matches leading closed‑source engines, achieving state‑of‑the‑art performance in identity‑consistent video generation.
Authors:Wenhao Yan, Fengjia Guo, Zhuoyi Yang, Jie Tang
Abstract:
Controlled character animation requires transferring motion from a driving sequence to a reference character. Prior works heavily rely on intermediate representations, including pose skeletons to represent motion or masked background to represent environment, which inevitably leads to information loss. To address this, we present SCAIL‑2, a framework that bypasses those intermediates and achieves end‑to‑end character animation. By directly concatenating driving videos to the sequence, the model can obtain all the required visual information from the input video. To address the lack of end‑to‑end data, we unify sub‑tasks of character animation with decoupled conditions and then curate a pipeline to synthesize MotionPair‑60K, an end‑to‑end motion transfer dataset containing heterogeneous tasks of character animation. To achieve the unification, we utilize in‑context mask conditioning and mode‑specific RoPE as soft guidance beyond textual instructions and raw visual information. To address synthetic discrepancy in detailed regions, we propose Bias‑Aware DPO to construct preference items to mitigate the errors. Extensive experiments demonstrate that our method substantially outperforms existing state‑of‑the‑art approaches in various character animation tasks. A large subset of synthetic data as well as model weights will be released at our project page: https://teal024.github.io/SCAIL‑2/.
Authors:Qiaoxin Li, Caini Pan, Pierre-Antoine Comby, Chaithya Giliyar, Philippe Ciuciu
Abstract:
Accelerated acquisition of fMRI enables enhanced detection of neurovascular (BOLD) activity in the brain, but image reconstruction becomes challenging with high k‑space undersampling: Task‑evoked BOLD signals are small in magnitude, which traditional anatomical MRI reconstruction methods fail to recover, as they favor spatial accuracy over temporal fidelity. We present DD‑INR, a Dynamics‑Driven Implicit Neural Representation framework tailored for accelerated fMRI that benefits from incoherent time‑varying sampling and a tailored spatiotemporal prior, outperforming traditional methods, demonstrated in simulation and in‑vivo acquisition, both in terms of image quality and retrieval of activation patterns. DD‑INR achieves this by splitting the fMRI data into a static background and a temporally varying dynamic component, representing only the dynamics with a dedicated INR, thereby focusing the model's capacity on activation‑relevant changes while remaining compact. In general, DD‑INR provides a promising framework for accelerated fMRI reconstruction, with the potential to improve the sensitivity and robustness of fMRI studies within practical scan time limits. The source code is available at https://github.com/JoosenLi/DD‑INR.
Authors:Ana Sofia Santos, André Ferreira, Gijs Luijten, Naida Solak, Lisle Faray de Paiva, Behrus Hinrichs-Puladi, Jens Kleesiek, Jan Egger, Victor Alves
Abstract:
The nnU‑Net has demonstrated continuous success in medical segmentation tasks, which heavily rely on the availability and diversity of annotated biomedical data. However, assembling medical imaging cohorts remains challenging due to numerous factors such as privacy regulations and annotation costs. As a result, data augmentation plays a crucial role in increasing data availability while maintaining anatomical feasibility. Hence, we propose the ++nnU‑Net, a novel data augmentation module based on image registration that operates prior to preprocessing and training take place. Our framework was evaluated across five different 2D datasets. In this workflow, image data go through a two‑stage registration process, generating new warped images. The transformations are then applied to the respective segmentation. In addition, the pipeline computes available disk space, generates supplementary binary synthetic masks and generates checkpoints. We demonstrate that the ++nnU‑Net outperforms the nnU‑Net baseline, yielding improvements in Dice Similarity Coefficient scores. In the most prominent cases, we observe performance gains of approximately 22%. These findings highlight the effectiveness of registration‑based data augmentation, particularly for 2D medical imaging datasets and suggest that the ++nnU‑Net provides a practical and scalable approach for enhancing segmentation performance in data‑limited settings. The source code for the ++nnU‑Net is available at: https://github.com/sofia‑adelie/plusplusnnunet.git
Authors:Yinglong Yan, Yunkai Yang, Haoyi Wang, Wei Fu, Linshan Wu, Honghu Pan, Shaobo Xia, Shanghang Zhang, Hao Chen, Leyuan Fang
Abstract:
Remote sensing vector mapping aims to generate structured maps of geospatial entities, such as buildings, roads, and water bodies, from remote sensing imagery. In practice, vector maps usually contain multiple category layers and heterogeneous entity structures, requiring a unified model for diverse mapping needs. However, existing methods typically represent vector objects as polygons or graphs, making them suitable only for specific categories: polygons poorly capture topological relations, while graphs often blur instance boundaries. We observe that language, as a natural medium for human communication, offers a flexible and expressive representation that can accommodate heterogeneous map elements, including geometry, semantics, and topolog. Motivated by this insight, we propose Vector Map as Language (VecLang), a unified paradigm that reformulates multiclass vector mapping as structured text generation. VecLang encodes the common elements of different geospatial entities into a GeoJSON‑like vector language, enabling cross‑category modeling within a shared textual format. To generate this language reliably, we design a progressive vision‑language mapping framework that first localizes vectorization units and then generates structured map elements. We further introduce Hierarchical Vector Language Optimization, which uses reinforcement learning to improve syntax validity, content fidelity, and map executability. We also build VecMap‑Bench with 54K images and 800K instances, supporting training and evaluation across standard and generalization settings. Extensive experiments demonstrate that VecLang handles both single‑class and multiclass vector mapping while achieving strong cross‑dataset and open‑vocabulary generalization. The model and dataset are publicly available at https://github.com/yyyyll0ss/VecLang.
Authors:Qi Song, Yifei He, Chi Zhang, Zheng Fu, Xuhe Zhao, Mengmeng Yang, Kun Jiang, Rui Huang, Diange Yang
Abstract:
Forecasting the future evolution of dynamic scenes is crucial in autonomous driving. However, existing feed‑forward paradigms are primarily designed for interpolation. When extended to future extrapolation, they suffer from ghosting artifacts under large displacements and are constrained by simplified motion assumptions or strict future priors. To overcome these challenges, we propose Envision4D, a fully self‑supervised feed‑forward framework for pose‑free future extrapolation. Specifically, we introduce a Future Pose Prediction module that infers future camera parameters via an iterative denoising process. Furthermore, to capture non‑linear dynamics, we propose In‑layer Temporal Attention and employ Conditioned Motion Lifting, which transforms the highly uncertain extrapolation process into robust relational mappings. Finally, a Progressive Training Strategy is utilized to stabilize unsupervised motion learning against error accumulation. Extensive experiments demonstrate that Envision4D achieves state‑of‑the‑art performance, significantly outperforming existing methods in future view synthesis.
Authors:Hao Liu, Ruping Cao, Kun Wang, Zhiran Li, Fan Liu, Yupeng Hu, Liqiang Nie
Abstract:
In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structured chart data and generate faithful natural‑language summaries from chart images. To address the complementary requirements of accurate data extraction and factual narration, we propose ChartLens, a dual‑branch framework for chart data correction and summary refinement. ChartLens consists of two key modules: Structure‑Aware CSV Verification and Correction (SAVC) and Text‑Retention‑Guided Summary Refinement (TRSR). SAVC improves the reliability of structured data extraction through verification and correction, while TRSR enhances summary generation by preserving critical textual and numerical evidence from charts. By combining model adaptation, correction‑based generation, and OCR‑assisted evidence grounding, ChartLens improves both structured data recovery and summary factuality. On the test set, our final system achieves an overall score of 69.10 and ranks first in Track 2, demonstrating its effectiveness for accurate chart understanding. Our code will be released at: https://github.com/iLearn‑Lab/CVPRW26‑ChartLens.
Authors:Zhengxuan Wei, Yi Dong, Zonghui Li, Xianhui Lin, Xing Liu, Hong Gu, Shaofeng Zhang, Wenbin Li, Qi Fan
Abstract:
Low‑Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, existing LoRA merging techniques often suffer from severe parameter interference, causing destructive collisions in the shared parameter space. To address this, we propose Subspace Signal Routing (SSR), which resolves interference by routing internal signals instead of performing parameter‑space merge. Specifically, SSR first constructs a unified subspace by concatenating candidate LoRAs along the rank dimension. Next, SSR employs an inverse correlation matrix to decorrelate mixed signals within this space. Finally, a directional guide matrix steers these purified signals into their respective task‑specific subspaces. We provide a rigorous theoretical analysis proving that SSR aligns with the Ordinary Least Squares (OLS) solution, thereby ensuring mathematical optimality. We utilize the additivity of sufficient statistics to design a streaming algorithm. This enables on‑the‑fly updates that significantly reduce memory overhead and computation time. Extensive experiments validate that SSR significantly outperforms state‑of‑the‑art methods while maintaining comparable efficiency. Code is available at https://github.com/nagara214/SSR‑Merge.
Authors:Haoliang Han, Ziyuan Luo, Renjie Wan
Abstract:
3D Gaussian Splatting (3DGS) is a powerful technique for creating high‑fidelity 3D assets. However, the widespread sharing and iterative modification of 3DGS models across digital platforms create pressing challenges for intellectual property protection and forensic traceability. To address this, we propose GaussTrace, a novel framework for constructing directed provenance graphs for 3DGS models. GaussTrace formulates provenance analysis as an evidence‑based reasoning problem. It builds upon attribute‑wise statistical profiling of 3DGS parameters to capture intrinsic properties. Moreover, we introduce hypothesis‑driven editing simulations of common operations to provide auxiliary evidence for plausible transformation pathways. These statistical and simulated cues jointly enable a Large Language Model (LLM) to perform structured Chain‑of‑Thought (CoT) reasoning, yielding directional provenance inferences and explainable edge reasons. Experimental results demonstrate that GaussTrace effectively constructs evolutionary relationships among diverse 3DGS models, delivering accurate, interpretable, and robust provenance graphs without requiring model training or access to editing histories. Project page: https://haolianghan.github.io/GaussTrace.
Authors:Mao Chen, Xu Yang, Chuankai Liu, Xiangkai Zhang, Xiaoxue Wang, Zheng Bo, Zuoyu Zhang, Zhiyong Liu
Abstract:
Precise rover localization is a prerequisite for autonomous lunar exploration, yet the absence of Global Navigation Satellite System (GNSS) signals and the cumulative drift of local localization methods severely constrain long‑range missions. Cross‑view localization provides a promising drift‑free global solution by matching rover‑view and satellite‑view imagery. However, the lunar environment poses unique challenges for correspondence alignment, including inter‑entity entanglement, inter‑viewpoint divergence, and simulation‑to‑real domain shift. To address these challenges, we propose Warped Alignment of Reprojected Graphs (WARG), a framework that leverages unified graph learning and reprojected graph matching for robust cross‑view alignment. Pretrained on the synthetic LuSNAR dataset, WARG achieves an average test error of 0.32 m and demonstrates robust zero‑shot generalization to the synthetic lunar south pole region with an error of 3.63 m. More importantly, when validated on real‑world data from the YuTu‑2 rover, WARG achieves a localization error of 1.68 m within a 100 m x 100 m search area, corresponding to nearly one‑pixel precision in low‑resolution satellite imagery with a spatial resolution of 1.40 m/pixel. Beyond accuracy, WARG is computationally efficient, containing only 1.56M parameters, corresponding to 16.12% of previous lightweight models, and operating at 5.49 Hz on an NVIDIA RTX A6000 GPU, approaching GNSS‑level update frequency. Finally, we observe that WARG naturally develops low‑level spatial awareness, including semantic segmentation and structural reasoning, through cross‑view localization learning, highlighting its potential as a promising paradigm for spatial intelligence with minimal annotation cost. The source code is available at https://github.com/maochen‑casia/warg.
Authors:Haodong Lei, Hongsong Wang, Bingxuan Dai, Pan Zhou
Abstract:
The growing need for high‑resolution image generation in autoregressive text‑to‑image models has resulted in extended token sequences, significantly increasing computational costs and inference times. However, existing state‑of‑the‑art methods for accelerating autoregressive text‑to‑image models rely on chain‑structured draft token sequences, leading to inefficient draft token search and limited acceptance lengths. To address this, we propose parallel‑path cross‑relaxed speculative Jacobi decoding (PathSpec), a novel framework that enhances efficiency through a multi‑sequence draft tree structure. Our parallel‑path speculative Jacobi decoding (PathExplore) expands the token search space, achieving a higher speedup ratio without sacrificing image quality. Additionally, we introduce cross‑path relaxed verification (PathRelax) that exploits semantic similarities across sequences to further boost token acceptance rates. Evaluated on the Parti‑Prompts, MSCOCO2017, and T2ICompBench datasets, our method achieves a speedup ratio of 4.14 ×, 3.95×, and 4.18×, respectively. Remarkably, PathExplore, without any relaxed sampling, outperforms relaxed sampling methods in the speedup ratio, such as GSD and LANTERN. Moreover, PathRelax's relaxation mechanism can be seamlessly integrated with other relaxation techniques, enabling further acceleration and providing an efficient solution for real‑time text‑to‑image generation. Our code is available at https://github.com/Haodong‑Lei‑Ray/PathSpec.
Authors:Yifan Zhu, Can Lin, Hangjie Yuan, Zixiang Zhao, Pengfei Zhang, Tao Feng, Zhonghong Ou
Abstract:
Parameter‑Efficient Fine‑Tuning (PEFT) methods provide a streamlined and efficient tool for adapting large models to domain‑specific multimodal downstream tasks. Although these methods proved their tangible effects in practice, their principal aspects remain under‑explored. Therefore we remain curious about the underlying generalization mechanisms in various PEFT methods and how they can be further enhanced. In this paper, we reveal the flatness preference widely present in various PEFTs, where a small fraction of sharp dimensions dominates the generalization of PEFT. This finding suggests an appealing possibility: we may be satisfied with a better generalization by merely attending to this small fraction of sharp dimensions instead of all of them. Furthermore, we propose Flatness Preference Optimization (FlatPO) to flatten these key sharpness dimensions, leading various PEFTs toward better generalization. Extensive experiments demonstrate the effectiveness of our findings and the proposed method. Code is available at https://github.com/Can‑Lin/FlatPO.
Authors:Dahye Kim, Jaehyun Choi, Hyun Seok Seong, Seongho Kim, Donghun Lee, Sungwon Yi, Jang-Ho Choi
Abstract:
While existing AI‑generated image detectors report high performance, we identify that this is largely driven by a critical prediction asymmetry: a bias toward the real class that severely limits sensitivity to generated content, especially under standard post‑processing operations such as compression and resizing. We hypothesize that this stems from the model's reliance on spurious features, distracting signals that obscure true generative artifacts. To address this, we propose DEAR (Dissect and Prune), which leverages inpainted images to identify and prune these interfering components. Specifically, we find that features strongly aligned to either inpainted or non‑inpainted regions are less robust to post‑processing. By measuring the alignment between channel activations and inpaint masks, DEAR removes features at both extremes, retaining only those that capture genuine generative artifacts. Experimental results demonstrate that our approach significantly enhances robustness against unseen generators and post‑processing, effectively mitigating the prediction asymmetry. Our code is available at https://github.com/dahyedahye/dear.
Authors:Fen Peng, Taizo Suzuki, Seisuke Kyochi
Abstract:
In this study, we propose an overlapped wavelet diffusion framework for Low‑Light Image Enhancement (LLIE), which incorporates two complementary components to achieve blocking artifact‑free and detail‑preserving enhancement. Although recent diffusion‑based LLIE methods have demonstrated remarkable performance compared with traditional approaches, DiffLL still suffers from blocking artifacts caused by the Haar Wavelet Transform (WT) and blurred edges or over‑smoothed textures due to the limitations of its High‑Frequency Restoration Module (HFRM). To overcome these issues, we introduce an Overlapped WT (OWT) that incorporates correlations across neighboring regions, thereby structurally preventing blocking artifacts. Furthermore, we integrate a low‑frequency‑guided High‑Frequency Enhance Block (HFEBlock) to strengthen detail recovery, yielding sharper edges and more reliable textures. Extensive experiments on the LOLv1 and LOLv2‑real datasets demonstrate that our framework, termed OWDiff, consistently outperforms existing LLIE methods both qualitatively and quantitatively, achieving superior visual quality while maintaining computational efficiency. OWDiff effectively addresses the structural limitations of the Haar WT and the HFRM, achieving an average PSNR gain of 0.58 dB, along with a 1.64% relative improvement in SSIM and a 5.9% relative reduction in LPIPS, compared to DiffLL across both the LOLv1 and LOLv2‑real datasets.
Authors:Ghodsiyeh Rostami, Po-Han Chen, Mahdi S. Hosseini
Abstract:
Parameter‑efficient fine‑tuning (PEFT) aims to adapt pretrained models with a small trainable parameter subset, however, most existing methods choose this subset from fixed architectural heuristics rather than using dynamic, task‑aware criteria. We introduce FisherAdapTune, a Fisher‑guided Adaptive Fine‑Tuning framework that progressively selects parameter groups by tracking the temporal drift of their Fisher geometry. Starting from a PAC‑Bayesian view of fine‑tuning, we decompose the generalization error bound into Fisher‑weighted update costs and show that parameter groups whose curvature contribution has stabilized can be frozen to reduce the error bound without interrupting the remaining adaptation dynamics. FisherAdapTune formulates this criterion with a scale‑invariant Jensen‑Shannon distance between consecutive Fisher distributions, yielding an adaptive active parameter set. We evaluate our approach on a downstream segmentation task, and results show FisherAdapTune improves the in‑distribution performance and zero‑shot transfer in multiple settings, validating that Fisher structural drift is a useful signal for efficient, task‑aware adaptation. We release our \hrefhttps://github.com/AtlasAnalyticsLab/FisherAdapTunecode publicly to enable further application of our proposed approach.
Authors:Taehyoung Kim, Tim Schoenbrod, David Eckel, Henri Meeß
Abstract:
Recent learning‑based path planners use neural networks to process visual map representations and approximate heuristics for classical search algorithms, yielding near‑optimal paths with reduced search effort. However, these methods are tied to the shortest‑path objective implicit in their supervision, which limits their flexibility to accommodate alternative criteria. We introduce FlexPath, a two‑stage framework that decouples feasibility from preference. In Stage 1, we use imitation learning to acquire a task‑independent spatial prior over feasible paths from visual map inputs. In Stage 2, differentiable Path Shape Objectives (PSOs) adapt this prior toward task‑specific criteria without relearning path structure, requiring only efficient objective‑level adaptation. A single pretrained model can be adapted to multiple objectives. For shortest‑path planning, FlexPath reduces search effort on TMP by 14.3% compared to the state‑of‑the‑art TransPath, while also finding lower‑cost paths on average and demonstrating strong zero‑shot generalization across three unseen domains. For obstacle clearance with minimum clearance distance 2, it achieves 96.8% full obstacle avoidance while maintaining low search cost. The framework further extends to semantic‑aware avoidance and waypoint guidance via objective‑level adaptation, and remains compatible with classical planners at inference time. Data and code are available at https://github.com/FraunhoferIVI/FlexPath.
Authors:Nathan Molinier, Adrian A. Marth, Reto Sutter, Christoph Germann, Jacob A. Connolly, Mathieu Guay-Paquet, Nathan D. Schilaty, Kenneth A. Weber, Julien Cohen-Adad
Abstract:
Lumbar spine conditions are a leading cause of disability worldwide, yet reliable quantification of degeneration from MRI remains challenging. In clinical practice, analysis is predominantly performed in two dimensions (2D), as manual three‑dimensional (3D) assessment is time‑consuming. However, 2D measurements suffer from limited reproducibility, particularly when anatomical structures are not aligned with the imaging plane. Existing automated approaches are often restricted to 2D, rely on discrete grading, or lack robustness and interpretability. We introduce SpineReport, an open‑source, fully automated framework for comprehensive 3D morphometric analysis of lumbar spine MRI. Leveraging robust anatomical segmentations, the method extracts quantitative metrics from key structures, including the spinal canal, spinal cord, vertebrae, intervertebral discs, and foramina. These include both morphological and signal‑based features, enabling cross‑subject and longitudinal assessment. SpineReport further generates subject‑specific reports that allow comparison with cohort distributions, improving interpretability and objective characterization of spinal morphology. Clinical relevance was evaluated against radiologist‑reported severity grades for central canal, lateral recess, and foraminal stenosis. Metrics showed strong associations with central canal stenosis severity, with T2‑weighted CSF signal providing the highest performance (AUC = 0.95). Canal AP diameter and area ratios also demonstrated strong correlations and high discriminative ability (AUC > 0.80). For lateral recess stenosis, associations were moderate, with lateral CSF signal being the most informative (AUC = 0.73). No significant associations were observed for foraminal stenosis despite robust region‑of‑interest extraction. SpineReport is released as an open‑access tool: https://ivadomed.github.io/SpineReport/
Authors:Chong Liu, Luxuan Fu, Xuyu Feng, Zhen Dong, Bisheng Yang
Abstract:
The paradigm of digital twin cities is shifting from coarse visual mapping toward more precise and actionable digitization of urban assets. However, existing datasets predominantly focus on coarse visual perception, lacking the strict multi‑modal alignment and attribute and status diagnosis required for automated infrastructure maintenance. To bridge this gap, we introduce WHU‑Infra3D, a large‑scale, multi‑modal benchmark dataset dedicated to roadside infrastructure inventory. Covering 53.8 km across three cities, WHU‑Infra3D uniquely integrates panoramic imagery and LiDAR point clouds with rigorous 2D‑3D instance association and cross‑frame tracking. Comprising over 175k multi‑view 2D bounding boxes alongside thousands of 3D infrastructure instances, the dataset provides over 181k detailed attribute and status annotations (e.g., rust, occlusion) to empower operational health assessment. We establish comprehensive baselines across five core tasks: 2D detection, 2D cross‑view matching, 3D geo‑identification, 3D point cloud segmentation, and attribute recognition. Extensive evaluations expose significant cross‑city domain gaps and inherent vulnerabilities of current models on long‑tailed defective statuses, establishing WHU‑Infra3D as an essential testbed for advancing scalable, AI‑driven urban infrastructure inventory and lifecycle management. The WHU‑Infra3D dataset is available at https://github.com/WHU‑USI3DV/WHU‑Infra3D.
Authors:Yuheng Chen, Teng Hu, Yuji Wang, Qingdong He, Zhucun Xue, Qianyu Zhou, Jason Li, Lizhuang Ma, Jiangning Zhang, Dacheng Tao
Abstract:
The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems showremarkableabilitytogeneratecinematicnarratives, the progress of open‑source models remains limited by the scarcity of high‑quality training data. To bridge this gap, we introduce CineDance‑1M, a large‑scale, open research Text‑to‑Audio‑Video (T2AV) dataset designed specifically for multi‑shot, long‑form joint audio‑video generation. Averaging 92.8 seconds and 24.2 continuous shots per video, it provides configurable, structured annotations for both audio and video modalities. This exceptional quality is achieved through a rigorous three‑stage curation pipeline: i) diverse sourcing and comprehensive cleansing, ii) film‑theory‑inspired narrative parsing, and iii) hierarchical dual‑modal captioning. For a comprehensive assessment, we propose CineBench, featuring a diverse prompt suite and a six‑dimensional, human‑aligned metric system tailored for complex narrative audio‑video evaluation. Furthermore, we adapt LTX‑2.3 into CineDance, which demonstrates exceptional single‑modality quality alongside precise audio‑video alignment and robust subject and environment consistency, effectively validating our curation strategy and the high quality of CineDance‑1M. We anticipate that this work will serve as a solid foundation for accelerating future research in multi‑shot, long‑form joint audio‑video generation. Our project page is available at https://aliothchen.github.io/projects/CineDance/.
Authors:Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh, Rajeev Yasarla, Reza Pourreza, Litian Liu, Risheek Garrepalli, Roland Memisevic
Abstract:
Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large language models (LLMs) as task guidance assistants. A crucial capability for the real‑world success of a prospective task guidance assistant is it's ability to intervene proactively as soon as a mistake is apparent in order to guide the user. To evaluate this crucial capability, we introduce Ego‑MC‑Bench (Mistake Corrections), a benchmark for evaluating reactive, step‑by‑step task guidance in realistic cooking scenarios. Extensive experiments show that Ego‑MC‑Bench is highly challenging for state‑of‑the‑art video LLMs. We argue that a key reason is the limited availability of training data for fine‑tuning models on this task. Although there exists a wide range of cooking video datasets, existing datasets lack examples of mistakes along with appropriately timed interventions. To help address this data limitation, we also introduce Ego‑CoMist, a counterfactual synthetic dataset created by transforming non interactive cooking videos into supervised training examples showing proactive interventions. We show that fine‑tuning on Ego‑CoMist yields performance gains especially for smaller and more efficient video LLMs that are well suited for delivering assistance on edge devices.
Authors:Zibin Liu, Shunkun Liang, Banglei Guan, Yang Shang, Qifeng Yu, Ji Zhao
Abstract:
Despite the rapid advancements in event‑based motion estimation, current geometric methods primarily focus on velocity estimation. However, absolute pose estimation, which is equally crucial for key applications such as robotic navigation and augmented reality, remains relatively underexplored. Consequently, the simultaneous recovery of absolute pose and velocity from event streams remains an open and challenging problem. To address this gap, we propose a geometric framework for absolute pose and velocity estimation by leveraging 3D lines in the scene and the events they trigger. At the core of the framework lie two key geometric constraints: the orthogonality between a 3D line and the normal vector of its corresponding event plane, and the collinearity of an event with the 2D projection of its associated line. Based on these constraints, we present both linear and polynomial solvers for absolute pose estimation. The former enables efficient computation, while the latter provides a globally optimal solution for rotation. For velocity estimation, we develop an efficient linear solver and a more accurate optimization‑based solver to recover both angular and linear velocities. Notably, our methods require a minimum of three event‑line correspondences to determine the 6‑DoF absolute pose or velocities independently. Extensive experiments in simulation and on real‑world datasets demonstrate that our methods achieve state‑of‑the‑art performance, with significant improvements in accuracy and computational efficiency compared to existing methods. The demo code is publicly available at https://github.com/Zibin6/EventPoseVelocity.
Authors:Sheng-Wei Chan, Yung-Che Wang, Hsin-Jui Pan, Chia-Min Lin, Jen-Shiun Chiang
Abstract:
Document image binarization aims to separate foreground text from degraded backgrounds while preserving thin, broken, and low‑contrast strokes. Although deep learning methods have improved binarization performance, most existing approaches rely on convolutional, transformer‑based, or generative architectures, while Mamba‑based state space models remain largely unexplored for this task. In this work, we investigate Mamba‑based feature propagation and observe that direct state‑space propagation may dilute weak foreground cues during long‑range modeling, especially faint ink traces, fragmented characters, and boundary‑sensitive stroke details. To address this problem, we propose DeepMine‑Mamba, a Mamba‑based binarization framework equipped with a novel Anti‑Dilution Gate that estimates propagation‑induced feature changes and selectively restores stroke‑sensitive local responses while suppressing unnecessary background enhancement. Experiments on DIBCO/H‑DIBCO benchmarks under a strict leave‑one‑year‑out protocol show that DeepMine‑Mamba achieves competitive overall performance, with strong average FM and Fps across benchmark years. Ablation results further demonstrate that the Anti‑Dilution Gate improves stroke preservation and reduces perceptually significant binarization errors.
Authors:Md Mahfuzur Rahman Siddiquee, Fazle Rafsani, Jay Shah, Teresa Wu, Catherine D Chong, Todd J Schwedt, Baoxin Li
Abstract:
Abnormality detection is a crucial yet challenging task in medical image analysis. Distinguishing abnormalities from normal data by learning to reconstruct normal‑only data alleviates the reliance on labeled datasets. However, many studies, even if unsupervised, rely on a labeled validation set to select the best model for inference from multiple training iterations. For many diseases labeled data are unavailable and substantially time consuming to obtain. To address this, AUCp ‑ a novel metric that supports abnormality detection for unsupervised and self‑supervised methods is proposed. Instead of evaluating the realism of reconstructed images to select the best of model for inference, it focuses on actual detection performance and without requiring an annotated test set. Assuming the pseudo ground truth of all unannotated samples in the test set as abnormal/positive and using traditional AUC calculation, AUCp scores are derived. Given a large and representative training set of normal samples, we show mathematical and empirical evidence that model selection using AUCp scores improves disease detection in terms of unsupervised and self‑supervised methods over conventional metrics. Using two unsupervised methods for neurologic disease detection and self‑supervised methods on diverse datasets, our results demonstrate that the AUCp score effectively identifies the optimal model for inference, significantly enhancing abnormality and disease detection. The corresponding implementations are available in https://github.com/mahfuzmohammad/AUCp.
Authors:Syed Rifat Raiyan, Mohsinul Kabir, Hasan Mahmud, Md Kamrul Hasan
Abstract:
Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified account of the field's evolution, from early rule‑based math word problem (MWP) solvers and template‑driven geometry systems, through neural expression generation and LLM prompting, to contemporary reasoning models, multi‑agent systems, neuro‑symbolic theorem provers, and verified discovery workflows. We organize the landscape along four axes: (i) informal reasoning over text and diagrams, spanning MWP solving, multimodal geometry, and VLMs; (ii) formal reasoning in proof assistants, including autoformalization, tactic prediction, compiler‑guided repair, and proof search; (iii) mathematical discovery, where systems propose constructions, improve bounds, or assist attacks on open problems; and (iv) the inference and training‑time techniques, including CoT prompting, tool use, process reward models, and RLVR, that increasingly connect generation with verification. We catalog major benchmarks across grade‑school arithmetic, competition mathematics, geometry, formal proving, multimodal and multilingual reasoning, and expert evaluation, and we examine benchmark saturation, contamination, reporting mismatches, and the distinction between pass@1, majority voting, and verifier‑assisted pass@k. We critically assess failure modes: brittleness under perturbation, reward hacking, multimodal grounding failures, fragile formalization, and the energy cost of reasoning‑scale inference. Drawing on recent perspectives from working mathematicians, we identify future directions centered on verified‑discovery workflows, reasoning efficiency, and infrastructure to make AI‑assisted formalization broadly usable. Companion materials: https://github.com/Starscream‑11813/awesome‑AI4Math.
Authors:George Ling, Lijin Yang, Hao Yang, Zhongzhan Huang
Abstract:
We present BLUE, a minimal method for better language use in vision‑language‑action (VLA) models for autonomous driving (AD). Through extensive analysis, we reveal that language matters on only a small fraction of routes, but on those routes it can greatly improve or degrade performance. Generating language at every frame is therefore inefficient, since most computation is spent on frames that do not benefit from language. We further show that pretrained VLA hidden states potentially already encode whether language will benefit a given frame, even though scene complexity and kinematic features alone struggle to predict this. Based on this finding, BLUE trains a lightweight gate on frozen VLA hidden states to decide per frame whether to activate language generation or predict actions directly, without modifying the backbone or requiring additional human annotation. With just a 0.11M‑parameter gate, BLUE sets a new state of the art on both benchmarks, achieving 76.2% success rate on Bench2Drive and 36 driving score on Longest6 v2, while delivering 2.54x inference speedup and 8.9% success rate improvement over the backbone. BLUE provides a practical path toward efficient language‑augmented AD, showing that VLA models can retain the benefits of language at a fraction of the cost. Our code, data, logs and checkpoints are fully available on https://github.com/George‑Ling3/BLUE.
Authors:Danilo Danese, Angela Lombardi, Giuseppe Fasano, Matteo Attimonelli, Tommaso Di Noia
Abstract:
Large and demographically balanced datasets are essential for reliable neuroimaging biomarkers. Full‑resolution 3D brain MRI synthesis can support data augmentation in this setting, but existing approaches either incur prohibitive computational cost at volumetric scale or rely on lossy latent compression that may compromise anatomical detail. As a result, practical 3D generative augmentation often requires specialized compute infrastructure. We propose WaveDiT, a conditional flow matching framework operating in the coefficient space of a 3D Haar Discrete Wavelet Transform. The model combines factorized spatio‑depth attention with band‑wise heteroscedastic uncertainty modeling derived from higher‑order wavelet statistics. Predicted log‑variance is integrated directly into both the flow objective and conditioning pathway, enabling adaptive precision consistent with the heavy‑tailed and input‑dependent variance structure of anatomical detail. This formulation supports full‑resolution 3D synthesis under practical memory and time constraints on a single modern GPU. Evaluation on a multi‑site cohort demonstrates improved alignment between generated and real MRI distributions, together with enhanced downstream brain age prediction and region‑level anatomical agreement relative to diffusion, latent, and wavelet‑based baselines. Code is available at https://github.com/sisinflab/WaveDiT
Authors:Chenhan Jin, Shengze Xu, Qingsong Wang, Fan Jia, Dingshuo Chen, Tieyong Zeng
Abstract:
Data pruning (DP), as an oft‑stated strategy to alleviate heavy training burdens, reduces the volume of training samples according to a well‑defined pruning method while striving for near‑lossless performance. However, existing approaches, which commonly select highly informative samples, can lead to biased gradient estimation compared to full‑dataset training. Furthermore, the analysis of this bias and its impact on final performance remains ambiguous. To address these challenges, we propose OrderDP, a plug‑and‑play framework that aims to obtain stable, unbiased, and near‑lossless training acceleration with theoretical guarantees. Specifically, OrderDP first randomly selects a subset and then chooses the top‑q samples, where unbiasedness is established with respect to a surrogate loss. This ensures that OrderDP conducts unbiased training in terms of the surrogate objective. We further establish convergence and generalization analyses, elucidating how OrderDP affects optimal performance and enables well‑controlled acceleration while ensuring guaranteed final performance. Empirically, we evaluate OrderDP against comprehensive baselines on CIFAR‑10, CIFAR‑100, and ImageNet‑1K, demonstrating competitive accuracy, stable convergence, and exact control ‑‑ all with a simpler design and faster runtime, while reducing training cost by over 40%. Delivering both strong performance and computational efficiency, our method serves as a robust and easily adaptable tool for data‑efficient learning. The code is publicly available at https://github.com/shengze‑xu/OrderDP.
Authors:Changliang Xia, Chengyou Jia, Minnan Luo, Zhuohang Dang, Xin Shen, Bowen Ping
Abstract:
Although video virtual try‑on (VVT) has achieved significant progress, existing methods still exhibit two fundamental limitations: first, they are restricted to single‑garment transfer, rendering simultaneous multi‑object try‑on highly impractical; second, their heavy reliance on explicit external priors (e.g., garment masks) inevitably destroys crucial physical dynamics and degrades visual quality. To bridge this gap, this paper proposes the novel Try‑On Anything task, which aims to simultaneously transfer diverse wearable objects onto a person in a video in a single inference pass. To support and standardize this paradigm, we introduce TryAny‑Bench, a comprehensive benchmark encompassing a paired video dataset alongside a tailored evaluation protocol. Furthermore, we present OmniTryOn, an external‑prior‑free generative framework designed to tackle this task. Specifically, OmniTryOn employs a First Frame Wearable Cache strategy, which directly provides diverse wearable objects for the generation process through the initial video frame. To maintain consistency, we propose the Spatiotemporally Consistent RoPE (STC‑RoPE), which inherently establishes robust spatiotemporal anchors to strictly preserve complex human motions and background dynamics. Optimized by the proposed Gradual Try‑On (GTO) training strategy, our model progressively masters robust multi‑object synthesis. Extensive experiments on TryAny‑Bench demonstrate that OmniTryOn significantly outperforms existing specialized video virtual try‑on models and general video editing baselines, establishing a powerful new standard for the Try‑On Anything task. Our dataset, code, and models are available at https://github.com/xcltql666/OminTryOn.
Authors:Lianyu Hu, Xiaoyu Ma, Zeqin Liao, Yang Liu
Abstract:
Chain‑of‑thought (CoT) reasoning has proven effective for enhancing problem‑solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approaches suffer from a fundamental limitation: they perform reasoning entirely in text without accessing visual features during the reasoning process. After initial visual encoding, image information becomes inaccessible, forcing models to reason based solely on whatever was captured in the initial description, which forms a `vision‑blind reasoning' paradigm that limits fine‑grained visual extraction, error verification, and adaptive attention. We propose Text‑Visual Interleaved Chain‑of‑Thought (TVI‑CoT), a framework that enables explicit interleaving of textual reasoning and visual feature access through learnable control tokens <THINK>, <LOOK> and <ANSWER>. These tokens allow dynamic switching between reasoning and visual grounding, attending to relevant image regions conditioned on the evolving reasoning state. Experiments on eight benchmarks demonstrate state‑of‑the‑art results among MLLM‑based CoT methods and notable performance boost compared to the baseline: +6.1% on MMMU, +3.8% on MathVerse, +3.4% on MathVista, and +3.4% on ScienceQA. Code is available at https://github.com/hulianyuyy/TVI‑CoT.
Authors:Jamal Seyedmohammadi, Pai Chet Ng, Angelo Genovese, Zhixiang Chi, Jeannie Lee, Konstantinos N. Plataniotis
Abstract:
Palmprint modality offers a privacy‑preserving biometric solution, yet its deployment is hindered by the domain gap between controlled enrollment and unconstrained authentication. Existing datasets are largely restricted to controlled setups and fail to capture the compound variability of real‑world environments. In this paper, we introduce X‑Palm, a cross‑domain dataset comprising 6,006 palm images from 103 individuals (206 hands). To the best of our knowledge, X‑Palm is the first palmprint dataset providing novel paired‑identity acquisition specifically designed to bridge the gap between reliably controlled multispectral enrollment and unconstrained mobile authentication while encompassing a broad spectrum of in‑the‑wild variability. Unlike existing datasets that focus on single to a few variations, X‑Palm addresses the massive modality and environmental shifts encountered in practical deployments by capturing paired data for identities across two distinct domains: (1) a controlled Multispectral Palmprint setting using our custom‑developed scanner, and (2) an unconstrained smartphone palmprint setting that is participant‑driven, incorporating simultaneous variations in hardware, hand pose, illumination, background, camera‑to‑hand distance, perspective, and palm surface conditions (e.g., moisture and occlusions). Our extensive benchmarks of 12 SOTA models reveal that while existing methods achieve high performance on controlled data, they experience severe performance collapse on X‑Palm. Conversely, models trained on X‑Palm demonstrate consistent robustness across domains, positioning X‑Palm as a valuable resource for training a model towards real‑world, cross‑domain generalization. Data access instructions and the related benchmarking codes are publicly available at: https://github.com/X‑Palm/X‑Palm‑2026
Authors:Wenwei Huang, Jia Wei, Jianlong Zhou
Abstract:
Multi‑contrast brain MRI provide complementary soft‑tissue characteristics that aid in the screening and diagnosis of diseases. However, limited scanning time, image corruption and various imaging protocols often result in incomplete multi‑contrast images. While current approaches excel in image synthesis, they often struggle to synthesize critical tumor regions and exploit contextual information in multi‑contrast brain MRI effectively. To address this issue, we propose a synthesis‑centric, segmentation‑assisted closed‑loop framework with retrieval augmentation synthesis. Our method overall takes a generative adversarial architecture, which aims to synthesize missing contrasts from any combination of available ones with a single model. To explicitly capture tumor semantics and focus synthesis on tumor regions, we add an auxiliary segmentation branch that predicts tumor masks and feeds them back as semantic conditioning in synthesis branch, thereby learning tumor‑aware representations in the model and improving synthesis fidelity. Furthermore, we propose a dual‑bank retrieval augmentation strategy. It dynamically queries two external knowledge bases, namely a tumor masks memory bank for crucial tumor context and cross‑image contrast feature memory bank for global style information, to augment synthesis. Verified on two public multi‑contrast magnetic resonance brain datasets: BraTs2020 and UCSF‑BMSR, the proposed method is effective in handling medical brain images synthesis tasks and shows superior performance compared to previous methods. Code is available at:https://github.com/iBizzard/SSCF.git
Authors:Yufei Wei, Shuhao Ye, Chenxiao Hu, Yiyuan Pan, Dongyu Feng, Rong Xiong, Yue Wang, Yanmei Jiao
Abstract:
Recovering the relative 6‑DoF pose between two image groups underlies cross‑sequence relocalization and multi‑camera rig odometry. Each group carries known intra‑group geometry from visual odometry or rig calibration, and pretrained multi‑view backbones already fuse such geometry into visual features. Yet current models treat all views as an unstructured set, leaving cross‑group reasoning as the missing piece. We introduce \ours, which keeps the foundation model entirely frozen and adds three lightweight trainable modules to bridge the two groups: a perceiver resampler, a cross‑group bridge with merged self‑attention, and a multi‑frame pose head. The trainable footprint totals about 32M parameters, under 6% of the full model, and is supervised only by relative poses. Across four datasets that span indoor and outdoor simulation, real‑world cross‑season capture, and zero‑shot sim‑to‑real transfer, \ours attains state‑of‑the‑art accuracy on both tasks, while every baseline is retrained with its full original supervision. Code is available at https://github.com/WeiYuFei0217/G2G.
Authors:Qi Liu, Gang Yue, Mingyu Yin, Lisai Zhang, Yidi Wu, Yaole Wang, Yaohui Wang, Chang Yao, Jingyuan Chen, Lin Ma
Abstract:
Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task‑specific models. Building a unified framework that supports diverse video tasks remains an open challenge: existing unified attempts either require dedicated auxiliary encoders or lack explicit mechanisms to distinguish heterogeneous conditioning tokens, struggling when the number and type of visual conditions vary across tasks. We propose TIDE, a unified framework that integrates instruction‑based editing, reference‑guided editing, and multi‑reference generation. At its core, we introduce per‑token task embeddings that assign each input token a task‑specific identifier, enabling the model to explicitly disambiguate target, source, and reference tokens. To simultaneously capture high‑level semantic understanding and fine‑grained structural fidelity, we design a dual‑path conditioning scheme that couples a vision‑language model with a VAE latent path for complementary signals. We further devise a multi‑task progressive training strategy that incrementally introduces tasks of increasing complexity, effectively harmonizing diverse objectives and enabling smooth generalization across heterogeneous task distributions. Extensive experiments on multiple video editing and generation benchmarks demonstrate that TIDE achieves state‑of‑the‑art performance across all evaluated tasks. Our project page is available at https://LittleWork123.github.io/tide.
Authors:Jiangshuan Pang, Wangyang Tang, Jing Yan, Zhixuan Cheng, Youzhe He, Zhenkun Zhuang, Tao Zhou, Shiping Liu
Abstract:
MRI preprocessing defines the input distribution seen by brain MRI foundation models, yet it is usually treated as routine data cleaning rather than a modeling choice. We ask how much preprocessing is worth its computational cost for self‑supervised 3D MRI pretraining. Keeping the corpus, 3D ViT backbone, masking protocol, and downstream evaluations fixed, we compare a graded P0‑P7 preprocessing spectrum for masked autoencoding (MAE) and joint‑embedding predictive learning (JEPA) on 20,000 heterogeneous brain MRI volumes, then transfer the encoders to IDH prediction, MCI classification, brain age regression, and GLI/PED tumor segmentation. The results do not support a simple "more is better" rule. P0/P1 are numerically unstable, making P2 the lowest‑cost feasible level; beyond P2, choosing the best feasible preprocessing level improves aggregate utility by only 3.4 percentage points for MAE and 1.8 percentage points for JEPA, with most paired gains statistically unresolved. Stronger preprocessing is beneficial only in selected regimes: IDH improves modestly, AGE and GLI/PED are often near or best at P2, and MCI shows the clearest empirical P7 gain. Cross‑level MCI transfer further shows that much of the P7 advantage can be recovered by applying stronger preprocessing downstream, without requiring P7 throughout pretraining. These findings recast MRI preprocessing as a downstream‑aware cost‑utility decision rather than a default escalation pipeline. Code is available at https://github.com/PangJiangShuan/PreBrain.
Authors:Bingxuan Dai, Hongsong Wang, Jie Gui
Abstract:
Designing 3D metamaterial microstructures that meet the intended functions remains a major challenge, as it typically requires domain expertise, iterative simulations, and extensive manual tuning. Existing work on inverse design that automatically generates microstructures based on desired target properties often suffers from limited design diversity and faces challenges in ensuring the physical feasibility of the generated structures. To address this issue, a property‑informed diffusion‑based network is proposed that enables the generation of 3D microstructures directly from textual descriptions. Unlike traditional property conditioning methods, our approach leverages rich guidance in terms of semantics and physical properties in the text input to support diverse structure synthesis. To enforce consistency between the generated structures and the target textual prompts, a dual alignment strategy is adopted, including contrastive text‑structure alignment and test‑time reward‑guided alignment. Experimental results show that the model is capable of generating semantically meaningful and physically plausible structures across a wide range of material categories. Our approach has good potential for interactive microstructure design and opens up new directions for combining language‑based interfaces with inverse material discovery. Code is available at: https://github.com/hongsong‑wang/PropDiff‑TMG
Authors:Ruben Dario Florez-Zela
Abstract:
Vision‑based driver monitoring systems are increasingly deployed in safety‑critical intelligent transportation settings, yet they are almost always compared on classification accuracy alone. This paper argues that accuracy is insufficient to characterize a model's fitness for real‑world deployment, and proposes the Human‑Centered Benchmarking Framework (HCBF), which evaluates models across four dimensions: accuracy, explainability, efficiency, and robustness. The framework is applied to four representative lightweight architectures, MobileNetV3, ShuffleNetV2, EfficientNet‑B0, and DeiT‑Tiny, on the MRL Eye Dataset for eye‑state classification. While the models are nearly indistinguishable on clean‑set accuracy, each leads in exactly one dimension, and all four lie on the Pareto frontier. A Human‑Centered Score computed under three deployment‑oriented weighting scenarios ranks ShuffleNetV2 first throughout. However, this aggregate winner retains less than half of its performance under sensor noise and fails by classifying closed eyes as open, whereas the transformer remains robust. These findings show that aggregate ranking can mask dimension‑specific vulnerabilities that are operationally decisive, underscoring the value of multi‑dimensional, human‑centered evaluation.
Authors:Xiaoqian Wu, Yejie Guo, Xiaoyang Chen, Lixin Yang, Cewu Lu, Yong-Lu Li
Abstract:
We are surrounded by various objects with movable, articulated parts, e.g., box, handle, door. An accurate and generalizable perception of articulated parts is essential to enhance robotic manipulation capabilities. Building on this need, recent efforts in articulated parts perception have followed two main directions: One line of work uses pose‑based representation, which requires high manual cost; in parallel, affordance‑based methods extract future object motion from point tracking without additional manual efforts, but suffer from low‑quality data. In this paper, we propose a new representation of articulated parts, Geometric Primary Structure (GPS), an abstraction of the part geometry structure to balance scalability and quality. For efficient and scalable data collection, GPS is integrated with a portable Virtual Reality (VR) device and requires only one minute to annotate one object sequence. This direct human annotation provides higher quality than the estimated affordance. With this efficient VR‑GPS system, we collect 41K frames for 234 objects across six part classes, and train a generalizable GPS model with a single RGB‑D object image as input. For object manipulation, we deploy a heuristic policy based on GPS prediction. Without any in‑domain fine‑tuning, our method achieves an 73% success rate, covering 270 initial states for 9 objects. Our code, data and reusable tool are available at https://enlighten0707.github.io/gps.
Authors:Jianhui Wei, Jie Tan, Hengchuan Zhu, Xiaotian Zhang, Yan Zhang, Ziyi Chen, Daoan Zhang, Wei Xu, Zuozhu Liu
Abstract:
Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video generation, a long‑horizon multimodal task, remains underexplored. Unlike earlier video agents whose pipeline is handcrafted, these frameworks can build and refine their own workflows. We introduce VideoWeaver, an agent harness and benchmark that evaluates and evolves skills for long video generation, where an agent turns a single instruction into a long video by composing foundation skills into its own workflow rather than following a predefined pipeline. The benchmark has 16 task categories and 285 cases, with references spanning text, image, audio, video, and their combinations. Because errors can arise at any stage and not just in the final video, we propose an agent‑as‑judge that inspects both the execution trace and the final video, grounding its scores in evidence such as metadata and intermediate files. Using this feedback, we further design a skill evolution algorithm that refines and merges the agent's skills. Across multiple frameworks and models, we find that an explicit composition skill improves the generation process over using foundation skills alone, that skill evolution further improves output quality, and that performance varies notably across harness and model choices. The proposed agent‑as‑judge also aligns well with human judgments, especially on process metrics. Code and dataset is available at https://github.com/JianhuiWei7/VideoWeaver
Authors:Jiaqi Tang, Jianmin Chen, Youyang Zhai, Wei Wei, Runtao Liu, Mengjie Zhao, Xiangyu Wu, Qingfa Xiao, Qifeng Chen
Abstract:
Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real‑world visual corruptions. While existing robustness enhancement approaches exist, they are limited: black‑box feature alignment lacks interpretability, and white‑box text‑based reasoning cannot restore lost pixel‑level details. This work investigates a fundamental research question: Can MLLMs recover corrupted visual content by themselves? To address this, we propose Robust‑U1, a novel framework that equips MLLMs with explicit visual self‑recovery capability for robust understanding. The approach comprises three core stages: supervised fine‑tuning for initial reconstruction, reinforcement learning with dual rewards (pixel‑level SSIM and semantic‑level CLIP similarity) for aligning high visual quality, and multimodal reasoning that jointly considers both the corrupted input and the recovered image. Extensive experiments demonstrate that Robust‑U1 achieves state‑of‑the‑art robustness on the real‑world corruption benchmark and maintains superior performance under adversarial corruptions on general VQA benchmarks. Analysis confirms that high‑quality visual recovery directly enhances reasoning performance, establishing self‑recovery as a critical mechanism for robust visual understanding. The source code is available at https://github.com/jqtangust/Robust‑U1.
Authors:Zichen Zhu, Yuheng Sun, Mingxuan Zhu, Wenjie Ma, Situo Zhang, Zhexiang Wang, Ziyue Yang, Danyang Zhang, Kunyao Lan, Zihan Zhao, Dingye Liu, Siqi Xiang, Lu Chen, Kai Yu
Abstract:
Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes. Creations by generative models may contain artifacts, implausible details, or stylistic drift away from photorealism and offer little insight into why an edit was made. We propose IEA, a conversational Image Editing Agent that learns to operate parameterized tools in an explicit, interpretable action space. IEA is trained via a three‑stage multitask pipeline: (1) SFT on distilled expert edits, (2) GRPO with rewards for likeness improvement, tool usefulness, and intent summarization, and (3) large‑scale synthetic fine‑tuning to jointly master image editing, refinement, and user intent summarization. By manipulating 16 editing tools step by step, IEA produces transparent edit traces that can be inspected and debugged. In quantitative experiments, it attains a lower pixel distance on the edit task and a higher ROUGE‑L on the summary task than strong baselines. In user studies, it ranks best among tool‑calling methods for instruction following while surpassing generative methods in overall perceptual quality. Our results validate interpretable, tool‑centric VLMs as a reliable path to human instruction‑guided image retouching.
Authors:Swarna Chakraborty, Gabriel De Castro Araújo, Syeda Tasmi Faria, Marcelo M. Carvalho, Mylene C. Q. Farias
Abstract:
Point Cloud Quality Assessment (PCQA) methods typically predict scalar Mean Opinion Scores (MOS), which quantify overall perceptual degradation but do not reveal its causes. In contrast, human observers naturally reason in terms of specific distortions such as blur, color shifts, point density changes, missing regions, and geometric deformations. To close this gap, we introduce DAL‑PCQA, a distortion‑aware, language‑annotated dataset for PCQA. DAL‑PCQA augments benchmark point clouds with multi‑level distortion severity labels, discrete quality categories, and structured natural language descriptions aligned with human perception. We define a point‑cloud‑specific distortion taxonomy that covers both photometric and geometric artifacts. Statistical analysis reveals characteristic degradation patterns across distortion types and quality levels. To assess the utility of these annotations, we compare zero‑shot and fine‑tuned multimodal models for generating perceptual quality descriptions. Experiments show that distortion‑aware supervision substantially improves lexical and semantic alignment with ground‑truth descriptions. By enabling interpretable, distortion‑level reasoning, DAL‑PCQA facilitates language‑driven, explainable point cloud quality assessment. The dataset is publicly available at https://github.com/swarna96/DAL‑PCQA.
Authors:Siyang Song, Micol Spitale, Zijian Wu, Xiangyu Kong, Cheng Luo, Cristina Palmero, German Barquero, Sergio Escalera, Michel Valstar, Mohamed Daoudi, Fabien Ringeval, Andrew Howes, Elisabeth Andre, Hatice Gunes
Abstract:
In dyadic interactions, various human facial reactions could be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023, 2024 and 2025 challenge series, a body of generative deep learning (DL) models have been developed for the problem of multiple appropriate facial reaction generation (MAFRG). This year, we propose the REACT 2026 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can generate multiple personalised, appropriate, diverse, realistic and synchronised human‑style facial reactions expressed by a specific human listener for responding to each given speaker behaviour. As a key of the challenge, we continuously provide challenge participants with MARS dataset introduced by REACT 2025 but additionally provide individual‑level Big‑Five personality labels and EEG recordings. This introduces a new one‑to‑many personalised facial reaction generation setting combining human expressive behavioural, affective and neurophysiological signals, which remains largely unexplored in current dyadic interaction modelling. This paper also presents the challenge guidelines and new baselines on the four proposed sub‑challenges: Offline generic and personalised MAFRG as well as Online generic and personalised MAFRG, respectively, which are publicly available at https://github.com/reactmultimodalchallenge/baseline_react2026.
Authors:Didi Zhu, Changrui Chen, Stefanos Zafeiriou, Jiankang Deng
Abstract:
When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task‑critical visual evidence? Correct answers can coexist with flawed reasoning, making accuracy alone an incomplete test of grounding. We introduce VisualFLIP, a paired benchmark with 1,374 images arranged as same‑question perturbation pairs across cardinality, attribute, spatial, and logic tasks. Each pair keeps the question fixed but minimally changes the evidence so the gold answer deterministically flips. We evaluate 24 MLLMs with pair accuracy, which requires solving both sides of a pair, and Collapse Rate (CR), which measures how often a model that solves at least one side repeats the same non‑empty answer for both images. Together, these metrics show that paired correctness and evidence dependence are related but distinct: capable models can still fail to update after task‑critical visual changes, and collapse becomes more severe for some models when the edited image follows an earlier answer in a sequential setting. Further details are available on our project page: https://didizhu‑judy.github.io/VisualFLIP/
Authors:Maksim Shandybo, Ivan Bespalov, Daniil Yefimov, Marina Kosheleva, Alexander Loukianov
Abstract:
Parsing historical documents with complex, non‑standard layouts remains a fundamental bottleneck in large‑scale archival digitization. Unlike modern typography, historical newspapers exhibit severe physical degradation and highly irregular page structures that confound even state‑of‑the‑art vision‑language models, presenting severe out‑of‑distribution challenges. We address this gap with an automated pipeline specifically designed for parsing historical newspapers, documents characterized by particularly intricate multi‑column layouts. Our approach combines a fine‑tuned YOLO architecture for layout analysis and block detection, trained on 1,426 fully human‑annotated scanned pages, with a novel semantic assembly module that reconstructs articles by jointly modeling lexical‑semantic similarity via TF‑IDF, visual embeddings from our fine‑tuned YOLO, and geometric layout constraints. This multi‑modal integration yields state‑of‑the‑art performance, achieving an F1 score of 0.904 on block‑to‑article mapping. Notably, end‑to‑end evaluation against vision‑language models (Qwen3.6‑35B‑A3B and Qwen3.6‑Plus) demonstrates that PereStruct achieves substantially higher fidelity (BLEU approximately 0.96 vs 0.34), validating that modular architectures excel where generic VLMs fail on complex historical layouts. To support reproducibility and advance research in this domain, we release both the training corpus of 599 annotated pages and a curated PereStruct benchmark of 93 pages with expert‑verified ground‑truth block‑to‑article mappings. This framework establishes a robust foundation for high‑fidelity digitization and semantic reconstruction of complex archival materials.
Authors:Yaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu, Wenjie Du, Cheng Liang, Weijun Wang, Yuanchao Li, Guangyao Li, Hao Fei, Yuanchun Li, Henghui Ding, Yunxin Liu
Abstract:
Recent advances in Omni‑Multimodal Large Language Models (Omni‑MLLMs) have enabled strong integration of vision, audio, and language. However, their audio‑visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI‑Bench, a cognitively inspired benchmark that evaluates Omni‑MLLMs across three stages, perception, understanding, and reasoning, through cross‑modal tasks requiring joint audio‑visual interpretation. This design enables fine‑grained diagnosis of model capabilities and failure modes. To further assess robustness beyond familiar domains, we propose AVI‑Bench‑PriSe, an extension that probes models' primitive audio‑visual sensation using unfamiliar, low‑semantic stimuli, testing generalization beyond common training distributions. Extensive experiments on both open‑source and closed‑source models reveal substantial limitations in current Omni‑MLLMs. Based on these findings, we present a four‑level AVI taxonomy. Overall, AVI‑Bench provides a principled evaluation framework to guide the development of more robust and generalizable AVI. Project website: https://fudancvl.github.io/AVI‑Bench/
Authors:Lecheng Yan, Yichong Zhang, Ben Pan, Xiaoyu Zheng, Jiawei Qian, Anqi Wu, Wenxi Li, Chenyang Lyu
Abstract:
Editing a long‑form video from heterogeneous footage requires more than selecting clips: an agent must preserve narrative intent across material preparation, timeline construction, post‑production, and revision while leaving enough evidence to diagnose failures. We present Crayotter, an open‑source multimodal multi‑agent system for prompt‑driven video editing. Crayotter organizes production into three phases: coverage‑aware material preparation, artifact‑based editing research, and tool‑grounded timeline execution. Each phase externalizes inspectable artifacts, including coverage reports, multimodal analyses, editing blueprints, tool calls, and intermediate renders. These artifacts make an editing run traceable and allow failed segments to be diagnosed and selectively revised instead of requiring a full restart. We evaluate Crayotter on 23 editing themes against CapCut‑Mate and CutClaw. Under human evaluation, Crayotter achieves an average score of 3.40/5, compared with 2.44 and 1.70 for the two baselines, with consistent gains in theme alignment, narrative coherence, and editing smoothness. We additionally describe a replayable trajectory schema and verifiable reward design that prepare these workflows for future policy optimization. Code, traces, and examples are publicly available at https://github.com/idwts/Crayotter.
Authors:Guangzhi Sun, Yixuan Li, Yudong Yang, Chao Zhang
Abstract:
Audio‑visual large language models (LLMs) hold strong promise for long‑form video understanding, yet their long‑video inference is fundamentally limited by the linear growth of video tokens and key‑value (KV) caches. We present OmniMem, a memory‑efficient streaming framework designed specifically for audio‑visual LLMs. Unlike existing compression methods that treat all tokens uniformly, OmniMem introduces a modality‑aware memory allocation strategy that separately manages visual and audio contexts, addressing the severe token imbalance between the two modalities. OmniMem further preserves informative and non‑redundant KV states through perturbation‑aware memory selection, enabling compact memory without sacrificing long‑range understanding. To strengthen compression under realistic deployment constraints, we also explore budget‑aware fine‑tuning, which encourages the model to consolidate useful information into retained memory. Experiments on VideoMME Long, LVBench, and LVOmniBench with video‑SALMONN 2+ and Qwen‑2.5‑Omni show that OmniMem consistently improves over strong training‑free compression baselines by 2‑4% absolute accuracy under the same memory budgets, with an additional 1‑2% gain after fine‑tuning.
Authors:Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou
Abstract:
Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning. However, it remains unclear whether they can approximate subjective human responses, which depend not only on content comprehension but also on individuals' social contexts. To address this gap, we evaluate MLLMs as synthetic participants in an emerging task: assessing perceived sensory engagement with short videos. Grounded in the Perceived Message Sensation Value (PMSV) framework, we compare ratings from recruited human participants and profile‑conditioned MLLM simulations (n=673) using a 17‑item scale measuring emotional arousal, dramatic impact, and novelty. We find that even leading MLLMs (Gemini 3 Flash and Qwen 3 Omni) show limited agreement with human participants. The models exhibit distinct downward mean‑shift and central‑tendency biases in their rating distributions. They both introduce and flatten subgroup differences, while showing inconsistent sensitivity to participant profiles. Prompting strategies affect these metrics differently, modestly improving some aspects while worsening others. These results highlight both the challenges and opportunities of developing MLLMs as synthetic participants in video‑based research. Data and code: https://github.com/MINDLab25/mllm‑human‑simulation‑eval
Authors:Shengli Zhou, Xiangchen Wang, Guanhua Chen, Feng Zheng
Abstract:
Large language models (LLMs) have recently been applied to 3D vision‑language (3D‑VL) tasks, which require spatial reasoning to identify target objects relative to anchors. Scene graphs are commonly employed to represent such relations, but reasoning over complete graphs incurs high token costs and computational inefficiencies, motivating the need for pruning. Existing pruning methods primarily rely on spatial proximity and often remove task‑relevant relations, thereby undermining reliable spatial reasoning. To address these limitations, we derive a key requirement for scene graph pruning: preserving spatial relations that are most pertinent to the specific 3D‑VL task. Guided by this insight, we propose the Conceptual‑Adjacent Scene Graph Pruner (CAPruner). CAPruner integrates fuzzy semantic relevance with spatial proximity to estimate the importance of relations, enabling the selection of critical relations in a task‑specific context. Moreover, to avoid costly relation‑level annotations, CAPruner is trained by supervising the aggregated scores of each node's incident edges. Extensive experiments demonstrate that CAPruner effectively preserves relations essential for spatial reasoning, leading to substantial performance improvements of LLMs on 3D‑VL tasks. Code is available at https://github.com/fz‑zsl/CAPruner.
Authors:Meixi Song, Dizhe Zhang, Hao Ren, Ruiyang Zhang, Bo Du, Ming-Hsuan Yang, Lu Qi
Abstract:
In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of camera systems, from conventional perspective cameras to wide‑field‑of‑view, fisheye and omnidirectional panoramic settings. To overcome the pinhole‑specific assumptions of SHARP, our key idea is to align various images in a unified omnidirectional latent space. Thus, we propose UniSHARP, which performs implicit alignment in both feature and Gaussian spaces. Specifically, Gaussian primitives are arranged along rays and radial distances in a ray‑based universal representation, while 2D semantic and 3D spatial features extracted from UniK3D‑inspired encoders are jointly decoded to generate the complete Gaussian cloud. To comprehensively evaluate our method, we construct a benchmark covering diverse imaging systems across various scenes. The benchmark is further stratified by field of view (FoV) to enable fine‑grained assessment of the universal monocular rendering task. Extensive experiments on the proposed benchmark demonstrate the effectiveness of UniSHARP, outperforming alternative methods by a large margin. The project page can be found at: https://insta360‑research‑team.github.io/Unisharp‑website/
Authors:Hanhui Wang, Yiming Xie, Haiwen Feng, Zhaoyang Lv, Shenlong Wang, Huaizu Jiang
Abstract:
We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior video models that train separate models for different force types, assume fixed forces, or rely on non‑causal processing, StreamForce is a causal and unified model that responds instantly and coherently to both local and global, time‑varying forces. To achieve this, we design a unified force representation as a control signal and develop a distillation pipeline for force‑controllable video generation. Our model combines autoregressive efficiency with force responsiveness, sustaining stable photometric and dynamic realism. StreamForce runs at up to 16.6 FPS on a single GPU, achieving state‑of‑the‑art performance in both force adherence and motion realism. Project website: https://neu‑vi.github.io/StreamForce/
Authors:Johannes Theodoridis, Johannes Maucher, Andreas Schilling
Abstract:
We propose Differences in Detection (DnD), an intuitive method to compare two object detection models. Based on the same matching algorithm, it complements the standard metrics of mean Average Precision (mAP) and TIDE error analysis with the ability to compare two models directly. More specifically, we calculate the intersection of ground truth labels that are recognized by both models, followed by the corresponding difference sets and the complement set of ground truth labels that are missed by both models. The resulting comparison is more direct and intuitive than a comparison of independent summary statistics. It reveals individual and shared mistakes and becomes particularly interesting when combined with error types. In this case, the differences in detection errors can be analyzed naturally in a standard confusion matrix. While valuable in itself, we believe that one of the best applications of DnD is to guide explainability methods such as ODAM towards metric‑relevant examples, grounded in structured subsets. The code for our method is available here: https://github.com/JohannesTheo/differences‑in‑detection
Authors:Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
Abstract:
Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge‑intensive video scenarios. These scenarios require models to handle sparse evidence, long‑range dependencies, multimodal alignment, and reliable inference under limited computational budgets. This work presents a human‑view perspective on LLM‑based video understanding, organized around three functional abilities: watching, remembering, and reasoning. Rather than treating video tasks as isolated benchmarks, this view provides a unified structure for analyzing how video MLLMs acquire evidence, preserve context, and produce grounded outputs. We introduce a formulation that characterizes video understanding systems by their perceptual representations, memory states, reasoning traces, and final predictions. Based on this formulation, we identify challenges in spatio‑temporal perception, efficient long‑video processing, memory modeling, streaming understanding, and faithful reasoning. Representative methods are organized by their roles in video MLLM systems. Watching covers fine‑grained, comprehensive, audio‑visual, and efficient perception. Remembering includes offline and streaming memory, while reasoning covers text‑only reasoning and thinking with videos. We further examine application domains such as egocentric, sports, instructional, medical, and narrative videos, and cover training datasets and evaluation benchmarks across task types, supervision formats, modalities, and capability dimensions. Finally, we outline open problems and future directions for scalable, memory‑aware, and evidence‑grounded video intelligence. Related works will be continuously traced at https://github.com/marinero4972/Awesome‑HumanView‑VideoUnderstanding.
Authors:Yuang Shi, Simone Gasparini, Géraldine Morin, Wei Tsang Ooi
Abstract:
Streaming 3D Gaussian Splatting requires highly scalable, progressive representations. Existing progressive methods rely on discrete layering, accumulating separate splat sets for each level of detail. This structural independence between layers inherently leads to error accumulation, severe splat redundancy, and uncontrolled quality transitions. We propose EvoGS, the first continuous‑layering representation. Organized as an Evolution Tree, EvoGS generates finer details via an explicit, wavelet‑inspired parent‑child refinement. This empowers child nodes to structurally correct ancestral errors, yield inherently sparse and highly compressible inter‑layer signals. Extensive experiments show EvoGS eliminates splat redundancy from over 65% to under 25%. Compared to state‑of‑the‑art baselines, it reduces transmission payload and GPU VRAM footprint by up to 2.4× and 5.5×, respectively, and achieves smooth quality transitions optimal for real‑time adaptive streaming. Project page: https://yuang‑ian.github.io/evogs/
Authors:Duc Tri Tran, Trung Thanh Nguyen, Vijay John, Phi Le Nguyen, Yasutomo Kawanishi
Abstract:
Video Text Spotting (VTS) is essential for urban surveillance and intelligent transportation systems, enabling automated reading of street signs, vehicle markings, and scene text in video streams. However, reliable recognition remains challenging due to dynamic video factors common in surveillance scenarios, including motion blur, occlusion, and scale variation, which degrade frame‑level recognition. Existing VTS methods typically perform recognition independently on each frame, leading to inconsistent and inaccurate results across sequences. To address these limitations, we propose TraRA (Trajectory‑level Recognition Aggregation for VTS), a plug‑and‑play method that performs trajectory‑level text recognition by leveraging temporal and multimodal consistency. TraRA integrates two key modules: (1) the Temporal Clustering and (2) the Vision‑Language Aggregation. The former refines noisy trajectories by grouping temporally and visually coherent text instances, while the latter employs a Low‑Rank Adaptation‑enhanced Vision‑Language model to fuse visual cues with linguistic context across frames. By aggregating information over entire text trajectories, TraRA achieves robust text recognition even under challenging surveillance conditions. Extensive experiments on four public benchmarks, including road and urban scene datasets (RoadText, BOVText, ArTVideo, and ICDAR15), demonstrate that TraRA consistently improves tracking and recognition performance over state‑of‑the‑art VTS methods. The source code is available at https://github.com/trid2912/TraRA.
Authors:Okan Umur, Ali Emre Güşlü, Ibrahim Delibasoglu
Abstract:
The rapid advancement of video editing and generative artificial intelligence technologies has made realistic video manipulation increasingly accessible. Although existing datasets have significantly advanced research in deepfake detection, object removal, and video inpainting, they do not adequately model scenarios in which a short manipulated segment is inserted into an otherwise authentic video and the original video continues afterward. In this study, we review representative datasets from the literature, analyze their characteristics, and discuss their limitations with respect to temporally localized realistic manipulation detection. Based on this analysis, we motivate the need for a new dataset specifically designed for authentic videos containing short and highly realistic manipulated intervals. Finally, we evaluate two complementary approaches on our custom‑curated test set to establish an initial benchmark for this challenging scenario. The first employs a linear probe on DINOv3 features, assessed under three thresholding strategies. The second leverages DINOv3 features with a consecutive frame similarity‑based method to detect temporal manipulation boundaries. Together, these experiments provide an initial benchmark for partially manipulated video detection and highlight the need for content‑adaptive thresholding mechanisms. The dataset, code, and supplementary materials are publicly available at https://github.com/OkanUmur/temporally‑localized‑video‑manipulation‑detection.
Authors:Jilles S. van Hulst, Jakub M. Tomczak, W. P. M. H. Heemels, Duarte J. Antunes
Abstract:
Variational autoencoders (VAEs) learn low‑dimensional latent representations of high‑dimensional data. When the data lies on a manifold with non‑Euclidean topology, the standard Gaussian prior introduces a topological mismatch that degrades reconstruction quality and prevents faithful representation. We present a constructive mathematical framework that resolves this mismatch for all manifolds that admit a product covering space. These are manifolds expressible as products of elementary factors (circles, intervals, or lines) or as quotients of such products by a finite symmetry group. The class includes cylinders, tori, Möbius strips, Klein bottles, and real projective spaces. Factorized distributions over the elementary factors yield product topologies with closed‑form, decoupled KL divergences, so that each latent factor can be shaped independently while keeping training tractable. We catalogue reparametrizable encoder‑prior pairs for periodic, bounded, and unbounded supports, and provide coordinate transformations that allow standard neural networks to output non‑Euclidean parameters with smooth gradients. For quotient manifolds, the decoder receives group‑invariant features of the covering‑space coordinates, so that identified points produce identical outputs. Anchor constraints fix the coordinate system relative to the data or create soft topological holes. Experiments on synthetic manifolds and real‑image datasets (rotated and cyclically shifted MNIST) confirm that a topology‑matched prior aligns KL regularization with the data manifold. The resulting topology‑aware models outperform the Gaussian baseline at all practically relevant regularization strengths. The code is available at https://github.com/JvHulst/VAE‑Topology.
Authors:Zhenyu Yang, Zemin Du, Shengsheng Qian, Changsheng Xu
Abstract:
Zero‑Shot Composed Image Retrieval (ZS‑CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption without training samples. Existing ZS‑CIR datasets often suffer from complete irrelevance between reference and target images due to noisy image sources, and do not achieve a true zero‑shot scenario as they use public image datasets that models like CLIP have been trained on. To tackle these challenges, we introduce ZeroSight, a novel benchmark for ZS‑CIR. It includes a dataset with consistent reference‑target pairs sourced from videos, a data construction pipeline, and evaluation methods that consider the ranking of multiple positive and negative target images. We ensure visually and semantically consistent reference‑target pairs by extracting frames from a single video and generating relative captions using LLM‑assisted methods. To ensure a true zero‑shot scenario, we use video data published after March 31, 2022, ensuring it was not included in CLIP's pre‑training data. Additionally, we propose a training‑free MLLM‑driven method, SC4CIR (Symmetric Consistency for CIR), which can effectively identify hard negative targets through 3 symmetric consistency checks. This method is plug‑and‑play, seamlessly integrating with various CIR methods and significantly improving performance. Our experimental results from 27 methods reveal that current ZS‑CIR datasets and evaluation metrics result in inflated retrieval performance, exaggerating the capabilities of CIR methods. Our benchmark and models can be accessed at https://github.com/sotayang/ZeroSight.
Authors:Minseong Kim, Jinyeong Park, Sungho Park, Jibum Kim
Abstract:
Multi‑modal approaches used for 3D CAD generation require substantial computational resources, necessitating efficient training. To address this, we propose GuideCAD, which leverages semantically rich visual‑textual representations having only a small number of trainable parameters to generate 3D CAD models. Specifically, GuideCAD uses a mapping network that converts image embeddings into prefix embeddings, enabling a pretrained large language model to integrate visual and textual information. As a result, a transformer‑based decoder predicts the construction sequence using the visual‑textual embeddings in order to generate the 3D CAD model. For experimental evaluation, we construct a new dataset, referred to as GuideCAD, which consists of text‑image pairs. Each pair includes a text prompt that represents a 3D CAD construction sequence and its corresponding 3D CAD image. Our experimental results show that GuideCAD generates comparably high‑quality 3D CAD models while using approximately four times fewer parameters and achieving twice the training efficiency compared to fine‑tuning approaches. We have released the source code and dataset for our method at: https://github.com/mskimS2/GuideCAD
Authors:Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
Abstract:
Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution. We present DaX, a pathology vision foundation model that adapts DINOv3‑style self‑supervised learning to whole‑slide histopathology. DaX is initialized from natural‑image DINOv3 weights and incorporates continuous magnification training, cross‑scale tissue views, orientation‑agnostic and acquisition‑robust augmentation, multi‑input‑size training, and Gram‑anchored dense consistency. These designs aim to connect local cellular morphology with global tissue architecture while stabilizing dense token‑level representations across input scales. We further construct a WSI‑level benchmark comprising 161 clinically meaningful tasks from 44 public datasets, covering 28,182 patients and 34,394 slides across four clinical domains and nine task categories. All models are evaluated under a fixed patient‑level cross‑validation protocol with fold‑level statistical ranking, enabling reproducible comparisons that are less sensitive to split‑dependent variation. Across this benchmark, DaX achieves the highest mean performance across tasks and consistently strong task‑level ranking scores, with gains spanning diagnostic pathology, biomarker and molecular profiling, tissue/specimen context, and risk, response, and prognosis. These results support DaX as a transferable visual encoder for computational pathology and provide a standardized evaluation framework for future pathology foundation models. Project page: https://alibaba‑damo‑academy.github.io/DaX/benchboard/.
Authors:Sunoh Kim, Daeho Um
Abstract:
Vision‑language models (VLMs) such as CLIP achieve strong zero‑shot recognition but remain highly fragile under adversarial perturbations. Recent test‑time adaptation defenses improve robustness by leveraging many augmented views, but this leads to impractical slowdown and a clear robustness‑throughput trade‑off. To address this challenge, we present Stability and Suitability‑guided Test‑time Prompt Tuning (SS‑TPT), evaluating the quality of each augmented view via two complementary scores: (1) stability, measuring prediction invariance to weak augmentations, and (2) suitability, measuring feature‑space density among views. These stability and suitability (SS) scores guide both adaptation and inference through an SS‑guided consistency loss and an SS‑weighted prediction, amplifying trustworthy views while suppressing corrupted ones. Extensive experiments demonstrate that SS‑TPT significantly outperforms prior state‑of‑the‑art methods, achieving superior robustness‑throughput trade‑offs across diverse datasets and varying numbers of views, thereby demonstrating both strong practicality and generality. Our code is available at https://github.com/sunoh‑kim/SS‑TPT.
Authors:Sunoh Kim, Daeho Um
Abstract:
Vision‑language models such as CLIP have achieved remarkable zero‑shot recognition capabilities, yet their robustness against adversarial perturbations remains limited. Test‑time counterattack (TTC) was recently proposed to improve CLIP's robustness by perturbing an input image to steer it away from a corrupted state during inference. However, TTC remains fragile under strong attacks because its counterattack relies on a directly corrupted original view and employs a noise‑driven hard‑gating scheme that cannot adapt to varying corruption severity. To address these limitations, we introduce Multi‑view guided Adaptive Counterattack (MAC), which performs counterattacks for multi‑view with corruption‑aware soft weighting. Specifically, MAC first constructs augmented views of an input image to obtain diverse embeddings. It then performs counterattacks to refine corrupted embeddings of views. Next, MAC adaptively scales the counterattack intensity for each view based on its estimated corruption degree. Finally, the adaptively counterattacked views are aggregated to yield a robust final prediction. Extensive experiments across 20 datasets and diverse attack scenarios demonstrate that MAC substantially improves robustness while preserving high inference speed and memory efficiency with its tuning‑free design. Our code is available at https://github.com/sunoh‑kim/MAC.
Authors:Donggyu Lee, Youngbin Ki, Jeonghun Kang, Taehwan Kim
Abstract:
While highlight detection for long‑form videos is of great practical importance, most existing methods remain limited to short‑form content, largely due to the absence of a suitable benchmark. To bridge this gap, we introduce SVHighlights, to the best of our knowledge, the first benchmark for highlight detection in extremely long sports videos, each exceeding one hour in duration, across multiple sports categories. SVHighlights is constructed from pairs of full‑length sports videos and their corresponding official highlight videos using a dataset generation pipeline, enabling scalable label generation without conventional per‑clip saliency annotation. The benchmark comprises 320 videos with an average duration of 2.00 hours and a total of 640.18 hours, substantially exceeding previous datasets. Existing methods also face fundamental challenges on long videos: models trained on short clips fail to generalize to hour‑long content, and their clip‑level scoring lacks the broader context needed to identify highlights. To address this and provide a strong baseline, we present TF‑SELECTOR, a training‑free segment‑based approach that divides each video into context‑aware segments by merging adjacent shots sharing the same semantic content, and predicts segment‑level saliency scores using a large language model with multimodal inputs including visual captions, transcripts, and audio volume. Experiments demonstrate that TF‑SELECTOR achieves superior performance across most metrics compared to Video Temporal Grounding (VTG)‑tuned baselines, with improvements of +3.12 in HIT@1, +4.06 in HIT@K, and +2.95 in IoU. These results establish SVHighlights as a challenging testbed for long‑form highlight detection and demonstrate that a simple segment‑based strategy can effectively scale to hour‑long videos.
Authors:Wenhao Zhang, Ramin Ramezani, Tao Han, Kai Hwang, Minyi Guo
Abstract:
Modern image‑analysis pipelines often convert images into structured semantic variables, such as facial attributes, object concepts, and scene descriptors. Learning directed dependencies among these variables can produce interpretable visual semantic graphs, but continuous directed acyclic graph learning is limited by the cost of enforcing acyclicity. We present polyDAG, a polynomial acyclicity framework for efficient continuous causal discovery in visual semantic graphs. polyDAG replaces the matrix‑exponential acyclicity constraint with a finite polynomial trace constraint and proves that the new constraint is zero exactly for acyclic graphs. We further derive a geometric‑series implementation that avoids the explicit summation loop while preserving the same acyclicity condition. Experiments on synthetic Erdos‑Renyi graphs and CelebA facial visual attributes show that polyDAG improves efficiency and structure recovery. Averaged over the revised synthetic protocol with d in 100, 200, 500, polyDAG reduces mean structural Hamming distance from 318.4 to 285.4 and improves mean F1 score from 0.725 to 0.756. At 100 nodes, the geometric variant runs in 3.44 seconds compared with 5.16 seconds for the exponential baseline, corresponding to a 33.4 percent speedup. Code and data are publicly available at https://github.com/wenhaoz‑fengcai/polyDAG.
Authors:Pei Yang, Hai Ci, Yanzhe Chen, Qi Lv, Han Cai, Mike Zheng Shou
Abstract:
Vision‑language‑action (VLA) models have advanced rapidly across backbones, training recipes, and data scale, yet the action decoder, which converts the backbone's hidden state into a continuous control signal, has barely changed and remains a single‑point predictor across the majority of current VLAs. Whether implemented via autoregressive token bins, L1 regression, or flow‑matching denoising, the resulting decoder treats the action space as unstructured, leaving the geometric proximity of neighboring actions unexploited during training. To advance this, we introduce ActionMap, a voxel heatmap action head that drops into an existing VLA in place of its native action decoder. For each new action, the head predicts a voxel heatmap over the action space, where each voxel directly stores the probability of the corresponding action. Across LIBERO simulation and real‑world Franka manipulation, our heatmap head surpasses two architecturally distinct backbones at matched training steps (e.g., +8.2% over OpenVLA‑OFT's L1 regression head on the LIBERO four‑suite average), converges at comparable or faster rates on both backbones, and remains markedly more data‑efficient at low training data. The cross‑backbone consistency indicates that action representation is a real lever for VLA performance, distinct from further backbone or recipe scaling. Project Page: https://showlab.github.io/ActionMap/.
Authors:Yuan Zeng, Yujia Shi, Yuhao Yang, Dongxia Liu, Zongqing Lu, Wenming Yang, Qingmin Liao
Abstract:
Human image animation aims to generate a video from a static reference image, guided by pose information extracted from a driving video. Existing approaches often rely on pose estimators to extract intermediate representations, but such signals are prone to errors under occlusion or complex poses. Building on these observations, we present DirectAnimator, a framework that bypasses pose extraction and directly learns from raw driving videos. We introduce a Driving Cue Triplet consisting of pose, face, and location cues that captures motion, expression, and alignment in a semantically rich yet stable form, and we fuse them through a CueFusion DiT block for reliable control during denoising. To make learning dependable when the driving and reference identities differ, we devise a Same2X training strategy that aligns cross‑ID features with those learned from same‑ID data, regularizing optimization and accelerating convergence. Extensive experiments demonstrate that DirectAnimator attains state‑of‑the‑art visual quality and identity preservation while remaining robust to occlusions and complex articulation, and it does so with fewer computational resources. Our project page is at https://directanimator.github.io/.
Authors:Yuan Zeng, Yujia Shi, Zongqing Lu, QingMin Liao
Abstract:
Human Image Animation has seen significant advancements, primarily driven by diffusion models. However, existing methods typically demand substantial training data and resources to achieve high‑quality results, limiting generalization and accessibility. In this work, we introduce \emphFreeAnimate, a training‑free framework that leverages the inherent capabilities of image diffusion models to enable temporal consistency, identity preservation, and background stability. Our approach incorporates a novel preview generation strategy that provides temporal and structural priors from generated preview frames, effectively guiding pose alignment and background consistency without training. Additionally, FreeAnimate introduces Inversion‑Boosted Attention and Reference‑Anchored Self‑Attention modules to guarantee temporal consistency and identity preservation. Experimental results demonstrate that FreeAnimate outperforms existing training‑free competitors and training‑based baseline methods, achieving generation quality comparable to state‑of‑the‑art methods and offering robust generalization across diverse datasets. Our project page is at https://freeani.github.io/.
Authors:Kangjian Zhu, Haobo Jiang, Jianjun Qian, Jin Xie
Abstract:
In this paper, we propose a cross‑view fusion framework that enhances the robustness of 6‑DoF grasp pose estimation in corner views. Our framework alleviates occlusion by incorporating an auxiliary view and avoids the time‑consuming, task‑agnostic multi‑view reconstruction through a post‑fusion strategy. To enhance cross‑view fusion, we propose a self‑supervised contrastive learning strategy that leverages cross‑view associations to regularize point cloud features. In brief, a cross‑view point pair is considered a match if the two points correspond to the same 3D location, and a non‑match if they represent distinct grasp directions. The learning strategy significantly enhances the spatial consistency and direction distinctiveness of point features, thereby facilitating cross‑view fusion and improving estimation robustness. Furthermore, we propose a cross‑view‑aligned cylinder integration module to fuse grasp‑relevant geometry into a comprehensive representation. Specifically, the module first aligns the cross‑view points and features according to their similarity to enhance the robustness against noise. Subsequently, these points are registered into the cylindrical coordinate frame, emphasizing the rotation‑symmetric geometry which is important for grasping. Finally, local self‑attention and seed cross‑attention layers are alternately employed, respectively enabling interactions within single views and across views, which supports fine‑grained representation of grasp‑relevant geometry. Our framework achieves strong performance on the GraspNet‑1Billion benchmark and in real‑world applications. Code is available at https://github.com/KJZhuAutomatic/Cross‑view‑Grasp.
Authors:Xiang Yang, Feifei Li, Mi Zhang, Geng Hong, Xiaoyu You, Mi Wen, Min Yang
Abstract:
Diffusion transformers (DiTs) equipped with multimodal attention (MM‑Attn) have become a dominant paradigm for image generation. However, preventing the generation of harmful content remains a critical challenge, particularly in image‑to‑image (I2I) editing tasks. Existing safety mechanisms are primarily designed for text‑to‑image (T2I) synthesis or U‑Net‑based architectures, which limits their effectiveness for unified safety mitigation in DiT‑based frameworks. To bridge this gap, we propose Unified Visual Safety Regulator (UVR), a training‑free safe generation framework that regulates unsafe semantics in generated images. UVR is grounded in an analysis of attention dynamics from the perspective of information flow in MM‑Attn. We identify a task‑independent start‑up stage, during which unsafe semantics in output patches rapidly emerge and can be accurately localized, followed by task‑specific semantic amplification and interference stages, where harmful signals are further propagated and entangled with benign content. Based on these observations, UVR mitigates unsafe generation through unified, targeted attention modulation and explicit restriction of harmful information flow over the identified unsafe output patches. Experiments across various concepts show that UVR achieves state‑of‑the‑art safety performance by achieving 91% and 77% erase rate in image synthesis and editing tasks, while preserving visual quality and fidelity with minimal degradation. Code is available at https://github.com/deng12yx/UVR.
Authors:Yuan Zeng, Zilue Gao, Yujia Shi, Zongqing Lu, Wenming Yang, QingMin Liao
Abstract:
Estimating hand‑surface contact pressure from an egocentric view is crucial for AR/VR devices, robotic imitation, and ergonomic analysis. Existing methods often discretize pressure signal and process frames independently, leading to quantization errors and temporal inconsistencies. We present \emphEgoPressDiff, a conditional video diffusion framework that generates UV‑pressure maps from visual input. The core of our approach is a multi‑modal conditioning strategy, introducing a PoseNet and a Vertex Encoder to efficiently extract features from hand pose and 3D mesh vertices. These signals, along with depth information, guide the generative process to ensure the pressure fields are physically grounded. To effectively fuse these heterogeneous features, we further propose a Distribution‑Calibrated Spatial Layer, which aligns their statistical properties before combination. Evaluated on the EgoPressure ego‑view setting, EgoPressDiff achieves state‑of‑the‑art results, improving Volumetric IoU by over 34% relative to prior baseline, while reducing MAE and maintaining high temporal accuracy. Our project page is at https://egopressdiff.github.io/.
Authors:Jiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang, Yuhang Zang, Tianyi Wei, Xiaohang Zhan, Jiaqi Wang, Tong Wu, Xingang Pan, Dahua Lin
Abstract:
Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text‑to‑image (T2I) flow models with human preferences. However, we have identified that the learning loop of current flow‑based GRPO is fundamentally decoupled from the learner's current capability, suffering from critical blind spots at both prompt selection and advantage estimation: (i) Existing methods sample prompts randomly, overlooking the substantial impact of data selection on reinforcement learning (RL) efficacy‑‑a factor proven crucial in GRPO for large language models; (ii) They evaluate sample quality solely relying on intra‑group statistics, lacking a global perspective to accurately measure true policy improvement. To address these issues, we propose Adaptive GRPO (AdaGRPO), a novel capability‑aware RL algorithm tailored for flow models. Specifically, AdaGRPO consists of two principal components: (i) Online Curriculum Filtering Strategy: Dynamically tracks the model's proficiency and adaptively selects prompts that best match its current learning boundary; (ii) Cross‑Level Advantage Fusion: Synergistically integrates fine‑grained intra‑group advantages with macro‑level global advantages, providing a comprehensive and unbiased policy evaluation. As a lightweight, plug‑and‑play module, AdaGRPO can be seamlessly integrated with existing frameworks such as Flow‑GRPO, DanceGRPO, and Flow‑CPS. Extensive experiments demonstrate that AdaGRPO consistently drives performance gains while significantly stabilizes GRPO training for flow models.
Authors:Ming Dai, Sen Yang, Boqiang Duan, Boyuan Tong, Jiedong Zhuang, Wankou Yang, Jingdong Wang
Abstract:
Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achieve precise pixel‑level localization. Existing methods are limited to reasoning over fixed initial inputs and lack the capacity to actively acquire further visual evidence, which is often essential for resolving complex references in long or intricate videos. To address this, we propose VideoSEG‑O3, the first multi‑turn reinforcement learning framework for RVOS that emulates the human ``coarse‑to‑fine'' cognitive process. It employs a multi‑turn temporal‑spatial chain‑of‑thought to capture fine‑grained details by iteratively pinpointing critical intervals and keyframes. Additionally, to enable the policy to perceive segmentation quality beyond mere text probability of \texttt[SEG] during the RL stage, we introduce SEG‑aware logit calibration, which integrates pixel‑wise segmentation feedback directly into the token‑level logits. Furthermore, we design a decoupled thinking trace to hierarchically decompose the reasoning process into temporal, spatial, and linguistic dimensions, and construct VTS‑CoT, a specialized cold‑start dataset featuring comprehensive reasoning trajectories. The code and models will be released at https://github.com/Dmmm1997/VideoSEG‑O3.
Authors:Dahee Kwon, Haeun Lee, Jaesik Choi
Abstract:
Recent text‑to‑image models built on large‑scale Transformer backbones and flow‑based objectives deliver strong text‑image alignment and high visual quality, yet often produce overly similar samples under a fixed prompt. Existing diversity‑enhancement methods alleviate this issue, but typically require expensive sampling or auxiliary optimization, incurring non‑trivial overhead. To investigate the root cause of this homogeneity, we examine intermediate Transformer features and observe that the zero‑frequency spatial average (DC) component rapidly converges across seeds early in generation, causing early trajectory lock‑in that limits downstream variation. Building on this observation, we propose DC Attenuation for diVersity Enhancement (DAVE), a training‑free representation‑level intervention that selectively attenuates this component in the early regime. DAVE preserves the sampling pipeline with negligible overhead, improving prompt‑consistent diversity while maintaining competitive image quality.
Authors:Tang Li, Yanlin Chen, Mengmeng Ma, Xi Peng
Abstract:
Despite high accuracy, Vision Transformer (ViT) predictions can be driven by spurious cues, raising the need to understand their inner workings before safe deployment. Sparse autoencoders (SAEs) provide a promising lens for decomposing model representations into human‑interpretable concepts, yet adapting SAE‑based interpretation to ViTs remains challenging due to limited control over concept coverage and subjective, non‑scalable feature interpretation. To fill the gaps, motivated by neuroscience‑inspired principles, we propose ViSAE, a mechanistic interpretability toolbox for understanding ViT inner workings through concept circuits. ViSAE consists of three components: (1) A probing suite with 64K images and a 16K visually grounded concept vocabulary, improving concept coverage efficiency by 20x over ImageNet and interpretation accuracy by 28.7% over existing concept sets. (2) Top‑down concept reading and Bottom‑up circuit tracing algorithms that automatically recover ViT inner workings via concept circuits. (3) Applications for auditing and steering ViT behavior. Through concept editing, ViSAE improves the worst‑group accuracy on WaterBirds by 48.2%, outperforming existing methods by 23.8%. Our data and code: https://github.com/deep‑real/ViSAE.
Authors:Richard Li, Aditya Prakash, Andrew Wen, Saurabh Gupta, Yilun Du, Pulkit Agrawal
Abstract:
Human video datasets used for cotraining robot manipulation policies largely consist of curated demonstrations where motions are orchestrated to resemble robot behavior and 3D hand poses are captured with specialized hardware. A more plentiful source of data is everyday Internet video, but it is an open question what factors enable transfer from such videos to robots. We investigate this using a new dataset of 532 human videos with 28 hours of high‑quality triangulated hand labels and natural motions. We find that hand pose quality affects transfer, but even with accurate hands, the inherent motion gap hinders transfer unless the vision and policy networks specialize to each embodiment. Our cotraining recipe yields consistent improvements, with an absolute success rate gain of 29.7% in the low‑robot‑data regime across six manipulation tasks.
Authors:Jingbo Gong, Yikai Wang, Yushi Lan, Yuhao Wan, Ziheng Ouyang, Rui Zhao, Ming-Ming Cheng, Qibin Hou, Chen Change Loy
Abstract:
Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion‑based methods achieve high visual quality but formulate insertion as a simple 2D inpainting task, providing no explicit control over the object's 3D pose and limiting their practical applicability. We propose DIRECT (Decomposed Injection for Reference Composition and Target‑integration), a novel framework that integrates interactive pose manipulation with high‑fidelity 2D image synthesis to enable pose‑controllable object insertion. Our method decomposes the insertion conditions into three complementary components: appearance guidance capturing visual details from the reference object, geometry guidance derived from the user‑adjusted 3D proxy, and context guidance from the target background. By injecting them through separate pathways, DIRECT avoids feature entanglement and simultaneously preserves reference appearance, follows the user‑specified pose, and adapts the object to the target scene. We also introduce an automated data construction pipeline to improve the diversity and quality of training data. Experiments show that DIRECT outperforms previous methods in both geometric controllability and visual quality.
Authors:Zhida Sun, Yulin Zhang, Zheng Gu, Min Lu, Bongshin Lee, Daniel Cohen-Or, Hui Huang
Abstract:
Traditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generative framework for the automated synthesis of pictorial charts that bridges the gap between semantic expression and structural faithfulness. Rather than treating charts merely as images to be stylized, we frame the problem as a dual‑conditioned generation task guided by two parallel external control signals: a text prompt capturing the semantic context of the editing intent, and a context image providing the abstract statistical chart's global structure. To reinforce these controls within a Multi‑Modal Diffusion Transformer, we introduce two complementary feature‑level mechanisms: structural alignment to anchor spatial layouts to the input chart, and semantic alignment to transfer expressive textures from reference images. Generalizing across major visual channels (i.e., length, area, angle, and position) and diverse semantic domains, our method produces pictorial charts that are both artistically compelling and structurally consistent. Extensive quantitative evaluations and perceptual user studies demonstrate that our framework outperforms traditional controllable generation and image editing baselines, providing a foundation for high‑fidelity, data‑driven generative modeling in expressive visual storytelling. Project page: https://ssalign.github.io/.
Authors:Ümit Mert Çağlar, Alptekin Temizel
Abstract:
Semantic segmentation of remote sensing imagery requires models that capture both global context and local detail under tight computational budgets. Prior work typically optimizes for one of these axes: attention for global context, convolution for local detail, or compactness for efficiency. While hybrid approaches aim to capture both, they require architectural changes and encoder backbones with computational overhead, limiting efficiency and performance. We present LALE (Lightweight‑transformer Architecture for Land‑cover Estimation), an end‑to‑end remote sensing image segmentation architecture, that bifurcates its encoder by resolution: lightweight ConvMixer stages handle high‑resolution local features, while transformer stages handle low‑resolution global context, confining the quadratic cost of self‑attention to deep, downsampled feature maps. An all‑MLP multi‑scale decoder, together with RMSNorm and StarReLU throughout, further reduces compute and parameter count. On the large‑scale ARAS400k remote‑sensing segmentation benchmark, LALE establishes a strong efficiency‑performance trade‑off against CNN, transformer, and hybrid baselines. Our smallest variant, (just 1.6M parameters), reaches within 2.6 F1 points of the best baseline (UPerNet) while using 4.5x fewer parameters, 7x less storage, 17x fewer GMACs, and delivering 1.8x higher throughput. The codebase for LALE is publicly available at https://github.com/caglarmert/LALE.
Authors:Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra, Xuelin Chen, Erkut Erdem, Aykut Erdem, Duygu Ceylan
Abstract:
Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language‑driven, human‑centric camera framing in generative video. Our core insight is that professional filmmakers conceive shots not as world‑space trajectories but as framings defined relative to the actor, encoding shot size, angle, and composition as functions of human pose and motion. We formalize this intuition as a human‑centric camera parameterization and introduce a Domain‑Specific Language (DSL) that is convertible to standard 6‑DoF camera parameters. A fine‑tuned multimodal large language model then acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are deterministically interpolated into continuous camera trajectories, which are then provided as input to video generators. We train and evaluate Auteur on a new dataset of 34K aligned text, human motion, and DSL‑annotated camera trajectories drawn from procedural synthesis and real‑world movie footage from the CondensedMovies dataset. Auteur enables cinematographic framing of human‑centered scenes, a capability largely absent in prior generative models. To assess this behavior, we propose new framing‑focused metrics, and our experiments show that Auteur consistently outperforms existing methods. Project page is https://cyberiada.github.io/Auteur/
Authors:Muyi Bao, Yuxin Cai, Hang Xu, Zongtai Li, Jinxi He, Jingfan Tang, Chen Lv, Ji Zhang, Yaqi Xie, Wenshan Wang
Abstract:
Vision‑language models (VLMs) have become a common foundation for vision‑and‑language navigation in continuous environments (VLN‑CE). Yet most VLM‑based methods cast navigation as low‑level action prediction, an interface that is ambiguous, tied to short‑horizon motion primitives, and inefficient due to repeated VLM querying. We propose Goal2Pixel, a pure pixel‑based paradigm that reformulates VLN‑CE as navigable pixel grounding. Rather than predicting actions, Goal2Pixel uses the image plane as a unified spatial interface between VLM reasoning and robot motion: the model predicts a visible navigable pixel to the agent, which is back‑projected into a 3D waypoint for forward navigation. For non‑forward actions, we append auxiliary directive regions to the image plane, where the left/right/bottom regions are interpreted as turning left, turning right, and stopping, respectively. To enable long‑horizon navigation, we propose a visibility‑aware keyframe memory for compact and informative history representation. To adapt pretrained VLMs to navigable pixel grounding, we introduce semantic embeddings and coordinate‑aware auxiliary losses. Goal2Pixel achieves competitive state‑of‑the‑art performance while requiring fewer VLM inference calls than prior methods. On R2R‑CE Val‑Unseen it achieves 54.1% SR and 52.5% SPL with just 7.75 VLM calls per episode, 6x fewer than the 46.62 required by direct action prediction at 32.9% SR. The same trend holds on RxR‑CE.Project Page: https://baobao0926.github.io/Goal2Pixel/.
Authors:Žiga Kovačič, Kevin Ellis
Abstract:
To study the ability to infer physical dynamics from videos and extrapolate them forward in time, we assemble a dataset of 2D Material Point Method (MPM) physical simulations covering rich physical phenomena such as deformable objects, fluids, kinetic objects, and emitters. We study code generation and video diffusion approaches on this dataset, identifying their strengths and weaknesses by varying the amount of physically relevant side information. The code generation model, beyond giving a working demonstration of automatic synthesis of MPM simulations, reveals that such an approach struggles with inferring physical parameters from visual input, but relative to video diffusion, produces physically and temporally stable extrapolations forward in time, while the video diffusion model more strongly identifies geometric properties from visual input but produces physically implausible extrapolations.
Authors:Olaf Dünkel, Basavaraj Sunagad, Haoran Wang, David T. Hoffmann, Christian Theobalt, Adam Kortylewski
Abstract:
Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part‑level supervision. Semantic correspondence (SC) evaluates this capability by testing whether object parts can be matched across instances and categories under large variations in appearance, viewpoint, and geometry. To enable a systematic SC evaluation, we introduce SOCO, a new benchmark for Semantic Object Correspondence that introduces a taxonomy of correspondence types and provides consistent, functionally meaningful keypoint annotations across 100 categories and over 1M correspondence pairs. In addition, SOCO includes keypoint language descriptions, enabling the evaluation of large vision‑language models (LVLMs) and their fine‑grained part‑level understanding. Comprehensive experiments reveal that (i) vision foundation backbones encode strong semantic structure but transfer correspondences poorly across related categories and only partially capture object‑part position, (ii) LVLMs are stronger at text‑prompted part localization than at visual‑reference cross‑image matching, exposing a gap between language‑grounded localization and fine‑grained visual correspondence, and (iii) correspondence performance predicts performance on dense downstream tasks, including segmentation, tracking, 3D pose estimation, and 3D detection, more strongly than ImageNet classification. Together, these findings position SOCO as a benchmark for structured, part‑level representation quality in vision and multimodal foundation models.
Authors:Shangjie Xue, Jesse Dill, Dhruv Ahuja, Frank Dellaert, Panagiotis Tsiotras, Danfei Xu
Abstract:
We present Gaussian Splatting Anisotropic Visibility Field (GAVIS), a novel framework for uncertainty quantification and active mapping in 3DGS. Our key insight is that regions unseen from the training views yield unreliable predictions from the 3DGS. To address this, we introduce a principled and efficient method for quantifying the visibility field in 3DGS, defined as the anisotropic visibility of each particle with respect to the training views, and represented using spherical harmonics. The resulting visibility field is integrated into a Bayesian Network‑based uncertainty‑aware 3DGS rasterizer, enabling real‑time (200 FPS) uncertainty quantification for synthesized views. Active mapping is further performed within a maximum information gain framework building on this formulation. Extensive experiments across diverse environments demonstrate that GAVIS consistently and significantly outperforms prior approaches in both accuracy and efficiency. Moreover, beyond standalone use, our method can be applied post‑hoc to improve the performance of existing approaches.
Authors:Xiaoxuan Ma, Jiashun Wang, Nicolas Ugrinovic, Yehonathan Litman, Kris Kitani
Abstract:
Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation‑ready digital assets for applications such as immersive interaction and content creation. However, existing single‑image reconstruction methods fall short in capturing the physical structure of a scene. As a result, they often produce geometrically plausible but physically inconsistent results, including object floating and penetration, which lead to unstable behavior in physics simulations. Image‑conditioned scene generation methods improve physical plausibility but often rely on strong scene priors, yielding plausible yet inaccurate object arrangements that fail to match the input image. We propose REST3D, a single‑image reconstruction framework that can reconstruct physically stable 3D scenes by integrating physical scene understanding with physics‑constrained refinement. We first introduce an agentic physical scene understanding technique that constructs a scene‑tree representation capturing object physical states and inter‑object relationships from a gravity‑support perspective, providing a structural prior for reconstruction. Leveraging this structure, we initialize the scene using image‑to‑3D models, followed by scene‑tree‑guided alignment and physics‑constrained optimization to resolve physical violations while preserving visual consistency with the input image. Experiments show that our method significantly reduces physical errors and improves simulation stability on both synthetic and real‑world datasets while maintaining strong reconstruction quality. We further demonstrate the reconstructed scenes in VR‑based human‑object interaction, showing their potential for immersive applications.
Authors:Hadar Davidson, Noam Issachar, Sagie Benaim
Abstract:
Diffusion models achieve state‑of‑the‑art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low‑frequency global structures early and high‑frequency fine details later. Conventional stochastic differential equation (SDE) solvers fail to account for this dynamic, naively injecting uniform white noise throughout the entire process and misusing the finite energy budget. In this work, we establish a mathematical framework that reconsiders SDE inference as a targeted, frequency‑decoupled energy transfer. Leveraging this framework, we introduce Colored Noise Sampling (CNS), a novel, training‑free stochastic solver. Rather than injecting uniform white noise, CNS utilizes a dynamic, timestep‑ and frequency‑dependent schedule that more efficiently allocates injected energy toward structurally unresolved frequency bands. By actively exploiting the model's inherent spectral bias, CNS systematically steers the generated distribution toward the true data manifold. Extensive experiments demonstrate that CNS significantly outperforms standard ODE and SDE baselines as a strictly plug‑and‑play, inference‑time sampler substitution across diverse architectures (SiT, JiT, FLUX). Compared to standard sampling on ImageNet‑256, CNS achieves substantial unguided FID reductions, improving from 8.26 to 6.27 on SiT‑XL/2, 32.39 to 26.69 on JiT‑B/16, and 11.88 to 8.31 on JiT‑H/16, while yielding consistent relative FID improvements with Classifier‑Free Guidance. Project page is available at https://hadardavidson.github.io/CNS/.
Authors:Daniel Rho, Jun Myeong Choi, Matthew Thornton, Biswadip Dey, Roni Sengupta
Abstract:
Existing inverse physics methods recover physical parameters from multi‑view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, leading to severe scale ambiguity, inaccurate geometry, and weak coupling between appearance optimization and physical simulation. We propose MonoPhysics, a framework for monocular inverse physics estimation of deformable objects using differentiable MPM simulation and 3D Gaussian Splatting, which jointly optimizes geometry, appearance, and physical parameters from a single camera view. We address these challenges through three visual‑physical bridges: global scale alignment, physics‑aware geometry refinement, and a differentiable position map, which together enable accurate optimization from monocular observations alone. We evaluate on Vid2Sim and our new dataset of elastic and plastic objects, showing that MonoPhysics outperforms existing baselines in monocular settings and achieves performance comparable to multi‑view baselines using only a single camera. Our project page is available at https://daniel03c1.github.io/MonoPhysics/
Authors:Ruixiang Jiang, Chang Wen Chen
Abstract:
Portrait photography is largely decided before the shutter opens: the subject's pose, the camera configuration, and the lighting devices must be coordinated within the surrounding 3D scene. In contrast, most existing computational methods focus on post‑production in 2D image space, such as retouching, relighting, or editing images that already exist; pre‑capture photographic planning remains largely unexplored. We introduce 3D aesthetic portrait planning, the task of generating human pose, camera, lighting, and exposure plans that produce visually compelling portraits while satisfying geometric and photometric feasibility in a 3D scene. Our approach builds a Photographic Scene Graph that represents scene affordances, subject‑scene relations, and portrait‑relevant lighting structure. Built on this representation, we perform aesthetic‑guided comparative planning over previous attempts and current viewfinder observations. Experiments across diverse indoor and outdoor scenes show that our method produces portraits preferred by human raters and MLLM evaluators over competitive baselines, while maintaining high physical plausibility. Together, our results suggest a path from post‑capture correction toward pre‑capture computational portrait planning. Project repository: https://github.com/songrise/Before‑the‑Shutter
Authors:Chong Bao, Shichen Liu, Lijun Yu, David Futschik, Stylianos Moschoglou, Shefali Srivastava, Ziqian Bai, Feitong Tan, Guofeng Zhang, Zhaopeng Cui, Sean Fanello, Yinda Zhang
Abstract:
Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archon, a fully pretrained, human‑centric unified multimodal model for holistic avatar generation. Archon unifies seven modalities with modality‑specific tokenizers, and a native autoregressive unified multimodal model pretrained on synchronized modalities and 72 diverse tasks to model holistic joint distributions. To address the token explosion challenge in high‑fidelity talking videos, we introduce a memory‑efficient semantic video reparameterization, achieving 4x token reduction while preserving fine‑grained dynamics, coupled with a semantic‑driven video diffusion decoder. We further propose a "Thinking in Modality" that decomposes ambiguous cross‑modal tasks into stepwise thinking in an alternative chain of modality, progressively enhancing fidelity and controllability. Extensive experiments demonstrate that Archon achieves superior or comparable performance across diverse digital human generation tasks, validating the effectiveness of our unified framework. Project page: https://zju3dv.github.io/archon/.
Authors:Omer Benishu, Gal Fiebelman, Sagie Benaim
Abstract:
We address the task of generating physically accurate and visually faithful 4D Human‑Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize dynamic scenes where the human actively engages with the object through actions, such as punching or kicking, in accordance with a given input text. To this end, we introduce PhyGenHOI, a novel framework that couples generative human motion with an explicit physical object simulation. We model the human as a semantic agent driven by a Motion Diffusion Model (MDM) and the object as a physical agent simulated via the Material Point Method (MPM), utilizing 3D Gaussians as a unified, differentiable representation. We supervise their interaction through three coupled mechanisms: (1) A Windowed Attraction Loss that temporally synchronizes generative motion to intercept the object; (2) A Contact‑Driven Re‑simulation step that triggers physically consistent momentum transfer upon impact; and (3) A Masked Video‑SDS objective that injects video‑based priors to enhance contact fidelity. Experiments show PhyGenHOI generates physically consistent 4D HOI across diverse actions, humans, and objects, outperforming baselines. Project page and videos: https://omerbenishu.github.io/PhyGenHOI/
Authors:Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou, Yimin Chen, Wenqiang Sun, Kaiwen Zheng, Guande He, Xiao Yang, Chongxuan Li, Fan Bao, Jun Zhu
Abstract:
Recent video diffusion foundation models have achieved remarkable progress in high‑quality video generation, yet turning them into real‑time interactive video world models remains challenging. Interactive world models require controllable, causal, and low‑latency rollout, which in practice demands a full pipeline spanning data construction, controllable fine‑tuning, autoregressive training, few‑step distillation, and streaming inference. In this work, we present minWM, a full‑stack open‑source framework for building real‑time interactive video world models. minWM provides an end‑to‑end pipeline that converts existing bidirectional T2V/TI2V video foundation models into camera‑controllable few‑step autoregressive world models. Specifically, minWM first fine‑tunes a bidirectional video diffusion model with camera control, and then applies the Causal Forcing / Causal Forcing++ pipeline, including AR diffusion training, causal ODE or causal consistency distillation, and asymmetric DMD, to distill it into a few‑step autoregressive generator for low‑latency rollout. The framework is modular and architecture‑extensible: we instantiate it on representative open backbones, including Wan2.1‑T2V‑1.3B and HY1.5‑TI2V‑8B, covering both cross‑attention‑based condition injection and MMDiT‑style architectures. minWM also supports adapting existing video world models, such as HY‑WorldPlay, to new data distributions, training recipes, and latency targets. Beyond releasing runnable scripts, checkpoints, documentation, and inference code, we provide practical ablations on camera trajectory quality, controllability training steps, and minimal batch‑size requirements. We hope minWM serves as a reproducible and extensible recipe for building and adapting real‑time interactive video world models.
Project Page: [https://github.com/shengshu‑ai/minWM](https://github.com/shengshu‑ai/minWM)
Authors:Ziwen Xu, Haiwen Hong, Linsong Yu, Benglei Cui, Longtao Huang, Hui Xue, Ningyu Zhang
Abstract:
Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real‑world environments. While Low‑Rank Adaptation (LoRA) is widely used for such memory updates, existing studies mainly rely on qualitative downstream evaluations, leaving the quantitative capacity limits and underlying dynamics of exact parametric memory largely unexplored. To bridge this gap, we employ LoRA as a controlled memory capacity probe within the latent space to systematically quantify exact parametric memory. We introduce the Parametric Memory Law, a robust power law linking loss reduction Delta L to effective parameters and sequence length. At the token level, fine‑grained analysis reveals a deterministic phase transition, demonstrating that a prediction probability of p > 0.5 constitutes a sufficient condition for verbatim recall under greedy decoding. Driven by these insights, we introduce MemFT, a threshold‑guided optimization strategy that dynamically redistributes the training budget toward sub‑threshold tokens. Empirical evaluations demonstrate that MemFT can enhance memory fidelity and efficiency. Code will be released at https://github.com/zjunlp/ParametricMemoryLaw.
Authors:Ciara Rowles, Reshinth Adithyan, Nikhil Pinnaparaju, Vikram Voleti, Mark Boss
Abstract:
We present Stable‑Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine‑tuning a pretrained layer decomposition model using only feedback from a vision‑language model (VLM). Starting from Qwen‑Image‑Layered, we apply Flow‑GRPO with LoRA adaptation, sampling multiple candidate decompositions per image, scoring them with a VLM, and optimising the policy from group‑relative advantages. The key challenge lies in designing a reliable reward signal: VLMs scoring samples in isolation tend to compress their judgements into a narrow band, leaving GRPO with little within‑group variance to learn from. We address this with a two‑stage evaluation pipeline that pairs structured per‑sample scoring across five edit‑centric criteria with a grid‑based calibration step in which the VLM re‑scores all candidates side‑by‑side. Stable‑Layers produces decompositions with stronger layer separation, fewer blank or artifact‑heavy layers, and lower per‑layer reconstruction error on the Crello dataset compared to the base model.
Authors:Xin Dong, Weijian Deng, Lihan Zhang, Tianru Dai, Wenfeng Deng, Yansong Tang
Abstract:
This work addresses the problem of recovering complete, simulatable object geometry from reconstructed real‑world scenes, enabling physics‑based interaction with objects embedded in the scene. While modern multi‑view reconstruction methods can produce visually accurate environments, objects are often incomplete due to occlusions and limited observations, making them unsuitable for physics simulation. To address this limitation, we propose SAM3D‑Phys, a framework that integrates scene reconstruction with generative 3D priors of SAM3D to recover physically simulatable objects. Our approach first reconstructs the scene from multi‑view images to obtain scene geometry and partial observations of objects. We then leverage SAM3D to infer complete object geometry from these partial observations. To ensure that the recovered objects remain consistent with the reconstructed scene, we restore scene‑consistent object states through two complementary strategies: a physics‑constrained spatial optimization algorithm that iteratively aligns the recovered object to its original location, and a mask‑guided appearance distillation module that refines texture fidelity based on the observed images. By recovering complete object geometry and restoring its pose and appearance within the scene, SAM3D‑Phys produces clean object representations suitable for physics‑based simulation, enabling simultaneous and physically consistent interactive simulation of multiple objects within a reconstructed scene. Project page: https://chnxindong.github.io/sam3d‑phys/
Authors:Chun-Hsiao Yeh, Shengyi Qian, Manchen Wang, Yi Ma, Joseph Tighe, Fanyi Xiao
Abstract:
Vision‑Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine‑tuning with 3D visual question‑answering (VQA) datasets may overfit dataset‑specific biases, while integrating specialized 3D visual encoders is often inflexible and cumbersome. In this paper, we argue that genuine spatial understanding should emerge from learning fundamental geometric priors, not only from high‑level VQA supervision. We propose GASP (Geometric‑Aware Spatial Priors), a framework that injects these priors directly into the LLM's transformer layers. GASP employs a small correspondence head, applied as a deep supervision signal across all layers, and is trained with a dual objective leveraging ground‑truth geometry from large‑scale video scenes: a contrastive loss on ground‑truth point correspondences enforces 2D view‑invariance, while a depth consistency supervision resolves 3D geometric ambiguities. Our analysis first provides a diagnostic showing that standard VLMs' internal correspondence matching accuracy is very low (often below 5%). We then demonstrate that our training substantially improves this behavior, boosting peak layer‑wise correspondence to over 70% and maintaining over 85% temporal robustness while baselines remain below 5%. These internal improvements translate to significant gains on downstream spatial benchmarks including +18.2% on All‑Angles Bench and +29.0% on VSI‑Bench, all without training on any 3D VQA data. Our findings indicate that learning from fundamental geometric priors is a promising and generalizable pathway towards VLMs with more reliable 3D spatial reasoning.
Authors:Rongzhen Zhao, Zhiyuan Li, Ruonan Wei, Juho Kannala, Joni Pajarinen
Abstract:
Self‑supervised video Object‑Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self‑supervised Multi‑Object Tracking (MOT) focuses on associating pre‑defined object detections or segmentations. Although well‑established in MOT, Cycle Consistency (CC) cannot naively or explicitly apply to the latent slot space of OCL. Unlike the deterministic and ideal object representations in MOT, OCL slots are inherently stochastic and ambiguous due to non‑unique scene decompositions. Enforcing explicit cycle consistency (ECC) on slots imposes rigid mean seeking. This severely penalizes the model for exploring alternative but equally valid decompositions, thereby driving towards feature collapse. To resolve this dilemma, we propose Implicit Cycle Consistency (ICC), which shifts the cycle‑consistency constraint from the restrictive slot space to the continuous reconstruction manifold, encouraging slots to reach a soft consensus on collectively interpreting the visual scene rather than forcing rigid point‑to‑point feature alignment. Extensive experiments on complex video OCL benchmarks demonstrate that ICC avoids feature collapse and outperforms ECC baselines. Our source code, model checkpoints and training logs are provided on https://github.com/Genera1Z/ICC.
Authors:Matan Levy, Ran Margolin, Bar Cavia, Dvir Samuel, Yael Pritch, Shmuel Peleg, Alex Rav Acha, Ariel Shamir, Dani Lischinski
Abstract:
We introduce LiveSVG, a zero‑shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation methods struggle with complex motions: LLM‑based code synthesis fails to express fine, non‑rigid Bézier deformations, while Score Distillation Sampling (SDS) provides noisy gradients and often requires category‑specific priors like skeletons. In contrast, LiveSVG fits vector geometry directly to an explicitly generated target video. Given an input SVG image and a motion prompt, we generate a previewable target video using a frozen image‑to‑video model, then fit the original SVG to this video via differentiable rendering. Our fitting stage is skeleton‑free, utilizing a dual‑level motion representation that combines per‑group homographies for coarse articulation with per‑path Bézier control‑point offsets for local deformations. To resolve color‑induced correspondence ambiguities during pixel‑wise fitting, we introduce a novel sphere‑packing recolorization strategy. We also present ChallengeSVG, a benchmark of complex, multi‑object scenes that exposes the limitations of prior work. Evaluations demonstrate that LiveSVG significantly outperforms existing methods on both AniClipart and ChallengeSVG, establishing direct reference‑video fitting as a practical, robust route to prompt‑aligned and fully editable vector animation.
Authors:Cheolhong Min, Jaeyun Jung, Daeun Lee, Hyeonseong Jeon, Yu Su, Jonathan Tremblay, Chan Hee Song, Jaesik Park
Abstract:
Vision‑language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts in natural images. We introduce a representation‑level analysis framework that constructs minimal contrastive pairs to measure how spatial axes are organized and disentangled within VLM embeddings. Our analysis across multiple model families reveals a consistent vertical‑distance entanglement: models conflate vertical image position with distance, mirroring the perspective bias of natural photographs. This bias produces a significant accuracy gap between perspective‑consistent and counter‑heuristic examples, and intensifies under data scaling even as overall benchmark accuracy improves. We further show that models with similar benchmark scores can exhibit different internal representations, and that these differences predict accuracy and robustness across diverse spatial reasoning benchmarks. To isolate this bias from evaluation‑set skew, we introduce SpatialTunnel, a synthetic benchmark designed to expose spatial shortcut biases by removing common correlations present in natural images. Experiments confirm that the entanglement is model‑intrinsic, and that models with well‑separated spatial axes exhibit greater robustness, suggesting that well‑structured spatial representations lead to more reliable spatial reasoning across diverse benchmarks. Code and benchmark are available on the project page: https://cheolhong0916.github.io/whyfarlooksup.github.io/.
Authors:Zhuguanyu Wu, Ruihao Gong, Yang Yong, Yushi Huang, Xiangyu Fan, Lei Yang, Dahua Lin, Xianglong Liu
Abstract:
Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few‑step video diffusion models. However, DMD‑style video distillation faces two coupled challenges: the fake score must track a continuously evolving generator, making training costly when frequent updates are required, while reverse‑KL‑style matching can be mode‑seeking and conservative for preserving strong motion dynamics. To address these issues, we propose Score Gradient Matching Distillation (SGMD). SGMD adopts a fake‑score perspective by directly optimizing the fake score toward the teacher, while using teacher stop‑gradient Fisher as a stable distribution‑matching objective. We provide a gradient analysis that motivates this objective choice under ideal tracking. Building on this, SGMD introduces a pair of dual potentials: negative‑residual (NR) for outer‑loop correction and residual‑contraction (RC) for inner‑loop tracking. Empirically, compared to DMD2, SGMD achieves an approximately ~ 3× training speedup and substantially improves motion dynamics for 4‑step distilled models while preserving temporal consistency. A human study confirms that SGMD is preferred in motion quality and overall preference, while visual quality and text alignment remain comparable. Code is available at https://github.com/ModelTC/LightX2V.
Authors:Zhu Yu, Zhengyi Zhao, Runmin Zhang, Lingteng Qiu, Kejie Qiu, Yisheng He, Siyu Zhu, Zilong Dong, Si-Yuan Cao, Hui-Liang Shen
Abstract:
This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single‑view metric depth estimation with sparse observations. Without relying on complex architectural designs, LDCM generates metric‑accurate dense depth maps using a transformer. It outperforms existing approaches across diverse datasets and sparse observations. We achieve this from two key perspectives: (1) leveraging existing monocular foundation models to improve the quality of sparse depth inputs, and (2) reformulating training objectives to better capture geometric structure and metric consistency. Specifically, a Poisson‑based depth initialization strategy is first introduced to generate a uniform coarse dense depth map from diverse sparse observations, providing a strong structural prior for the network. Regarding the training objective, we replace the conventional depth head with a point map head that regresses per‑pixel 3D coordinates in camera space, enabling the model to directly learn the underlying 3D scene structure instead of performing pixel‑wise depth map restoration. Moreover, this design eliminates the need for camera intrinsic parameters, allowing LDCM to naturally produce metric‑scaled 3D point maps. Extensive experiments demonstrate that LDCM consistently outperforms state‑of‑the‑art methods across multiple benchmarks and varying sparsity levels in both depth completion and point map estimation, showcasing its effectiveness and strong generalization to unseen data distributions.
Authors:Longbin Ji, Guan Wang, Xuan Wei, Chenye Yang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Jingzhou He
Abstract:
Joint audio‑video generation aims to synthesize temporally synchronized and semantically coherent visual‑acoustic content. However, existing open‑source methods mainly rely on either dual‑tower designs with posterior alignment or fully unified tri‑modal designs that mix textual context, audio and video in one shared space. The former weakens fine‑grained audio‑video co‑evolution, while the latter couples semantic conditioning with low‑level synchronization. To address these limitations, we propose NAVA, a Native Audio‑Visual Alignment framework for joint audio‑video generation. NAVA is built upon context‑conditioned native audio‑visual alignment: it first establishes audio‑video correspondence in a dedicated interaction space, and then uses external context to condition the joint denoising process. Specifically, NAVA is instantiated with an Align‑then‑Fuse MMDiT architecture, which transitions from modality‑aware audio‑video alignment to modality‑shared joint denoising. Furthermore, we introduce Timbre‑in‑Context Conditioning to associate reference timbre cues with corresponding speech spans to achieve controllable speech timbre. Experiments on Verse‑Bench and Seed‑TTS, together with a user study, demonstrate that NAVA achieves superior video quality, precise audio‑visual synchronization, competitive audio quality, and stronger reference‑timbre controllability using only 6.3B parameters.
Authors:Yuqing Chen, Lin Liu, Haisu Wu, Xiaopeng Zhang, Yaowei Wang, Yujiu Yang, Qi Tian
Abstract:
Video object removal frequently struggles to simultaneously eliminate target objects and their associated physical effects (e.g., smoke, reflections, light, and ripples) in out‑of‑domain scenarios due to complex spatiotemporal ambiguities. While existing methods primarily rely on spatial masks, they often fail to capture weakly correlated effects, and the potential of explicit textual guidance remains underexplored. Furthermore, a fundamental optimization conflict exists in removal models between high‑level semantic generalization and precise pixel‑level background preservation. To address these challenges, we propose GenEraser, a novel framework for generalized and high‑fidelity video object and effect removal. First, we introduce a Multi‑Conditional Mixture‑of‑Experts (MC‑MoE) paired with Bipartite Text guidance to fully exploit the multimodal priors of Diffusion Transformers, significantly enhancing the identification of complex effects. Second, a Learnable Deep ``CFG'' Fusion mechanism (LD‑CFG) is developed to adaptively balance the relative dominance of mask and textual conditions across diverse scenarios. Finally, we propose a Decoupled Expert Architecture, comprising a Locator and a Preserver, to mitigate the inherent trade‑off between semantic generalization and pixel alignment. Extensive experiments demonstrate that our GenEraser surpasses recent state‑of‑the‑art approaches, achieving significant quantitative improvements (e.g., 2.16 dB and 1.44 dB on the ROSE Benchmark and VOR‑Eval, respectively) while maintaining exceptionally robust generalization in open‑world scenarios. https://cyqii.github.io/GenEraser.github.io/
Authors:Jaa-Yeon Lee, Yeobin Hong, Taesung Kwon, Jong Chul Ye
Abstract:
Diffusion models generate highly realistic images but often struggle with precise text‑image alignment. While recent post‑training methods improve alignment using external rewards or human preference signals, their performance heavily depends on reward quality and does not directly address alignment within the diffusion process itself. Recent reward‑free approaches such as SoftREPA demonstrate that optimizing soft text tokens via contrastive learning can effectively improve text‑image representation alignment, outperforming standard parameter‑efficient fine‑tuning baselines. However, the contrastive formulation can excessively penalize negative pairs, which manifests as characteristic failure cases such as over‑counting and repetition. To address this issue, we propose a lightweight, reward‑free post‑training method that refines soft tokens by integrating contrastive alignment guidance directly into the score‑matching objective of diffusion models. By assigning alignment directions at the score level, our approach mitigates these limitations and yields more coherent and semantically faithful generations. Experiments show that our method matches SoftREPA while substantially improving its failure cases, achieving over 35% improvement in counting accuracy on the GenEval benchmark. Our method is seamlessly applicable to existing diffusion backbones (SD1.5, SDXL, and SD3), and is complementary to existing RL‑based diffusion post‑training methods. Project page: https://jaayeon.github.io/AGSM
Authors:Hesong Wang, Xin Jin, Lu Lu, Chenhaowen Li, Jian Chen, Qiang Liu, Huan Wang
Abstract:
Video large language models (Video‑LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amounts of visual tokens. Although recent approaches achieve extremely low token retention ratios while maintaining accuracy comparable to full‑token baselines, most of them perform compression only at the late stage of prefilling, leaving the efficiency of the vision encoder unoptimized. In this paper, we first show that vision encoding contributes a large portion to the time‑to‑first‑token (TTFT). Therefore, instead of compressing visual tokens only after the vision encoder, performing compression inside the encoder still leaves substantial room for exploration. Based on this insight, we propose EarlyTom, a training‑free token compression framework that performs early‑stage visual token compression inside the vision encoder, enabling significantly better TTFT reduction and higher throughput. In addition, we introduce a decoupled spatial token selection strategy that improves the overall compression effectiveness. EarlyTom reduces TTFT by up to 2.65x and FLOPs by up to 61% on a single NVIDIA A100 GPU for the LLaVA‑OneVision‑7B model, while maintaining accuracy comparable to the full‑token baseline. These improvements substantially enhance the practicality of deploying Video‑LLMs in real‑world production scenarios.
Authors:Muhammed Furkan Dasdelen, Fatih Ozlugedik, Ilaria Looser, Rao Muhammad Umer, Christian Pohlkamp, Carsten Marr
Abstract:
Multimodal alignment of histopathology encoders with transcriptomic and genomic data has been shown to significantly improve performance in downstream diagnostic tasks. Hematological cytology is unique in that visual single‑cell evaluation is often paired with cytogenetics and molecular genetics for blood cancer diagnosis. In this study, we present a framework to align single white blood cell images with chromosomal aberrations (karyotype) and somatic mutations from targeted gene panels. Our training strategy follows a two‑stage approach: (i) self‑supervised, vision‑only pretraining of a transformer aggregator using an iBOT head on a cohort of over 1500 patients, and (ii) genetic alignment via supervised contrastive loss on acute myeloid leukemia patients. Our genetically aligned patient encoder improves hematological diagnostic tasks, outperforming slide‑level histopathology foundation models. Additionally, the model provides off‑the‑shelf retrieval capabilities for diseases and genetic alterations. Incorporating genetic data into patient encoders increases the quality of patient representations, providing a framework that aligns with clinical diagnostic workflows and paves the way for future multimodal hematology‑specific AI. The code and model weights are available at https://github.com/marrlab/GenBloom.
Authors:David Hagerman, Roman Naeem, Jakob Lindqvist, Carl Lindström, Fredrik Kahl, Lennart Svensson
Abstract:
Sparse vision transformers have gained popularity as efficient encoders for medical volumetric segmentation, with Swin emerging as a prominent choice. Swin uses local attention to reduce complexity and yields excellent performance for many tasks but still tends to overfit on small datasets. To mitigate this weakness, we propose a novel architecture that further enhances Swin's inductive bias by introducing Inception blocks in the feed‑forward layers. The introduction of these multi‑branch convolutions enables more direct reasoning over local, multi‑scale features within the transformer block. We have also modified the decoder layers in order to capture finer details using fewer parameters. We demonstrate a performance improvement on eleven different medical datasets through extensive experimentation. We specifically showcase advancements over the previous state‑of‑the‑art backbones on benchmark challenges like the Medical Segmentation Decathlon and Beyond the Cranial Vault. By showing that the existing inductive bias in Swin can be further improved, our work presents a promising avenue for enhancing the capabilities of sparse vision transformers for both medical and natural image segmentation tasks. Code and pre‑trained weights can be accessed at https://github.com/Eiphodos/SwInception.
Authors:Cheng Sun, Jaesung Choe, Min-Hung Chen, Ryo Hachiuma, Yu-Chiang Frank Wang
Abstract:
Recent Large View Synthesis Models (LVSMs) advocate an encoder‑decoder architecture that separates reconstruction and rendering into distinct networks. We re‑examine this design. Through controlled experiments, we show that a decoder‑only architecture, which represents scenes implicitly as a KV‑cache, outperforms encoder‑decoder variants while using fewer parameters at identical rendering complexity. Further analysis shows that sharing weights between the color‑input reconstruction network and the camera‑only rendering network better aligns their features at the same viewpoint, facilitating image synthesis. Building on this finding, our model, dubbed DVSM, further incorporates foundation model priors and stage‑wise patch sizing for an improved efficiency‑quality tradeoff. Our results establish a new state of the art for novel‑view synthesis across multiple benchmarks, in some cases even outperforming per‑scene‑optimized 3DGS under dense input views.
Authors:Hongyu Long, Jiaxuan Liu, Rui Cao
Abstract:
As a widespread form of informal settlements, urban villages present significant challenges for sustainable urban development and governance. Precise mapping of their infrastructure is essential, however, existing remote sensing datasets primarily focus on formal urban environments, lacking fine‑grained annotated data for the high‑density building patterns and narrow road networks typical of urban villages. To address this gap, we introduce the DenseUIS dataset, the first high‑resolution remote sensing dataset specifically designed for building and road extraction in extremely dense urban informal settlements, covering 126 urban villages across Shenzhen and Guangzhou in China. Furthermore, we conduct a comprehensive evaluation of state‑of‑the‑art deep learning models on this dataset. Experimental results reveal the limitations of existing methods in handling the unique morphological patterns of dense informal settlements, underscoring the need for specialized approaches. DenseUIS therefore provides a robust benchmark for advancing fine‑grained urban mapping in complex and high‑density informal environments. The dataset is publicly available at https://github.com/rui‑research/DenseUIS.
Authors:Leyi Qi, Yiming Li, Siyuan Liang, Zhengzhong Tu, Dacheng Tao
Abstract:
Large‑scale text‑to‑image (T2I) diffusion models have enabled unprecedented creative applications, but their unauthorized use has raised serious intellectual property concerns, making model ownership verification (MOV) increasingly critical. We find that existing backdoor‑based diffusion watermarking methods often (implicitly) assume a "faithful" verification process, namely, that the verifier can query a suspicious model and obtain the faithful watermark response to complete MOV. However, in practice, adversaries may intentionally or unintentionally damage potential watermark signals, significantly degrading verification reliability. To address this issue, we propose Cert‑LAS, the first certified MOV method for T2I models based on layer‑adaptive smoothing. In general, Cert‑LAS embeds specified watermarks using diffusion classifiers and an LFS‑guided layer‑adaptive noise, and verifies ownership by examining whether the suspected model exhibits significantly stronger watermark responses compared to unwatermarked references through hypothesis testing. We further prove that, under certain conditions, our Cert‑LAS can still achieve reliable verification even in the presence of malicious removal attacks. Extensive experiments validate the effectiveness of Cert‑LAS and its resistance to adaptive attacks. Our code is available at https://github.com/Leyi‑Qi/Cert‑LAS.
Authors:Shuai Yi, Yixiong Zou, Yuhua Li, Ruixuan Li
Abstract:
Vision‑Language Models (VLMs) such as CLIP demonstrate strong zero‑shot generalization, but their performance significantly degrades in cross‑domain scenarios with scarce target‑domain training data (Cross‑Domain Few‑Shot Learning, CDFSL). In this paper, we focus on the target‑domain few‑shot finetuning in the CLIP‑based CDFSL task. Prevailing finetuning paradigms uniformly align all image patch tokens with their corresponding textual embeddings. However, we find a counterintuitive phenomenon: actively pushing away certain low‑similarity image tokens, termed "tail tokens", from their textual embeddings consistently improves target‑domain performance. We delve into this phenomenon and provide a novel interpretation: under great domain shifts and scarce training data, the model can hardly extract semantic information from visual inputs; therefore, the common belief of alignment is valid only for tokens already containing sufficient semantic information; for tail tokens, forcing the alignment would lead to excessive overfitting to the scarce training, while breaking the alignment is more useful. Motivated by this, we propose Adaptive Tail‑Head Alignment (ATHA), a novel fine‑tuning strategy for CLIP that transforms the conventional uniform alignment paradigm to an adaptive alignment paradigm, with both alignment strengthening and weakening. Extensive experiments on four challenging CDFSL benchmarks validate our state‑of‑the‑art performance. Our code is available at https://github.com/shuaiyi308/ATHA.
Authors:Boyuan Zhang, Huanshan Huang, Yifei Cao
Abstract:
Reliable semantic segmentation for mobile robots requires both accurate dense prediction and robust uncertainty estimation under distribution shift. Strong uncertainty baselines such as Monte Carlo Dropout often require repeated stochastic forward passes and are difficult to deploy on edge platforms.
We propose Energy‑Aware NECO, a single‑pass pixel‑wise out‑of‑distribution (OOD) detector for semantic segmentation. The method combines a centered NECO‑style geometric ratio computed from decoder features with a logit‑based Energy score. Both components are standardized using statistics fitted on a pure in‑distribution validation split and fused through a convex combination.
We evaluate the method on the miniMUAD subset using true pixel‑level OOD labels. The proposed hybrid score achieves an AUROC of 0.8539, outperforming NECO‑only (0.8280), Energy‑only (0.8171), and an ensemble predictive‑entropy baseline (0.8124). Additional qualitative and operating‑point analyses show that the hybrid detector improves overall ranking performance while preserving the efficiency advantages of a single‑pass design.
Code is available at https://github.com/boyuan‑zhangx/Energy‑Aware_NECO
Authors:Dario Pisanti, Georgios Georgakis
Abstract:
Aerial navigation on Mars requires vision‑based pipelines that are robust to the diverse illumination conditions and terrain morphology of the Martian surface. A key bottleneck for training and evaluating such methods is the scarcity of large‑scale, annotated aerial datasets. We present MARTIAN, an open‑source Blender‑based rendering framework that leverages real HiRISE orbital map products to synthesize realistic aerial views of the Martian terrain under controllable lighting conditions and at varying altitudes. MARTIAN generates observations with accurate pose annotations, directly addressing the scarcity of training data for vision‑based navigation on Mars. The framework has been validated through its deployment in concurrent work on map‑based localization systems for Ingenuity and future Mars rotorcraft, where synthetically trained deep image matchers were successfully evaluated on real Mars imagery. MARTIAN is publicly available at: https://github.com/nasa‑jpl/martian.
Authors:Yilun Qiu, Jiahe Wang, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Chun Yuan
Abstract:
Cross‑Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence distributed across multiple videos. Current Multimodal Large Language Models (MLLMs) often struggle with CVR, as simple single‑pass strategies encode multiple videos into a shared compressed context, potentially obscuring rare but critical evidence. In this paper, we propose AgentCVR, a multi‑agent framework that treats CVR as an active evidence‑acquisition task. AgentCVR employs a Master Agent to iteratively coordinate specialized Visual and Audio Agents for targeted evidence extraction. To ensure efficient training, we introduce Script‑Simulated RL, which optimizes the agent's policy with LLM‑generated semantic scripts and a lightweight text‑based simulator, bypassing costly multimodal inference during online exploration. Experimental results on a comprehensive CVR benchmark show that AgentCVR outperforms single‑pass baselines and achieves comparable performance to state‑of‑the‑art closed‑source systems, particularly in complex cross‑video alignment and localization. To ensure reproducibility, our code is available at https://github.com/wang‑jh24/AgentCVR.
Authors:NamGyu Jung, Chang Choi
Abstract:
In scene graph generation, a central challenge is modeling polysemous predicates whose meanings shift across contexts. Prior approaches address this issue by decomposing predicates into multiple static prototypes or retrieving semantically similar exemplars. However, these strategies keep predicate representations static and cannot reorganize semantics to reflect image‑specific evidence, leading to systematic confusions in ambiguous contexts. We propose AlignG, which learns context‑conditioned predicate semantics via prototype feedback. AlignG infers context‑conditioned predicate semantics from the relation candidates within each image and feeds the adapted semantics back to recalibrate relation representations. The learning objective anchors this adaptation to global semantic centers, preventing semantic drift while still allowing selective reorganization when the scene provides consistent relational cues. Experiments on VG‑150 and GQA‑200 show consistent improvements over state‑of‑the‑art baselines, with F@100 improvements of +1.4 on VG‑150 and +2.7 on GQA‑200 under SGDet. We further visualize per‑image prototype similarity shifts and observe coherent context‑dependent reorganization where prototypes selectively merge or separate predicates according to scene evidence. The code is available at https://github.com/Namgyu97/AlignG‑SGG.pytorch.
Authors:Shizhe Zhou, Bohan Jia, Kai Wu, Yan Shen, Tongyun Li, Yuyang Wu, Shaohui Lin
Abstract:
While multimodal large language models (MLLMs) have achieved rapid progress in vision‑language understanding, they remain prone to multimodal hallucinations, producing responses that are inconsistent with the visual input. Existing benchmarks predominantly focus on detecting hallucination outcomes rather than evaluating the underlying causes of these failures. Moreover, many benchmarks rely on simplistic scenarios and limited evaluation formats that no longer challenge state‑of‑the‑art models. To address these limitations, we introduce ReactBench, a cause‑driven hallucination benchmark featuring multiple tasks and an exam‑style evaluation format. By generating adversarial images and hallucination‑inducing queries, ReactBench introduces four targeted tasks: Relational Erasure, Counterfactual Attribute, Alteration Tracing, and Dense Counting. These tasks systematically expose co‑occurrence bias, language priors, cross‑image comparative perception deficiencies, and fine‑grained perceptual bottlenecks. Beyond standard accuracy‑based evaluation, we leverage Chain‑of‑Thought reasoning to identify fine‑grained sub‑causes of hallucination within each task. Extensive evaluations reveal that current MLLMs remain notably vulnerable to cause‑specific hallucination triggers, demonstrating the value of ReactBench as a systematic and interpretable testbed for diagnosing and improving multimodal model robustness. The project page is available at https://reactbench.github.io/.
Authors:Yanyan Chen, Ruigang Fu, Yu Song, Ping Zhong
Abstract:
Severe image degradation under low‑light nighttime conditions constitutes a core bottleneck preventing all‑day applications for UAV‑based single object tracking. Existing image enhancement methods often struggle to distinguish between target and background regions, which can easily lead to amplified background noise or compromise target features. To overcome this limitation, we propose TAE, a target‑aware low‑light enhancement framework tailored for nighttime object tracking. Guided explicitly by weak supervisory signals from tracking bounding boxes, the framework performs region‑aware enhancement to ensure operations focus on the target area. It further adopts an adaptive RGB multi‑curve fusion mechanism to achieve refined modeling and adaptive adjustment across different regions. To facilitate research in this domain, we also contribute DarkSOT, a new benchmark for nighttime UAV tracking, comprising 268 sequences across 9 target categories. Experimental results on the DarkSOT and UAVDark135 demonstrate that TAE significantly improves tracking performance in low‑light nighttime scenarios, exhibiting strong robustness and generalization. The DarkSOT dataset is available at https://github.com/Fu0511/DarkSOT‑Dataset.
Authors:Zekang Zhang, Guangyu Gao, Youyun Tang, ChengJing Wu, Xiaochao Qu, Chi Harold Liu, Jianbo Jiao, Yunchao Wei, Luoqi Liu, Ting Liu
Abstract:
LLM‑conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we identify a persistent failure mode in current propose‑then‑select pipelines. Although high‑quality mask candidates are often generated, the final prediction may fail to match the given linguistic condition. This failure arises because language semantics are typically used as static prompts or post‑hoc matching signals, rather than participating in the iterative mask generation process. Through systematic analysis, we show that many errors stem from semantic misalignment rather than poor mask quality. To address this issue, we propose FlowSeg, which introduces dynamic semantic guidance via a bidirectional semantic flow between intermediate decoding states and LLM‑derived condition embeddings throughout the generation process. Language conditions actively guide mask refinement at each stage, while condition embeddings are progressively updated by emerging visual evidence. This design yields semantically grounded mask representations and visually aligned language conditions, enabling more reliable matching. We further incorporate a lightweight boundary‑aware refinement to selectively enhance uncertain regions without perturbing confident interiors. Extensive experiments on referring expression segmentation and reasoning segmentation tasks demonstrate that FlowSeg consistently improves language‑mask alignment and achieves state‑of‑the‑art performance. Project page: https://zkzhang98.github.io/FlowSeg_page
Authors:Zehao Wang, Guanglei Yang, Yihan Zeng, Hang Xu, Hongzhi Zhang, Wangmeng Zuo, Chun-Mei Feng
Abstract:
Federated fine‑tuning of foundation models with Low‑Rank Adaptation (LoRA) provides an efficient solution for reducing communication and computation costs while preserving data locality. However, the direct combination of FedAvg and LoRA suffers from three key issues: limited update space, which restricts the model's effective learning capacity; inter‑round state mismatch, which disrupts cross‑round local optimization continuity; and a client‑agnostic starting state, which slows local convergence on clients. Although recent methods mitigate the limited update space issue by merging LoRA updates into the backbone across communication rounds, inter‑round state mismatch and the client‑agnostic starting state remain insufficiently addressed. To address these issues, we propose FedSmoothLoRA, a federated LoRA tuning framework that preserves the enlarged update space, improves cross‑round local optimization continuity, and provides a client‑aware starting state for local training. At each communication round, FedSmoothLoRA constructs the local LoRA initialization using two matrices: a Round‑Matching matrix that preserves cross‑round local state continuity, and a Gradient‑Aligned matrix that provides client‑specific optimization guidance from gradient signals estimated on local data. Together, these designs enable smoother and faster convergence. Extensive experiments on image classification and natural language generation tasks demonstrate that FedSmoothLoRA consistently outperforms existing federated LoRA tuning methods. Code: https://github.com/wangzehao0704/FedSmoothLoRA
Authors:Tianpeng Bu, Xin Liu, Qihua Chen, Hao Jiang, Shurui Li, Hongtao Duan, Lu Jiang, Lulu Hu, Bin Yang, Minying Zhang
Abstract:
While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real‑world deployment. To bridge this gap at both the evaluation and data levels, we introduce GUI‑RobustEval and propose Robustness‑driven Trajectory Synthesis. GUI‑RobustEval contains 1,216 executable test cases that systematically measure error recovery capabilities across a broad and realistic spectrum of error modes. At the data level, RoTS is a scalable synthesis framework that creates 800k high‑quality data via a tree‑based pipeline that proactively discovers diverse error modes and synthesizes corresponding recovery steps. Our two models, RoTS‑7B and RoTS‑32B, fine‑tuned on our dataset, both demonstrate significant gains on GUI‑RobustEval and traditional GUI benchmarks. Notably, RoTS‑32B achieves state‑of‑the‑art performance on OSWorld, with a 47.4% success rate and a 33.8% All‑Pass@4 score, suggesting that improved long‑horizon error recovery ability contributes to both robustness and overall performance. Our code is available at https://github.com/AlibabaResearch/RoTS.
Authors:Sanghyun Jo, Seo Jin Lee, Seohyung Hong, Yoorim Gang, Hyeongsub Kim, Hyungseok Seo, Kyungsu Kim
Abstract:
Cell instance segmentation models trained on cell‑specific datasets suffer severe performance drops on out‑of‑distribution cell types, while interactive foundation models overcome this through per‑instance prompting at a cost that is prohibitively expensive for histopathology images containing hundreds to thousands of densely packed instances. We introduce Group Prompting, a new paradigm that shifts interactive segmentation from per‑instance O(N) to per‑type O(T), where a single click per cell type suffices to segment all instances of that type. Our key observation is that the frozen image encoder of the Segment Anything Model (SAM) already clusters same‑type cells in its feature space before any prompt is given. Exploiting this property, we propose Chain‑of‑Prompts (CoP), a training‑free framework that recursively expands a single user click by (1) identifying reliable same‑type locations through non‑parametric gating of multi‑scale encoder features, and (2) selecting the most spatially distant reliable point as the next prompt to maximize coverage. On three cell‑type‑annotated benchmarks, CoP with one click per type retains over 90% of per‑instance performance and surpasses fully‑supervised methods without any additional training. On four morphologically homogeneous benchmarks, a single click retains over 99%. Project Page: https://shjo‑april.github.io/Chain‑of‑Prompts/
Authors:Hesam Asadollahzadeh, Feng Liu, Christopher Leckie, Sarah M. Erfani
Abstract:
Mainstream strategies for finetuning pretrained multimodal models often degrade out‑of‑distribution (OOD) robustness, a phenomenon known as catastrophic forgetting. In this paper, we develop a theoretical framework for multimodal contrastive finetuning, yielding closed‑form solutions and a geometric decomposition for each strategy. This framework shows that self‑distillation is more effective than other regularization approaches to retain the knowledge of the pretrained model. Our analysis reveals a largely overlooked limitation: standard Exponential Moving Average (EMA) teachers, widely used in robust finetuning, suffer from collapse. To solve this, we prove that a Weighted Moving Average (WMA) teacher maintains a persistent regularizing force over finite horizons and yields bias‑free convergence in the task subspace while preserving orthogonal knowledge. These insights motivate TRACER (Trajectory‑Robust Anchoring for Contrastive Encoder Regularization), which combines contrastive learning with WMA‑guided multi‑perspective distillation. Extensive experiments on CLIP finetuning demonstrate consistent OOD accuracy and calibration gains across three backbone architectures, and comprehensive ablations confirm that TRACER is both principled and robust to hyperparameter choices. Code is available at [https://github.com/HesamAsad/TRACER](https://github.com/HesamAsad/TRACER).
Authors:Fumiya Tatematsu, Fumihiko Takahashi
Abstract:
We present the 1st‑place solution to the ACCIDENT challenge at the CVPR 2026 AUTOPILOT Workshop, which asks for zero‑shot prediction of accident timing, impact centroid, and collision type from CCTV footage. On a frozen Qwen3‑VL‑32B‑Instruct checkpoint we build a three‑stage pipeline (full‑video joint prediction, time refinement, and single‑frame grounding of the impact centroid), run the same pipeline a second time on a 235B Mixture‑of‑Experts sibling, blend the two outputs 9:1, and finally snap each predicted point onto the nearest vehicle detection. The final system reaches Public LB 0.55469 / Private LB 0.57080, roughly +0.21 over the strongest host baseline (Molmo‑7B, 0.358) and wins the challenge. We ablate each component, report the negative results that shaped the final design, and release the code at https://github.com/fuumin621/cvpr2026‑accident‑1st‑place‑solution.
Authors:Xinyu Liu, Darryl Cherian Jacob, Yang Zhou, Jindong Wang, Pan He
Abstract:
Recent reinforcement learning (RL) post‑training approaches primarily optimize the final output policy using sparse outcome‑level rewards, while largely overlooking predictive signals encoded in intermediate representations. In this paper, we introduce a new paradigm called on‑policy internal self‑distillation and propose the OISD framework, which improves reasoning by transferring on‑policy predictive signals from the final layer to intermediate representations. During rollout and Group Relative Policy Optimization (GRPO) optimization, the final layer acts as both the policy and a detached internal teacher for selected intermediate layers, which are guided to align with it through two complementary mechanisms: logit alignment, which transfers high‑level reasoning behaviors (how to think), and attention alignment, which enforces consistent attention patterns (where to look) from the final layer to the selected intermediate layer, both without requiring external privileged information. Our OISD, together with GRPO, employs signed advantage‑weighted Jensen‑‑Shannon alignment to distill informative intermediate representations while preserving policy consistency under a unified acting policy. Experimental results demonstrate the effectiveness of OISD, with substantial and consistent improvements over strong reasoning RL baselines across four mathematical reasoning tasks. The code will be released at https://github.com/THE‑MALT‑LAB/OISD
Authors:Yurong Gao, Zicheng Zhang, Congying Han, Tiande Guo, Xinmin Qiu
Abstract:
Diffusion bridge models offer a powerful framework for connecting two data distributions, such as in image restoration and translation. Many existing methods learn this bridge by mimicking the score‑matching formulation of standard diffusion models. In this work, we find that this way leads to an anomalous underfitting phenomenon near the target endpoint, as the process approaches the target distribution (t \to 0). This underfitting, characterized by significant drift in the predicted variance and direction, results from an excessively large discrepancy in noise levels between the network's input and its regression target.To resolve this issue, we propose the Noise‑Aligned Diffusion Bridge (NADB).Our approach reformulates the diffusion bridge by first employing a mean network to provide a cleaner conditional target, and then introducing a novel, noise‑aligned mapping relationship. This new formulation resolves the noise mismatch and corrects the underfitting near the target endpoint. Experimental validation across multiple image restoration and image translation tasks demonstrates the effectiveness of our approach. Code is available at https://github.com/gyr02/NADB.
Authors:Haiwen Diao, Jiahao Wang, Penghao Wu, Yuhao Dong, Yuwei Niu, Yue Zhu, Zhongang Cai, Weichen Fan, Linjun Dai, Silei Wu, Xuanyu Zheng, Mingxuan Li, Yuanhan Zhang, Bo Li, Hanming Deng, Huchuan Lu, Quan Wang, Lei Yang, Lewei Lu, Dahua Lin, Ziwei Liu
Abstract:
Current vision‑language models (VLMs) typically stitch together separate image encoders and language decoders via multi‑stage alignment, a modular framework that inevitably fragments pixel‑level signals across frames and scatters early pixel‑word interactions. In parallel, native VLMs, despite impressive performance on single images, remain largely unexplored in multi‑image, video understanding, and spatial intelligence. Hence, we introduce NEO‑ov, a native foundation model that learns cross‑frame and pixel‑word correspondence end‑to‑end, without any external encoders, auxiliary adapters, or post‑hoc fusion. By eliminating module boundaries entirely, NEO‑ov enables fine‑grained and unified spatiotemporal modeling to emerge natively inside the model. Notably, NEO‑ov largely narrows the gap to modular counterparts while excelling at fine‑grained visual perception, validating that native "one‑vision" architectures are not only feasible but competitive at scale. Beyond empirical performance, we unveil systematic architectural analyses and detailed training recipes to facilitate subsequent native multimodal modeling. Our code and models are publicly available at: https://github.com/EvolvingLMMs‑Lab/NEO.
Authors:Zhen-Hao Xie, Yu-Cheng Shi, Da-Wei Zhou
Abstract:
Class‑Incremental Learning (CIL) is important in building real‑world learning systems. In CLIP‑based CIL, the model performs classification by comparing similarity between visual and textual embeddings obtained from template prompts, e.g., ``a photo of a [CLASS]''. This seemingly monolithic matching process can be decomposed into two conceptually distinct stages: attribute extraction and attribute aggregation. For example, a model may recognize cat using attributes such as fur texture and whiskers. When learning a new class like car, the model must extract additional attributes like wheels and adjust how they are aggregated in the shared representation space. However, since only data from the current task is available, incremental updates can bias both attribute extraction and aggregation toward new classes, leading to catastrophic forgetting. Therefore, we propose AREA for attribute extraction and aggregation in CLIP‑based CIL. To stabilize extraction, we anchor class‑level visual and textual attributes on the hyperspherical embedding space via principal geodesic analysis. To stabilize aggregation, we learn lightweight task‑specific experts with scoring and residual refinement, regularized by a variational information bottleneck objective. During inference, we perform routing over task attribute manifolds via optimal transport for more concise prediction. Experiments show that AREA consistently outperforms SOTA methods. Code is available at https://github.com/LAMDA‑CL/ICML2026‑AREA.
Authors:Viet Nguyen, Thao Nguyen, Vishal M. Patel, Yuheng Li
Abstract:
Long‑term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text‑centric. Even when images are included, the user‑specific information needed for later questions is typically recoverable from text alone, and most memory systems reduce image turns to generic captions. Yet images often carry personal information that text rarely states ‑‑ both explicit evidence, such as recurring user‑associated entities, and implicit evidence, such as latent user facts inferred from visual or multimodal cues. We introduce a benchmark for personal visual memory that targets both forms of evidence, and propose VisualMem, a hybrid visual‑‑text architecture that augments a text‑memory backend with a structured personal visual memory module. Rather than collapsing images into captions, VisualMem uses conversational context to resolve identity, ownership, and durable user facts. Experiments show that VisualMem substantially outperforms prior memory systems on our benchmark while remaining competitive on standard text‑memory benchmarks, indicating that personal visual memory is a distinct and important component of long‑term memory for personalized AI agents.
Authors:Xinchen Zhang, Bowei Liu, Jiale Liu, Chufan Shi, Yizhen Zhang, Junhong Liu, Youliang Zhang, Zhiheng Li, Yujiu Yang, Ling Yang
Abstract:
Visual outcomes are increasingly central to multimodal large language models, making reliable and fine‑grained verification essential for scaling generalist foundation models. In this work, we investigate multimodal meta‑verification, which leverages verifier‑generated rationales rather than decision‑only signals, and explore how to effectively incorporate meta‑verification feedback into multimodal verifier training. We identify two key findings. First, symbolic verifier outputs (e.g., bounding boxes) outperform textual explanations as meta‑verification rationales, enabling efficient rule‑based reinforcement learning rewards while avoiding reliance on model‑based rewards from auxiliary judge models. Second, decoupling reinforcement learning objectives for binary judgment and meta‑verification substantially outperforms joint reward optimization, due to intrinsic differences in output structure and learning dynamics. Based on these insights, we train OmniVerifier‑M1, a generalist visual verifier leveraging symbolic meta‑verification and decoupled reinforcement learning. OmniVerifier‑M1 provides robust verification and fine‑grained error localization, and further enables M1‑TTS, a verifier‑driven agentic generation system achieving dynamic region‑level self‑correction. This approach paves the way for more reliable, interpretable, and fine‑grained multimodal verification, supporting safer and more controllable foundation model deployment.
Authors:Xinyu Wang, Mingze Li, Sicheng Lyu, Dongxiu Liu, Kaicheng Yang, Ziyu Zhao, Yufei Cui, Xiao-Wen Chang, Peng Lu
Abstract:
Vision‑Language‑Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi‑billion‑parameter backbones and diffusion‑based action heads make on‑device deployment prohibitively expensive. Prior quantization efforts offer only partial solutions, compressing the LLM backbone while leaving the DiT action head at full precision, or resorting to mixed‑precision schemes, driven by the belief that uniformly quantizing the action head is inherently unstable. We challenge this assumption with Omega‑QVLA, the first training‑free post‑training quantization framework that compresses both the language backbone and the entire diffusion action head of a VLA model to a uniform W4A4 precision, eliminating the need for mixed‑precision allocation. Omega‑QVLA combines a composite SVD‑Hadamard rotation that equalizes per‑channel weight energy while diffusing residual activation outliers with per‑step DiT activation scaling quantization that absorbs dynamic‑range drift across denoising steps. On LIBERO, Omega‑QVLA compresses Pi 0.5 and GR00T N1.5 to W4A4 with 98.0% and 87.8% task success rates, matching or exceeding their FP16 references of 97.1% and 87.0%, while reducing the static memory footprint by 71.3%. Real‑world manipulation experiments further confirm smooth, accurate manipulation where prior methods fail. Code is available at https://github.com/UCMP13753/Omega‑QVLA.
Authors:Thomas Vitry, Kieran Edgeworth, Stefan Wermter, Jae Hee Lee
Abstract:
Vision classifiers can exploit spurious correlations, achieving high in‑distribution accuracy yet failing under distribution shift. Existing approaches to bias mitigation and analysis often depend on curated datasets, spurious‑attribute or group labels, or retraining, which may be infeasible once a model is deployed or the relevant bias is unknown. We present a bias‑label‑free, post‑hoc method for identifying spurious concepts in frozen vision models, relying only on standard class labels from a held‑out audit dataset. For each target class, we collect patches from inputs predicted as that class and apply non‑negative matrix factorization to intermediate activations to obtain a bank of interpretable concept vectors. Candidate concepts are then ranked with a bias estimator derived from their interaction with backpropagated gradients on misclassified examples: bias concepts tend to get activated when correcting false negatives and suppressed when correcting false positives. On Colored MNIST and Waterbirds the method recovers concepts aligned with the known spurious cue, and on CelebA it surfaces decision‑relevant directions that only partially coincide with the annotated gender attribute; suppressing the top‑ranked concepts at inference time improves worst‑group accuracy by up to 17.9 percentage points on Waterbirds and 10.4 on CelebA without any retraining or parameter updates. Our method identifies decision‑relevant spurious directions that need not coincide with annotated ones, providing both an interpretable auditing tool and an actionable debiasing handle for frozen vision models. Code is available at https://github.com/vitryt/label‑free‑bias‑identification.
Authors:Hongyu Wen, Jia Deng
Abstract:
Transparent objects are common in daily life, and it is important to understand their multilayer depth, including the transparent surface and the objects behind it. Existing methods for multilayer depth typically extend single‑layer prediction. They define layers by the front‑to‑back ordering of 3D points and predict the layers sequentially. However, as layered geometry can admit multiple valid groupings of 3D points into layers, a predefined grouping strategy is inherently restrictive. In this work, we propose SeeGroup, a multi‑layer depth estimation method that avoids imposing a predefined grouping and allows the model itself to adaptively assign surfaces to depth maps. We formulate per‑pixel multi‑layer depth as a point process, treating depth layers as unordered events along each camera ray. This induces a permutation‑invariant likelihood over the observed depth layers, yielding a loss that naturally supports arbitrary layer groupings. Experiments demonstrate that our method significantly advances the state of the art of multi‑layer depth estimation, improving quadruplet relative depth accuracy on LayeredDepth benchmark from 61.34% to 70.09%. Code is available at https://github.com/princeton‑vl/SeeGroup.
Authors:Peiyuan He, Hainuo Wang, Hengxing Liu, Mingjia Li, Xiaojie Guo
Abstract:
Self‑supervised low‑light image enhancement (LLIE) is highly appealing as it eliminates the reliance on external paired data. However, the lack of external references causes networks to struggle with decoupling entangled illumination, delicate textures, and amplified noise. To resolve this challenge, we propose an Internally Referenced LLIE framework that extracts reliable physical and structural references from the degraded input image itself. First, we introduce a local exposure‑simulated scheme to extract a low‑frequency pseudo ground‑truth. This serves as an internal physical reference to guide global illumination estimation and correct color casts. Second, we propose a dual‑domain preservation strategy with spatial and spectral constraints to construct internal structural references. Specifically, an Illumination‑Aligned Perceptual loss preserves global structures under illumination shifts, while a Shift‑Invariant Spectral Correlation loss captures fine‑grained local structures and suppresses high‑frequency noise. Finally, we propose a Gain‑Adaptive Feature Modulation (GAFM) mechanism to address highly spatially‑variant residual noise. By transforming the self‑estimated illumination map into an internal spatial gain prior, GAFM dynamically guides a blind‑spot network for spatially‑aware denoising. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance, delivering superior noise suppression and textural fidelity. Code will be publicly released at https://visonj.github.io/IRLE/.
Authors:Yang Gao, Wuyang Li, Po-Chien Luan, Alexandre Alahi
Abstract:
Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human‑centric, nonrigid agents. However, existing weakly supervised occupancy prediction frameworks predominantly assume rigid‑body motion and rely on simple frame‑to‑frame offsets, limiting their ability to capture fine‑grained deformations and maintain temporal coherence. To address this issue, we propose DeGO, a deformable Gaussian occupancy framework that unifies decoupled Gaussian deformation with factorized 4D foundation‑model distillation. DeGO disentangles rigid and nonrigid motion, enabling each Gaussian primitive to evolve through both deformation and offset‑based updates. In parallel, a factorized 4D distillation strategy transfers cross‑camera and cross‑frame knowledge from the VGGT foundation model, producing foundation‑aligned features that enhance temporal consistency. Experiments on the Occ3D‑NuScenes benchmark demonstrate that our method achieves state‑of‑the‑art performance under weak supervision, delivering 13.5% gains on human‑centric instances and 10.9% overall improvements. These results highlight the effectiveness of deformation‑aware and foundation‑guided occupancy modeling for dynamic scene understanding. The code is publicly available: https://github.com/vita‑epfl/DeGO
Authors:Ruowen Zhao, Bangguo Li, Zuyan Liu, Yinan Liang, Junliang Ye, Fangfu Liu, Diankun Wu, Zhengyi Wang, Xumin Yu, Yongming Rao, Han Hu, Jun Zhu
Abstract:
Embodied Vision‑Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision‑Language‑Action frameworks. However, a significant gap remains between the high‑level semantic focus of standard text‑guided pre‑training paradigms and the low‑level spatial and physical knowledge critical for execution in embodied environments. In this paper, we introduce GEM, a Generative‑supervised Embodied vision‑language Model designed to bridge this divide. We propose integrating a depth map generation task directly into the VLM pre‑training phase. By training this generative objective jointly with the main model, we observe substantial improvements in embodied intelligence, significantly enhancing both semantic understanding and physical operation capabilities. To support this paradigm, we curate and release GEM‑4M, a comprehensive large‑scale dataset featuring a mixture of grounding, reasoning, and planning data paired with high‑quality depth supervision. Extensive experiments demonstrate that GEM achieves state‑of‑the‑art results across diverse embodied benchmarks. Furthermore, our deployed action model, GEM‑VLA, exhibits vastly superior task execution abilities in both simulation environments and real‑world evaluations. Code, models, and datasets are available at https://zhaorw02.github.io/GEM/
Authors:Changxuan Li, Nadine Berner, Nassir Navab, Federico Tombari, Stefano Gasperini
Abstract:
Self‑supervised depth estimation from monocular sequences relies on the joint learning of a depth and a pose network. Despite abundant research done to improve the depth network, efforts on the pose remain limited. In this context, even when depth is estimated up to scale, we highlight the importance of the alignment between the scene scales estimated by the pose and depth nets. Then, we introduce SA4Depth, an approach to improve this alignment and boost the depth predictions while keeping the inference time unchanged. Our proposed method uses the depth estimated during training to reproject learnable visual features across consecutive frames and refine the pose estimates by reducing feature alignment residuals. With our method, the estimated scene scales by the separate depth and pose networks are aligned, and the prediction scale consistency is improved across different sequences. Our differentiable refinement integrates seamlessly into existing self‑supervised pipelines and substantially improves their depth estimates. We demonstrate this with extensive experiments both outdoors and indoors on KITTI, Cityscapes, and NYUv2. Additionally, results on KITTI Odometry confirm the effectiveness of our pose refinement. Our code is available at https://github.com/Runningchauncey/SA4Depth .
Authors:Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro
Abstract:
Existing Visual Speech Recognition (VSR) systems commonly rely on left‑to‑right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose DLLM‑VSR, to the best of our knowledge, the first Diffusion Large Language Model (DLLM)‑based VSR framework, formulating transcription as iterative masked denoising with flexible‑order decoding. With confidence‑based unmasking, DLLM‑VSR commits high‑confidence positions early and uses the committed tokens as bidirectional context to refine ambiguous ones. To adapt DLLMs to VSR, we introduce a two‑stage masked‑denoising training strategy that separates visual‑to‑text content alignment from length modeling. We further observe a performance gap with oracle‑length decoding, which assumes access to the true transcript length, indicating that reducing target‑length uncertainty can improve DLLM‑based VSR. To reduce this gap, we develop length‑guided candidate decoding, which uses video duration to construct plausible transcript‑length hypotheses, decodes under multiple hypotheses, and reranks candidates using length plausibility and decoding confidence. The proposed method achieves a state‑of‑the‑art WER of 19.5% on LRS3 using only its labeled training data.
Authors:Peng Cui, Jiahao Zhang, Lijie Hu
Abstract:
While Contrastive Learning (CL) has revolutionized self‑supervised representation learning, its latent representations remain highly entangled and opaque, limiting their interpretability in safety‑critical applications. We identify that a fundamental cause of this entanglement is the reliance on deterministic similarity measures, which treat all feature dimensions equally. In compositional scenes, this creates an Optimization Conflict: common background features, such as, "blue sky", are encouraged to align in positive pairs but simultaneously repelled in negative pairs, causing gradient oscillations that hinder precise semantic disentanglement. To address this, we propose BayesNCL (Bayesian Gated Non‑Negative Contrastive Learning). Unlike standard approaches, BayesNCL introduces a probabilistic gating mechanism that dynamically filters out task‑irrelevant, high‑frequency common features while selectively retaining discriminative semantics. By formalizing feature selection as a variational inference problem with a sparse Bernoulli prior, our method effectively resolves the optimization conflict. Empirical experimental results on Imagenet‑100 demonstrate that BayesNCL achieves a remarkable 142.1% improvement in semantic consistency compared to state‑of‑the‑art baselines, yielding highly interpretable representations without compromising downstream task performance. Code is available at https://github.com/Cui‑Peng‑624/BayesNCL.
Authors:Bohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng, Yanlin Weng, Kun Zhou
Abstract:
Real‑time synthesis of high‑fidelity 3D character motion from audio is a pivotal component for next‑generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio sequences or are constrained to specific domains, rarely handling both speech and music effectively. In this paper, we introduce a novel framework designed to generate continuous, coherent full‑body motion from streaming speech and music with low latency. Central to our approach is a unified streaming architecture capable of synthesizing continuous motion from incremental audio inputs. We employ a robust training strategy that enforces strong audio dependency, allowing the model to seamlessly generalize across conversational speech and rhythmic music without requiring explicit domain labels or mode switching. Additionally, we explored Reinforcement Learning to refine the quality of online generation. Furthermore, we bridge reactive animation with intent‑driven behavior via a tool‑call interface that allows upstream Large Language Models to inject explicit semantic control. By combining this controllability with stream audio‑driven synthesis, our framework serves as a plug‑and‑play solution for transforming voice agents into interactive humanoid avatars. Extensive experiments demonstrate that our method outperforms state‑of‑the‑art realtime baselines in motion quality and synchronization while maintaining the flexibility required for live deployment. Our code, pre‑trained models, and videos are available at https://robinwitch.github.io/EchoAvatar‑Page.
Authors:Leonhard Sommer, Emil Akopyan, Adam Kortylewski
Abstract:
Estimating the 9D pose of everyday objects from a single real‑world image remains challenging. This is largely due to the lack of large‑scale supervision. Most existing datasets either rely heavily on synthetic renderings or provide limited coverage of real‑world objects: the largest real‑world 9D pose dataset to date contains only 17K annotated objects across 9 categories. We address this gap with Every9D‑21M, a dataset of 9D pose annotations for 21.8M real‑world images from 109K object‑ centric videos spanning 700 everyday object categories ‑ two orders of magnitude larger than prior real‑world 9D pose benchmarks in both image and category count. To achieve this scale, we leverage object‑centric videos by reconstructing object‑ level point clouds via multi‑view geometry and aligning similar instances into a shared canonical coordinate frame. Canonical poses are manually annotated for only a small set of reference objects (fewer than 0.01% of all images) and propagated to the remaining instances via cross‑instance alignment. All propagated canonical poses are then verified from multiple viewpoints. We further introduce cross‑category orientation rules that induce category‑level symmetries, enabling symmetry‑aware evaluation. Beyond establishing dedicated training and evaluation splits as a benchmark for 9D pose foundation models, we show that training on Every9D‑21M improves performance on ImageNet3D and PASCAL3D+, and generalizes to HANDAL substantially better than training on ImageNet3D. Data and code are available at https://github.com/GenIntel/Every9D.
Authors:Leiyue Zhao, Tianyu Shi, Daniel Reisenbuchler, Xinzi He, Junchao Zhu, Tianyuan Yao, Yuechen Yang, Yanfan Zhu, Junlin Guo, Gelei Xu, Haichun Yang, Yuankai Huo, Mert R. Sabuncu, Yihe Yang, Ruining Deng
Abstract:
Instance‑level quantification of kidney functional units is essential for morphometric analysis, yet most publicly available pathology datasets provide only semantic segmentation annotations, where adjacent structures of the same class are merged into single regions. This prevents reliable instance‑level analysis and limits downstream quantitative studies. Existing heuristic post‑processing methods often yield suboptimal instance separation, particularly in crowded and adherent regions, while deep learning‑based instance segmentation approaches typically require intensive instance‑level annotations that are costly and labor‑intensive to obtain. We propose MORI‑Seg, a deep learning framework that enables instance segmentation without requiring instance‑level annotations. Instead of heuristic splitting or instance supervision, MORI‑Seg learns morphology‑aware geometric representations directly from semantic masks by jointly modeling object‑centric distance fields and boundary‑band representations to encode interior structure and contact interfaces. A class‑conditioned feature disentanglement module further promotes intra‑instance coherence and inter‑instance separation. Under semantic‑only supervision, MORI‑Seg decomposes connected semantic regions into distinct instance masks in an end‑to‑end manner. Experiments demonstrate improved instance separation accuracy and more reliable morphometric quantification compared with classical post‑processing pipelines and representative semantic‑to‑instance learning approaches. The official implementation is publicly available at https://github.com/ddrrnn123/MORI‑Seg.
Authors:Leonhard Sommer, Artur Jesslen, Basavaraj Sunagad, Adam Kortylewski
Abstract:
Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category‑level pose estimation, current representations fail to capture the fine‑grained semantics needed for reasoning about object parts, functions, and interactions. In this work, we study category‑level 3D correspondence in camera space ‑‑ predicting, from a single image, 3D locations that remain consistent across instances within a category ‑‑ and show that it can emerge without explicit correspondence supervision by learning a shared morphable object prior. To enable research in this direction, we introduce HouseCorr3D, the first large‑scale benchmark for monocular category‑level 3D correspondence with 178k images across 50 household object categories, 280 unique instances, and 3D keypoint annotations directly on CAD models. Crucially, HouseCorr3D provides amodal correspondence labels for occluded regions and explicit symmetry annotations, addressing key limitations of existing datasets. We further propose Morpheus, a method that learns morphable category‑level shape priors by disentangling canonical shape, deformation, and object pose. Through this shared canonical grounding, semantically meaningful 3D correspondences in camera space emerge implicitly. These emerging 3D correspondences set a new state of the art on HouseCorr3D, demonstrating that semantic 3D object understanding can arise without direct correspondence supervision. Data and code are publicly available at https://github.com/GenIntel/HouseCorr3D.
Authors:Rui Lin, Chuanming Wang, Huadong Ma
Abstract:
With the rapid development of pre‑training technologies, adapting large‑scale Vision‑Language Models (VLMs) for video understanding \emph\ie image‑to‑video transfer learning has become a dominant paradigm. To achieve superior performance, it raises as an effective strategy among recent advances to employ Mixture‑of‑Experts (MoE) to enhance VLMs' temporal modeling capabilities. However, conventional MoE designs suffer from expert homogenization, where all experts act as identical generalists, inefficiently learning spatio‑temporal features from undifferentiated video streams. To overcome this problem, we propose VidPrism, a novel heterogeneous temporal Mixture‑of‑Experts framework. VidPrism pioneers a division of labor by deploying functionally specialized experts, each assuming a role ranging from spatial understanding to temporal modeling. To feed these specialists appropriately, we introduce a content‑aware, multi‑rate sampling module that dynamically generates streams ranging from semantically rich to motion‑focused representations, providing specialized inputs for experts. Furthermore, a dynamic, bidirectional fusion mechanism enables synergistic information exchange between these pathways, leading to a comprehensive video representation. Extensive experiments on various video recognition benchmarks demonstrate that VidPrism achieves state‑of‑the‑art performance and effectively fosters expert specialization. Our source code is available at \hrefhttps://github.com/Lrrrr549/VidPrism.githttps://github.com/Lrrrr549/VidPrism.git.
Authors:Shurui Xu, Siqi Yang, Weiping Ding, Hui Wang, Mengzhen Fan, Yuyu Sun, Shuyan Li
Abstract:
Clinical diagnosis of meniscus injuries requires radiologists to integrate volumetric MRI evidence with patient context (e.g., sex, age, BMI) and to produce structured diagnostic reports. Existing knee MRI benchmarks are typically unimodal and rely on coarse labels, limiting their ability to evaluate holistic clinical reasoning. We introduce MeniOmni, a structured multimodal benchmark for meniscus injury assessment, consisting of 746 multi‑center MRI studies with tri‑planar volumetric inputs, Clinical Priors, and expert‑annotated clinical text. MeniOmni supports two tasks: (1) fine‑grained Stoller severity grading and (2) diagnostic report generation. We further propose risk‑aware ordinal evaluation and a semantic consistency metric (Meni‑Score) to better reflect clinical relevance. Baseline experiments show that incorporating Clinical Priors improves grading performance and reduces severe errors, highlighting the value of multimodal context for safer assessment. Code and data are available at https://github.com/ShuruiXu/MeniOmni.
Authors:Haozhan Shen, Tiancheng Zhao, Kangjia Zhao, Jianwei Yin
Abstract:
Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre‑training schemes are now widely used as foundation backbones: Vision‑Language Models (VLMs), which use language supervision to align visual observations with semantic concepts, and Video Generation Models (VGMs), which learn from temporally evolving visual worlds. However, it still remains unclear which pre‑training scheme provides a better representation substrate for spatial intelligence. In this paper, we present the first systematic frozen‑feature probing study of VLMs and VGMs across three representative axes of spatial intelligence: semantic tagging, instance grouping, and 3D geometry prediction. Using the lightweight probe, our framework enables a controlled comparison of what information is already encoded in frozen representations from two model families. Experimental results reveal a clear complementarity: VLMs are stronger at semantic tagging and instance grouping, while VGMs provide more accessible signals for dense geometry and camera motion. Moreover, a naive fusion of the two already yields a representation that excels at both geometry and semantics, suggesting a promising direction for building stronger spatial‑intelligence backbones by effectively integrating features from both model families. Our code is available at \hrefhttps://github.com/om‑ai‑lab/Probing‑VLM‑VGMhttps://github.com/om‑ai‑lab/Probing‑VLM‑VGM.
Authors:Hongtao Yang, Bineng Zhong, Qihua Liang, Yaozong Zheng, Xiantao Hu, Yuanliang Xue, Shuxiang Song
Abstract:
Given the real‑time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and degrades performance in complex scenarios. To alleviate this issue, we propose EATrack, an efficient and asymmetric UAV tracking framework centered around a teacher‑guided dual‑branch distillation strategy that enhances the feature expressiveness of the lightweight student model. Specifically, EATrack investigates two complementary perspectives of knowledge transfer: spatially focused feature‑level distillation that compensates for weakened representations by guiding the student to learn strong target representations, and prediction‑level distillation that enhances spatial localization by learning the teacher's capability for accurate target localization. Furthermore, to enhance robustness against appearance variations, we introduce a fine‑grained target‑aware distillation strategy that selectively transfers the teacher's target modeling capacity to the student. A temporal adaptation module is incorporated at inference to enhance robustness over time. Experiments on five UAV benchmarks demonstrate that EATrack achieves a favorable balance between accuracy and speed. Code: https://github.com/GXNU‑ZhongLab/EATrack
Authors:Seunghyeok Shin, Minwoo Kim, Dabin Kim, Hongki Lim
Abstract:
Diffusion posterior sampling conditions diffusion priors on measurements, but data‑consistency updates are typically scaled by hand‑tuned guidance weights and can destabilize sampling under stiff, operator‑dependent curvature. We replace scalar guidance with a per‑noise‑level damped Gauss‑‑Newton correction computed in diffusion‑state coordinates. The correction pulls likelihood gradients back through the denoiser, uses a one‑sided curvature model that avoids forward denoiser Jacobians, and applies diffusion‑calibrated rank‑one damping aligned with the denoiser residual. Each correction is solved with matrix‑free GMRES using automatic differentiation, and sampling proceeds with a variance‑preserving Langevin transition with a closed‑form drift/noise split. On FFHQ and ImageNet across inverse problems, it achieves competitive PSNR/SSIM/LPIPS while running markedly faster than most of the compared baselines; on accelerated MRI reconstruction, it achieves the best PSNR/SSIM among the compared baselines.
Authors:Hongyu Ding, Sizhuo Zhang, Ziming Xu, Jinwen Guo, Hongxiu Liu, Xingzhi Cheng, Zixuan Chen, Haifei Qi, Duo Wang, Hao Xu, Jieqi Shi, Yifan Zhang, Jing Huo, Jian Cheng, Yang Gao, Jiebo Luo
Abstract:
Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dominant approach has been to scale vision‑language‑action (VLA) foundation models on ever‑larger collections of robot trajectories. This paper argues that, for navigation specifically, generality can be obtained structurally, not only through data scale. The underlying decision structure of navigation reduces to a single Language‑Vision‑Robot Actions Translation. The language action emits semantic‑level directional command and the vision action emits a pixel‑level visual target. Both outputs lie inside the natural output manifold of pretrained multimodal large language models (MLLMs), so the task can be reasoned about by an agent rather than learned from robot data. Therefore, we present Uni‑LaViRA, a unified agentic architecture that extends the same insight to four task families (VLN‑CE, ObjectNav, EQA, and Aerial‑VLN) and to four heterogeneous real robots (Wheeled, Quadruped, Humanoid robot, and a self‑built UAV) in a zero‑shot manner. Two agent‑loop mechanisms make this unification practical. TODO List Memory (TDM) rewrites a structured checklist of pending sub‑goals at every step, reciting the unfinished items back into the agent's most recent attention window. Second Chance Backtrack (SCB) rolls the robot back to the pre‑error state and conditions the agent's next plan on the failed sub‑trajectory, turning single‑pass navigation into a self‑correcting process. With zero training effort, Uni‑LaViRA reaches 60.7% SR on VLN‑CE R2R, 51.3% on VLN‑CE RxR, 77.7% on HM3D‑v2, 60.0% on HM3D‑OVON, 54.7% on MP3D‑EQA, and 40.0% on OpenUAV, matching or even surpassing recent training navigation foundation models that consume millions of samples and thousands of GPU‑hours.
Authors:Chung-Ta Huang, Leopold Das, Jeffrey Zhou, Faizaan Siddique, Julia Seungjoo Baek, Serena Liu, Andrew Rusli, Todd Y. Zhou, Freddy Yu, Sinclair Hansen, Ziling Hu, Arnav Sharma, Mengyu Wang
Abstract:
AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always‑on sensor, the head‑mounted Inertial Measurement Unit (IMU), detects only motion primitives such as walking or standing. We push beyond motion primitives to behavioral‑level recognition, defining five categories that balance AR application need with sensor observability. To this end, we construct a 160K‑sample Ego4D dataset with a four‑tier quality assurance framework spanning 8 activity scenarios, and propose HiT‑HAR, a 703K‑parameter hierarchical model that outperforms prior head‑mounted IMU models on five‑class action and eight‑class scenario recognition. We further map the observability frontier of head‑mounted IMU through per‑class separability analysis, identifying which behavioral categories are reliably observable (Locomotion), which benefit from temporal context (Object Transfer, Task Operation), and where scenario‑dependent signal overlap poses remaining challenges. Our results indicate that architectural choices exploiting temporal context and scenario structure outperform simply scaling model size. The code and dataset are publicly available at https://github.com/Harvard‑AI‑and‑Robotics‑Lab/HiT‑HAR.
Authors:Zixiao Hu, Tianyu Li, Guoqing Wang, Wei Li, Guoguo Xin, Xun Liu, Peng Wang
Abstract:
Single‑frame atmospheric turbulence mitigation is inherently ill‑posed due to spatially varying blur coupled with non‑rigid geometric distortion. Existing end‑to‑end approaches trained on flat‑field simulations often struggle to balance texture recovery with geometric rectification. To overcome this limitation, we propose D^2Turb, a unified framework that bridges physics‑grounded simulation with explicitly decoupled restoration. First, we introduce a Depth‑Aware Turbulence Synthesis protocol that incorporates scene depth into the phase‑to‑space formulation. This generates physically consistent, depth‑dependent degradations and provides a crucial intermediate tilt supervision signal for disentangled learning. Building upon this simulation engine, D^2Turb decomposes restoration into two interactive stages: texture deblurring and geometric rectification. The texture deblurring stage employs a deblurring backbone to recover fine‑grained details while preserving geometric distortion for the subsequent rectification stage. To mitigate the information fragmentation commonly observed in cascaded designs, we further propose an Adaptive Structural Prior Injection (ASPI) mechanism that dynamically transfers deep structural representations from the deblurring module to guide dense flow prediction for spatial unwarping. Extensive experiments demonstrate that D^2Turb achieves state‑of‑the‑art performance on both synthetic and real‑world datasets, with consistent improvements in both texture recovery and geometric fidelity. Our code and pre‑trained models are publicly available at https://github.com/HertzDot222/D2Turb.
Authors:Arijit Ghosh, Aritra Bandyopadhyay, Chiranjeev Bindra, Jingfen Qiao
Abstract:
Multimodal alignment is critical for bridging the semantic gap in information retrieval. However, traditional pairwise strategies introduce a geometric blind spot: while they align anchor modalities (e.g., text) with others, they lack constraints to enforce mutual consistency between peripheral modalities (e.g., video and audio). The TRIANGLE framework addresses this by minimizing the area of modality triplets on a hypersphere to enforce holistic alignment. In this reproducibility study, we verify the robustness of this geometric objective for retrieval tasks. We confirm that TRIANGLE outperforms pairwise baselines in zero‑shot settings, achieving Recall@1 gains of up to +8.7 points, though benefits are domain‑dependent. However, we fail to reproduce the reported learning‑from‑scratch results. Analysis using a synthetic toy dataset attributes this to instability when jointly optimizing geometric alignment with Data‑Text Matching (DTM) loss. Furthermore, we find that cosine regularization primarily stabilizes text‑to‑video retrieval, and fine‑tuning with domain supervision amplifies geometric benefits but reduces cross‑dataset generalization. Our findings support the efficacy of geometric alignment while highlighting critical optimization sensitivities. Code available at https://github.com/ARIJIT00171/RE‑TRIANGLE.
Authors:Jing Hao, Siyuan Dai, Yongxin Zhang, Yuci Liang, Jiamin Wu, Jiahao Bao, Yuxuan Fan, Zanting Ye, Yanpeng Sun, Xinyu Zhang, Ming Hu, Liang Zhan, James Kit Hon Tsoi, Linlin Shen, Junjun He, Kuo Feng Hung
Abstract:
Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have produced dental AI models for specific tasks and individual imaging modalities, their isolated designs limit practical use in real‑world clinical workflows. In this paper, we present OralAgent, the first dental‑specialized AI agent that unifies multimodal reasoning, tool‑based decision‑making, and knowledge‑grounded retrieval within an end‑to‑end automated framework. It integrates 22 visual analysis tools and 368 widely‑used classical dental textbooks, enabling autonomous reasoning, planning, tool use, knowledge retrieval, and multi‑step workflow execution. Furthermore, we introduce OralCorpus, a large‑scale, high‑quality bilingual textual resource containing 134.8M tokens curated for dental retrieval‑augmented generation (RAG). To evaluate models' multidisciplinary dental knowledge, we construct OralQA‑ZH, a Chinese multiple‑choice question benchmark consisting of 798 items across eleven oral subspecialties. Extensive experiments demonstrate that OralAgent achieves state‑of‑the‑art performance on the MMOral‑Uni, MMOral‑OPG, and OralQA‑ZH benchmarks, highlighting its effectiveness, interpretability, and adaptability in real‑world clinical settings. The code and models are publicly available at https://github.com/isjinghao/OralAgent.
Authors:Jiawei Weng, Saining Zhang, Zhenxin Diao, Peishuo Li, Henghaofan Zhang, Junhao Chen, Hao Zhao
Abstract:
3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large‑scale feedforward generative paradigms, 3D AI generation remains dominated by training‑free editing pipelines. A central challenge of feedforward 3D editing lies in the lack of high‑quality paired supervision. Editable 3D assets require simultaneous preservation of geometry, multi‑view consistency, structural coherence, and localized edit controllability. Existing 3D editing datasets often rely on independently generated assets, image‑mediated reconstruction or narrow edit taxonomies, leading to inaccurate localization, weak preservation, blurred edit boundaries, and limited semantic consistency. In this work, we introduce a new perspective: scalable feedforward 3D editing should be learned from semantic‑part transformations. Based on this insight, we propose Pxform, a high‑quality 3D editing dataset with over 100K consistent before/after editing pairs across seven edit types. Instead of treating objects as unstructured shapes, our pipeline grounds edits directly in semantic 3D parts. Built upon Pxform, we further propose PartFlow, a feedforward 3D editing network that injects source‑aware latent control into pretrained 3D generative priors. PartFlow introduces mask‑aware velocity preservation and render‑space consistency supervision to jointly improve edit fidelity and source preservation, while requiring no 3D edit mask during inference. Extensive experiments demonstrate that high‑quality semantic‑part supervision substantially improves scalable 3D editing, enabling PartFlow to achieve state‑of‑the‑art performance on both geometric and appearance editing benchmarks.
Authors:Zihui Zhang, Zhixuan Sun, Yafei Yang, Jinxi Li, Jiahao Chen, Bo Yang
Abstract:
We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene‑level human annotations during training. Existing methods are typically constrained to identifying simple objects, primarily due to insufficient object priors in the learning process. In this paper, we present FoundObj, a novel framework featuring a superpoint‑based object discovery agent that incrementally merges suitable neighboring superpoints, guided by our innovative semantic and geometric reward modules. These modules synergistically leverage semantic and geometric priors from self‑supervised 2D/3D foundation models, providing complementary feedback to the object discovery agent and enabling robust identification of multi‑class objects through reinforcement learning. Extensive experiments on diverse benchmarks demonstrate that our approach consistently outperforms existing baselines. Notably, our method exhibits strong generalization in zero‑shot and long‑tail scenarios, underscoring its potential for scalable, label‑free 3D object segmentation.
Authors:Yingxin Lai, Yafei Zhou, Fucai Zhu, Siyu Zhu, Weihao Yuan
Abstract:
While rule‑based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile‑language models primarily rely on supervised or contrastive objectives, which limits their capacity to ground predictions in physical evidence or rectify misleading visual priors. Tactile reasoning introduces two modality‑specific challenges: the ordinal nature of physical attributes (e.g., hardness, roughness) and the cross‑sensor distribution shifts inherent in optical tactile hardware. In this work, we introduce TouchReason‑1M, a large‑scale multimodal dataset comprising over 1M synchronized tactile pairs across four distinct sensors, and TouchReason‑Bench, a rigorous framework for evaluating tactile perception and visual‑tactile conflict resolution. Building upon these, we propose Touch‑R1, a tactile reasoning MLLM based on Qwen2.5‑VL‑7B. Touch‑R1 is trained via a tactile‑grounded GRPO objective that combines ordinal‑aware accuracy, cross‑sensor physical consistency, structured‑format control, and an input‑side tactile grounding objective. Specifically, the tactile‑use reward assigns credit only when authentic tactile inputs yield superior correctness relative to counterfactual controls where the tactile stream is removed, shuffled, or noise‑masked. On TouchReason‑Bench, Touch‑R1‑7B outperforms Octopi‑13B by 18.4% and GPT‑4o by 24.7% on average. Its structured reasoning traces reveal emergent behaviors of probing, comparison, and revision, demonstrating that R1‑style reasoning can be effectively grounded in physical contact.
Authors:Qida Tan, Hongyu Yang, Wenchao Du
Abstract:
Appearance‑based gaze estimation always suffers from poor generalization due to limited annotated samples and insufficient dataset diversity. Leading approaches adopt weakly supervised learning to generate large‑scale pseudo‑labeled data from unconstrained real‑world scenarios, aiming to mitigate the domain shifts. In this work, we devise a simple yet effective semi‑supervised learning architecture that leverages unlabeled data to enhance domain generalization, thereby reducing reliance on labor‑intensive manual annotations. Our key insight is to impose Jacobian regularization to disentangle feature representations into discriminative subspaces dedicated to specific gaze components, such as pitch and yaw angles. We further exploit the intrinsic ordinal ranking within each subspace for contrastive learning, enabling the model to learn robust gaze representations from a small set of labeled samples and an abundance of unlabeled ones. This ultimately yields our Disentangled Subspace Contrastive Learning (DSCL) framework. Extensive experiments on multiple benchmarks verify that the proposed DSCL is plug‑and‑play, achieving competitive performance using only 20%, 10%, and even 5% of the annotated data under both in‑domain and cross‑domain evaluation settings. The public code is available at \hrefhttps://github.com/da60266/DSCLhttps://github.com/da60266/DSCL.
Authors:Yuqi Liu, Yufei Chen, Wei Fu, Xiaodong Yue, Shuo Li
Abstract:
Accurate pancreas segmentation is critical for early cancer diagnosis, where annotation scarcity necessitates Semi‑Supervised Learning (SSL). However, due to significant inter‑sample morphological variability, existing SSL methods face severe generalizability limitations under sparse supervision, leading to the Supervision Bias problem. To address this, we propose Structural Consensus‑based KAN Prototype Learning (SCKAN), which constructs the first cross‑sample structural consensus learning with Kolmogorov‑Arnold Networks (KANs), to achieve more generalizable and accurate segmentation. Specifically, SCKAN contains two key designs: Structure‑constrained Prototype Consistency Learning (SPCL), which prompts unbiased structural representation by enforcing cross‑sample consistency via prototype‑level contrastive optimization, and Consensus‑based Kolmogorov‑Arnold Fusion (CKaF), which reduces morphology‑specific bias by aggregating stable consensus and filtering sample‑wise noise via KAN's adaptive B‑spline nonlinearity. Extensive experiments on two public pancreas datasets demonstrate the effectiveness of SCKAN. Code is at https://github.com/rhodaliu17/SCKAN.
Authors:Muye Huang, Lin Wu, Lingling Zhang, Hang Yan, Zhiyuan Wang, Yumeng Fu, Zesheng Yang, Jun Liu
Abstract:
Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts, but real‑world charts are often dynamic and interactive. Key information may only appear after actions such as hovering, clicking, zooming, or dragging. Dynamic chart understanding therefore requires models to identify visible content, choose proper interactions, and reason over changing chart states. To evaluate this ability, we propose ChartAct, an interactive benchmark for dynamic chart understanding. ChartAct collects and filters 673 dynamic charts from 8 real chart websites, covers 7 common chart types, and constructs 1,440 high‑quality question‑answer samples. Each sample is instantiated in two environments, Dynamic Chart and Dashboard Chart, to evaluate dynamic chart understanding under different contexts. Based on ChartAct, we systematically evaluate 11 advanced multimodal models and GUI agents. Experimental results show that existing models still have clear limitations in dynamic chart understanding. The strongest model, Claude‑Opus‑4.7, achieves an average success rate of 84.5%, while most models remain below 60%. We also conduct detailed failure attribution and case analysis. ChartAct provides a new benchmark for studying chart understanding in real interactive environments. Codes at https://github.com/wulin‑wulin/OSWorld_Chart
Authors:Yujie Lin, Kaidi Jia, Jiayao Ma, Chengyi Yang, Jinsong Su
Abstract:
Vision‑language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work, we present the first systematic survey and robustness analysis of VLM unlearning. We provide a comprehensive taxonomy and review of existing VLM unlearning methods, together with unified evaluations under multiple prompt settings. We then propose three attack paradigms to examine whether forgotten multimodal knowledge can be reactivated through contextual prompting or downstream retraining. Extensive experiments show that many existing methods remain vulnerable under these attacks, indicating that current approaches often hide rather than fully remove target knowledge. Our study provides new insights into the robustness and limitations of current VLM unlearning methods and highlights the need for more reliable multimodal unlearning strategies. Code is available at https://github.com/XMUDeepLIT/VLM‑UnL‑Attack.
Authors:Oussama Messai, Abbass Zein-Eddine, Abdelouahid Bentamou, Mickael Picq, Nicolas Duquesne, Stéphane Puydarrieux, Yann Gavet
Abstract:
In this paper, we address the problem of detecting small, dense, and overlapping objects, a major challenge in computer vision. Our focus is on reviewing proposed methods based on deep learning supervised approaches. We provide a detailed comparison of these systems on a new dataset of more than 10k images and 120k instances, highlighting their performance, accuracy, and computational efficiency in the industrial recycling process use case. Through this comparative analysis, we identify the most reliable systems currently available and the specific challenges they are designed to tackle. Furthermore, we explore the benefits of data augmentation and synthetic images. Based on our analysis, we also propose potential future directions and innovative solutions that could enhance the effectiveness of small, dense and overlapped object detection systems. The scope of our investigations encompasses object detection, length measurement, and anomaly detection within the context of the recycling process. The anomaly detection strategy is robust against variations in image resolution and zoom levels, ensuring reliable performance in industrial applications. The repository of the proposed dataset, methods and evaluation codes can be found at: https://github.com/o‑messai/SDOOD
Authors:Dingkun Wei, Zehong Shen, Yan Xia, Georgios Pavlakos, Yujun Shen, Xiaowei Zhou
Abstract:
Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate. We observe that this limitation stems from the absence of reliable high‑order temporal cues ‑‑ velocity and acceleration ‑‑ which are essential for reconstructing motion that exhibits realistic momentum, timing, and high‑frequency detail. We introduce HTD‑Refine, a post‑processing framework that augments existing Human Motion Recovery (HMR) pipelines using explicitly estimated high‑order temporal dynamics. At the core of our system is PVA‑Net, a temporal transformer that infers per‑joint 2D positions, 3D velocities, and 3D accelerations directly from a monocular video. These predicted dynamics serve as soft yet informative constraints in a global optimization procedure that refines world‑space trajectories, significantly reducing jitter, suppressing over‑smoothing, and restoring physically plausible motion. Extensive experiments on challenging in‑the‑wild benchmarks show that HTD‑Refine consistently improves state‑of‑the‑art HMR methods, yielding more accurate global trajectories and substantially more natural motion dynamics. Our results highlight the critical role of high‑order temporal modeling in advancing monocular human motion recovery.
Authors:Chenxu Peng, Chenxu Wang, Yimian Dai, Yongxiang Liu, Ming-Ming Cheng, Xiang Li
Abstract:
Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scene diversity, low semantic granularity, and poor structural continuity, restricting their generalization across environments. To address these challenges, we introduce WorldRoadSeg‑360K, the largest and most diverse road segmentation dataset to date, comprising 366,947 high‑resolution images collected from 38 countries and 223 cities across various terrains and continents. WorldRoadSeg‑360K serves as a comprehensive benchmark and reveals key challenges in handling diverse and structurally complex scenes. Automated approaches often struggle to preserve road connectivity, while current interactive methods lack efficient, topology‑sensitive tools for real‑world road editing. To this end, we present RoadGIE, establishing a novel interactive paradigm for road extraction in remote sensing. Unlike prior point‑ or box‑based prompting strategies, RoadGIE supports connectivity‑aware prompts, including clicks and scribbles, which inherently align with the topology of road networks. To improve structural consistency and mitigate performance degradation during iterative interactions, RoadGIE integrates an expert‑guided prompting strategy and adapts the skeleton‑based recall loss for interactive scenarios. RoadGIE achieves state‑of‑the‑art performance in both segmentation accuracy and topological consistency on WorldRoadSeg‑360K and other benchmarks, while maintaining efficient operation with only 3.7M parameters. The code are publicly available at: https://github.com/chaineypung/RoadGIE
Authors:Yong Li, Furong Jia, Dacheng Yin, Kang Rong, Fengyun Rao, Jing Lyu, Fan Zhang
Abstract:
Image geo‑localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts typically solve it through an iterative workflow: they inspect informative regions, form location hypotheses, seek external evidence, and revise their judgments as new clues appear. Existing methods only partially capture this process: direct prediction methods bypass evidence acquisition altogether, while retrieval‑augmented methods introduce external evidence but usually provide limited supervision on the intermediate decisions of where to search, how to query, and how to filter noisy results. We present REVERSE, a framework that reinforces the interplay between evidence search and verification to enable multi‑turn agentic reasoning. REVERSE teaches three intermediate decisions: where to look, what to query, and what evidence to trust. To support this, we construct tool‑grounded trajectories with annotated region selections, search observations, and geo‑informative evidence labels, and introduce process rewards for visual grounding, query utility, and evidence discrimination. An offline search cache makes retrieval observations stable and reusable during reinforcement learning, enabling dense supervision over noisy search results. With a 4B model, REVERSE outperforms strong retrieval‑augmented baselines and rivals substantially larger models on Im2GPS3k and YFCC4k. Code is available at https://github.com/yonglleee/REVERSE.
Authors:Regina Kurkova, Maxim Popov, Sergey Kolyubin
Abstract:
Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evaluation is still largely tied to fixed benchmark datasets with limited coverage of manipulation‑relevant corner cases. In this work, we extend OSMa‑Bench toward controllable benchmarking with prompt‑generated synthetic indoor scenes. Our pipeline automatically generates scene descriptions, synthesizes corresponding environments with SceneSmith, and adapts the resulting assets into an OSMa‑Bench‑compatible simulation format. This adaptation requires a nontrivial intermediate layer, including semantic normalization, material and texture repair, shader fallback policies, floor handling, navigation setup, and controlled lighting configuration. A key advantage of the proposed setup is that the original scene‑generation prompt is known in advance and can therefore serve as an auxiliary semantic specification of the intended scene. We use this property to extend the VQA component of OSMa‑Bench with a prompt‑grounded question category. The resulting framework supports targeted stress‑testing of semantic scene representations under conditions such as clutter, small objects, partial occlusions, and lighting variation, and makes benchmarking more extensible and better aligned with downstream manipulation requirements. Our code is available at https://github.com/be2rlab/OSMa‑Bench‑v2.
Authors:Pascal Herrmann, Maarten Bieshaar, Dennis Mack, Robert Herzog, Juergen Gall
Abstract:
Human motion generation has made tremendous progress in recent years, with state‑of‑the‑art approaches surpassing ground truth data in leading evaluation benchmarks. However, visual inspection of the generated motions paints a different picture. Even state‑of‑the‑art approaches generate motions frequently containing self‑intersections, i.e., body parts interpenetrating, which are strong artifacts, severely limiting the perceived motion quality. We introduce a novel loss, which explicitly penalizes self‑intersections, to the training of human motion generation methods. We base our loss on a sphere proxy of human geometry, which allows us to calculate a self‑intersection loss 98% faster and uses 83% less memory than comparable methods based on triangular meshes. The loss is agnostic to the specific approach, and we add it to the training of the recent human motion generation methods human motion diffusion model (MDM) and MoMask. Our extensive experiments show a reduction of self‑intersections in generated motions of up to 49% while improving other evaluation metrics. The code is available at https://github.com/boschresearch/humansphereproxy .
Authors:Tomohisa Takeda, Yu-Chieh Lin, Yuji Nozawa, Youyang Ng, Osamu Torii, Yusuke Matsui
Abstract:
Existing Multi‑Turn Composed Image Retrieval (MTCIR) datasets lack dialogue‑history consistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In CIRCLED, the query at each turn progressively approaches the target image. Data are generated via a CIReVL‑based retrieval pipeline and curated with multiple filters on retrieval success, turn length, consistency, and information redundancy to ensure quality. In total, we collect 22,608 multi‑turn sessions across nine subsets, substantially exceeding Multi‑turn FashionIQ (11,505 sessions) in both scale and generality. We further apply multiple baseline methods and quantitatively assess retrieval accuracy on CIRCLED. Our work provides a practical, high‑quality benchmark to facilitate future research on multi‑turn CIR. The dataset and code are publicly available at https://huggingface.co/datasets/tk1441/CIRCLED and https://github.com/mti‑lab/circled.
Authors:Peng Zhang, Guanghao Zhang, Wanggui He, Longxiang Zhang, Mushui Liu, Yan Xia, Zhenhao Peng, Weilong Dai, Jinlong Liu, Haobing Tang, Le Zhang, Hao Jiang, Pipei Huang
Abstract:
Recent video multimodal large language models (MLLMs) increasingly couple step‑by‑step reasoning with on‑demand visual evidence retrieval, allowing models to revisit relevant video segments during inference. However, two structural gaps remain in existing thinking‑with‑video systems. (i) Sampling density is not a learnable decision: existing methods may let the model decide where to look, but the per‑window frame rate is largely fixed. As a result, fine‑grained evidence is often recovered through repeated retrieval calls, which increases inference context length and training difficulty. (ii) Retrieval and answer generation are usually optimized with a single trajectory‑level advantage, so the "where to look" tokens and the "how to answer" tokens receive the same credit even when one is correct and the other is not. To address these gaps, we present DynFrame, a framework that emits the temporal window and the sampling density as native tokens within a single autoregressive pass. This learnable span‑density retrieval enables acquiring multi‑granularity evidence with a single retrieval step. Based on the above tokenized retrieval interface, we further introduce Segment‑Decoupled GRPO (SD‑GRPO), which splits each rollout at the retrieval boundary and assigns role‑specific token‑level advantages, separately crediting the sampling decision and the answer. Trained on the curated DM‑CoT‑74k and DM‑RL‑45k, DynFrame‑4B is competitive with strong 7B‑8B baselines across six benchmarks (NExT‑GQA, Charades‑STA, ActivityNet‑MR, Video‑MME, MLVU, LVBench), and DynFrame‑8B sets new state‑of‑the‑art on most metrics. Code is available at https://github.com/zhangguanghao523/DynFrame.
Authors:Sirojbek Safarov, Jaewoo Park, Yoon Gyo Jung, Kuan-Chuan Peng, Wonchul Kim, Seongdeok Bang, Octavia Camps
Abstract:
Anomaly detection (AD) under data contamination is critical for deploying unsupervised defect detection in industrial environments, where curating perfectly clean training sets is impractical. However, existing methods are sensitive to contamination, suffering significant performance degradation as the noise ratio increases. In this paper, we propose Memory‑Distilled Selection (MeDS), a training algorithm based on data selection. MeDS constructs an ensemble of partial memories via random subsampling, where the resulting sparsity acts as a low‑pass filter that captures nominal patterns across a wide range of noise ratios, enabling coarse‑level identification of contaminated samples. The aggregated distances to the bootstrapped memories are then distilled into a reconstruction score network, which is subsequently fine‑tuned on clean data filtered using scores from the distilled model, enabling fine‑grained localization of anomalies. MeDS is robust across a wide range of noise ratios without requiring noise‑ratio‑specific hyperparameter tuning, achieving 99.16% image‑level AUROC on MVTecAD at a 40% noise ratio, and attaining state‑of‑the‑art performance on both VisA and Real‑IAD under noisy settings. We thoroughly verify the efficacy of MeDS on industrial AD benchmarks under noisy data scenarios, accompanied by in‑depth empirical analyses.
Authors:Yunze Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen
Abstract:
Unified multimodal embedding spaces have become the standard interface for cross‑modal retrieval and multimodal RAG, and recent audio‑video‑text (AVT) encoders extend this setting to three modalities. Such encoders can produce a joint (T,V,A) embedding whenever all three modalities are available, but standard pairwise InfoNCE objectives leave this signal unused during training. We close this gap with fusion‑as‑teacher distillation, which treats a stop‑gradient copy of the fused embedding as a teacher signal for the single‑modal embeddings, paired with a Tuple‑InfoNCE term that supervises the fused embedding directly. We instantiate this objective as OmniRetriever‑7B. Across six zero‑shot retrieval benchmarks, OmniRetriever‑7B surpasses the closed‑source Gemini Embedding 2 by 13.3‑18.0 R@1 on Clotho and SoundDescs, and reaches the contemporary zero‑shot specialist band of open video‑text encoders on MSR‑VTT and MSVD. To stress‑test joint representations, we further release OmniRetriever‑Bench, a 12‑direction AVT retrieval benchmark totaling 3782 triples; on it OmniRetriever‑7B attains AVG‑all 34.84, improving over Gemini Embedding 2 by 1.72 and over the best prior open‑source AVT method by 8.03.
Authors:Lanqing Liu, Ruize Cui, Jialun Pei, Diandian Guo, Tiffany Y. So, Pheng-Ann Heng, Jing Qin
Abstract:
Liver surface landmark detection is a fundamental prerequisite for anatomical guidance in laparoscopic liver surgery. However, it remains unreliable in practice due to two pervasive challenges: illumination attenuation in underexposed regions and the structural mismatch between pixel‑wise localization and continuous curvilinear geometry. To address these limitations, we propose A2ONet, an attenuation‑resilient alternating optimization network for robust liver landmark detection. To mitigate illumination attenuation, A2ONet embraces an illumination field compensation (IFC) block that adaptively enhances dark regions while preserving structural consistency. Meanwhile, we introduce a lightweight frequency‑orientation selective filter (FOSF) to suppress repetitive texture interference and preserve salient curvilinear cues. Building upon these resilient representations, we design an alternating seg‑curve optimization (ASCO) decoder that iteratively couples dense segmentation with explicit curve modeling, enabling mutual guidance to optimize both structural continuity and endpoint localization. Extensive evaluations on L3D‑2K, L3D, and P2ILF demonstrate consistent improvements over competitive methods, establishing a more reliable foundation for intraoperative anatomy guidance. Our code will be available at https://github.com/hyperiondk115/A2ONet.
Authors:Zhenhua Du, Zhen Tan, Haoyu Zhang, Dewen Hu, Shuaifeng Zhi, Peidong Liu
Abstract:
While 3D Gaussian Splatting has achieved remarkable success in photorealistic novel view synthesis, its pursuit of fast and high‑fidelity 3D reconstruction has long been constrained by a trade‑off between geometric accuracy and optimization efficiency. Methods specialized in image rendering converge quickly at the cost of imperfect geometry caused by superfluous primitives overfitting training views, while methods integrating neural signed‑distance field (SDF) for better geometry incur prohibitive training costs. In this paper, we attempt to strike a better trade‑off by tethering scaffold‑anchored Gaussians to a jointly optimized sparse voxel scaffold. This hybrid Gaussian‑Voxel representation explicitly confines anchored Gaussians to a narrow band around surfaces defined by voxelized SDFs, which effectively improves representation efficiency and condenses floating Gaussians without sacrificing geometry quality. An implicit surface tethering loss further pulls individual Gaussian primitives closer to SDF‑induced surfaces in a mutually regularized manner for improved reconstruction accuracy. Extensive experiments on diverse real‑world indoor scenes from ScanNet++, ScanNetv2, and DeepBlending datasets demonstrate that our method achieves state‑of‑the‑art surface reconstruction quality as well as superior novel view synthesis against leading baselines, while maintaining fast training convergence and real‑time rendering. Code will be available at https://github.com/duzh11/VoxelGS.
Authors:Amey Sunil Kulkarni
Abstract:
Style transfer with pre‑trained diffusion models has advanced rapidly, but a core question remains underexplored: where in the model should style injection be strongest? StyleID, the leading training‑free method, uses a single global parameter (gamma) uniformly across all layers and timesteps, which forces a fixed tradeoff between style quality and content preservation. We show this tradeoff is unnecessarily rigid. We systematically explore four dimensions of control: varying style injection strength across decoder layers, across denoising timesteps, and scheduling ControlNet geometric conditioning along both axes. The pattern is consistent everywhere: decreasing schedules, with stronger structural signal injection in shallower layers and earlier timesteps, reliably outperform the reverse. Beyond direction, schedule shape matters: cosine and square‑root timestep schedules outperform linear. Most importantly, we find that gamma scheduling and ControlNet conditioning are nearly independent. The resulting combined configurations expand the Pareto frontier, offering superior tradeoffs between style fidelity and content preservation compared to any single baseline setting. Our best balanced configuration achieves ArtFID of 27.036 versus StyleID's 28.801 ‑ a 6.1% relative improvement, with consistent gains across the full style‑content tradeoff frontier. Results are validated across 35 configurations totaling over 28,000 stylized images using four complementary metrics. These findings generalize across SD backbones with identical rank ordering. All modifications are training‑free, parameter‑free, and require only a few lines of scheduling code; code is available at https://github.com/ameyskulkarni/scheduled_style_injection.
Authors:Jiahe Huang, Sihan Xu, Sharvaree Vadgama, Rose Yu
Abstract:
Generative models have emerged as a powerful paradigm for solving physics systems and modeling complex spatiotemporal dynamics. However, achieving high physical accuracy without incurring high computational cost remains a fundamental challenge, as existing approaches face a critical speed‑fidelity trade‑off. In this work, we introduce Recursive Flow Matching (RecFM), a generative framework for forecasting complex spatiotemporal dynamics. RecFM enforces self‑consistency to align trajectories across discretization scales, reducing discretization errors and improving performance across metrics for physics‑based tasks. To our knowledge, this is the first method to achieve high‑fidelity one‑ and few‑step (2‑4 step) dynamic generation for scientific systems with performance comparable to state‑of‑the‑art multi‑step solvers. Across challenging scientific benchmarks, RecFM achieves up to a 20× speedup over leading diffusion‑based emulators while improving predictive accuracy. Furthermore, RecFM reduces mean squared error by over 15% compared to vanilla flow matching, offering a scalable and efficient solution for real‑time scientific emulation.
Authors:Akide Liu, Jinbo Xing, Chaojie Mao, Ye Li, Zeyu Zhang, Yefei He, Weijie Wang, Zihan Wang, Yu Liu, Gholamreza Haffari, Bohan Zhuang
Abstract:
Minute‑scale cinematic video generation is a central challenge for generative video models. Existing paradigms address only fragments of this challenge: single‑shot extrapolation preserves an anchor but lacks cinematic structure, while multi‑shot storytelling imposes structure yet remains free to invent its visual states rather than continue an observed one. We define Multi‑Shot Video Extrapolation (MSVE), a task that extends an observed frame or clip into a sequence of cinematically structured shots while preserving anchor state and advancing narrative intent. This setting operates under the finite per‑call generation budget of short‑video models. We identify three coupled bottlenecks: (1) global planners over‑specify unsupported details from full screenplays; (2) shot‑level prompts dilute task‑relevant state when carrying the complete story; and (3) temporal chaining turns generated frames into a lossy memory in which identity, scene, object, and action state decay. MSVE reveals that long‑video failure is not merely a limitation of context length, but a failure of context allocation. We propose Recursive Context Allocation (ReCA), an inference‑time framework that allocates context hierarchically across planning and generation. ReCA recursively decomposes MSVE into context‑bounded subproblems, invokes frozen generators at leaf nodes, and propagates structured state updates across time. To evaluate this setting, we further propose MSVE‑Bench and NB‑Q, a source‑grounded protocol with prompts purpose‑built for 3 to 5 minute long‑video generation, a regime not addressed by existing short‑clip benchmarks. Compared to previous methods, ReCA improves average normalized score by 8 to 16 percent over the strongest competing controller and improves multi‑shot consistency metrics by 28 to 43 percent. View the project page at https://reca.vmv.re.
Authors:Yuxu Lu, Dong Yang, Xiaoyu Li, Mengwei Bao, Congcong Zhao
Abstract:
Maritime intelligent transportation systems (MITS) are essential for ensuring navigation safety and efficiency in busy waterways. However, accurate vessel trajectory prediction remains challenging due to the limitations of single‑source data. Automatic identification system (AIS) data is often sparse or unavailable for small vessels, while closed‑circuit television (CCTV) data alone cannot fully capture dynamic vessel behavior. To mitigate these challenges, we propose a cross‑modal interaction‑based vessel trajectory prediction (named CmIVTP) framework to model the intricate interactions between vessel dynamics and environmental constraints. Specifically, we introduce a target‑aware scene encoder to extract scene semantic features, effectively capturing vessel‑environment interactions and enhancing trajectory prediction accuracy. In addition, we propose a cross‑modal interaction transformer, which integrates AIS‑derived motion features, CCTV‑based environmental features, and scene representations. It leverages cross‑modal attention mechanisms to simultaneously capture intra‑modal semantics and inter‑modal interactions, ensuring dynamically consistent and environmentally feasible predictions. Furthermore, we construct a vessel group trajectory bank by clustering historical AIS trajectories into representative motion patterns, providing an efficient and scalable approach for candidate trajectory generation. Additionally, we introduce the maritime multimodal dataset plus (named Maritime‑MmD^+), a large‑scale dataset that synchronizes AIS data and CCTV video data, providing robust support for multimodal trajectory prediction research. Extensive experiments demonstrate that CmIVTP achieves better performance on multimodal‑driven vessel trajectory prediction benchmarks. The code resources for this work can be available at https://github.com/LouisYxLu/CmIVTP.
Authors:Meituan LongCat Team, Xunliang Cai, Meng Cheng, Feng Gao, Zhe Kong, Jiamu Li, Le Li, Weiheng Li, Hongyu Liu, Shuai Tan, Xiaoming Wei, Tianyu Yang, Yong Zhang
Abstract:
Despite advances in audio‑driven video generation, achieving commercial‑grade stability remains challenging. We present LongCat‑Video‑Avatar 1.5, an upgraded open‑source framework prioritizing systematic engineering and production‑readiness over architectural novelty. By upgrading the audio encoder to Whisper Large and meticulously scaling our training recipes, v1.5 achieves accurate lip‑synchronization, full‑body temporal stability, and robust long‑video generation with strict identity consistency. Through rigorous data curation and RLHF Training, the model readily generalizes to stylized domains such as anime and animals, and natively handles complex real‑world conditions, such as multi‑person interactions and object handling. Furthermore, addressing the practical demands of industrial deployment, we employ advanced step distillation to accelerate inference to an optimal 8 NFE, achieving a favorable trade‑off between serving efficiency and visual fidelity. The superiority of our approach is validated through extensive quantitative metrics and a rigorous human evaluation conducted on a comprehensive benchmark of over 500 diverse test cases. Results show that v1.5 achieves competitive or superior performance compared to leading closed‑source systems (e.g., HeyGen, OmniHuman 1.5, Kling Avatar 2.0) across human‑likeness ratings and expert‑level quality assessments on our benchmark. With its open‑source release, LongCat‑Video‑Avatar 1.5 narrows the gap between academic research prototypes and commercial‑grade deployment.
Authors:Xudong Lu, Xueying Li, Annan Wang, Yang Bo, Jinpeng Chen, Zengliang Li, Nianzu Yang, Rui Liu, Xue Yang, Jingwen Hou, Hongsheng Li
Abstract:
We introduce OmniInteract, a streaming benchmark for real‑time omnimodal large language models evaluated through native online inference over audio‑visual streams. Unlike offline video understanding or text‑prompted streaming QA, OmniInteract preserves the original audio‑visual stream and requires models to process it online, without access to future content. User queries and ambient sounds are embedded in the audio track, requiring models to detect multimodal triggers, decide when to respond, and answer while the stream unfolds. OmniInteract contains 250 videos with 1,430 temporally grounded response slots: 1,062 1Q1A slots across real‑time, proactive, and nested scenarios, and 368 1QnA slots for continuous task monitoring and step guidance. Each slot includes a trigger, response window, and target answer. We evaluate response correctness, timing, invalid outputs, interruption handling, and context continuity using Interaction‑Aware Quality‑Timeliness F1, Interruption Diagnostic Suite, and Nested Chain Completion Score. Experiments show that current models remain weak in streaming interaction, with the best overall IA‑QTF1 reaching only 0.368 and the best 1QnA IA‑QTF1 only 0.052. Further study on mathematical reasoning in full‑duplex settings shows that offline capability does not necessarily transfer to online interaction. Code and datasets will be made publicly accessible at https://github.com/Lucky‑Lance/OmniInteract.
Authors:Zhiyao Cui, Chenxu Wang, Shuyue Hu, Yiqun Zhang, Wenqi Shao, Qiaosheng Zhang, Zhen Wang
Abstract:
Producing presentation slides automatically entails coordinating narrative structure with page‑level graphic design under strict spatial constraints. For such structured multimodal tasks, a well‑organized design process is essential to ensure the final quality of slides. Existing approaches rely on fixed templates or directly emit executable code, thereby both limiting the creative layout‑design capabilities of LLMs and bypassing the essential slide‑page design step. To address these limitations, this paper (1) proposes a hierarchical slides generation workflow, DeepSlides, that systematically organizes slide design tasks without any predefined template or style, decoupling slide‑page design from implementation; (2) introduces SlideDesign, a dataset tailored specifically for slides generation tasks; and (3) presents a multi‑agent reinforcement learning training paradigm and trains a couple of models, SlideQwens, for slide design and implementation. Experimental results demonstrate that our proposed framework outperforms baseline methods on evaluated metrics and achieves superior performance in human preference evaluations. The dataset and code are available at https://github.com/sxswz213/DeepSlides.
Authors:Jiangbei Hu, Weichao Song, Shibo Yu, Mohan Wang, Zihan Yi, Rui Wu, Mingkang Xiang, Na Lei, Shengfa Wang, Zhongxuan Luo, Ying He
Abstract:
Underwater scene reconstruction is essential for immersive exploration of aquatic environments, yet remains challenging due to complex participating‑media effects such as absorption and scattering, as well as the limited field of view (FoV) of conventional cameras. Although combining panoramic imaging with 3D Gaussian Splatting (3DGS) offers a promising direction for photorealistic underwater rendering, traditional 3DGS struggles with both spherical projection distortion and underwater medium degradation. In this paper, we propose Underwater360, a physics‑informed omnidirectional 3DGS framework for underwater panoramic scene reconstruction. First, we introduce an Omnidirectional Gaussian Splatting module that performs ray casting directly in spherical camera space instead of relying on 2D projection approximations, thereby reducing geometric distortions under 360^\circ FoV. Second, we design a physics‑based appearance‑medium modeling architecture with pose‑conditioned appearance embeddings to explicitly decouple intrinsic scene radiance from depth‑dependent backscatter and attenuation, enabling physically grounded scene appearance restoration. Finally, we establish a new panoramic underwater benchmark dataset containing both synthetic and real‑world scenes. Extensive experiments demonstrate that Underwater360 achieves superior performance in underwater novel view synthesis and scene appearance restoration, delivering improved rendering quality and cross‑view consistency in complex underwater environments. The code and datasets are released at https://github.com/SwcK423/Underwater360
Authors:Qiaomu Miao, Haoyu Wu, Jingyi Xu, Minh Hoai, Dimitris Samaras
Abstract:
Understanding human gaze behavior is essential for complex scene comprehension and human‑computer interaction. Traditional gaze following models are typically restricted to pure spatial localization, lacking the high‑level capacity to reason about semantic targets or complex social contexts. Furthermore, these models often process individuals sequentially, requiring redundant computations over the same scene image for multi‑person inference. While recent Vision‑Language Models (VLMs) offer the exceptional semantic reasoning needed to address gaze‑related semantic tasks, their reliance on discrete text generation inherently limits precision in continuous spatial tasks like gaze localization. To bridge this gap, we propose OmniGF, a unified vision‑language framework that adapts foundational VLMs for highly scalable multi‑person gaze reasoning. The model adopts a dual‑branch decoding strategy: a structured language branch generates discrete reasoning states, while a continuous spatial branch directly taps into the VLM's dense hidden states. Supervising these extracted representations with high‑resolution gaze target heatmaps effectively overcomes the spatial bottleneck of text‑only coordinate generation. Furthermore, to explicitly ground the model in multi‑person scenes, we augment the input with head embeddings encoded from cropped head images, providing fine‑grained appearance and orientation cues for all individuals simultaneously. By modeling all individuals and leveraging the strong semantic capability of VLMs, OmniGF seamlessly integrates precise spatial gaze target estimation, semantic gaze prediction, and complex social gaze reasoning. Extensive experiments demonstrate that our framework establishes new state‑of‑the‑art performance across multiple standard benchmarks. Code is available at https://github.com/cvlab‑stonybrook/omnigf.
Authors:Mengchen Fan, Baocheng Geng, Xi Xiao, Tianyang Wang, Siyuan Mei, Pulin Che, Xiaoqian Jiang, Qizhen Lan
Abstract:
Deploying high‑performing 3D medical image segmenters (e.g., nnU‑Net) is often limited by memory footprint and inference latency. Compression is therefore necessary, but compact 3D encoders tend to lose fine structural cues (small lesions and sharp boundaries) as downsampling repeats across multi‑resolution stages. We propose Detail Consistent Distillation (DCD), a stage‑wise distillation framework that preserves structural detail across scales by aligning teacher‑student features in a wavelet‑decomposed representation. At each encoder stage, DCD distills directional detail components in the wavelet domain while leaving the coarse approximation comparatively unconstrained, avoiding over‑regularization of global semantics. DCD is used only during training and introduces no inference‑time overhead. Experiments on the BraTS 2024 and ISLES 2022 benchmarks demonstrate that our approach achieves superior performance in MRI segmentation using 3D multi‑modal data. Code and implementation details for DCD are publicly available at https://github.com/ClinicaAlpha/DCD‑3D‑MedSeg.
Authors:Junlin Yang, Tian Yu, Nicha C. Dvornek, Yuexi Du, Peiyu Duan, Annabella Shewarega, Lawrence H. Staib, James S. Duncan, Julius Chapiro
Abstract:
Hepatocellular carcinoma (HCC) is biologically heterogeneous, shaped by the interplay between hepatic functional reserve and tumor‑related oncologic factors; thus, similar survival outcomes may reflect fundamentally different underlying biological processes. Prognostic modeling in HCC is informed by rich multimodal information from multiparametric MRI and radiology reports from routine clinical practice. Existing prognostic vision‑language models (VLMs) learn a single entangled latent representation that blends hepatic and tumor‑related factors, limiting both accuracy and biological interpretability. We present BioFact‑MoE, a biologically factorized Mixture of Experts (MoE) framework that explicitly decomposes liver and tumor factors via biologically supervised experts within a residual MoE survival architecture. On a HCC cohort of N=588 patients (pretrained on 4,582 3D MRI image‑report pairs), BioFact‑MoE consistently improves survival prediction over all baselines across time horizons, achieving 12‑, 18‑, and 24‑month AUCs of 75.33%, 75.85%, and 73.96%. Beyond scalar risk prediction, gated expert weights enable phenotype‑aware risk stratification. Pathway‑informed gating uncovers clinically meaningful treatment‑associated survival heterogeneity. In held‑out validation, hepatic and tumor embeddings show selective associations with liver function and tumor burden markers, respectively (p<0.05), without supervision. The code is available at https://github.com/jy‑639/BioFact‑MoE.
Authors:Vukasin Bozic, Isidora Slavkovic, Dominik Narnhofer, Nando Metzger, Denis Rozumny, Konrad Schindler, Nikolai Kalischek
Abstract:
Geometry estimation from perspective images has greatly advanced, maturing to the point where off‑the‑shelf foundation models are able to reconstruct 3D scene structure not only from multi‑view imagery, but even from a single view. A natural extension is 3D reconstruction from panoramas, with the exciting prospect of recovering a full 360‑degree scene from a single panoramic image. In this work, we introduce PaGeR (Panoramic Geometry Reconstruction), a framework to lift powerful 3D foundation models designed for perspective imagery to the panorama domain. Our strategy is to start from a pre‑trained transformer for 3D reconstruction and turn it into a unified high‑performance model that predicts scale‑invariant depth, metric depth, surface normals, and sky masks from both perspective and omnidirectional images, in a single forward pass. By keeping architectural changes to a minimum and mixing perspective and panoramic images during training, PaGeR retains the rich 3D prior of the underlying foundation model while learning to also estimate geometrically consistent 360‑degree scenes from single panoramas. We extensively test our method in both indoor and outdoor environments and find that it delivers state‑of‑the‑art performance and excellent zero‑shot performance across a wide range of scenes. Code, data and models are available \hrefhttps://github.com/prs‑eth/PaGeR\texthere.
Authors:Xinran Liang, Esin Tureci, Prachi Sinha, Ye Zhu, Vikram V. Ramaswamy, Olga Russakovsky
Abstract:
Different visual patterns appear with different frequencies in the world: e.g., beach balls appear on sand more often than they do on a road. These statistics are reflected in vision datasets, and as a result trained models more easily recognize objects in common scenarios. However, recognizing a beach ball on a road may arguably be even more important than recognizing it on sand. We study how to mitigate this discrepancy. Since collecting uncommon images in the real world may be difficult, we explore whether generating images with less frequent contexts can serve as effective training augmentation. A key challenge is guiding generations to remain close to the original dataset distribution while creating diverse images with uncommon contexts. We introduce Decoupling Contextual Patterns with Generations (DecoupleGen), a method that personalizes text‑to‑image diffusion models to facilitate coherent synthesis of images with rare contexts while preserving original visual details. The generated images contain semantically meaningful content and remain visually aligned with the original datasets. We further apply verification constraints to ensure relevance of the augmented data. We evaluate our approach on object classification and recognition tasks on complex scene datasets. Our experiments demonstrate consistent improvements over previous approaches, and our analyses identify factors underlying these improvements.
Authors:Chuhan Chen, Tianshu Huang, Akarsh Prabhakara, Chaithanya Kumar Mummadi, Zhongxiao Cong, Anthony Rowe, Matthew O'Toole, Deva Ramanan
Abstract:
Radars are an ideal complement to cameras: both are inexpensive, solid‑state sensors, with cameras offering fine angular resolution, while radars provide metric depth and robustness under adverse weather. However, radar data is more difficult to interpret than camera images and varies significantly between sensors, necessitating increased reliance on simulation for prototyping sensors and processing pipelines. Recent work treating radar reconstruction as a novel view synthesis problem has shown great promise in reconstructing radar‑relevant geometry and simulating low‑level radar data. However, such methods are constrained by the low spatial resolution of the underlying radar. To address this, we propose a unified differentiable renderer, RadarSim, which leverages the high angular resolution of RGB cameras to generate Doppler radar range images from a camera‑initialized neural field. Using a novel data set of calibrated radar camera recordings from a custom hand‑held rig, we demonstrate that RadarSim produces sharper geometry and Doppler range frames than radar‑only reconstructions.
Authors:İsmail Emre Canıtez, Özgür Erkent
Abstract:
Semantic segmentation in complex environments such as urban driving scenes remains challenging under adverse lighting conditions, where RGB images alone provide insufficient information. RGB‑Thermal fusion leverages the complementary strengths of visible and infrared imagery to improve scene understanding; however, effectively integrating these heterogeneous modalities at varying levels of feature abstraction remains an open problem. In this paper, we propose a multi‑modal fusion architecture built upon dual ConvNeXt V2 backbones that employs stage‑wise, modality‑adaptive fusion strategies. For early‑stage features, we introduce a Frequency‑Based Fusion Module that decomposes infrared features into low‑ and high‑frequency components via Gaussian filtering, applies dual‑branch spatial attention to selectively emphasize thermal patterns and fine‑grained boundaries, and integrates them with RGB features through a confidence‑gated residual mechanism. For late‑stage features, we design a semantic fusion module with cross‑modal attention and multi‑scale depthwise convolutions to capture semantic correspondences across modalities. The fused features are decoded via a PANet‑style bidirectional decoder with deep supervision. Experiments on MFNet and PST900 demonstrate that our lightest variant achieves 61.73% and 86.24% mIoU, respectively, with only 35.43M parameters, outperforming recent methods while using substantially fewer parameters and lower computational cost. Code is available at https://github.com/ismailemrecntz/VISIBLE‑INFRARED‑SENSOR‑FUSION
Authors:Xiangye Lin, Hongxin Zhang, Ruxi Deng, Qinhong Zhou, Chuang Gan
Abstract:
In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmental constraints across city‑scale outdoor domains. We introduce Sentinel Challenge, a benchmark where multiple decentralized embodied agents must communicate in natural language to agree on a mutually safe and convenient meeting point within large, city‑scale outdoor environments. Each agent must then navigate safely while avoiding dynamic sentinels patrolling the area, using a tool that provides coarse spatial information. To address this, we propose CoSaR (Cooperative Spatial Reasoning and Planning), a framework that bridges the high‑level communication and planning abilities of foundation models with the precision of classical spatial navigation algorithms. CoSaR enables agents to exchange situational updates, reason over evolving spatial constraints, and collaboratively replan trajectories. Evaluated across 14 city‑level scenes with 3‑5 agents, CoSaR consistently leads to faster gathering, shorter path lengths, and improved safety. Our results demonstrate that integrating dynamic communication with spatial reasoning is essential for robust multi‑agent cooperation. By formalizing this new setting and providing a scalable benchmark, we aim to build a foundation for advancing cooperative spatial intelligence in embodied multi‑agent systems. Code and challenge are available at https://github.com/UMass‑Embodied‑AGI/Sentinel.
Authors:Weijie Wang, Zimu Li, Jinchuan Shi, Zeyu Zhang, Botao Ye, Marc Pollefeys, Donny Y. Chen, Bohan Zhuang
Abstract:
Sparse‑view 3D reconstruction is increasingly addressed with feed‑forward splatting networks that predict explicit primitives directly from images. Yet most existing methods remain centered on Gaussian primitives and expose surfaces only indirectly: extracting a usable mesh for downstream simulation, physics reasoning, or embodied interaction still requires expensive post‑hoc steps that break the feed‑forward promise. This limitation is especially pronounced in pose‑free settings, where scene structure and camera parameters must be estimated jointly from sparse observations. We present TriSplat, a feed‑forward reconstruction network that represents scenes with oriented triangle primitives and directly exports simulation‑ready mesh scenes from a single forward pass. Given input images, the network predicts local 3D point maps, triangle attributes, camera poses, and optional intrinsics. Rather than regressing triangle orientation as an unconstrained latent variable, our approach constructs geometry normals from the predicted point maps, refines them with an image‑conditioned normal head, and converts them into stable local frames for triangle parameterization. A mono‑normal bootstrap schedule further stabilizes early training, while opacity and blur scheduling progressively sharpens the learned surface representation for direct mesh extraction. Experiments on RealEstate10K and DL3DV show that this representation produces more geometry‑faithful reconstructions than Gaussian feed‑forward baselines while maintaining competitive novel‑view rendering quality. Because the rendering primitives are themselves surface triangles, the output can be directly ingested by physics engines, collision detectors, and standard rendering pipelines without any conversion, making it a practical simulation‑ready solution for feed‑forward 3D scene reconstruction.
Authors:Shuhong Zheng, Aashish Kumar Misraa, Yu-Teng Li, Yu-Jhe Li, Igor Gilitschenski
Abstract:
Subject‑driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. This limits cross‑modal reasoning abilities and causes copy‑paste artifacts. Recent frameworks that connect multimodal models and diffusion models improve instruction following, but largely overlook identity preservation. To address these limitations, we condition diffusion models on Multimodal Large Language Models (MLLMs) that jointly encode text and reference images, and augment it with VAE‑based identity conditioning. A novel Dual Layer Aggregation (DLA) module is designed to aggregate multi‑level MLLM features for optimal conditioning, and a multi‑stage denoising strategy is applied to progressively balance the semantic information from MLLM and fine‑detail identity from VAE during inference. Extensive experiments demonstrate that our approach harmonizes multimodal understanding with identity preservation, mitigates copy‑paste issues, and achieves superior performance regarding human preference on subject‑driven image generation. Our project website is available at https://zsh2000.github.io/squeeze‑mllm‑subject‑gen/.
Authors:Jun-Tao Tang, Yu-Cheng Shi, Zhen-Hao Xie, Da-Wei Zhou
Abstract:
Multimodal Large Language Models (MLLMs) achieve versatility by reformulating diverse tasks into a unified instruction‑following framework via instruction tuning. However, real‑world deployment requires continuous adaptation to emerging tasks, motivating Multimodal Continual Instruction Tuning (MCIT). Despite its growing importance, current MCIT research is hindered by severe engineering bottlenecks. Existing methods are typically implemented by directly modifying the base MLLM codebase, which imposes substantial implementation overhead and yields method‑specific architectures that severely limit code reuse and fair comparison. To address this, we introduce Prism, a plug‑in reproducible codebase specifically designed for scalable MCIT research. It separates algorithmic development from the backbone implementation via a lightweight plugin registration mechanism, enabling new strategies to be integrated as independent plugins without modifying the underlying MLLM codebase, thereby eliminating structural fragmentation and accelerating method development. Prism natively supports widely used large‑scale training pipeline, thereby enabling reproducible and scalable MCIT experimentation. Code is available at https://github.com/LAMDA‑CL/Prism.
Authors:Jiraphon Yenphraphai, Jianqi Chen, Jian Wang, Gordon Qian, Sergey Tulyakov, Rameen Abdal, Raymond A. Yeh, Peter Wonka, Chaoyang Wang
Abstract:
Current video‑to‑4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Helix4D, a dynamic mesh generation framework by inheriting the expressive representation of Trellis2, adapting it from image‑to‑3D to video‑conditioned 4D generation. Our design arises from two key questions: (a) how to enable Trellis2's frame‑local attention to share information across frames while preserving its pretrained quality on rare cases such as transparent objects and inner surfaces, and (b) how to inject temporal information into a purely 3D positional encoding without breaking pretrained capabilities. We address (a) with a sliding‑window cross‑frame attention and anchor on the first frame. The first frame is generated by the base Trellis2 model and injected into our model, letting it inherit Trellis2's quality in rare cases through cross‑frame attention. We address (b) with a 4D temporal encoding that repurposes redundant low‑frequency spatial RoPE bands for time, extending the encoding from 3D with no additional parameters. Extensive experiments show the effectiveness of Helix4D for high‑quality dynamic mesh generation on ActionBench and our own challenging complex dynamics set.
Authors:Yushi Huang, Xiangxin Zhou, Ruoyu Wang, Chi Zhang, Jun Zhang, Tianyu Pang
Abstract:
Recent advances in few‑step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains challenging. We propose Reward‑Tilted Distribution Matching Distillation (RTDMD), a two‑stage framework that unifies distribution matching distillation with reward‑guided reinforcement learning for few‑step flow generators. We show that minimizing the KL divergence to a reward‑tilted teacher distribution naturally decomposes into a distribution matching term and a reward maximization term. In the first stage, we introduce Ambient‑Consistent Distribution Matching Distillation (AC‑DMD), which performs subinterval‑wise distribution matching and augments the fake score objective with a consistency regularizer to help the fake score model track the shifting generator distribution under limited updates. In the second stage, we jointly optimize both terms: for the reward maximization term, we derive a hybrid policy gradient that combines a GRPO‑style estimator for the stochastic intermediate transitions with direct reward backpropagation through the deterministic final step, and further introduce step‑subset GRPO (SubGRPO) to reduce variance. Experiments on SD3, SD3.5, and FLUX.2 demonstrate that RTDMD establishes new state‑of‑the‑art results across preference, aesthetic, and compositional metrics with only 4 inference steps, outperforming previous few‑step text‑to‑image generation methods. Code and models are available at https://github.com/Harahan/RTDMD.
Authors:Linfei Pan, Johannes Schönberger, Marc Pollefeys
Abstract:
Structure‑from‑Motion ‑‑ the process of simultaneously estimating camera poses and 3D scene structure from a collection of images ‑‑ remains a central challenge in computer vision, with many open problems yet to be solved. Recent advances in feedforward 3D reconstruction have made significant strides in overcoming persistent failure cases of classical SfM methods, particularly in scenarios characterized by low texture, limited overlap, and symmetries. However, while feedforward approaches excel in these challenging conditions, they often face limitations regarding scalability, accuracy, or robustness, and typically fall short of classical methods in standard reconstruction settings. In this work, we systematically analyze these limitations and propose a new Structure‑from‑Motion pipeline by combining the respective strengths of classical and feedforward methods. Extensive experiments across multiple datasets show the benefits of our approach, achieving state‑of‑the‑art results across a wide range of scenarios. We share our system as an open‑source implementation at https://github.com/colmap/gluemap.
Authors:Xinrui Shi, Kai Liu, Ziqing Zhang, Jianze Li, Anqi Li, Yulun Zhang
Abstract:
Lightweight vision‑language models perform competitively on standard benchmarks yet fail systematically in dense‑scene reasoning, where multiple objects, attributes, and relations must be jointly grounded and resolved through multi‑step inference. Such capability is critical for real‑world applications where models must reliably interpret cluttered environments. Yet existing training signals provide no explicit grounding between reasoning steps and the underlying visual entities and relations, leaving lightweight models free to generate fluent but visually unanchored reasoning chains. To address this gap, we first introduce DRBench, a benchmark of 14,573 questions across 2,943 images, organized into five task categories spanning three progressive reasoning layers. Building on DRBench, we propose DRScaffold, a supervised fine‑tuning framework that decomposes the supervision target into four causally ordered stages, enforcing grounded reasoning without architectural modification. Experiments on three lightweight VLMs demonstrate substantial gains on DRBench while preserving or improving performance on general‑purpose benchmarks. Notably, Qwen2.5‑VL‑3B trained with DRScaffold surpasses the frozen Qwen2.5‑VL‑32B on DRBench, demonstrating that structured supervision can substitute for a significant portion of model scale in dense‑scene reasoning. Our code and models are available at https://github.com/irene‑shi/DRScaffold .
Authors:Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold
Abstract:
Light sheet fluorescence microscopy (LSM) enables high‑resolution, three‑dimensional (3D) imaging of biological specimens, providing rich volumetric data for studying cellular organization, pathology, and vascular networks. However, the size, dimensionality, and annotation burden of LSM data make supervised deep learning approaches costly and difficult to scale. Additionally, despite the abundance of unannotated LSM volumes, foundation models for this modality remain underexplored due to computational challenges and the complexity of volumetric representation learning. In this work, we introduce a 3D foundation model for LSM data, pretrained on a large curated collection of 3D images spanning multiple organisms, stains, and imaging protocols. We learn transferable volumetric representations by jointly optimizing for masked reconstruction and image‑text alignment. The pretrained backbone drastically reduces the annotation burden, enabling efficient, few‑shot adaptation for varied downstream tasks. We evaluate this approach on downstream segmentation, classification, and deblurring. Our results demonstrate consistent improvements over baselines, (1) when measured using standard evaluation metrics and (2) when rigorously assessed by domain experts. This highlights the potential of foundation model pretraining to reduce annotation requirements while improving performance across diverse LSM analysis tasks. Pretrained model weights and code for pretraining and finetuning are publicly available: https://github.com/AdinaScheinfeld/lsm_fm_public_repo.git.
Authors:Haoyang Peng, Qian Hu, Songan Zhang, Ming Yang
Abstract:
Predicting the interaction between pedestrian and vehicle is essential for autonomous driving safety in unstructured and semi‑structured scenarios; however, this task is severely hindered by the scarcity of public datasets that feature dense pedestrian‑vehicle interactions. Most current studies rely on structured road data, leaving the complex, heterogeneous interactions found in unstructured environments insufficiently represented and researched. In this paper, we propose a dataset annotation framework based on video data from uncalibrated surveillance cameras and present PINNS (Pedestrian‑vehicle Interaction dataset from uNcalibrated cameras in uNstructured Scenes). The dataset covers multiple countries and regions, includes diverse typical traffic scenarios, and considers variations in seasons, lighting conditions, and weather. It focuses on complex scenes with dense pedestrian‑vehicle interactions and is designed to be easily extensible. The dataset is constructed and annotated according to the standard issued by the Chinese Association of Automation, providing both trajectory data and corresponding scene‑level information. Furthermore, this paper analyzes current challenges and research directions in heterogeneous agent trajectory prediction, shows the necessity and usefulness of the proposed dataset. We hope our framework and dataset will facilitate research on trajectory prediction and autonomous driving in complex mixed traffic scenarios. PINNS is publicly available at https://github.com/Songan‑Lab.
Authors:Ruiqiang Xiao, Zhaohu Xing, Yijun Yang, Zhenyan Han, Weiming Wang, Kaishun Wu, Lei Zhu
Abstract:
Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent promptable foundation models enable point‑guided segmentation, but their direct deployment in ultrasound remains unreliable: a single point provides insufficient spatial context to resolve scale ambiguity, and greedy memory updates amplify early errors into severe temporal drift. We present EchoPilot, a training‑free framework for ultrasound video segmentation under sparse first‑frame interaction, requiring only a single point click and an anatomical category name. EchoPilot orchestrates a frozen medical vision‑language model (VLM) for semantic localization, a vision foundation model (VFM) for dense geometric feature extraction, and a promptable video segmentor for mask prediction and propagation. To resolve initialization ambiguity, we propose Scale‑Space Semantic Prompting, which first selects an optimal contextual view via a parameter‑free S.E.E.D. (Semantic Energy‑Entropy Density) criterion, and then synthesizes geometrically precise auxiliary point prompts from dense foundation features without additional user interaction. To reduce propagation drift, a Reliability‑Gated Memory update is further introduced to selectively freeze the segmentor's memory bank under uncertain predictions, preventing error accumulation. We also contribute the first dynamic fetal placenta ultrasound video segmentation dataset with 671 annotated frames. Across three ultrasound video datasets, EchoPilot achieves state‑of‑the‑art performance under the sparse‑interactive setting, consistently outperforming training‑free baselines and finetuned specialists.
Authors:Benjamin Herb, Steve Göring, Alexander Raake, Rakesh Rao Ramachandra Rao
Abstract:
Recent video super‑resolution (VSR) approaches use deep neural networks to enhance low‑quality input videos and recover visual detail, with diffusion‑based methods in particular showing promising results. In this paper, we investigate whether existing video quality models can be used to assess the performance of these diffusion‑based VSR methods, by comparing model predictions with results from a subjective test. The study compares six upscaling methods (Lanczos, Rhea, SCST, DOVE, SeedVR2, Starlight Mini) applied to both compressed (AV1 and DCVC‑RT) and uncompressed low‑resolution videos considering the play‑out on a UHD‑1/4K screen. A range of full‑ and no‑reference quality models are used to assess their applicability to this new type of quality degradation, focusing on within‑sequence performance. The results highlight that CNN‑based full‑reference models, such as LPIPS, DISTS, and CVQA‑FR show significantly higher correlation coefficients than both conventional full‑ as well as the tested no‑reference models. Most overestimate the overly sharp results of SCST, with VMAF mainly failing due to spatial inconsistencies introduced by Starlight Mini. None of the tested video quality models reach sufficient accuracy so as to replace complementary subjective testing. The reference, degraded and upscaled videos, as well as the user ratings and model scores are made available with the paper at https://github.com/Telecommunication‑Telemedia‑Assessment/AVT‑VQDB‑UHD‑1‑VSR as open data.
Authors:Denis Gridusov, Maxim Popov, Sergey Kolyubin
Abstract:
Reconstructing and predicting dynamic 3D scenes from multi‑view videos is a foundational task for robotics, AR/VR, and digital twins. Recent physics‑informed Gaussian Splatting methods achieve impressive future frame extrapolation but lack semantic awareness and suffer from large computational overhead. We introduce R5DGS, a framework that augments a physics‑driven 4D Gaussian representation with compact Identity Encoding vectors, enabling precise Gaussian‑to‑object association. By constructing an offline CLIP‑based object lookup table, we support open‑vocabulary text prompting to retrieve and render object‑specific Gaussians across arbitrary timestamps and viewpoints. Furthermore, we propose a rigid‑body inference constraint that predicts and integrates physical dynamics exclusively for object centroids, propagating motion to associated Gaussians via relative transformations. This optimization yields a 11 FPS speedup during extrapolation without compromising trajectories plausibility.
Authors:Cuong Huynh, Maxim Popov, Denis Gridusov, Sergey Kolyubin
Abstract:
3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language descriptions. Recent zero‑shot methods leverage 2D vision‑language models (LVLMs). However, they often rely on existing sets of multi‑view images and struggle with the limited semantic and spatial details provided by standard 3D segmentation tools. We present AgentGrounder, a zero‑shot 3D visual grounding framework that operates directly on colored point clouds without task‑specific 3D training. Our approach follows a two‑stage design: (1) an offline stage that applies 3D model to build an Object Lookup Table (OLT) with instance IDs, semantic labels, 3D bounding boxes; and (2) an online tool‑driven agent that decomposes each query, retrieves only relevant candidates from the OLT, performs geometric scoring, and triggers image rendering on demand when additional visual evidence (e.g., color, material, or viewpoint‑sensitive cues) is required. Compared with fixed anchor‑target matching pipelines, this design reduces cascading matching errors and improves context‑window efficiency by avoiding prompts overloaded with irrelevant objects. We evaluate on ScanRefer and Nr3D under a zero‑shot setting and observe consistent improvements over SeeGround in our setup, including +2.5% Acc@0.5 on ScanRefer and +6.3% on Nr3D, with a notable +6.3% gain on Nr3D view‑independent queries. These results show that combining selective retrieval, geometric reasoning, and adaptive visual inspection yields a practical and robust foundation for open‑vocabulary 3D grounding. Our code is available at https://github.com/be2rlab/AgentGrounder.
Authors:Kaining Ying, Hengrui Hu, Siyu Ren, Jiamu Li, Fengjiao Chen, Ziwen Wang, Xuezhi Cao, Xunliang Cai, Henghui Ding
Abstract:
Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill this gap, we introduce WBench, a comprehensive multi‑turn benchmark for interactive world model evaluation along five dimensions, namely video quality, setting adherence, interaction adherence, consistency, and physics compliance. WBench contains 289 test cases and 1,058 interaction turns, where each case specifies a world setting and a multi‑turn interaction sequence, covering diverse scenes, styles, subjects, and both first‑ and third‑person perspectives, together with four interaction types, including navigation, subject action, event editing, and perspective switching. For navigation, WBench unifies text, 6‑DoF pose, and discrete‑action control, enabling evaluation of models with different native input interfaces. Evaluation uses 22 automatic sub‑metrics that combine specialist vision models with large multimodal models, and all metrics are validated against human judgments. Across 20 state‑of‑the‑art models, we find that no single model performs strongly across all dimensions. We provide detailed diagnostic insights into the characteristic strengths, weaknesses, and open challenges of each model. Code and data are available at https://github.com/meituan‑longcat/WBench.
Authors:Yunqi Gao, Leyuan Liu, Yuhan Li, Changxin Gao, Jingying Chen
Abstract:
3D human mesh recovery and 3D clothed human reconstruction are inherently related, yet they have long been studied in isolation, thereby overlooking the potential gains of joint optimization. To overcome this limitation, we propose to address these two tasks within a unified framework, which allows their mutual dependencies to be effectively exploited. Building on this idea, we propose MuNet, a mutualistic network for joint 3D human mesh recovery and 3D clothed human reconstruction from single images. First, we adopt 2‑manifold graphs as a unified representation for all 3D models, enabling consistent modeling across 3D human mesh recovery and clothed human reconstruction. Second, we design an end‑to‑end graph convolutional network that progressively deforms an initial graph into a 3D human mesh and refines it into a detailed 3D clothed human model. Third, we introduce a mutualistic mechanism that allows reciprocal interaction between the two tasks during training, where 3D human mesh recovery provides guidance for 3D clothed human reconstruction, and reconstruction feedback refines the 3D human mesh recovery. We extensively evaluate MuNet on six benchmark datasets for 3D human mesh recovery and 3D clothed human reconstruction, including Human3.6M, 3DPW, MPI‑INF‑3DHP, THuman2.0, CAPE, and RenderPeople. Experimental results demonstrate that MuNet achieves state‑of‑the‑art performance on both tasks across all datasets. The code of MuNet is released for research purposes at https://github.com/starVisionTeam/MuNet.
Authors:Akang Wang, Xili Deng, Zhanxuan Hu, Yi Zhao, Yonghang Tai, Huafeng Li
Abstract:
Vision‑Language Models such as CLIP exhibit strong zero‑shot recognition capability by aligning images with textual concepts, yet they often underperform on multi‑label recognition where multiple objects co‑exist. A key bottleneck is that the [CLS] token, as a single global visual representation, is insufficient to faithfully encode diverse targets with varying scales, contexts, and co‑occurrence patterns. To address this limitation, we present a new multi‑label image recognition framework, termed PIAA, which formulates prediction as Patch‑level Inference followed by Adaptive Aggregation. Specifically, we first enhance patch‑wise predictions from two complementary perspectives: (i) mitigating semantic entanglement in the visual encoder to obtain more discriminative patch representations, and (ii) learning an unsupervised visual classifier to narrow the vision‑language modality gap. We then introduce an adaptive aggregation module that consolidates patch‑level scores into the final multi‑label prediction. Notably, the entire pipeline is fully training‑free, requiring no gradient updates or parameter fine‑tuning. Experiments show that our method achieves strong improvements with minimal extra computation, exceeding a 6% mAP gain on the challenging NUS‑WIDE benchmark over representative baselines. Code is available at https://github.com/akang‑wang/PIAA.
Authors:Shuai Yi, Yixiong Zou, Yuhua Li, Ruixuan Li
Abstract:
Vision‑language models (VLMs) like CLIP have shown impressive generalization capabilities, yet their potential for Cross‑Domain Few‑Shot Learning (CDFSL) remains underexplored, where the model needs to transfer source‑domain information to target domains with scarce training data. While the attention sink phenomenon has been observed in VLMs for certain tasks, its role in CDFSL scenarios has not been studied. In this paper, we uncover a critical issue overlooked by prior works: standard target‑domain few‑shot fine‑tuning in CDFSL significantly exacerbates the attention sink problem, leading to poor discriminability across classes. To understand this phenomenon, through extensive experiments, we interpret it as the model's shortcut learning for domain adaptation: to overcome the huge domain gap between the source and target domains, the model shows a high tendency to push tokens that are initially closer to target‑domain classes (i.e., simple tokens) to be even closer to these classes, exacerbating the attention sink and wasting the capability of learning other discriminative but initially further tokens (i.e., hard tokens). To address this, we propose a novel approach to dynamically re‑weight tokens according to their relevance with target‑domain classes during the target‑domain finetuning, which explicitly suppresses the model's reliance on these simple tokens and enhances the learning of hard tokens, reducing sink tokens and enhancing discriminability. Extensive experiments on four benchmark datasets validate the rationale of our method, demonstrating new state‑of‑the‑art performance. Our codes are available at https://github.com/shuaiyi308/TIR.
Authors:Bokai Zhao, Yiyang Zhang, Yuanchi Zhu, Hanqing Chao, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang
Abstract:
Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they are typically benchmarked through downstream clinical endpoints. While such task level evaluations are indispensable, they offer limited insight into what the representations themselves encode, particularly whether PFM embeddings can distinguish meaningful tissue regions and capture their spatial relationships. We present SpaPath‑Bench, a representation level benchmark designed to diagnose spatial representation capability in PFMs. SpaPath‑Bench formulates spatial domain identification (SDI) on paired whole slide image and spatial transcriptomics (ST) data as a diagnostic task. It curates 42 public paired WSI and ST slides, enables large scale evaluation across 19 encoders and seven SDI methods, and measures partition quality using three complementary criteria: unsupervised spatial coherence, transcriptomics referenced agreement, and expert referenced agreement. Across 83K runs, SpaPath‑Bench reveals that different pretraining paradigms capture distinct aspects of tissue spatial architecture, and it provides practical guidance for building the next generation of spatially aware computational pathology models. Code and data pipelines are publicly available at https://bokai‑zhao.github.io/SpaPath‑benchboard/.
Authors:Shipeng Cao, Biao Qian, Haipeng Liu, Yang Wang, Meng Wang
Abstract:
Text‑to‑image synthesis has made significant progress, benefiting from the strong generative capabilities of diffusion models. However, these models struggle to achieve precise text‑to‑image alignment within cross‑attention maps during the denoising process. Existing works primarily focus on inter‑subject‑token activations (i.e., cross‑attention scores) overlap for different subjects, overlooking the intra‑subject‑token activations scattering issue for identical subjects. In this paper, we propose an Aggregating‑and‑Isolating cross‑attention approach to diffusion models for Text‑to‑Image synthesis, dubbed AI‑T2I. Technically, to address the scattering issue, we devise an aggregation loss to identify and consolidate the scattered intra‑token activations, which implicitly helps mitigate the potential overlap issue. Upon that, an isolation loss is further introduced to push the inter‑token activations apart, thus fulfilling precise text‑to‑image alignment. Extensive experiments on various benchmarks demonstrate the superiority of AI‑T2I over the state‑of‑the‑art works for text‑to‑image synthesis. Furthermore, our AI‑T2I exhibits excellent generalization across other tasks, e.g., controllable layout generation and personalized generation. Our code is available at https://github.com/Hatter77/AI‑T2I.
Authors:Chuyu Zhong, Keyan Chen, Qinzhe Yang, Bowen Chen, Zhengxia Zou, Zhenwei Shi
Abstract:
Pixel count and geographical coverage are two key characteristics of remote sensing images. Existing remote sensing image segmentation methods typically focus on images with either a small pixel count or a large pixel count but limited geographical coverage. In this paper, we introduce a novel segmentation task targeting ultra‑wide area (UWA) remote sensing images, characterized by both a large pixel count and extremely wide geographical coverage. The core challenges of UWA segmentation lie in simultaneously handling ground objects with significantly varying scales and maintaining long‑range contextual semantic continuity. To address these challenges, we propose the Scale‑Frustum Representation Network (SFR‑Net). Inspired by the viewing frustums of remote sensing images captured from different altitudes, we construct scale‑frustum representations, enabling unified modeling of ground objects and contextual features at different scales. Furthermore, we design a cascaded cross‑scale fusion mechanism to effectively integrate these representations, enhancing local semantic understanding while ensuring long‑range contextual continuity. Experimental results on GID and FBPS demonstrate that SFR‑Net achieves state‑of‑the‑art performance, improving mIoU by 1.72% and 4.29%, respectively, over the strongest competing methods. In addition, the proposed scale‑frustum representations can be integrated into generic segmentation networks to improve both segmentation accuracy and convergence speed. The implementation code will be publicly available at https://github.com/ChuyuZhong/SFR‑Net.
Authors:Zongjian Wu, Lei Zhang
Abstract:
Referring expression comprehension (REC) aims to localize a target object within an image based on a given expression. Although recent advances in vision‑language models have led to substantial improvements in REC tasks, current REC benchmarks often hold simple scenarios and the assumption that each expression maps to a unique object. These limitations hinder the deployment of REC models in open‑world environments. To fill this gap, we introduce OpenRef, a new benchmark for REC in complex visual and linguistic scenarios. OpenRef features three key advancements: 1) Diverse visual scenarios: spanning diverse visual domains, including ground views, drone views, dark scenes and adverse weather conditions; 2) Variable target counts: breaking the single‑target limitation with multi‑target and none‑target samples; 3) Rich vocabulary types: incorporating proper nouns, polysemous words and ordinal terms to fit a wider range of expression needs. Furthermore, as traditional metrics are insufficient for open‑world setting, we leverage F1 to measure grounding accuracy and propose N3R (Negative Relative Rejection Reliability) to assess relative rejection reliability against negative expressions. Finally, we introduce Multi‑task Consistency Checker (MCC), a training‑free but plug‑and‑play strategy that enhances model performance with one click by enforcing consistency self‑verification. Extensive experiments demonstrate that this work significantly advances the performance of existing REC models in complex scenarios, paving the way for open‑world REC. Project page: https://zongjianwu.github.io/openref
Authors:Florent Tariolle, Florian Yger
Abstract:
Black‑box adversarial attacks that minimize only the ground‑truth confidence suffer from class drift: perturbations wander through the feature space without committing to a specific adversarial class, wasting queries on diffuse, undirected progress. We introduce Opportunistic Target Selection (OTS), a lightweight wrapper that switches an untargeted attack to a targeted objective early in its trajectory, locking onto whichever non‑true class currently leads. OTS requires no architectural modification to the underlying attack, no gradient access, and no a priori target‑class knowledge.
We validate OTS on three score‑based attacks (SimBA, Square Attack with cross‑entropy loss, and Bandits) across five standard ImageNet classifiers (4,500 runs). On random‑search attacks, OTS closely tracks oracle performance, with gains up to +27 pp in success rate and 43% relative reduction in censored‑mean iterations on ResNet‑50. On gradient‑estimation attacks (Bandits) and attacks with margin loss, OTS is redundant, a negative result that reinforces our interpretation of OTS as a margin‑loss surrogate. On adversarially‑trained models, a bimodal difficulty distribution eliminates the regime where targeting helps.
Authors:Jaxon Zhang, Binxin Yang, Hubery Yin, Chen Li, Jing Lyu
Abstract:
Current mainstream methods of aligning diffusion models with human preferences typically employ VLM‑based reward models. However, these reward models, pre‑trained for semantic alignment, struggle to capture the essential perceptual qualities‑such as aesthetics, composition, and visual harmony. In this work, we argue that a model capable of high‑fidelity generation must possess a profound understanding of these visual attributes. Based on this insight, we introduce the Diffusion‑based Reward Model (DRM), a novel paradigm that use the pre‑trained diffusion model as a powerful evaluative backbone. A key advantage of the DRM is its unique ability to assess not only the final image but also the noisy intermediate latents at any stage of the generative process. We leverage this step‑wise evaluative capacity in two ways. First, we propose Step‑wise GRPO, a reinforcement learning algorithm that provides dense, per‑step rewards to resolve the imprecise credit assignment problem in GRPO algorithm, leading to more stable and effective alignment. Second, we introduce Step‑wise Sampling, a novel inference strategy that employs the DRM as a dynamic guide to evaluate multiple generation paths at each step, steering the process towards higher‑quality outcomes. Extensive experiments confirm that our approach significantly enhances the final quality of generated images. Code: https://github.com/jjaxonx/DRM.
Authors:Sizhe Zhao, Shengping Zhang, Shuo Yang, Weiyu Zhao, Shuigen Wang, Xiangyang Ji
Abstract:
Existing embodied control research demonstrates remarkable performance improvements by scaling training data and model size. We instead explore inference‑time strategy as an alternative axis. Non‑deterministic generative models, such as diffusion and autoregressive models, have been widely adopted in the field of embodied control. However, the single‑shot inference paradigm limits their performance. In this paper, we propose TapSampling, a plug‑and‑play framework for inference‑time sampling. First, we introduce an Action‑VAE that represents actions in a low‑dimensional latent space by mapping policy‑generated initial actions into a compressed posterior distribution, from which any number of latent samples can be drawn and decoded into candidate actions that approximate the true action distribution. Second, we formulate action verification as task‑progress outcome prediction, using the intrinsic sequential structure of robotic datasets to train a semantically grounded verifier for interpretable action selection. Furthermore, TapSampling is a policy‑agnostic framework. Extensive experiments in both simulated and real‑world environments demonstrate that our method substantially improves multiple generalist policies without further policy finetuning. Code and models are available at the project page.
Authors:Jiayi Kong, Xuhui Chen, Chen Zong, Fei Hou, Junhui Hou, Wenping Wang, Ying He
Abstract:
Neural Signed Distance Functions (SDFs) excel at reconstructing watertight manifolds but fail on thin structures and open boundaries due to strict inside‑‑outside constraints. Conversely, Unsigned Distance Fields (UDFs) accommodate general geometries but suffer from gradient singularities at the zero‑level set, hindering optimization and extraction. We introduce Metric‑‑Phase Fields (MPFs), a decoupled implicit representation that separates metric proximity from topological phase. Given an unoriented point cloud, MPFs learn (i) an unsigned metric field r and (ii) a smooth phase field θ, for which we derive a bounded phase indicator P=\tanh(βθ) that provides soft inside‑‑outside cues where they are meaningful. We couple the two fields via a gated‑metric formulation with a residual phase injection to obtain a signed implicit function with stable near‑surface gradients. The phase coefficient β is learnable, allowing MPFs to adaptively control the sharpness of the phase transition and the degree of saturation of the soft sign indicator. Experiments on both synthetic and scanned thin‑shell and thin‑plate shapes demonstrate that MPFs preserve thin and layered structures more faithfully than recent SDF‑based methods, while also enabling more robust training and more reliable surface extraction than UDF‑based approaches. Check out \hrefhttps://github.com/JIAYI‑Scarlett/ICML2026‑MPFMPFs‑GitHub for source code and test models.
Authors:Ting-Hsuan Chen, Ying-Huan Chen, Tao Tu, Jie-Ying Lee, Cho-Ying Wu, Fangzhou Lin, Hengyuan Zhang, David Paz, Xinyu Huang, Yuliang Guo, Yu-Lun Liu, Yue Wang, Liu Ren
Abstract:
Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial‑temporal consistency constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi‑view trajectories, amplifying cross‑view inconsistency and temporal drift. We argue that 360° video generation offers a natural solution: panoramic coverage simplifies trajectory design and provides a strong global context for maintaining coherence. We introduce Pantheon360: Taming Digital Twin Generation via 3D‑Aware 360° Video Diffusion, a controllable 360° video generation framework that synthesizes high‑fidelity videos from sparse 360° inputs. The key idea is an explicit 3D Cache, reconstructed from the input, which serves as a geometric scaffold for any user‑defined camera path. This allows the diffusion model to focus on photorealistic texture refinement while the 3D Cache enforces global geometric consistency. Experiments show that Pantheon360 achieves superior visual quality and unmatched geometric coherence, enabling reliable and flexible 360° scene generation for downstream simulation and digital‑twin applications.
Authors:Eyal Hanania, Nadav Kirsch, Daniel Arkushin, Jonathan Benvenisti, Amos Bercovich, Elie Zemmour, Sahar Froim
Abstract:
Detecting laughter in video is essential for affective computing and narrative understanding, yet existing approaches treat it as coarse clip‑level classification, failing to capture precise temporal boundaries of brief, transient laughter events. We address this gap with two complementary contributions.
First, we introduce UR‑FUNNY‑Temporal and SMILE‑Temporal, fully annotated temporal laughter datasets extending two widely‑used humor benchmarks. Our annotations cover over 11,053 videos (78.8 hours) and provide precise onset/offset boundaries for each laughter event, along with rich metadata distinguishing speaker vs. audience laughter, modality dominance (acoustic, visual, or both), and intensity levels.
Second, we propose a lightweight weakly‑supervised framework for temporal laughter localization. Our architecture combines fixed HuBERT and MAE encoders with temporal softmax pooling and adaptive modality gating, learning fine‑grained temporal grounding from clip‑level labels without requiring frame‑level annotations during training. Experiments across three datasets demonstrate that our approach substantially outperforms multimodal foundation models including Gemini 3 Flash, achieving 99% F1 and 68.1% localization precision on sports broadcast data. Ablations validate each architectural component. Furthermore, our precise temporal tags improve downstream laughter reasoning by 227% on CIDEr, enabling GPT‑3.5 to outperform GPT‑4o. The code, UR‑FUNNY‑Temporal and SMILE‑Temporal datasets are publicly available at https://github.com/WSCSports/MTLLFM‑temporal‑laughter‑localization.
Authors:Chunzheng Zhu, Jianxin Lin, Feng Wang, Cheng Jiang, Guanghua Tan, Zhenyu Zhou, Shengli Li, Kenli Li
Abstract:
Reliable quality control (QC) of ultrasound images is essential for both real‑time acquisition guidance and retrospective clinical audit, yet existing approaches rely heavily on per‑plane annotations, or employ pseudo‑labeling prone to systematic bias under spatial deformations inherent in clinical acquisition. We present STRIQ, a registration‑driven framework that recasts annotation‑free US plane quality control as a subspace‑guided consistency measurement problem. Specifically, STRIQ introduces a Latent Registration Aligner (LRA) to establish hierarchical feature space correspondences between query images and variance‑driven anchors, which are autonomously distilled from unlabeled data via a variance spectrum criterion to serve as structurally stable prototypes. To further disambiguate anatomical planes and mitigate negative knowledge transfer, we propose an Orthogonal Knowledge Subspace (OKS) module. The OKS decomposes plane‑specific representations into mutually orthogonal subspaces, enabling fine‑grained expert collaboration while preventing inter‑plane interference, ensuring that the quality metric is grounded in principled subspace proximity. Extensive experiments on the in‑house US4QA and public CAMUS datasets demonstrate that STRIQ achieves state‑of‑the‑art correlation with clinical quality scores, establishing a new paradigm for annotation‑free, real‑time reliable ultrasound quality control. Our code is available at https://github.com/zhcz328/STRIQ.
Authors:Fangtai Wu, Hailong Guo, Shijie Huang, Jiayi Song, Yubo Huang, Mushui Liu, Zhao Wang, Yunlong Yu, Jiaming Liu, Ruihua Huang
Abstract:
Customized image editing aims to equip pre‑trained diffusion models with specific visual effects using limited paired data, typically via Low‑Rank Adaptation (LoRA). As the number of desired effects grows, storing and dynamically loading numerous these effect LoRAs significantly increases deployment overhead. Furthermore, current pipelines typically cascade these effect LoRAs with acceleration modules for fast generation, which triggers severe parameter interference and results in concept bleeding and style degradation. We propose CollectionLoRA, a multi‑teacher on‑policy distillation framework capable of distilling the concepts of up to 50 different effect LoRAs along with few‑step generation capabilities into a single LoRA. This fundamentally resolves the feature interference issue and significantly reduces deployment costs. Specifically, the method introduces (i) a Probabilistic Dual‑Stream Routing mechanism that enables the model to randomly switch between data sources during training, effectively enhancing its generalization in unseen scenarios; (ii) an Asymmetric Orthogonal Prompting strategy to achieve concept isolation within the prompt space; (iii) a Coarse‑to‑Fine Distillation Objective to mitigate the distribution gap between the teacher and student models. Extensive evaluations show that CollectionLoRA distills all customized effects and few‑step generation into a single LoRA, reducing deployment overhead while achieving concept fidelity comparable to or better than independently trained teacher models. Code: https://github.com/Qwen‑Applications/CollectionLoRA
Authors:Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai, Yiyan Huang, Ranjie Duan, Tianle Zhang, Xu Yang, Ziyi Ye, Xingjun Ma
Abstract:
Large Vision‑Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts. Existing mitigation methods either rely on costly external interventions, such as instruction tuning and retrieval, or use internal mechanisms that remain limited by flawed attention weights and entangled hidden representations. We propose Adversarial Orthogonal Disentanglement (AOD), a latent geometric framework for mitigating LVLM hallucinations. AOD learns a hallucination‑related direction through a minimax objective: a classifier concentrates hallucination signals into the projected component, while an adversary removes them from the orthogonal residual space via a Gradient Reversal Layer. The learned direction enables a training‑free dual‑forward‑pass contrastive decoding strategy that suppresses hallucinations while preserving general capabilities. Experiments on three LVLMs across four hallucination and four utility benchmarks show that AOD consistently outperforms strong baselines. It improves POPE accuracy by over 6% on average, boosts AMBER by 6%, and maintains strong performance on utility tasks such as MMMU. Further analysis shows robust transfer across datasets, suggesting that AOD captures general hallucination‑related biases rather than dataset‑specific artifacts. Our source code and datasets are available at https://github.com/Hunter‑Wrynn/AOD.
Authors:Longteng Guo, Yifan Wang, Pengkang Huo, Tailai Chen, Yuze Wu, Jing Liu, Xinxin Zhu
Abstract:
Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning directly grounded in visual evidence. We introduce VisReason, a benchmark for vision‑centric reasoning in everyday scenarios where perception and inference are tightly coupled. VisReason contains 1,505 questions across 10 categories spanning perceptual, structural, and conceptual reasoning. Our evaluation shows that VisReason poses a qualitatively different challenge from existing benchmarks, exposing substantial gaps between humans and current MLLMs and revealing limited benefits from test‑time reasoning strategies. VisReason offers a focused diagnostic for evaluating vision‑centric reasoning beyond language.
Authors:Renjie Lu, Xulong Zhang, Xiaoyang Qu, Shangfei Wang, Jianzong Wang
Abstract:
Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge that lies in inductive biases induced by distinct supervision signals: generation branch prefers high‑fidelity, fine‑grained representations capable of reconstruction, while the understanding favours semantically discriminative embeddings that remain invariant to task‑irrelevant factors. Consequently, optimizing these complementary but non‑equivalent objectives within a monolithic backbone leads to mutual impairment instead of enhancement. In this paper, we first analyze the root cause of this interference in unified backbones and reveal a complementary structure in their internal representations. Motivated by the observation, we propose DIVA, a self‑improved post‑training framework that transforms the representation divergence into interior synergy. By explicitly factorizing the visual representation into shared and unique components based on two complementary information flow, DIVA enables both the understanding and generation branches to achieve beneficial transferring while preserving the integrity of unique information from cross‑flow interference via mutual information estimation. Despite its generality, our method consistently achieves improvements across visual understanding (+7.82%) and generation (+8.46%). The official code is available at: https://github.com/Jayyy‑H/DIVA.
Authors:Xiaoyang Lyu, Muxin Liu, Xiaoshan Wu, Ruicheng Wang, Yi-Hua Huang, Yang-Tian Sun, Shaoshuai Shi, Xiaojuan Qi
Abstract:
Consistent 3D geometry estimation from streaming RGB input is crucial for real‑world applications such as autonomous driving, embodied AI, and large‑scale reconstruction. While modern monocular geometry foundation models achieve strong single‑image accuracy, they exhibit severe temporal inconsistency on continuous input, notably dominated by scale‑‑shift drifting. Through targeted empirical analysis, we trace this instability to its root cause: fluctuations in latent feature statistics, whose mean and variance directly determine the predicted depth's scale and shift. Building on this insight, we introduce Dynamic Feature Normalization (DyFN), a lightweight, causal recurrent module that dynamically and robustly modulates feature statistics to maintain stable geometry over time. We adapt powerful pretrained monocular geometry models for streaming by finetuning only DyFN, a mere 2% additional parameters, while keeping the backbone frozen, thereby achieving temporal consistency without compromising single‑image accuracy. Extensive experiments across four benchmarks show that DyFN effectively eliminates temporal artifacts such as disjointed layering and positional jitter, and achieves state‑of‑the‑art temporal stability, improving over prior streaming methods by up to 14% and even outperforming heavier non‑causal video baselines. Project Page: https://shawlyu.github.io/DyFN
Authors:Aviral Chharia, Fernando De la Torre
Abstract:
High‑fidelity 3D Gaussian head avatar generation is critical for applications such as AR/VR, telepresence, and digital humans. Existing methods depend on multi‑view datasets, 3D captures, or intermediate 2D view synthesis. In contrast, we learn both conditional and unconditional 3D head models from randomly sampled 2D images alone, without using multi‑view data, 3D supervision, or intermediate view generation. We introduce MVCHead, a single‑shot state space model that enforces multi‑view consistency (MVC) directly in the 3D representation while regressing 3D Gaussians under these constraints. At its core, we propose a Hierarchical State Space (HiSS) block that progressively refines Gaussians from coarse to fine, while capturing long‑range dependencies. Within each HiSS block, we modify Mamba's standard unidirectional scan with the proposed Hierarchical Bi‑directional State Scan (HiBiSS) that aligns recurrence with the axes along which multi‑view inconsistencies are strongest. Finally, we design an SE(3) Multi‑view Critic that judges whether a set of self‑renders arises from a single underlying 3D configuration, rewarding cross‑view pixel alignment without observing real multi‑view pairs. MVCHead achieves state‑of‑the‑art perceptual quality, surpasses prior methods in both texture and geometric consistency, and maintains comparable shape consistency. To demonstrate scalability, we release FaceGS‑10K, the first large‑scale dataset of ready‑to‑use 3D Gaussian head assets for training and evaluation of 3D head models. Project Page and code: https://humansensinglab.github.io/MVCHead/
Authors:Bohai Gu, Taiyi Wu, Yueyang Yuan, Jian Liu, Xiaocheng Lu, Dazhao Du, Jie Zhang, Jinxiang Lai, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo
Abstract:
Recent video‑based world models have made pixel‑space environments interactive at the camera level: users can navigate viewpoints while the model generates coherent visual continuations. Yet their action spaces remain incomplete: users can move the camera, but cannot act on individual objects. Since real‑world interaction is inherently object‑centric, such models remain closer to passive scene observers than truly manipulable environments. We present WorldCraft, a framework that expands interactive video world models from camera navigation to object‑level trajectory actions. Given a user click and a sketched path, WorldCraft generates future frames in which the selected object follows the prescribed trajectory while the camera continues to navigate the scene. WorldCraft achieves this through a trajectory‑centric control pipeline: First, Normalized World Trajectory (NWT) represents user‑drawn motion in a camera‑invariant world coordinate system and dynamically re‑projects it under the current camera pose, separating object motion from camera‑induced screen‑space displacement; Spatial‑Pathway LoRA (SP‑LoRA) then injects this world‑space signal through the model's spatial‑control pathway, adding object manipulation capability while preserving the pretrained camera controller; finally, Trajectory‑Anchored State Persistence (TASP) treats the world trajectory as a persistent spatial state and refreshes autoregressive memory after trajectory‑conditioned generation, allowing moved objects to reappear at their updated positions after leaving the camera view. Experiments show that WorldCraft enables accurate object control, preserves the video‑based world model's camera fidelity under camera‑only evaluation, and maintains object state across long autoregressive rollouts with off‑camera excursions.
Authors:Ruoyu Wang, Yong Liu, Sheng Tao, Yuhang Lin, Yukai Ma
Abstract:
Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally constructs dense spatial representations on the fly. However, recent Gaussian‑centric methods struggle with structural boundary fidelity and rely heavily on predefined scene‑size priors, fundamentally limiting their operational efficiency. In this work, we present VEOcc, a voxel‑centric framework formulated as a recursive perception‑and‑assimilation paradigm. By eliminating the need for initial scale estimation, VEOcc enables highly streamlined, open‑ended map expansion. Furthermore, to robustly aggregate noisy temporal observations within the discrete voxel space, we propose a Spatio‑Temporal‑Aware Online Update Strategy. It integrates Cross‑Temporal Logit Aggregation (TLA) for temporal consistency, Reliability‑Aware Confidence Modulation (RCM) for spatial uncertainty calibration, and Confidence‑Driven Incremental State Update (CSU) for robust global state assimilation. % Extensive experiments on Occ‑ScanNet and EmbodiedOcc‑ScanNet demonstrate that VEOcc establishes new state‑of‑the‑art performance in both local and embodied settings, providing an accurate and efficient solution for real‑world exploration. Extensive experiments on Occ‑ScanNet and EmbodiedOcc‑ScanNet demonstrate that VEOcc establishes new state‑of‑the‑art performance in both local and embodied settings. Notably, zero‑shot evaluations on self‑collected video sequences further confirm its robust out‑of‑distribution generalization capability in completely unseen real‑world environments. Ultimately, our framework provides an accurate and highly efficient solution for autonomous exploration. Code and supplementary visualizations are available on our project page: https://wryzju.github.io/VEOcc/.
Authors:Jun-Wei Hsieh, Meng-Yu Kao, Ghufron Wahyu Kurniawan, Kuan-Chuan Peng
Abstract:
YOLO‑series and DETR‑based detectors struggle with tiny‑object detection. YOLO‑style models benefit from efficient dense prediction, but their large‑stride backbones may suppress tiny instances in deep feature maps and make grid assignment ambiguous. DETR‑based models remove hand‑crafted post‑processing through set prediction, yet they reason over coarse token grids, where tiny objects occupy only a few weak tokens and are easily overlooked during matching. To address these limitations, we propose TinyFormer, a unified YOLO‑‑DETR hybrid real‑time detector that combines ViT representations, NMS‑free set prediction, and a YOLO‑style pyramid neck for accurate small‑object detection. TinyFormer introduces a Parallel Bi‑fusion Module (PBM), which builds high‑resolution shortcuts from shallow stages to the feature pyramid, preserving fine spatial details during multi‑scale fusion. We further design a Spatial Semantic Adapter (SSA) to compensate for the spatial loss caused by coarse tokenization. SSA extracts high‑resolution cues from early stages and injects them into transformer token embeddings, improving tiny‑object localization without sacrificing the global modeling ability of DETR. Experiments on MS COCO show that TinyFormer consistently outperforms recent YOLO‑series detectors and the strong DEIMv2 baseline. TinyFormer‑X achieves 58.4% AP even without PBM, while adding PBM improves the overall AP to 58.5% and brings a 1.6% AP gain on small objects. With Objects365 pre‑training, TinyFormer‑X‑PBM reaches 60.2% AP, surpassing RF‑DETR and other Objects365‑pretrained detectors with fewer parameters and lower computation. These results demonstrate that TinyFormer bridges dense YOLO‑style feature fusion and DETR‑style set prediction, providing a strong accuracy‑efficiency trade‑off for real‑time tiny‑object detection. Code is available at https://github.com/mmpmmpmmpjosh/TinyFormer.
Authors:Bangrui Xu, Ziyang Miao, Xuanhe Zhou, Yiming Lin, Zirui Tang, Xiaomeng Zhao, Fan Wu, Cheng Tan, Fan Wu, Bin Wang, Conghui He
Abstract:
VLM‑based OCR models have become the de facto choice for document parsing, as they can accurately extract page‑level elements (e.g., paragraphs within individual pages) together with their bounding boxes and textual content. However, downstream applications such as RAG require coherent document‑level information, whereas these models often break cross‑page continuity and fail to recover disrupted structures, such as paragraphs and tables truncated by page boundaries. Such relationships are not confined to a single page; instead, they require joint analysis of titles, paragraphs, tables, and images spanning multiple pages. A natural solution is therefore to reuse existing OCR outputs and reconstruct document‑level logical structures through post‑processing.
To this end, we propose MinerU‑Popo, a lightweight and universal framework for POst‑Processing OCR outputs, which converts page‑level results from diverse parsers into coherent document‑level structures. MinerU‑Popo decomposes the problem into four focused subtasks: text truncation recovery, table truncation recovery, title hierarchy reconstruction, and image‑text association. To address these effectively, we build a task‑oriented data engine with task‑specific input filtering, and use the generated data (30K) to fine‑tune a lightweight post‑processing model (Qwen3‑VL‑4B). To support long documents, we introduce dynamic chunking with overlap‑based synchronization, which aligns chunk‑level outputs from the fine‑tuned model and preserves global consistency. Finally, we assemble the aligned outputs into a tree‑structured document representation, further enriched with node chunking and summaries for downstream retrieval and analysis. Empirical results show MinerU‑Popo improves title‑hierarchy TEDS by at least 20% across all five tested OCR models, improves RAG accuracy and reduces per‑query latency.
Authors:Ibrahim Delibasoglu
Abstract:
The rapid evolution of generative models has enabled the creation of hyper‑realistic facial deepfakes, exposing a critical vulnerability in modern digital forensics: the inability of detectors to generalize to unseen manipulation techniques. Traditional networks suffer from representation collapse, overfitting to localized artifact fingerprints of specific training generators. This work investigates whether modern Vision Foundation Models can serve as generalizable, out‑of‑the‑box feature extractors capable of tracking forensic anomalies across entirely unseen generative manifolds. We conduct a systematic cross‑domain evaluation comparing three foundational learning paradigms: fully supervised macro‑semantic features (RoPE‑ViT), pure self‑supervised geometric features (DINOv3), and multi‑teacher agglomerative representations (NVIDIA C‑RADIOv4‑H). By deploying frozen backbones subjected to downstream linear probing, we map the performance limitations of these architectures on the challenging DF40 benchmark. Our empirical findings expose the intrinsic trade‑offs between pre‑training paradigms and parameter scale, proving that while foundation models retain high discriminative capabilities for entire face synthesis, localized face editing techniques expose fundamental boundaries in linear probe evaluation structures. Source code and model weights are available in http://github.com/mribrahim/deepfake
Authors:Jianrui Zhang, Hyun Jung Lee, Sukanta Ganguly, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee
Abstract:
Multimodal retrieval relies heavily on single‑vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discard fine‑grained, local evidence critical for dense retrieval tasks. Multi‑vector approaches were introduced as a solution, but they strictly require training and many ignore the necessity of a globally summarizing representation. To address this, we introduce SMART, a framework that unlocks the latent multi‑vector capabilities of standard single‑vector models. We first demonstrate that standard contrastive training on the pooled embedding implicitly shapes the retrieval geometry of preceding hidden states via gradient flow. By applying direct late‑interaction over these frozen hidden states during inference, SMART acts as a plug‑and‑play upgrade that consistently improves performance across diverse modalities, improving even the state‑of‑the‑art models further on MMEB‑V2. We also reveal SMART's superior performance, as simple lightweight post‑training not only saves time and compute, but also brings forth further improvement on Visual Document retrieval, allowing a single‑vector model to outperform SoTA multi‑vector counterparts. Ultimately, SMART offers both a highly efficient inference enhancement and a powerful finetuning technique for multimodal retrieval. We open source our code and weights at https://github.com/HanSolo9682/SMART.
Authors:Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos
Abstract:
Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challenging due to the embodiment gap between human and robot in both visual appearance and kinematics. We present HumanEgo, a framework that bridges the embodiment gap by lifting each human demonstration to an entity‑level representation of hand‑object interaction, and training a flow matching policy with dense auxiliary objectives that amplify supervision from every trajectory. HumanEgo is robot‑data‑free, hardware‑agnostic, data‑efficient, and zero‑shot human‑to‑robot transferable. With only 30 minutes of human videos per task, HumanEgo achieves 92.5% average success across four real‑world tasks (75% with just 15 minutes), outperforms matched‑time robot teleoperation by 41%, and robustly transfers zero‑shot across novel robots, cameras, and environments. We release HumanEgo as an easy‑to‑use, open‑source framework for learning robot policies directly from human data: https://github.com/TX‑Leo/HumanEgo
Authors:Yuanye Liu, Siyuan Zhou, Ke Zhang, Lei Li, Wei Chen, Xiahai Zhuang
Abstract:
Pre‑trained Vision Transformers (ViTs) are increasingly deployed for medical image classification. However, correcting their inevitable failure cases in dynamic clinical scenarios poses a critical challenge. Conventional fine‑tuning approaches inherently suffer from catastrophic forgetting, severely degrading previously acquired diagnostic capabilities. Such instability fundamentally compromises clinical safety. Addressing this vulnerability requires an active, controllable, and reliable intervention mechanism that is both theoretically grounded and inherently interpretable. To this end, we propose X‑Edit (eXact, eXplicit, and eXplainable Editing), an efficient null‑space model editing framework. X‑Edit transitions the editing process from iterative gradient‑based optimization to a theoretically grounded, closed‑form solution. Specifically, we first explicitly localize the influential layers via causal tracing governing the erroneous prediction. Subsequently, we construct an orthogonal null‑space projection matrix from a curated anchor set. By geometrically constraining the exact parameter update strictly within this null space, we provide mathematical guarantees that the intervention rectifies targeted errors without perturbing established diagnostic representations. Extensive evaluations on six medical imaging benchmarks demonstrate that X‑Edit comprehensively suppresses catastrophic forgetting while achieving superior edit success rates. Our code is available at https://github.com/HenryLau7/X‑Edit.
Authors:Hui Lin, Jiayi Li, Jing Wang, Shenghui Rong
Abstract:
Sonar imaging is the primary modality for underwater target detection, yet small targets remain difficult to detect due to insufficient pixel coverage, low acoustic contrast, and scale ambiguity across imaging ranges. CNN‑based detectors extract local features efficiently but cannot suppress noise‑induced false alarms without global acoustic context. Transformer‑based methods capture long‑range dependencies at quadratic computational cost. Existing Mamba‑based vision models offer efficient linear‑cost scanning but lack multi‑scale semantic alignment across pyramid levels, multi‑receptive‑field fusion, and small‑target‑aware training supervision needed for reliable sonar detection.
This letter proposes Mamba Dilated‑Scale Fusion (MambaDSF), a hybrid framework addressing these limitations through three contributions: a Mamba Enhanced Feature Pyramid (MambaEFP) backbone that jointly captures local echo cues and global acoustic context at linear complexity; a Dilate Fusion Mamba (DFMamba) encoder that enforces multi‑scale feature alignment across pyramid levels; and Scale‑Adaptive Weighted IoU (SA‑WIoU) and Cross‑Scale Coherence (CSC) losses that stabilize small‑target training. MambaDSF achieves 91.5% mAP50 on the UATD forward‑looking sonar benchmark with 28.7 million parameters, surpassing all compared detectors. On a small‑target subset the gain reached +2.2 percentage points, and cross‑domain evaluation on FLS and MD‑FLS confirms the generalization of the proposed architecture. The codes are publicly available at https://github.com/IDontKnowAAA/MambaDSF.
Authors:Zijie Cao, Weijie Tu, Yao Xiao, Weijian Deng, Liang Lin, Pengxu Wei
Abstract:
Detecting AI‑generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods are trained on large datasets, their performance still degrades when generation settings change, indicating that data scale alone is insufficient and that limited coverage of generative variations during training is a key factor. Studies on generative model editing show that small changes in internal representations can produce diverse and meaningful image variations, many of which are not explored under standard sampling. Leveraging this insight, we propose PROBE (Probing Robustness via Boundary Exploration), a framework that improves detector generalization by actively exploring challenging regions of the generative process. Instead of treating the generator as a fixed data source, PROBE uses the detector as a critic to steer the generator through manifold‑level modifications, producing realistic samples that are difficult to classify. These samples expose failure cases that are uncommon under standard data sampling strategies and are used to refine the detector. Experimental results across multiple benchmarks indicate that PROBE enhances generalization to unseen generators, resulting in more generalizable AIGI detection performance. Code and models are available at https://github.com/Amamiya‑C/PROBE‑AIGI‑Detection
Authors:Tyler Rust, Dara McNally, Kyle O'Donnell, Colin Kelly, Chandra Kambhamettu
Abstract:
Building upon the SAM2 vision foundation model for downstream segmentation, this study introduces Boundary Enhanced Depth (BED)‑SAM2. The SAM2 Hiera encoder architecture is modified to directly encode monocular depth information from RGB images, thereby providing geometric cues that enhance object boundary delineation and facilitate the extraction of camouflaged object shapes. BED‑SAM2 demonstrates competitive state‑of‑the‑art performance across multiple salient and camouflaged object detection tasks with as few as five training epochs.
Authors:Mingyu Liang, Dingkun Xu, Jingwei Xu
Abstract:
Diffusion Transformers require repeated denoiser evaluations during iterative sampling, making inference computationally expensive. Cache‑based acceleration reduces this cost by reusing intermediate representations across denoising steps, but can introduce representation deviations and degrade generation quality. In this paper, we analyze these deviations and show that effective calibration should consider both the direct mismatch caused by reuse and the subsequent trajectory shift induced by earlier corrections. To address this challenge, we propose Trajectory‑Consistent Calibration (TCC), a training‑free method that calibrates cached representations toward their full‑computation counterparts. Specifically, rather than estimating all calibration priors from a single uncorrected cache trajectory, TCC uses an offline iterative procedure so that each prior accounts for the trajectory shift induced by preceding calibrations. Experiments on PixArt‑alpha and DiT‑XL/2 show that TCC consistently improves FID across representative cache‑based acceleration methods while preserving their underlying reuse policies. Notably, in a representative PixArt‑alpha cache‑acceleration setting based on FORA, TCC reduces FID from 29.83 to 27.35, slightly surpassing the full‑computation baseline.
Authors:Ligong Bi, Tao Huang, Jianyuan Guo, Chang Xu
Abstract:
Visual Autoregressive (VAR) models have emerged as a powerful paradigm for image synthesis by performing hierarchical next‑scale prediction. However, VAR models are inherently prone to cascading error propagation, where subtle coarse‑scale mispredictions are amplified across the hierarchy, ultimately distorting the final synthesis. To mitigate this, we propose AID‑VAR, a plug‑and‑play framework that enhances pre‑trained VARs through Adversarially Injected Diagnosis. Instead of a standard passive generation, AID‑VAR introduces a proactive error‑correction mechanism inspired by the adversarial feedback in GANs. We deploy a discriminator to diagnose fidelity gaps at each scale transition, coupled with a lightweight guidance injector. This module operates as a non‑invasive adapter that refines the feature manifold of a frozen VAR backbone, effectively steering the generation toward the distribution of real images without destabilizing the pre‑trained latent space. Furthermore, to rigorously evaluate this cross‑scale progression, we introduce the Inter‑Scale Consistency Score (ISCS), a novel metric that quantifies the fidelity and structural alignment between consecutive resolution scales. Experimental results across various backbones demonstrate that AID‑VAR delivers sharper textural details and fewer structural distortions with negligible overhead. For instance, AID‑VAR‑d20 achieves a 16% improvement in FID with only a 3% increase in parameters. These results establish AID‑VAR as a highly efficient and scalable pathway for upgrading large‑scale VAR generators, enhancing global coherence and local detail without altering training data, base architectures, or sampling schedules. Code is available at https://github.com/bijiw515/AID‑VAR.
Authors:Jian Lang, Rongpei Hong, Ting Zhong, Fan Zhou
Abstract:
Deploying multimodal systems in real‑world environments often entails handling modality‑missing scenarios, where one or more modalities are unavailable. While recent studies address this challenge for the general Multimodal Transformer (MT) architecture via prompt tuning, we identify a fundamental limitation in these methods: the Implicit Modality‑Reduction bottleneck. By conditioning prompts solely on the observed modalities, they inadvertently restrict the reasoning scope of MTs to the modality‑reduced subspace, cutting off access to the latent information sources of the missing modalities. To overcome this limitation, we propose AOEPT, which pioneers a novel modal‑contextualized prompting fashion. Specifically, we introduce lightweight Modal‑Contextualized Prompts (MCPs) that distill global modality‑wise priors from training data, serving as latent repositories of the information sources for missing modalities. Conditioned on the remaining modalities, these MCPs are instantiated into instance‑aware prompts that selectively augment missing‑modality information for each sample, thereby restoring the reasoning scope of MTs beyond the observed‑modality‑only subspace. Experiments across various multimodal benchmarks and backbones confirm the strong performance of AOEPT, with minimal computational overhead.
Authors:Jie-En Yao, Hong-En Chen, C. -C. Jay Kuo
Abstract:
Deep neural networks trained with backpropagation have achieved outstanding performance in vision tasks but remain biologically implausible, computationally demanding, and difficult to interpret. The Forward‑Forward (FF) algorithm offers a promising alternative by training each layer independently through local goodness objectives. However, its purely local optimization lacks hierarchical coordination across layers, and the decoupling of goodness from features leaves the representations unconstrained and semantically ambiguous. We propose a Hierarchical and Contrastive Learning FF framework (HCL‑FF) to address these limitations. HCL‑FF introduces (1) a coarse‑to‑fine hierarchical learning strategy that guides representations from low‑level cues to high‑level semantics, and (2) a supervised contrastive objective that enforces class‑discriminative alignment after goodness decoupling. Experiments on CIFAR‑10, CIFAR‑100, and Tiny‑ImageNet demonstrate that HCL‑FF achieves new state‑of‑the‑art performance among FF‑based methods, with notable accuracy gains of +5.46%, +17.00%, and +12.51%, respectively.
Authors:Zihao Zhu, Kuan-Ru Huang, Zhaoming Xu, Renjie Li, Bo Wu, Ruizheng Bai, Mingyang Wu, Sayak Paul, Zhengzhong Tu
Abstract:
High‑resolution datasets are essential for advancing super‑resolution (SR) and text‑to‑image (T2I) diffusion research. However, current publicly available datasets lack both the native 4K resolution and the extensive scale necessary for training state‑of‑the‑art models. To address this gap, we introduce a 4K Large Scale Dataset and Benchmark (4KLSDB), a large‑scale, diverse dataset consisting of 129,484 carefully curated 4K resolution images spanning multiple categories such as nature, urban scenes, people, food, artwork, and CGI, alongside distinct validation and test sets containing 2,000 and 1,984 images respectively. Images were sourced from established open datasets including Photo Concept Bucket, Laion2B, and PD12M. 4KLSDB underwent rigorous multi‑stage automated filtering and annotation pipelines involving both human annotators and Large Multimodal Models (LMMs) to ensure high aesthetic quality and dataset consistency. We demonstrate 4KLSDB's effectiveness by training representative super‑resolution and diffusion models, observing significant improvements in performance on native 4K benchmarks. Comprehensive experiments illustrate a positive correlation between training on true 4K resolution data and improved fidelity in image restoration task, especially on 4K resolution. We provide the research community a valuable resource to drive progress toward genuinely high‑fidelity image synthesis and restoration by providing 4KLSDB. Our project page is available at: https://4klsdb.github.io/.
Authors:Ismail Lamaakal
Abstract:
Neural network weights are increasingly a bottleneck for deployment, yet most compression pipelines treat layers independently and overlook cross‑layer redundancy induced by function‑preserving symmetries. We propose Motion‑Compensated Weight Compression (MCWC), a weight‑only codec that aligns permutation‑symmetric blocks (e.g., hidden units and attention heads) to maximize cross‑layer correspondence, turning depth into a predictable sequence. In the aligned coordinate system, MCWC uses a lightweight layer‑sequential predictor with periodic keyframes and encodes only quantized prediction residuals using a learned entropy model trained under a rate distortion objective. A simple decoder reconstructs deployable weights by entropy decoding, dequantization, predictor‑driven reconstruction, and inverse alignment, enabling fast weight materialization for inference. Across Transformer language modeling and vision classification, MCWC improves the rate accuracy Pareto frontier over strong quantization and learned weight‑codec baselines, while maintaining competitive decode time. Ablations confirm that alignment, prediction, entropy modeling, and keyframe scheduling are each necessary for the full gains. Our code is available via https://github.com/Ism‑ail11/MCWC.
Authors:Ruyi Chen, Lu Zhou, Xiaogang Xu, Chiyu Zhang, Jiafei Wu, Liming Fang
Abstract:
Text‑to‑Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation methods typically address only single‑dimensional biases, lacking perspectives to uncover model biases at social‑related deeper semantic levels. We introduce HoloFair, a comprehensive benchmark framework for multidimensional demographic bias analysis. Built upon our large‑scale fairness‑oriented dataset and the SpaFreq (Spatial‑Frequency) attribute classifier, this framework proposes the Multi‑attribute, Group‑wise Bias Index (MGBI) metric, designed to assess both intrinsic diversity and conditional biases. Beyond evaluation, we further introduce Fair‑GRPO, a reinforcement‑learning‑based debiasing method that alters the distribution of generative models through a designed multi‑objective reward function. E.g., experiments on the SD3.5‑Medium model demonstrate that Fair‑GRPO significantly improves multidimensional fairness while maintaining high image quality. We also analyze potential reward hacking phenomena and provide corresponding mitigation strategies. Code and dataset are available at https://github.com/1059684669/HoloFair
Authors:Sol Park, Soobin Um
Abstract:
Minority sampling aims to generate low‑density instances on a data manifold and is of central importance in applications such as medical diagnosis, anomaly detection, and creative AI. Existing approaches, however, define minority samples relative to generative priors learned from training data, confining rarity to model‑specific notions that may poorly reflect real‑world semantics. In this work, we propose a world‑centric perspective on minority sampling, which defines rarity with respect to real‑world priors rather than generator‑induced densities. To this end, we introduce JEPA guidance, a diffusion sampling framework guided by a Joint‑Embedding Predictive Architecture (JEPA) ‑‑ a class of world models that encode broad, semantically rich representations. JEPA guidance steers diffusion trajectories toward low‑density regions under the implicit density induced by the JEPA, thereby aligning generated minorities with real‑world semantic rarity. To make JEPA guidance computationally practical, we develop principled approximation strategies accompanied by theoretical error bounds, significantly reducing the overhead of guidance computation. Extensive experiments across unconditional, class‑conditional, and text‑to‑image generation demonstrate that JEPA guidance consistently improves the fidelity and semantic validity of minority samples, outperforming generator‑centric baselines in capturing real‑world notions of rarity. Code is available at https://github.com/soobin‑um/jepa‑guidance.
Authors:Toufiq Musah, Salvatore Calcagno, Federica Proietto Salanitri, Xiaomeng Li, Maruf Adewole, Marawan Elbatel
Abstract:
Ultra‑low‑field (ULF) MRI offers portable and accessible neuroimaging but suffers from reduced signal‑to‑noise ratio and limited spatial resolution compared to high‑field (HF) systems. Acquiring paired ULF‑HF data for supervised enhancement is often difficult, particularly in resource‑limited settings. We introduce ULF‑Synth, a framework that combines: (i) acquisition‑based synthesis of realistic ULF images from HF volumes to create large‑scale paired training data, (ii) a spatial‑frequency domain objective that prioritizes recovery of high‑frequency anatomical detail. This formulation is architecture‑agnostic, consistently improving structural similarity and perceptual fidelity across encoder‑decoder, adversarial, and diffusion‑based translation models. When trained exclusively on synthetic data, the resulting models generalize effectively to real 64mT ULF acquisitions, improving downstream multiclass brain segmentation and achieving higher radiologist preference and diagnostic acceptability in a blinded reader study. These findings demonstrate that synthetic paired supervision provides a practical and scalable pathway for enhancing ULF MRI without requiring real paired acquisitions. Code, Models and Dataset: https://github.com/toufiqmusah/ULF‑Synth
Authors:XiaoWan Hu, Jing Yang, HeNan Liu, HuaQiu Li, Mai Xu
Abstract:
Zero‑shot image restoration provides a flexible way to handle diverse degradations without task‑specific training. However, existing methods typically rely on stacked layers or pre‑trained features to enhance degradation expression, while overlooking physically consistent priors. The insufficient degradation prompts impose the heavy training burden and high sampling costs during zero‑shot diffusion. Moreover, the fixed inference trajectory often collapses to suboptimal solutions under complex corruptions. We observe that heterogeneous degradations can be reparameterized into a minimal set of physically coherent parameters for compact representation. Based on this insight, we first propose a unified physical zero‑shot image restoration (UP‑ZeroIR) framework that explicitly models heterogeneous degradations into a homogeneous all‑in‑one distribution. The distribution can be optimized directly in the latent space, enabling principled solution exploration and effective prompt adaptation. Besides, we introduce a dynamic quality‑refinement strategy that adaptively adjusts the diffusion trajectory for robust globally optimal convergence. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance across both single and mixed degradations. Our code is available at https://github.com/yangjinglyy/UP‑ZeroIR
Authors:Sattam Altuuaim, Lama Ayash, Muhammad Mubashar, Naeemullah Khan
Abstract:
Despite the central role of optimization in deep learning, most optimizers rely on update structures whose functional form is fixed before training begins. This static design can limit their ability to respond to changing gradient behavior across the loss landscape, where training may shift between stable, noisy, and inconsistent regimes.
This study proposes PILOT (Policy‑Informed Learned OpTimizer), an online optimizer that adapts its update behavior during training. Rather than using a fixed balance between momentum, normalization, and sign‑based updates, PILOT uses gradient‑direction agreement as a signal of local training stability. Conditioning the update rule on this agreement signal allows the optimizer to adjust its behavior when gradients become stable, noisy, or inconsistent.
Experiments on FashionMNIST and CIFAR‑10 show that PILOT consistently achieves the highest accuracy among the evaluated optimizers across convolutional settings. On the CNN architecture, PILOT reaches 94.13% on FashionMNIST and 81.94% on CIFAR‑10. On ResNet‑18, it further improves performance, reaching 95.71% on FashionMNIST and 93.42% on CIFAR‑10. These results suggest that learning how to adapt the update structure during training can improve performance across both compact and deeper convolutional models while preserving a simple first‑order optimization framework.
The implementation of PILOT is publicly available at https://github.com/SattamAltwaim/PILOT.git
Authors:Alif Tri Handoyo, Vincent C. S. Lee, Rizka Widyarini Purwanto, Alex M. Lechner, Deanna Kemp, Muhamad Risqi U. Saputra
Abstract:
Automatically mapping and segmenting global mining footprints using remote sensing and deep learning is critical for monitoring the socio‑environmental risks and impacts of mining, yet its progress is hindered by the scarcity of fine‑grained annotated data. Although large‑scale datasets with coarse boundaries are widely available, leveraging them to improve fine‑grained segmentation is challenging due to significant domain shift. To address this, we propose MineC2FNet, a coarse‑to‑fine domain incremental learning framework that exploits abundant coarse data to enhance fine‑grained mining footprint segmentation. MineC2FNet adopts a teacher‑student architecture with attentive distillation at both the feature and prediction levels, selectively transferring generalized knowledge from the coarse domain while enabling boundary refinement using limited fine‑grained data (fine domain). We further introduce an expertly validated dataset of 219 images with precise boundary annotations across diverse geographies and commodities. Extensive experiments against state‑of‑the‑art approaches, including domain adaptation and domain incremental learning methods, demonstrate that MineC2FNet achieves superior performance while effectively handling domain shift. The dataset and code are publicly available at https://github.com/risqiutama/MineC2FNet.
Authors:Bill Psomas, Dionysis Christopoulos, Thanasis Petropoulos, Nikos Efthymiadis, Ioannis Kakogeorgiou, Ondřej Chum, Yannis Avrithis, Giorgos Tolias, Konstantinos Karantzalos
Abstract:
Remote sensing composed image retrieval (RSCIR) enables search in large satellite image archives using composed queries that combine a reference image with a textual modifier. Although RSCIR offers a flexible interface for expressing targeted retrieval intent, the transferability of modern composition methods to Earth observation (EO) imagery and their relevance to operational EO workflows remain underexplored. We address this gap through a unified benchmark and an application‑oriented study. First, we systematically adapt and evaluate representative composed image retrieval methods with six vision‑language backbones on PatternCom under a standardized protocol, analyzing their behavior across backbones, composition strategies, and query types. Second, we introduce xView2‑CIR, a change‑centric dataset for disaster and damage monitoring, where retrieval is conditioned on scene identity and a target post‑event state. Our results show that training‑free composition methods provide strong and scalable baselines for EO retrieval, while change‑centric retrieval presents different challenges from attribute‑based retrieval, particularly due to the need to preserve scene identity. Overall, this study establishes a practical benchmark for RSCIR and positions composed retrieval as a complementary tool for remote sensing image retrieval, archive exploration, and change analysis. The dataset and code are available at https://github.com/billpsomas/rscir.
Authors:Ruoyu Wang, Jingke Wang, Yukai Ma, Yuehao Huang, Shuangming Lei, Guanglin Xu, Aixue Ye, Yong Liu
Abstract:
Recently, world models have made significant progress in enhancing end‑to‑end driving systems through both future situation forecasting and improved scene understanding. However, existing driving world models are typically built upon dense scene representations, causing high computational costs and redundant information. In this paper, we present SparseWorld, a lightweight world model that focuses on predicting only the critical layout of the scene, enabling efficient future forecasting for end‑to‑end driving systems. SparseWorld first performs autoregressive rollout to forecast future map elements and surrounding agents, enabling the model to learn how driving scenarios evolve over time. It then leverages these predicted futures to refine downstream motion prediction and trajectory planning. Specifically, we propose a Sparse Dreamer that anticipates future instances in the latent space through joint temporal and spatial attention. By interacting with predicted future instances, the motion planner captures more accurate motion patterns and generates more informed and safety‑aware trajectories. Extensive experiments demonstrate that SparseWorld significantly reduces collision risk and achieves state‑of‑the‑art performance on the open‑loop planning metrics of the nuScenes dataset with a collision rate of 0.05%. Moreover, it substantially outperforms the baseline method in closed‑loop planning metrics on the Bench2Drive benchmark. Supplementary material is available at the project page: https://wryzju.github.io/SparseWorld/.
Authors:Gunjan Shrivastava, Saad Nadeem
Abstract:
Accurate cell segmentation in pathology images typically requires dense pixel‑wise annotations, which are costly and time‑consuming to obtain. This challenge is especially important for emerging biological imaging modalities and multiplexed datasets with variable channel configurations, where expert‑labeled data are scarce. In this work, we introduce ImPartial, a deep learning framework designed to achieve state‑of‑the‑art segmentation performance in low‑annotation regimes using sparse scribbles and limited supervision. ImPartial augments the segmentation objective via self‑supervised multi‑channel quantized imputation. This approach leverages the observation that perfect pixel‑wise reconstruction or denoising of the image is not needed for accurate segmentation, and thus, introduces a self‑supervised classification objective that better aligns with the overall segmentation goal. We demonstrate that ImPartial achieves performance at par with fully supervised models while requiring substantially fewer annotations. Extensive experiments on benchmark multiplexed cellular imaging and single‑plex clinical brightfield immunohistochemistry datasets show consistent improvements over strong baselines with only partial annotations. All benchmark datasets and code are available via our Github: https://github.com/nadeemlab/ImPartial.
Authors:Kevin Richard, Alphin Varghese, Colin Pham, David Oh, Srijan Das
Abstract:
Single‑vehicle Vision‑Language Models (VLMs) are fundamentally constrained by sensor occlusions. While Vehicle‑to‑Everything (V2X) systems mitigate this, current benchmarks lack the cooperative reasoning required for resolving ambiguities in complex environments. We introduce D2‑V2X, a spatially‑aware Question‑Rationale‑Answer (QRA) benchmark featuring 8,500 triplets derived from multimodal vehicle and infrastructure sensors. We additionally establish a baseline that aligns 3D LiDAR features with the VLM's latent space. By enforcing natural language Chain‑of‑Thought rationales prior to structured JSON outputs, our model is forced to explicitly articulate spatial relations. Our experiments demonstrate that grounding VLMs in cooperative LiDAR achieves 24.4% recall in identifying occluded hazards compared to near‑zero in zero‑shot models and reduces spatial estimation error for visible objects by 77% compared to the zero‑shot baseline. While the model achieves a functional decision‑making F1‑score of 53.5, we identify 3D‑to‑2D projection as a fundamental bottleneck in current VLM architectures, establishing a new baseline for future innovation. Data, code, and trained models available at https://github.com/KevinRichard1/D2‑V2X
Authors:Ilia Indyk, Ignat Penshin, Ivan Sosin, Maxim Monastyrny, Aleksei Valenkov, Ilya Makarov
Abstract:
Fisheye cameras are increasingly adopted in robotics for near‑field manipulation, navigation, and immersive perception, yet indoor depth benchmarks with accurate ground truth are still missing. To address this, we introduce WideDepth ‑ the first indoor dataset for fisheye depth estimation, featuring 101 scenes containing 5K high‑resolution stereo pairs labeled with millimeter‑level ground truth depth and disparity. Our dataset also includes paired pinhole and fisheye samples across varying fields of view and baselines in both horizontal and vertical stereo setups. We further propose a method to adapt pinhole‑trained stereo models to fisheye images and introduce a novel stereo fisheye image generation pipeline based on high‑resolution LiDAR scans. Leveraging these methods, we thoroughly evaluate state‑of‑the‑art monocular depth, stereo matching, and depth completion models on our benchmark. Additionally, we provide 18K LiDAR‑derived sparse depth training samples, achieving up to a 62% performance boost on fisheye data when fine‑tuning pinhole‑based stereo models. In summary, the high precision and versatility of our benchmark set a strong foundation for advancing research in fisheye depth estimation and robotics perception. Project page: https://ilyaind.github.io/WideDepth
Authors:Farhat Shaikh, Ayan Banerjee, Sandeep Gupta
Abstract:
We introduce EMMA, a physics‑informed multimodal framework that recovers all identifiable dynamical parameters of a system directly from raw video, audio, and image‑based time‑series observations. Unlike prior video‑only approaches that struggle with occluded states, hidden actuation inputs, or assumptions about known initial conditions and coordinate frames, EMMA performs joint inference of explicit parameters, implicit dynamical components, and calibration invariants within a unified continuous‑time model. EMMA leverages a Liquid Time‑Constant (LTC) network to learn latent dynamics from heterogeneous modalities while a physics‑constrained loss enforces consistency with the governing differential equations. A unified feature pipeline enables consistent alignment across video trajectories, acoustic signatures, and chart‑derived measurements, allowing EMMA to estimate parameters under forced, implicit, and multivariate dynamics without requiring segmentation masks, differentiable rendering, or specialized sensors. Across 100+ scenarios including five standard dynamical benchmarks (75 Delfys videos), real‑world rover and quadrotor systems with hidden inputs, and simulation‑chart case studies spanning biological and chaotic systems, EMMA delivers robust multi‑parameter recovery and significantly outperforms existing single‑modality and equation‑discovery baselines. Our results establish EMMA as a general, scalable solution for physics‑consistent model extraction from opportunistic multimodal data. Code and data are available at: https://github.com/ImpactLabASU/EMMA‑CVPR2026
Authors:Youwei Pang, Changsheng Gao, Dong Liu, Huchuan Lu, Weisi Lin
Abstract:
Large models have delivered remarkable performance across a wide range of perception and generation tasks, yet practical deployment is increasingly constrained by computational and memory budgets, as well as privacy requirements. Split execution alleviates these constraints by partitioning computation across devices, but it inevitably introduces intensive transmission and storage of intermediate features. Unlike conventional feature coding for CNNs that typically targets homogeneous spatial activation maps, modern large models generate heterogeneous features with varying statistical distributions and compression tolerances, e.g., multi‑level/multi‑modal representations and autoregressive context caches. These characteristics necessitate treating large model feature coding (LaMoFC) as a fundamental system component and call for a systematic evaluation framework. In this paper, we present a comprehensive benchmark and evaluation framework for LaMoFC. We first build the feature dataset LaMoFCBench, covering diverse task requirements across 4 categories and 16 scenarios while integrating widelyadopted architectures and various split‑computing settings. We then specify representative split points according to practical application scenarios to extract intermediate features, establishing a unified pipeline for fair and reproducible comparisons. Finally, we benchmark mainstream universal feature codecs, exposing the profound misalignment between existing coding paradigms and the heterogeneous nature of large model features. These findings reveal that LaMoFC demands a fundamental departure from existing paradigms, and LaMoFCBench provides the shared empirical foundation to drive this transition. The data and code will be available at https://github.com/lartpang/LaMoFCBench.
Authors:Zhengqi Sun, Yiwen Sun, Boxuan Liu, Tailai Chen, Tianxu Guo, Jiabin Liu
Abstract:
Large language models (LLMs) are promising for autonomous driving, but semantics‑only decision policies can yield physically unsafe behavior in dynamic traffic. Existing methods either perform online language reasoning without explicit dynamics verification or use world models mainly in offline pipelines, leaving a gap between semantic intent and physical feasibility at decision time. We propose Reason‑‑Imagine‑‑Act (RIA), a closed‑loop framework that couples an LLM reasoner with an action‑conditioned world model for online safety verification. At each step, the LLM proposes an action template and candidate sub‑actions, the world model performs short‑horizon rollouts, and a safety scorer selects the safest executable action with feedback to the next reasoning step. Under a unified CARLA point‑goal protocol (1000 episodes), RIA achieves 80.05% route completion, 51.10% arrival rate, and 0.20% collision rate. Under the same closed‑loop interface, RIA consistently outperforms training‑free baselines, including CARLA TM and MADA, on core closed‑loop metrics. For reproducibility, code is available at https://github.com/pku‑smart‑city/source_code/tree/main/RIA.
Authors:Chi Kit Wong, Yan Liu, Haowen Yan
Abstract:
We present a brain‑to‑image system that decodes visual stimuli from EEG signals recorded during natural image viewing. Our system addresses two tasks: (1) EEG‑to‑image retrieval, which ranks the correct stimulus image among 200 candidates given an EEG segment, and (2) EEG‑to‑image reconstruction, which generates an image consistent with the perceived stimulus. For retrieval, we implement a multi‑level blurring approach improved with biologically inspired EVNet features and trained with the InfoNCE loss. Evaluated over 10 random seeds for a single subject, the retrieval model achieves a mean final‑epoch Top‑1 accuracy of 86.30% and Top‑5 accuracy of 98.55%. For reconstruction, we implement CognitionCapturerPro, which aligns EEG representations to multi‑modal CLIP embeddings, including image, text, depth, and edge embeddings, and synthesizes images with SDXL‑Turbo conditioned via IP‑Adapter. Averaged over 10 seeds, the reconstruction model achieves a CLIP score of 0.903 using ViT‑H‑14, a CLIP score of 0.870 using ViT‑L/14, and an SSIM of 0.409. These results demonstrate the feasibility of decoding rich visual representations from EEG signals using modern multi‑modal alignment and generative modeling techniques.
Authors:Siqiao Huang, Partha Kaushik, Michael Chen, Hengkai Pan, Kaiwen Geng, Omar Chehab, Fernando Moreno-Pino, Max Simchowitz
Abstract:
World models have become a central paradigm for learning predictive simulators that support generation, planning, and decision‑making. Yet, despite rapid progress in industry‑scale interactive video generation, the broader research community still lacks compact, reproducible, and easily extensible implementations for studying the design choices underlying modern world models. We introduce Nano World Models, a minimalist codebase for future video prediction centered around diffusion forcing. Nano World Models provides a unified interface for generative objectives, model scales, action‑conditioning mechanisms, latent observation spaces, datasets, evaluation protocols, and long‑horizon rollout procedures. This design enables controlled studies of world‑modeling components that are often entangled across separate implementations. Through experiments across simple control environments, game simulation, and real‑robot data, we examine how prediction parameterization, architecture scale, action injection, sampling budget, and domain complexity affect video prediction quality and autoregressive rollout behavior. By releasing code, configurations, evaluation scripts, and pretrained checkpoints, Nano World Models aims to provide a compact yet extensible experimental substrate for open, reproducible, and scientific world‑model research.
Authors:Sam Earle, Kai Arulkumaran, Andrew Dai, Akarsh Kumar, Julian Togelius, Sebastian Risi
Abstract:
We are in the midst of large‑scale industrial and academic efforts to automate the processes of scientific, technological and creative production through AI‑driven assistants. Historically, a fundamental property of these processes in their human form has been their open‑endedness: their capacity for generating a seemingly endless supply of novel and meaningful new forms. Do artificial agents have any capacity for such fruitful unguided discovery? To answer this question, we turn to Picbreeder, the canonical exemplar of human‑driven open‑ended search, in which users collaboratively generated a diverse library of images through interactive evolution of small neural networks. We replicate Picbreeder, replacing human users with frontier Vision Language Models (VLMs). We observe clear qualitative differences between the output of our system and the historical human baseline, and attempt to characterize them using metrics of phylogenetic complexity and visual and semantic salience and novelty. In an effort to identify some of the causal factors contributing these differences, we study the addition of exploratory noise to the agents' selection process, of behavioral diversity between agents, and of narrative momentum in the form of memory of past actions. We make our code available at https://github.com/smearle/picbreeder‑vlm.
Authors:Beichen Zhang, Yuhong Liu, Jinsong Li, Yuhang Zang, Jiaqi Wang, Dahua Lin
Abstract:
Multimodal Large Language Models have advanced visual reasoning, yet a purely textual chain of thought remains a bottleneck for questions that require fine‑grained focus or view transformations. The ''think with images'' paradigm narrows this gap, but existing approaches are either constrained by fixed predefined toolkits or produce noisy intermediate images from unified multimodal methods. We pursue a third option: using a dedicated image editing model and decouple it with an understanding model. However, off‑the‑shelf image editors fail as reasoning assistants with two complementary gaps: a language‑side gap, where editors trained as passive instruction‑followers cannot map an abstract question to an appropriate visual transformation, and a generation‑side gap, where edit correctness degrades as reasoning depth grows. Guided by this analysis, we introduce ETCHR (Editing To Clarify and Harness Reasoning), a question‑conditioned, reasoning‑aware image editor decoupled from the downstream understanding model and trained with a two‑stage recipe targeted at the two gaps: Reasoning Imitation via supervised fine‑tuning on edit trajectories, followed by Reasoning Enhancement with VLM‑derived rewards for edit correctness and downstream reasoning accuracy. Since the editor is decoupled, ETCHR plugs into different open‑ and closed‑source MLLMs in a training‑free manner. Across five task families (fine‑grained perception, chart understanding, logic reasoning, jigsaw restoration, and 3D understanding), ETCHR raises average Pass@1 from 55.95 to 60.77 (+4.82) with Qwen3‑VL‑8B, from 65.08 to 70.55 (+5.47) with Gemini‑3.1‑Flash‑Lite, and from 76.55 to 81.16 (+4.61) with the 1T‑parameter MoE model Kimi K2.5.
Authors:Shuhong Zheng, Michael Oechsle, Erik Sandström, Marie-Julie Rakotosaona, Federico Tombari, Igor Gilitschenski
Abstract:
Visual geometry transformers have become powerful architectures for multi‑view 3D reconstruction, enabling joint prediction of multiple 3D attributes in a feed‑forward manner. However, their computational cost grows quadratically with the input sequence length due to the global attention layers inside these models. This limits both their scalability and efficiency. In this work, we address this challenge with a simple yet general strategy: restricting the number of key/value tokens that each query interacts with during global attention. To achieve effective token selection, we introduce a two‑stage framework. First, an inter‑frame selection step operates at the frame level to identify frames that should be preserved. Second, an intra‑frame selection step further discards more redundant tokens within the selected frames. Our analysis highlights the advantage of a diversity‑based strategy for inter‑frame selection, which ensures broad coverage of the scene. For intra‑frame selection, we show that layer‑aware sparsification is necessary, with the selection process guided by the entropy of the global attention pattern. Our approach offers a superior speed‑accuracy trade‑off compared to existing solutions. Extensive experiments show that it accelerates visual geometry transformers by over 85% for scenes with 500 images while maintaining, or even improving, baseline performance, which hints that how our token selection strategy can play a crucial role in future applications of visual geometry transformers. Our project website is available at https://zsh2000.github.io/good‑token‑hunting.github.io.
Authors:Chong Cheng, Peilin Tao, Nanjie Yao, Guanzhi Ding, Xianda Chen, Yuansen Du, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Zhengqing Chen, Hao Wang
Abstract:
Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded‑memory constraints. Existing methods often suffer from drift, jitter, or collapse on long sequences. We trace these failures to a fundamental mismatch. Streaming geometry is inherently temporally heterogeneous, with evidence ranging from short‑lived correspondences to persistent global scale. However, current architectures impose uniform and pathological influence patterns. For example, sliding windows enforce hard cutoffs, while ungated recurrence and causal attention cause cache saturation and spike‑like attention sinks. To resolve this, we formalize geometric propagation as an \emphevidence influence kernel and propose HorizonStream, a long‑horizon Transformer that explicitly factorizes this kernel. For the long‑range temporal factor, Geometric Linear Attention learns channel‑wise decay rates to enable bounded, multi‑timescale propagation of geometric evidence. For the short‑range spatial factor, Geometric Local Attention with Spatiotemporal RoPE performs reliable 3D matching while suppressing attention sinks. Finally, Metric Readout Tokens recover stable scale and rigid pose directly from the persistent geometric state. Extensive experiments show that HorizonStream, trained on only 48‑frame clips, generalizes stably to sequences exceeding 10,000\ frames with constant memory and linear time, achieving state‑of‑the‑art streaming 3D reconstruction performance. Project Page: https://3dagentworld.github.io/horizonstream/
Authors:Katharina Schmid, Nicolas von Lützow, Jozef Hladký, Angela Dai, Matthias Nießner
Abstract:
We introduce a new approach to high‑fidelity 3D scene reconstruction from multi‑view RGB images that tightly couples reconstruction with a strong generative 3D prior. We cast scene reconstruction as conditional 3D generation over a set of spatially‑localized, overlapping chunks that together tile the scene, scaling generation to large scene extents. Crucially, we inherit the fidelity and completeness of state‑of‑the‑art generative shape models ‑‑ we use Trellis.2 as an example ‑‑ which we generalize to the scene level. To this end, we propose a projection‑based conditioning mechanism that lifts posed multi‑view image features into a coherent 3D representation aligned with the generative model, independent of view ordering and spatially anchored to the scene, yielding high‑fidelity, multi‑view consistent generated geometry. This enables lifting the strong object‑level prior of Trellis.2 to multi‑view, scene‑scale generation, producing faithful, editable PBR mesh reconstructions of indoor environments. As a result, we obtain high‑fidelity results that outperform cutting‑edge reconstruction methods by 16%.
Authors:Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang, Derek Hoiem
Abstract:
Keyframe selection is a direct way to provide verifiable visual evidence for long‑video question answering (QA). Queries differ in what they require, and finding the right frames depends on knowing what to look for. Existing keyframe selectors either score every frame against a single query, or decompose the query into a fixed schema evaluated by a single visual tool. We propose ToolMerge, a keyframe retrieval method based on decomposition and merging: an Large Language Model (LLM) based planner decomposes the query into tool calls and specifies how their per‑tool rankings are merged using boolean operators. To evaluate retrieval directly, we construct Molmo‑2 Moments (M2M), a benchmark in which every question is anchored to a specific time interval by construction. Across QA, question retrieval, and caption retrieval, ToolMerge is competitive with prior keyframe selectors, most notably on caption retrieval, outperforming other methods by 5%. Code and data can be found at https://github.com/michalsr/ToolMerge .
Authors:Bo Peng, Jie Lu, Guangquan Zhang, Zhen Fang
Abstract:
Aiming at identifying unexpected inputs from unknown classes, out‑of‑distribution (OOD) detection has emerged as a pivotal approach to enhancing the reliability of machine learning models. This paper focuses on the burgeoning paradigm of post‑hoc OOD detection with pre‑trained vision‑language models (VLMs), where a popular pipeline is to detect OOD inputs by examining their affinities between ID labels and negative labels, i.e., those semantically different from ID labels. Due to the unavailability of target OOD labels, existing works predominantly rely on heuristic rules to mine negative labels from unlabeled wild corpus data. Despite the empirical success, we argue that the power of VLM‑based OOD detection has yet to be fully unleashed since the notorious false negative problem is far from addressed in the literature. With this motivation, we are interested in addressing the challenge of mining true negative labels for OOD scoring. To this end, we develop a theoretical framework for correcting the sampling bias of negatives labels by indirectly approximating the distribution of negative labels. Perhaps surprisingly, we show that the debiased negative mining can be naturally converted into Monte‑Carlo sampling based on ID labels and the unlabeled wild corpus data. Extensive experiments empirically manifest that our method establishes a new state‑of‑the‑art in a variety of OOD detection setups. Code is publicly available at \hrefhttps://github.com/60pen9/Debiased‑Negative‑Mining‑Improves‑OOD‑Detection‑with‑Pre‑trained‑VLMs\textcolorredhere.
Authors:Chenyu Wu, Wanhua Li, Zhu-Tian Chen, Hanspeter Pfister
Abstract:
Reconstructing dynamic 3D scenes from monocular videos is a fundamental yet highly challenging task, as real‑world motions often involve both long‑term smooth transformations and short‑term complex deformations. Existing methods either struggle to maintain temporal consistency or fail to capture high‑frequency dynamics due to limited motion modeling capacity. In this work, we present Rigid‑aware 4D Gaussian Splatting (RiGS), which simultaneously captures motions across multiple temporal scales. Specifically, RiGS introduces three types of Gaussian primitives: static, rigid, and transient, which represent static backgrounds, long‑term low‑frequency motions, and short‑term high‑frequency dynamics, respectively. An object‑wise dynamic mask is proposed to aggregate long‑range spatiotemporal motion information and guide the decomposition of static and dynamic regions. To jointly model motion across scales, rigid Gaussians are allowed to transition into transient Gaussians based on their temporal duration, and both are optimized under scene flow guidance, providing dense 3D motion supervision. Extensive experiments demonstrate that RiGS achieves state‑of‑the‑art performance on novel view synthesis benchmarks. Code is available at \hyperlinkhttps://github.com/ladvu/RiGShttps://github.com/ladvu/RiGS.
Authors:Liupeng Li, Haoqian Kang, Zhenyu Lu, Jinpeng Wang, Bin Chen, Ke Chen, Yaowei Wang
Abstract:
High‑resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade‑off between coverage and efficiency. Visual expert‑assisted search is efficient but prone to blind spots when proposals fail, whereas scan‑based search guarantees coverage at the cost of computational redundancy and semantic fragmentation. To address this dilemma, we introduce CVSearch, a training‑free adaptive framework that dynamically schedules search strategies via an Assess‑then‑Search workflow. Specifically, CVSearch first invokes expert‑assisted search when global information is insufficient, and only triggers a novel semantic‑aware scanning mechanism upon failure. Distinct from rigid grid partitioning, this efficient scanning paradigm incorporates Semantic Guided Adaptive Patching to decompose images into semantically consistent regions, effectively mitigating object fragmentation. Furthermore, we devise a Dynamic Bottom‑Up Search strategy driven by a Visual Complexity prior to enable efficient and precise iterative exploration of local details. Extensive experiments on HR benchmarks demonstrate that CVSearch achieves state‑of‑the‑art accuracy while substantially improving search efficiency. Code is released at https://github.com/liliupeng28/ICML26‑CVSearch.
Authors:Chanho Lee, Seunghee Koh, Yunho Jeon, Junmo Kim
Abstract:
The DEtection TRansformer (DETR) is a powerful end‑to‑end object detector, yet its one‑to‑one matching strategy suffers from slow convergence and low recall. A common approach to address this issue is to use one‑to‑many label assignment to provide more positive samples. However, existing methods that use one‑to‑many matching as an auxiliary objective lead to increased training costs, with their auxiliary decoders discarded during inference. To address this limitation, we propose MDS‑DETR, which leverages both one‑to‑one and one‑to‑many supervision within a single decoder. Specifically, we introduce a Masked Duplicate Suppressor (MDS) that injects asymmetry into self‑attention via confidence‑based causal masking. MDS filters out the duplicates generated by the one‑to‑many supervised layer, enables explainable, duplicate‑free predictions in a fully end‑to‑end framework. MDS‑DETR outperforms existing one‑to‑many DETR variants such as MS‑DETR, MR.DETR and Relation‑DETR, without relying on any additional queries or auxiliary decoders. Under a 12‑epoch training schedule on MS COCO with a ResNet‑50 backbone, MDS‑DETR achieves a +2.8 mAP improvement over Deformable‑DETR with only a 5% increase in training time, and outperforms the state‑of‑the‑art MR.DETR by +0.3 mAP while being even 20% faster in training. Our code and models are available at \hrefhttps://github.com/dcholee/mds‑detrhttps://github.com/DChoLee/MDS‑DETR.
Authors:Jongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin Yoon
Abstract:
Dataset distillation compresses large training sets into compact synthetic datasets while preserving downstream performance. As modern systems increasingly operate on paired vision‑language inputs, multimodal distillation must preserve representation quality and cross‑modal alignment under tight compute and memory budgets, yet prior methods often require heavy computes and overlook their correlations. To address this, we present Multimodal Distribution Matching (MDM), a geometry‑aware framework for efficient and generalizable multimodal distillation. Specifically, MDM integrates complementary components at the data, model, and loss levels. At the data level, it initializes synthetic image‑text pairs by sampling from clusters in the joint embedding space. At the model level, it forms a mixed teacher by interpolating independently fine‑tuned models in weight space according to their angular deviation from the pretrained anchor. At the loss level, it matches joint distributions on the unit hypersphere using a geometry‑aware matching objective that exploits the joint features in the cross‑modal agreement and discrepancy directions along with symmetric contrastive learning. Across image‑text retrieval benchmarks with cross‑architecture evaluation, MDM yields compact synthetic sets that preserve multimodal semantics, substantially reduce distillation cost, and remain robust across architectures.
Authors:Jiaqi Feng, Justin Cui, Yuanhao Ban, Cho-Jui Hsieh
Abstract:
Recent advances have substantially improved real‑time interactive video generation in the autoregressive regime. However, most existing few‑step autoregressive video generation methods, often distilled from a corresponding many‑step teacher, default to a 4‑step sampling configuration, which still incurs considerable latency during deployment and suffers from severe quality degradation when the number of sampling steps is further reduced, particularly in the one‑step setting. Trajectory‑style consistency distillation methods often produce videos with weak dynamics, while DMD‑based approaches, such as Self‑Forcing, tend to yield blurry frames. To address this challenge, we propose One‑Forcing, a simple yet effective approach which augments the DMD objective with an auxiliary GAN loss for high‑quality and efficient one‑step video generation. Experiments on VBench show that One‑Forcing achieves a total score of 83.76, establishing state‑of‑the‑art performance among one‑step causal video generation methods and remaining competitive with strong many‑step approaches. We further demonstrate that one‑step framewise autoregressive generation can be achieved stably with merely one‑third of the training cost of the chunkwise model, a setting that prior methods have failed to achieve successfully.
Authors:Jie Hu, Zixiang Gao, Yutong He, Kun Yuan
Abstract:
Diffusion transformers have achieved remarkable success in high‑quality video generation, yet their reliance on spatiotemporal 3D full attention incurs prohibitive computational cost due to the quadratic complexity of attention. Block sparse attention is a common approach to mitigate this by focusing computation on important regions. However, attention maps in DiTs exhibit inherently dynamic and fine‑grained sparsity, which causes existing block sparse attention methods to degrade significantly in quality, especially at high sparsity ratios. In this paper, we revisit block sparse attention and derive a theoretical lower bound on attention recall to characterize the key factors governing its effectiveness. Guided by these insights, we propose DFSAttn, a training‑free sparse attention framework that enables dynamic, fine‑grained sparsification efficiently. DFSAttn incorporates three core designs: Hilbert curve‑based token reordering to achieve fine‑grained sparsity while preserving efficient GPU execution, hierarchical block scoring for accurate block importance estimation, and sparse mask caching with adaptive ratios to balance accuracy and efficiency. Experimental results demonstrate that DFSAttn consistently outperforms prior methods under high sparsity, achieving up to 2.1× end‑to‑end speedup while maintaining high generation quality. Our code is open‑sourced and available at https://github.com/jessica‑hujie/DFSAttn.
Authors:Eunwoo Heo, Kyeongkook Seo, Jaejun Yoo
Abstract:
The explosive growth of open‑source model repositories has created a Model Jungle, where checkpoints are frequently shared without adequate documentation or metadata. While weight‑space learning offers a pathway to identify and analyze these models directly from their parameters, processing full‑scale weights is computationally prohibitive. Probing‑based methods have emerged as a lightweight alternative, extracting permutation‑equivariant representations via learnable probe vectors. However, existing probing methods are limited by a single‑view design: they capture first‑order structures but fail to encode the rich, higher‑order correlation patterns inherent in row‑column interactions. To bridge this gap, we introduce MVProbe, a multi‑perspective probing framework that synthesizes first‑order signals with interaction‑aware (Gram‑based) views. Our approach is theoretically grounded; we analyze the scaling laws of different probing orders to derive a principled standardization and fusion strategy that ensures balanced contributions from all branches. On the Model Jungle benchmark, MVProbe consistently outperforms the state‑of‑the‑art ProbeX across diverse architectures, including discriminative backbones (ResNet, SupViT, MAE, DINO) and large‑scale generative LoRA adapters (Stable Diffusion LoRA).
Authors:Zizhao Tong, Yeying Jin, Hongfeng Lai, Zeqing Wang, Zhaohu Xing, Kexu Cheng, Haoran Xu, Zhao Pu, Shangwen Zhu, Ruili Feng, Jian Zhao, Yan Zhang, Hao Tang, Ling Shao
Abstract:
Interactive world models for first‑person shooter (FPS) games must resolve high‑frequency overlapping control signals at every frame without disrupting unaffected regions. Existing methods inject actions globally and train on single titles, failing under dense FPS inputs. We observe that FPS actions are spatially selective: discrete events such as firing or reloading affect only a localized region around the weapon (the scope), while continuous camera and movement signals govern stable surroundings. We propose SCOPE, which inserts a conditioning module into each transformer block of a pretrained video diffusion model. It reshapes features into per‑pixel temporal sequences so that each position computes its action response from local visual content. This separates in‑scope effects from out‑of‑scope generation without segmentation labels. We also introduce CrossFPS, the first multi‑game FPS dataset with frame‑aligned action telemetry. It comprises 69K clips from 7 titles with 10‑DoF controller signals, curated to remove gameplay bias. The model learns general visual‑to‑action mappings rather than game‑specific patterns, enabling zero‑shot transfer to unseen scenes. Experiments confirm strong action responsiveness, precise scope separation, and effective cross‑game generalization.
Authors:Yilong Liu, Wanhua Li, Chen Zhu-Tian, Hanspeter Pfister
Abstract:
We present LangFlash, a feed‑forward framework for 3D Language Gaussian Splatting that reconstructs 3D scenes parameterized by Gaussian primitives enriched with language‑aligned semantic features from sparse unposed multi‑view images. Unlike optimization‑based 3D methods, LangFlash directly predicts the geometry and semantics in a single forward pass, enabling low‑latency 3D reconstruction and language‑consistent scene understanding. To support large‑scale training, we enriched the RealEstate10k dataset with coherent and dense semantic information for 3D semantic supervision. Furthermore, we propose a sparse semantic encoding scheme that combines a global semantic dictionary with locally varying per‑primitive weights, preserving high‑level linguistic information, while reducing representation complexity. Experimental results show that LangFlash achieves superior novel view synthesis and semantic consistency compared with previous methods. This study establishes a new paradigm for pose‑free, language‑grounded 3D scene reconstruction, advancing generalizable 3D vision and multimodal scene understanding. Demo is available at https://liylo.github.io/langflash.github.io/.
Authors:Shaoqing Duan, Haofei Song, Xintian Mao, Qingli Li, Yan Wang
Abstract:
Defocus deblurring in pathological microscopy remains challenging due to the spatially varying and locally discontinuous nature of optical blur induced by a position‑dependent integral imaging process.
Existing deep learning methods, constrained by shift‑invariance assumptions and limited interpretability, are not well suited to such heterogeneous blur patterns.
Neural operators provide a principled alternative by modeling defocus formation directly as an integral operator, offering a new perspective on defocus deblurring.
However, most existing neural operator architectures for low‑level vision rely on globally parameterized kernels that assume smoothness and stationarity, limiting their ability to model heterogeneous and locally discontinuous blur patterns.
To address this limitation, we propose the Discontinuous Galerkin Neural Operator (DGNO), which parameterizes the integral kernel using a discontinuous Galerkin formulation with element‑local volume operators and interface numerical fluxes.
DGNO provides a principled combination of locality,
heterogeneity modeling, and global coherence while preserving the underlying physics
of optical image formation.
Extensive and insightful experiments demonstrate that DGNO surpasses state‑of‑the‑arts, delivering sharper reconstructions, robust handling of spatially varying blur, and scalable high‑resolution performance. The code will be released at https://github.com/DeepMed‑Lab‑ECNU/Single‑Image‑Deblur.
Authors:Xiyang Wang, Xinlin Wang, Tingguang Zhou, Gong Chen, Xingtai Gui, Zhi Xu, Xiaolei Wu, Feiyang Tan, Hangning Zhou, Mu Yang
Abstract:
Current end‑to‑end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction‑aware temporal dependencies via causal factorization, but their step‑wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints, making them unreliable in interactive and safety‑critical scenarios. This dichotomy reveals a deeper issue: existing methods treat causal modeling and global optimization as separate paradigms, without a principled way to unify them within a single trajectory distribution. To address this, we propose ChainFlow‑VLA, which unifies causal generation and global refinement within a unified probabilistic framework. We formulate planning as a mixture over AR‑induced modes and learn Vision‑Language Model (VLM)‑conditioned residual distributions over these modes. An autoregressive generator (Chain) produces a discrete set of causal trajectory modes, followed by a diffusion‑based refiner (Flow) that leverages VLM hidden states as semantic priors to perform mode‑conditioned correction in residual space while preserving causal structure. This straightforward conditioning seamlessly injects high‑level scene understanding into fine‑grained trajectory adjustments. Experiments demonstrate that ChainFlow‑VLA achieves robust planning in ambiguous and long‑tail scenarios, achieving a state‑of‑the‑art score of 94.85 on the NAVSIM v1 leaderboard, matching human‑level performance (94.8). Code will be available at https://github.com/AFARI‑Research/ChainFlow‑VLA.
Authors:Mengke Li, Haiquan Ling, Lihao Chen, Yang Lu, Yiqun Zhang, Hui Huang
Abstract:
Learning from real‑world data is frequently hindered by the compound challenge of long‑tailed class distributions and noisy annotations. Existing methods partially address these issues but typically ignore the non‑uniform impact of label noise across classes, resulting in ineffective correction for tail classes and over‑regularization for head classes. To address this issue, we propose Class‑Adaptive Rectification with Experts (CARE), a parameter‑efficient framework that leverages three complementary supervision sources from vision‑language models (VLM): observed noisy labels, VLM text embeddings, and visual features. CARE introduces a class‑adaptive expert consensus mechanism that enforces stricter agreement for tail classes and more permissive agreement for head classes based on class frequency. By aggregating high‑confidence predictions across these sources, CARE filters unreliable signals and recalibrates class distributions, yielding more reliable rectification under long‑tailed distributions. Extensive experiments on both synthetic and real‑world benchmarks demonstrate that CARE consistently outperforms state‑of‑the‑art methods, achieving up to 3.0% performance gains. The source code is available at https://github.com/qwq123‑study/CARE.
Authors:Jean-Guillaume Durand, Panagiotis Kouvaros, Maxime Gariel, Alessio Lomuscio
Abstract:
The adoption of vision neural networks in regulated industries requires formal robustness guarantees, especially in safety‑critical domains such as healthcare, autonomous vehicles, and aerospace. However, current approaches are confined to incomplete statistical verification or robustness to \ell_p‑norm and affine transforms, which cover only a narrow subset of perturbations to the image formation process. In particular, robustness to camera motion remains an open problem despite being key to deploy many vision applications. We present a formal verification approach that targets robustness against 3D motion perturbations of the capturing camera. We first establish a closed‑form mapping from camera pose to pixel values. By analyzing the continuity properties of the resulting homographies, we show that recent work on Lipschitz optimization and piecewise continuity can be extended to derive tight linear bounds on perturbed pixel values. Our approach applies to scenes with predominantly planar structure, such as ground planes in augmented reality, road markings and traffic signs in autonomous driving, or planar workspaces in robotic manipulation. This enables the first formal verification of projective geometry transforms, without complex simulation, surrogate networks, or explicit image‑formation models. We validate our implementation and show up to 89% speedup and 7% tighter bounds over prior work. We then evaluate our method on the VNN‑COMP benchmark and reveal systematic weaknesses to projective perturbations. Finally, we demonstrate a real‑world case study on a safety‑critical runway classifier, highlighting practical vulnerabilities to camera motion, and addressing a key challenge in the certification of learned models. Data and code are publicly available at https://github.com/jeangud/homography‑verification .
Authors:Wenxuan Peng, Bharath Hariharan, Hadar Averbuch-Elor
Abstract:
Despite recent progress, text‑to‑image models still struggle to generate semantically diverse and compositionally accurate multi‑person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded interactions. In this work, we bridge this gap by introducing a dual pose‑image representation that brings person‑centric structural priors into pretrained diffusion transformers. Our model jointly predicts a 2D pose visualization image and its corresponding RGB image, enabling structure and appearance to co‑evolve during learning. At its core, a cross‑modal alignment scheme binds text, pose, and image representations, ensuring consistent grounding across modalities. Furthermore, we design an iterative scene construction scheme, progressively generating complex multi‑human interactions while effectively decomposing the overall generation complexity. Extensive experiments demonstrate that our method substantially improves prompt alignment and scene diversity in multi‑person image generation.
Authors:Jun Seong Lee, Samyeul Noh, Changki Sung, Hyun Myung
Abstract:
Remote photoplethysmography (rPPG) enables non‑contact measurement of physiological signals from facial videos, offering strong potential for remote healthcare and daily health monitoring. Driven by this potential, various deep learning‑based rPPG methods have been proposed to improve rPPG estimation. However, previous deep learning‑based rPPG methods have paid little attention to the quality of training labels and their impact on model learning. Contact‑based PPG signals used as training labels often contain noise and variability caused by motion artifacts, inconsistent sensor contact, and morphological distortions. Such label inconsistency can lead models to overfit to the label noise and variability and consequently degrade generalization performance. To address this issue, we propose LQ‑rPPG, a label‑quantized coarse‑to‑fine learning framework for robust rPPG estimation. LQ‑rPPG consists of a label quantization module and a coarse‑to‑fine rPPG estimation model. The label quantization module transforms continuous PPG signals into multi‑bit quantized pseudo labels with reduced noise and variability. The coarse‑to‑fine estimation model progressively refines rPPG signals under hierarchical supervision guided by the multi‑bit pseudo labels. This design alleviates overfitting to label‑specific variations and enables the model to learn structured and consistent representations. As a result, LQ‑rPPG achieves robust and generalizable rPPG estimation even under challenging conditions. Experiments on multiple benchmark datasets demonstrate that LQ‑rPPG achieves strong performance in both intra‑ and cross‑dataset evaluations, while reducing parameters and multiply‑accumulate operations by 88% and 29%, respectively, and increasing throughput by 191%. The code is available at https://github.com/Anonymous‑repo‑code/LQ‑rPPG.
Authors:Chenxu Wang, Yuxuan Li, Yunheng Li, Xiang Li, Jingyuan Xia, Qibin Hou
Abstract:
Existing language‑image pre‑training for remote sensing object detection is constrained by Monolithic Label Learning, which relies on exhaustively enumerating open‑set categories via black‑box data to acquire fine‑grained representations, creating a dependency incompatible with the domain's inherent data scarcity. To transcend this bottleneck, we propose SLIP‑RS, establishing a Structured‑Attribute Decoupling Paradigm that maps the open‑ended category space into a finite, physically meaningful attribute space, unlocking fine‑grained discriminability via explicit structural logic. This paradigm is realized via two technical pillars: (1) Structured‑Attribute Contrastive Learning, which enforces the learning of decoupled intrinsic visual logic via combinatorial attribute augmentation; and (2) Conformal Attribute Reliability Engine, which leverages conformal prediction theory to rigorously distill high‑fidelity supervision from noisy sources, yielding RS‑Attribute‑15M, the largest dataset with over 15 million attribute annotations. Extensive experiments demonstrate that SLIP‑RS establishes unprecedented performance in fine‑grained detection and cross‑domain generalization, validating structured attributes as a vital foundation for remote sensing. Code: https://github.com/facias914/SLIP‑RS.
Authors:Jiahe Meng, Weiming Zeng, Yueyang Li, Bo Chai, Hongjie Yan, Zhiguo Zhang, Wai Ting Siok, Nizhuan Wang
Abstract:
Electroencephalography (EEG) visual decoding remains challenging due to the modality gap between low‑SNR neural signals and highly structured vision‑‑language spaces, making direct cross‑modal alignment unstable. To address this, we propose STAMBRIDGE, a versatile two‑stage framework that sequentially tackles feature conditioning and cross‑modal alignment. First, we introduce a Spectral‑Temporal Amplitude‑aware Modulation (STAM) to extract well‑conditioned EEG representations. By replacing hard frequency masking with amplitude‑derived soft channel weighting and multi‑scale temporal convolutions, STAM explicitly preserves frequency‑aware transients while reducing the risk of time‑domain ringing artifacts. Building upon these robust neural features, we further introduce a model‑agnostic Mid‑Feature Semantic Bridge (MFSB) that constructs a regularized intermediate space through directed cross‑modal interactions, enabling staged distillation and more stable semantic alignment. Experiments on the THINGS‑EEG benchmark show competitive 200‑way zero‑shot retrieval performance, with 34.50% Top‑1 and 65.95% Top‑5 accuracy. In addition, embeddings learned by STAMBRIDGE produce semantically coherent image reconstructions with a diffusion model, demonstrating robust EEG‑to‑vision semantic alignment. The code is available at: https://github.com/thabeatmjh/STAMBRIDGE.
Authors:Yannick Kirchhoff, Maximilian Rokuss, Daniel Philipp Mertens, David Füller, Benjamin Hamm, Andreas Schreyer, Oliver Ritter, Klaus Maier-Hein
Abstract:
Tracking tumor lesions across serial CT scans is essential for oncological response assessment. Existing automated methods face a fundamental trade‑off: end‑to‑end trackers achieve high automation but offer no opportunity to correct silent tracking failures, while decoupled registration‑segmentation pipelines permit user verification yet discard the lesion's prior appearance, limiting accuracy in ambiguous cases. In this work, we propose a Verified Tracking paradigm: a clinician verifies a registration‑proposed prompt, which the model leverages alongside the baseline lesion appearance to resolve segmentation ambiguities. We present a unified framework combining early spatial prompt fusion with latent temporal difference weighting for longitudinally‑informed segmentation. To address data scarcity, we leverage large‑scale synthetic pretraining, proving essential for exploiting longitudinal context, improving performance by up to 4.5 Dice points over training from scratch. Our approach secured first place in the MICCAI autoPET IV challenge. We further curate and release PanTrack, a new longitudinal pancreatic cancer benchmark, to assess out‑of‑distribution generalization. Experiments show that our model outperforms prior work in both fully automatic and the proposed verified tracking setting offering a clinically safe middle ground between automation and control. Code, model and dataset will be released at https://github.com/MIC‑DKFZ/LongiSeg
Authors:Hyeongmuk Lim, Youngbum Hur
Abstract:
Existing Video Anomaly Detection (VAD) methods typically rely on task‑specific training, leading to strong domain dependency and high training costs. Moreover, most existing methods output only scalar anomaly scores, providing limited insight into why specific events are considered abnormal. Recent advances in Vision‑Language Models (VLMs) have enabled both anomaly detection and human‑interpretable reasoning. However, many VLM‑based approaches still require additional training steps (e.g., instruction tuning or verbalized learning) or external Large Language Models (LLMs), incurring further training costs and inference overhead. To address these challenges, we propose CoReVAD, a contextual reasoning framework for training‑free video anomaly detection that operates with a single frozen VLM. CoReVAD directly generates anomaly scores and temporal descriptions from the VLM. To mitigate noise in generative outputs, we introduce a Local Response Cleaning (LRC) module based on local vision‑text alignment. Furthermore, global temporal context and progression are incorporated through softmax‑based refinement, Gaussian smoothing, and position weighting. Experiments on UCF‑Crime and XD‑Violence demonstrate that CoReVAD achieves competitive performance among training‑free methods while providing reliable and interpretable explanations. Our official code is available at: https://github.com/Muk‑00/CoReVAD
Authors:Chengyi Zhang, Zi Ye, Ziyang Wang
Abstract:
Reliable visual understanding in robot‑assisted and minimally invasive surgery (RMIS/MIS) demands more than accurate masks: in clinical practice, clinicians pose language‑like questions about procedural context, visibility, artefacts, and the presence of anatomical structures and surgical instruments, often under degraded views caused by occlusion, smoke, bleeding, and specular highlights. We present RoboSurg‑VQA, a segmentation‑aware visual question answering (VQA) benchmark built by repurposing public surgical segmentation datasets under a shared schema. Each frame is paired with a fixed set of clinically motivated questions spanning procedure context, anatomy (including region), imaging modality/view, surgical artefacts, image quality, and basic visibility and spatial attributes, with closed answer sets to enable consistent evaluation. To scale annotation, we generate candidate answers via constrained prompting with automatic validity and consistency checks, followed by human auditing to improve plausibility and label consistency. We report benchmark statistics, sanity baselines, and common evaluation challenges under challenging surgical conditions. The code will be available on https://github.com/ziyangwang007/Robosurg‑VQA.
Authors:Borja Carrillo-Perez
Abstract:
This report presents a lightweight modification to the DETR‑based fusion transformer baseline for the MaCVi 2026 Vision‑to‑Chart data association challenge. The challenge baseline decoder receives per‑buoy queries encoding world‑space distance and bearing, forcing the transformer to implicitly learn the complex geometric projection from world coordinates to image pixels. Instead, this work trains an additional dedicated MLP, QueryMLP, to explicitly predict the buoy's waterline contact point in the image from chart measurements and IMU orientation data. The predicted pixel coordinates are appended to the baseline decoder query vector, providing a direct spatial prior per buoy and reducing the geometric reasoning burden on the transformer decoder. On the challenge leaderboard, the presented approach achieves an Overall score of 0.7386, with F1 = 0.8055 and mIoU = 0.6718, on the held‑out test set, placing second among all submissions.
Authors:Jongseo Lee, Hyuntak Lee, Sunghun Kim, Sooa Kim, Jihoon Chung, Jinwoo Choi
Abstract:
Video Large Language Models (Video‑LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed image‑plane motion direction. On simple videos of a single object moving left, right, up, or down, most Video‑LLMs perform near chance, with above‑chance cases largely attributable to prediction biases rather than genuine direction understanding. We call this failure directional motion blindness. We localize the failure by tracing motion direction information through the Video‑LLM pipeline. Motion direction remains linearly accessible from the vision encoder, projector, and LLM hidden states, but the readout fails to bind this signal to the correct verbal answer option, revealing a direction binding gap. Although synthetic motion direction instruction tuning reduces this gap on the source domain, motion direction concept vector analysis shows that visual complexity weakens the signal magnitude and limits out‑of‑domain generalization. We introduce MoDirect, a dataset family for motion direction instruction tuning and evaluation, and DeltaDirect, a diagnosis‑driven, projector‑level objective that predicts normalized 2‑D motion vectors from adjacent‑frame feature deltas. On MoDirect‑SynBench, instruction tuning with DeltaDirect improves motion direction accuracy from 25.9% to 85.4%. On MoDirect‑RealBench, DeltaDirect improves real‑world motion direction accuracy by 21.9 points over the vanilla baseline without real‑world tuning data, while preserving standard video‑understanding performance. Code: https://github.com/KHU‑VLL/DeltaDirect
Authors:Wenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li, Hang Yin, Huangxing Chen, Wenzhao Zheng, Jianjiang Feng, Jie Zhou, Jiwen Lu
Abstract:
Vision‑and‑Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state‑of‑the‑art methods leverage the reasoning capabilities of Vision‑Language Models (VLMs) for end‑to‑end action prediction, they often lack an explicit and explainable understanding of the relationships between the agent, the instruction, and the scene. Conversely, explicitly building a scene map for heuristic planning is intuitively appealing but relies on additional 3D sensors and hinders large‑scale vision‑language pre‑training. To bridge this gap, we propose AwareVLN, a novel framework that equips the navigation model with a self‑aware reasoning mechanism, enabling it to understand the agent's state and task progress in a fully end‑to‑end and data‑driven manner. Our approach features two key innovations: (1) a structural reasoning module that fosters spatial and task‑oriented self‑awareness, and (2) an automatic data engine with progress division for effective training. Extensive experiments on various datasets in Habitat simulator show our AwareVLN significantly outperforms previous state‑of‑the‑art vision‑language navigation methods. Project page: https://gwxuan.github.io/AwareVLN/.
Authors:Wenxuan Guo, Ziyuan Li, Meng Zhang, Yichen Liu, Yimeng Dong, Chuxi Xu, Yunfei Wei, Ze Chen, Erjin Zhou, Jianjiang Feng
Abstract:
Vision‑Language‑Action (VLA) models have shown strong potential for general‑purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily rely on textual instructions and struggle to resolve spatial ambiguity in complex scenes with multiple similar objects. To address this limitation, we introduce gesture as a parallel instruction modality and propose a Gesture‑aware Vision‑Language‑Action model (GesVLA). Our approach encodes gesture features directly into the latent space, enabling them to participate in both high‑level reasoning and low‑level action generation, and adopts a dual‑VLM architecture to achieve tight coupling between gesture representations and action policies. At the data level, we construct a scalable gesture data generation pipeline by rendering hand models onto real‑world scene images. This reduces the sim‑to‑real visual gap while producing rich data with diverse motion patterns and corresponding pointing annotations. In addition, we employ a two‑stage training strategy to equip the model with both gesture perception and action prediction capabilities. We evaluate our approach on multiple real‑world robotic tasks, including a controlled block manipulation task for validation and more practical scenarios such as product and produce selection. Experimental results show that incorporating gesture consistently improves target grounding accuracy and human‑robot interaction efficiency, especially in complex and cluttered environments. Project page: https://gwxuan.github.io/GesVLA/.
Authors:Jung Yi, Minjae Kim, Paul Hyunbin Cho, Wooseok Jang, Sangdoo Yun, Seungryong Kim
Abstract:
Autoregressive video diffusion models have enabled real‑time, action‑conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains an open problem. Full KV‑cache attention preserves this consistency but breaks real‑time constraints: memory footprint and attention cost grow linearly with rollout length. Sliding window inference restores throughput but discards long‑term consistency. We propose WorldKV, a training‑free framework with two components: World Retrieval and World Compression. World Retrieval stores evicted KV‑cache chunks in GPU/CPU memory and selectively retrieves scene‑relevant chunks via camera/ action correspondence, inserting them back into the native attention window without re‑encoding. World Compression prunes redundant tokens within each chunk via key‑key similarity to an anchor frame, halving per‑chunk storage to fit 2x more history under a fixed budget. On Matrix‑Game‑2.0 and LingBot‑ World‑Fast, WorldKV matches or exceeds full‑KV memory fidelity at roughly 2x the throughput, and is competitive with memory‑trained baselines without any fine‑tuning. Project Page: https://cvlab‑kaist.github.io/WorldKV/
Authors:Yannick Porto, Renato Martins, Thomas Chalumeau, Cedric Demonceaux
Abstract:
Robustness to domain changes is a key capability for effective deployment of human action recognition systems in real‑world scenarios, where action categories at inference can present important domain shifts or even unseen actions from training. In this context, improving the recognition capabilities of Zero‑Shot Action Recognition models (ZSAR), without requiring strong annotation efforts, remains a central challenge. Most ZSAR approaches assume that actions are observed under geometric conditions similar to those seen during training. In practice, variations in human body orientation and camera viewpoint add a significant domain gap in ZSAR, substantially limiting generalization to novel action‑motion combinations. In this context, this paper presents a novel orientation‑aware action recognition approach with improved cross‑domain capabilities. Our approach combines motion cues of multiple camera viewpoints and text descriptions of human actions in the training phase. We present a new orientation‑aware motion encoding network to learn different motion features, and adapt a specific orientation‑aware text prompt to match the corresponding features at inference. Extensive experiments demonstrate that the proposed method consistently improves ZSAR performance across different recognition benchmarks, outperforming recent state‑of‑the‑art zero‑shot approaches on NTU‑RGB+D, BABEL, NW‑UCLA, and on two surveillance datasets. In addition, the learned representations exhibit strong transfer learning capabilities, yielding competitive performance on both cross‑domain and same‑domain recognition of seen actions. Code and trained models are available at: https://icb‑vision‑ai.github.io/OrientationAware‑HAR
Authors:Yannick Porto, Renato Martins, Thomas Chalumeau, Cedric Demonceaux
Abstract:
Viewpoint change invariance and action temporal consistency are critical aspects for the effective deployment of human action detection of untrimmed videos. Existing appearance‑based video detection methods often struggle with limited viewpoint diversity during training, while motion‑based detection approaches frequently fail to model fine‑grained temporal relationships across consecutive motion windows. This paper introduces a novel two‑stage action detection approach designed to improve both view‑invariance and global temporal coherence properties. In the first stage, we extract motion features from augmented virtual viewpoints, solely used at training. Then, the second stage introduces a new view‑invariant, multi‑scale temporal encoder based on selective state‑space sequence modelling to aggregate information across viewpoints and time scales. Experiments on PKU‑MMD and BABEL benchmarks demonstrate that this approach significantly outperforms state‑of‑the‑art methods in all considered splits. Code and trained models are available at: https://icb‑vision‑ai.github.io/HydraView‑TAD
Authors:Javad Rajabi, Kimia Shaban, Koorosh Roohi, David B. Lindell, Babak Taati
Abstract:
Diffusion transformers (DiTs) have emerged as a dominant architecture for text‑to‑image generation, yet their performance drops when generating at resolutions beyond their training range. Existing training‑free approaches mitigate this by modifying inference‑time attention behavior, often through Rotary Position Embeddings (RoPE) extrapolation combined with attention scaling. However, these strategies apply a uniform and content‑agnostic scaling across RoPE components with distinct frequency characteristics, inducing a trade‑off between preserving global structure and recovering fine detail. We introduce SEGA, a training‑free method that dynamically scales attention across RoPE components according to the latent's spatial‑frequency structure at each denoising step. This adaptive scaling improves both structural coherence and fine‑detail fidelity. Experiments show that SEGA consistently improves high‑resolution synthesis across multiple target resolutions, outperforming state‑of‑the‑art training‑free baselines.
Authors:Zhenyu Lu, Liupeng Li, Jinpeng Wang, Haoqian Kang, Yan Feng, Ke Chen, Yaowei Wang
Abstract:
While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoning to visual perception. Current methods, such as latent query alignment, are end‑to‑end yet opaque "black boxes". Conversely, textual localization readout is merely readable, not truly interpretable, often functioning as an unconstrained post‑hoc step. To bridge this interpretability gap, we propose SegCompass, an end‑to‑end model that leverages a Sparse Autoencoder (SAE) to forge an explicit, interpretable, and differentiable alignment pathway. Given an image‑instruction pair, SegCompass first generates a chain‑of‑thought (CoT) trace. The core of our method is an SAE that maps both the CoT and visual tokens into a shared, high‑dimensional sparse concept space. A query codebook selects salient concepts from this space, which are then spatially grounded by a slot mapper into a multi‑slot heatmap that guides the final mask decoder. The entire model is trained jointly, unifying reinforcement learning for the reasoning path with standard segmentation supervision. This SAE‑driven interface provides a "white‑box" connection that is significantly more traceable than latent queries and more coherent than textual readouts. Extensive experiments on five challenging benchmarks demonstrate that SegCompass matches or surpasses state‑of‑the‑art performance. Crucially, our visual and quantitative analyses show a strong correlation between the quality of the learned sparse concepts and final mask accuracy, confirming that SegCompass achieves superior results through its enhanced and inspectable alignment. Code is available at https://github.com/ZhenyuLU‑Heliodore/SegCompass.
Authors:Erjian Zhang, Yatong Hao, Liejun Wang, Zhiqing Guo
Abstract:
While multi‑task learning based automatic radiology report generation (RRG) is widely adopted to ensure clinical consistency, most focus on architectural designs yet remain limited to coarse linear scalarization strategies. These strategies cannot effectively balance the hard constraints of discriminative clinical supervision with the smoothness requirements of report generation. To address these problems, we analyze the failure mechanism of linear scalarization from the perspective of gradient dynamics, utilizing the stochastic differential equation (SDE) framework to characterize it as a "Double Dilemma" of drift term deviation and diffusion term decay. Based on this, we propose a backbone‑agnostic optimizer named Conflict‑Averse Magnitude‑Enhanced Gradient Descent (CAME‑Grad). Through conflict‑averse direction rectification and magnitude‑enhanced energy injection, the algorithm not only ensures geometric validity, but also avoids local optimal solutions. Then, the adaptive gradient fusion mechanism is used to establish a dynamic balance between the theoretical optimal direction and the task‑specific inductive bias. Experiments show that as a universal plug‑and‑play optimizer, CAME‑Grad brings substantial and consistent improvements across eight diverse RRG methods, elevating overall clinical efficacy performance by an average of 2.3% on MIMIC‑CXR and 1.9% on IU X‑Ray. Our code is available at https://github.com/vpsg‑research/CAME‑Grad.
Authors:Junhyeong Cho, Ruojin Cai, Hadar Averbuch-Elor
Abstract:
Many public buildings provide floorplans with a "you are here" indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observations were captured within a floorplan. However, existing methods typically assume controlled small‑scale environments and precise vectorized floorplans, limiting their ability to operate in large‑scale buildings and rasterized floorplans. In this work, we present an approach for performing floorplan localization in the wild by grounding the task in a reconstructed 3D representation of the scene. Given an unconstrained image collection, our method reconstructs a gravity‑aligned 3D scene and projects it into a 2D density map that serves as a floorplan proxy. Floorplan localization is then formulated as aligning this proxy with the input floorplan via a 2D similarity transform. To bridge the appearance gap between density maps and architectural floorplans, we adapt a 2D foundation model to learn cross‑modal correspondences, introducing a fine‑tuning scheme that encourages semantically aligned matches while preserving structural consistency. Extensive experiments demonstrate substantial improvements over prior methods, including in extremely sparse settings with as little as a single input image. Our code and data will be publicly available.
Authors:Jinho Park, Youbin Kim, Hogun Park, Eunbyung Park
Abstract:
Spatio‑temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio‑temporal reasoning benchmark datasets primarily rely on static image sets or passively curated video data, which limits the evaluation of fine‑grained reasoning capabilities. In this paper, we introduce VGenST‑Bench, a video benchmark that employs generative models to actively synthesize highly controlled and diverse evaluation scenarios. To construct VGenST‑Bench, we propose a multi‑agent pipeline incorporating a human quality control stage, ensuring the quality of all generated videos and QA pairs. We establish a comprehensive 3x2x2 video taxonomy, encompassing Spatial Scale, Perspective, and Scene Dynamics to span diverse scenarios. Furthermore, we design a hierarchical task suite that decouples low‑level visual perception from high‑level spatio‑temporal reasoning. By shifting the paradigm from passive curation to active synthesis, VGenST‑Bench enables fine‑grained diagnosis of spatio‑temporal understanding in MLLMs.
Authors:Francesco Benedetto, Roberto Basla, Luca Magri, Giacomo Boracchi
Abstract:
Training Deep Neural Networks for tracking individual cells in biomedical videos requires a large amount of annotated data. The annotation of videos for cell tracking is very time consuming and often requires domain expertise; this explains the limited availability of public annotated data to address important medical problems like tissue repair or cancer treatment. Generating synthetic videos along with their Ground Truth annotations is a promising solution that relies, as a foundational first step, on the synthesis of single cell annotations (or phantoms). Phantoms need to be time consistent, as they have to replicate biological processes that are specific to the cell types. In this work, we propose a novel framework for generating videos of cell phantoms in the Elliptical Fourier Descriptors (EFDs) domain, a compact and geometrically interpretable representation for 2D closed contours. We represent the cell phantom evolution as a multivariate time series of EFD coefficients, introducing a strong prior for cell morphology and enabling the efficient generation of sequences that evolve coherently in time. Our experimental validation proves that modelling the temporal evolution in EFD space enables the generation of biologically plausible phantom videos. Our method can be used in generative pipelines for synthesizing annotated data for cell tracking, thus strongly mitigating the annotation effort for creating new datasets. Our code is available for download here: https://github.com/FrancescoBenedetto99/efd‑cell‑video‑gen.
Authors:Deshui Miao, Xingsen Huang, Yameng Gu, Xin Li, Haijun Zhang, Ming-Hsuan Yang
Abstract:
Spatio‑temporal reasoning in vision‑language models requires visual representations that preserve physical geometry rather than merely semantic appearance. Recent multimodal models incorporate geometric information through structural branches, 3D‑aware supervision, reasoning‑stage fusion, or long‑horizon memory. While these approaches demonstrate the importance of geometry for spatial intelligence, they typically treat geometric cues as a shared signal across all visual tokens. We note that this overlooks a finer‑grained challenge: different visual tokens require different geometric evidence depending on their spatial roles. To address this limitation, we introduce GeoWeaver, a pre‑reasoning geometric grounding framework that treats geometry as a representational prerequisite for spatio‑temporal reasoning. GeoWeaver constructs a multi‑level geometry bank from a frozen geometry encoder and performs token‑adaptive geometric evidence allocation, enabling each visual token to retrieve the most relevant geometric abstractions. The selected evidence is incorporated into visual tokens via a residual grounding operation prior to language modeling, yielding geometry‑grounded representations for downstream reasoning. Extensive evaluations on spatial reasoning benchmarks demonstrate that GeoWeaver consistently enhances geometry‑aware reasoning while retaining general multimodal capabilities. This indicates that geometric information yields the greatest benefit not as a late‑fusion auxiliary signal but as a fundamental prerequisite that shapes the representational foundation on which large language models perform reasoning. All source code and models will be released at https://github.com/yahooo‑m/GeoWeaver .
Authors:Haokun Wen, Xuemeng Song, Xinghao Xie, Xiaolin Chen, Xiangyu Zhao, Weili Guan
Abstract:
Fashion image retrieval is a cornerstone of modern e‑commerce systems. A unified framework that supports diverse query formats and search intentions is highly desired in practice. However, existing approaches focus on narrow retrieval tasks and do not fully capture such diversity. Therefore, in this work, we aim to develop a unified framework capable of handling diverse realistic fashion retrieval scenarios, achieving truly versatile fashion image retrieval. To establish a data foundation, we first introduce U‑FIRE, a comprehensive benchmark that consolidates fragmented fashion datasets into a unified collection, supplemented by two manually curated datasets for testing generalization. Building upon this, we propose FashionLens, a unified framework based on Multimodal Large Language Models. To handle divergent matching objectives, we design a Proposal‑Guided Spherical Query Calibrator that dynamically shifts query representations into task‑aligned metric spaces via adaptive spherical linear interpolation. Additionally, to mitigate the optimization imbalance caused by varying task complexities and data scales, we develop a Gradient‑Guided Adaptive Sampling strategy that automatically re‑weights tasks based on realtime learning difficulty and the data scale prior. Experiments on U‑FIRE show that FashionLens achieves state‑of‑the‑art performance across diverse retrieval scenarios and generalizes robustly to unseen tasks. The data and code are publicly released at https://github.com/haokunwen/FashionLens.
Authors:Deyi Zhu, Yuji Wang, Yong Liu, Yansong Tang, Bingyao Yu, Jiwen Lu, Jie Zhou
Abstract:
Traditional visual object tracking (VOT) methods typically rely on task‑specific supervised training, limiting their generalization to unseen objects and challenging scenarios with distractors, occlusion, and nonlinear motion. Recent vision foundation models, exemplified by SAM 2, learn strong video understanding priors from large‑scale pretraining and offer a promising foundation for building more robust and generalizable trackers. However, directly applying SAM 2 to VOT remains suboptimal, as it does not explicitly model target motion dynamics or enforce geometric and semantic consistency across frames, both of which are essential for reliable tracking. To address this issue, we propose SAMOSA, a new tracking framework that adapts SAM 2 to complex VOT scenarios by explicitly leveraging motion, geometry, and semantic cues. Specifically, we introduce a lightweight nonlinear motion predictor to model target dynamics and guide mask selection as well as memory filtering. We further exploit semantic cues to detect target shifts and recover from tracking failures, while geometric cues are incorporated as structural constraints to improve tracking stability. In this way, SAMOSA bridges the gap between the implicit video understanding prior of SAM 2 and explicit tracking‑oriented modeling. Extensive experiments show that SAMOSA consistently outperforms state‑of‑the‑art SAM 2‑‑based approaches on general benchmarks, demonstrates stronger generalization than supervised VOT methods, and achieves substantial gains on anti‑UAV datasets, which typify complex nonlinear motion scenarios. Our code is available at https://github.com/DurYi/SAMOSA.
Authors:Xiang Ji, Guixu Lin, Zhengwei Yin, Jiancheng Zhao, Yinqiang Zheng
Abstract:
Motion degradation, manifested as blur in global shutter (GS) images or rolling shutter (RS) distortion in RS counterparts, remains a fundamental challenge in computational imaging, especially under fast motion or low‑light conditions. While prior works have treated blur decomposition and RS temporal super‑resolution as separate tasks, this separation fails to exploit their intrinsic complementarity. In this paper, we propose a unified framework to invert motion degradation and reenact imaging moment by jointly leveraging the complementary characteristics of GS blur and RS distortion. To this end, we introduce a novel dual‑shutter setup that captures synchronized blur‑RS image pairs and demonstrate that this combination effectively resolves temporal and spatial ambiguities inherent in both modalities. For allowing flexible performance‑cost trade‑offs, we further extend this dual‑shutter setup to a stereo Blur‑RS configuration with a narrow baseline. In addition, we construct a triaxial imaging system to collect a real‑world dataset with aligned GS‑RS pairs and ground‑truth high‑speed frames, enabling robust training and evaluation beyond synthetic data. Our proposed network explicitly disentangles motion into context‑aware and temporally‑sensitive representations via a dual‑stream motion interpretation module, followed by a self‑prompted frame reconstruction stage. Extensive experiments validate the superiority and generalizability of our approach, establishing a new paradigm for realistic high‑speed video reconstruction under complex motion degradations. Codes and more resources are available at https://jixiang2016.github.io/dualBR_site/.
Authors:Laziz Hamdi, Amine Tamasna, Pascal Boisson, Thierry Paquet
Abstract:
Table structure recognition (TSR) requires both table‑level coherence (row/column counts, headers, spanning cells) and precise separator localization. We introduce FastTab, a grid‑centric TSR model that avoids autoregressive HTML decoding by combining (i) a lightweight Tiny Recursive Module (TRM) for global reasoning and (ii) axial 1D Transformer encoders that capture long‑range dependencies along rows and columns. The model predicts row/column counts, header rows, and separators to construct a grid, then infers rowspan/colspan using ROI‑aligned cell features. Across four benchmarks (PubTabNet, FinTabNet, PubTables‑1M, and SciTSR), FastTab achieves competitive structure recovery performance while operating at low‑latency inference. We further study robustness under pixel‑level anonymisation and show an extension to curved separators for camera‑captured documents. The source code will be made publicly available at https://github.com/hamdilaziz/FastTab .
Authors:Yandi Wang, Libin Zhan, Ziwei Huang, Tiancheng Luo, Yuxuan Jiang, Wang Dong, Leilei Gan, Jun Chen
Abstract:
Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multimodal Large Language Models (MLLMs) have shown promising capabilities, existing benchmarks suffer from critical limitations in scale and realism, lack semantic granularity, and fail to cover diverse document types. To bridge this gap, we introduce ReceiptBench, a large‑scale, human‑annotated benchmark consisting of 10k diverse receipts, organizing information extraction into four hierarchical sub‑tasks: (1) Basic Perception for raw text spotting, (2) Format Normalization for strictly following standardization instructions, (3) Semantic Reasoning for inferring implicit attributes from context, and (4) Structure Parsing for handling nested line items. Furthermore, we propose a two‑stage training framework incorporating Metric‑Aware Group Relative Policy Optimization (GRPO), which translates rigorous evaluation constraints into reinforcement learning signals to enhance structural consistency. Extensive experiments demonstrate that our method yields state‑of‑the‑art performance, surpassing leading proprietary models on complex reasoning tasks. We release our datasets and code at https://github.com/wwwT0ri/ReceiptBench.
Authors:Corentin Dumery, David Colmenares, Alexander Fix, Pascal Fua, Ali Behrooz, Jogendra Kundu
Abstract:
Eye tracking (ET) is a foundational technology for advanced AR/VR applications. However, training ET models for every new ET device is challenging: real data collection is costly and time‑consuming, while existing synthetic data generation methods lack realism. To remove the need for additional data collection while maintaining data quality, we introduce a data‑driven 3D prior that models the distribution of human eyes across diverse identities, gaze directions, and light settings. This model, which we coin GazePrior, then enables sparse‑input 3D reconstruction of annotated data collected with previous ET devices, which can in turn be rendered from the cameras of any target ET device. Our approach synthesizes data with the realism, diversity and ground‑truth accuracy of real data collection without its prohibitive costs. Our experiments demonstrate that ET models trained with our synthesized data outperform previous zero‑shot methods, achieving higher accuracy and robustness.
Authors:Jose Edgar Hernandez Cancino Estrada, Mauro Díaz Lupone, Žiga Emeršič, Vitomir Štruc, Peter Peer, Darian Tomašević
Abstract:
Identity‑conditioned diffusion models enable high‑quality and identity‑consistent face generation, but they also raise severe privacy concerns, as models may continue to synthesize individuals despite their right to be forgotten. While machine unlearning has been extensively studied for concept and data removal, identity unlearning remains largely unexplored, particularly in models conditioned directly on identity embeddings rather than text prompts. In this work, we study identity unlearning in Arc2Face, a state‑of‑the‑art identity‑conditioned latent diffusion model for face generation, and introduce Proximity‑guided Identity Unlearning (PIU), an anchor‑guided framework for identity unlearning. Specifically, we formulate identity removal as an identity replacement objective that reassigns the source identity to a selected anchor identity in the learned identity space, and we complement it with a proximity‑based anchor selection strategy motivated by the geometry of ArcFace representations. We further show that effective unlearning can be achieved through localized fine‑tuning of a small subset of identity‑sensitive cross‑attention layers. Experiments across many target identities show that our framework effectively suppresses generation of the target identity while preserving realism and identity consistency for retained identities, as validated by improved performance on unlearning and image‑quality metrics, together with qualitative evaluation. The source code for the PIU framework is publicly available at https://github.com/edgarcancinoe/piu_unlearning .
Authors:Junbin Xiao, Jiajun Chen, Tianxiang Sun, Xun Yang, Angela Yao
Abstract:
Long streaming video QA remains challenging due to growing visual tokens and limited reasoning length of large language models (LLMs). KV‑caching stores the Key‑Value (KV) of the historical tokens via LLM prefill and enables more efficient streaming QA. However, existing methods cache every one or two frames, causing redundant memory usage and losing fine‑grained spatial details within frame or temporal contexts across frames. This paper proposes MuKV, a method that features a multi‑grained KV cache compression module and a semi‑hierarchical retrieval approach to improve both efficiency and accuracy for long streaming VideoQA. For the offline KV cache, MuKV extracts visual representations at patch‑, frame‑, and segment‑levels. The multiple levels of granularity preserve both local cues and global temporal context, while maintaining efficiency with a dual signal token compression mechanism guided by self‑attention and frequency. For online QA, MuKV designs a semi‑hierarchical retrieval method to retrieve relevant KV caches for answer generation. Experiments on long‑streaming VideoQA benchmarks show that MuKV significantly improves answer accuracy, without sacrificing memory and online QA efficiency. Moreover, our compression mechanism alone brings consistent benefits across answer accuracy, memory, and QA efficiency over baselines, showcasing highly effective contribution.
Authors:Jinming Chai, Libo Yan, Licheng Jiao, Fang Liu
Abstract:
This report presents our solution for the WeatherProof Dataset Challenge, namely CVPR 2026 8th UG2+ Challenge Track 2: Semantic Segmentation in Adverse Weather. For the semantic segmentation task under adverse weather conditions, we propose a semi‑supervised segmentation pipeline. Our method is trained exclusively on the WeatherProof dataset, without using any additional external data. Specifically, we adopt UniMatch V2 as the baseline model and treat all degraded‑weather images as unlabeled data for semi‑supervised training, thereby fully exploiting the data distribution provided by the challenge. During inference, we further apply test‑time augmentation to improve the robustness and segmentation accuracy of the final predictions. The code is publicly available at: https://github.com/ylb888/weatherproof‑challenge‑unimatchv2.
Authors:Matteo Balice, Yanik Kunzi, Chenyangguang Zhang, Matteo Matteucci, Marc Pollefeys, Sungwhan Hong
Abstract:
Recent feed‑forward 3D gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method jointly addresses dynamic content, multi‑view input, and unknown camera poses in a single feed‑forward pass. Methods that handle dynamics either require accurate camera poses or accept only monocular input; pose‑free multi‑view methods address only static scenes; and per‑scene optimization methods bridge some of these gaps but at minutes‑to‑hours cost per scene. We introduce NoPo4D, the first feed‑forward system that addresses this empty quadrant. Building on a pretrained geometry backbone and recent 4D Gaussian frameworks, NoPo4D introduces a velocity decomposition that splits Gaussian motion into per‑pixel image‑plane shifts and depth changes, allowing direct supervision from pseudo ground‑truth optical flow on the 2D component. This sidesteps both the differentiable rendering that couples prior posed methods to pose accuracy and the 3D motion ground truth that prior pose‑free methods require. The system is rounded out by a bidirectional motion encoder for cross‑view and cross‑frame feature aggregation, and view‑dependent opacity that mitigates cross‑view and cross‑timestep Gaussian misalignments. On four multi‑view dynamic benchmarks, NoPo4D consistently outperforms prior feed‑forward baselines, and with an optional post‑optimization stage surpasses per‑scene optimization methods, while running orders of magnitude faster.
Authors:Senyan Xu, Zhijing Sun, Kean Liu, Xin Lu, Ruixuan Jiang, Mingyang Huang, Xueyang Fu, Zheng-Jun Zha
Abstract:
Event‑based low‑light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking the essential global illumination in images and the inherent noise sensitivity of event signals in real‑world scenarios. To address these issues, we propose EIC‑LIE, an event‑illumination collaborative LIE framework. Concretely, we first design an Event‑Illumination Collaborative Interaction (EICI) module, which contains two key processes: forward gathering, which gathers HDR features across varying lighting conditions, and backward injection, which provides complementary content for illumination and event representations. Next, we introduce an Illumination‑aware Event Filter (IAEF) that dynamically reduces event noise based on brightness statistics derived from images. Additionally, we build a beam‑splitter‑based hybrid imaging system to collect high‑quality event‑image pairs with temporal synchronization from dynamic scenes, providing the first high‑resolution, real‑world event‑based LIE dataset. Extensive experiments show that our EIC‑LIE outperforms state‑of‑the‑art methods on five real‑world and synthetic datasets, significantly surpassing previous methods with improvements of up to 1.24dB in PSNR and 0.069 in SSIM. The code and dataset are released at https://github.com/QUEAHREN/EIC‑LIE.
Authors:Vipul Arya, S. H. Shabbeer Basha, Srikrishna U N, Sunainha Vijay, Snehasis Mukherjee
Abstract:
Deep learning models, including Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have achieved state‑of‑the‑art performance on various computer vision tasks such as object classification, detection, segmentation, generation, and many more. However, these models are data‑hungry as they require more training data to learn millions or billions of parameters. Especially for supervised learning tasks, curating a large number of labeled samples for model training is an expensive and time‑consuming task. Active Learning (AL) has been used to address this problem for many years. Existing active learning methods aim at choosing the samples for annotation from a pool of unlabeled samples that are either diverse or uncertain. Choosing such samples may hinder the model's performance as we pool based on one dimension, i.e., either diverse or uncertain. In this paper, we propose four novel hybrid sampling methods for pooling both easy and hard samples, which are also diverse. To verify the efficacy of the proposed methods, extensive experiments are conducted using high and low‑confidence samples separately. We observe from our experiments that the proposed hybrid sampling method, Least Confident and Diverse (LCD), consistently performs better compared to state‑of‑the‑art methods. It is observed that selecting uncertain and diverse instances helps the model learn more distinct features. The codes related to this study will be available at https://github.com/XXX/LCD.
Authors:Bingjun Luo, Tony Wang, Chaoqi Chen, Xinpeng Ding
Abstract:
Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual tokens required. To improve efficiency, existing methods primarily reduce redundancy by pruning or merging tokens based on importance or similarity. However, these approaches largely overlook a critical dimension of video content, i.e., changes and turning points, and they lack a collaborative model for spatio‑temporal relationships. To address this, we propose a new perspective: similarity is for identifying redundancy, while difference is for capturing key events. Based on this, we designed a training‑free framework named ST‑SimDiff. We first construct a spatio‑temporal graph from the visual tokens to uniformly model their complex associations. Subsequently, we employ a parallel dual‑selection strategy: 1) similarity‑based selection uses community detection to retain representative tokens, compressing static information; 2) temporal difference‑based selection precisely locates content‑changing points to preserve tokens that capture key dynamic shifts. This allows it to preserve both static and dynamic content with a minimal number of tokens. Extensive experiments show our method significantly outperforms state‑of‑the‑art approaches while substantially reducing computational costs. Our code is available in https://github.com/bingjunluo/ST‑SimDiff.
Authors:Senyan Xu, Shuai Chen, Chuanfu Shen, Kean Liu, Zhijing Sun, Chengzhi Cao, Xueyang Fu
Abstract:
Gait recognition enables non‑intrusive, privacy‑preserving identification but suffers in uncontrolled environments due to illumination and motion sensitivity of conventional cameras. In this work, we explore gait recognition using event cameras, which offer microsecond temporal resolution and high dynamic range, naturally capturing robust dynamic cues and suppressing static noise. Existing event‑based approaches typically aggregate event streams into event images over long time windows, thereby discarding fine‑grained motion dynamics critical for gait recognition. Therefore, we propose EventGait, an end‑to‑end dual‑stream framework that separately models motion and shape while preserving the advantages of events. Our dynamic stream leverages a Mixture of Spiking Experts (MoSE) with diverse neuron constants for robust dynamic perception across complex motion and illumination scenes, while the static stream learns dense shape representations via Cross‑modal Structure Alignment (CroSA) with large vision foundation models. To address the absence of large‑scale event‑based gait datasets, we introduce a synthesis pipeline and release two new benchmarks: SUSTech1K‑E and CCGR‑Mini‑E. Extensive experiments have shown that event‑based gait recognition not only achieves results comparable to camera‑based gait recognition under normal conditions but also significantly outperforms it in low‑light scenarios. Our approach sets a new state of the art on both synthesized and real‑world event‑based gait benchmarks, highlighting the robustness and potential of event‑driven gait analysis. The code and datasets are released at https://github.com/QUEAHREN/EventGait.
Authors:Tianxiang Du, Hulingxiao He, Yuxin Peng
Abstract:
In everyday photography, aesthetically appealing moments are often captured with structural flaws (e.g., composition, camera viewpoint, or pose) that existing retouching and portrait enhancement methods cannot fix. We formulate Aesthetic Photo Reconstruction (APR) as improving a photo's aesthetic quality via structural reconstruction while preserving subject identity and scene semantics. Although recent advances in image editing models make APR feasible, they often lack aesthetic understanding, yielding edits that are semantically plausible yet aesthetically weak. To address this, we propose AesFormer, a two‑stage framework that decouples aesthetic planning from image editing. In Stage 1, an aesthetic action model (AesThinker) analyzes the input along seven progressive photographic dimensions and outputs executable editing actions; we further apply GRPO‑A to encourage broad exploration over diverse action plans beyond SFT. In Stage 2, an action‑conditioned editor (AesEditor) performs structural edits guided by these actions. To support APR, we build a video‑based corpus‑mining pipeline (VCMP) and construct AesRecon, a benchmark of 9,071 strictly aligned (poor, good) image pairs. Experiments show that AesFormer substantially improves APR performance and is competitive with Nano Banana Pro. Code is available at https://github.com/PKU‑ICST‑MIPL/AesFormer_ICML2026.
Authors:Zhiqing Hong, Zelong Li, Xiubin Fan, Guang Yang, Baoshen Guo, Haotian Wang, Tian He, Desheng Zhang
Abstract:
Human Activity Recognition (HAR) has shown remarkable effectiveness in various applications, such as smart healthcare and intelligent manufacturing. However, a major challenge faced by HAR is the distribution shift across different sensor data domains, which often leads to decreased performance when deployed for real‑world applications. To address this issue, this paper introduces GenHAR, a novel framework designed to mitigate the domain gap by learning domain‑invariant sensor representations. GenHAR aims to enhance the generalization capabilities of HAR on target domains purely with data from the source domain. The key novelty of GenHAR lies in two aspects. Firstly, GenHAR tokenizes sensor data and learns correlations among frequency sensor channel dimensions to improve the robustness of HAR models. Secondly, GenHAR improves the efficiency via selective masking and an efficient attention mechanism. We conduct a systematic analysis of GenHAR by comparing it with state‑of‑the‑art HAR methods on real‑world human activity datasets. Results show that GenHAR outperforms state‑of‑the‑art methods by 9.97% in accuracy, and reduces Floating Point Operations by 6.4 times. Moreover, we deploy GenHAR at a leading logistics company in 4 cities, and have detected 2.15 billion real‑time activities. We release our code at: https://github.com/Sensor‑FoundationModel/GenHAR.
Authors:Bingjun Luo, Tony Wang, Hanqi Chen, Xinpeng Ding
Abstract:
Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficiently compressing visual tokens while preserving spatiotemporal interactions. Existing methods, such as LLaVA family, utilize simplistic pooling or interpolation techniques that overlook the intricate dynamics of visual tokens. To bridge this gap, we propose ST‑GridPool, a novel training‑free visual token enhancement method designed specifically for Video LLMs. Our approach integrates Pyramid Temporal Gridding (PTG), which captures multi‑grained spatiotemporal interactions through hierarchical temporal gridding, and Norm‑based Spatial Pooling (NSP), which preserves high‑information visual regions by leveraging the correlation between token norms and semantic richness. Extensive experiments on various benchmarks demonstrate that ST‑GridPool consistently enhances performance of Video LLMs without requiring costly retraining. Our method offers an efficient and plug‑and‑play solution for improving visual token representations. Our code is available in https://github.com/bingjunluo/ST‑GridPool.
Authors:Hyeseong Kim, Geonhui Son, Deukhee Lee, Dosik Hwang
Abstract:
Novel view synthesis from sparse‑view inputs poses a significant challenge in 3D computer vision, particularly for achieving high‑quality scene reconstructions with limited viewpoints. We introduce TWINGS, a framework that enhances 3D Gaussian Splatting (3DGS) by directly addressing point sparsity. We employ Thin Plate Splines (TPS), a smooth non‑rigid deformation model that minimizes bending energy to estimate a globally coherent warp from control‑point correspondences, to align backprojected points from estimated depth with triangulated 3D control points, yielding calibrated backprojected points. By sampling these calibrated points near the control points, TWINGS provides a fast and geometrically accurate initialization for 3DGS, ultimately improving structural detail preservation and color fidelity in reconstructed scenes. Extensive experiments on DTU, LLFF, and Mip‑NeRF360 demonstrate that TWINGS consistently outperforms existing methods, delivering detailed and accurate reconstructions under sparse‑view scenarios.
Authors:Junhyub Lee, Seunghun Chae, Hyosu Kim
Abstract:
We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained granularity and flexibility. Specifically, we provide the foundation benchmark for this new paradigm with the following three key contributions. First, we overcome the prohibitively high cognitive and physical bottlenecks of manual annotation by developing a fully automated generation pipeline that synergizes the semantic reasoning of Large Vision‑Language Models (LVLMs) with the precise geometric grounding of SAM 3. Second, leveraging this pipeline, we construct COCOTree, a massive‑scale benchmark featuring over 21K images and 1.8M structural nodes. By embracing an open‑vocabulary space of over 3.5K unique labels, it successfully captures the long‑tail distribution of complex physical assemblies. Notably, rigorous human evaluation confirms our generated annotations demonstrate strong alignment with human structural judgment. Third, we establish a standardized evaluation protocol by proposing the Open Tree Quality (OTQ) metric, which jointly assesses mask precision, label accuracy, and structural consistency. We release our dataset and benchmark code at https://github.com/melonkick3090/COCOTree.
Authors:Anthony Song, Boyan Zhou, Mayank Golhar, Marisa Morakis, Alex Baras, Nicholas Durr
Abstract:
Three‑dimensional (3D) histopathology of unprocessed tissues has the potential to transform disease management by enabling volumetric characterization of tissue microarchitecture and in‑vivo assessment. Back‑illumination Interference Tomography (BIT) is a new phase microscopy technology that provides rapid, non‑destructive volumetric imaging of unprocessed tissues. However, translating BIT volumes into clinically interpretable H&E images remains challenging, particularly due to shift‑variant contrast and the absence of quantitative validation benchmarks. We introduce HistoBIT3D, the first voxel‑wise paired BIT and fluorescence‑labeled nuclei dataset, enabling quantitative evaluation of structural preservation in unsupervised virtual staining against ground‑truth nuclear distributions. Using this dataset, we present a novel virtual staining framework that translates BIT volumes with shift‑variant contrast into realistic H&E volumes by leveraging bidirectional multiscale content consistency and cross‑domain style reuse to enhance structural fidelity and perceptual realism. Our method achieves state‑of‑the‑art realism metrics while significantly improving 3D nuclei segmentation accuracy and boundary preservation under zero‑shot Cellpose evaluation. Together, these contributions establish a quantitatively validated, structurally faithful, and scalable pipeline for 3D virtual H&E staining, advancing the paradigm of slide‑free, volumetric computational histopathology. Our data and code are available at: https://github.com/aasong113/HistoBIT3D_VirtualStaining.
Authors:Dazhao Du, Jian Liu, Jialong Qin, Tao Han, Bohai Gu, Fangqi Zhu, Yujia Zhang, Eric Liu, Xi Chen, Song Guo
Abstract:
Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single‑frame cues and language priors rather than by tracking spatiotemporal dynamics. This issue is exacerbated in RL post‑training, where correctness‑only rewards can further reinforce shortcut policies that obtain high reward without tracking video dynamics. We address this by asking a controlled counterfactual question: if the visual world changed while the question remained fixed, should the answer change or stay the same? Based on this view, we propose Counterfactual Relational Policy Optimization (CRPO), a dual‑branch RL framework for improving \emphspatiotemporal sensitivity. CRPO constructs counterfactual videos through horizontal flips and temporal reversals, trains on both original and counterfactual branches, and introduces a Counterfactual Relation Reward (CRR) between their answers. CRR encourages answers to change for dynamic questions and remain unchanged for static questions. This cross‑branch constraint makes it difficult for shortcut policies to be consistently rewarded across both branches. To evaluate this property, we introduce DyBench, a paired counterfactual video benchmark with 3,014 videos covering reversible dynamics, moving direction, and event sequence, together with a strict pair‑accuracy metric that prevents fixed‑answer shortcuts from inflating scores. Experiments show that CRPO outperforms prior RL methods on spatiotemporal‑sensitive evaluations while maintaining competitive general video performance. On Qwen3‑VL‑8B, CRPO improves DyBench P‑Acc by +7.7 and TimeBlind I‑Acc by +8.2 over the base model, indicating improved spatiotemporal sensitivity rather than stronger reliance on static shortcuts. The project website can be found at https://ddz16.github.io/crpo.github.io/ .
Authors:Le Zhang, Ning Mang, Aishwarya Agrawal
Abstract:
Flow matching with x‑prediction ‑‑ regressing the clean data point rather than the ambient velocity ‑‑ is known to exploit low‑dimensional manifold structure effectively in pixel space \citeli2025back. We ask whether a pretrained representation space, while containing a low‑dimensional data manifold of comparable intrinsic dimensionality, offers a distribution more favorable for flow‑matching learning. Comparing pixel, SD‑VAE, and DINOv2 features along four geometric axes, we find that pixel and DINOv2 share nearly identical intrinsic dimensionalities (both \hatd\!\approx\!33) yet DINOv2 exhibits 7.3× higher effective rank, 35× better covariance conditioning, 11.5× lower excess kurtosis, and 1.7× lower on‑manifold interpolation error; SD‑VAE latents are consistently intermediate, indicating that the advantage stems from representation‑learning objectives rather than mere compression. These statistical properties render the flow‑matching regression well‑conditioned and remove the need for the specialized prediction heads or Riemannian transport used by prior DINOv2 diffusion methods. We propose the \emphRepresentation Image Transformer (RiT): a vanilla Diffusion Transformer trained by x‑prediction on frozen DINOv2 features, augmented only by a dimension‑aware noise schedule and joint \texttt[CLS]‑patch modeling. On ImageNet 256×256, RiT attains FID 1.45 without guidance and 1.14 with classifier‑free guidance, outperforming DiT^\textDH‑XL with 19% fewer parameters (676M vs.\ 839M). The resulting ODE is efficiently solvable at coarse discretizations: with classifier‑free guidance, 5 Heun steps already reach FID 2.0 and 10 steps reach 1.25, without distillation or consistency training. Code at https://github.com/lezhang7/RiT.
Authors:Dazhao Du, Liao Duan, Jian Liu, Tao Han, Yujia Zhang, Eric Liu, Xi Chen, Song Guo
Abstract:
Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs) understand not only what happens but also when it happens. Although modern MLLMs describe video content fluently, their timestamp predictions remain unreliable, while existing remedies either require costly post‑training on temporal annotations or rely on coarse training‑free heuristics. In this work, we probe the cross‑modal attention of MLLMs and uncover a perception‑generation gap. Our key finding is that MLLMs often know the target interval during prefill, but lose this signal when generating the final answer. In the prefill stage, a sparse set of attention heads, which we call \emphTemporal Grounding Heads (TG‑Heads), concentrates query‑to‑video attention on the ground‑truth interval. During autoregressive decoding, however, the answer tokens shift attention away from this interval toward visually salient but query‑irrelevant segments. This observation motivates an inference‑time read‑then‑regenerate framework. We first convert TG‑Head prefill attention into a debiased frame‑level relevance signal and extract the high‑attention interval it highlights. We then re‑invoke the MLLM with visual context restricted to this interval, using video cropping or attention masking to suppress distractors. Without parameter updates and architectural changes, our framework consistently improves MiMo‑VL‑7B, Qwen3‑VL‑8B, and TimeLens‑8B on three VTG benchmarks, with gains of up to +3.5 mIoU. The project website can be found at https://ddz16.github.io/mllmsknowwhen.github.io/.
Authors:Shiqi Huang, Ziyue Wang, Zhongrong Zuo, Han Qiu, Qi She, Bihan Wen
Abstract:
Recent Video Large Language Models (Video‑LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, existing RL pipelines rely heavily on human‑annotated tasks and solutions, making them costly to scale and fundamentally constrained by human expertise. Self‑evolving frameworks have recently emerged as a promising alternative through autonomous Questioner‑Solver self‑play. Unfortunately, these approaches are primarily designed for static modalities such as text and images, fundamentally failing to capture the temporal dynamics that are central to video reasoning. In this work, we propose EvoVid, a temporal‑centric self‑evolving framework that enables Video‑LLMs to improve directly from raw, unannotated videos. Specifically, we introduce two complementary temporal‑centric rewards: a temporal‑aware Questioner reward that encourages temporally dependent question generation through temporal perturbation sensitivity, and a temporal‑grounded Solver reward that provides automatic temporal supervision via inherent video segment localization. Extensive experiments across four base models and six benchmarks demonstrate consistent improvements over both base models and existing self‑evolving baselines, achieving competitive performance with supervised methods. These results highlight temporal‑centric self‑evolution as an effective and scalable paradigm for video understanding and reasoning.
Authors:Yuheng Li, Yuan Gao, Haoyu Dong, Yuxiang Lai, Shansong Wang, Mojtaba Safari, James E. Baciak, Xiaofeng Yang
Abstract:
Computed tomography (CT) is a central to three‑dimensional medical imaging, yet CT‑based artificial intelligence remains fragmented across task‑specific models for segmentation, classification, registration, and report analysis. Here we present FlexiCT, a family of CT foundation models trained by agglomerative continual pretraining on 266,227 CT volumes from 56 publicly available datasets, forming a large‑scale public resource for CT representation learning. FlexiCT uses agglomerative pretraining across three stages: two‑dimensional axial pretraining, three‑dimensional anatomical pretraining and report‑guided semantic alignment. This training strategy supports slice‑level, volume‑level and vision‑language analysis. Across five downstream task families (segmentation, classification, registration, vision‑language understanding and clinical retrieval), FlexiCT matches or exceeds prior task‑specific approaches on multiple benchmarks. Its embeddings further organize CT scans along gradients associated with various tumor stages, suggesting that CT foundation models can capture imaging features relevant to disease phenotype characterization. Project page and code are available at: https://ricklisz.github.io/flexict.github.io and https://github.com/ricklisz/FlexiCT.
Authors:Rusiru Thushara, Yasiru Ranasinghe, Jay Paranjape, Vishal M. Patel
Abstract:
Vision‑language models (VLMs) often fail under low illumination because their visual grounding is learned predominantly from RGB imagery, whereas thermal infrared preserves complementary scene structure when visible cues degrade. We present Thermo‑VL, a wavelength‑aware VLM that augments a frozen Molmo‑7B backbone with a trainable thermal encoder and a text‑guided dual‑attention fusion module. Given aligned RGB tokens, thermal tokens, and prompt embeddings, the fusion module conditions thermal features on both language and RGB context, then injects a gated residual into the frozen RGB stream so thermal evidence can be incorporated without disrupting Molmo's pretrained RGB‑language interface. We train the model with the standard language‑modeling objective together with auxiliary alignment and regularization losses that improve cross‑modal grounding and reduce over‑reliance on RGB. We also introduce a pixel‑aligned RGB‑thermal instruction‑tuning dataset and Thermo‑VL‑Bench, a manually screened RGB‑thermal VQA benchmark for low‑light and cross‑spectrum reasoning. Experiments show strong gains on challenging thermal‑only and RGB+thermal reasoning tasks, highlighting the value of prompt‑conditioned multispectral fusion. Our dataset and code are publicly available at: https://thusharakart.github.io/Thermo‑VL
Authors:Yuting He, Chenyu You, Shuo Li
Abstract:
Multi‑modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non‑IID feature statistics across heterogeneous imaging modalities. Monolithic self‑supervised optimization on such data induces conflicting gradients, driving representations to collapse toward modality‑dominant shortcuts. This work reframes this failure as an imbalance between specialization and coordination in emergent modularity, and proposes Director‑Experts (DEX), a modular network that explicitly regulates these dynamics in stacked modules. Each DEX module comprises a pool of experts, dynamically adapted by our image‑wise activation strategy, autonomously specializing in modality‑dominant statistics, together with a director, updated via our group exponential moving average, which distills multi‑expert knowledge into a shared space for semantic integration across modalities, thus driving the emergence of modular representations. We curate a new benchmark, Medical Vision Universe, over 4 million images across 10 modalities, which provides a FM‑level pre‑training with the broadest coverage of distinct imaging modalities to our DEX. Extensive evaluations on 26 downstream tasks demonstrate improved optimization behavior and transferability, indicating DEX as a principled step toward general‑purpose multi‑modality medical AI. Our code and dataset will be opened at https://github.com/YutingHe‑list/DEX.
Authors:Zhi Liu
Abstract:
Vision‑Language‑Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete‑token autoregression (e.g. OpenVLA) and continuous‑action flow‑matching (e.g. pi‑0.5). Yet preference alignment via Direct Preference Optimisation (DPO) ‑‑ the de‑facto post‑training step in language models ‑‑ has been studied almost exclusively on autoregressive VLAs. We present CrossVLA, an empirical study of cross‑paradigm VLA post‑training. Three contributions: (i) a surrogate flow‑matching log‑probability estimator that lets DPO operate on continuous‑action backbones without probability‑flow ODE integration; (ii) a head‑to‑head comparison of LoRA and DoRA as the parameter‑efficient layer for VLA DPO, finding DoRA improves over OpenVLA SFT by a mean +10.4 pp across LIBERO 4‑suite (600 trials, 3 seeds) ‑‑ per‑suite +20.0 Object, +11.0 Long‑horizon, +8.0 Goal, +2.7 Spatial ‑‑ with zero seed variance on Object (38/50 on each of 3 seeds); (iii) an inference‑time anatomy showing the denoise loop dominates 78.6% of sample_actions latency and prefix‑K/V caching a la VLA‑Cache caps at a 21% acceleration ceiling ‑‑ both chunk‑level and token‑level cache strategies degrade success rate to 0‑80% in our benchmarks. We further pretrain a multi‑view + temporal projection head on 6000 LIBERO frames, achieving 99.5% k‑NN recall@1 for same‑task retrieval (36x over random), available as a downstream initialisation. All code, ckpts, training logs, and reproduction scripts are open at https://github.com/lz‑googlefycy/vla‑lab.
Authors:Xiaofeng Liu, Qianru Zhang, Thibault Marin, Menghua Xia, Chi Liu, Georges El Fakhri, Jinsong Ouyang
Abstract:
The synergistic interpretation of anatomical information from computed tomography (CT) and metabolic information from positron emission tomography (PET) is important to oncologic imaging. However, existing deep learning methods for PET/CT remain largely task‑specific, are often trained on single‑center cohorts, or adopt dual‑branch fusion schemes that delay cross‑modal interaction and underutilize early spatial correspondence between PET and CT. To address these limitations, we present an open‑source, multi‑center, whole‑body FDG PET/CT foundation model utilizing 4,997 harmonized scans from four public datasets. Our framework employs hierarchical UNet‑shaped backbones with early channel‑wise concatenation, enabling anatomical and metabolic features to interact from the first embedding layer onward. We further introduce a masked autoencoding objective based on zero‑mean imputation, combined with a weighted global reconstruction loss. This design avoids non‑physical intensity discontinuities at masked‑region boundaries that arise from learnable mask tokens. On downstream AutoPET lesion segmentation, the proposed models demonstrate strong label efficiency: with only 10% of the labeled training data, they achieve performance comparable to models trained from scratch on the full dataset. Under extreme 5‑shot linear probing, joint PET/CT pretraining also achieves higher Dice scores than separated‑modality pretraining. This multi‑center foundation model demonstrates label efficiency and cross‑modality representation learning for PET/CT tumor segmentation. It provides a robust, open‑source basis for advancing automated oncologic imaging, significantly reducing the need for large‑scale manual annotations in clinical practice.
Authors:Li Ma, Mingming He, Xueming Yu, David M. George, Ahmet Levent Taşel, Paul Debevec, Julien Philip
Abstract:
Being able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject‑specific video diffusion‑based framework for relighting full‑body human performances in a temporally consistent way. Our model is trained on a hybrid dataset of pixel‑aligned video relighting pairs, covering a diverse combination of lighting conditions, performances and viewpoints. To acquire such dataset, we combine traditional static One‑Light‑at‑a‑Time (OLAT) capture and a novel dynamic performance capture in which two smoothly varying lighting sequences are rapidly interleaved. Because the lighting operates above the human flicker‑fusion threshold, the interleaving does not appear to strobe. We train our video relighting model from a pretrained text‑to‑video model to fully leverage the generative priors for producing high quality videos. To achieve accurate lighting control, we introduce a new lighting conditioning method that represents each light source as a token. We further condition on sequences of lighting using masked attention to support dynamic lighting control. Together with a carefully designed data augmentation pipeline, we achieve photorealistic, robust, and temporally consistent video relighting of subject‑specific human performances.
Authors:Ritik Shah, Marco F. Duarte
Abstract:
Hyperspectral super‑resolution (HSR) reconstructs a high‑spatial‑resolution hyperspectral image by fusing a low‑resolution hyperspectral image (LR‑HSI) with a high‑resolution multispectral image (HR‑MSI). In the absence of real‑world paired data, HSR methods are evaluated almost exclusively on synthetic experiments derived from hyperspectral datasets through Wald's protocol. Despite the protocol's widespread adoption, its practical implementation varies markedly across research works, typically relying on a single (usually Gaussian) or very few point spread functions (PSFs), one or two spectral response functions (SRFs), and a couple of spatial downsampling factors. As a result, reported performance figures are difficult to compare across the literature, in addition to being often difficult to reproduce; furthermore, they may not generalize across realistic sensing conditions. We introduce HyperBench, a unified and extensible framework that standardizes synthetic experimentation for HSR. HyperBench supports diverse degradation configurations spanning ten PSFs, four SRFs derived from operational multispectral sensors, configurable spatial downsampling factors, and matched additive white Gaussian noise; its goal is to automate large‑scale evaluation and structured logging. By decoupling model development from experimental design, the framework enables reproducible, apples‑to‑apples cross‑method comparison with minimal friction. We use HyperBench to evaluate six recently proposed HSR methods across a 70‑configuration sweep on four widely used hyperspectral scenes and observe that the inter‑method PSNR spread widens from approximately 5 dB on the easiest PSF to over 13 dB on the hardest ‑ a fragility that is structurally invisible to the prevailing single‑configuration evaluation protocol. HyperBench code is available at https://github.com/ritikgshah/HyperBench .
Authors:Jinghang Li, Tales Santini, Courtney Clark, Bruno de Almeida, Cong Chu, Salem Alkhateeb, Andrea Sajewski, Jacob Berardinelli, Hecheng Jin, Tobias Campos, Jeremy J. Berardo, Joseph Mettenburg, Ariel Gildengers, Howard J. Aizenstein, Minjie Wu, Tamer S. Ibrahim
Abstract:
Hippocampal subfield segmentation requires high‑resolution T2w turbo spin echo (TSE) MRI, yet this sequence is susceptible to motion artifacts, leading to substantial data loss. We developed a conditional generative model (MRecover) that synthesizes routinely acquired T1w images to create TSE images with autoregressive slice conditioning for volumetric consistency. Trained on 7T MRI data (n=577), the model achieved high in‑domain fidelity (n=148, SSIM=0.84, FSIM=0.94) and generalized well to out‑of‑domain 3T data: subfield volumes from synthesized and the as‑acquired images closely matched: (n=416, r=0.87‑0.97) and yielded 31.8% more analyzable subjects in the motion‑affected ADNI3 dataset after quality control (593 vs 450). The synthesized images also achieved larger effect sizes due to increasing the sample size for diagnostic group differences in hippocampal subfield atrophy (whole hippocampus ε^2= 0.121‑0.100 vs. 0.086‑0.062, left‑right hemispheres). Project page: https://jinghangli98.github.io/MRecover/
Authors:Sixiang Chen, Zhaohu Xing, Tian Ye, Xinyu Geng, Yunlong Lin, Jianyu Lai, Xuanhua He, Fuxiang Zhai, Jialin Gao, Lei Zhu
Abstract:
Open‑ended image generation is no longer a simple prompt‑to‑image problem. High‑quality generation often requires an agent to combine a model's internal generative ability with external resources. As requests become more diverse and demanding, we aim to develop a general image‑generation agent that can self‑evolve through trajectories and use tools more effectively across varied generation challenges. To this end, we propose GenEvolve, a self‑evolving framework based on Tool‑Orchestrated Visual Experience Distillation. In GenEvolve, each generation attempt is modeled as a tool‑orchestrated trajectory, where the agent gathers evidence, selects references, invokes generation skills, and composes them into a prompt‑reference program. Unlike existing agentic generation methods that mainly rely on image‑level scalar rewards, GenEvolve compares multiple trajectories for the same request and abstracts best‑worst differences into structured visual experience, provided only to a privileged teacher branch. Inspired by on‑policy self‑distillation, Visual Experience Distillation provides dense token‑level supervision, helping the student internalize better search, knowledge activation, reference selection, and prompt construction. We further construct GenEvolve‑Data and GenEvolve‑Bench. Experiments on public benchmarks and GenEvolve‑Bench show substantial gains over strong baselines, achieving state‑of‑the‑art performance among current image‑generation frameworks. Our website is as follows: https://ephemeral182.github.io/GenEvolve/
Authors:Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, Chong Luo, Jianmin Bao, Ji Li, Lei Shi, Qinhong Yang, Xiuyu Wu, Xuelu Feng, Yan Lu, Yanchen Dong, Yitong Wang, Yunuo Chen
Abstract:
We introduce Lens, a 3.8B‑parameter T2I model that achieves performance competitive with, and in several cases surpassing, state‑of‑the‑art models with more than 6B parameters across various benchmarks, while requiring significantly less training compute. For example, Lens requires only about 19.3% of the training compute used by Z‑Image. The training efficiency of Lens stems from two key strategies beyond its compact model size. First, we maximize data information density per training batch by (i) training on Lens‑800M, a dataset of 800M densely captioned image‑text pairs whose captions are generated by GPT‑4.1 and contain approximately 109 words on average, providing richer semantic supervision than conventional short captions, and (ii) constructing each batch from images with multiple resolutions and diverse aspect ratios, thereby enlarging the effective visual coverage of each optimization step. Second, we improve convergence speed through careful architectural choices, including adopting a semantic VAE that provides better latent representations and employing a strong language encoder that accelerates optimization while enabling multilingual generalization from English‑only training data. After pre‑training, we apply RL with taxonomy‑driven prompts (Lens‑RL‑8K) and structured reward rubrics to suppress artifacts and improve visual quality, a reasoner module with training‑free system prompt search to better align user requests with the model, and distillation‑based acceleration for 4‑step inference. Through efficient training and systematic optimization, Lens generalizes to arbitrary aspect ratios from 1:2 to 2:1 and resolutions up to 1440^2, and supports prompts in several commonly used languages. Thanks to its compact size, Lens generates a 1024^2 image in 3.15 seconds on a single NVIDIA H100 GPU, while its distilled turbo version performs 4‑step generation in 0.84 seconds.
Authors:Rony Abecidan, Vincent Itier, Jérémie Boulanger, Patrick Bas, Tomáš Pevný
Abstract:
Steganalysis models excel on benchmark datasets but struggle in the wild when analyzed images are produced by a processing pipeline unseen during training. This problem known as Cover Source Mismatch (CSM) is particularly hard in realistic settings where practitioners (1) have access to only a small, unlabeled dataset, (2) are unsure of the processing techniques applied to these images, and (3) lack information on the proportion of covers and stegos in that set. To answer this challenge, we introduce TADA (Target Alignment through Data Adaptation), a framework learning to emulate the unknown processing pipeline from a small unlabeled target set. This architecture is trained with a loss combining residual covariance alignment, residual distribution matching, and a \ell^2 loss constraining the emulator to produce realistic images. Across toy and operational targets, TADA yields substantial gains in robustness to CSM and improves operational generalization compared to strong holistic and atomistic baselines. Additional resources are available at this link: https://github.com/RonyAbecidan/TADA
Authors:Dian Zheng, Manyuan Zhang, Hongyu Li, Hongbo Liu, Kai Zou, Kaituo Feng, Hongsheng Li
Abstract:
Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi‑task training. Due to inherent task conflicts, such strategy requires complex multi‑stage pipelines, massive data mixing, and balancing tricks, merely resulting in a performance trade‑off rather than true mutual reinforcement. To break this paradigm, we propose Uni‑Edit, an intelligent image editing task that serves as the first general task for UMM tuning. Unlike complex mixed pipelines, Uni‑Edit improves performance across all three abilities at once using only one task, one training stage, and one dataset. Specifically, we first identify image editing as an inherently ideal general task, as it naturally demands both visual understanding and generation. However, existing editing data relies on simplistic instructions that severely underutilize a model's understanding capacity. To address this, we introduce the first automated and scalable data synthesis pipeline for intelligent editing, transforming diverse VQA data into complex and effective editing instructions with embedded questions and nested logic. This yields Uni‑Edit‑148k, pairing diverse reasoning‑intensive instructions with high‑quality edited images. Extensive experiments on BAGEL and Janus‑Pro demonstrate that tuning solely on Uni‑Edit achieves comprehensive enhancements across all three capabilities without any auxiliary operations.
Authors:Guanlong Jiao, Chenyangguang Zhang, Jia Jun Cheng Xian, Zewei Zhang, Renjie Liao
Abstract:
Although existing video editing methods are generally feasible, they often require many costly iterations and still struggle to deliver high‑quality yet satisfying editing results. We attribute this limitation to the prevalent data‑to‑data paradigm, which is less compatible with modern generative models than noise‑to‑data generation. To address this gap, we revisit video editing from a noise‑to‑data perspective and propose Streaming‑Generation‑based Video Editing (StreamGVE), which preserves few‑step sampling while seamlessly injecting source‑video conditions. Built on pre‑trained streaming generation models, StreamGVE introduces dual‑branch fast sampling with a self‑attention bridge and cross‑attention grounding/boosting to satisfy both sampling and conditioning requirements. We further propose source‑oriented guidance to improve target‑generation quality, and a visual prompting strategy to enhance editing flexibility and practicality. The method is effective, robust, and generalizable across different models. Extensive experiments on diverse video editing tasks show that StreamGVE consistently outperforms existing approaches, even in few‑step settings with minimal time cost.
Authors:Amaya Gallagher-Syed, Costantino Pitzalis, Myles J. Lewis, Michael R. Barnes, Gregory Slabaugh
Abstract:
We introduce ProtoPathway, an interpretable‑by‑design multimodal framework for cancer survival prediction that unifies whole slide imaging and transcriptomics through encoders producing biologically grounded representations on both sides of the fusion. On the histopathology side, K learnable morphological prototypes, trained end‑to‑end with the survival objective, serve as the slide representation itself: patches flow into prototype tokens via soft assignment, compressing variable‑length patch sets into fixed task‑adaptive tokens. On the genomic side, a bipartite graph neural network encodes gene expression within the Reactome pathway hierarchy, producing pathway embeddings that reflect both constituent genes and their broader biological context through bidirectional message passing over a shared gene‑‑pathway graph. Cross‑modal attention then operates over a compact prototype × pathway matrix in which prototypes query pathways, modeling the biological direction in which molecular programs give rise to tissue morphology. Because both axes carry stable task‑learned identity, the attention matrix is itself an interpretability output, yielding native inference‑time attribution across the full biological hierarchy, from genes through pathways and prototypes to spatial tissue maps. We evaluate on five TCGA cancer cohorts, demonstrating competitive or superior survival prediction with substantially improved biological interpretability and reduced computational cost, with interpretability claims validated through fold‑stratified rank‑based population‑level analysis. Our source code, model weights, and Reactome pathways, together with a unified codebase reimplementing all multimodal survival baselines under identical preprocessing and evaluation, are available at: https://github.com/AmayaGS/ProtoPathway.
Authors:Jun Zheng, Zhengze Xu, Mengting Chen, Jing Wang, Jinsong Lan, Xiaoyong Zhu, Kaifu Zhang, Bo Zheng, Xiaodan Liang
Abstract:
Video Virtual Try‑On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to non‑interactive scenarios where models merely showcase garments. This limitation overlooks a crucial aspect of real‑world apparel presentation: active human‑garment interaction. To bridge this gap, we introduce and formalize a new challenging task: Interactive Video Virtual Try‑On (Interactive VVT), where subjects in the video actively engage with their clothing. This task introduces unique challenges beyond simple texture preservation, including: (1) resolving the semantic ambiguity of interactions from standard pose information, and (2) learning complex garment deformations from video where interactive moments are sparse and brief. To address these challenges, we propose iTryOn, a novel framework built upon a large‑scale video diffusion Transformer. iTryOn pioneers a multi‑level interaction injection mechanism to guide the generation of complex dynamics. At the spatial level, we introduce a garment‑agnostic 3D hand prior to provide fine‑grained guidance for precise hand‑garment contact, effectively resolving spatial ambiguity. At the semantic level, iTryOn leverages global captions for overall context and time‑stamped action captions for localized interactions, synchronized via our novel Action‑aware Rotational Position Embedding (A‑RoPE). Extensive experiments demonstrate that iTryOn not only achieves state‑of‑the‑art performance on traditional VVT benchmarks but also establishes a commanding lead in the new interactive setting, marking a significant step towards more dynamic and controllable virtual try‑on experiences.
Authors:Shizhe Chen, Paul Pacaud, Cordelia Schmid
Abstract:
Vision‑Language‑Action (VLA) models have shown strong potential for general‑purpose robotic manipulation by leveraging large pretrained vision‑language backbones. However, most existing VLAs rely primarily on 2D visual representations, which limit their ability to reason about fine‑grained geometry and spatial grounding ‑ capabilities that are essential for precise and robust manipulation in 3D environments. In this paper, we propose PointACT, a dual‑system 3D‑aware VLA policy that integrates hierarchical 3D point cloud representations directly into the action decoding process. PointACT employs a multi‑scale point‑action interaction mechanism with efficient bottleneck window self‑attention, enabling evolving action tokens to densely attend to both local geometric detail and global scene structure. We evaluate PointACT on the LIBERO and RLBench benchmarks and systematically compare it against monolithic and dual‑system VLA baselines, including variants augmented with point cloud inputs. PointACT achieves consistent improvements across both benchmarks, increasing success rates by 10% on the challenging RLBench‑10Tasks suite over state‑of‑the‑art pretrained VLAs, with even larger gains when the vision‑language backbone is frozen and the action expert is trained from scratch. Extensive ablation studies demonstrate that tightly coupling hierarchical 3D geometry with pretrained 2D semantic representations is critical for robust and spatially grounded robot control. Our results also highlight the promise of pretrained 3D representations for 3D‑aware VLA policies.
Authors:Abhishek Dinkar Jagtap, Sanath Tiptur Sadashivaiah, Andreas Festag
Abstract:
Cooperative perception enabled by Vehicle‑to‑Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental representation through shared sensory data. While recent works have advanced multi‑agent fusion for improved perception, uncertainty quantification in such cooperative frameworks remains largely unexplored. This paper introduces Hyper‑V2X, a hypernetwork‑based framework for estimating both epistemic and aleatoric uncertainties in V2X‑based perception. Specifically, we propose a partial weight generation scheme and V2X context embedding module that conditions a Bayesian hypernetwork on fused multi‑agent features to generate weight distributions for stochastic Bird's‑Eye‑View (BEV) segmentation. Unlike existing deterministic BEV models, Hyper‑V2X enables efficient uncertainty estimation with little computation overhead. Our approach is architecture‑agnostic, and can be seamlessly integrating with modern cooperative backbones such as CoBEVT. Experiments on the OPV2V benchmark demonstrate that Hyper‑V2X provides accurate, well‑calibrated uncertainty estimates and improves overall perception reliability. Our code and benchmark are publicly available under an open‑source license: https://github.com/abhishekjagtap1/Hyper‑V2X
Authors:Robin Louiset, Edouard Duchesnay, Benoit Dufumier, Antoine Grigis, Pietro Gori
Abstract:
In biomedical Subgroup Discovery, practitioners are interested in discovering interpretable and homogeneous subgroups within a group of patients. In this paper, assuming that healthy subjects (i.e., controls) share common but irrelevant factors of variation with the patients, we motivate and develop a Contrastive Subgroup Discovery method, entitled Deep UCSL. By contrasting patients with controls, Deep UCSL identifies subgroups driven solely by pathological factors, ignoring common variability shared with healthy subjects. Our framework employs a deep feature extractor to learn a discriminative representation space. Mathematically, we derive a novel loss based on the conditional joint likelihood of latent clusters and patient/control labels, optimized via an Expectation‑Maximization strategy alternating between subgroup inference and feature encoder updates. A regularization term further encourages representations to capture disease‑specific variability while ignoring variability shared with controls. Compared to previous related works, our approach quantitatively improves the quality of the estimated subgroups, as demonstrated on a MNIST example and four distinct real medical imaging datasets. Code and datasets are available at: https://github.com/rlouiset/deep_ucsl.
Authors:Yifan Wang, Yijia Ma, Wen Li, Chenyu You
Abstract:
High‑fidelity EEG generation is critical for alleviating data scarcity and addressing privacy constraints in large‑scale neural modeling. Despite recent progress, most existing approaches formulate EEG generation via discrete denoising objectives, which inadequately reflect the inherently continuous temporal dynamics and spectral structure of neural activity. As a result, these methods often struggle to preserve long‑range temporal dependencies and exhibit mismatches in the spectral and temporal structure of the generated signals. In this work, we argue that effective EEG generation requires models that operate directly on the continuous evolution of neural signals. We introduce Just EEG Transformer (JET), a generative framework based on conditional flow matching that models EEG as raw sequences evolving along continuous trajectories. By learning a smooth vector field that transports noise to the EEG data distribution, JET captures temporal continuity and transient dynamics without relying on discretized denoising schemes or domain‑specific representations. To ensure that the learned dynamics remain consistent with key properties of EEG signals, we introduce principled constraints that preserve spectral structure, temporal stationarity, and signal‑level statistics. Across three large‑scale benchmarks, JET consistently achieves state‑of‑the‑art performance, reducing TS‑FID by over 40% compared to strong baselines. Extensive analyses show that JET captures key structural properties of neural dynamics, providing a scalable and principled approach to EEG generation. Project page: https://y‑research‑sbu.github.io/JET/ .
Authors:Xiaoyu Zhou, Jianwei Fei, Peipeng Yu, Jingchang Xie, Chong Cheng, Zhihua Xia
Abstract:
The rapid evolution of generative AI, from GANs to modern diffusion models, has resulted in increasingly subtle discriminative clues. These fine‑grained signals are often overshadowed by dominant, high‑fidelity image content (e.g., the main subject), limiting the reliability of existing detectors that predominantly rely on global representations. To address this challenge, we propose the Peak‑Guided Calibration (PGC) framework. PGC introduces a novel strategy that aggregates salient features via a peak‑focusing mechanism. Specifically, by employing a peak‑sensitive aggregation that accentuates the most discriminative local clues, PGC leverages these critical signals to calibrate the global decision. This approach recovers subtle patterns that would otherwise be submerged in the global context. Furthermore, to better simulate real‑world threats, we introduce the CommGen15 dataset, a challenging benchmark comprising samples from 15 commercial models. Extensive experiments demonstrate that PGC achieves state‑of‑the‑art performance. Specifically, it improves mean accuracy by +12.3% on our CommGen15 dataset, and sets new records on standard benchmarks, including GenImage (+2.1%), AIGI (+3.5%), and UniversalFakeDetect (+1.7%). Code is available at https://github.com/xiaoyu6868/PGC.
Authors:Jeonghun Baek, Atsuyuki Miyai, Shota Onohara, Hikaru Ikuta, Kiyoharu Aizawa
Abstract:
Manga is a culturally distinctive multimodal medium and one of the most influential forms of Japanese popular culture. As AI systems increasingly target manga understanding, OCR, and translation, Manga109 has become a foundational dataset for manga‑related AI research. However, the current Manga109 dataset contains transcription errors and coarse annotations, which do not align well with modern OCR and multimodal manga understanding tasks. In this work, we revisit the dialogue text annotations of Manga109 and identify five categories of annotation issues, including transcription errors, missing text regions, overlapping dialogue and onomatopoeia, and under‑segmented speech balloons. To address these issues, we combine OCR‑based issue detection and manual revision to construct Manga109‑v2026, revising approximately 29,000 dialogue annotations. Our revisions better align Manga109 with modern OCR and multimodal manga understanding systems while preserving expressive structures characteristic of manga.
Authors:Kesong Li, Yixuan Xu, Kuo-kun Tseng, Weiyi Lu, Kan Liu, Tao Lan
Abstract:
Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text‑to‑image generation. Existing studies are confined to denoising diffusion models while overlooking flow‑matching, and suffer from an objective mismatch when applying discrete NLP‑based DPO to regression‑based generative tasks.\ In this paper, we derive a generalized DPO objective that covers both diffusion and flow‑matching via a unified reverse‑time SDE framework, and point out from a gradient perspective that the standard DPO objective is suboptimal for text‑to‑image generation. Consequently, we propose Linear‑DPO, which replaces the aggressive sigmoid‑based utility function with a sustained linear utility and incorporates an EMA‑updated reference model. Qualitative and quantitative experiments on diffusion models (SD1.5, SDXL) and flow‑matching model (SD3‑Medium) demonstrate the superiority of our approach over existing baselines.
Authors:Yuanhan Wang, Yifei Chen, Beining Wu, Mingxuan Liu, Xiaotian Hu, Chunbo Jiang, Yijin Li, Changmiao Wang, Feiwei Qin, Qiyuan Tian
Abstract:
Accurate estimation of the Angle of Progression (AoP) from intrapartum transperineal ultrasound is critical for objective assessment of labor progression, yet remains highly sensitive to imaging noise, boundary ambiguities, and the geometric amplification of local segmentation errors. We propose R2AoP, a reliable and robust AoP estimation framework that integrates structurally informed segmentation and confidence‑guided geometric modeling to achieve stable and reproducible measurements. A three‑branch local‑structure‑enhanced backbone improves the delineation of the pubic symphysis (PS) and fetal head (FH), while confidence‑weighted contour fitting explicitly suppresses the influence of unreliable boundary points in AoP computation. To further improve performance under heterogeneous acquisition conditions, we introduce a lightweight geometry‑reliable test‑time adaptation strategy as an auxiliary component, enabling stable inference without target annotations. Extensive evaluations on multi‑center benchmarks demonstrate consistent reductions in AoP error and boundary metrics compared with state‑of‑the‑art AoP methods. Our source code is available at https://github.com/baiyou1234/R2AoP.
Authors:Yiheng Lin, Siyu Jiao, Xiaohan Lan, Wei Zhou, Qi She, Fei Yu, Heyun Chen, Zhengwei Wang, Jinghuan Chen, Moran Li, Yingchen Yu, Zijian Feng, Yao Zhao, Yunchao Wei, Yujie Zhong
Abstract:
Recent advances in Multimodal Large Language Models (MLLMs) and diffusion‑based generative models have substantially improved prompt‑driven image editing. However, scene text editing remains challenging, as it requires models to precisely modify textual content while preserving visual realism and non‑target regions. Current open‑source models still lag behind proprietary systems, largely due to the scarcity of high‑quality training data and the lack of standardized benchmarks tailored to text editing. To address these challenges, we present TextSculptor, a comprehensive framework for data construction and evaluation of scene text editing. We first develop an automated data construction pipeline that combines text‑aware image synthesis with programmatic text rendering and compositing. Based on this pipeline, we build TextSculpt‑Data, a large‑scale dataset containing 3.2M training samples, including 1.2M OCR‑verified text‑to‑image samples and 2M paired text editing samples with naturally aligned source‑target images and strong background consistency. We further introduce TextSculpt‑Bench, a benchmark covering four fundamental text editing tasks: text addition, text replacement, text removal, and hybrid editing. To support reliable evaluation, we design a tailored protocol that measures text accuracy, visual quality, and background preservation through OCR‑based text alignment, multimodal judgment, and background‑region similarity. Extensive experiments show that TextSculptor improves open‑source text editing performance and narrows the gap to proprietary models. The data and benchmark are available at https://github.com/linyiheng123/TextSculptor.
Authors:Zhiyi Zhou, Libo Zhu, Zihan Zhou, Yulun Zhang, Xiaokang Yang
Abstract:
Capturing digital screens with smartphones frequently induces severe banding due to hardware synchronization mismatches. Existing video restoration methods struggle with these structured, periodic luminance fluctuations, often resulting in residual artifacts or over‑smoothed textures. We firstly construct DeViD, a real‑world dataset in various scenes to deal with the lack of available datasets. Then we propose VDFP (Video Deflickering with Flicker‑banding Priors), a novel perception‑guided generation framework. First, we introduce a Degradation Field Modeling Based on Rolling Shutter Mechanism (DFM) capable of synthesizing complex multi‑banding scenarios. Second, we present a spatial‑temporal continuous prior perception (CPP). Unlike traditional binary segmentation, this module is optimized via a Flicker‑Aware Mean Squared Error (FA‑MSE) to capture the luminance transitions. By zero‑initializing an augmented input layer, our model preserves pre‑trained generative priors as well as spatial‑temporal prior perception. Extensive experiments demonstrate that VDFP significantly outperforms other methods, eliminating complex banding with high‑fidelity spatial details and temporal consistency. Our dataset and code will be released at https://github.com/ZhiyiZZhou/VDFP.
Authors:Siao Tang, Xinyin Ma, Gongfan Fang, Xingyi Yang, Xinchao Wang
Abstract:
Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real‑time interactive video generation and world modeling. Despite their potential, the substantial inference cost of ARVDs remains a major obstacle to practical deployment, making model quantization a natural direction for improving efficiency. However, quantization for ARVDs remains largely unexplored. Our empirical analysis shows that directly applying existing quantization schemes developed for standard diffusion transformers to ARVDs leads to suboptimal performance, revealing quantization behaviors that differ from those observed in bidirectional diffusion models. In this paper, we identify two critical challenges in quantizing ARVDs: (C1) Highly unbalanced frame‑wise quantization sensitivity. Error accumulation during autoregressive generation can induce severely skewed quantization sensitivity across frames, following an exponential‑like decay pattern. (C2) Prominent and heterogeneous outlier patterns in weights. Weight distributions exhibit pronounced outlier channels, whose patterns vary substantially across layer types and block depths. To address these issues, we propose Q‑ARVD, a novel framework for accurate ARVD quantization. (S1) To tackle the highly unbalanced frame‑wise sensitivity, Q‑ARVD incorporates a final‑quality aware frame‑weighting mechanism into the quantization objective. (S2) To prevent heterogeneous outliers from degrading performance, Q‑ARVD introduces an outlier‑aware adaptive dual‑scale quantization, which automatically detects the presence and quantity of outlier channels for an arbitrary layer, and isolates them to protect normal channels. Extensive experiments demonstrate the superiority of Q‑ARVD.
Authors:Bo Ye, Xinyu Cui, Jian Zhao, Tong Wei, Min-Ling Zhang
Abstract:
Autoregressive long video generation often adopts bounded‑memory streaming for efficiency, typically combining local windows for short‑term continuity with static early‑frame sinks as long‑range anchors. However, this fixed allocation keeps early frames cached even when the current visual state has substantially diverged from them, while discarding potentially more relevant intermediate history. As a result, the retained long‑range context may become less adaptive and bias generation toward outdated cues; in severe cases, RoPE‑induced phase re‑alignment can homogenize inter‑head attention and cause sink collapse, where content regresses toward sink frames. We propose DySink, a retrieval‑based framework that maintains a compact memory bank and selects visually relevant historical frames as dynamic frame sinks. DySink couples adaptive retrieval with a sink anomaly gate, which detects excessive inter‑head consensus over retrieved context and suppresses collapse‑prone context. Experiments on minute‑long videos show that DySink consistently improves dynamic degree over strong baselines while also achieving higher temporal quality. The code and model weights will be released at https://github.com/yebo0216best/DySink.
Authors:Daniel Eskandar, Berna Kabadayi, Garvita Tiwari, Gerard Pons-Moll
Abstract:
Existing 3D clothed avatar reconstruction methods achieve high visual fidelity but ignore geometric structure and physical plausibility. They either model clothed humans as a single deformable surface or attempt garment disentanglement without enforcing geometric constraints, resulting in ambiguous garment boundaries and no control over stacking or layer ordering. To address these limitations, we introduce DAMA (Disentangled body‑Anchored Gaussians for Controllable Multi‑layered Avatars), a 3D avatar reconstruction method that produces physically plausible clothed avatars through a dedicated representation and reconstruction method. At the representation level, we bind Gaussians to SMPL‑X faces using barycentric in‑plane coordinates and a positive normal offset. Based on this parameterization, the reconstruction method lifts 2D segmentations to body‑anchored Gaussians, refines layers using topology‑guided correction, and jointly optimizes geometry and appearance. DAMA is the first Gaussian avatar reconstruction method from multi‑view images to achieve physically plausible layering, clean garment separation, and explicit stacking control. On the full 4D‑DRESS dataset (82 scans), it achieves state‑of‑the‑art performance in geometry reconstruction, garment separation, penetration rate, and penetration depth. The representation further supports user‑defined garment reordering and fast conversion of body‑conforming garments to simulation‑ready meshes. Project Page: https://danieleskandar.github.io/dama/
Authors:Yutong Xie, Zhenglin Hua, Ran Wang, Wing W. Y. Ng, Xizhao Wang, Yuheng Jia
Abstract:
Large Vision‑Language Models (LVLMs) have shown remarkable performance on a wide range of vision‑language tasks. Despite this progress, they are still prone to hallucination, generating responses that are inconsistent with visual content. In this work, we find that LVLMs tend to hallucinate when they pay insufficient attention to the correct visual evidence and gradually forget it during the generation process. We empirically find that although LVLMs overall attend insufficiently to visual evidence, they exhibit sensitivity to the correct visual evidence in specific layers, with notable inter‑layer discrepancy. Motivated by this observation, we propose a novel hallucination mitigation method that enhances visual evidence based on Inter‑Layer Visual Attention Discrepancy (ILVAD). Specifically, we obtain the attention weights from early generated tokens to visual tokens across layers and identify the tokens that are repeatedly activated as visual evidence, forming a saliency map. We then enhance attention to visual evidence during generation through the saliency map to reduce visual forgetting. In addition, we leverage the saliency map to obtain attention scores of generated text to visual evidence, in order to select and emphasize text tokens that are strongly grounded in visual evidence. Our method is training‑free and plug‑and‑play. Multiple benchmark evaluations conducted on five recently released models show that our method can consistently mitigate hallucinations in different LVLMs over various architectures. Code is available at https://github.com/ytx‑ML/ILVAD.
Authors:Zhangchi Hu, Wenzhang Sun, Xiangchen Yin, Jiahui Yuan, Chunfeng Wang, Hao Li, Kun Zhan, Xiaoyan Sun
Abstract:
Existing 4D‑driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source‑observed regions while synthesizing disoccluded or out‑of‑view content. We identify Evidence‑Role Mismatch: reliable source‑backed evidence, unreliable rendered cues, and unsupported regions are entangled in a single conditioning signal, causing preservation drift, ghosting, and unstable extrapolation. We propose PREX (Preserve, Reveal, Expand), a region‑aware framework that decomposes the target spatiotemporal volume into Preserve, Reveal, and Expand roles according to observation support and scene extent. PREX builds observation‑backed appearance cues with calibrated confidence and injects them into a frozen video diffusion backbone through a region‑aware adapter, trained with proxy tasks without requiring paired edited videos. We further introduce PREBench, a diagnostic benchmark with curated edits, region‑role masks, and human‑aligned metrics that complement global video‑quality and 4D‑control evaluations. Experiments show that PREX reduces region‑structured failures while maintaining strong visual quality and 4D edit control capability. Project Page: https://ricepastem.github.io/PREX‑Open
Authors:Tao Wang, Lei Jin, Zhihua Wu, Qiaozhi He, Jiaming Chu, Yu Cheng, Junliang Xing, Jian Zhao, Shuicheng Yan, Li Wang
Abstract:
Text‑to‑motion generation, which translates textual descriptions into human motions, faces the challenge that users often struggle to precisely convey their intended motions through text alone. To address this issue, this paper introduces DrawMotion, an efficient diffusion‑based framework designed for multi‑condition scenarios. DrawMotion generates motions based on both a conventional text condition and a novel hand‑drawing condition, which provide semantic and spatial control over the generated motions, respectively. Specifically, we tackle the fine‑grained motion generation task from three perspectives: 1) freehand drawing condition. To accurately capture users' intended motions without requiring tedious textual input, we develop an algorithm to automatically generate hand‑drawn stickman sketches across different dataset formats; 2) multi‑condition fusion. We propose a Multi‑Condition Module (MCM) that is integrated into the diffusion process, enabling the model to exploit all possible condition combinations while reducing computational complexity compared to conventional approaches; and 3) training‑free guidance. Notably, the MCM in DrawMotion ensures that its intermediate features lie in a continuous space, allowing classifier‑guidance gradients to update the features and thereby aligning the generated motions with user intentions while preserving fidelity. Quantitative experiments and user studies demonstrate that the freehand drawing approach reduces user time by approximately 46.7% when generating motions aligned with their imagination. The code, demos, and relevant data are publicly available at https://github.com/InvertedForest/DrawMotion.
Authors:Jiawen Dai, Yue Song
Abstract:
Oscillations and synchronization are widely believed to play a fundamental role in representation and computation. However, existing machine learning approaches based on synchronization dynamics have largely been confined to specialized settings such as object discovery, with limited evidence of scalability to standard vision benchmarks or logic reasoning tasks. We propose the Winfree Oscillatory Neural Network (WONN), a dynamical neural architecture based on generalized Winfree dynamics. WONN evolves representations on the torus (S^1)^d through structured oscillatory interactions, combining phase‑based inductive biases with flexible and hierarchical interaction mechanisms instantiated as either fixed trigonometric mappings or learnable neural networks. We evaluate WONN on image recognition and complex reasoning tasks, including CIFAR, ImageNet, Maze‑hard, and Sudoku. Across these domains, WONN achieves competitive or superior performance with strong parameter efficiency. In particular, WONN is, to our knowledge, the first synchronization‑based oscillatory architecture to scale competitively to ImageNet‑1K. Furthermore, on Maze‑hard, WONN achieves 80.1% accuracy using only 1% of the parameters of prior state‑of‑the‑art models. These results suggest that structured oscillatory dynamics provide a scalable and parameter‑efficient alternative to conventional neural architectures.
Authors:Chaoran Xu, Yingmao Miao, Pengfei Zhang, Hao Dou, Lei Sun, Xiangxiang Chu
Abstract:
Vision‑language models (VLMs) have achieved strong multimodal reasoning capabilities, but further improving them still relies heavily on large‑scale human‑constructed supervision for post‑training. Such supervision is costly to obtain, especially for reasoning‑intensive multimodal tasks where questions, answers, and feedback signals must be carefully designed. This motivates self‑evolving learning, where a model improves itself through a dual‑role closed loop: a questioner autonomously poses questions and a solver learns to solve them. However, we observe that current VLM self‑evolving methods still face three major challenges: coarse‑grained role alternation delays the interaction between question generation and solver adaptation; generated questions can progressively degrade in quality; and question types may collapse toward a narrow distribution. These issues limit the efficiency and reliability of self‑evolution. Thus, we propose RISE, a reliable self‑evolving framework for vision‑language models. RISE is built on three complementary designs: fine‑grained role alternation, which shortens the feedback loop between the questioner and the solver to improve efficiency; a quality supervisor, which improves question validity and pseudo‑label reliability; and skill‑aware dynamic balancing, which mitigates mode collapse and maintains broad skill coverage during evolution. Together, these components enable more reliable and effective self‑evolution from unlabeled images. Experiments on two VLM backbones across seven benchmarks show that RISE consistently improves the base models, yielding broad and sustained gains. Our code is publicly available at https://github.com/AMAP‑ML/RISE.
Authors:Qiaohui Chu, Haoyu Zhang, Yisen Feng, Meng Liu, Weili Guan, Dongmei Jiang, Liqiang Nie
Abstract:
We propose JFAA, a JEPA‑based Future Action Anticipation method for the EPIC‑KITCHENS‑100 (EK‑100) Action Anticipation task. Inspired by the representation learning and future prediction ability of V‑JEPA 2.1, JFAA uses a frozen encoder and predictor to extract observed context features and near‑future latent tokens. A lightweight attentive probe is then trained to predict verb, noun, and action logits with separate task queries. To improve robustness, we further build a field‑aware ensemble over selected epoch‑level predictions, allowing each output field to benefit from its most reliable candidates. Experimental results on the official challenge server show that JFAA achieves first place in the EgoVis 2026 EK‑100 Action Anticipation Challenge. Our code will be released at https://github.com/CorrineQiu/JFAA.
Authors:Qiaohui Chu, Haoyu Zhang, Yisen Feng, Meng Liu, Weili Guan, Dongmei Jiang, Liqiang Nie
Abstract:
We propose VISTA, a V‑JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short‑Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric video timestamp, the task requires anticipating the next human‑object interaction, including the future active object's bounding box, noun category, verb category, time‑to‑contact, and confidence score. VISTA follows a StillFast‑style design that combines object‑centric spatial detection with short‑horizon temporal context. Specifically, a COCO‑pretrained Faster R‑CNN ResNet‑50 FPN detector generates object proposals from the last observed high‑resolution frame, while a frozen V‑JEPA 2.1 temporal branch extracts clip‑level egocentric context from the observed video. The temporal representation is injected into the detection pathway through feature modulation and ROI‑level context fusion. The fused proposal features are then passed to multi‑head STA predictors for box refinement, noun classification, verb classification, time‑to‑contact regression, and interaction confidence estimation. For the final submission, we further ensemble complementary predictions to improve robustness. Experimental results on the official challenge server show that VISTA achieves first place in the EgoVis 2026 Ego4D STA Challenge. Our code will be released at https://github.com/CorrineQiu/VISTA.
Authors:Huayi Wang, Haochao Ying, Yuyang Xu, Qiyao Zheng, jun wang, Cheng Zhang, Ying Sun, Jian Wu
Abstract:
Multimodal survival prediction, a crucial yet challenging task, demands the integration of multimodal medical data (\eg Whole Slide Images (WSIs) and Genomic Profiles) to achieve accurate prognostic modeling. Given the inherent heterogeneity across modalities, the feature decoupling‑fusion paradigm has emerged as a dominant approach. However, these methods have the following shortcomings: (1) fail to reduce the redundant information of modality features before decoupling, which negatively affects the feature decoupling and fusion effect;(2) lack the ability to model the fine‑grained relationships of the features and capture the local information interactions between intra‑ and inter‑modality features. To address these issues, we propose a \underlineHierarchical \underlineDecoupling‑Fusion \underlineMixture‑\underlineof‑\underlineExperts (HDMoE) framework with two levels of MoE and \underlineRandom \underlineFeature \underlineReorganization (RFR) modules.In the first‑level MoE, shared experts and routed experts are employed to remove redundant information and extract fine‑grained specific features within each modality, while the second‑level MoE facilitates fine‑grained inter‑modality feature decoupling. Besides, we design two RFR modules following each level of MoE to finely fuse intra‑ and inter‑modality features, which can help the model capture more fine‑grained relationships between modalities. Extensive experimental results on our private Liver Cancer (LC) and three TCGA public datasets confirm the effectiveness of our proposed method. Codes are available at https://github.com/ZJUMAI/HDMoE.
Authors:Hiroyuki Deguchi, Ryosuke Hori, Kotaro Amaya, Tsubasa Maruyama, Mitsunori Tada, Hideo Saito
Abstract:
Monocular egocentric human pose estimation is essential for ubiquitous activity monitoring. However, understanding the user's absolute location within the environment remains a challenge. Existing methods primarily focus on relative motion from an initial position, and tend not to account for the wearer's absolute location within an environment. Furthermore, inherent scale ambiguity in monocular vision leads to severe translational drift, limiting long‑term tracking without specialized multi‑sensor hardware. To address this, we propose MapMonoEgo, a novel framework achieving globally consistent human pose estimation solely from a monocular camera by leveraging a pre‑scanned 3D point cloud. We also introduce AIST‑Living dataset, a new dataset pairing egocentric video with ground‑truth motion in a scanned environment. Experiments demonstrate that our approach significantly outperforms the state‑of‑the‑art baseline, proving its utility for practical monitoring tasks without specialized hardware.
Authors:Jiae Yoon, Ue-Hwan Kim
Abstract:
In this work, we address the challenge of Scene Change Detection (SCD), where the goal is to identify variations between two images of the same location captured at different times. Existing SCD models often overlook the varying importance of features across layers, employ single‑step decoders that confine refinement, and provide limited insight into encoder pretraining strategies. We propose TERDNet, a Transformer Encoder‑Recurrent Decoder Network designed to overcome these limitations. TERDNet consists of a transformer‑based encoder that extracts multi‑level representations, a feature fusion module that integrates correlation volumes with these features, a recurrent 3‑gate‑GRU decoder that performs iterative refinement, and a combined convolution‑interpolation upsampler that restores fine‑grained resolution. Extensive experiments on four public benchmarks show that TERDNet consistently outperforms prior approaches and produces more accurate and detailed change masks. Ablation studies confirm the benefit of segmentation‑based pretraining and the effectiveness of our fusion design. In addition, robustness tests under viewpoint misalignment confirm TERDNet's potential for deployment in real‑world robotic systems, where reliable perception is critical. Our code is available at https://github.com/AutoCompSysLab/TERDNet.
Authors:Zhaojie Zeng, Yuesong Wang, Yawei Luo, Tao Guan
Abstract:
2D Gaussian splatting provides an efficient explicit representation for image reconstruction, but existing methods still require costly per‑image iterative optimization or rely on handcrafted priors for primitive allocation. We present AIR, a self‑supervised feed‑forward framework that amortizes iterative Gaussian fitting into a single network pass, eliminating per‑image test‑time optimization. AIR adopts a stage‑wise residual architecture that progressively predicts additional Gaussian primitives from reconstruction residuals, together with an explicit Stage Control mechanism that activates new primitives only in under‑reconstructed regions. A Predict‑‑Optimize‑‑Distill training strategy stabilizes multi‑stage prediction by distilling short‑horizon optimized Gaussian increments back into the predictor. The stabilized predictor is then jointly finetuned across stages and equipped with an image‑adaptive quantizer for compact Gaussian storage. Experiments on Kodak and DIV2K show that AIR achieves better reconstruction quality than representative Gaussian‑based baselines while reducing encoding time to 160‑‑300\,ms. Code: https://github.com/whoiszzj/AIR.git
Authors:Yisen Feng, Leigang Qu, Haoyu Zhang, Qiaohui Chu, Meng Liu, Xuemeng Song, Weili Guan, Liqiang Nie
Abstract:
In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accurately localizing temporal segments from long untrimmed egocentric videos. To address these tasks, we propose a reranking‑based framework that effectively leverages the strong video‑language reasoning capability of multimodal large language model (MLLM) while preserving the efficiency and candidate recall of conventional localization pipelines. Specifically, we first obtain a set of candidate segments from existing localization model OSGNet, and then employ MLLM to select the segment that best matches the given query, thereby refining the final prediction. Ultimately, our method achieved first place in both the Natural Language Queries and GoalStep tracks. Our code can be found at https://github.com/iLearn‑Lab/CVPR25‑OSGNet.
Authors:Jinjin Zhang, Xiefan Guo, Di Huang
Abstract:
Modern ultra‑high‑resolution image synthesis relies heavily on the robust generative capacity of large‑scale pre‑trained Latent Diffusion Models (LDMs). While recent representation alignment methods have proven effective by distilling visual priors from foundation models (e.g., SAM or DINO) into generative latent features, scaling these approaches to pre‑trained LDMs at extreme resolutions exposes a critical learnability‑fidelity conflict. Specifically, forcing direct patch‑wise feature distillation inherently perturbs the pre‑trained latent manifold, ultimately leading to generation degradation. To address this bottleneck, we propose Spatial Gram Alignment (SGA), a novel framework that explicitly leverages the representation priors of vision foundation models while preserving the native generative capacity of LDMs. Moving beyond restrictive direct alignment, SGA imposes a non‑invasive spatial constraint by aligning the internal self‑similarities of the generative features with those of the foundation priors. This spatial constraint effectively establishes macroscopic structural coherence, while the native generative objectives retain the microscopic pixel‑level fidelity inherent to the original LDMs. Notably, this versatile strategy integrates seamlessly across both intermediate diffusion features and VAE latents within pre‑trained LDMs. Extensive experiments demonstrate that SGA achieves state‑of‑the‑art performance for ultra‑high‑resolution text‑to‑image synthesis, yielding an effective reconciliation between global structural integrity and fine‑grained visual details. Code is available at https://github.com/zhang0jhon/SGA.
Authors:Haozhe Jia, Pengyu Yin, Wenshuo Chen, Shaofeng Liang, Lei Wang, Bowen Tian, Xiucheng Wang, Nanqian Jia, Yutao Yue
Abstract:
Physics‑informed diffusion models typically enforce PDE constraints only on final outputs, leaving intermediate representations unconstrained and prone to shortcut learning under shifted boundary conditions. We introduce REPA‑P, a teacher‑free, architecture‑agnostic framework that aligns intermediate features with physical states using first‑principles residuals. REPA‑P attaches lightweight 1×1 projection heads to selected layers, decodes hidden activations into physical quantities, and applies PDE residual losses during training. These heads are discarded at inference, introducing zero overhead. Across four PDE tasks, including Darcy flow, topology optimization, electrostatic potential, and turbulent channel flow, REPA‑P accelerates convergence by up to 2×, reduces physics residuals by up to 66.4%, and improves out‑of‑distribution robustness by up to 49.3%, with consistent gains on both U‑Net and Diffusion Transformer backbones. Ablations show that supervising a small set of intermediate layers captures most benefits and complements output‑level physics losses. Code is available at [https://github.com/Hxxxz0/REPA‑P](https://github.com/Hxxxz0/REPA‑P).
Authors:Manogna Sreenivas, Rohit Kumar, Soma Biswas
Abstract:
Visual storytelling with diffusion models has made impressive strides in maintaining character consistency across narrative scenes. However, a critical gap remains: while these methods ensure a character remains consistent across scenes, they provide no systematic method to ensure if fine‑grained attributes such as color and textures of clothing, accessories are faithfully rendered in the generated images. Towards this goal, we introduce AttriStory, a benchmark enabling attribute realization in visual storytelling. We curate 200 multi‑scene stories across 10 distinct artistic styles using Large Language Model. Each scene is constructed with detailed attribute specifications to enable rich visual narratives. Further, to address attribute realization, we propose a plug‑and‑play latent optimization module that operates during early denoising steps, when the model establishes structural and semantic content. We achieve this through AttriLoss objective designed to maximize alignment between the cross‑attention maps for desired attribute‑object pairs while suppressing spurious associations, guiding models to localize attributes correctly. This approach operates orthogonally to existing consistency mechanisms, integrating seamlessly with current story generation pipelines without requiring architectural modifications. Our experiments demonstrate consistent improvements on incorporating AttriLoss across all baselines. This work positions attribute realization as a distinct, complementary dimension of visual storytelling, alongside character consistency, advancing the field toward fine‑grained attribute‑controlled story generation. Project‑page:https://manogna‑s.github.io/attristory/
Authors:Jiayi Chen, Benteng Ma, Zehui Liao, Winston Chong, Yasmeen George, Jianfei Cai
Abstract:
While medical Multimodal Large Language Models (MLLMs) have shown promise in assisting diagnosis, they still frequently generate hallucinated responses that appear linguistically plausible but lack visual evidence. Such hallucinations pose risks to clinical decision‑making and necessitate effective detection. Existing introspective detection methods primarily perform uncertainty estimation or logical verification by analyzing model responses conditioned on original or perturbed inputs. However, such external perturbations are often heuristic and context‑agnostic, which overlooks the internal cross‑modal dependency between generated tokens and related visual tokens during decoding. To address this issue, we propose VIHD, a Visual Intervention‑based Hallucination Detection method that leverages targeted visual token masking to calibrate semantic entropy for more effective hallucination detection. VIHD locates visually dominant decoder layers via Visual Dependency Probing (VDP), executes Visual Intervention Decoding (VID) via token masking to calibrate the semantic distribution, and quantifies the resulting Calibrated Semantic Entropy (CSE) as a reliable hallucination signal. Extensive experiments on three medical VQA benchmarks with two medical MLLMs demonstrate that VIHD consistently outperforms state‑of‑the‑art methods, underscoring the importance of fine‑grained visual dependency for hallucination detection. The code will be available at https://github.com/Jiayi‑Chen‑AU/VIHD
Authors:Zhu Liu, Yuanhang Yao, Ping Qian, Zihang Chen, Risheng Liu
Abstract:
Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two coupled bottlenecks: unstable pseudo‑label evolution in cluttered, low‑contrast infrared imagery and severe sample‑distribution imbalance. In this paper, we present a more adaptive and stable framework to address these issues. Leveraging the intrinsic consistency between thermal radiation patterns and heat diffusion, we propose a physics‑induced annotation strategy that expands single‑point labels into reliable pseudo‑masks. To further enhance supervision and alleviate sample imbalance, we develop a bi‑level dual‑update framework that jointly optimizes detector weights, sample weights, and diffusion parameters. A meta‑classifier dynamically predicts sample‑wise loss weights, while a differentiable diffusion module refines pseudo‑labels with detection feedback, enabling adaptive interaction between training and hyperparameter optimization. Extensive experiments across multiple datasets demonstrate five‑fold annotation acceleration, superior detection accuracy, and comparable performance with 30% of the training data, validating the efficiency and practicality of our approach. Our code is available at https://github.com/yuanhang‑yao/diffuse‑to‑detect.
Authors:Xuehui Yu, Fucheng Cai, Meiyi Wang, Xiaopeng Fan, Harold Soh
Abstract:
Inference‑time guided sampling steers state‑of‑the‑art diffusion and flow models without fine‑tuning by interpreting the generation process as a controllable trajectory. This provides a simple and flexible way to inject external constraints (e.g., cost functions or pre‑trained verifiers) for controlled generation. However, existing methods often fail when composing multiple constraints simultaneously, which leads to deviations from the true data manifold. In this work, we identify root causes of this off‑manifold drift and find that the approximation error scales severely with gradient misalignment. Building on these findings, we propose Conflict‑Aware Additive Guidance (g^\textcar), a lightweight and learnable method, which actively rectifies off‑manifold drift by dynamically detecting and resolving gradient conflicts. We validate g^\textcar across diverse domains, ranging from synthetic datasets and image editing to generative decision‑making for planning and control. Our results demonstrate that g^\textcar effectively rectifies off‑manifold drift, surpassing baselines in generation fidelity while using light compute. Code is available at https://github.com/yuxuehui/CAR‑guidance.
Authors:Yaoteng Zhang, Qing Zhou, Junyu Gao, Qi Wang
Abstract:
Remote sensing imagery typically arrives in the form of continuous data streams. Traditional detectors often forget previously learned categories when learning new ones; therefore, research on Remote Sensing Incremental Object Detection (RS‑IOD) is of great significance. However, existing methods largely overlook the intra‑class scale variations prevalent in remote sensing scenes, which undermines the effectiveness of knowledge transfer and old knowledge preservation. Moreover, RS‑IOD also suffers from missing annotations, which cause the model to misclassify old‑class instances as background. To address these challenges, we propose a novel framework, STAR‑IOD. First, we introduce a Subspace‑decoupled Topology Distillation (STD) module to transfer structural knowledge, explicitly aligning inter‑class topological relationships and mitigating intra‑class representation discrepancies induced by scale shifts. Furthermore, we introduce the Clustering‑driven Pseudo‑label Generator (CPG), a plug‑and‑play module that leverages K‑Means clustering to dynamically identify class‑specific thresholds, thereby guaranteeing an accurate distinction between true positive targets and background noise and alleviating the issue of missing annotations for old classes. We also constructed two Remote Sensing Incremental Object Detection datasets, DIOR‑IOD and DOTA‑IOD to facilitate research on RS‑IOD. Extensive experiments demonstrate that our method outperforms state‑of‑the‑art approaches by 1.7% and 2.1% mAP on DIOR‑IOD and DOTA‑IOD, respectively, effectively alleviating catastrophic forgetting while preserving strong detection performance on both base and novel classes. The code and dataset are released at: https://github.com/zyt95579/STAR‑IOD.
Authors:Siqi Wei, Hongbin Xu, Feng Xiao, Tian Lan, Chun Li, Ming Li, Qiuxia Wu
Abstract:
Existing approaches for unsupervised 3D point cloud segmentation predominantly rely on a purely visual similarity‑based learning‑by‑clustering paradigm, which suffers from a fundamental limitation: long‑tail ambiguity. In such a paradigm, features of minor classes are consistently absorbed by dominant clusters, leading to severely imbalanced predictions. To address this issue, we propose LangTail, a language‑guided hierarchical learning framework that leverages the balanced world knowledge encoded in language models to mitigate long‑tail ambiguity in unsupervised 3D segmentation. The key idea is to establish multi‑level associations between language‑derived semantic priors and visually underrepresented minor classes, thereby compensating for the biased attention of purely visual clustering toward dominant classes. Specifically, LangTail first constructs an entity‑level semantic prior from language models, capturing balanced and fine‑grained world knowledge across categories. These priors are injected into a hierarchical clustering framework via contrastive alignment. This guides multi‑granularity semantic structure formation and prevents minor classes from being absorbed by dominant clusters, yielding more discriminative representations for underrepresented categories. Extensive experiments on ScanNet‑v2, S3DIS, and nuScenes demonstrate that LangTail consistently outperforms existing methods by significant margins, \ie, +13.5, +12.9, and +8.9 mIoU, respectively. These results demonstrate the effectiveness of language priors in improving the representation of minority classes in 3D point clouds. The code will be released at: https://github.com/Whisky0129/langtail_official.
Authors:Chia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang, Noranart Vesdapunt, Stefan Stojanov, Bharath Hariharan, Oleksandr Obiednikov, Ning Zhou
Abstract:
Inferring rigid‑body physical states and properties from monocular videos is a fundamental step toward physics‑based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, making them unable to generalize to complex real‑world settings. We introduce ΔYNAMICS, a vision‑language framework that uses language as a unified representation of rigid‑body dynamics. Instead of directly predicting parameters, ΔYNAMICS generates scene configurations in a structured text format for physics simulation. We enhance the model's generalization by integrating natural language motion reasoning and leveraging optical flow as a semantic‑agnostic input. On the CLEVRER dataset, ΔYNAMICS achieves a segmentation IoU of 0.30, a 7x improvement over leading VLMs (InternVL3‑8B, Qwen2.5‑VL‑7B and Claude‑4‑Sonnet). Additionally, test‑time sampling and evolutionary search further boost performance by 27% and 120% in segmentation IoU, respectively. Finally, we demonstrate strong transfer to a new dataset of 235 real‑world rigid‑body videos, highlighting the potential of language‑driven physics inference for bridging perception and simulation.
Authors:Xu Han, Mohammad Aminul Islam, Lei Wang, Zekun Long, Guanmanyi Fu, Wangshu Cai, Kuldip K. Paliwal, Jun Zhou
Abstract:
Hyperspectral imagery encodes rich material properties that can improve tracking robustness under appearance ambiguity, illumination change, and background clutter. However, due to the limited availability of hyperspectral video data, many existing methods adapt pretrained RGB trackers via spatial or channel fusion strategies, largely neglecting the intrinsic material information in hyperspectral imagery. Moreover, the few material‑aware approaches typically rely on external spectral unmixing pipelines that are decoupled from the tracking objective, limiting effective optimization of material representations for target localization. To address these limitations, we formulate hyperspectral object tracking as a joint optimization problem of material decomposition and target localization, coupling the two tasks via a weighted target‑oriented unmixing loss that explicitly aligns material representations with localization accuracy. Specifically, we propose a material representation decomposition module for deep learning‑based spectral unmixing with adaptive frequency decomposition. Building on the decomposed material representations, we further introduce a dual‑branch wavelet‑enhanced material prompt module that learns low‑ and high‑frequency material prompts through efficient spatial‑material interactions in the frequency domain. The framework is model‑agnostic and can be seamlessly generalized to different unmixing backbones. Extensive experiments on standard hyperspectral tracking benchmarks demonstrate state‑of‑the‑art performance and validate the effectiveness of the proposed end‑to‑end material‑aware tracking framework. Code is available at https://github.com/han030927/E2EMPT.
Authors:Doguhan Yeke, Elif Su Temirel, Ananth Shreekumar, Brandon Lee, Dongyan Xu, Z Berkay Celik
Abstract:
Vision‑language models (VLMs) are used as high‑level planners for embodied agents, translating natural language instructions and visual observations into action plans. While prior work has studied abstention in LLMs, existing benchmarks are largely text‑only and do not capture the perceptual grounding and physical constraints inherent to embodied robotics environments. In such settings, abstention requires recognizing when instructions are ambiguous, physically infeasible, based on false premises, or otherwise unresolvable given the available sensory modalities and context. To address this gap, we introduce a taxonomy to categorize abstention in the context of embodied robotics and present RoboAbstention, a scalable and auditable framework for generating abstention instructions grounded in images gathered from five robotics datasets. RoboAbstention instantiates the taxonomy through a three‑phase pipeline: (1) structured visual grounding, (2) deterministic constraint derivation, and (3) controlled instruction generation via category‑specific templates. This enables the construction of a diverse dataset with verifiable abstention conditions. We evaluate several frontier VLMs and find that all models exhibit significant weaknesses in abstention, including those with advanced reasoning capabilities. The best‑performing model, Gemini 2.5 Flash, abstains on only 39.0% of our 6,069 benchmark instructions, while the embodied planner Gemini Robotics ER 1.6 Preview abstains on just 16.5%. We further explore methods for improving abstention in VLM planners, such as defensive prompting and in‑context learning, and find that these interventions substantially improve performance, reaching 93.6% abstention rate for Gemini Robotics ER 1.6 Preview and 88.6% for GPT 5.4 Mini, yet no approach fully solves the problem. We open‑source RoboAbstention at https://purseclab.github.io/RoboAbstention/.
Authors:Huan Huang, Michele Esposito, Chen Zhao
Abstract:
Accurate vessel segmentation is essential for medical image analysis, yet remains challenging due to complex vascular patterns and imaging ambiguity. Most deep models rely on single‑pass prediction, limiting their ability to refine uncertain or disconnected regions during inference. To address this limitation, we propose Uncertainty‑Guided Conservative Propagation (UGCP), a general plug‑in module for vessel segmentation. Instead of directly using a one‑shot output as the final prediction, UGCP performs a small number of logit‑space update steps to refine the segmentation through local predictions interaction. Predictive uncertainty guides reliable regions to support ambiguous regions, while structure‑aware modulation and source‑based stabilization reduce unreliable propagation and excessive drift. The module is differentiable and can be trained end‑to‑end with different segmentation networks. We evaluate UGCP on four public vessel segmentation datasets covering 2D and 3D tasks, including retinal vessel, coronary artery, and cerebral vessel segmentation. Experiments with convolutional neural network‑based and Transformer‑based backbones show consistent improvements in Dice similarity coefficient, centerline Dice, and 95th percentile Hausdorff distance. Further analysis demonstrates that UGCP reduces vessel disconnections and improves structural consistency with limited additional computation. The code will be made available at https://github.com/chenzhao2023/UGC_PR.
Authors:Longchao Da, Mithun Shivakoti, Xiangrui Liu, T Pranav Kutralingam, Yezhou Yang, Hua Wei
Abstract:
Urban heat exposure is becoming an increasingly critical challenge due to the intensifying urban heat island effect. Fine‑grained shade patterns, especially those induced by urban buildings, strongly influence pedestrians' thermal exposure and outdoor activity planning. However, accurately modeling and analyzing urban shade at scale remains difficult because of the lack of large‑scale datasets and systematic evaluation frameworks. To address this challenge, we present ShadeBench, a comprehensive dataset and benchmark for urban shade understanding. ShadeBench contains geographically diverse urban scenes with temporally varying simulated shade maps and textual descriptions, together with aligned satellite imagery, building skeleton representations, and 3D building meshes. Built upon this multimodal dataset, ShadeBench supports a range of downstream tasks, including shade generation, shade segmentation, and 3D building reconstruction. We further establish standardized evaluation protocols and baseline methods for these tasks. By enabling scalable and fine‑grained shade analysis, ShadeBench provides a foundation for data‑driven urban climate research and supports future studies in heat‑resilient urban planning and decision‑making. The code and dataset are publicly available at https://darl‑genai.github.io/shadebench/.
Authors:Pablo Marcos-Manchón, Rishi Jha, Lluís Fuentemilla
Abstract:
The Strong Platonic Representation Hypothesis suggests that representational convergence in artificial neural networks can be harnessed constructively: embeddings can be translated across models through a universal latent space without paired data. We ask whether an analogous geometry can be recovered across human brains. Using fMRI data from the Natural Scenes Dataset, we propose a self‑supervised encoder that learns subject‑specific embeddings from brain data alone by exploiting repeated stimulus presentations. We show that these independently learned spaces can be translated across subjects using unsupervised orthogonal rotations, without paired cross‑subject samples or intermediate model representations. Synchronizing pairwise rotations into a single shared latent space further improves cross‑subject retrieval, indicating that subject‑specific spaces are mutually compatible with a common coordinate system. These results provide evidence for a shared neural geometry in the human visual cortex: subject‑specific fMRI representations are approximately isometric across individuals and can be translated through purely geometric transformations.
Authors:Xinqi Xiong, Andrea Dunn Beltran, Junmyeong Choi, Sarah K. McGill, Marc Niethammer, Roni Sengupta
Abstract:
Accurate polyp size stratification guides surveillance decisions, with lesions larger than 5 mm typically requiring closer follow‑up. However, monocular colonoscopy lacks a reliable metric reference. We present a diagnostic audit of binary polyp size classification (<=5 mm vs. >5 mm) across multiple public multi‑center datasets, model families, and patient‑stratified cross‑validation. Across architectures and input modalities, including RGB appearance, relative depth, and photometry, model performance is moderately consistent, suggesting reliance on cues correlated with examination behavior rather than true metric scales. By providing ground‑truth scale at varying granularities, we quantify the potential improvement from perfect scale information and show that current depth estimation and global calibration offer limited gains. We further demonstrate that segmentation errors under distribution shift eliminate most of this potential, with oracle scale under predicted masks recovering only baseline performance. These results highlight metric scale and mask robustness as two independent bottlenecks and provide reusable evaluation tools such as oracle scale ladders, shortcut partitions, and mask substitution for auditing future polyp sizing pipelines. Our code is publicly accessible at https://github.com/anaxqx/polyp‑sizing‑audit.
Authors:Iason Skylitsis, Dimitrios Karkalousos, Ivana Išgum
Abstract:
Class imbalance is a fundamental challenge in medical image segmentation, where frequent classes typically dominate training at the expense of rare classes. Loss‑based approaches mitigate imbalance by reweighting the per‑pixel loss within the batch, while sampling strategies control which images enter the batch. Yet neither explicitly controls which classes appear within the batch, leaving rare‑class exposure only partially rebalanced. In this work, we adopt episodic sampling from few‑shot learning to promote class‑balanced batch construction in a fully supervised setting. We decouple episodic sampling from its conventional metric‑learning context and evaluate it in body composition segmentation in CT. We compare episodic sampling against random and weighted sampling on nine muscle and adipose tissues, derived from 210 scans of the public SAROS dataset. Training is performed under full‑ and low‑data regimes, with additional comparisons under matched training iteration budgets. Under full‑data training, all three strategies performed comparably (mean Dice 0.882 for episodic, 0.878 for random and weighted). Under low‑data training, episodic sampling outperformed random and weighted (0.787 vs. 0.758 and 0.762), driven by a 12‑fold difference in training iterations. Under matched training budgets, random and weighted overfit earlier, while episodic improved for approximately three times more iterations before plateauing. Our findings identify the training iteration budget as under‑recognized confound in sampling strategies, motivating iteration‑aware evaluation protocols for small datasets. Furthermore, the residual advantage of episodic sampling is consistent with an implicit regularization effect of class‑balanced batches, offering a low‑cost, model‑agnostic strategy for class‑imbalanced medical image segmentation. Code is available at https://github.com/iasonsky/episodic‑sampling.
Authors:Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara, Lorenzo Torresani
Abstract:
Predicting how a person's first‑person view will evolve (what action will follow, what plan completes a task, whether an in‑progress shot will score) is fundamentally under‑specified: the same context admits many plausible futures, and a model trained to minimize prediction error is forced to hedge or average across them, getting it wrong either way. Two findings shape our approach. First, the future camera trajectory, the path the head carves through space, lets the model commit to one of those futures: it carries the operator's intent in a form fine enough to determine how an action will unfold, substantially outperforming language as a conditioning signal. Second, this same intent makes the trajectory itself partially predictable from the context at hand, enough that trajectory need not be observed at test time to recover most of the gain. We instantiate these findings as TrajPilot, a model that predicts candidate future trajectories from egocentric context and uses them to pilot action prediction in an action‑aligned embedding space where language shapes the structure but is never used as a conditioning input. TrajPilot beats VLM and structured‑planner baselines on procedural planning across Ego‑Exo4D atomic, Ego‑Exo4D Keystep, Ego4D GoalStep, and EgoPER, with the trajectory advantage widening with horizon (exactly where prior planners collapse) and holding under RGB‑only camera‑pose estimation. With the goal masked at inference, the same model performs goal‑free anticipation, beating VLM baselines on Ego‑Exo4D atomic and extending to EPIC‑Kitchens‑100 and basketball shot‑outcome prediction.
Authors:Tianshu Wu, Xiangqi Kong, Yue Chen, Qize Yu, Hang Ye, Jia Li, Yizhou Wang, Hao Dong
Abstract:
Building humanoid robots capable of generalizable whole‑body loco‑manipulation in the real world remains a fundamental challenge. Existing methods either rely on laborious task‑specific reward engineering, rigidly replay reference motions that fail to generalize, or depend on costly teleoperation that limits scalability. While human videos capture diverse human behaviors, motion priors inferred from them are inherently imperfect, suffering from occlusion, contact artifacts, and retargeting errors that render them unsuitable for direct policy learning. To address this, we present SUGAR, a scalable data‑driven framework that converts diverse human videos into deployable humanoid loco‑manipulation skills, without any task‑specific reward engineering or reference‑motion conditioning at inference. SUGAR proceeds in three stages. First, a fully automated pipeline extracts kinematic interaction priors including human‑object motion trajectories and contact labels from unstructured human videos. Second, a privileged physics‑based refiner uses a unified mimic reward and progressive state pool to transform imperfect priors into physically feasible, high‑fidelity skills. Third, refined skills are distilled into a hierarchical autonomous policy consisting of a command generator and a command tracker. We evaluate SUGAR on six representative loco‑manipulation tasks in simulation and real‑world humanoid hardware. Our method substantially outperforms reference‑tracking baselines, and performance scales clearly with the amount of human video data. It also achieves zero‑shot real‑world transfer with reliable closed‑loop execution, autonomous failure recovery, and stable long‑horizon performance under external perturbations. Project Page: https://tianshuwu.github.io/sugar‑humanoid/
Authors:Irem Ulku, Ö. Özgür Tanrıöver, Erdem Akagündüz
Abstract:
Multimodal semantic segmentation benefits remote sensing analysis by combining complementary information from different sensor modalities. In real‑world remote sensing applications, one or more modalities may be unavailable due to sensor failures, adverse atmospheric conditions, or data acquisition problems. Even with pretrained multimodal representations and existing fine‑tuning or adaptation strategies, performance may remain limited because all modality availability scenarios are typically treated as equally informative during training. In this paper, we propose a novel training strategy that learns a scenario sampling distribution directly from the pretrained latent space. Instead of relying on uniform random modality dropout, the proposed method guides fine‑tuning toward more informative modality availability scenarios. More specifically, we quantify the effect of each scenario independently based on the distortion it induces in the shared latent representation. We then capture scenario relations using a radial basis function kernel and derive refined scenario scores through a regularized kernel smoothing. These scores are then converted into a probability distribution during scenario sampling for fine‑tuning. We evaluate this strategy on three remote sensing image sets, namely DSTL, Potsdam, and Hunan, using CBC‑SLP, CBC, and CMX backbones. The experimental results with different image sets and backbones show that our method outperforms standard fine‑tuning and LoRA‑based adaptation. These findings suggest that the pretrained latent representation can serve as an effective basis for sampling during missing modality fine‑tuning. Code is available at https://github.com/iremulku/Latent‑Space‑Guided‑Scenario‑Sampling
Authors:Zuhao Yang, Kaichen Zhang, Sudong Wang, Keming Wu, Zhongyu Yang, Bo Li, Xiaojuan Qi, Shijian Lu, Xingxuan Li, Lidong Bing
Abstract:
Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video‑processing tools (e.g., cropping) has become a promising route to long‑video understanding. However, existing native‑RL methods dispatch tool calls sequentially (i.e., one per turn): a single wrong crop propagates errors without peer correction, multi‑turn tool calls corrupt context, and inference cost scales linearly with the number of turns. We introduce ParaVT, the first multi‑agent end‑to‑end RL‑trained framework for Parallel Video Tool calling, dispatching multiple time‑window crops in a single turn for cleaner context and better fault tolerance. Yet applying standard RL to ParaVT reveals an obstacle we term the Tool Prior Paradox: the pretrained tool priors that enable tool exploration also destabilize cold‑started structural format and expose the skip‑tool reward shortcut under temperature sampling. A cross‑model contrast on a weaker‑prior LMM supports this claim: format stays stable but RL elicits zero tool calls, indicating that prior strength is the shared driver of both format collapse and tool exploration. We propose PARA‑GRPO (Parseability‑Anchored and Ratio‑gAted GRPO), which augments standard RL with two complementary mechanisms: (i) a targeted format reward applied only at the structural‑token positions most prone to collapse, and (ii) a per‑prompt frame‑budget randomization that creates training prompts where calling the tool yields a measurable reward signal over skipping it. Across six long‑video understanding benchmarks, ParaVT improves over the Qwen3‑VL baseline by +7.9% on average, with PARA‑GRPO lifting training‑time format compliance from 0.13 to 0.64. As tool capabilities become increasingly internalized in modern LMMs, RL must cooperate with the resulting priors, and ParaVT offers a general recipe for agentic RL. Code, data, and model weights are publicly available.
Authors:Eric Tillmann Bill, Enis Simsar, Alessio Tonioni, Thomas Hofmann
Abstract:
Modern text‑to‑image diffusion models encode rich visual priors, but expose them only through one‑way text‑conditioned generation. Existing unified vision‑‑language models derived from them recover bidirectional capability through large‑scale joint pretraining or substantial retraining of the text pathway, discarding the strong image prior the text‑to‑image backbone already encodes. We introduce \emphFullFlow, a parameter‑efficient recipe that upgrades a pretrained rectified‑flow text‑to‑image model into a bidirectional vision‑‑language generator by training only LoRA adapters and lightweight text heads. FullFlow keeps images in their native continuous flow and adds a discrete insertion process for text. Separate image and text timesteps turn inference into trajectory selection in a two‑dimensional generative space, enabling text\rightarrowimage, image\rightarrowtext, joint sampling, and partial‑text prediction with a single backbone. On Stable Diffusion 3 (SD3) under an identical trainable‑parameter count and matched LoRA rank, FullFlow improves text\rightarrowimage FID from 62.7 to 31.6 and image\rightarrowtext CIDEr from 2.0 to 99.4 over a LoRA equivalent following the previous SOTA formulation (Dual Diffusion) at matched wall‑clock training time, while reducing peak VRAM from ~84\,GB to ~38\,GB and raising throughput by ~8× on two RTX A5000 GPUs in under 24 hours, training only ~5% of the backbone parameters. The same recipe transfers to FLUX.1‑dev and supports downstream VQA through partial‑text generation. These results show that strong bidirectional vision‑‑language capability can be unlocked from pretrained text‑to‑image flow models without full multimodal pretraining.
Authors:Xinlei Liu, Tao Hu, Jichao Xie, Peng Yi, Hailong Ma, Baolin Li
Abstract:
Gradient‑based attacks are important methods for evaluating model robustness. However, since the proposal of APGD, it has been difficult for such methods to achieve significant breakthroughs. To achieve such an effect, we first analyze the issue of "high‑loss non‑adversarial examples" that degrades attack performance in previous methods, and prove that this issue arises from inappropriate objectives for adversarial example generation. Subsequently, we reconstruct the objective as "maximizing the difference between the non‑ground‑truth label probability upper bound and the ground‑truth label probability", and proposes a novel and powerful gradient‑based attack method named Sequential Difference Maximization (SDM). SDM establishes a three‑layer optimization framework of "cycle‑stage‑step". It adopts the negative probability loss function and the Directional Probability Difference Ratio (DPDR) loss function in the initial and subsequent optimization stages, respectively, and approaches the ideal objective of adversarial example generation via stage‑wise sequential optimization. Experiments demonstrate that compared with previous state‑of‑the‑art methods, SDM not only achieves stronger attack performance but also exhibits superior cost‑effectiveness. The code is available at https://github.com/X‑L‑Liu/ICML‑SDM.
Authors:Panagiotis Koromilas, Theodoros Giannakopoulos, Mihalis A. Nicolaou, Yannis Panagakis
Abstract:
Supervised classification has a theoretical optimum, Neural Collapse (NC), yet neither of its two dominant paradigms reaches it in practice. Cross entropy (CE) leaves radial degrees of freedom unconstrained and converges to a degenerate geometry, while supervised contrastive learning (SCL) drives features toward NC during pretraining but discards this structure in a post hoc linear probing phase. We show that both paradigms are different appearances of the same method that contrasts prototypes on the unit hypersphere, and that closing the gap requires fixing each at its point of failure. From the CE side, we propose NTCE and NONL, two normalized losses that import contrastive optimization's missing ingredients into classifier learning: a large effective negative set and decoupled alignment and uniformity terms. From the SCL side, we prove that SCL's objective already optimizes throughout training for a principled classifier whose weights are the class mean embeddings, making linear probing both redundant and harmful. Empirically, on four benchmarks including ImageNet‑1K, NTCE and NONL surpass CE accuracy, closely approximate NC (\geq 95%), and match CE's converged NC on 4/5 metrics in under 7.5% of its iterations, while SCL with fixed prototypes matches linear probing without the hours‑long classifier training phase. The learned geometry yields +5.5% mean relative improvement in transfer learning, up to +8.7% under severe class imbalance, and improved robustness to corruptions on ImageNet‑C. Our work recasts supervised learning as prototype learning on the hypersphere, with NC reached by design.
Authors:Ziyuan Gao
Abstract:
Medical image segmentation faces a fundamental challenge in continual learning: data arrives sequentially from heterogeneous sources, yet effective continual learning requires discovering which tasks share sufficient structure to benefit from joint learning. Existing methods either apply uniform constraints across all tasks, causing catastrophic forgetting when tasks conflict, or require predefined task groupings that cannot anticipate future task diversity. We introduce MedCRP‑CL, a framework that performs online task structure discovery and structure‑aware continual learning. Leveraging the Chinese Restaurant Process (CRP), our method dynamically infers task groupings from clinical text prompts as tasks arrive, without requiring predefined cluster counts or access to future tasks. We term these discovered groupings semantic modalities, as they capture finer‑grained structure than physical imaging modalities by integrating anatomical region and pathological context. Guided by this discovered structure, we maintain semantic modality‑specific LoRA adapters regularized by intra‑modality EWC, ensuring parameter isolation across dissimilar task groups while facilitating knowledge transfer within similar ones. The framework is also replay‑free, storing only aggregate statistics rather than raw patient data. Experiments on 16 medical segmentation tasks across four imaging modalities demonstrate that MedCRP‑CL achieves 73.3% Dice score with only 4.1% forgetting, outperforming the best baseline by 8.0% while requiring 6× fewer parameters. Code is available at https://github.com/zygao930/MedCRP‑CL.
Authors:Xin Zhang, Yabo Chen, Yijie Fang, Wanying Qu, Haibin Huang, Chi Zhang, Feng Xu, Xuelong Li
Abstract:
Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing video generation methods provide minimal physical control, and single‑image‑to‑3D conversion approaches often suffer from object interpenetration. Furthermore, physics‑based scene‑level 3D generation methods exhibit spatial misalignment, stylized artifacts, and inconsistencies with the input data, restricting their use in realistic interactive video synthesis. We propose TelePhysics, a training‑free framework that converts a single image into a physically consistent and controllable video through holistic scene‑level 3D reconstruction. By representing the full scene geometry in a unified spatial coordinate system, TelePhysics resolves object penetration and alignment ambiguity. Unlike prior methods, this formulation enables accurate scenelevel multi‑object interactions and introduces richer, complex control types for advanced mechanicsbased manipulation. By decoupling simulation from rendering, TelePhysics bypasses latency‑heavy priors, achieving real‑time physical interaction previews paired while preserving photorealistic visual fidelity. Experimental results demonstrate that TelePhysics substantially outperforms prior methods in physical fidelity, spatial coherence, and controllability. The open‑source code is available at https://github.com/xinzhang007/TelePhysics.
Authors:Tianwei Lin, Zhongwei Qiu, Jie Cao, Jiang Liu, Wenjie Yan, Bo Zhang, Yu Zhong, Wenqiao Zhang, Yingda Xia, Ling Zhang
Abstract:
Medical vision‑language models (VLMs) have rapidly advanced as general‑purpose multimodal assistants, yet their deployment in 3D Computed Tomography (CT) analysis remains constrained by a persistent mismatch between optimization objectives and clinical rigor. Current Reinforcement Learning (RL) paradigms still rely on lexical proxy signals that induce ``Evaluation Hallucinations'', where models optimize linguistic fluency rather than factual clinical correctness, leading to diagnostically critical errors. To bridge this gap, we introduce the Clinical Abnormality Benchmarking Substrate (CABS), a structured system that decomposes radiology reports into verifiable clinical semantic units. Using CABS, we identify a ``Mechanistic Divergence'' in standard RL, where surface‑similarity rewards drive policy gradients to bypass medical facts. We therefore propose Trajectory‑Integral Feedback GRPO (TIF‑GRPO), a novel framework integrating control‑theoretic principles into policy optimization. By formulating clinical reasoning as a pseudo‑temporal trajectory for anomaly discovery, TIF‑GRPO regulates anatomy‑aware rewards via an integral feedback loop that penalizes persistent omissions as cumulative state errors and suppresses hallucinations as excessive control effort. Experiments on 3D CT benchmarks demonstrate that our approach significantly enhances abnormality detection and clinical faithfulness, establishing a new paradigm for fine‑grained regulation in medical VLMs. Our project is available at \hrefhttps://github.com/ZJU4HealthCare/TIF‑GRPOGitHub.
Authors:Juncheng Wu, Hardy Chen, Haoqin Tu, Xianfeng Tang, Freda Shi, Hui Liu, Hanqing Lu, Cihang Xie, Yuyin Zhou
Abstract:
Recent advances in vision‑language models (VLMs) emphasize long chain‑of‑thought reasoning; yet, we find that their performance on visual tasks is primarily limited by a lack of visual perception as opposed to reasoning itself. In this work, we systematically study the interplay between perception and reasoning in VLM post‑training by decomposing their capabilities into three separate training stages: visual perception, visual reasoning, and textual reasoning, incorporating specialized training data. We demonstrate that visual perception (a) requires targeted optimization with specialized data; (b) serves as a fundamental scaffold that should be solidified through staged training before refining visual reasoning; and (c) is more effectively learned via RL than caption‑based SFT. Our experiments across multiple VLMs demonstrate that staged training consistently improves both visual perception and reasoning performance over merged training. Notably, models trained with our approach achieve 1.5% higher reasoning accuracy with 20.8% shorter reasoning traces, suggesting that superior perception reduces the need for excessive reasoning. Furthermore, we show that this capability‑based staging represents a new curriculum dimension orthogonal to traditional difficulty‑based curricula, and combining both yields further additive gains. Our staged‑training models achieve superior performance among open‑weight VLMs, establishing advanced results on several visual math and perception (e.g., +5.2% on WeMath and +3.7% on RealWorldQA) tasks compared with the base counterpart.
Authors:Hsiang-Wei Huang, Junbin Lu, Kuang-Ming Chen, Jianxu Shangguan, Cheng-Yen Yang, Jenq-Neng Hwang
Abstract:
Vision‑Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect genuine spatial intelligence. We show that existing spatial VLMs lack basic camera motion understanding, a key component of spatial cognition. We propose the Spatial Narrative Score (SNS), an evaluation framework that requires VLMs to generate explicit spatial narratives capturing both scene semantics and camera motion, followed by reasoning with a frozen proxy LLM. Under SNS, state‑of‑the‑art spatial VLMs exhibit significant performance degradation despite high direct question answering accuracy. To address this gap, we introduce CaMo, a camera motion grounded VLM that achieves consistent performance across SNS evaluation and direct spatial question answering accuracy. Our results highlight the importance of explicit spatial narrative externalization for evaluating VLMs with transferable 3D spatial understanding. Our code, data, and model is available at https://github.com/hsiangwei0903/CaMo
Authors:Guangzhi Xiong, Qiao Jin, Sanchit Sinha, Zhiyong Lu, Aidong Zhang
Abstract:
Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clinical trustworthiness. While visual attribution methods are widely used to explain LVLM predictions, whether these explanations actually reflect the visual evidence underlying the model's decision is largely unverified, since ground‑truth annotations for internal model reasoning are typically unavailable. We address this question for chest X‑ray (CXR) reasoning by developing a causal evaluation framework that retains only CXR‑VQA samples for which the expert‑annotated region is verified, via counterfactual editing, to be causally responsible for the model's prediction. Using this framework across 11 attribution methods, six open‑source LVLMs, and two output modes (direct answer and step‑by‑step reasoning), we find that existing attribution methods often fail to identify the evidence used by LVLMs. To address this failure, we propose MedFocus, a concept‑based attribution method that localizes clinically meaningful anatomical regions via unbalanced optimal transport and measures their causal effect on model outputs through targeted interventions. MedFocus produces spatial, concept‑level, and token‑level attributions and substantially outperforms prior methods, taking a step toward more trustworthy attribution for medical LVLMs. Our data and code are available at https://github.com/gzxiong/medfocus/.
Authors:Chonghao Zhong, Linfeng Shi, Hua Chen, Tiecheng Sun, Hao Zhao, Binhang Yuan, Chaojian Li
Abstract:
Training 3D Gaussian Splatting (3DGS) at billion‑primitive scale is fundamentally memory‑bound: each Gaussian primitive carries a large attribute vector, and the aggregate parameter table quickly exceeds GPU capacity, limiting prior systems to tens of millions of Gaussians on commodity single‑GPU hardware. We observe that 3DGS training is inherently sparse and trajectory‑conditioned: each iteration activates only the Gaussians visible from the current camera batch, so GPU memory can serve as a working‑set cache rather than a persistent parameter store. Building on this insight, we introduce TideGS, an out‑of‑core training framework that manages parameters across an SSD‑CPU‑GPU hierarchy via three synergistic techniques: block‑virtualized geometry for SSD‑aligned spatial locality, a hierarchical asynchronous pipeline to overlap I/O with computation, and trajectory‑adaptive differential streaming that transfers only incremental working‑set deltas between iterations. Experiments show that TideGS enables training with over one billion Gaussians on a single 24 GB GPU while achieving the best reconstruction quality among evaluated single‑GPU baselines on large‑scale scenes, scaling beyond prior out‑of‑core baselines (e.g., approximately 100M Gaussians) and standard in‑memory training (e.g., approximately 11M Gaussians).
Authors:Haojun Chen, Haoyang He, Chengming Xu, Qingdong He, Junwei Zhu, Yabiao Wang, Zhucun Xue, Xianfang Zeng, Zhennan Chen, Xiaobin Hu, Hao Zhao, Yong Liu, Jiangning Zhang, Dacheng Tao
Abstract:
Text‑to‑Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imaging technology, the demand for Ultra‑High‑Resolution (UHR) image generation has grown significantly. However, UHR image generation poses great challenges due to the scarcity and complexity of high‑resolution content. In this paper, we first introduce PixVerve‑95K, a high‑quality, open‑source UHR T2I dataset curated with a carefully designed data pipeline, which contains 95K images across diverse scenarios (each image has a minimum pixel‑count of 100M) and seven‑dimensional annotations. Based on our large‑scale image‑text dataset, we take a pioneering step to extend various T2I foundation models to native 100MP generation with three training schemes. Finally, leveraging both conventional metrics and multimodal large language model‑based assessments, our proposed PixVerve‑Bench benchmark establishes a comprehensive evaluation protocol for UHR images encompassing visual quality and semantic alignment. Extensive experimental results on our benchmark and the constructive exploration of training strategies collaboratively provide valuable insights for future breakthroughs.
Authors:Zhiping Yu, Chenyang Liu, Jinqi Cao, Qinzhe Yang, Siwei Yu, Zhengxia Zou, Zhenwei Shi
Abstract:
Multi‑modal remote sensing images are vital for Earth observation, yet complete paired observations are often scarce in practice. Existing generative methods commonly address this problem through isolated pairwise modality translation, but their versatility and scalability remain limited as the number of modalities and generation tasks increases. Here, we develop a generative foundation model MetaEarth‑MM for multi‑modal remote sensing imagery, enabling paired joint generation and any‑to‑any translation across five modalities within a unified model. Recognizing the intrinsic scene consistency underlying multi‑modal observations, we introduce a scene‑centered joint modeling paradigm in MetaEarth‑MM. Unlike previous methods that rely on direct appearance‑level cross‑modal mapping, our model organizes the generation around the underlying scene content. Specifically, MetaEarth‑MM adopts a decoupled architecture that first infers a latent scene representation from available observations, and then generates target modalities conditioned on this intermediate state. To support training, we further construct EarthMM, a large‑scale dataset comprising 2.8 million multi‑resolution global images with 2.2 million aligned pairs. Extensive experiments demonstrate that MetaEarth‑MM not only exhibits strong generative capability and robust generalization across diverse generation tasks, but also supports downstream tasks at both data and representation levels, highlighting its potential as a general foundation model for cross‑modal Earth observation. The code and dataset will be available at https://github.com/YZPioneer/MetaEarth‑MM.
Authors:Zijie Xin, Jie Yang, Ruixiang Zhao, Tianyi Wang, Fengyun Rao, Jing Lyu, Xirong Li
Abstract:
Omni‑modal large language models (om‑LLMs) achieve unified audio‑visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window level. However, processing these dense non‑textual tokens throughout the LLM incurs substantial computational overhead. Although training‑free token selection can reduce this cost, existing methods either focus on visual‑only inputs or prune om‑LLM tokens only before the LLM with fixed per‑modality ratios, failing to capture how cross‑modal token importance evolves across layers. To address this limitation, we first analyze the layer‑wise token dependency of om‑LLMs. We find that visual and audio dependencies follow a block‑wise pattern and gradually weaken with depth, indicating that many late‑layer non‑textual tokens become redundant after cross‑modal fusion. Motivated by this observation, we propose SEATS, a training‑free, stage‑adaptive token selection method for efficient om‑LLM inference. Before the LLM, SEATS removes spatiotemporal redundancy via attention‑weighted diversity selection. Inside the LLM, it progressively prunes tokens across blocks and dynamically allocates the retention budget from temporal windows to modalities using query relevance scores. In late layers, it removes all remaining non‑textual tokens once cross‑modal fusion is complete. Experiments on Qwen2.5‑Omni and Qwen3‑Omni demonstrate that SEATS effectively improves inference efficiency. Retaining only 10% of visual and audio tokens, it achieves a 9.3x FLOPs reduction and a 4.8x prefill speedup while preserving 96.3% of the original performance.
Authors:Xinyi Wang, Angeliki Katsenou, Junxiao Shen, David Bull
Abstract:
Short‑form video poses new challenges to the quality assessment of user‑generated content (UGC) due to its complex generation pipeline, rapid content variation, and mixed distortions. To address this challenge, we propose an end‑to‑end video quality assessment (VQA) framework that employs a dense visual encoder based on CLIP, and incorporates compression priors derived from the frequency domain to generate artifact‑ and structure‑aware weight maps for feature aggregation. By explicitly decomposing artifact, structure, and original visual feature branches and adaptively fusing them over time through a learned gating module, the proposed method achieves accurate and efficient quality prediction. Experimental results show that our method achieves strong performance on short‑form video datasets in terms of average rank and linear correlation (SRCC: 0.736, PLCC: 0.787), while maintaining efficient inference runtime. The code and additional results are available at: https://github.com/xinyiW915/FGSVQA.
Authors:Hongji Yang, Songlian Li, Yucheng Zhou, Xiaotong Zhao, Alan Zhao, Chengzhong Xu, Jianbing Shen
Abstract:
Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, leading to poor performance in professional production workflows such as storyboard sketches and clay render conditions. Existing video generation models, either inject conditions through adapters or couple a generic vision‑language model (VLM) within a diffusion backbone, leaving a capability gap and failing to produce the videos that align with the user's creative intent. We present CogOmniControl, a reasoning‑driven framework that factorizes controllable video generation into creative intent cognition and generation. Specifically, we train a specialized CogVLM using authentic anime production data. Compared to generic VLMs, it generates more professional and clear outputs, accurately cognizing user creative intent from sparse and abstract conditions and tuning these cues into dense reasoning output. Besides, CogOmniDiT unifies the controls from various conditions through in‑context generation and is aligned to the CogVLM reasoning outputs via reinforcement learning. Furthermore, leveraging CogVLM's robust capability in guiding video generation, we release its potential in planning specific evaluators and enable a Best‑of‑N selection for the generated videos. This integration transforms the entire framework into a closed‑loop "harness‑like" architecture. We further introduce CogReasonBench and CogControlBench, built from professional workflows data that carry genuine creative intent rather than simulated ones. Experiments on two benchmarks show that CogOmniControl surpassed the existing open‑source models. The project website: https://um‑lab.github.io/CogOmniControl/
Authors:He-Yang Xu, Pengyuan Zhang, Zongyuan Ge, Xiaoshuai Hao, Serge Belongie, Xin Geng, Yuxin Peng, Xiu-Shen Wei
Abstract:
Fine‑grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high‑fidelity spatial perception, and constraint‑respecting motor execution. However, current embodied AI benchmarks collapse these capacities into binary success rates, systematically inflating reported capabilities by up to 70% and masking the architectural bottlenecks that impede real‑world deployment. We introduce MetaFine, a diagnostic meta‑evaluation framework that disentangles manipulation competency along three axes: understanding, perception, and controlled behavior. Built on a compositional task graph, MetaFine absorbs heterogeneous external benchmarks and reconstructs them into diagnostic scenarios of varying complexity under a unified protocol. Evaluating state‑of‑the‑art vision‑language‑action (VLA) models through this lens exposes severe dimension‑specific failures invisible to conventional metrics. Through targeted causal intervention, we identify the visual encoder's ability to preserve local spatial structure as a key bottleneck for fine‑grained precision: improving it directly unlocks previously inaccessible manipulation capabilities without modifying downstream policies. MetaFine further supports hybrid real‑sim validation, using limited paired real‑world rollouts to calibrate scalable simulation‑based estimates for more stable physical benchmarking. By shifting evaluation from ranking to diagnosis, MetaFine turns benchmarking into an actionable compass for repairing the layered capacities underlying genuine physical dexterity. The MetaFine framework, benchmarks, and supporting resources will be publicly released at our project page: https://metafine.github.io/.
Authors:Ziqi Wang, Xu Zhang, Laibin Chang, Shi Chen, Jiaqi Ma, Huan Zhang
Abstract:
Low‑Light Image Enhancement (LLIE) has long been a challenging problem in low‑level vision, as insufficient illumination often leads to low contrast, detail loss, and noise. Recent studies show that deep learning‑based Retinex theory can effectively decouple illumination and reflectance. However, existing methods frequently suffer from over‑enhancement or color distortion, and often assume uniform noise or ideal lighting. To address these limitations, we propose InterLight, a novel framework that systematically excavates and operationalizes intrinsic illumination priors for LLIE.Our core insight is that robust enhancement requires not just estimating illumination, but constructing an illumination‑aware pipeline. We first inject sensor‑level illumination‑response priors via physics‑guided augmentation, then represent the degradation through adaptive prompts conditioned on the scene's latent illumination state. This explicit representation directly guides a luminance‑gated intrinsic memory mechanism to selectively compensate for information loss, prioritizing reconstruction in dark regions while preserving fidelity in bright ones. Finally, the entire process is regularized by a self‑supervised consistency objective that distills illumination‑invariant features. By deeply exploiting intrinsic illumination priors, our method achieves clearer textures and more visually coherent enhancement results. Extensive experiments across multiple benchmarks demonstrate the effectiveness of our approach. Code is available at: https://github.com/House‑yuyu/InterLight.
Authors:Jia-Wei Hai, Yijun Wang, Xiu-Shen Wei
Abstract:
Vision‑Language Models (VLMs), such as CLIP, have achieved significant zero‑shot performance on downstream tasks with various fine‑tuning adaptation methods. However, recent studies have proven that adversarial attacks can significantly degrade the inference ability of VLMs, posing substantial risks to their practical applications. Prevalent test‑time adaptation methods typically rely on multi‑view augmentation to implement various fine‑tuning strategies, which struggle to identify semantic information and are prone to destroying discriminative regions in fine‑grained scenarios. To address these limitations, we propose Attention‑Guided Test‑Time Prompt Tuning (A‑TPT), a semantics‑preserving method designed for test‑time adaptation. We first refine the gradient attention rollout mechanism to identify semantically meaningful regions surviving under adversarial attacks. Furthermore, we leverage them to guide the spatially varying augmentation intensities and multi‑view ensemble for prompt tuning and inference. Extensive experiments demonstrate that A‑TPT outperforms existing test‑time adaptation methods on both adversarial and clean data. Codes are available at https://github.com/SEU‑VIPGroup/A‑TPT .
Authors:Yi Zhong, Haotong Qin, Xindong Zhang, Lei Zhang, Guolei Sun
Abstract:
Low‑bit post‑training quantization (PTQ) is a pivotal technique for deploying Vision‑Language Models (VLMs) on resource‑constrained devices. However, existing PTQ methods often degrade VLMs' accuracy due to the heterogeneous activation distributions of text and vision modalities during quantization. We find that this cross‑modal heterogeneity is distributed unevenly across channels: a small subset of channels contains most modality‑specific outliers, and these outliers typically reside in different channels for each modality. Motivated by this, we propose SplitQ, a channel‑Splitting‑driven post‑training Quantization framework. At its core, SplitQ introduces a novel Modality‑specific Outlier Channel Decoupling (MOCD) module that effectively isolates salient modality‑specific outlier channels with minimal overhead. To further address the remaining cross‑modal distribution discrepancies, we design an Adaptive Cross‑Modal Calibration (ACC) module that employs dual lightweight learnable branches to dynamically mitigate modality‑induced quantization errors. Extensive experiments on popular VLMs demonstrate that SplitQ significantly outperforms existing approaches across 6 popular multi‑modal datasets under all evaluated quantization settings, including W4A8, W4A4, W3A3, and W3A2. Notably, SplitQ preserves 93.5% of FP16 performance under the challenging W3A3 setting (69.5 vs. 74.3), pushing the efficiency frontier for deploying advanced VLMs. Our code is available at https://github.com/EMVision‑NK/SplitQ
Authors:Ananth Sriram, Neel Mokaria, Rajveer Singh
Abstract:
Construction remains the deadliest industry sector in the United States, with 1,055 fatal worker injuries recorded in 2023, and the majority preventable. Existing monitoring approaches are expensive, require real‑time human operators, or address only a narrow subset of violations. This paper presents a passive, end‑of‑shift construction safety monitoring pipeline processing video from POV body‑worn and fixed wall‑mounted cameras through a three‑stage architecture: (1) fine‑tuned YOLO11 for primary PPE and hazard detection, (2) SAM 3 for segmentation refinement and worker deduplication, and (3) Qwen3‑VL‑8B‑Instruct with a method‑prompted, persona‑scaffolded three‑pass adversarial chain‑of‑thought protocol for compliance verification and hallucination control. The principal contribution is the Stage 3 prompt design: professional persona backstories following the method‑actor framing drive an observed 12% precision improvement over single‑pass prompting in an informal three‑author review of the 12‑video Ironsite development corpus, with the largest gains on hallucination‑prone violation categories. Structural message isolation enforces observational independence between a generator, discriminator, and reconciliation pass governed by asymmetric rules encoding priors about human observation versus automated detection reliability. The system maps violations to OSHA standards, performs REBA‑inspired ergonomic risk scoring from pose keypoints, and produces per‑worker safety reports with timestamped evidence. An evaluation harness is released for future reproduction.
Authors:Giacomo Astolfi, Matteo Bianchi, Riccardo Campi, Antonio De Santis, Marco Brambilla
Abstract:
Concept‑based Explainable Artificial Intelligence (XAI) interprets deep learning models using human‑understandable visual features (e.g., textures or object parts) by linking internal representations to class predictions, thereby bridging the gap between low‑level image data and high‑level semantics. A major challenge, however, is the reliance on large sets of labeled images to represent each concept, which limits scalability. In this work, we investigate the use of zero‑shot Text‑to‑Image (T2I) generative models as a source of synthetic concept datasets for concept‑based XAI methods. Specifically, we generate concepts using predefined prompts and evaluate their faithfulness to real ones through four complementary analyses: (1) comparing synthetic vs. real concept images via concept representation similarity; (2) evaluating their intra‑similarity by comparing pairs of subsets of the same concept with progressively increasing size; (3) evaluating their performance for downstream explanation tasks using relevant class images; (4) evaluating how removing a concept from tested class images affects explanations of generated concepts. While current T2I generative models promise a shortcut to concept‑based XAI, our study highlights challenges and raises open questions about the use of synthetic data generated by zero‑shot pipelines in model analyses. The resulting dataset is available at https://github.com/DataSciencePolimi/ZeroShot‑T2I‑Concepts.
Authors:Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh, Hung-Ting Su, Winston H. Hsu
Abstract:
Vision‑Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine‑grained comprehension crucial for real‑world applications requiring nuanced interpretation of human actions and interactions. While some recent human‑centric benchmarks evaluate aspects of model behaviour such as fairness/ethics, emotion perception, and broader human‑centric metrics, they do not combine long‑form videos, very dense QA coverage, and frame‑level spatial/temporal grounding at scale. To bridge this gap, we introduce FineBench, a human‑centric video question answering (VQA) benchmark specifically designed to assess fine‑grained understanding. FineBench comprises 199,420 multiple‑choice QA pairs densely annotated across 64 long‑form videos (15 minutes each), focusing on detailed person movement, person interaction, and object manipulation, including compositional actions. Our extensive evaluation reveals that while proprietary models like GPT‑5 achieve respectable performance, current open‑source VLMs significantly underperform, struggling particularly with spatial reasoning in multi‑person scenes and distinguishing subtle differences in human movements and interactions. To address these identified weaknesses, we propose FineAgent, a modular framework that enhances VLMs by leveraging a Localizer and a Descriptor. Experiments show that FineAgent consistently improves the performance of various open VLMs on FineBench. FineBench provides a rigorous testbed for future research into fine‑grained human‑centric video understanding, while FineAgent offers a practical approach to enhance such reasoning in current VLMs. Project page and code at https://joslefaure.github.io/assets/html/finebench.html.
Authors:Jiaxin Wang, Muwei Jian, Hui Yu, Junyu Dong, Yifan Xia
Abstract:
Facial Expression Recognition (FER) in the wild is still challenging due to uncontrolled variations in pose, occlusion, and illumination. Most existing attention‑based methods primarily rely on visual appearance cues, suffering from attention redundancy and instability, which limits their performance in complex scenarios. To address these issues, we propose a novel landmark‑guided contrastive learning network with vision‑language enhancement for FER (LaCoVL‑FER), which integrates geometric priors from facial landmarks and semantic priors from a vision‑language model. Specifically, a Landmark‑Guided Adaptive Encoder (LGAE) is designed to introduce geometric priors through a Bi‑branch Gated Cross Attention (BGCA) mechanism, which achieves adaptive fusion of landmark‑based geometric and visual appearance features to produce expression‑relevant features, thereby focusing on key facial regions and suppressing noise interference. In parallel, a Vision‑Language Enhancement Strategy (VLES) is presented to leverage the expression‑relevant features to refine the generalizable visual features extracted by the frozen pretrained CLIP image encoder, yielding expression‑specific visual representations. Based on these representations, an Expression‑Conditioned Prompting (ECP) mechanism is utilized to further adapt the textual features of fixed class‑level prompts from the frozen pretrained CLIP text encoder, generating more instance‑aware textual representations. These visual‑textual representations are aligned as semantic priors to enhance the robustness and generalization of FER. Quantitative and qualitative experiments demonstrate that our LaCoVL‑FER outperforms state‑of‑the‑art methods on three representative real‑world FER datasets, including RAF‑DB, FERPlus, and AffectNet. The code is available at https://github.com/ylin06804/LaCoVL‑FER.
Authors:Hyojun Go, Hyungjin Chung, Prune Truong, Goutam Bhat, Li Mi, Zhaochong An, Zixiang Zhao, Dominik Narnhofer, Serge Belongie, Federico Tombari, Konrad Schindler
Abstract:
For practical use, diffusion‑ or flow‑based generative models must be aligned with task‑specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challenging because the reward is defined for clean output images, but the alignment procedure requires value function estimates at noisy intermediate latents. Existing methods resort to Tweedie‑style or Monte Carlo approximations, trading off estimator bias against computational cost: Tweedie estimates are efficient but biased, while Monte Carlo estimates are more accurate but require expensive rollouts. A natural alternative would be a learned value function, but it remains an open question how to effectively train a strong and general value model specifically for noisy latents. Here, we propose StitchVM, a model stitching framework that efficiently transfers reward models pretrained for clean images to the noisy latent regime. StitchVM starts from an existing, truncated pixel‑space reward model and attaches a frozen diffusion backbone to it as its head. From the pixel‑space model, the resulting hybrid retains a carefully pretrained, robust reward capability; from the diffusion backbone, it inherits its native ability to handle noisy latents. The stitching procedure is exceptionally lightweight, e.g., stitching and finetuning CLIP ViT‑L and SD 3.5 Medium takes only 10 GPU‑hours. By lifting powerful pixel‑space reward models to latent space, StitchVM opens up a new style of diffusion alignment: instead of rough, yet costly per‑sample approximation of the value function, the correct function for the actual, noisy latents is constructed once and then amortized over many samples and iterations. We show that this approach yields improvements across a broad range of downstream steering and post‑training methods: DPS becomes 3.2× faster while halving peak GPU memory, and DiffusionNFT becomes 2.3× faster.
Authors:Tonghao Zhuang, Shanglong Hu, Yongsheng Luo, Zhiqi Zhang, Yu Li
Abstract:
We present a semi‑supervised framework for joint segmentation and classification of fetal cardiac ultrasound images. Built upon the EchoCare multi‑task backbone, our method integrates SAM‑Med2D for boundary refinement and leverages DINOv3 to enhance pseudo‑label quality. We introduce view‑specific hard masking along with a two‑stage optimization strategy: an EMA phase to consolidate segmentation capabilities, followed by a Classification Fine‑Tuning phase that freezes segmentation parameters and resets the classification head to recover classification performance without compromising segmentation gains.
Evaluated on the FETUS 2026 leaderboard, our method achieves a Dice Similarity Coefficient at 79.99%, Normalized Surface Distance at 61.62%, and F1‑score at 41.20%, validating the effectiveness of our approach for prenatal congenital heart disease screening. Source code is publicly available at: https://github.com/2826056177/zcst_fetus2026.
Authors:Oserebameh Augustine Beckley
Abstract:
Digital Twin (DT) technology holds immense potential for surgical planning and personalized medicine. However, generating interactive, patient‑specific anatomical twins currently relies on computationally heavy Server‑Side Rendering (SSR) or expensive local workstations, creating significant barriers to deployment, especially in resource‑constrained settings (RCS). This paper presents a decentralized, client‑side WebGPU architecture that democratizes access to high‑fidelity anatomical Digital Twins. By bypassing standard server‑side rendering pipelines, the framework executes deterministic single‑pass raymarching and morphological gradient calculations directly on low‑cost integrated edge GPUs. Eliminating the network latency inherent to cloud‑rendered solutions, the system achieves a Time to First Pixel (TTFP) of under 920.0ms and maintains stable interactivity at >= 82.0 FPS. Continuous Interaction Fidelity is maintained via uniform buffers, enabling zero‑latency manipulation of tissue parameters for dynamic clinical decision‑making. By proving that complex 3D medical simulations of patient‑specific MRI scan can be executed natively in the browser without deep learning or external computational dependencies, this architecture provides a scalable, affordable foundation for the widespread clinical adoption of healthcare Digital Twins.
Authors:Hyunsoo Han, Sangyeop Yeo, Jaejun Yoo
Abstract:
We demonstrate that in knowledge distillation for diffusion models, the teacher network's highly complex denoising process ‑ stemming from its substantially larger capacity ‑ poses a significant challenge for the student model to faithfully mimic. To address this problem, we propose a coarse‑to‑fine distillation framework with LInear FiTtingbased distillation (LIFT) and Piecewise Local Adaptive Coefficient Estimation (PLACE). First, LIFT decomposes the objective into a "coarse" alignment and a "fine" refinement. The student is then trained on coarse alignment before proceeding to hard refinement. Second, PLACE extends LIFT to address spatially non‑uniform errors by partitioning outputs into error‑based groups, providing locally adaptive guidance. Our experiments show that LIFT and PLACE is effective across diffusion spaces (image/latent), backbones (U‑Net/DiT), tasks (unconditional/conditional), datasets, and even extends to flow‑based models such as MMDiT (SD3). Furthermore, under extreme compression with a 1.3M‑parameter student (only 1.6% of the teacher), conventional KD fails to provide sufficient guidance for stable training, with FID scores often degrading to 50‑200+, but our method remains stably convergent and achieves an FID of 15.73.
Authors:Kylian Ronfleux-Corail, Guillaume Bernard, Mickaël Coustaty, Nicolas Sidère
Abstract:
Document manipulation localization models achieve strong performance on public benchmarks yet fail to generalize to operational document workflows. We identify a critical and overlooked source of this gap: the mismatch between the narrow distribution of JPEG quantization tables used during training restricted to standard libjpeg quality factors and the heterogeneous compression profiles encountered in real‑world insurance document pipelines. To isolate this factor, we conduct a controlled factorial study comparing two architectures with contrasting levels of quantization table awareness FFDN [2] and Mesorch [20] each trained under either standard quality factor augmentation (Standard‑QT ) or operationally calibrated quantization tables sampled from DocQT, a quantization‑table bank derived from a MAIF operational image corpus (Real‑QT ), and evaluated under three recompression conditions. Training under Real‑QT yields substantial localization gains on DocTamper [15] and significantly reduces the pixel‑level false positive rate on authentic operational documents, but only for architectures that explicitly ingest the quantization table as input. The released DocQT quantization‑table dataset and compression‑reproduction material are directly available at https://github.com/Kyliroco/Improving‑Document‑Forgery‑Localization‑Robustness‑via‑Diverse‑JPEG‑Quantization‑Tables. These results demonstrate that standard quality factor augmentation does not adequately proxy operational compression diversity, and that architectural choices explicitly conditioning on the quantization table provide a meaningful robustness advantage for real‑world deployment.
Authors:Matias Turkulainen, Akshay Krishnan, Filippo Aleotti, Mohamed Sayed, Guillermo Garcia-Hernando, Juho Kannala, Arno Solin, Gabriel Brostow, Daniyar Turmukhambetov
Abstract:
We present Cross‑View Splatter, a feed‑forward method that predicts pixel‑aligned Gaussian splats for outdoor scenes captured at ground level AND by satellite. Faithful reconstructions require good camera coverage, but ground imagery is time‑consuming and hard to capture at scale for large outdoor scenes. Fortunately, satellite imagery can provide a global geometric prior that is easy to access via public APIs. Cross‑View Splatter fuses orthorectified satellite views with GPS‑tagged ground photos to predict Gaussian splats in a unified 3D coordinate frame. By aligning ground and bird's‑eye feature representations, our model improves scene coverage and novel‑view synthesis, compared to ground imagery alone. We train on curated georeferenced datasets and paired satellite‑terrain data, mined from open mapping services. We evaluate our method on a new benchmark for novel‑view synthesis with georeferenced imagery allowing comparison to prior state‑of‑the‑art methods. Our code and data preparation will be available at https://nianticspatial.github.io/cross‑view‑splatter/.
Authors:Junjie Wang, Xinghua Lou, Jason Li, Ye Tian, Keyu Chen, Yulin Li, Bin Kang, Jacky Mai, Yanwei Li, Zhuotao Tian, Liqiang Nie
Abstract:
Text‑to‑Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single‑pass generation paradigm limits their ability to handle complex prompts requiring iterative refinement. To enable multi‑round Reflective Visual Generation (RVG), we formalize the Reason‑Reflect‑Rectify (R^3) loop as a core framework and introduce R^3‑Bench, a benchmark of over 600 expert‑annotated instances that quantifies iterative reasoning and rectification capabilities. Evaluation on R^3‑Bench reveals a critical gap: while state‑of‑the‑art models can identify generation errors, they fail to generate actionable rectification instructions. To bridge this gap, we propose R^3‑Refiner, a dual‑stage framework leveraging Group Relative Policy Optimization (GRPO) and a Hierarchical Reward Mechanism (HRM) to better align rectification with reflective reasoning. Experiments show that R^3‑Refiner achieves significant improvements on R^3‑Bench (+12.0% in Reflective Verdict Score, +9.0% in Rectification Score), and can be seamlessly integrated with various MLLMs to enhance the generation quality of different T2I models on GenEval++ and T2I‑CompBench. Code is available at https://github.com/xiaomoguhz/R3‑Bench.
Authors:Gabriele Rosi, Fabio Cermelli, Carlo Masone, Barbara Caputo
Abstract:
Segmenting images is critical for visual understanding but demands extensive pixel‑level annotations. Foundational models have enabled new paradigms for predicting new classes guided by textual prompts, without annotations from the target domain. Yet, on specialized target domains, far from the original pre‑training, their performance degrades. We study the errors of existing methods under such domain‑shift, finding that misclassification rather than mask generation is the main culprit. To address this, we introduce the novel problem of Few‑Shot Visual Adaptation for text‑prompted Segmentation. This kind of adaptation has been largely studied for image classification, but it remains unexplored for segmentation. We tackle this task with Prototype Adaptation (PrAda), a novel, parameter‑efficient method that adapts a frozen text‑prompted segmentation model. Our approach learns class‑specific prototypes by combining fine‑grained pixel features and high‑level transformer representations, which are then fused with the original text‑based predictions through a learned importance factor. This preserves the model's zero‑shot potential while enabling strong adaptation to new domains. Experiments across semantic, instance, and panoptic segmentation on five benchmarks demonstrate that PrAda yields significant improvements over state‑of‑the‑art and proposed baselines.
Authors:Shuwei Li, Lei Tan, Robby T. Tan
Abstract:
Color constancy aims to keep object colors consistent under varying illumination. Cross‑camera generalization in color constancy remains challenging because learning‑based models often overfit to the color response characteristics of the training camera, resulting in degraded performance on images captured by other cameras. We propose VLM‑CC, a feedback‑guided framework that formulates color constancy as an iterative refinement process. Instead of directly estimating the illuminant from raw input, VLM‑CC performs iterative correction driven by vision‑language model (VLM)‑based evaluation. At each iteration, the image is white‑balanced using the current estimate and converted to pseudo‑sRGB. A lightweight LoRA‑tuned VLM then assesses the corrected image, identifying the dominant residual color cast and providing qualitative feedback. This feedback is mapped to a residual illumination direction (red, green, or blue) and used to update the illuminant estimate until convergence. Our key idea is to reframe color constancy as an iterative perceptual feedback problem, leveraging VLM evaluation instead of direct RGB regression. By replacing direct RGB estimation with VLM‑guided perceptual feedback, VLM‑CC achieves state‑of‑the‑art robustness in cross‑camera color constancy across multiple datasets. Code will be available at https://github.com/NothingIknow/VLM‑CC.
Authors:Soyeon Kim, Seongwoo Lim, Kyowoon Lee, Jaesik Choi
Abstract:
Integrated Gradients (IG) is a widely adopted feature attribution method that satisfies desirable axiomatic properties. However, the choice of integration path significantly affects the quality of attributions, and the standard straight‑line path introduces all input features simultaneously, often accumulating noisy gradients along the way. To address this limitation, we propose Spectral Integrated Gradients, which constructs integration paths based on singular value decomposition (SVD) of the baseline‑to‑input difference. By progressively activating singular components from largest to smallest, SIG introduces global structure before fine‑grained details, naturally following a coarse‑to‑fine progression. Through extensive evaluation across diverse image classification datasets, we demonstrate that SIG produces cleaner attribution maps with reduced noise and achieves improved quantitative performance compared to existing path‑based attribution methods. Our code is available at https://github.com/leekwoon/sig/.
Authors:Mengyuan Liu, Ziyi Wang, Peiming Li, Junsong Yuan
Abstract:
RGB camera‑based surveillance systems enable human action recognition for public safety and healthcare, yet raise serious privacy concerns. Existing methods rely on post‑capture algorithms, which fail to protect privacy during data acquisition. We propose Lens Privacy Sealing (LPS), a simple hardware solution that physically obscures camera lenses with adjustable laminating film, providing pre‑sensor privacy protection at minimal cost. Unlike software methods or expensive engineered optics, LPS achieves strong privacy through stochastic multi‑layer scattering that is physically irreversible. We introduce the P^3AR dataset for privacy‑preserving action recognition, featuring both large‑scale replay‑captured (P^3AR‑NTU, 114K videos) and real‑world collected (P^3AR‑PKU) subsets with privacy attribute annotations. To handle video degradation from LPS, we propose MSPNet, a single‑stage framework incorporating Inter‑Frame Noise Suppressor (IFNS) and Cross‑Frame Semantic Aggregator (CFSA), enhanced by contrastive language‑image pre‑training for robust semantic extraction. Extensive experiments demonstrate that MSPNet with IFNS and CFSA nearly doubles action recognition accuracy compared to baseline methods while suppressing identity recognition to low levels. Comprehensive validation shows LPS achieves a superior privacy‑utility trade‑off compared to state‑of‑the‑art hardware methods, resists reconstruction attacks including PSF inversion and data‑driven recovery, and generalizes robustly across optical configurations and challenging environments. Code is available at https://github.com/wangzy01/MSPNet.
Authors:Yang Dai, Dian Jiao, Tianwei Lin, Wenqiao Zhang
Abstract:
The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine‑grained hand‑object interactions, track object state changes over time, and reason about manipulative processes in dynamic environments from a first‑person perspective. However, existing egocentric video benchmarks suffer from limited grounded rationale evaluation, offering limited support for fine‑grained operation‑centric reasoning and rarely examining whether model rationales are grounded in explicit spatio‑temporal evidence. To address this gap, we introduce EgoCoT‑Bench, a fine‑grained egocentric benchmark for grounded and verifiable operation‑centric reasoning with explicit step‑by‑step rationale annotations. Overall, EgoCoT‑Bench comprises 3,172 verifiable QA pairs over 351 egocentric videos separated into four task groups for a total of 12 sub‑task groups, encompassing perception and retrospection, anticipation, and high‑level reasoning. The benchmark is constructed through a spatio‑temporal scene graphs (STSG) guided generation framework and is further refined by human annotators to ensure correctness, egocentric relevance and fine‑grained quality. Experimental results show continuing difficulties with egocentric fine‑grained reasoning and further reveal that many multimodal models produce explanations that are answer‑correct, but have evidence that is inconsistent with the answer. We hope EgoCoT‑Bench can serve as a useful testbed for grounded and verifiable reasoning in egocentric video understanding. Project page and supplementary materials are available at: https://dstardust.github.io/EgoCoT/.
Authors:Zihao Zhu, Wenyuan Zhao, Nuo Chen, Chao Tian, Zhiwen Fan
Abstract:
Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images. However, in current feed‑forward designs, their predicted confidence scores are heuristic, lack probabilistic interpretation, and often fail to indicate where and how much the predicted geometry can be trusted. To address this gap, we present Trust3R, a lightweight evidential uncertainty framework for feed‑forward 3D reconstruction. Trust3R combines gated residual mean refinement with a Normal‑Inverse‑Wishart evidential head, yielding a closed‑form multivariate Student‑t distribution for per‑point geometric uncertainty. This design provides probabilistically grounded pointmap uncertainty estimates while adding moderate inference overhead. We evaluate on diverse indoor and outdoor benchmarks and compare against MASt3R's built‑in confidence map as well as common uncertainty‑aware baselines spanning single‑pass heteroscedastic regression and sampling‑based methods such as MC dropout and deep ensembles. Experimental results show that Trust3R consistently improves risk‑coverage and sparsification, and generally improves geometric accuracy. These gains are reflected in stronger uncertainty ranking across benchmarks, with 25% lower AURC and 41% lower AUSE on ScanNet++, providing a practical reliability signal for uncertainty‑aware weighting in downstream geometry pipelines. The project page and code are available at https://trust3r‑z.github.io/.
Authors:Zhangjian Ji, Shaotong Qiao, Kai Feng, Wei Wei
Abstract:
Occluded person re‑identification focuses on matching partially visible pedestrians across multiple camera views. However, occlusions disrupt body‑region cues, thereby complicating cross‑view matching. Most person ReID methods built on pretrained vision‑language models only focus on enhancing prompt‑based feature learning while ignoring the semantic information of occluders. Based on the success of CLIP‑ReID, we propose a novel Dual Prompt Learning ReID (DPL‑ReID) model for occluded person ReID. It incorporates a Dual Prompt Learning (Dual‑PL) strategy, which can utilize textual cues to capture complete pedestrian semantics and keep robustness against occlusion, and a Real‑World Occlusion Augmentation (RWOA) method that realistically simulates occlusion scenarios encountered in real word to enrich occluded samples. In addition, we also design a Weighted Gated Feature Fusion (WGFF) method, which in corporates LSNet to capture global information and act as a feature‑gating mechanism. This mechanism can effectively guide the CLIP visual encoder toward generating more comprehensive feature representations. Extensive experiments on several benchmark occluded ReID datasets show that our proposed DPL‑ReID achieves the state‑of‑the art performance. The occlusion instance library are available at https://github.com/stone‑qiao/DPL‑ReID.
Authors:Jiusong Ge, Yingkang Zhan, Wenjie Zhao, Di Zhang, Ke Wang, Jiashuai Liu, Chunze Yang, Chengzu Li, Jian Zhang, Yuxin Dong, Ni Zhang, Qidong Liu, Mireia Crispin-Ortuzar, Huazhu Fu, Chen Li, Zeyu Gao
Abstract:
Traditional whole slide image (WSI) analysis methods typically rely on the multiple instance learning (MIL) paradigm, which extracts patch‑level features at high magnification and aggregates them for slide‑level prediction. However, such exhaustive patch‑level processing is computationally expensive, severely limiting the efficiency and scalability of WSI analysis. To address this challenge, we propose PathCTM (a Pathology‑oriented Continuous Thought Model) that enables token‑efficient scale‑space continuous reasoning for gigapixel WSIs. PathCTM formulates diagnostic inference as a dynamic sequential information pursuit. It progressively transitions from low‑magnification global to high‑magnification local inspection, and adaptively terminates inference when sufficient evidence is gathered to effectively bound decision uncertainty. Specifically, it uses conditional computation for dynamic scale switching with attention‑guided region pruning, coupled with confidence‑aware early stopping. Extensive experiments demonstrate that, compared with standard MIL‑based methods, PathCTM reduces the number of required image patches by 95.95% and shortens inference time by approximately 95.62%, while maintaining AUC without degradation. Code is available at https://github.com/JSGe‑AI/PathCTM.
Authors:Ahmed Heakl, Abdelrahman M. Shaker, Youssef Mohamed, Rania Elbadry, Omar Fetouh, Fahad Shahbaz Khan, Salman Khan
Abstract:
When a model produces a correct solution under reinforcement learning with verifiable rewards (RLVR), every token receives the same reward signal regardless of whether it was a decisive reasoning step or a grammatical filler. A natural fix is to condition the model on the correct answer as a teacher, identifying tokens it would have generated differently had it known the answer. Prior work shows this either corrupts training by leaking the answer into the gradient, or produces a weak signal that cannot distinguish decisive steps from filler, since both look equally surprising relative to the model's baseline. We propose Contrastive Evidence Policy Optimization (CEPO), which asks a sharper question at every token: not just "does the correct answer favor this token?" but "does the correct answer favor it while the wrong answer disfavors it?" A token satisfying both is a genuine reasoning step; one satisfying neither is filler. The wrong‑answer teacher is constructed from rejected rollouts already in the training batch, incurring no additional sampling cost. We prove CEPO inherits all structural safety guarantees of the prior state of the art while strictly sharpening credit at decisive tokens, with the improvement vanishing exactly at filler positions. Empirically, CEPO achieves 43.43% and 60.56% average accuracy across five multimodal mathematical reasoning benchmarks at 2B and 4B scale, respectively, versus 41.17% and 57.43% for GRPO under identical training budgets. Distribution‑matching self‑distillation methods (OPSD, SDPO) fall below the untrained baseline, empirically confirming the information leakage our theory predicts. Our code is available at https://github.com/ahmedheakl/CEPO.
Authors:Maya Yanko, Yoli Shavit
Abstract:
Visual Place Recognition (VPR) is critical for autonomous navigation, yet state‑of‑the‑art methods lack well‑calibrated uncertainty estimation. Standard pipelines cannot reliably signal when a query is ambiguous or a match is likely incorrect, posing risks in safety‑critical robotics. We propose KappaPlace, a principled framework for learning uncertainty‑aware VPR representations. Our core contribution is a Prototype‑Anchored supervision strategy that leverages latent class representatives as targets for a probabilistic objective. By modeling image descriptors as von Mises‑Fisher (vMF) variables, we learn a lightweight module to predict the concentration parameter as a direct proxy for aleatoric uncertainty. While existing VPR uncertainty methods are typically restricted to a query‑centric view, we derive a novel match‑level formulation to quantify the reliability of specific query‑reference pairs. Across five diverse benchmarks, KappaPlace reduces Expected Calibration Error (ECE@K) by up to 50% compared to existing methods while maintaining or improving retrieval recall. We provide both a joint‑training variant and a post‑training extension for frozen backbones. Our results demonstrate that KappaPlace provides a robust, stable, and well‑calibrated signal that enables reliable decision‑making within the VPR pipeline. Our code is available at: https://github.com/mayayank95/UncertaintyAwareVPR
Authors:Chaoyue Li, Yongxue Xu, Jie Feng, Jiayu Ding
Abstract:
Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spatiotemporal dynamic reasoning. To study this capability gap, we formulate trajectory‑grounded multi‑turn spatiotemporal dialogue, a new task in which a model must answer spatiotemporal queries while returning structured 3D target trajectories over an entire short clip or a specified segment of a longer clip, and introduce Track4D‑Bench, a benchmark with 526 clip‑level dialogue samples spanning 23.5k frames and 7.5k object annotations, for training and evaluation. Building on this task, we propose LMM‑Track4D, which combines RTGE (Ray‑‑Time Geometry Encoding), a dedicated streaming state token TRK for long‑horizon dynamic propagation, and an Object‑Slot Kinematic, Residual‑Anchor (OSK‑RA) decoder for stable 4‑step 3D state estimation under occlusion and viewpoint variation. Experiments on Track4D‑Bench show consistent improvements over strong baselines, suggesting that explicit dynamic state modeling is a useful design principle for eliciting 4D dynamic reasoning in LMMs. Our code and dataset will be publicly available at https://github.com/mikubaka88/LMM‑Track4D.
Authors:Chenyu Lian, Hong-Yu Zhou, Chun-Ka Wong, Jing Qin
Abstract:
Vision‑language alignment using chest X‑rays and radiology reports has emerged as an advanced paradigm for zero‑shot classification and grounding of chest X‑ray findings. However, standard contrastive learning typically treats radiographs and reports from different patients simply as negative pairs. This assumption introduces noisy negatives, as different patients frequently exhibit similar findings. Such noisy negatives cause semantic ambiguity and degrade performance in zero‑shot understanding tasks. To address this challenge, we propose CoNNS, a concept‑guided noisy‑negative suppression framework. To support the negative suppression mechanism, unlike previous methods that use raw reports or templatized texts, we construct a hierarchical concept ontology using large language models. The ontology structures 41 key clinical concepts by explicitly modeling presence, attributes (location and characteristics), and texts (evidential segment and presence statement). Leveraging this ontology, we implement a cross‑patient pair relabeling strategy comprising three steps: (1) Fine‑Grained Breakdown to categorize pairs based on finding presence; (2) Noisy Negative Filtering to resolve semantic conflicts by removing false negatives; and (3) Hard Negative Mining to identify subtle attribute discrepancies using a lightweight language model. Finally, we propose a Concept‑Aware NCE loss to align visual features with text while suppressing the identified noisy negatives. Extensive experiments across multi‑granularity zero‑shot grounding tasks and five zero‑shot classification datasets validate that CoNNS outperforms existing state‑of‑the‑art models. The code is available at https://github.com/DopamineLcy/conns.
Authors:Halil Ibrahim Gulluk, Olivier Gevaert
Abstract:
Deep learning methods have demonstrated promising results in predicting BI‑RADS scores from mammography images. However, the interpretation of these images can vary, leading to discrepancies even among radiologists. Given the inherent complexity of mammograms, training classification models solely on image labels often yields limited performance. To address this challenge, we curated 2313 mammogram images and their corresponding captions from two mammography atlases. Our proposed approach employs a multi‑modal model that uses a pretrained PubMedBERT as the language component. By training this model on image‑text pairs with contrastive learning, we enable the vision encoder to absorb the rich information contained in the captions, thereby improving its understanding of mammography findings. We then fine‑tune the vision encoder on two datasets for BI‑RADS prediction, achieving superior performance compared with models trained without this pretraining, particularly when labeled samples are scarce. The improvement in the 3‑class average F1 score ranges from +1% to +14%: a +1% increase with 40K training samples, and a +14% increase with 1K samples. Furthermore, our experiments reveal that 2K image‑text pairs from mammography atlases can be more informative than 2K labeled samples for label prediction, with an average margin of +1.1% when more than 10K training samples are available. Overall, our work provides a vision‑language model for mammography and highlights the value of textual information from mammography atlases. In addition, we publicly release preprocessed mammography images of the TEKNOFEST dataset. The training code, pre‑trained model weights, data extraction scripts, and the released dataset are publicly available at: https://github.com/igulluk/MAM‑CLIP
Authors:Soojin Choi, Seokhyeon Hong, Chaelin Kim, Junghyun Nam, Junhyuk Jeon, Junyong Noh
Abstract:
Retargeting motion across characters with varying body shapes while preserving interaction semantics, such as self‑contact and near‑body proximity, remains a challenging problem. While recent geometry‑aware approaches address this by maintaining spatial relationships between predefined corresponding regions, their reliance on static correspondences often struggles when the target character exhibits exaggerated body proportions. In this paper, we present a geometry‑aware motion retargeting framework that preserves interaction semantics by performing proximity matching over spatially adaptive anchors. Unlike prior methods with static anchor definitions, the proposed method dynamically repositions anchors to reachable regions on the target character. This is achieved via a Transformer‑based anchor refinement strategy that predicts anchor displacements and constrains the translated anchors to remain on the target character geometry through differentiable soft projection. By incorporating pose‑dependent spatial structures from the source character, the adapted anchors provide structurally coherent guidance for interaction‑aware retargeting. Conditioned on these anchors, a graph‑based autoencoder predicts target skeletal motion that preserves the spatial configuration of the source. To encourage task‑aligned optimization between anchor adaptation and motion retargeting, we adopt an alternating training scheme in which each module is optimized in turn. Through extensive evaluations, we demonstrate that our method outperforms state‑of‑the‑art approaches in preserving interaction fidelity across diverse character geometries.
Authors:Yilmaz Korkmaz, Vishal M. Patel
Abstract:
MRI reconstruction is an inherently ill‑posed inverse problem, since incomplete measurements admit many plausible solutions. This ambiguity becomes more severe under high acceleration, where pixel‑domain continuous predictors tend to average over feasible reconstructions and suppress high‑frequency anatomy. We address this limitation by moving reconstruction to discrete multi‑scale latent space and posing it as autoregressive next‑acceleration‑scale prediction. Leveraging discrete priors proven effective in visual autoregressive modeling, our method restricts the solution to compact sequences of codebook tokens, enabling sharp reconstructions even from extremely sparse measurements. This discrete autoregressive formulation also aligns naturally with modern large language model post‑training techniques. Building on this observation, we introduce on‑policy privileged information distillation for visual autoregressive modeling, where a teacher is provided training only privileged context that is unavailable at inference, in our case fully sampled acquisitions, and supervises a student trained on its own rollouts, leading to consistent reconstruction gains. Through extensive experiments on the fastMRI benchmark, we show that our approach delivers improved reconstruction performance across diverse sampling patterns under extreme undersampling. Project website is \hrefhttps://yilmazkorkmaz1.github.io/discrete‑mri‑reconstruction‑opd/here.
Authors:Xuezhi Cui, Dongbo Zhou, Wang Guo, Zeyuan Wang, Ziyu Li, Gaozhi Zhou, Xian Li, Ling Zhao, Wentao Yang, Chao Tao, Haifeng Li
Abstract:
Vision‑Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter‑Efficient Fine‑Tuning mitigates catastrophic forgetting, assigning isolated modules per task leads to parameter explosion. Conversely, recent similarity‑driven sharing mechanisms falsely equate superficial visual similarity with underlying alignment consistency. This fundamental mismatch triggers severe negative transfer between visually similar but logically distinct tasks and fails to exploit alignment reuse across visually diverse ones. We argue thatalignment sharing is fundamentally a geometric problem of overlapping optimization trajectories within shared low‑rank subspaces. Grounded in this insight, we propose iGSP, a novel framework that achieves efficient adaptation via implicit gradient subspace projection. Leveraging the early convergence of MoE routers to establish the subspace basis, iGSP bifurcates the adaptation process into two phases. First, the Subspace Identification phase introduces candidate experts via basis pre‑expansion, applies a novel subspace‑constrained regularization to implicitly project new task gradients onto the historical subspace, and precisely prunes redundant dimensions by treating routing probabilities as gradient flow indicators, ultimately to maximize knowledge reuse. Second, the Orthogonal Subspace Fine‑Tuning phase fixes this structural basis and removes the regularization to rapidly fit the task‑specific residual loss. Extensive experiments on the MTIL benchmark demonstrate that iGSP achieves state‑of‑the‑art accuracy while significantly improving training efficiency, reducing the average trainable parameters by 42.7% compared to current SOTA methods, and decreasing the final total parameters by 86.9% relative to counterparts. The source code is available at https://github.com/GeoX‑Lab/iGSP.
Authors:Jinjin Zhang, Xiefan Guo, Yizhou Jin, Nan Zhou, Di Huang
Abstract:
Driven by rapid advances in large‑scale generative models, synthetic data has emerged as a promising solution for visual understanding. While modern diffusion models achieve remarkable photorealistic image synthesis, their potential in complex visual segmentation tasks remains underexplored. In this work, we conduct a systematic analysis of synthetic images from state‑of‑the‑art diffusion models to uncover the factors governing their utility. In particular, synthetic images characterized by dense scene composition and fine instance fidelity demonstrate distinctive benefits, yielding significantly more discriminative spatial representations. Building on these insights, we propose SENSE, a unified framework that leverages flexible and scalable synthetic data to substantially enhance segmentation performance. Notably, SENSE is model‑agnostic, compatible with diverse architectures (e.g., DPT and Mask2Former), and scales effectively across models with varying parameter capacities. Extensive experiments on Cityscapes, COCO, and ADE20K validate the effectiveness and generalization capability of our approach. Code is available at https://github.com/zhang0jhon/SENSE.
Authors:Svetlana Orlova, Niccolò Cavagnero, Gijs Dubbelman
Abstract:
Video foundation models achieve strong performance across many video understanding tasks, but typically require large‑scale pre‑training on massive video datasets, resulting in substantial data and compute costs. In contrast, modern image foundation models already provide powerful spatial representations. This raises an important question: can competitive video models be built by reusing these spatial representations and pre‑training only for temporal reasoning? We take initial steps toward exploring a lightweight training paradigm that freezes a pre‑trained image foundation model and trains only a recurrent temporal module to process streaming video. By reusing an image foundation model as a spatial encoder, this approach could significantly reduce the amount of video data and compute required compared to end‑to‑end video pre‑training. In this work, we explore the feasibility of this approach before investing in computing for video pre‑training. Our empirical findings across multiple video understanding tasks suggest that strong temporal performance can emerge without large‑scale video pre‑training, motivating future work on recurrent video foundation models obtained by pre‑training a temporal module on top of a frozen image foundation model. Code: https://github.com/tue‑mps/towards‑video‑image‑frozen .
Authors:Mahesh Bhosale, Abdul Wasi, Vishvesh Trivedi, Pengyu Yan, Akhil Gorugantu, David Doermann
Abstract:
Grounded multi‑video question answering over real‑world news events requires systems to surface query‑relevant evidence across heterogeneous video archives while attributing every claim to its supporting source. We introduce CRAFT (Critic‑Refined Adaptive Key‑Frame Targeting), a query‑conditioned pipeline that combines dynamic keyframe selection, per‑video ASR with multilingual fallback, and a hybrid critic loop to iteratively verify and repair claims before consolidation. The pipeline integrates UNLI temporal entailment, DeBERTa‑v3 cross‑claim screening, and a Llama‑3.2‑3B adjudicator, with a final citation‑merging stage that emits each fact once with all supporting source identifiers. On MAGMaR 2026, CRAFT achieves the best overall average (0.739), reference recall (0.810), and citation F1 (0.635). We further evaluate on a MAGMaR‑style conversion of WikiVideo with 52 non‑overlapping event queries, where CRAFT also performs strongly (0.823 Avg), showing that its claim‑centric evidence aggregation generalizes beyond MAGMaR. Ablations show that atomic claims, ASR, and the critic loop drive the main gains over the vanilla query‑conditioned baseline. Code and implementation details are publicly available at https://github.com/bhosalems/CRAFT.
Authors:Ehsan Ahmadi, Hunter Schofield, Behzad Khamidehi, Fazel Arasteh, Jinjun Shan, Lili Mou, Dongfeng Bai, Kasra Rezaee
Abstract:
Supervised open‑loop training has been widely adopted for training traffic simulation models; however, it fails to capture the inherently dynamic, multi‑agent interactions common in complex driving scenarios. We introduce RLFTSim, a reinforcement‑learning‑based fine‑tuning framework that enhances scenario realism by aligning simulator rollouts with real‑world data distributions and provides a method for distilling goal‑conditioned controllability in scenario generation. We instantiate RLFTSim on top of a pre‑trained simulation model, design a reward that balances fidelity and controllability, and perform comprehensive experiments on the Waymo Open Motion Dataset. Our results show improvements in realism, achieving state‑of‑the‑art performance. Compared with other heuristic search‑based fine‑tuning methods, RLFTSim requires significantly fewer samples due to a proposed low‑variance and dense reward signal, and it directly addresses the realism alignment issue by design. We also demonstrate the effectiveness of our approach for distilling traffic simulation controllability through goal conditioning. The project page is available at https://ehsan‑ami.github.io/rlftsim.
Authors:Zachary Yahn, Fatih Ilhan, Tiansheng Huang, Selim Tekin, Sihao Hu, Yichang Xu, Margaret Loper, Ling Liu
Abstract:
Photos of faces uploaded online are vulnerable to malicious actors who can scrape facial images from online sources and intrude on personal privacy via unauthorized use of facial recognition models. This paper presents FaceCloak, a novel personalized face privacy protection system, which can generate defensive identity‑specific universal face privacy masks from a single image of a user, causing facial recognition to fail. FaceCloak introduces a three‑stage personalized face perturbation learning methodology: (1) It generates a small set of high‑variety synthetic face images of a person based on a single image of the person. (2) It learns face cloaking by adding more protection to key facial‑identity leakage regions through iterative perturbation generation over the small set of synthetic images, effectively shifting a user's identity embedding towards a distant anchor identity and away from a similar one. (3) It generates a personalized identity‑protective mask in the form of pixel‑wise cloaking, which is light‑weight and can be efficiently applied to any facial image of a user while maintaining good perceptual quality. Extensive experiments on three popular face datasets across ten recognition models show the effectiveness of FaceCloak compared to 29 other existing representative methods. Code is available at https://github.com/zacharyyahn/FaceCloak
Authors:Xiangxiang Cui, Tianjin Huang, Yifang Wang, Lijie Hu, Lu Yin
Abstract:
Medical foundation models have achieved remarkable clinical performance, yet their robustness under real‑world perturbations remains underexplored. We present a robustness benchmark comprising 40 perturbation types (12 base, 28 medical‑specific) across eight imaging modalities, evaluating five VLMs (LLaVA‑Med, MedGemma, MedGemma‑1.5, Gemini‑2.5‑flash and GPT‑4o‑mini) on VQA, visual grounding, and captioning, alongside two segmentation models (MedSAM, SAM‑Med2D) with five fine‑tuning strategies. Our findings reveal: (1) Fine‑tuning strategy dominates robustness, with LoRA exhibiting nearly double the degradation of full fine‑tuning, while SAM‑Med2D's Adapter offers favorable efficiency‑robustness trade‑off. (2) Medical‑specific perturbations disproportionately damage segmentation, with 9 of 15 top corruptions being domain‑specific. (3) LoRA‑tuned visual grounding drops over 40 points, whereas zero‑shot captioning remains stable (<7% drop). Zero‑shot VQA shows model‑dependent robustness‑‑medical models drop under 20% while Gemini‑2.5‑flash drops 54%. General‑purpose VLMs achieve higher VQA accuracy but fail on grounding; among medical VLMs, MedGemma demonstrates the best overall stability. These results provide deployment guidelines and underscore the necessity of domain‑specific robustness evaluation for medical AI. Our code is available at: https://abnerai.github.io/MedFM‑Robust.
Authors:Ahmad Yehia, Abduallah Mohamed, Tianyi Wang, Jiseop Byeon, Kun Qian, Junfeng Jiao, Christian Claudel
Abstract:
Accurately forecasting human trajectories from an egocentric perspective plays a central role in applications such as humanoid robotics, wearable sensing systems, and assistive navigation. However, progress in this direction remains limited due to the scarcity of egocentric trajectory datasets collected in real‑world environments. Addressing this need, we introduce EgoTraj, an egocentric multimodal open dataset recorded using Meta Quest Pro (MQPro). EgoTraj contains 75 sequences of human navigation collected from multiple MQPro wearers in real‑world urban environments. Each recording provides synchronized RGB video along with ground‑truth data, including continuous time‑synchronized 6‑degree‑of‑freedom head poses, per‑frame 3D eye gaze vectors, scene annotations. To the best of our knowledge, EgoTraj differs from typical egocentric trajectory datasets by capturing long‑horizon, self‑directed navigation across diverse urban routes with broad participant diversity. To demonstrate the potential of the dataset, we benchmark several state‑of‑the‑art methods for egocentric trajectory prediction and conduct ablation studies to analyze the contributions of gaze, scene, and motion cues. The results highlight the utility of EgoTraj for AR‑based perception, navigation, and assistive systems. The EgoTraj dataset, code, and EgoViz Dashboard are publicly available at https://github.com/yehiahmad/EgoTraj.
Authors:Gyubin Lee, Junwon Lee, Juhan Nam
Abstract:
We investigate Counterfactual Video Foley Generation, which aims to adopt a sound‑source identity that contradicts the visual evidence while remaining temporally synchronized to a silent video. Existing Video&Text‑to‑Audio (VT2A) models struggle with this, often remaining anchored to the visually implied sound source when video and text contents disagree. We present ConterFlow, an inference‑time dual‑phase sampling scheme for pretrained flow‑matching VT2A models. Phase 1 builds a video‑derived temporal structure while suppressing the visually implied source; Phase 2 drops video conditioning to focus entirely on shaping audio timbre toward the target prompt. ConterFlow substantially improves counterfactual Video Foley generation compared to naive negative prompting and state‑of‑the‑art baselines. To evaluate replacement quality, we propose a metric leveraging a text‑audio co‑embedding space to measure both target‑prompt evidence and residual visually implied source leakage. Video demonstrations and code are available at https://gyubin‑lee.github.io/counterflow‑demo/
Authors:Cheng Luo, Zefan Cai, Junjie Hu
Abstract:
Attention Residuals replace standard additive residual connections with learned softmax attention over previous layer outputs, enabling selective cross‑layer routing. However, standard Attention Residuals still attend over cumulative hidden states in previous layers, which are highly redundant. We show that this redundancy leads to routing collapse in deeper layers: attention weights become low‑contrast and closer to uniform (max weight \approx0.2), limiting the model's ability to select informative states in previous layers. This raises a key but underexplored design question: what layer‑wise representations should be routed in Attention Residuals? To answer this question, we propose Delta Attention Residuals, which attend over deltas ‑‑ the change introduced by each sublayer (\mathbfv_i = \mathbfh_i+1 ‑ \mathbfh_i) ‑‑ instead of cumulative states. Delta representations are structurally diverse and yield higher‑contrast attention distributions (max weight \approx0.6), enabling more selective and effective routing across layers. This principle applies at both per‑sublayer and block granularity. Across all tested scales (220M‑‑7.6B), Delta Attention Residuals consistently outperform both standard residuals and Attention Residuals, with 1.7‑‑8.2% validation perplexity gains. Delta Attention Residuals also enables converting pretrained checkpoints into Delta Attention Residuals via standard fine‑tuning. Code is available at https://github.com/wdlctc/delta‑attention‑residuals‑code.
Authors:Soumava Paul, Prakhar Kaushik, Alan Yuille
Abstract:
Multiview 3D evaluation assumes that the images being scored are observations of one static 3D scene. This assumption can fail in NVS and sparse‑view reconstruction: inputs or generated outputs may contain artifacts, outlier frames, repeated views, or noise, yet still receive high 3D consistency scores. Existing reference‑based metrics require ground truth, while ground‑truth‑free metrics such as MEt3R depend on learned reconstruction backbones whose failure modes are poorly characterized. We study this reliability problem by comparing neural reconstruction priors with classical geometric verification. We introduce \benchmark, a controlled robustness benchmark for multiview 3D consistency, and a parametric family that decomposes neural metrics into backbone, residual, and aggregation components. This family recovers MEt3R and yields variants up to 3× more robust. Our analysis shows that VGGT, MASt3R, DUSt3R, and Fast3R can hallucinate dense geometry and cross‑view support for unrelated scenes, repeated images, and random noise. We introduce COLMAP‑based metrics that use matches, registration, dense support, and reconstruction failure as failure‑aware consistency signals. On real NVS outputs and a structured human study, these metrics achieve up to 4× higher correlation with human judgments than MEt3R.
Authors:Feiyan Zhou, Luyuan Wang, Shoufa Chen, Zhe Wang, Zhiheng Liu, Yuren Cong, Xiaohui Zhang, Fanny Yang, Belinda Zeng
Abstract:
Modern audio generation predominantly relies on latent‑space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generates high‑fidelity audio directly in raw waveform space without intermediate representations. To overcome the inherent difficulties of modeling high‑dimensional and low‑energy signals, we reshape audio into 2D token grids through waveform patchify and introduce amplitude lifting to align signal scales, enabling stable optimization via direct x‑prediction in flow matching. To capture complex semantic alignment and temporal synchronization, we leverage an automated data pipeline to curate 5 million high‑quality video‑text‑audio triplets, allowing the model to learn fine‑grained acoustic patterns from scratch. Experimental results show that WavFlow achieves competitive performance on the video‑to‑audio benchmark VGGSound (FD_PaSST: 59.98, IS_PANNs: 17.40, DeSync: 0.44) and the text‑to‑audio benchmark AudioCaps (FD_PANNs: 10.63, IS_PANNs: 12.62), matching or exceeding the performance of established latent‑based methods. Our work demonstrates that intermediate compression is not a prerequisite for high‑quality synthesis, offering a simpler and more scalable alternative for multimodal audio generation.
Authors:Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao, Zhenghong Zhou, Hang Hua, Wei Xiong, Jiebo Luo
Abstract:
Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and reference images, and one set of weights covers replacement, removal, style transfer, and reference‑driven insertion. The design is flexible, but it assumes that the user already provides model‑ready text, reference images, and spatial grounding for local edits, which real requests often omit. We present Aurora, an agentic video editing framework that pairs a tool‑augmented vision‑language model (VLM) agent with a unified video diffusion transformer. The VLM agent maps a raw user request to a structured edit plan aligned with the transformer's conditioning channels, thereby resolving textual and visual underspecification before generation. We train the VLM agent with supervised data for complete edit planning and reference‑image selection, together with preference pairs for robust tool use and instruction refinement. We introduce AgentEdit‑Bench to evaluate agent‑enhanced video editing under textual and visual underspecification. Experiments on AgentEdit‑Bench and two existing video editing benchmarks show that Aurora improves over instruction‑only baselines and that the VLM agent transfers to compatible frozen video editing models. Project page: https://yeates.github.io/Aurora‑Page
Authors:Yukang Chen, Luozhou Wang, Wei Huang, Shuai Yang, Bohan Zhang, Yicheng Xiao, Ruihang Chu, Weian Mao, Qixin Hu, Shaoteng Liu, Yuyang Zhao, Huizi Mao, Ying-Cong Chen, Enze Xie, Xiaojuan Qi, Song Han
Abstract:
We present LongLive‑2.0, an NVFP4‑based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing speed and memory bottlenecks. For training, we introduce sequence‑parallel autoregressive (AR) training, instantiated as Balanced SP, which co‑designs the efficient teacher‑forcing layout with SP execution by pairing clean‑history and noisy‑target temporal chunks on each rank, enabling a natural teacher‑forcing mask with SP‑aware chunked VAE encoding. Combined with NVFP4 precision, it reduces GPU memory cost and accelerates GEMM computation during training, the proportion of which increases as video length grows. Moreover, we show that a high‑quality infrastructure and dataset enable a remarkably clean training pipeline. Unlike existing Self‑Forcing series methods that rely on ODE initialization and subsequent distribution matching distillation (DMD), LongLive‑2.0 directly tunes a diffusion model into a long, multi‑shot, interactive auto‑regressive (AR) diffusion model. It can be further converted to real‑time generation (4 to 2 denoising steps) with standalone LoRA weights. For inference on Blackwell GPUs, we enable W4A4 NVFP4 inference, quantize KV cache into NVFP4 for memory savings, and boost end‑to‑end throughput with asynchronous streaming VAE decoding. On non‑Blackwell GPU architectures, we deploy SP inference to match the speed on Blackwell GPUs, while the quantized KV cache can lower inter‑GPU communication of SP. Experiments show up to 2.15x speedup in training, and 1.84x in inference. LongLive‑2.0‑5B achieves 45.7 FPS inference while attaining strong performance on benchmarks. To our knowledge, LongLive‑2.0 is the first NVFP4 training and inference system for long video generation.
Authors:Miguel Farinha, Ronald Clark
Abstract:
We present PIXLRelight, a feed‑forward approach for physically controllable single‑image relighting. Existing methods either provide limited lighting control (e.g. through text or environment maps), accumulate errors when chaining inverse and forward rendering, or require costly per‑image optimization. Our key idea is to bridge physically based rendering (PBR) and learned image synthesis through a shared intrinsic conditioning that can be obtained from either real photographs or PBR renders. At training time, paired multi‑illumination photographs are decomposed into albedo, diffuse shading, and non‑diffuse residuals, which condition the model. At inference time, the same conditioning is computed from a path‑traced render of a coarse 3D reconstruction of the input under user‑specified PBR lights. A transformer‑based neural renderer then applies the target illumination to the source photograph, preserving fine image detail through a per‑pixel affine modulation. PIXLRelight enables arbitrary PBR‑style lighting control, achieves state‑of‑the‑art relighting quality, and runs in under a tenth of a second per image. Code and models are available at https://mlfarinha.github.io/pixl‑relight/.
Authors:Ruiping Liu, Junwei Zheng, Yufan Chen, Di Wen, Shaofang Quan, Chengzhi Wu, Jiaming Zhang, Kailun Yang, Kunyu Peng, Rainer Stiefelhagen
Abstract:
Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial‑temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoMem, the first benchmark for cross‑view memory reasoning over synchronized egocentric and exocentric videos. EgoExoMem contains 2.6K high‑quality MCQs across eight temporal, spatial, and cross‑view QA types. To support dual‑view retrieval, we propose E^2‑Select, a training‑free frame selection method for synchronized ego‑exo videos. It combines relevance‑based budget allocation with per‑view k‑DPP sampling to handle view asymmetry and cross‑view temporal consistency. Experiments show that ego and exo views provide complementary memory cues, while existing MLLMs remain far from solving the benchmark: the best model reaches only 55.3%. E^2‑Select achieves state‑of‑the‑art performance of 58.2% over frame‑selection and RAG‑based memory baselines. Further analysis reveals systematic view‑preference conflicts between question framing and answer grounding, underscoring the novelty and challenge of cross‑view memory reasoning.
Authors:Jinzhuo Liu, Jiangning Zhang, Wencan Jiang, Yabiao Wang, Dingkang Liang, Zhucun Xue, Ran Yi, Yong Liu
Abstract:
Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long‑term inconsistency and memory degradation. Most existing solutions either compress historical frames using predefined strategies or retrieve keyframes based on coarse implicit attention signals, both of which fail to handle evolving prompts with shifting entity references, leading to identity drift, character duplication, and attribute loss. To address this, we propose IAMFlow, a training‑free identity‑aware memory framework that explicitly models and tracks persistent entity identities, enabling consistent generation across prompt transitions. Specifically, an LLM extracts entities with visual attributes from each prompt and assigns unique global IDs for identity‑aware memory, while a VLM asynchronously verifies and refines attributes from rendered frames, enabling explicit entity tracking in place of implicit similarity‑based matching. To keep the proposed framework computationally practical, we design a systematic inference acceleration pipeline, including asynchronous visual verification, adaptive prompt transition, and model quantization, which achieves faster generation than existing baselines. Furthermore, we introduce NarraStream‑Bench, a benchmark for narrative streaming video generation that features 324 multi‑prompt scripts spanning six dimensions and a three‑dimensional evaluation protocol that integrates both traditional metrics and multimodal large language model‑based assessments. Extensive experiments show that IAMFlow, despite being training‑free, achieves the best overall performance on NarraStream‑Bench, outperforming the strongest baseline by 2.56 points, while achieving a 1.39× speedup over the most efficient baseline in the 60‑second multi‑prompt setting.
Authors:Komal Kumar, Ankan Deria, Abhishek Basu, Fahad Shamshad, Hisham Cholakkal, Karthik Nandakumar
Abstract:
Diffusion models have been widely studied for removing unsafe content learned during pre‑training. Existing methods require expensive supervised data, either unsafe‑text paired with safe‑image groundtruth or negative/positive image pairs, making them impractical to scale. Furthermore, offline reinforcement learning and supervised fine‑tuning approaches that generate synthetic data offline suffer from catastrophic forgetting, degrading generation quality. We propose a novel online reinforcement learning framework that addresses both data scarcity and model degradation through post‑training with Group Relative Policy Optimization (GRPO) on both negative and positive text prompts. To eliminate the need for fine‑tuning specialized safe/unsafe reward models, we introduce a steering reward mechanism that exploits an inherent property of CLIP embeddings: steering text representations toward positive safety directions and away from negative ones in the embedding space. Our online‑policy approach enables the model to learn from diverse prompts, including explicit unsafe content, without catastrophic forgetting. Extensive experiments demonstrate that our method reduces inappropriate content to 18.07% (vs. 48.9% for SD v1.4) and nudity detections to 15 (vs. 646 baseline) while improving compositional generation quality from 42.08% to 47.83% on GenEval. Remarkably, these safety gains generalize to out‑of‑domain unsafe prompts across seven harm categories, achieving state‑of‑the‑art performance without supervised paired data or reward tuning. Github: https://github.com/MAXNORM8650/SafeDiffusion‑R1.
Authors:Songsong Yu, Yuxin Chen, Ying Shan, Yanwei Li
Abstract:
Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently optimize understanding via sparse text signals and generation through dense pixel objectives. Such a decoupled strategy yields misaligned representation spaces, isolating visual understanding from generation and hindering their mutual reinforcement. This work presents the first systematic investigation into generative post‑training, where we formulate hierarchical visual tasks as generative proxies to bridge the isolation in UMMs. Our empirical investigation reveals that high‑level semantic tasks, particularly image segmentation, serve as optimal proxies. Unlike low‑level tasks that distract models with texture details, segmentation provides structural semantics that significantly enhance both vision‑centric perception and generative layout fidelity. Building upon these insights, we introduce Semantic Generative Tuning (SGT), a novel paradigm that leverages segmentation as a generative proxy to align and synergize multimodal capabilities. Mechanistic analyses further demonstrate that SGT fundamentally improves feature linear separability and optimizes visual‑textual attention allocation pattern. Extensive evaluations show that SGT consistently improves both multimodal comprehension and generative fidelity across mainstream benchmarks. Our code is available on the https://song2yu.github.io/SGT/.
Authors:Edwin Arkel Rios, Augusto Christian Surya, Oswin Gosal, Fernando Mikael, Mary Madeline Nicole, Kisoon Jang, Bo-Cheng Lai, Min-Chun Hu
Abstract:
Prior work on fine‑grained image recognition (FGIR) has established the importance of the backbone selection, but has neglected the accuracy‑vs‑cost trade‑offs under different training and evaluation settings. In this work we conduct a large‑scale study with over 2000 experiments across 6 training and evaluation settings, 9 pretrained backbones, and 17 datasets. Preliminary observations on the effectiveness of data augmentation for fine‑grained training motivate us to extend Counterfactual Attention Learning (CAL), a state‑of‑the‑art method based on data‑aware cropping and masking augmentations, with cross‑image discriminative region mixing augmentation. We also propose an efficient evaluation‑only variant that maintains competitive accuracy while reducing inference costs by forfeiting the forward pass on discriminative crops that is normally used by CAL and similar FGIR methods. Our results show that data‑aware augmentations during training only can enable a model to achieve excellent accuracy even without crops, significantly reducing inference costs. To support future research we share our code and checkpoints at: \urlhttps://github.com/arkel23/FGIR‑Backbones
Authors:Fengyi Fu, Mengqi Huang, Shaojin Wu, Yunsheng Jiang, Yufei Huo, Hao Li, Yinghang Song, Fei Ding, Jianzhu Guo, Qian He, Zheren Fu, Zhendong Mao, Yongdong Zhang
Abstract:
We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text‑image‑dominant designs, Lance explores a practical paradigm for unified multimodal modeling via collaborative multi‑task training. It is grounded in two core principles: unified context modeling and decoupled capability pathways. Specifically, Lance is trained from scratch and employs a dual‑stream mixture‑of‑experts architecture on shared interleaved multimodal sequences, enabling joint context learning while decoupling the pathways for understanding and generation. We further introduce modality‑aware rotary positional encoding to mitigate interference among heterogeneous visual tokens and boost cross‑task alignment. During training, Lance adopts a staged multi‑task training paradigm with capability‑oriented objectives and adaptive data scheduling to strengthen both semantic comprehension and visual generation performance. Experimental results demonstrate that Lance substantially outperforms existing open‑source unified models in image and video generation, while retaining strong multimodal understanding capabilities. The homepage is available at https://lance‑project.github.io.
Authors:Arslan Artykov, Tom Ravaud, Nicolás Violante-Grezzi, Vincent Lepetit
Abstract:
Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex video setups or cues such as long‑term point tracking or wide‑baseline matching, but are frequently brittle under severe occlusions, rapid camera ego‑motion, or weak local features. Learning‑based methods, meanwhile, struggle to generalize beyond their training categories. We propose a category‑agnostic optimization framework that treats articulated object understanding as a primitive‑fitting problem. Geometric primitives serve as a proxy representation that avoids the pitfalls of unstable point tracks; a novel mechanism organizes them into coherent parts constrained by revolute and prismatic joints. Our formulation jointly optimizes part segmentation and joint parameters, recovering complex kinematics from a single casually captured video. A visibility‑aware procedure handles partial observations and occlusions inherent to real‑world data. We also propose the AiP‑synth and AiP‑real benchmarks, featuring significant camera motion and heavy occlusions, and outperform existing methods. Project page: https://aartykov.github.io/Articulation‑in‑Prime/
Authors:Dongyao Zhu, Zhen Wang, Xi Xiao, Han Jiang, Saeed Vahidian, Wei-Lun Chao, Tanya Berger-Wolf, Yu Su, Raju Vatsavai, Jianyang Gu
Abstract:
Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of these latent tokens at inference remains ambiguous. We show that replacing latent tokens with random noise or removing them completely causes little performance degradation across spatial reasoning benchmarks. Reinforcement learning further diminishes the latent generation behavior after post‑training. These observations raise a central question: Is latent visual reasoning still meaningful? We argue that its value should be measured by how effectively latent tokens guide learning, rather than whether they persist as an inference‑time format. Our analysis shows that latent reasoning is unevenly favorable across question types, yet hard task‑level routing for applying latent generation is brittle. Motivated by these findings, we propose an attention‑based reward that encourages generated latent tokens to interact with later text tokens during RL. This reward promotes latent utilization when the latent mode is activated while preserving the flexibility to use pure‑text reasoning. Experiments show that our method improves performance across perception and visual reasoning benchmarks, even when latent tokens are rarely generated after post‑training. Our results highlight that, without explicit expression at inference, latent visual reasoning can shape better visual grounding and more accurate textual reasoning in silence. Our code and trained models are publicly available at \hrefhttps://github.com/ddydyd32/silent‑lvr/tree/masterGitHub and \hrefhttps://huggingface.co/collections/cornuHGF/silent‑lvrHugging Face.
Authors:Wencan Jiang, Jiangning Zhang, Jianbiao Mei, Jinzhuo Liu, Yu Yang, Xiaobin Hu, Zhucun Xue, Yong Liu, Dacheng Tao
Abstract:
Long‑horizon multimodal agents in open‑world games must stay goal‑directed across many low‑level interactions under tight token and latency budgets. Existing approaches often trade off costly per‑step reasoning against reactive execution that can drift, repeat failures, and recover poorly. Our key idea is to reuse strategic reasoning across locally stable segments and reinvoke it at event boundaries. We present SPIKE, an adaptive dual controller framework for cost‑efficient long‑horizon game control. Its Strategic Controller performs low‑frequency global planning, failure analysis, and recovery, while its Reactive Controller handles fast local execution under a strict token budget. An Event Trigger monitors visual change, task progress, repeated actions, and failure signals to decide when control should stay reactive or escalate to strategic reasoning. Hierarchical Memory separates short‑term experience reuse in the State‑Action Memory Bank (SA‑MB) from structured evidence in the State Action Knowledge Graph (SA‑KG), allowing each controller to retrieve the context it needs. This design reuses strategic proposals over multiple reactive steps, supports local override when plans become stale, and reserves expensive reasoning for moments where extra deliberation is useful. On the Lite‑100 split of StarDojo, SPIKE improves Lite‑100 success rate (SR) by 5.0 percentage points (38.5% relative) over the strongest Lite‑100 baseline and Budgeted SR by 9.3 points (75.6% relative) over the strongest budgeted baseline. It also reduces token consumption by 54.9% and latency by 40.8%. Ablations show that event triggering, reactive override, and heterogeneous memory each contribute to success and recovery, supporting selective reasoning rather than reasoning at every step.
Authors:Wei Wang, Yuqian Yuan, Tianwei Lin, Wenqiao Zhang, Siliang Tang, Jun Xiao, Yueting Zhuang
Abstract:
Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single‑view perception and reason consistently about objects, visibility, geometry, and interactions across multiple viewpoints. However, progress in cross‑view reasoning remains limited by three major gaps: the scarcity of large‑scale well‑annotated training data, the lack of comprehensive benchmarks for systematic evaluation, and the absence of explicit alignment mechanisms that establish object‑level consistency across views. To address these gaps, we thoroughly develop CrossView Suite across three coordinated components: CrossViewSet, CrossViewBench, and CrossViewer. Firstly, we introduce a multi‑agent data engine to meticulously curate a large‑scale, high‑quality cross‑view instruction dataset, termed CrossViewSet, covering 17 fine‑grained task types with 1.6M samples. Second, we meticulously create a scene‑disjoint CrossViewBench to comprehensively assess the cross‑view spatial understanding capability of an MLLM, evaluating it across various aspects. Finally, we propose CrossViewer, a progressive three‑stage framework for cross‑view spatial reasoning in MLLMs, following a Perception ‑> Alignment ‑> Reasoning paradigm. Our method equips an adaptive spatial region tokenizer to capture fine‑grained object representations, and then aligns the multi‑view objects explicitly, and thus fuses aligned features for boosting the cross‑view inference capacity for MLLMs. Extensive experiments and analyses show that large‑scale training data, systematic evaluation, and explicit cross‑view alignment are all critical for advancing MLLMs from single‑view perception toward real‑world spatial intelligence. The project page is available at https://github.com/Thinkirin/Crossview‑Suite.
Authors:Ziyu Wei, Luting Wang, Chen Gao, Li Wen, Si Liu
Abstract:
Most existing vision‑language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft robotic arms offer an appealing alternative due to their deformability, but confront challenges such as unreliable proprioception and distributed low‑level actuation. To investigate these challenges, we introduce \ManiSoft, a benchmark for vision‑language manipulation with soft arms. ManiSoft features a tailored simulator that couples realistic soft‑body dynamics with contact‑rich interactions via an elastic force constraint. On this basis, ManiSoft defines four tasks, each highlighting distinct aspects of deformable control, from basic end‑effector coordination to obstacle avoidance. To support policy training and evaluation, \ManiSoft includes an automated pipeline that generates 6,300 diverse scenes and corresponding expert trajectories. To produce high‑quality trajectories at scale, we first employ a high‑level planner to decompose each task into a sequence of waypoints, followed by a low‑level reinforcement learning policy that generates torque commands to track waypoints. Benchmarking three representative policy models shows relatively promising results in clean scenes but substantial performance drop under randomization. Visualization analysis indicates that failures stem primarily from inaccurate visual estimation of proprioceptive state and limited exploitation of deformability for adaptive obstacle avoiding. We anticipate ManiSoft to serve as a valuable testbed, bridging the gap between rigid and soft arms in the context of vision‑language manipulation. Out codes and datasets are released at https://buaa‑colalab.github.io/ManiSoft.
Authors:Zhilin Zhu, Yabin Wang, Zhiheng Ma, Yaguang Song, Yaowei Wang, Xiaopeng Hong
Abstract:
Continual Test‑Time Adaptation (CTTA) aims to empower perception systems to handle dynamic distribution shifts encountered after deployment. Existing methods predominantly follow a backward‑alignment paradigm, which rigidly aligns incoming data with supervisory surrogates derived from the source domain. Consequently, they struggle with unreliable supervision and evolving distribution shifts. To overcome these limitations, we introduce a novel forward‑facilitation paradigm through a method termed Dynamic Style Bridging. Prior to deployment, we construct a compact knowledge base of generated class exemplars. During test time, to mitigate inherent generative bias and adapt these proxies to incoming data, we propose a multi‑level bridging mechanism. This mechanism dynamically injects the proxies with incoming data styles at the input, statistical, and representation levels, while preserving the original semantics of the proxies. These high‑fidelity proxies are then used to provide reliable, on‑demand supervisory signals, enabling stable adaptation under continual shifts. Extensive experiments across standard CTTA benchmarks demonstrate that our method achieves consistent and substantial improvements over recent state‑of‑the‑art approaches. Code is available at \hrefhttps://github.com/z1358/DAS.
Authors:Yihang Wu, Yihang Sun, Shaofeng Zhang, Zuxuan Wu, Junchi Yan, Xiaosong Jia, Yu-gang Jiang
Abstract:
Transformer‑based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS‑LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Plücker rays) into a shared feature space. Since Plücker rays naturally carry lattice‑like spatial structure, these designs can make the spatial bias interfere with appearance representation and degrade rendering fidelity. To this end, we propose to decouple the representation of feedforward NVS transformers into separate semantic and spatial tokens. The decoupled design keeps semantic and spatial information explicit in their branches while preserving cross‑branch interaction through shared attention routing. Built on this design, we introduce optional categorized supervision and bidirectional modulation: the former provides branch‑specific training signals, while the latter improves interaction between the two branches. Notably, the base decoupled design introduces virtually zero additional inference latency due to its architectural design. The proposed designs achieve consistent improvements, demonstrating effectiveness across decoder‑only and encoder‑decoder feedforward NVS models.
Authors:Ruixiang Zhao, Jie Yang, Zijie Xin, Tianyi Wang, Fengyun Rao, Jing LYU, Xirong Li
Abstract:
Omni‑proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio‑visual streams, is an emerging capability of omni‑modal large language models. Existing benchmarks fall short in three key aspects: they rely primarily on visual signals, adopt polling or fixed‑timestamp protocols instead of true proactive evaluation, and cover only a limited range of tasks, preventing reliable assessment and differentiation of omni‑proactive streaming models. We present OmniPro, the first benchmark to jointly evaluate omni‑modal perception, proactive responding, and diverse video understanding tasks. It comprises 2,700 human‑verified samples spanning 9 sub‑tasks and 3 cognitive levels, covering 6 basic video understanding capabilities. Notably, 84% of samples require audio signals (speech or non‑speech), and each sample is annotated with modality‑isolation labels to enable fine‑grained multimodal analysis. We further introduce a dual‑mode evaluation protocol: Probe mode assesses content understanding by querying the model before and after each ground‑truth trigger, while Online mode evaluates full proactive ability by requiring models to autonomously decide when to respond in streaming input. Evaluating 11 representative models reveals three key findings: (1) audio provides consistent gains but with highly variable utilization across models, (2) performance degrades significantly over time, indicating limited long‑horizon robustness, and (3) non‑speech audio perception remains the weakest dimension.
Authors:Huajian Zeng, Chaohua Yao, Yuantai Zhang, Jiaqi Yang, Rolandos Alexandros Potamias, Xingxing Zuo
Abstract:
Recovering world space 4D motion of two interacting hands from egocentric video is a fundamental capability for supervising robot policy learning, where wrist trajectories track the end‑effector and finger articulations specify the grasp pose. Two major challenges arise in this setting: hands frequently leave the camera view for extended periods due to head motion, and persistent hand‑object interactions cause severe occlusions of one or both hands. Existing methods uniformly condition on noisy hand motion observations without accounting for their per‑frame reliability, leading to substantial performance degradation. Our key insight is that accurate world space hand motion estimation is tightly coupled with the quality of per‑frame hand observations. To this end, we decompose the quality of hand motion observations extracted from an off‑the‑shelf hand pose estimator into four channels: wrist global translation and finger articulations for both hands. We propose StableHand, a quality‑aware flow‑matching framework conditioned on these four‑channel quality signals, which are predicted by a learned quality network. We naturally incorporate the quality signals into the flow‑matching process through a per‑channel forward schedule, a quality‑adjusted velocity target, AdaLN modulation of the DiT denoiser, and a quality‑aware ODE initialization. This unified generative process preserves high‑quality observations while reconstructing unreliable ones using a learned bimanual motion prior. Experiments on HOT3D and ARCTIC, two egocentric benchmarks featuring long missing‑hand spans and persistent hand‑object occlusions, show that StableHand achieves state‑of‑the‑art performance across all reported metrics, reducing W‑MPJPE by 20‑25% compared to the strongest baseline, with the largest gains on heavily occluded ARCTIC sequences.
Authors:Jingyun Fu, Zhiyu Xiang, Na Zhao
Abstract:
Due to the difficulty of obtaining ground‑truth data for 4D radar scene flow estimation, previous methods typically rely on either self‑supervised losses or cross‑modal supervision using 3D LiDAR data, 2D images, and odometry. However, self‑supervised approaches often yield suboptimal results due to radar's inherently low‑fidelity measurements, while existing cross‑modal supervised methods introduce complex multi‑task architecture and require costly LiDAR sensors to generate pseudo radar scene flow labels from pretrained 3D tracking models. To overcome these limitations, we propose a task‑specific iterative framework for weakly supervised radar scene flow learning, using only images and odometry for auxiliary supervision during training. Specially, we establish two novel instance‑aware self‑supervised losses by exploiting off‑the‑shelf 2D tracking and segmentation algorithms to obtain tracked instance masks, which are back‑projected into 3D space to provide instance‑level semantic guidance; for static regions, we integrate vehicle odometry with radar's intrinsic motion cues to construct a rigid static loss. Extensive experiments on the real‑world View‑of‑Delft (VoD) dataset demonstrate that our method not only surpasses state‑of‑the‑art cross‑modal supervised approaches that rely on 3D multi‑object tracking on dense LiDAR point clouds but also outperforms existing fully supervised scene flow estimation methods. The code is open‑sourced at \hrefhttps://github.com/FuJingyun/IterFlowhttps://github.com/FuJingyun/IterFlow.
Authors:Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Hao Shi, Yi Zhou, M. Saquib Sarfraz, Danda Pani Paudel, Luc Van Gool
Abstract:
Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study this problem through multi‑robot cooperative dynamic spatial reasoning, where a model must answer spatial, temporal, visibility, and coordination questions by integrating synchronized egocentric videos from a team of moving robots. To support this setting, we introduce CoopSR, the first benchmark for this task, together with EgoTeam, a multi‑robot egocentric QA dataset. EgoTeam contains 114,227 QA pairs spanning 19 question types, four difficulty tiers, and three team sizes in Habitat and iGibson, along with a real‑world test set of around 2,326 QAs collected using two quadruped robots. We further propose SP‑CoR (Spectral and Physics‑Informed Cooperative Reasoner), an MLLM framework for fine‑grained cooperative spatial reasoning. SP‑CoR combines dynamics‑aware multi‑robot frame sampling, spectral‑ and physics‑guided view fusion, and physics‑aligned prompt distillation, enabling the model to benefit from privileged robot‑pose supervision during training while requiring only egocentric videos at test time. Across 22 MLLM baselines, SP‑CoR consistently improves cooperative reasoning, outperforming the strongest fine‑tuned baseline by +3.87% on Habitat and +7.12% on iGibson. It also shows stronger generalization to unseen team sizes and real‑world robot tests. Code can be found at https://github.com/KPeng9510/seeing‑together.git.
Authors:Yuxiang Feng, Juncheng Wang, Chao Xu, Yijie Qian, Huihan Wang, Wenlong Hou, Yang Liu, Baigui Sun, Yong Liu, Shujun Wang
Abstract:
Video generation models produce visually compelling results but systematically violate physical commonsense ‑‑ on VideoPhy‑2, the best model achieves only 32.6% joint accuracy. We identify a specification bottleneck: text prompts are lossy compression of the physical world, omitting the parameters that fully determine dynamics, and no amount of model scaling can recover what was never specified. From this diagnosis we derive three properties that physics conditioning must satisfy ‑‑ sufficiency, dynamism, and verifiability ‑‑ and show that no existing approach satisfies all three. We present NEWTON, in which video generation is demoted from the system output to one action inside an agent's toolbox: a learned planner orchestrates physics‑aware tools (keyframe generation, scientific computation, prompt refinement) to construct rich conditioning, and a verifier closes the loop for iterative re‑planning. The planner is the sole trainable component, optimized on‑policy via Flow‑GRPO inside the live multi‑turn loop. On VideoPhy‑2, NEWTON improves joint accuracy from 21.4% to 29.7% on LTX‑Video and from 30.7% to 37.4% on Veo‑3.1, without modifying either generator. Our project page: https://Newton026.github.io/newton
Authors:Beizhen Zhao, Yifan Zhou, Gaochao Song, Ziran Yin, Hao Wang
Abstract:
While 3D Gaussian Splatting (3DGS) has revolutionized real‑time photorealistic view synthesis, its fundamental reliance on symmetric Gaussian distributions introduces visual artifacts that hinder accurate spatial data exploration. Specifically, symmetric kernels struggle to capture shape and color discontinuities , which cause blurriness and primitive redundancy that mislead human perception during visual analysis. To address these visualization barriers, we introduce 3D Skew Gaussian Splatting (3DSGS), a novel framework that significantly enhances the structural fidelity and compactness of explicit scene representations. Our key insight lies in extending the standard primitive to a general Skew Gaussian counterpart. This generalized primitive inherits the highly efficient rasterization properties of standard Gaussians while gaining intrinsic asymmetric modeling capabilities. We couple this with an enhanced opacity representation to better handle complex transparency, alongside a depth‑aware densification strategy that intelligently manages primitive allocation. Furthermore, to make these advancements actionable for real‑world visual analytics, we re‑derive the CUDA rasterization pipeline to universally support both symmetric and skew Gaussians, integrating it into a decoupled, free‑camera interactive visualization engine. Extensive experiments demonstrate that 3DSGS achieves superior rendering quality and structural compactness, particularly in regions with intricate details, while maintaining the real‑time frame rates necessary for fluid interactive exploration. Supplementary derivations and visual results are available at https://3d‑skew‑gs.github.io/.
Authors:Luca Hagen, Johanna P. Müller, Weitong Zhang, Mengyun Qiao, Bernhard Kainz
Abstract:
Small vision‑language models (2‑8B) are well‑suited for clin‑ ical deployment due to privacy constraints, limited connectivity, and low‑latency requirements favouring on‑device or on‑premise inference. However, their limited capacity exacerbates the generation of plausible but incorrect outputs. We extend game‑theoretic decoding, previously restricted to text‑only, closed‑ended NLP tasks, to vision‑language mod‑ els for open‑ended Medical VQA. We introduce a semantically aware Wasserstein stopping criterion that replaces lexical order matching, en‑ abling convergence based on semantic consensus among near‑synonymous candidate answers and avoiding unnecessary iterations caused by clini‑ cally equivalent ranking swaps. On VQA‑RAD and PathVQA, we ob‑ tain consistent, statistically significant improvements over greedy and discriminative baselines. On VQA‑RAD, we improve Qwen3‑VL‑2B by +3.5 percentage points (p < 0.01), surpassing the greedy 4B model, with similar trends at larger scales. On PathVQA, Gemma‑3‑4B with BDG matches MedGemma‑4B under greedy decoding despite no domain‑ specific fine‑tuning. At accuracy parity with classic BDG, the Wasser‑ stein criterion reduces average convergence iterations by approximately 20%, improving inference efficiency while preserving the game‑theoretic equilibrium behaviour. Code is available at https://github.com/luca‑hagen/ Wasserstein‑BDG‑medical‑VQA.
Authors:Yiyang Fu, Chubin Zhang, Shukai Gong, Yufan Deng, Kaiwei Sun, Qiyang Min, Qibin Hou, Yansong Tang, Jianan Wang, Daquan Zhou
Abstract:
It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision‑Language‑Action (VLA) models when encountering unseen real‑world visual disturbances, particularly under imperfect visual conditions. In this work, we conduct a systematic study based on recent state‑of‑the‑art VLA models and reveal a significant performance drop when visual disturbances absent from the training data are introduced. To mitigate this issue, we propose a lightweight adapter module grounded in information theory, termed the Information Bottleneck Adapter (IB‑Adapter), which selectively filters potential noise from visual inputs. Without requiring any extra data or augmentation strategies, IB‑Adapter consistently improves over the baseline by an average of 30%, while adding fewer than 10M parameters, demonstrating notable efficiency and effectiveness. Furthermore, even with a 14x smaller backbone (0.5B parameters) and no pre‑training on the Open X‑Embodiment dataset, our model StableVLA achieves robustness competitive with 7B‑scale state‑of‑the‑art VLAs. With negligible parameter overhead (<10M), our approach maintains accuracy on long‑horizon tasks and surpasses OpenPi under both synthetic and physical visual corruptions.
Authors:Longtao Jiang, Jianmin Bao, Zhendong Wang, Xin Tao, Pengfei Wan, Zhihui Li, Xiaojun Chang
Abstract:
Normalizing flows (NFs) provide exact likelihoods and deterministic invertible sampling, but have historically lagged behind diffusion models for large‑scale image generation. We identify a key obstacle: NFs are required to learn a single invertible transport over the full ambient space, making them highly sensitive to high‑dimensional representations. This leads to a semantic‑capacity mismatch in modern visual representation spaces, where semantic information is compact but encoded in overcomplete features. We propose SRC‑Flow, which introduces a Semantic Representation Compressor (SRC) to compact high‑dimensional RAE features into a low‑dimensional semantic space before flow modeling and preserve reconstruction through the frozen RAE decoder. This compact space reduces the modeling burden of NFs and enables effective likelihood‑based generation in semantic representation space. We further adopt constant noise regularization tailored to the fixed unconditional bijection learned by flows. On ImageNet 256 × 256 and 512 × 512, SRC‑Flow achieves state‑of‑the‑art generation quality among normalizing flow methods, with gFID scores of 1.65 and 2.07 under classifier‑free guidance, while retaining exact likelihood computation in the compact semantic representation space and deterministic invertible sampling at the flow level. Codes and models will be available at https://github.com/longtaojiang/SRC‑Flow.
Authors:Ji Shi, Xianghua Ying, Bowei Xing, Ruohao Guo, Wenzhen Yue
Abstract:
3D Gaussian Splatting (3DGS) enables real‑time novel view synthesis with high visual quality. However, existing methods struggle with semi‑transparent specular surfaces that exhibit both complex reflections and clear transmission, often producing blurry reflections or overly occluded transmission. To address this, we present RT‑Splatting, a framework that disentangles each Gaussian's geometric occupancy from its optical opacity. This factorization yields a unified surface‑volume scene representation with a single set of Gaussian primitives. Our hybrid renderer interprets this representation both as a surface to capture high‑frequency reflections and as a volume to preserve clear transmission. To mitigate the ambiguity in jointly optimizing reflection and transmission, we introduce Specular‑Aware Gradient Gating, which suppresses misleading gradients from highly specular regions into the transmission branch, effectively reducing distracting floaters. Experiments on challenging semi‑transparent scenes show that RT‑Splatting achieves state‑of‑the‑art performance, delivering high‑fidelity reflections and clear transmission with real‑time rendering. Moreover, our factorization naturally enables flexible scene editing. The project page is available at https://sjj118.github.io/RT‑Splatting.
Authors:Zeyu Chen, Jie Li, Kai Han
Abstract:
Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross‑modal information discrepancies and data scarcity, leading to suboptimal alignment spaces that overlook modality‑unique features. We propose CodeBind, a framework that optimizes multimodal representation spaces through a modality‑shared‑specific codebook design. By incrementally aligning target and bridging modalities, CodeBind bypasses the need for fully paired data. Unlike traditional hard alignment, CodeBind decomposes features into shared components for semantic consistency and specific components for modality‑unique details. This design utilizes a compositional vector quantization scheme, where a shared codebook bridges modality gaps and modality‑specific codebooks mitigate representation bias by preventing dominant modalities from overshadowing others. Validated across nine modalities (text, image, video, audio, depth, thermal, tactile, 3D point cloud, EEG), CodeBind achieves state‑of‑the‑art performance in multimodal classification and retrieval tasks.
Authors:X. Feng, J. Zhu, M. Wu, C. Chen, F. Mao, H. Guo, J. Wu, X. Chu, K. Huang
Abstract:
Without incurring significant computational overhead, train‑free long video generation aims to enable foundation video generation models to produce longer videos. Frame‑level autoregressive frameworks, e.g., FIFO‑diffusion, offer the advantage of generating infinitely long videos with constant memory consumption. However, the mismatch between training and inference, coupled with the challenge of maintaining long‑term consistency, limits the effective utilization of foundation models. To mitigate these concerns, we propose MIGA, a novel infinite‑frame long video generation method. Firstly, we propose an effective two‑stage alignment mechanism that mitigates the training‑inference gap by reducing the excessive noise span fed to the model. We then introduce an innovative dual consistency enhancement mechanism, where the self‑reflection approach corrects early high‑noise frames and the long‑range frame guidance approach leverages later low‑noise frames with broad coverage to steer generation, jointly improving temporal consistency. Extensive experiments on VBench and NarrLV demonstrate the state‑of‑the‑art performance of MIGA. Our project page is available at https://xiaokunfeng.github.io/miga_homepage/.
Authors:Md Hasan, Nyvenn Castro, Daiqi Liu, Lukas Mulzer, Jana Hutter, Jonghye Woo, Moritz Zaiss, Andreas Maier, Paula A. Perez-Toro
Abstract:
Real‑time magnetic resonance imaging (rtMRI) of speech production enables non‑invasive visualization of dynamic vocal‑tract motion and is valuable for speech science and clinical assessment. However, rtMRI is fundamentally constrained by trade‑offs among spatial resolution, temporal resolution, and acquisition speed, often leading to undersampled k‑space measurements and degraded reconstructions. We propose SIREM, a speech‑informed MRI reconstruction framework that uses synchronized speech as a cross‑modal prior. The central idea is that vocal‑tract configurations during speech are correlated with the produced acoustics, making part of the image content predictable from audio. SIREM models each frame as a fusion of an audio‑driven component and an MRI‑driven component through a spatial weighting map. The audio branch predicts articulator‑related structure from speech, while the MRI branch reconstructs complementary content from measured k‑space data. We further introduce a learnable soft weighting profile over spiral arms, enabling a differentiable study of how k‑space arm usage interacts with speech‑informed fusion. This yields a unified multimodal formulation that combines audio‑driven prediction, MRI reconstruction, and sampling adaptation. We evaluate SIREM on the USC speech rtMRI benchmark against standard baselines, including gridding, wavelet‑based compressed sensing, and total variation. SIREM introduces a speech‑informed reconstruction paradigm that operates in a substantially higher‑throughput regime than iterative methods while preserving anatomically plausible vocal‑tract structure. These results establish an initial benchmark for multimodal speech‑informed rtMRI reconstruction and highlight the potential of synchronized speech as an auxiliary prior for fast reconstruction. The source code is available at https://github.com/mdhasanai/SIREM
Authors:Itai Lang, Dongwei Lyu, Dale Decatur, Rana Hanocka
Abstract:
Finding correspondences is a fundamental and extensively researched problem in computer vision and graphics. In this work, we examine the underexplored task of estimating segmentation‑to‑segmentation correspondence between images in the wild and untextured 3D shapes. This task is highly challenging due to substantial differences in appearance, geometry, and viewpoint. Our approach bridges the cross‑modality gap by linking pixels in the image segment to vertices in the corresponding semantic part of the 3D shape. To achieve this, we first distill deep visual features from a 2D vision model onto the 3D shape surface, allowing for the computation of feature similarity between image pixels and shape vertices. Then, we identify Best Segmentation Buddies, vertices whose most similar image pixel lies within the image segmentation region, enabling the reliable discovery of vertices in semantically corresponding shape parts. Finally, we leverage distilled 3D features from the 2D image segmentation model to segment the shape directly in 3D, bootstrapping the correspondence process. We demonstrate the generality and robustness of our approach across a wide range of image‑shape pairs, showcasing accurate and semantically meaningful correspondences. Our project page is at https://threedle.github.io/bsb/.
Authors:Quan Zhang, Zeqiang Cai, Peiming Zhao, Jingze Wu, Cailun Wu, Hongbo Chen, Jianhuang Lai
Abstract:
Aerial‑Ground Person Re‑Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view‑invariant paradigm, aligning shared features across views to achieve robustness. However, view‑invariant inherently enforces part‑level alignment, which ignores view‑specific cues and discriminative identity information. To this end, this work proposes ViSA (View‑aware Semantic Alignment), a view‑aware framework that achieves cross‑view semantic consistency containing an Expert‑driven Token Generation Module (ETGM) and a Dual‑branch Local Fusion Module (DLFM). Technically, the former constructs a set of view‑aware experts to generate adaptive semantic queries that perceive viewpoint‑specific patterns, while the latter leverages graph reasoning to extract and align local regions responsive to different experts. Extensive experiments on three AGPReID benchmarks including AG‑ReID.v2, CARGO and LAGPeR demonstrate that ViSA consistently achieves superior performance, with a notable 10.06% mAP improvement on the challenging CARGO cross‑view protocol. The code is available at \hrefhttps://github.com/Cat‑Zero/ViSAhttps://github.com/Cat‑Zero/ViSA.
Authors:Haoyu Zhang, Qiaohui Chu, Yisen Feng, Meng Liu, Weili Guan, Yaowei Wang, Liqiang Nie
Abstract:
This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participants must answer 185 closed‑form questions over the CASTLE 2024 dataset. In contrast to prior single‑video egocentric benchmarks, CASTLE requires reasoning over four days of activity, 15 synchronized perspectives, official transcripts, and multiple auxiliary modalities, including personal photos, auxiliary videos, gaze, thermal imagery, and heartrate measurements. MARS therefore treats the task as an agentic evidence‑selection problem over multimodal sources rather than a purely text‑only pipeline. MARS first follows the official CASTLE directory organization to build evidence memories from two primary sources, videos and transcripts, and four auxiliary sources, gaze, heartrate, photos, and thermal imagery. Long videos are converted into captions and DeepSeek‑based summaries only because CASTLE videos are too long to fit directly into the model context for every question; this step compresses temporal evidence while keeping photos and other auxiliary media available as source‑specific evidence. At inference time, a GPT‑5.4 decision agent repeatedly chooses whether to continue reasoning, request a specific missing modality, produce an answer, or fall back to a random option when the evidence remains insufficient. The resulting system achieved second place on the final CASTLE Challenge leaderboard. Our codes are available at https://github.com/Hyu‑Zhang/MARS.
Authors:Yiwei Guo, Shaobin Zhuang, Zhipeng Huang, Canmiao Fu, Chen Li, Jing Lyu, Yali Wang
Abstract:
Building a unified visual tokenizer is essential for bridging the gap between visual understanding and generation. Yet existing approaches struggle with the inherent conflict between these tasks, as a single token space is forced to support both high‑level semantic abstraction and low‑level pixel reconstruction. We propose WinTok, a concise hybrid tokenizer that achieves a win‑win performance by explicitly decoupling the two objectives. WinTok supplements pixel tokens with a set of learnable semantic tokens, effectively mitigating cross‑task interference without incurring the computational overhead of dual tokenizers. To further enhance understanding capability, we introduce an asymmetric token distillation mechanism: the semantic tokens are guided by pretrained semantic embeddings from any visual foundation model, enabling them to inherit strong discriminative power while maintaining flexibility. Across 10 challenging benchmarks, WinTok delivers consistent improvements in reconstruction, understanding, and generation. Trained on only 50M open‑source data, WinTok surpasses the strong baseline UniTok by 11.2% in classification accuracy and achieves a competitive reconstruction rFID of 0.41, despite using substantially less training data. Code is released at https://github.com/markywg/WinTok.
Authors:ZhiYuan Feng, Yu Deng, Ruichuan An, Zhenhua Liu, Qixiu Li, Keming Wu, Zhiying Du, Weijie Wang, Haoxiao Wang, Shuang Chen, Sicheng Xu, Yaobo Liang, Jiaolong Yang, Baining Guo
Abstract:
In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such requests require agents to identify task‑relevant entities, recover intended task conditions, and resolve ordering constraints from the surrounding scene context. We formalize this capability as full‑scene household reasoning: given a complete household scene and a situated household request, an agent must infer executable task structure before producing a grounded skill‑level action sequence. This setting is challenging because complete household scenes contain substantial task‑irrelevant information, making direct complete‑scene prompting inefficient and error‑prone. In practical deployment, this challenge is further amplified by privacy and local compute constraints, which favor compact open‑weight models with limited long‑context reasoning ability. We propose TaskGround, a training‑free and model‑agnostic Ground‑Infer‑Execute framework that grounds complete scenes into compact task‑relevant scene slices, infers executable task structure, and compiles it into grounded skill‑level action sequences. To evaluate this setting, we introduce FullHome, a human‑validated evaluation suite of 400 household tasks spanning diverse home‑scale environments and both goal‑oriented and process‑constrained requirements. On FullHome, TaskGround improves task success rates by large margins across both proprietary and open‑weight models. Notably, it makes Qwen3.5‑9B competitive with GPT‑5 under direct complete‑scene prompting while reducing total input‑token cost by up to 18x. Our results identify executable task‑structure inference as a central bottleneck in full‑scene household reasoning and show that structured grounding can make compact local models substantially more effective for practical household deployment.
Authors:Kailai Sun, Mingyi He, Heye Huang, Can Rong, Alok Prakash, Baoshen Guo, Shenhao Wang, Jinhua Zhao
Abstract:
Urban Building Energy Modeling plays a critical role in achieving the United Nations' Sustainable Development Goals 7 and 11. Although existing studies based on satellite imagery and deep learning have achieved remarkable progress, many challenges exist: most existing studies are inherently predictive, failing to reflect the generative nature of urban planning; although generative AI and diffusion models have seen explosive growth in satellite imagery, they lack the urban functional generation (e.g., energy layer); third, aligned high‑quality high‑resolution building energy data with satellite imagery is limited and scarce. Here we propose SENSE (Satellite‑based ENergy Synthesis for Sustainable Environment), a unified generative UBEM framework that jointly synthesizes realistic urban satellite imagery and aligned high‑quality building energy consumption and height maps. By conditioning on road networks and urban density metrics, SENSE, based on a controllable diffusion model, leverages the knowledge learned by large vision models to generate urban building energy consumption and height information (annotations) in the latent space. Experiments across four cities (New York City, Boston, Lyon, Busan) demonstrate that SENSE achieves high visual fidelity and strong physical consistency, satisfying the ASHRAE standard metric. Experiments demonstrate that SENSE can generate enough annotated synthetic data using less than 20% labeled energy data, boosting downstream prediction performance by 10% IoU. Compared to SOTA urban energy prediction methods, SENSE significantly reduced prediction error (reduced 3%‑11% NMBE and 1%‑9% CVRMSE). This study offers an energy‑efficiency urban planning and physical generation solution for urban science, energy science and building science. The dataset and code: https://huggingface.co/datasets/skl24/MUSE and https://github.com/kailaisun/GenAI4Urban‑Energy/.
Authors:Corentin Dumery, Niki Amini-Naieni, Shervin Naini, Pascal Fua
Abstract:
Object counting is a foundational vision task with over a decade of dedicated research, yet state‑of‑the‑art models still fail systematically in the mixed‑object setting that dominates real‑world applications such as industrial inspection and product sorting. We show that this gap is strongly driven by limitations in existing training and evaluation data: real counting datasets are prohibitively expensive to annotate and suffer from labeling noise, while existing synthetic alternatives lack diversity and realism. We address this with MixCount, a dataset and benchmark for mixed‑object counting designed to target the failure modes of current counting models. To overcome the high cost of constructing and labeling such data, we develop an automatic generation pipeline that synthesizes images, fine‑grained textual descriptions, and pixel‑perfect counting annotations at scale, eliminating the labeling ambiguity that plagues prior datasets. Evaluating state‑of‑the‑art counting models on MixCount exposes severe degradation in the mixed‑object setting. More importantly, training these models on our synthesized data yields substantial gains on real‑world benchmarks, reducing MAE by 20.14% on FSC‑147 and by 18.3% on PairTally. These results establish MixCount as both a benchmark and a training dataset for fine‑grained counting, and demonstrate that our pipeline, which produces effectively unlimited labeled data, helps address a long‑standing bottleneck in counting models.
Authors:Espen Uri Høgstedt, Christian Schellewald, Annette Stahl, Rudolf Mester
Abstract:
Salmon re‑identification in commercial net‑pens is challenging due to large populations, which impose strict accuracy requirements and make large‑scale labeled data acquisition infeasible. Trajectory IDs can be used as proxy labels, but this introduces trajectory‑ID bias. To address these challenges, we propose a patch‑based re‑identification framework that fuses patch‑level predictions into a salmon identity decision. A key component is the prediction of the salmon's lateral line, enabling extraction of texture‑anchored patches and patch slices. To enable realistic evaluation, we introduce an experimental setup using multiple cameras placed 6 m apart, allowing the same fish to be recorded in different trajectories. This enables the construction of a cross‑camera test set through manual match confirmation. Our ensemble approach outperforms the full‑image baseline in same‑trajectory validation (0.932 to 0.965 mAP) and cross‑camera testing (0.609 to 0.860 mAP). The substantial improvements in the cross‑camera setting demonstrate improved generalizability and robustness. Code and data: https://github.com/espenbh/salmon‑reid‑patch‑ensemble.
Authors:Emmanuel G. Maminta, Rowel O. Atienza
Abstract:
Multimodal product retrieval (MPR) underpins checkout‑free retail and automated inventory systems, yet it demands fine‑grained SKU discrimination that standard vision‑language benchmarks fail to capture. We present the first systematic zero‑shot evaluation of 190 open‑source VLMs on the MPR task of the GroceryVision Challenge, isolating pre‑training data, architecture, and input resolution. Our analysis yields three actionable findings. (1) Data quality trumps scale. Switching from raw web‑scrapes to filtered datasets delivers up to 16.6% accuracy gains, exceeding the benefit of doubling model parameters. (2) Efficient models can win. MobileCLIP‑B (150M parameters) outperforms 351M counterparts trained on noisy data. We introduce semantic power density (ϕ), an efficiency metric that penalizes sub‑threshold accuracy. (3) A precision gap persists. State‑of‑the‑art models achieve 94.5% Recall@5 but suffer a 17.5% drop at Recall@1, revealing that contrastive embeddings cluster categories effectively but fail to rank visually similar SKUs. Code and evaluation scripts are available at \urlhttps://github.com/upeee/openmpr.
Authors:Boyuan Sun, Bowen Yin, Yuanming Li, Xihan Wei, Qibin Hou
Abstract:
We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine‑grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual prompts, such as masks or points, SWIM leverages mask supervision only during training to guide cross‑modal attention, allowing the model to automatically attend to the user‑specified object at inference. Our cross‑attention analysis of pretrained multimodal large languagemodels (MLLMs) reveals a systematic discrepancy: Attribute words produce sharp, localized activations in the visual modality, whereas object nouns yield diffuse and scattered patterns due to semantic reference bias and distributed high‑level representations. To address this misalignment, we construct NL‑Refer, an enriched dataset, in which each object mask is paired with a precise natural language referring expression. SWIM extracts multi‑layer cross‑attention maps from object nouns and enforces spatial consistency with ground‑truth masks. Experimental results demonstrate that SWIM substantially improves text‑visual alignment and achieves superior performance over visual‑prompt‑based methods on fine‑grained object understanding benchmarks. The code and data are available at \hrefhttps://github.com/HumanMLLM/SWIMhttps://github.com/HumanMLLM/SWIM.
Authors:Chang Sun, Hui Yuan, Shiqi Jiang, Chongzhen Tian, Guanghui Zhang, Raouf Hamzaoui
Abstract:
Because LiDAR sensors acquire point clouds with a fixed angular resolution, the resulting data can be systematically parameterized and efficiently compressed in the spherical coordinate system. Traditional spherical coordinate‑based point cloud compression methods have demonstrated strong rate‑distortion (RD) performance, with the predictive geometry coding (PredGeom) method in the geometry‑based point cloud compression (G‑PCC) standard being a prominent example. Although PredGeom includes an inter‑frame prediction mode, it relies on a simple linear model, which limits its ability to capture complex motion patterns and structural dependencies. Meanwhile, existing learning‑based compression methods in the spherical domain do not exploit inter‑frame correlations to reduce geometry redundancy. To address these limitations, we propose a learning‑based inter‑frame predictive coding method, termed Inter‑LPCM. For azimuth prediction, we employ a delta coding strategy based on the predefined angular resolution. To improve radius compression, we introduce an inter‑frame radius predictive (Inter‑RP) model that estimates the current point's radius using neighboring points from both the current frame and the registered reference frame. In addition, we design a lightweight attention‑based prediction (LAEP) model to predict elevation angles by capturing long‑range geometric correlations across different coordinates. For quantization, we propose an RD‑optimized method to select quantization steps in the spherical coordinate system. For entropy coding, we design distinct models for each spherical coordinate component. These models are adapted to the statistical priors of each coordinate, enabling more accurate probability estimation. Our source code is publicly available at https://github.com/SDUChangSun/Inter‑LPCM
Authors:Pei Zhang, Shijie Lin, Zhou Ge, Jinpeng Chen, Wei Pu
Abstract:
Quasi‑bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to clear silhouettes, binarization uses a minimal language to convey essential vision cues for maximum downstream efficiency. The catch is that frame‑based imaging often struggles on mobile platforms like drones, self‑driving cars, and underwater vehicles. In these dynamic scenes, rapid motion and harsh lighting can make it blind, causing severe motion blur and erasing crucial details. To overcome the limits, neuromorphic vision via event cameras, featuring microsecond‑level temporal resolution and high dynamic range, steps in as a natural solution. Building upon this event‑driven sensing paradigm, we introduce a simple yet effective dual‑modal approach that harnesses the synergy between frames and events to achieve real‑time, high‑frame‑rate binarization on CPU‑only devices. Extensive evaluations present that it earns competitive performance against leading techniques in reducing motion blur, while delivering impressive improvements under challenging illumination. Besides, our asynchronous workflow bypasses event scarcity that breaks traditional time‑binning reconstruction, maintaining clear target shapes even at extreme kilohertz frame rates. Its binary results further serve as reliable representations that facilitate a range of downstream tasks. This work paves the way towards lightweight perception and interaction in embodied intelligence on resource‑constrained edge platforms.
Authors:Hyun Lee, Hyemin Jeong, Yejin Kim, Hyungwook Choi, Hyunsoo Cho, Soo Kyung Kim, Joonseok Lee
Abstract:
Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optimized to operate on discrete, semantically meaningful tokens, while prevailing visual projectors transform an image into a long stream of continuous and highly correlated embeddings. This causes the visual tokens to behave differently from the word‑like units that LLMs are originally trained to understand. We propose a novel Disentangled Visual Tokenization (DiVT) that clusters patch embeddings into coherent semantic units, so each token corresponds to a distinct visual concept instead of a rigid grid cell. DiVT further adapts its token budget to image complexity, providing an explicit accuracy‑compute trade‑off modifying neither the vision encoder nor the language model. Across diverse multimodal benchmarks, DiVT matches or surpasses baselines with significantly fewer visual tokens, demonstrating robustness under limited token budgets, significantly reducing memory cost and latency while making visual inputs more compatible with LLMs. Our code is available at https://github.com/snuviplab/DiVT.
Authors:Xiang Yang, Yongli Wang, HaiFeng Li, Yunsheng Zhang
Abstract:
Feed‑forward 3D reconstruction has advanced rapidly, but current models remain unreliable in UAV photogrammetric acquisition. We argue that this failure is caused not only by appearance‑domain shift, but also by UAV‑specific camera‑geometry variations, especially oblique views and HFOV‑height ambiguity. Existing UAV datasets mainly emphasize scene diversity and provide limited coverage of camera configurations, which restricts robustness evaluation and UAV‑domain adaptation. To address this gap, we introduce UAVFF3D, a geometry‑aware real‑synthetic benchmark for feed‑forward UAV 3D reconstruction. UAVFF3D contains more than 170k real UAV images and more than 370k synthetic images rendered from high‑quality textured 3D models, covering diverse HFOVs, flight altitudes, viewing directions, and acquisition patterns. It also includes a controlled HFOV‑height test subset for diagnosing projection‑geometry ambiguity. We further propose an evaluation protocol that jointly assesses camera‑geometry estimation and dense scene reconstruction under a shared global alignment, avoiding the bias caused by separate camera and geometry alignments. Experiments on representative feed‑forward reconstruction models show that UAVFF3D‑based domain adaptation consistently improves camera and geometry estimation, reducing Ray Error by up to 84.2%, Pose ATE by up to 76.0%, and Chamfer Distance by up to 41.1%. In oblique scenes, adaptation reduces the oblique‑nadir rotation gap by up to 90.7%. Under HFOV‑height ambiguity, it improves robustness across HFOV‑height configurations and yields more stable performance across HFOV settings. Incorporating camera priors further improves reconstruction under UAV‑specific acquisition geometries. The dataset and evaluation code are available at https://github.com/yanxian‑ll/UAVFF3D .
Authors:Pan Wang, Yihao Hu, Xiujin Liu, Jingchu Yang, Hang Wang, Zhihao Wen
Abstract:
Vision‑language model (VLM) agents increasingly rely on memory‑augmented reinforcement learning to reuse experience across long‑horizon tasks, yet most existing frameworks store memory as text and depend on proprietary teacher models to summarize or refine it. This design is poorly matched to spatial decision making: geometric priors are compressed into lossy language, and sparse interaction is often supervised through delayed textual feedback rather than dense visually grounded signals. We argue that reusable experience for VLM agents should remain visually grounded. Based on this insight, we propose AtlasVA, a teacher‑free visual skill memory framework that organizes memory into three complementary layers: spatial heatmaps, visual exemplars, and symbolic text skills. AtlasVA further evolves danger and affinity atlases directly from trajectory statistics and lightweight grid heuristics, and reuses these self‑evolving atlases as potential‑based shaping rewards for reinforcement learning. This unifies perception, memory, and optimization without external LLM supervision. Experiments on \textscSokoban, \textscFrozenLake, 3D embodied navigation, and 3D robotic manipulation benchmarks show that AtlasVA consistently outperforms text‑centric memory baselines and competitive VLM agents, with especially strong gains on spatially intensive tasks. Homepage: https://wangpan‑ustc.github.io/AtlasvaWeb
Authors:Jinrang Jia, Zhenjia Li, Yijiang Hu, Yifeng Shi
Abstract:
Generating a consistent whole‑house VR tour from a floorplan and style reference requires both photorealistic panoramas and cross‑view spatial coherence. Pure 2D generators produce appealing single panoramas but re‑imagine geometry and materials when the viewpoint changes, whereas monolithic 3D generation becomes expensive and loses fine texture at multi‑room scale. We introduce PanoWorld, a generative spatial world model that treats whole‑house synthesis as autoregressive generation of node‑based 360‑degree panoramas, matching the discrete navigation used by real VR tour products. PanoWorld uses a floorplan‑derived 3D shell as a global geometric proxy and a dynamic 3D Gaussian Splatting cache as renderable spatial memory. A feed‑forward panoramic LRM designed for metric‑scale multi‑room 360‑degree inputs lifts generated panoramas into local 3DGS updates, while Room‑aware Group Attention suppresses cross‑room feature interference. A topology‑aware progressive caching strategy fuses these local updates without repeatedly reconstructing the full history. By decoupling shell‑based geometry guidance from cache‑rendered visual memory, PanoWorld preserves high‑frequency 2D synthesis quality while improving cross‑node layout and material consistency. The project link is https://jjrcn.github.io/PanoWorld‑project‑home/
Authors:Diandian Guo, Xikai Yang, Ruiyang Li, Jialun Pei, Pheng-Ann Heng
Abstract:
Surgical Video Question Answering (VideoQA) provides a promising paradigm for dynamic intraoperative interpretation, enabling real‑time decision support and context‑aware retrieval in clinical environments. Nevertheless, existing approaches are predominantly restricted to images or short clips, limiting their ability to model long‑range procedural dynamics and causal dependencies across extended surgical workflows. To address this challenge, we propose SurgLQA, a unified long‑horizon VideoQA framework for scalable surgical reasoning. This framework incorporates Faithful Temporal Consolidation (FTC), which leverages intrinsic temporal cues to construct compact long‑range representations while preserving fine‑grained temporal fidelity. Further, we develop Temporally‑Grounded Multi‑Policy Scaling (TMS), an adaptive test‑time inference paradigm that strategically adjusts policy‑level reasoning capacity within temporally grounded contexts. To facilitate systematic evaluation, we restructured a long‑duration colonoscopy VideoQA benchmark, Colon‑LQA, and conducted extensive experiments on Colon‑LQA and REAL‑Colon‑VQA. Experimental results demonstrate that our approach achieves consistent performance gains in long‑range reasoning with temporally grounded inference. Code link: https://github.com/RascalGdd/SurgLQA.
Authors:Yang Li, Weize Li, Quan Yuan, Congzhang Shao, Guiyang Luo, Yunqi Ba, Xuanhan Zhu, Xinyuan Ding, Xiaoyuan Fu, Jinglin Li
Abstract:
By sharing intermediate features, collaborative perception extends each agent's sensing beyond standalone limits, but real‑world feature modality heterogeneity remains a key barrier to effective fusion. Most existing methods, including direct adaption and protocol‑based transformation, typically rely on training adapters for newly emerging feature modalities and often require additional retraining or fine‑tuning. Such repeated training is costly and is often infeasible across manufacturers due to model and data privacy constraints, limiting real‑world scalability. To address this issue, we propose UniTrans, a universal any‑to‑any feature modality translation model that instantiates translators on the fly for arbitrary modalities.
UniTrans pretrains a bank of translator expert parameters and learns their combination coefficients as a function of source‑to‑target modality mapping. The mapping is measured in a modality‑intrinsic latent space, where an intrinsic encoder extracts modality‑specific yet scene‑invariant codes from single‑frame intermediate features, enabling UniTrans to instantiate translators in a zero‑shot manner.
Experiments on OPV2V‑H and DAIR‑V2X demonstrate that UniTrans consistently outperforms state‑of‑the‑art methods in both simulated and real‑world settings, enabling efficient any‑to‑any translation through a universal model. The code is available at https://github.com/CheeryLeeyy/UniTrans.
Authors:Yixing Yong, Jian Wang, Ming Lei, Lijun He, Fan Li
Abstract:
Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlike in the RGB domain, where attacks rely on color texture, infrared attacks must manipulate thermal signatures, making the geometry shape of heat‑blocking materials the primary adversarial information carrier. Current shape‑based methods suffer from a fundamental trade‑off between representational capability and optimization power, limiting their attack effectiveness.In this work, we overcome this dilemma by introducing learnable Fourier shapes to the infrared domain. We utilize an end‑to‑end differentiable framework where a compact set of Fourier coefficients, defining the shape boundary, is analytically mapped to a pixel‑space mask via the winding number theorem. This enables efficient gradient‑based optimization to generate potent shapes that cause human targets to evade detection. Extensive digital and physical experiments provide a comprehensive evaluation and validate our superior performance. Our resulting physical patch achieves striking robustness, successfully evading detectors across diverse distances, angles, poses, and individuals, and achieves over 88% attack success rate at distances greater than 25m (conf.=0.5). Code is available at https://github.com/Yongyx99/Fourier‑shape‑attack.
Authors:Xinpeng Liu, Hiroaki Santo, Yosuke Toda, Fumio Okura
Abstract:
Accurate estimation of plant skeletal structures (e.g., branching structures) from images is essential for smart agriculture and plant science. Unlike human skeletons with fixed topology, plant skeleton estimation presents a unique challenge, i.e., estimating arbitrary tree graphs from images. To address this problem, we introduce PlantPose, a universal plant skeleton estimator via tree‑constrained graph generation. PlantPose combines learning‑based graph generation with traditional graph algorithms to enforce tree constraints during the training loop. To enhance the model's generalization capability, we curate a large and diverse dataset comprising real‑world and synthetic plant images, along with simplified representations (e.g., sketches and abstract drawings). This dataset enables the generalized model to adapt to diverse input styles and categories of plant images while preserving topological consistency. Our approach demonstrates robust and accurate plant skeleton estimation across multiple domains, including previously unseen out‑of‑domain scenarios. Further analyses highlight the method's strengths and limitations in handling complex, heterogeneous data distributions. All implementations and datasets are available at https://github.com/huntorochi/PlantPose/.
Authors:Yinyi Luo, Wenwen Wang, Hayes Bai, Marios Savvides, Jindong Wang
Abstract:
Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We observe that this issue does not stem from a lack of shared representations, but from the absence of explicit alignment between the transformations that map into and out of the latent space. As a result, generation and re‑encoding can follow inconsistent trajectories, leading to semantic drift under modality transitions. In this work, we propose LatentUMM, a framework that constructs an enhanced shared latent space to explicitly align these transformations and improve cross‑modal consistency. LatentUMM consists of two stages. First, dual latent alignment enforces consistency at both the modality and capacity levels: cross‑modal alignment uses a stronger embedding model to impose structured cross‑modal semantics, while dual capacity alignment enforces bidirectional consistency under generation and re‑encoding. Second, latent dynamics stabilization improves robustness via stochastic latent rollouts and preference optimization, favoring trajectories that better preserve semantic consistency. Experiments show that LatentUMM consistently improves multimodal consistency across diverse architectures. Code is available at: https://github.com/AIFrontierLab/TorchUMM/tree/main/src/umm/post_training/LatentUMM.
Authors:U. V. B. L. Udugama, George Vosselman, Francesco Nex
Abstract:
Autonomous agile robots need more than metric geometry: they must understand objects, rooms, places, and spatial relations for search, inspection, exploration, and human robot interaction. Conventional metric maps support localization and collision avoidance, but do not provide this semantic and relational structure. 3D scene graphs address this gap by connecting geometry with object level and room level understanding. Building such representations on agile platforms remains difficult because aerial and lightweight robots operate under strict payload, power, and compute limits, making RGB‑D cameras and LiDAR sensors impractical for many onboard settings. We present Mono‑Hydra++, a real time monocular RGB plus IMU pipeline for indoor metric semantic mapping and hierarchical 3D scene graph construction. The system combines M2H‑MX, a DINOv3 based multi‑task model for depth and semantics, with a deep feature visual inertial odometry front end, sparse predicted depth constraints in the VIO derived pose graph, semantic masking for dynamic regions, and pose aware temporal alignment before volumetric fusion in the Mono‑Hydra backend. On the Go‑SLAM ScanNet evaluation subset, Mono‑Hydra++ achieves 1.6% lower average trajectory error than the strongest RGB‑D baseline in our comparison, while using only monocular RGB plus IMU input. On calibrated 7‑Scenes, it improves average ATE by 29.8% over the strongest competing calibrated baseline. We further validate Mono‑Hydra++ in a real ITC building deployment using RealSense RGB plus IMU and demonstrate embedded feasibility by deploying the ONNX/TensorRT FP16 M2H‑MX‑L perception model at 25.53 FPS on a Jetson Orin NX 16GB. These results show that Mono‑Hydra++ can provide real time metric semantic mapping and scene graph construction for resource constrained robotic platforms without relying on active depth sensors.
Authors:Debashish Chakraborty, Dengjia Zhang, Jialiang Jin, Hanting Liu, Katherine Guerrerio, Hanxiang Qin, Tyler Skow, Alexander Martin, Reno Kriz, Benjamin Van Durme
Abstract:
Retrieval‑augmented generation from videos requires systems to retrieve relevant audiovisual evidence from large corpora and synthesize it into coherent, attributed text. Current approaches struggle at both ends: retrieval methods fail on complex, multi‑faceted queries that cannot be captured by a single embedding, while generation methods lack the high‑level reasoning needed to synthesize across multiple videos and face memory constraints over long, multi‑video contexts. We present MARQUIS: a three‑stage pipeline that addresses these limitations through (1) query expansion, fusion, and reranking, (2) calibrated structured evidence extraction, and (3) article generation from extracted evidence, optionally controlled by an RLM. On the MAGMaR2026 shared task, we improve retrieval performance from 0.195 to 0.759 (nDCG@10). For article generation, ITER‑QA‑BASE improves average human score from 3.09 to 3.83 over the CAG baseline, while MARQUIS‑RLM achieves a human score of 3.30 and the strongest citation recall among non‑QA systems.
Authors:Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Abstract:
Open‑vocabulary segmentation models such as SAM3 perform well across broad categories via text prompting, yet degrade when target classes are visually underrepresented in pretraining or depart from canonical depictions‑limitations text prompts cannot resolve spatially. We present SegRAG, a training‑free retrieval‑augmented segmentation framework that grounds SAM3 with class‑specific point prompts derived from a curated DINOv3 feature bank. Offline, dense patch‑level descriptors are extracted from annotated references and filtered by Intra‑Class Cohesion Distillation (ICCD), retaining only prototypes that reliably retrieve within‑class foreground. At inference, Topographic Similarity Grounding (TSG) computes a cosine‑similarity landscape against retrieved prototypes, identifies coherent high‑confidence regions via connected‑component analysis, and extracts peak locations through non‑maximum suppression. The resulting point prompts are delivered jointly with class‑name text in a single SAM3 forward pass. On four standard benchmarks, SegRAG consistently outperforms the text‑only baseline, gaining up to +3.92 mIoU on LVIS. On AgML agricultural benchmarks under zero‑shot domain transfer, it raises mean IoU from 25.27 to 59.24 (+33.97) and recovers individual classes from zero to over 95 mIoU. Ablations confirm that ICCD, TSG, and joint prompting each contribute independently and compound when combined. Code is available at (https://github.com/boudiafA/SegRAG).
Authors:Miquel Martí i Rabadán, Alessandro Pieropan, Hossein Azizpour, Atsuto Maki
Abstract:
We investigate the potential of invariant and equivariant semi‑supervised learning for addressing the challenges of training multi‑task models on partially labeled datasets with differently structured output tasks. Specifically, we use the popular FixMatch method for invariant semi‑supervised learning and its equivariant extension Dense FixMatch. We evaluate their performance on the Cityscapes and BDD100K datasets in the context of the prevalent object detection and semantic segmentation tasks in computer vision. We consider varying sizes of the subsets annotated for each task and different overlaps among them. Our results for both invariant and equivariant semi‑supervised learning outperform supervised baselines in most situations, with the most significant improvements observed when fewer labeled samples are available for a task and generally better results for the latter approach. Our study suggests that invariant/equivariant learning is a promising general direction for multi‑task learning from limited labeled data.
Authors:Bin Kang, Shaoguo Wen, Yang Fan, Shunlong Wu, Junjie Wang, Yulin Li, Junzhi Zhao, Junle Wang, Zhuotao Tian
Abstract:
While existing text‑to‑speech (TTS) models exhibit high expressiveness, fine‑grained control over composite instructions remains challenging due to the structural mismatch between discrete textual intents and continuous acoustic realizations. Inspired by human cognitive decoupling, we introduce AgentSteerTTS, a multi‑agent closed‑loop framework designed for intent‑faithful expressive control of composite instructions. First, in our framework, an adversarial disentanglement agent mitigates speaker‑emotion leakage by learning separable identity and emotion‑prosody subspaces with leakage‑suppressing regularization. Next, a Dual‑Stream Anchoring Controller grounds abstract intents using a large‑scale acoustic prototype library: a Retrieval Agent selects expressive anchors, while a Synthesis Agent fuses them into continuous control vectors via gated attention. Finally, a Fast‑Slow Feedback Agent refines output intensity through latent gradient correction and resolves semantic‑acoustic mismatches using high‑level perceptual critique. Experiments on a composite‑instruction benchmark and public test sets show that AgentSteerTTS yields consistent and significant improvements to the baselines, demonstrating the effectiveness of the proposed method.
Authors:Jeongeun Park, Janghyeok Han, Geonung Kim, Hyun-Seung Lee, Kyuha Choi, Youngseok Han, Sunghyun Cho
Abstract:
Video outpainting generates plausible visual content beyond the original spatial extent of a video, playing a key role in adapting videos to diverse display formats. To support such use cases, it must enable large spatial extrapolation over long sequences. However, most existing methods address only one of these challenges or lack explicit mechanisms for ensuring global spatio‑temporal consistency, leading to notable limitations. In this paper, we propose HL‑OutPaint, a high‑resolution video outpainting framework for long sequences. Our approach follows a coarse‑to‑fine strategy with a two‑stage pipeline. We first construct Global Coarse Guidance (GCG), a low‑resolution representation that captures global structure and dominant motion across the video. Unlike naive downsampling, GCG is built via a novel global‑local frame swapping mechanism that couples sparse global keyframes with local temporal windows and exchanges information during sampling. This enables GCG to encode both long‑term structural consistency and short‑term temporal dynamics in a unified representation. Guided by this representation, HL‑OutPaint then performs high‑resolution outpainting to generate spatially detailed and temporally consistent content. By separating global structure modeling from fine‑grained synthesis, our framework achieves stable, coherent generation for large spatial expansion and long video sequences. Extensive experiments show that HL‑OutPaint outperforms existing methods in challenging scenarios involving wide spatial extrapolation and long video sequences.
Authors:Yuting Yang, Haichao Jiang, Tianming Liang, Quan Zhang, Jian-Fang Hu
Abstract:
Referring segmentation aims to segment the target objects in images or videos based on the textual query. Despite remarkable progress over the past years, existing works always assume that the user‑provided queries are already precise and clear. However, this assumption is impractical. In real‑world scenarios, it is unrealistic to expect all users to thoroughly review their visual content and carefully ensure their queries are unique and unambiguous. When encountering such cases, existing segmentation models tend to arbitrarily guess the user preferences, often resulting in undesired outcomes. To address this limitation, we propose IC‑Seg, a novel agentic framework that proactively clarifies user intent through multi‑turn conversation before segmentation. To effectively incentivize this capability, we further introduce Hi‑GRPO, a new hierarchical optimization strategy that injects dense and informative supervision signals at the trajectory, turn, and step levels. This strategy encourages efficient intent clarification, effectively eliminating redundant interactions and improving overall dialogue quality. For evaluation, we establish Ambi‑RVOS, a referring video object segmentation benchmark with ambiguous user queries. Extensive experiments demonstrate that IC‑Seg not only outperforms existing methods by a large margin in resolving ambiguous queries, but also maintains state‑of‑the‑art performance on standard reasoning segmentation benchmarks. Code and data will be released at https://github.com/iSEE‑Laboratory/IC‑Seg.
Authors:Erdi Sarıtaş, Eren Onaran, Vitomir Štruc, Hazım Kemal Ekenel
Abstract:
Face Image Quality Assessment (FIQA) is a crucial control step in biometric pipelines. It ensures only reliable samples are processed to maintain system accuracy. State‑of‑the‑art FIQA methods achieve high utility but typically operate as "black boxes." They produce scalar scores without human‑interpretable justifications. This lack of transparency limits their effectiveness in human‑in‑the‑loop scenarios, such as automated border control, where actionable feedback is essential. In this paper, we investigate the potential of off‑the‑shelf Vision‑Language Models (VLMs) to bridge this gap by performing FIQA in a zero‑shot setting. We present a comprehensive evaluation framework for assessing VLM performance. This involves benchmarking traditional FIQA methods through error‑versus‑reject curves. Additionally, using a diverse set of datasets, ranging from surveillance‑oriented to synthetically generated, we analyzed their interpretability, consistency, and robustness to prompt changes. Our results show biometric utility performance depends significantly on architecture, not merely on parameter count. Most VLMs' outputs align with those of traditional methods. We also find that VLM ranking performance and the generated scores may vary across prompts. Our synthetic ablation study shows that while increasing the parameter count can improve internal consistency, it yields worse degradation‑detection performance than smaller models. These findings suggest that zero‑shot FIQA score estimation using VLMs is promising and could effectively complement conventional FIQA pipelines as an interpretability module. The codes are available at https://github.com/ThEnded32/VLM4FIQA.git.
Authors:Wentong Li, Zhiyuan Qi, Zichen Zhao, Kai Zhang, Lei Zhang
Abstract:
Pre‑trained vision foundation models (VFMs) provide strong semantic representations, yet their patch‑level features are inherently coarse, limiting their effectiveness on tasks requiring fine‑grained localization, dense prediction, and point‑wise correspondence. In this work, we revisit feature upsampling for VFMs from the perspective of inverse problem and propose Weighted Reverse Convolution (WRC), a spatially adaptive inverse operator for densifying high‑level visual descriptors. Specifically, we formulate feature upsampling as a weighted Tikhonov‑regularized least‑squares problem, where spatially varying weights modulate both data fidelity and prior strength at each spatial location. This allows WRC to adapt the reconstruction to spatially varying feature characteristics, thereby preserving critical structures while mitigating over‑smoothing. Moreover, WRC retains an efficient, fully differentiable closed‑form FFT solution, making it a practical drop‑in upsampling operator. Integrated into a lightweight self‑supervised densification framework, WRC consistently improves dense feature quality across various downstream benchmarks, including segmentation, depth estimation, video object segmentation, object discovery, and keypoint correspondence, while maintaining high computational efficiency.
Authors:Hanli Zhao, Binhao Wang, Shihao Zhao, Tao Wang, Kaihao Zhang, Wanglong Lu
Abstract:
Image super‑resolution (SR) aims to reconstruct high‑quality, high‑resolution (HR) images from low‑resolution (LR) inputs and plays a critical role in various downstream applications. Despite recent advancements, balancing reconstruction fidelity and computational efficiency remains a fundamental challenge, particularly in resource‑constrained scenarios. While existing lightweight methods attempt to expand receptive fields, many of them either incur substantial computational overhead, naively scale up kernel sizes, or lack mechanisms for coherent multi‑scale integration, limiting their overall effectiveness and scalability. To address these limitations, we propose EchoSR, an efficient context‑harnessing framework for lightweight image super‑resolution, which unifies multi‑scale receptive field modeling and hierarchical context fusion. EchoSR decouples feature learning into disentangled local, multi‑scale, and global modeling stages through an efficient context‑harnessing strategy, and further promotes seamless cross‑scale integration via a cross‑scale overlapping fusion mechanism. Extensive experiments have shown that EchoSR consistently outperforms state‑of‑the‑art lightweight super‑resolution methods across multiple benchmarks, while also achieving a faster speed (~ 2×). The source code is available at https://github.com/funnyWang‑Echoes/EchoSR.
Authors:Zhipeng Deng, Jiale Zhou, Wenhan Jiang, Haolin Wang, Xun Lin, Yafei Ou, Yefeng Zheng
Abstract:
Deploying multi‑sequence magnetic resonance imaging (MRI) segmentation models to new clinical environments is challenging due to variations in scanners and acquisition protocols. Although existing TTA methods handle basic per‑modality shifts, they often fail under a fundamental dual‑shift problem, as their adaptation signals fail to capture modality‑interaction shifts that disrupt inter‑sequence consistency. To address this, we propose Variance‑gated Inter‑Sequence Test‑time Adaptation (VISTA), a source‑free framework that tackles modality‑interaction shifts. First, we design an Inter‑Sequence Intervention Generator (ISIG) that generates a set of consistency probes by swapping low‑frequency spectra and entropy‑localized patches across sequences, preserving anatomical semantics while challenging inter‑sequence dependencies. Second, we introduce Cross‑View Disagreement‑Aware Pseudo Labeling (CDPL), which establishes a voxel‑wise reliability metric using cross‑view disagreement variance to dynamically gate self‑training and enforce interventional consistency, encouraging the network to rely on robust anatomical semantics. Extensive experiments adapting from standard adult MRI (BraTS‑GLI‑Pre) to African low‑field (BraTS‑SSA) and pediatric (BraTS‑PED) cohorts show improved performance over competing methods under clinical shifts, achieving absolute Dice improvements of +1.89% (SSA) and +2.82% (PED) over the source model. The code is available at https://github.com/dzp2095/VISTA.
Authors:Decheng Liu, Bin Hu, Xinbo Gao, Dawei Zhou, Chunlei Peng, Nannan Wang, Ruimin Hu
Abstract:
Different from existing cross‑modality identification tasks (e.g., heterogeneous face recognition, sketch re‑identification, etc.), we introduce a novel yet practical setting for these related identification tasks, named sketch biometric identification, which aims to continually train a unified model across different data domains, even diverse identification tasks. Sketch biometric identification faces challenges, including scarce real sketch data, high annotation costs, privacy risks, and insufficient generalization ability of cross‑task models. Existing methods usually rely on limited real data or single‑task optimization, making it difficult to effectively address the joint challenges of cross‑modality and cross‑task. This paper proposes a unified framework that integrates efficient synthetic sketch generation and task‑sequential continual learning. First, we design an efficient pipeline to generate a large‑scale and high‑quality synthetic person and face sketch data, which significantly reduces costs and avoids privacy risks. Meanwhile, we enhance the model's robustness by fusing real data. Second, we construct a universal unified framework for sketch biometric identification, which adopts a task‑sequential training strategy: the model first completes sketch person re‑identification learning on the person dataset; subsequently, it maintains the acquired person recognition capability through a trusted sample replay technique and seamlessly performs incremental training on the face dataset. This enables a single model to simultaneously handle the cross‑task capabilities of multiple sketch biometric identification tasks. To support the study of the mentioned sketch biometric identification, we built a new large‑scale benchmark, SketchUnified‑BioID, with several practical evaluation protocols.
Authors:Xianke Chen, Daizong Liu, Yushuo Lou, Xin Tan, Xun Yang, Shuhui Wang, Xun Wang, Jianfeng Dong
Abstract:
Different from traditional text‑to‑image retrieval tasks, chat‑based image retrieval allows the human‑interactive system to iteratively clarify and refine user intent through multi‑round dialogue, thereby achieving more fine‑grained retrieval results. The key challenge in this task lies in dynamically understanding and updating the user's query intent across dialogue rounds. Although existing works have achieved great performance on this new task, they simply handle history query information either by directly concatenating all previous queries into a long textual sequence or by relying on large language models to reconstruct the current query from history. Such strategies are computationally redundant and easily lead to inconsistent intent representations as the dialogue progresses. To alleviate these issues, this paper proposes a novel and efficient memory‑based user intent updating framework for the chat‑based image retrieval task, called Memory‑Augmented Query Intent Understanding (MAQIU). It introduces a lightweight memorization module that dynamically aggregates and evolves the semantic representation of query intent across dialogues, while a memory recall mechanism is further employed to prevent intent forgetting and enhance long‑term semantic integrity. In addition, MAQIU also integrates historical image retrieval results as visual guidance, allowing the model to strengthen cross‑round correlations and refine current visual understanding. Extensive experiments demonstrate that MAQIU achieves substantial performance gains while maintaining high computational efficiency, reducing dialogue encoding FLOPs by 86.4% compared with the prior baseline ChatIR. Source code is available at https://github.com/HuiGuanLab/MAQIU.
Authors:Xinyao Liu, Zhipeng Deng, Wenhan Jiang, Haolin Wang, Xun Lin, Yafei Ou, Yefeng Zheng
Abstract:
The release of public 3D medical image segmentation (MIS) datasets accelerates clinical research but simultaneously heightens risks of unauthorized AI model training. While Unlearnable Examples (UE) offer protection by injecting imperceptible perturbations to prevent effective model learning, existing methods primarily target 2D scenarios. They neglect the volumetric spatial correlations and inter‑slice anatomical consistency inherent in 3D medical volumes, which serve as critical learning priors for 3D segmentation networks. To bridge this gap, we propose VoxShield, a UE framework that explicitly targets the volumetric inductive biases of 3D networks. Our core insight is that by systematically dismantling the cross‑slice continuity that 3D architectures rely on, we can fundamentally impair their spatial aggregation process. Specifically, we introduce an Inter‑Slice Frequency Consistency Disruption mechanism that maximizes the spectral divergence between adjacent slices, injecting structural incoherence along the z‑axis. Complementing this structural attack, a Semantic Prediction Disruption module is incorporated. By maximizing the \ell_1 divergence between clean and perturbed logits, it forces the injected noise to penetrate the entire network and corrupt the final semantic mapping. Experiments on BraTS19 and FLARE21 demonstrate that VoxShield successfully degrades 3D segmentation performance, reducing the DSC from 80.0% to near 0.0% and from 88.6% to 6.8%, respectively. All protections are achieved with minimal perturbation (ε=4/255) to preserve high visual fidelity. The code is available at https://github.com/KK266299/VoxShield.
Authors:Yuantai Zhang, Jiaqi Yang, Huajian Zeng, Changhao Chen, Haoang Li, Liang Li, Dezhen Song, Xingxing Zuo
Abstract:
Fast and reliable initialization is critical for monocular visual‑inertial navigation systems (VINS), as it establishes the starting conditions for subsequent state estimation. Despite steady progress, most existing methods heavily rely on visual feature correspondences and require 3‑4 seconds of sensory data for successful initialization, which limits their applicability and efficiency. With the advent of feed‑forward 3D models that can directly predict point clouds from images, we revisit the visual‑inertial initialization problem from a concise perspective. In this work, we propose a feature‑free initialization framework that leverages up‑to‑scale point clouds predicted by a feed‑forward 3D model, thereby obviating the need for visual feature tracking and estimation. This design substantially reduces system complexity and improves the reliability of initialization. Experiments on public datasets demonstrate that the proposed feature‑free initialization method achieves the highest success rate, exceeding 90%, and significantly reduces the data duration required for successful initialization, typically to under 1.2 s. We further validate our method on a self‑collected dataset covering various indoor and outdoor scenarios, demonstrating robust performance, particularly in visually degraded environments where existing methods often fail. The code and dataset are available at https://github.com/Yuantai‑Z/FF‑VIO‑Init.
Authors:Yuyao Zhang, Alexander Huang-Menders, Yu-Wing Tai
Abstract:
High‑resolution image editing is essential for professional and creative applications, yet existing multimodal diffusion‑based editors remain computationally inefficient and constrained to relatively low resolutions. Current approaches redundantly process the entire image canvas or rely on large‑scale high‑resolution datasets, resulting in substantial training and inference costs. We introduce HierEdit, a region‑aware hierarchical diffusion framework designed for efficient and scalable high‑resolution image editing. Our method first performs edits on a low‑resolution proxy using an off‑the‑shelf editing model to generate a reference and to localize the modified regions. A hierarchical local‑window diffusion model (Local‑Window MMDiT) that refines only edited regions within the original high‑res image, while reusing the unaltered regions as conditioning inputs. The low‑resolution proxy further provides structural guidance and intermediate denoising supervision (Inference Acceleration) , ensuring consistent global semantics and stable generation without the need for full‑resolution attention computation. This targeted and hierarchical design enables fast, high‑fidelity editing of images up to 4K resolution without any specialized high‑resolution training data. Extensive experiments demonstrate that HierEdit achieves competitive visual quality on commodity‑resolution datasets while significantly accelerating inference and extending seamlessly to ultra‑high‑resolution 4K editing. Please check our \hrefhttps://peteryyzhang.github.io/HierEdit‑page/Project Page.
Authors:Jun Ma, Zhenye Yang, Ruichen Zhou, Pei Zhang, Huan Li, Jinpeng Chen
Abstract:
Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudden lighting changes and sensor noise, spatial‑domain models struggle to disentangle authentic gaze cues from irrelevant visual attributes. In this paper, we propose LISA, a Language‑guided Interference‑aware Spatial‑Frequency Attention framework that combines frequency‑domain priors with vision‑language knowledge. Observing that the amplitude spectrum remains relatively stable even under spatial perturbations, we design a dual‑domain fusion mechanism. It integrates stable low‑frequency semantics into high‑frequency details, employing spatial attention to precisely target ocular regions. To reduce semantic ambiguity, we also introduce a training‑time disentanglement strategy. Using a frozen CLIP encoder and orthogonal regularization, we explicitly separate gaze features from appearance interference. Experiments on two benchmarks show that LISA achieves state‑of‑the‑art performance, with significantly improved robustness against occlusions and lighting variations. The code repository is available at https://github.com/Mason‑bupt/LISA.
Authors:Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Xiangyu Liu, Fengchao Xiong, Jianfeng Lu, Chengrong Chen, Jun Zhou
Abstract:
While hyperspectral imaging provides rich spatial‑spectral information across hundreds of narrow wavelength bands for precise material identification, ground‑based hyperspectral pre‑trained backbones remain absent, constrained by varying spectral configurations across sensors, the scarcity and inconsistency of labels, and the limited scale and scene diversity of existing datasets. To address these challenges and enable universal perception, we propose HyperVision, the first ground‑based hyperspectral pre‑trained backbone. First, to handle varying spectral configurations, HyperVision adopts a channel‑adaptive dynamic embedding mechanism to map heterogeneous inputs into a unified token space. Second, we develop an unsupervised representation learning framework. Specifically, to address label scarcity and inconsistency, a multi‑source pseudo‑labeling method is introduced to fuse spatial structures from SAM2 and fine‑grained spectral material information from HyperFree. Furthermore, to enrich scene diversity and compensate for limited dataset scale, a cross‑modal knowledge distillation mechanism is utilized to transfer rich semantic representations from a pre‑trained RGB vision model to our backbone. Pre‑trained on a collection of 15k images from 26 diverse ground‑based datasets, HyperVision demonstrates exceptional generalization. Requiring only efficient head‑only adaptation without adjusting backbone parameters, it achieves state‑of‑the‑art performance compared to task‑specific methods across three downstream tasks under varying sensor configurations, yielding up to a 16.3% relative improvement in hyperspectral semantic segmentation \mathrmAcc_\mathrmM, a 2.1% relative gain in object tracking AUC, and a 35.5% reduction in salient object detection MAE. The source code and pre‑trained model will be publicly available on https://github.com/lronkitty/HyperVision .
Authors:Chenmin Yu, Liu Yu, Daiqing Wu, Gengluo Li, Zeyu Chen, Yu Zhou
Abstract:
Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored, tracking text in videos is essential for dynamic text manipulations such as segmentation, removal, and editing. To fill this gap, this paper formalizes this specific task as Scene Text Tracking and presents the first systematic work for it. We identify three primary challenges in this task: 1) severe geometric distortions from perspective shifts, 2) high visual ambiguity across different instances, and 3) high sensitivity to fine‑grained structural details. To address these issues, we propose SymTrack, a unified detection‑free framework with synergistic dual‑branch design. It integrates a Cross‑Expert Calibration mechanism to reduce semantic bias, along with a Predictive Token Rectification mechanism to correct structural imbalances, complemented by an Adaptive Inference Engine that stabilizes predictions under motion constraints. Considering the lack of dedicated benchmarks for this task, we utilize three datasets from video text spotting to construct a benchmark with high‑quality annotations. Extensive experiments demonstrate that SymTrack sets the new state‑of‑the‑art on all three benchmarks, outperforming previous best trackers by up to 11.97% AUC on \textBOVText_\textSOT . Overall, our work promotes efficient and thorough text tracking, paving the way toward more generalized video text manipulation.
Authors:Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin, Junhwa Hur, Deqing Sun, Bohyung Han, Ming-Hsuan Yang, Boqing Gong
Abstract:
The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long‑form video lies in managing the explosion of visual‑token context length. Existing strategies predominantly focus on "post‑hoc" token reduction ‑‑ reducing visual tokens after feature extraction to alleviate the LLM's computational overhead. While these methods effectively reduce the number of visual tokens, we observe that the primary latency bottleneck then shifts from the LLM to the expensive per‑frame processing of the vision encoder. To address this, we introduce LiteFrame, a strong, yet highly efficient video encoder backbone for Video LLMs. To train LiteFrame, we propose Compressed Token Distillation (CTD), a novel training framework that teaches a compact student vision encoder to directly predict information‑dense, spatio‑temporally compressed representations produced by a large teacher vision model, effectively bypassing redundant computation. When coupled with further Language Model Adaptation (LMA), this approach results in a new latency‑accuracy Pareto frontier ‑‑ compared with InternVL3‑8B, LiteFrame provides a 35% reduction in end‑to‑end latency while processing 8× more frames and improves average video understanding accuracy across multiple benchmarks. Our results demonstrate a new potential path to unlocking longer‑form video understanding under fixed compute budgets.
Authors:Hoda Osama Elkhodary, Sherin Mostafa Youssef, Marwa Elshenawy, Dalia Sobhy
Abstract:
The rapid advancement of Deepfake technologies and video manipulation tools poses a critical challenge to multimedia forensics, judicial evidence integrity, and information authenticity. Current detectors rely on single‑modality signals, treating appearance, geometry, and motion independently. However, advanced generators maintain within‑modality consistency while producing cross‑modal contradictions, which are forensically discriminative but invisible to any single‑modal detector. We propose CAM‑VFD, a Cross‑Attention Multimodal Video Forgery Detection framework that models cross‑modal contradiction as a directional forensic signal. The framework uses a cross‑attention fusion mechanism in which CLIP‑based appearance representations serve as queries against VideoMAE motion features and MiDaS depth features, enabling the identification of contradictions between visual, temporal, and geometric evidence. We examine this design through cross‑modal attention discrepancy analysis, observing statistically separable real and fake distributions (p<0.001, Cohen's d=0.68). Experimental results on two generative video benchmarks indicate consistent performance, with 95.31% Top‑1 accuracy on GenVidBench and 93.43% accuracy, 90.63% F1‑score, and 96.56% AUROC on GenVideo. Moreover, CAM‑VFD demonstrates stable performance under compression, noise, blur, and adversarial perturbations, suggesting that cross‑modal reasoning may improve robustness in media forensics. The code is publicly available at \urlhttps://github.com/Hoda‑Osama/CAM‑VFD/tree/main.
Authors:Minhas Kamal, Hiranya Garbha Kumar, Balakrishnan Prabhakaran
Abstract:
Point cloud stands as the most widely adopted format for representing 3D shapes and scenes due to its simplicity and geometric fidelity. However, its inherent unordered and irregular nature, exacerbated by sensor noise and occlusions, introduces unique challenges for machine learning based methodologies. To combat these issues, diverse strategies have been developed, including converting to a format that has orderliness, extracting local geometry, and permutation‑invariant or self‑attention‑based processing. In this paper, our focus is directed towards deep learning models for three fundamental tasks in 3D vision: point cloud classification, part segmentation, and semantic segmentation. We begin by formally defining point cloud data, followed by an in‑depth discussion on its structural characteristics. Then, we categorize notable works based on their backbone structure and evaluate their performance on popular benchmarks. Beyond empirical comparison, we offer insights into architectural innovations and limitations. We also outline open challenges and promising future directions for 3D point cloud understanding.
Authors:Lixin Xue, Chengwei Zheng, Georgios Paschalidis, Chen Guo, Manuel Kaufmann, Juan Zarate, Dimitrios Tzionas
Abstract:
Reconstructing people, objects, and their interactions in 3D is a long‑standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill‑posed; depth is ambiguous, humans and objects occlude each other, and camera and object motion entangle to create apparent motion. Most prior work addresses humans or objects in isolation, ignoring their interplay, or assumes known 3D shapes or cameras, which is impractical for real‑world applications. We develop RHINO (Reconstructing Human Interactions with Novel Objects), a three‑step framework that recovers in 3D a human, novel (unseen) manipulated object, and static scene in a common world frame from a monocular RGB video. First, we leverage 3D‑aware foundation models to obtain cues that stabilize Structure‑from‑Motion (SfM) even for low‑texture regions; this yields a coarse shape and apparent motion of a manipulated object from foreground pixels, and a coarse scene shape and camera motion from background pixels. Second, we estimate a human in the camera frame via an off‑the‑shelf method, and subtract the camera motion from apparent motion to extract the object motion; this registers the human, object, and coarse scene shapes into a common world frame. Third, we refine shapes using a compositional neural field with per‑component signed‑distance fields. The latter further enables differentiable contact priors that attract surfaces while penalizing interpenetration, improving the physical plausibility of the final reconstruction. For evaluation, we capture a new dataset of handheld monocular videos synchronized with a volumetric 4D capture stage, providing ground‑truth shape and camera motion. RHINO outperforms state‑of‑the‑art baselines on novel‑view synthesis and 4D reconstruction. Ablations show that each stage contributes substantially. Code and data are available at https://lxxue.github.io/RHINO.
Authors:Pu Li, Huafeng Li, Yafei Zhang, Wen Wang, Neng Dong, Jie Wen
Abstract:
In open‑world settings, thermal infrared (TIR) image degradations continuously emerge and evolve, while most existing all‑in‑one restoration methods are built on a closed‑set assumption and struggle to continually adapt to novel degradations. To address this, we propose ECMRNet, an Expandable, Compressible, and Mineable Restoration Network for open‑world TIR restoration from a continual learning perspective. Conceptually, ECMRNet unifies continual degradation learning as an "expand‑compress‑mine" closed‑loop process, enabling sustained adaptation to new degradations with controllable evolution. Structurally, ECMRNet decomposes intermediate representations into group‑isolated subspaces, and achieves strict parameter isolation and fast adaptation to new degradations by freezing historical groups and isomorphically expanding new ones. To curb model growth as tasks accumulate, we present Structural Entropy Pruning, which identifies and removes redundant channel groups via two‑dimensional structural entropy minimization, achieving information contribution‑driven adaptive compression. Moreover, we design a Sub‑degradation Knowledge Mining Module that dynamically retrieves and recombines transferable components from historical representations to improve restoration under compound degradations. Experimental results demonstrate that ECMRNet achieves superior overall performance across diverse single and compound degradations while using fewer parameters and lower computational cost. The source code is available at https://github.com/Kust‑lp/ECMRNet.
Authors:Jinjie Shen, Zheng Huang, Yuchen Zhang, Yujiao Wu, Yaxiong Wang, Lechao Cheng, Shengeng Tang, Tianrui Hui, Nan Pu, Zhun Zhong
Abstract:
Existing vision‑language forgery detection and grounding methods operate under a closed‑world paradigm, assuming verification can be completed by the model alone. However, self‑contained MLLMs are constrained by finite parametric knowledge, static training corpora, and limited perceptual resolution, creating a practical ceiling in dynamic open‑world forensics ‑‑ particularly for real‑time event verification requiring external clues and forgery segmentation demanding fine‑grained scrutiny of local manipulations. To address these limitations, we shift from scaling up the self‑contained model toward reaching beyond it. We propose OmniVL‑Guard Pro, a tool‑augmented agent that extends unified forensics from closed‑world prediction to open‑world clues‑driven reasoning. OmniVL‑Guard Pro integrates a tool environment spanning real‑time event search, local cropping and zooming, edge‑anomaly screening, face detection, video frame extraction, and SAM3‑based segmentation. To generate high‑quality tool‑reasoning trajectories, we introduce Tree‑Structured Self‑Evolving Tool Trajectory Generation, which produces diverse trajectories through seed guidance, guider‑free self‑evolution, and weakly‑hinted hard sample synthesis, yielding the Full‑Spectrum Tool Reasoning (FSTR) dataset for training. We further propose Checker‑Guided Agentic Reinforcement Learning (CGARL), which provides process‑level supervision to penalize cases where the answer is correct but the reasoning is distorted. Extensive experiments demonstrate that OmniVL‑Guard Pro achieves state‑of‑the‑art performance across various tasks, and exhibits strong zero‑shot generalization. The FSTR dataset and code for OmniVL‑Guard Pro will be publicly released at https://github.com/shen8424/OmniVL‑Guard‑Pro.
Authors:Saeed Firouzi Daghigh, Majid Iranpour Mobarekeh, Mostafa Alavi, Mehdi Bagheri
Abstract:
We present HighSync, an end‑to‑end diffusion‑based framework for high‑fidelity lip synchronization that generates photorealistic talking‑face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state‑of‑the‑art performance on both fronts. Source code, pre‑trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync
Authors:Danyang Li, Tianhao Wu, Bin Li, Zhenyuan Chen, Yang Zhang, Yuxuan Li, Ming-Ming Cheng, Xiang Li
Abstract:
Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open set of object categories encountered in the real world. However, traditional closed‑set segmentation approaches struggle to adapt to complex open world scenarios, while foundation segmentation models such as SAM exhibit notable discrepancies between their strong segmentation capabilities and relatively weaker semantic understanding. To bridge these discrepancies, we propose WOW‑Seg, a Word‑free Open World Segmentation model for segmenting and recognizing objects from open‑set categories. Specifically, WOW‑Seg introduces a novel visual prompt module, Mask2Token, which transforms image masks into visual tokens and ensures their alignment with the VLLM feature space. Moreover, we introduce the Cascade Attention Mask to decouple information across different instances. This approach mitigates inter‑instance interference, leading to a significant improvement in model performance. We further construct an open world region recognition test benchmark: the Region Recognition Dataset (RR‑7K). With 7,662 classes, it represents the most extensive category‑rich region recognition dataset to date. WOW‑Seg attains strong results on the LVIS dataset, achieving a semantic similarity of 89.7 and a semantic IoU of 82.4. This performance surpasses the previous SOTA while using only one‑eighth the parameter count. These results underscore the strong open world generalization capabilities of WOW‑Seg. The code and related resources are available at https://github.com/AAwcAA/WOW‑Seg‑Meta.
Authors:Xi Liu, Weiwei Sun, Zhou Ren, Chris Broaddus, Siyu Huang, Laurent Guigues
Abstract:
Diffusion priors have recently demonstrated strong capability in enhancing the quality of sparse‑view 3D reconstruction by augmenting training views at novel viewpoints, but they inevitably introduce hallucinated content ‑‑ artifacts inconsistent with the input views ‑‑ into the final 3D model. To address this challenge, we propose Hallucination‑Aware Diffusion prior (HAD), which estimates pixel‑wise hallucination score maps for augmented images by leveraging multi‑view reasoning capabilities from a feedforward novel view synthesis (NVS) network pre‑trained on large‑scale 3D data. These hallucination scores enable selective masking of unreliable pixels during the progressive 3D reconstruction procedure, preventing the introduction of non‑existent artifacts into the 3D model. To further enhance performance, we create multiple versions of augmented images at each novel view by conditioning the diffusion prior on different input views, which are then fused into a final image that leverages the broader context across all input views. We show that our method substantially reduces hallucination artifacts in diffusion‑assisted 3D reconstruction, thereby achieving state‑of‑the‑art performance across multiple benchmarks on novel view synthesis. Our project are publicly available at \hrefhttps://xiliu8006.github.io/HAD‑Project‑website/project website.
Authors:Yachan Guo, JoseLuis Gomez Zurita, Danna Xue, Yi Xiao, AntonioManuel Lopez Pena
Abstract:
Although large‑scale visual foundation models (VFMs) achieve remarkable performance in semantic understanding, they still underperform in instance‑aware dense prediction tasks. They exhibit different biases in representation: for instance, promptable segmentation models (e.g., SAM2) focus on fine‑grained region boundaries, while self‑supervised models (e.g., DINOv3) emphasize object‑level structure. This observation highlights the potential of combining complementary features from different VFMs to enhance downstream dense prediction tasks. However, naive multi‑VFM fusion seldom leads to reliable gains, and interpretable principles for leveraging their complementary features are still underexplored. In this work, we propose a metric‑guided approach that effectively selects and aggregates complementary features from different VFMs based on explicit assessment scores. Specifically, we design a suite of label‑free metrics in feature space across two aspects, Structural Coherence and Edge Fidelity, to assess features of VFM encoders. Guided by these scores, we identify complementary edge‑strong and structure‑strong encoder pairs, and integrate them via a master‑auxiliary fusion scheme. This feature fusion requires no complex architectural changes and is trained only in a single stage. Our model shows consistent performance gains across multiple dense prediction tasks compared with the baselines, with better object‑level semantics and more accurately localized boundaries. The code is available at https://github.com/gyc‑code/metric‑guided‑fusion.
Authors:Wei Zhang, Songhua Li, Yihang Wu, Qiang Li, Qi Wang
Abstract:
3D change detection from multi‑view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods predominantly operate in the 2D domain, where viewpoint variations are mistaken for physical changes and depth is unavailable. While visual geometry foundation models like VGGT rapidly produce dense point clouds from unposed images, independent per‑epoch reconstruction encounters fundamental obstacles: unpredictable inter‑epoch scale ambiguity, registration‑change paradox where scene changes corrupt alignment, and pervasive edge‑flying noise. To address these challenges, we present VGGT‑CD, a training‑free pipeline decoupling cross‑temporal registration from dynamic‑change interference. In the Coarse Stage, sparse keyframe joint inference establishes a unified metric space and yields an initial Sim(3) prior. In the Fine Stage, dense reconstructions are purified by isolating static‑background correspondences. A closed‑form centroid alignment refines the translation while locking scale and rotation, using a residual self‑check to mathematically guarantee non‑degradation. Evaluated on an 11‑scene benchmark from the World Across Time dataset, VGGT‑CD reduces Absolute Trajectory Error by 44% outdoors and 59% indoors. It completes registration over 6 times faster, producing high‑purity 3D change maps without task‑specific training.
Authors:Yifei Pei, Ying Liu, Nam Ling
Abstract:
Learned image compression has achieved competitive rate‑distortion performance, but very‑low‑bitrate reconstruction remains difficult because the transmitted representation often cannot preserve fine textures and local structures. Perceptual and generative codecs address this problem by using learned reconstruction priors, and controllable codecs allow one model to cover different bitrate and reconstruction preferences. However, controllability alone does not resolve the decoder‑side reconstruction‑prior problem: under severe bit constraints, the decoder must infer missing details from limited transmitted information, while existing codebook‑based controllable designs generally rely on single‑codebook token‑based priors. This paper proposes Adaptive Fused Prior Transfer for Controllable Generative Image Compression (AFP‑GIC), a controllable codec that transfers an adaptive fused prior from a frozen pretrained AdaCode model. Encoder‑side fused‑prior features guide latent formation, while the decoder predicts a compatible fused prior from the compressed representation and selected control variables, enabling prior‑guided reconstruction without transmitting the fused prior itself. A motivating analysis relates decoder‑side fused‑prior alignment to a reconstruction‑error upper bound and shows that the fused‑prior family contains single‑codebook choices as special cases. Under the unified benchmark, AFP‑GIC reduces decoder latency by 18.1% and the overall parameter count by 31.10 million (20.5%) relative to DC‑VIC. Experiments on Kodak, CLIC2020, and DIV2K show competitive PSNR, with the clearest perceptual gains in NIQE scores and very‑low‑bitrate visual comparisons.
Authors:Darshana Rathnayake, Dulanga Weerakoon, Meera Radhakrishnan, Archan Misra
Abstract:
LiDARs are widely used for 3D depth reconstruction, but their performance is often limited by inherent hardware constraints that impose trade‑offs between range, spatial resolution, and frame rate. Many LiDAR systems typically operate at low frame rates (e.g., 5‑10 Hz), prioritizing long‑range sensing over responsiveness to rapid scene changes. We present NeuroLiDAR, an adaptive depth sensing framework that achieves effective frame rates of up to \approx66 Hz by fusing temporally sparse LiDAR data with temporally dense inputs from neuromorphic event cameras. NeuroLiDAR integrates two components: event‑based keyframe detection and event‑guided depth extrapolation, to dynamically adjust the sensing rate in response to scene dynamics. To evaluate our approach, we introduce ELiDAR, a dataset spanning outdoor and indoor scenarios, and show that NeuroLiDAR reduces depth reconstruction error by \approx29% in RMSE while achieving adaptive frame rates between 27.8‑47.3 Hz. Our code and dataset are available at https://github.com/darshanakgr/neurolidar.
Authors:Hwidong Kim, Yunho Kim, Tae-Kyun Kim
Abstract:
Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works such as PhysGen3D tackle single image‑to‑3D physics through mesh reconstruction and Physically‑Based Rendering, but challenges remain in modeling fluid dynamics, multi‑object interactions and photorealism. This work introduces 3DPhysVideo, a novel training‑free pipeline that generates physically realistic videos from a single image. We repurpose an off‑the‑shelf video model for two stages. First, we use it as a novel view synthesizer to reconstruct complete 360‑degree 3D scene geometry by guiding the image‑to‑video (I2V) flow model with rendered point clouds. Second, after applying physics solvers to this geometry, the physically simulated point cloud is used to guide the same I2V flow model to synthesize final, high‑quality videos. Consistency‑Guided Flow SDE, which decomposes the predicted velocity of the I2V flow model into denoising and consistency bias, enforces consistency to the conditional inputs, allowing us to effectively repurpose the model for both 3D reconstruction and simulation‑guided video generation. In the diverse experiments including multi‑objects, and fluid interaction scenes, our method successfully bridges the gap from single‑images to physically plausible videos, while remaining efficient to run on a single consumer GPU. It outperforms state‑of‑the‑art baselines on GPT‑based scores, VideoPhy benchmark and human evaluation.
Authors:Arpan Kusari
Abstract:
Hyperdimensional (HD) computing offers an attractive alternative to deep networks for edge learning due to its simplicity, fast prototype‑based inference, and compatibility with online updates. However, standard pixel‑based HD encoders are brittle: small distribution shifts such as rotation, noise, or occlusion can drastically reduce accuracy. We extract discrete topological primitives‑most notably holes‑from binarized shapes and pair them with rotation/translation/scale (RTS)‑invariant shape signatures. Our method constructs RTS‑stable descriptors for (i) the outer shape using a spatial‑pyramid variant of Zernike moments and (ii) each hole using an intrinsic Fourier descriptor of its radial signature together with RTS‑canonical relative geometry. Each primitive is mapped to a bipolar hypervector via randomized projection and role binding, and variable‑cardinality hole sets are aggregated by permutation‑invariant bundling to form a single image hypervector. To avoid over‑weighting any cue, we learn nonnegative reliability weights for the Zernike and hole channels on a validation set via late fusion of cosine similarities. Experiments on MNIST and EMNIST under controlled corruptions (rotation, Gaussian noise, salt‑and‑pepper, cutout, zoom) show that Topology‑guided HD computing substantially improves robustness compared with a naive HD baseline, maintaining high accuracy across multiple corruption families and benefiting from lightweight online training. Compared with a compact CNN trained on clean data, our method achieves competitive clean accuracy while offering markedly stronger robustness to several pixel‑level corruptions, demonstrating that explicit topological structure is a practical route to robust HD representations. The code is provided at https://github.com/arpan‑kusari/Topological‑HDC.
Authors:Mingyang Zhao, Sipu Ruan, Xiaohong Jia
Abstract:
This work presents a novel method for fitting superquadrics to point clouds under the contamination of noise and outliers, which has many applications for shape modeling across diverse fields. Unlike prior approaches that either exclusively focus on fitting rigid or deformable superquadrics, or suffer from robustness and numerical instability issues, our method redefines the problem from a new unsupervised clustering perspective, enabling the holistic fitting of both rigid and deformable superquadrics within a unified framework. Central to our approach is a stable optimization function inspired by unsupervised clustering analysis, where we formulate the point cloud data and samples from the potential parametric surface as clustering members and centroids, respectively. Then, the clustering process with dynamic updates to centroid locations serves as a direct proxy for optimizing superquadric parameters, establishing a principled link between geometric fitting and clustering dynamics. We further derive the relationship between pairwise computations of clustering centroids and clustering members to orthogonal distances, effectively eliminating the need for the time‑consuming surface sampling process. Moreover, our formulation provides closed‑form analytical solutions for both the fuzzy membership degree vector and the covariance matrix, ensuring efficient iteration optimization and enabling more effective handling of geometric deformations. In addition, we provide a theoretical certificate of convergence analysis and demonstrate that the clustering‑inspired fitting method can escape local minima by inherently increasing the convexity of the objective function. The implementation is publicly available at https://github.com/zikai1/SuperquadricFitting.
Authors:Feng Gao, Zhilin Jin, Yanhai Gan, Junyu Dong, Qian Du
Abstract:
Semantic segmentation of multi‑source remote sensing images is a fundamental task for Earth observation applications. Existing methods often struggle with insufficient multi‑scale context modeling and suboptimal cross‑modal feature fusion, limiting their performance in complex high‑resolution scenes. To this end, we propose Axial‑Relation Guided Fusion Mamba (ARG‑Mamba), a state space model‑based framework for optical‑elevation remote sensing image segmentation. Specifically, we introduce a Multi‑Scale State Space Module to capture both fine‑grained local details and global contextual dependencies with linear computational complexity. Moreover, an Axial‑Relation Guided Fusion Module is designed to explicitly model global cross‑modal correlations along horizontal and vertical axes, enabling efficient feature fusion between optical and elevation modalities. Extensive experiments conducted on the ISPRS Vaihingen and Potsdam datasets demonstrate that our ARG‑Mamba consistently outperforms state‑of‑the‑art methods while maintaining favorable computational efficiency. The code will be made publicly available at \urlhttps://github.com/oucailab/ARG‑Mamba.
Authors:Baogui Huan, Chuanzheng Gong, Dezhong Chen, Feng Gao, Junyu Dong, Qian Du
Abstract:
Convolutional neural networks (CNNs) have been extensively and successfully applied to the task of synthetic aperture radar (SAR) image change detection. However, conventional convolutional layers are inherently limited by their local receptive fields, which mainly capture spatially localized patterns while neglecting the global context that is often crucial for accurately distinguishing subtle or large‑scale changes in SAR imagery. To address these limitations, we propose a novel Global Dynamic Context‑Aware Network (GDNet) specifically tailored for SAR image change detection. At the core of our approach lies a novel global dynamic convolution module, which adaptively modulates convolution kernel weights according to the global semantic information extracted from the input features. By dynamically incorporating long‑range dependencies, this mechanism enables the network to integrate both local detail and global context, thus improving its ability to detect diverse change patterns. In addition, we introduce a carefully designed two‑stage Mixup strategy for model training. Unlike conventional single‑stage Mixup, our two‑stage design generates more diverse and informative training samples, effectively regularizing the model and yielding more stable and reliable classification results even under limited data scenarios. Extensive experiments on three SAR datasets demonstrate the superiority of the proposed GDNet compared to other state‑of‑the‑art methods. These findings highlight the potential of global dynamic modeling and advanced data augmentation strategies for advancing SAR image interpretation. Source codes are available at \urlhttps://github.com/oucailab/GDNet.
Authors:Anuska Roy, Pravin Nair
Abstract:
Flow and diffusion models achieve high‑fidelity, high‑resolution image synthesis, but often require many function evaluations (NFEs) at sampling time. Existing acceleration methods either require additional training through distillation or rely on training‑free high‑order solvers, and both can degrade sample quality at low NFE budgets. We propose CAB (Corrected Adams‑Bashforth), a training‑free sampler that accelerates both flow and diffusion models. CAB first transforms the sampling dynamics to a common rectified coordinate system, and then applies a multistep Adams‑Bashforth predictor augmented with a simple correction term based on past velocity evaluations and therefore incurs no additional NFEs. The resulting method is simple, has the same algorithmic form across model classes, and has at least third‑order local truncation error and second‑order global error. Experiments on pretrained flow and diffusion models, including class‑conditional and large‑scale text‑to‑image benchmarks, show that CAB improves quality‑NFE trade‑offs in the low‑step regime of 6‑20 NFEs. It also remains competitive with strong training‑free samplers at higher step counts across most tested models. The official implementation is available at https://github.com/Anuska‑Roy/CAB.
Authors:Marawan Elbatel, Mohamed Ghonim, Jiaji Mao, Zhuosheng Lin, Katharina Eckstein, Andrés Martínez Mora, Jonathan Deissler, Maximilian Rokuss, Constantin Ulrich, Zdravko Marinov, Wenhui Deng, Baoxun Li, Huijun Hu, Jun Shen, Mohanad Ghonim, Khadiga Omar Nassar, Mariam Elbakry, Menna Dyab, Amr Muhammad Abdo Salem, Nouran Elghitany, Noha Elghitany, Yi Qin, Xuanqi Huang, Haonan Wang, Shao-Woo Yen, Ahmed Elghamry Saba, Salma Ahmad, Xinyan Fang, Jiahao Zhang, Xiaodi Wang, Xinghua Ma, Gongning Luo, Jessica C. Delmoral, João Manuel R. S. Tavares, Ankan Deria, Adinath Dukre, Yutong Xie, Imran Razzak, Dongwook Kim, Matthew Choi, Hanxiao Zhang, Minghui Zhang, Xin You, Abdul Qayyum, Steven A. Niederer, Moona Mazher, Rachika E. Hamadache, Ricardo Montoya-del-Angel, Robert Martí, Xavier Lladó, Toufiq Musah, Livingstone Eli Ayivor, Enrique Almar-Munoz, Agnes Mayr, Kaouther Mouheb, Esther E. Bron, Stefan Klein, Ahmed Abouelhoda, Amira Adel, Susan Adil Ali, Rainer Stiefelhagen, Klaus H. Maier-Hein, Fabian Isensee, Aya Yassin, Xiaomeng Li
Abstract:
Automated segmentation of liver lesions on non‑contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly in low‑resource settings across Africa and Asia where contrast agents are frequently unavailable. Progress has been limited by the absence of annotated NCCT benchmarks. Here we describe the TriALS challenge for automated liver lesion segmentation under contrast‑limited conditions, supported by a multi‑centre dataset of 150 cases with four‑phase CT acquisitions (600 volumes) from Egyptian and Chinese institutions. Algorithms were evaluated on 70 cases from three institutions, including an independent external cohort. The top‑performing method achieved a mean venous‑phase Dice of 0.754, consistent with human‑level performance, yet dropped to 0.57 on NCCT. On external validation, the leading method outperformed off‑the‑shelf models by up to 28% in Dice on NCCT. Algorithm performance was most strongly predicted by training data scale and pre‑training strategy. A cross‑year comparison exposed a persistent perceptual barrier on NCCT that scaling pre‑training alone cannot overcome. Data, annotations, and code are available at https://github.com/xmed‑lab/TriALS.
Authors:Ssharvien Kumar Sivakumar, Akwele Johnson, Anirudh Dhingra, Yannik Frisch, Ghazal Ghazaei, Anirban Mukhopadhyay
Abstract:
Realistic surgical simulation plays a crucial role in training novice surgeons and in the development of autonomous agents. World models can scale such simulation environments to realistic and diverse procedures by predicting future patient states conditioned on current observations and surgical actions. However, current state‑of‑the‑art approaches often fail to satisfy key criteria required for clinical applicability, including visual realism, physically grounded interactions, and the ability to simulate scenarios beyond the training distribution. Hence, we introduce SWoMo, a neuro‑symbolic world model for cataract surgery simulation that decouples motion generation from visual realism. The symbolic component, consisting of a rule‑based simulator and scene graph representations, models motion dynamics and tool‑tissue interactions, while a diffusion model produces realistic visual appearance, including textures and tissue deformations. We propose an inverse pairing strategy that reconstructs real surgical videos in the simulator to obtain paired simulated and real videos, which are then used to train our video diffusion model for the reverse objective of sim‑to‑real translation. Our experiments show both qualitative and quantitative improvements over prior work. We demonstrate that our simulator further satisfies the key criteria, including generalisation to unseen interaction geometries, improvements in downstream phase detection, and unsupervised video style transfer. The code, data, and model weights are available at: https://ssharvienkumar.github.io/SWoMo/
Authors:Zhuoyu Wu, Wenhui Ou, Lexi Zhang, Pei-Sze Tan, Dongjun Wu, Junhe Zhao, Wenqi Fang, Raphaël C. -W. Phan
Abstract:
Accurate polyp segmentation in colonoscopy is essential for early colorectal cancer detection, yet real‑world clinical environments pose persistent challenges such as motion blur, specular reflections, and illumination instability. Most existing methods are optimized on clean benchmark images and suffer noticeable performance degradation when deployed in authentic surgical scenarios. We propose DepthPolyp, a lightweight and robust segmentation framework based on pseudo‑depth‑guided multi‑task learning and efficient feature modulation. The architecture combines hierarchical Ghost factorization for compact feature generation, Interleaved Shuffle Fusion for low‑cost cross‑scale interaction, and Dynamic Group Gating for adaptive group‑wise feature weighting. Extensive experiments demonstrate that DepthPolyp achieves strong cross‑dataset generalization when trained on degraded data and evaluated on both clean and noisy target domains, consistently outperforming lightweight baselines and remaining competitive with substantially larger models. In real surgical video evaluation on PolypGen, DepthPolyp achieves better segmentation performance than models up to 20× larger while preserving real‑time inference speed. With only 3.57M parameters and 0.86 GMACs, the proposed method runs at over 180 FPS on mobile devices, making it well suited for real‑time deployment in resource‑constrained clinical environments. Code and pretrained weights are available at: https://github.com/ReaganWu/DepthPolyp/
Authors:Amin Karimi Monsefi, Abolfazl Meyarian, Mridul Khurana, Shuheng Wang, Pouyan Navard, Cheng Zhang, Anuj Karpatne, Wei-Lun Chao, Rajiv Ramnath
Abstract:
Animals are described as effectively camouflaged when they blend seamlessly with their surrounding, yet no standardized quantitative measure of this seamlessness exists. We address this gap by framing camouflage evaluation as a visual localization problem: a well‑camouflaged animal is one that remains difficult to detect even when its category is known. We introduce SeamCam (Seamless Camouflage), a metric that quantifies how detectable an animal is from the available visual evidence. Given an image and a target species, SeamCam generates category‑conditioned detection proposals, extracts segmentation masks, and identifies the subset whose collective union yields the highest IoU with the ground‑truth mask. The SeamCam score is one minus this maximum recoverable localization signal, where a higher score indicates stronger camouflage (i.e., lower detectability). In a human two‑alternative forced‑choice study with 94 participants and 2,390 comparisons, SeamCam achieves 78.82% agreement with human camouflage difficulty judgments, outperforming state‑of‑the‑art by about 25%. We then demonstrate SeamCam's utility as a preference signal for Direct Preference Optimization (DPO) to fine‑tune a diffusion‑based inpainting model for camouflage generation. This offers an affordable training approach with an objective explicitly suited for camouflage generation, unlike typical diffusion models. To support rigorous benchmarking, we further introduce CamFG‑1.5k, a curated dataset of 1,521 high‑resolution images in which animals are fully visible prior to camouflage generation, enabling unbiased evaluation by controlling for occlusion artifacts present in existing datasets. https://7amin.github.io/SeamCam/
Authors:Aiden Yiliu Li, Nels Numan, Anthony Steed
Abstract:
Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps it searchable over long horizons, and grounds later reasoning in recoverable observations rather than compressed latent state alone. We propose Visual Agentic Memory (VAM), a training‑free framework with three components. Online Indexing supports selective evidence retention under streaming constraints. Hierarchical Memory organises retained evidence in a Parallel Representation that aligns temporal context with spatial observations. Agentic Retrieval searches, inspects, and verifies candidate evidence before producing a grounded answer. On OVO‑Bench, VAM achieves the highest RT+BT average (68.41) across all reported baselines, improving over end‑to‑end use of the same underlying MLLM (Gemini 3 Flash, 67.46). On the month‑scale split of MM‑Lifelong train@month (105.6 hours over 51 days), VAM reaches 17.11%, second only to ReMA with GPT‑5 (17.62%). These results suggest that long‑horizon video understanding benefits from treating visual memory as an explicit, inspectable, and queryable substrate. Code is available at https://github.com/yiliu‑li/Visual‑Agentic‑Memory.
Authors:Youngin Kim, Ray Sun, Inho Kim, Bumsoo Park, Hyun Oh Song
Abstract:
Token‑based transformer world models have shown strong performance in visual reinforcement learning, but often suffer from temporal inconsistency in long‑horizon rollouts, including object duplication, disappearance, and transmutation. A key reason is that most existing approaches treat next‑frame prediction purely as a token generation problem, without considering the persistence of tokens across time. We introduce Identifiable Token Correspondence (ITC), a decoding step for token‑based transformer world models that formulates next‑frame prediction as a structured assignment problem with latent token correspondence variables: each next‑frame token is explained either by copying a token from the previous frame or by generating a new one. ITC leaves the transformer architecture and training procedure unchanged and can be added on top of existing backbones. Our experiments show state‑of‑the‑art performance on 4 challenging benchmarks. The proposed method achieves a return of 72.5% and a score of 35.6% on the Craftax‑classic benchmark, significantly surpassing the previous best of 67.4% and 27.9%. We release our source code on https://github.com/snu‑mllab/Identifiable‑Token‑Correspondence.
Authors:Yousra Nabila Taifour, Marouane Tliba, Zuheng Ming, Marie Luong, Nour Aburaed, Aladine Chetouani, Gorkem Durak, Alessandro Bruno, Faouzi Alaya Cheikh, Habib Zaidi, Ulas Bagci, Azeddine Beghdadi
Abstract:
Computed tomography (CT) images are frequently degraded by acquisition artifacts, including noise, blur, streaking, aliasing, and metal artifacts. Yet CT enhancement is still largely evaluated using image quality metrics with limited perceptual and clinical validity, while existing datasets remain focused on isolated restoration tasks, hindering unified benchmarking across diverse degradation types. We present CT‑DegradBench, a dataset and benchmark for CT degradation detection and severity estimation under controlled single‑ and mixed‑artifact settings. CT‑DegradBench enables systematic evaluation across multiple degradation families and severity levels within a common experimental framework. We further propose SeSpeCT (Semantic‑Spectral CT degradation estimation), a framework that combines semantic priors from medical vision‑language models with complementary frequency‑domain cues for artifact analysis. SeSpeCT constructs a training‑free semantic quality axis in the multimodal embedding space using radiology‑informed text prompts, without task‑specific fine‑tuning, and combines it with spectral features that capture degradation‑specific frequency patterns. The resulting representation enables joint prediction of artifact type and severity. Experimental results show that SeSpeCT consistently outperforms the evaluated baselines under both single‑ and mixed‑degradation settings. The framework is available at https://github.com/yousranb/CT‑DEGRADBENCH.
Authors:Haoren Zhao, Tianyi Chen, Zhen Wang
Abstract:
Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single‑layer interfaces. This paper identifies a critical reliability gap: state‑of‑the‑art agents face distinct robustness challenges in real‑world desktop environments characterized by multi‑window stacking, occlusion, and visual clutter. To address this, we introduce WinDeskGround, a novel benchmark and synthesis framework tailored for evaluating GUI grounding robustness. Unlike static datasets, our framework parametrically generates complex desktop scenarios by controlling window occlusion, layout density, and semantic similarity, thereby simulating the distribution shifts of authentic workflows. We construct a diverse meta‑dataset of 1,356 high‑fidelity instruction‑target pairs and conduct comprehensive evaluations of five leading MLLMs. Our results demonstrate that while top‑tier agents excel in simplified settings, their accuracy declines under partial occlusion. WinDeskGround provides a valuable benchmark to facilitate the assessment and advancement of GUI agent robustness in realistic environments. The code is available at https://github.com/ZZZhr‑1/WinDeskGround.
Authors:Jinhao Jing, Zheng Ma, Jinwei Liang, Qiannian Zhao, Shawn Chen, Jing Yang, Por Lip Yee, Prayag Tiwari, Jingjing Bai, Benyou Wang, Lewei Lu, Zhan Su
Abstract:
Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain‑of‑Thought (CoT) data. To address this, we propose the GeoSym Engine, an automated and scalable neuro‑symbolic framework. By leveraging a type‑conditional grammar and an analytic SymGT Solver, it derives exact symbolic ground truths and seamlessly integrates with a robust rendering pipeline to produce high‑precision geometric diagrams. Using this engine, we construct GeoSym127K, a difficulty‑stratified dataset featuring 51K high‑resolution images, 127K questions with symbolic ground truths, and 55K answer‑verified CoT QA pairs. We also introduce GeoSym‑Bench, an expert‑curated suite of 511 complex samples for rigorous evaluation. Through extensive supervised fine‑tuning (SFT), we demonstrate that GeoSym drives concentrated improvements specifically on diagram‑dependent and multi‑step geometry tasks. Our Qwen3‑VL‑8B model gains an absolute +22.21% on the MathVerse Vision‑Only subset and reaches 61.52% (+6.19% improvement) on WeMath, mitigating long‑horizon logic fragmentation and outperforming advanced closed‑source models like Doubao‑1.8. Furthermore, applying Reinforcement Learning with Verifiable Rewards (RLVR) via GRPO reveals that initializing from structural SFT checkpoints substantially elevates the performance ceiling over zero‑shot RL. Driven by deterministic exact‑match signals, this showcases the robust scaling potential of our verifiable reasoning synthesis. Datasets and code are available at https://huggingface.co/datasets/Tomie0506/GeoSym127K and https://github.com/Tomie56/GeoSym127K.
Authors:Chang Che, Ziqi Wang, Hui Ma, Cheems Wang, Zenglin Shi
Abstract:
Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models to incrementally acquire new abilities. However, existing CVIT methods operate under a restrictive task‑incremental setting, where each training phase corresponds to a single, predefined task. This does not reflect real‑world conditions, where data arrives as a continuous stream of interleaved and dynamically evolving tasks. To bridge this gap, we introduce Streaming CVIT (StrCVIT), a more general and realistic setting where models learn from a stream of data chunks containing a dynamic mixture of tasks. In StrCVIT, a model must simultaneously acquire new abilities, reinforce recurring abilities, and mitigate forgetting. Existing CVIT methods fail here as they cannot reliably distinguish or adapt to the heterogeneous task samples within each chunk. We therefore propose StrLoRA, a regularized two‑stage expert routing framework. StrLoRA first performs task‑aware expert selection using the textual instruction to activate a sparse subset of relevant experts, reducing cross‑task interference. It then applies token‑wise expert weighting within this subset, where contribution weights are computed via cross‑modal attention between local visual tokens and the global instruction representation. To maintain stability across the non‑stationary stream, a routing‑stability regularization aligns current routing distributions with a historical exponential moving average reference. Extensive experiments on a newly developed StrCVIT benchmark show that StrLoRA substantially outperforms existing methods, effectively enhancing model's abilities from continuously evolving data streams. The code is available at https://github.com/chanceche/StrCVIT.
Authors:Ruichen Zheng, Biao Zhang, Michael Birsak, Mikhail Skopenkov, Peter Wonka
Abstract:
We introduce Patchwork, a new general‑purpose shape representation capable of modeling 2D and 3D geometry with a small number of parameters. Patchwork is grounded in a rigorous mathematical framework, providing provable complexity bounds and the ability to approximate arbitrary shapes with arbitrary precision in any dimension. We propose an efficient gradient‑based optimization scheme to fit Patchwork representations to 2D and 3D data, along with a novel regularization loss that progressively prunes redundant elements, yielding high compactness after convergence. Our approach offers fast fitting performance, a fraction of the required parameters compared to existing alternatives, and native support for inside‑outside classification, making it a versatile and compact representation for geometric learning and reconstruction tasks, with future potential for 3D generation. Our implementation is available at: https://github.com/Ankbzpx/patchwork‑experiment.
Authors:Yuqi Wu, Tianyu Hu, Wenzhao Zheng, Yuanhui Huang, Haowen Sun, Jie Zhou, Jiwen Lu
Abstract:
Reconstructing coherent 3D geometry and appearance from unposed multi‑view images is a fundamental yet challenging problem in computer vision. Most existing visual geometry foundation models predict explicit geometry by regressing pixel‑aligned pointmaps, often suffering from redundancy and limited geometric continuity. We propose IVGT, an Implicit Visual Geometry Transformer that implicitly models continuous and coherent geometry from pose‑free multi‑view images. This formulation learns a continuous neural scene representation in a canonical coordinate system and supports continuous spatial queries at any 3D positions, retrieving local features to predict signed distance (SDF) values and colors using lightweight decoders. It allows direct extraction of continuous and coherent surface geometry, enabling rendering of RGB images, depth maps, and surface normal maps from arbitrary viewpoints. We train IVGT via multi‑dataset joint optimization with 2D supervision and 3D geometric regularization. IVGT demonstrates generalization across scenes and achieves strong performance on various tasks, including mesh and point cloud reconstruction, novel view synthesis, depth and surface normal estimation, and camera pose estimation.
Authors:Xinyue Liu, Jianyuan Wang, Biao Leng, Shuo Zhang
Abstract:
Few‑shot Generalist Anomaly Detection requires models to generalize to novel categories without retraining, posing significant challenges in real‑world scenarios with scarce samples and rapidly changing categories. Existing CLIP‑based methods face two major challenges: coarse‑grained unified text prompts struggle to adapt to fine‑grained foreground‑background differences, causing cross‑granularity mismatch; and fine‑tuning on auxiliary datasets disrupts CLIP's inherent open‑world generalization due to domain shift, leading to cross‑category generalization degradation. To address these, we propose to shift multimodal alignment entirely into a unified residual space, where residual representations naturally eliminate fine‑grained normal feature differences across regions and class‑specific biases, simultaneously resolving both problems. Based on this insight, Res^2CLIP, the first residual‑to‑residual alignment framework that symmetrically bridges visual and text modalities within CLIP's residual space, is designed. The framework is developed from a residual perspective into three branches: a text prompt‑based branch, a visual prompt‑based branch, and a novel residual‑to‑residual alignment branch. All learnable optimizations are constrained within the residual domain, and the residual alignment optimization objectives are designed to force the model to focus on relative anomaly deviations rather than optimizing class‑specific features. Experiments on multiple datasets demonstrate the effectiveness of our architecture. The code is available at https://github.com/hito2448/Res2CLIP.
Authors:Zhipei Xu, Xuanyu Zhang, Youmin Xu, Qing Huang, Shen Chen, Taiping Yao, Shouhong Ding, Jian Zhang
Abstract:
Diffusion‑based image synthesis has made AI‑generated images (AIGI) increasingly photorealistic, raising urgent concerns about authenticity in applications such as misinformation detection, digital forensics, and content moderation. Despite the substantial advances in AIGI detection, how to correct detected AI‑generated images with visible artifacts and restore realistic appearance remains largely underexplored. Moreover, few existing work has established the connection between AIGI detection and artifact correction. To fill this gap, we propose GenShield, a unified autoregressive framework that jointly performs explainable AIGI detection and controllable artifact correction in a closed loop from diagnosis to restoration, revealing a mutually reinforcing relationship between these two tasks. We further introduce a Visual Chain‑of‑Thought based curriculum learning strategy that enables self‑explained, multi‑step ``diagnose‑then‑repair'' correction with an explicit stopping criterion. A high‑quality dataset with large‑scale ``artifact‑restored'' pairs is also constructed alongside a unified evaluation pipeline. Extensive experiments on our correction benchmark and mainstream AIGI detection benchmarks demonstrate state‑of‑the‑art performance and strong generalization of our method. The code is available at https://github.com/zhipeixu/GenShield.
Authors:Yiming Zhao, Yu Zeng, Wenxuan Huang, Zhen Fang, Qing Miao, Qisheng Su, Jiawei Zhao, Jiayin Cai, Lin Chen, Zehui Chen, Yukun Qi, Yao Hu, Xiaolong Jiang, Feng Zhao
Abstract:
Large Vision‑Language Models (LVLMs) have shown significant progress in video understanding, yet they face substantial challenges in tasks requiring precise spatiotemporal localization at the instance level. Existing methods primarily rely on text prompts for human‑model interaction, but these prompts struggle to provide precise spatial and temporal references, resulting in poor user experience. Furthermore, current approaches typically decouple visual perception from language reasoning, centering reasoning around language rather than visual content, which limits the model's ability to proactively perceive fine‑grained visual evidence. To address these challenges, we propose VideoSeeker, a novel paradigm for instance‑level video understanding through visual prompts. VideoSeeker seamlessly integrates agentic reasoning with instance‑level video understanding tasks, enabling the model to proactively perceive and retrieve relevant video segments on demand. We construct a four‑stage fully automated data synthesis pipeline to efficiently generate large‑scale, high‑quality instance‑level video data. We internalize tool‑calling and proactive perception capabilities into the model via cold‑start supervision and RL training, building a powerful video understanding model. Experiments demonstrate that our model achieves an average improvement of +13.7% over baselines on instance‑level video understanding tasks, surpassing powerful closed‑source models such as GPT‑4o and Gemini‑2.5‑Pro, while also showing effective transferability on general video understanding benchmarks. The relevant datasets and code will be released publicly.
Authors:Mingqiang Wu, Weilun Feng, Zhefeng Zhang, Haotong Qin, Yuqi Li, Guoxin Fan, Xiaokun Liu, Zhulin An, Libo Huang, Yongjun Xu, Chuanguang Yang
Abstract:
Autoregressive video diffusion models enable open‑ended generation through local attention and KV caching. However, existing training‑free long‑video optimization methods mainly focus on stable extension under a single prompt, making them difficult to handle interactive scenarios involving prompt switching, old scene forgetting, and historical scene recall. We identify the core bottleneck as the functional entanglement of historical KV states: stable anchors and recent dynamics are handled by the same cache policy, leading to outdated background contamination, delayed response to new prompts, and loss of long‑range memory. To address this issue, we propose Echo‑Forcing, a training‑free scene memory framework specifically designed for interactive long video generation with three core mechanisms: (1) Hierarchical Temporal Memory, which decouples stable anchors, compressed history, and recent windows under relative RoPE; (2) Scene Recall Frames, which compresses historical scenes into spatially structured KV representations to support long‑term recall; and (3) Difference‑aware Memory Decay, which adaptively forgets conflicting tokens according to the discrepancy between old and new scenes. Based on these designs, Echo‑Forcing uniformly supports smooth transitions, hard cuts, and long‑range scene recall under a bounded cache budget. Extensive evaluations on VBench‑Long further demonstrate that Echo‑Forcing achieves the best overall performance in both long‑video generation and interactive video generation settings. Our code is released in https://github.com/mingqiangWu/Echo‑Forcing
Authors:Baining Zhao, Jiacheng Xu, Weicheng Feng, Xin Zhang, Zhaolu Wang, Haoyang Wang, Shilong Ji, Ziyou Wang, Jianjie Fang, Zhiheng Zheng, Weichen Zhang, Yu Shang, Wei Wu, Chen Gao, Xinlei Chen, Yong Li
Abstract:
Aerial vision‑language navigation (VLN) requires agents to follow natural‑language instructions through closed‑loop perception and action in 3D environments. We argue that aerial VLN can be formulated as a prediction‑driven world‑action problem: the agent should anticipate latent world evolution and act according to the predicted consequences. To this end, we propose WorldVLN, the first autoregressive world action model for aerial VLN. Unlike full‑sequence video‑generation world models that generate an entire visual clip, WorldVLN adapts a latent autoregressive video backbone to predict short‑horizon world‑state transitions and directly decodes them into executable waypoint actions. After each action segment is executed, newly received observations are encoded back into the autoregressive context, enabling closed‑loop world‑action prediction. We further introduce a two‑stage training framework that first grounds the video prior in instruction‑conditioned navigation dynamics and then develops Action‑aware GRPO, the first reinforcement learning method tailored to autoregressive WAMs, to optimize waypoint decisions through their downstream rollout consequences. On public outdoor and indoor benchmarks, WorldVLN consistently outperforms existing Vision‑Language‑Action baselines with 12%+ success‑rate gains and larger advantages on challenging cases. It further transfers zero‑shot to real drone deployment, suggesting that the proposed WorldVLN offers a promising route for spatial action tasks. Demos and code are available at https://embodiedcity.github.io/WorldVLN/.
Authors:Fabian Morelli, Arnas Uselis, Ankit Sonthalia, Seong Joon Oh
Abstract:
Large‑scale pre‑trained vision‑language models like CLIP demonstrate remarkable zero‑shot performance across diverse tasks. However, fine‑tuning these models to improve downstream performance often degrades robustness against distribution shifts. Recent approaches have attempted to mitigate this trade‑off, but often rely on computationally expensive text‑guidance. We propose a novel method for robust fine‑tuning, SAE‑FT, which operates only on the model's visual representations. SAE‑FT regularizes changes to these representations by penalizing the addition and removal of semantically meaningful features identified by a Sparse Autoencoder trained on the pre‑trained model. This constraint prevents catastrophic forgetting and makes the fine‑tuning process interpretable, enabling direct analysis of semantic changes. SAE‑FT is both mechanistically transparent and computationally efficient, matching or exceeding state‑of‑the‑art performance on ImageNet and its associated distribution shift benchmarks. Code is publicly available at: https://github.com/Fabian‑Mor/sae‑ft.
Authors:Yuyuan Liu, Yiping Ji, Anjie Le, Jiayuan Zhu, Jiazhen Pan, Can Peng, Jiajun Deng, Fengbei Liu, Junde Wu
Abstract:
Finetuning Large Vision‑Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object‑level grounding. However, existing methods, mainly based on GRPO, assign rewards at the response level. Such sparse reward, often criterion‑induced, leads to minimal learning signals when all candidate responses fail in challenging scenarios. In this work, we propose a group‑revision optimisation paradigm that enhances learning on hard cases. It begins with a sampled initial response and generates a set of revised candidates to explore improved grounding outcomes. Inspired by reward shaping, we introduce a consolidation process that quantifies each candidate's improvement over the initial attempt and converts it into informative shaping signals. These signals are used to both refine the reward and modulate the advantage, amplifying the influence of high‑quality revisions. Our method achieves consistent gains across referring and reasoning segmentation, REC, and counting benchmarks compared with prior GRPO‑based models. Our code is available at https://github.com/yyliu01/GroupRevision.
Authors:Leyang Chen, Junyi Wu, Zhiteng Li, Yulun Zhang
Abstract:
Streaming 3D reconstruction from long monocular video sequences requires maintaining a key‑value (KV) cache that grows linearly with sequence length, creating a severe memory bottleneck. Existing approaches either truncate the cache to a fixed set of anchor frames, leading to reconstruction quality degradation, or rely on attention‑score heuristics that are agnostic to 3D scene structure, failing to preserve geometrically valuable tokens. To address these problems, we present GHOST (Geometry‑Hierarchical Online Streaming Token Eviction), a training‑free KV cache management framework that exploits the model's own 3D geometry outputs to evict redundant tokens online. GHOST introduces three mutually reinforcing innovations: a hierarchical dual‑level importance scoring scheme, a privilege mechanism that protects special tokens from eviction, and a cosine‑similarity‑guided layer‑wise budget allocation. Experiments on various benchmarks show that GHOST preserves excellent reconstruction quality while cutting the KV cache by nearly half and delivering 1.75x faster inference compared to state‑of‑the‑art methods. Our code is available at https://github.com/lokiniuniu/GHOST.
Authors:Jichen Hu, Jiawei Guo, Jiazhong Cen, Chen Yang, Sikuang Li, Wei Shen
Abstract:
Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outputs are typically static monolithic assets with limited editability and physical interaction. This restricts their use in immersive content creation and embodied simulation, where generated worlds must be actively modified and manipulated. To tackle this challenge, we present WorldAct, a framework that converts static generated 3D worlds into editable and interaction‑ready scenes. WorldAct uses a multimodal agent to guide scene decomposition, identify actionable objects, reconstruct geometrically aligned object‑level meshes for interaction, and restore the residual background via 3D inpainting. The resulting scenes support object‑level editing, collision‑aware manipulation, and embodied task execution while preserving global scene coherence. Experiments show that WorldAct enables richer interaction scenarios than the original generated scenes, suggesting a practical path toward editable and interactive 3D world models.
Authors:Yipu Zhang, Jintao Cheng, Weilun Feng, Jiehao Luo, Chuanguang Yang, Zhulin An, Yongjun Xu, Wei Zhang
Abstract:
Feed‑forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera pose prediction, and point cloud reconstruction in a single forward pass. They have been widely adopted in 3D vision applications, but their billion‑scale parameters bring substantial memory and computation overhead, posing challenges for on‑device deployment. Post‑Training Quantization (PTQ) is an effective technique to reduce this overhead. Existing PTQ methods for feed‑forward 3D models mainly focus on handling heavy‑tailed activation distributions and constructing diverse calibration datasets. However, we observe that feed‑forward 3D models predict multiple geometric attributes through a shared backbone, where different transformer blocks and hidden channels contribute distinctly to each task, resulting in substantially different sensitivities to quantization errors across tasks, blocks, and channels. Consequently, treating all tasks equally over‑emphasizes insensitive tasks and causes significant accuracy loss on the sensitive ones. To address this issue, we propose Fisher‑Guided Quantization (FGQ) for feed‑forward 3D reconstruction models. Specifically, FGQ uses the diagonal Fisher information matrix to quantify the different sensitivities across tasks, blocks, and channels, and incorporates these sensitivities into the Learnable Affine Transformation during calibration to better preserve the channels and blocks most critical to each task. Extensive experiments across camera pose estimation, point map reconstruction, and depth estimation show that FGQ consistently outperforms state‑of‑the‑art quantization baselines on VGGT, achieving up to 39% relative improvement under the 4‑bit quantization. Code is available at https://github.com/ypzhng/FGQ.
Authors:Quanjian Song, Yefeng Shen, Mengting Chen, Hao Sun, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Liujuan Cao
Abstract:
Human‑centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low‑latency and interactive garment control, which is crucial for applications such as e‑commerce and content creation. This paper studies how to achieve interactive multi‑garment video customization while preserving motion coherence using only single‑garment video data. We present FashionChameleon, a real‑time and interactive framework for human‑garment customization in autoregressive video generation, where users can interactively switch garment during generation. FashionChameleon consists of three key techniques: (i) Instead of training on multi‑garment video data, we train a Teacher Model with In‑Context Learning on a single reference‑garment pair. By retaining the image‑to‑video training paradigm while enforcing a mismatch between the reference and garment image, the model is encouraged to implicitly preserve coherence during single‑garment switching. (ii) To achieve consistency and efficiency during generation, we introduce Streaming Distillation with In‑Context Learning, which fine‑tunes the model with in‑context teacher forcing and improves extrapolation consistency via gradient‑reweighted distribution matching distillation. (iii) To extend the model for interactive multi‑garment video customization, we propose Training‑Free KV Cache Rescheduling, which includes garment KV refresh, historical KV withdraw, and reference KV disentangle to achieve garment switching while preserving motion coherence. Our FashionChameleon uniquely supports interactive customization and consistent long‑video extrapolation, while achieving real‑time generation at 23.8 FPS on a single GPU, 30‑180× faster than existing baselines.
Authors:Junho Kim, Xu Cao, Houze Yang, Bikram Boote, Ana Jojic, Fiona Ryan, Bolin Lai, Sangmin Lee, James M. Rehg
Abstract:
Understanding social interactions requires reasoning over subtle non‑verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with whom in multi‑person videos. We introduce GRASP, a large‑scale social reasoning dataset that connects high‑level social QA with fine‑grained gaze and deictic gesture events. GRASP contains 290K question‑‑answer pairs over 46K videos totaling 749 hours, organized by a 16‑category taxonomy spanning gaze, gesture, and joint gaze‑‑gesture reasoning, together with GRASP‑Bench for evaluation. Unlike prior resources that focus on either isolated cues or high‑level social QA, GRASP builds questions from identity‑consistent gaze trajectories, deictic gestures, and their joint compositions into social events. Moreover, we propose Social Grounding Reward (SGR), a learning signal that uses these social events to encourage models to reason about the participants involved in each interaction. Experiments show that SGR improves performance on GRASP‑Bench while maintaining zero‑shot performance on related social video QA benchmarks.
Authors:Naama Pearl, Stefano Esposito, Haofei Xu, Amit Peleg, Patricia Gschossmann, Lorenzo Porzi, Peter Kontschieder, Gerard Pons-Moll, Andreas Geiger
Abstract:
3D Gaussian Splatting (3DGS) optimization is most commonly performed using standard optimizers (Adam, SGD). While stable across diverse scenes, standard optimizers are general‑purpose and not tailored to the structure of the problem. In particular, they produce independent parameter updates that do not capture the structural and spatial relationships within a scene, leading to inefficient optimization and slow convergence. Recent works introduced learned optimizers that predict correlated updates informed by inter‑parameter and inter‑Gaussian dependencies. However, these methods are trained for a fixed number of optimization iterations and rely on manually scheduled learning rates to avoid degradation. In this paper, we introduce a learned optimizer for 3DGS that avoids degradation over extended optimization horizons without auxiliary mechanisms. To enable this, we propose a meta‑learning scheme that extends the optimization horizon via a checkpoint buffer and an optimizer rollout strategy, combined with an architecture that encodes gradient scale information in its latent states. Results show improved early novel view synthesis quality while remaining stable over long horizons, with zero‑shot generalization to unseen reconstruction settings. To support our findings, we introduce the first unified framework for training and evaluating both learned and conventional optimizers across sparse and dense view settings. Code and models will be released publicly. Our project page is available at https://naamapearl.github.io/learn2splat .
Authors:Cheng Zhang, Yuer Liu, Zhiyu Zhou, Hongxia Xie, Wen-Huang Cheng
Abstract:
Multimodal large language models (MLLMs) can produce fluent artwork emotion explanations, but they often suffer from attribute flooding: they enumerate many visible formal attributes without identifying which cues actually support the affective judgment. We therefore formulate artwork emotion understanding as Attribute‑Grounded Selective Reasoning (AGSR), where predefined formal attributes serve as evidence units and only emotionally operative attributes should enter the final interpretation. To make this problem measurable, we extend EmoArt, originally introduced at ACM MM 2025 as a 132,664‑artwork resource with content, formal‑attribute, valence‑arousal, and emotion annotations, by adding a 1,400‑artwork human salience extension annotated by 15 art‑trained annotators. This extension provides instance‑level supervision for distinguishing attributes that are merely present from those that are emotionally salient. We further propose FAB‑G (Formal‑Attribute Bottleneck‑Guided reasoning), a supervised multi‑agent framework that first predicts attribute‑level salience and then constrains downstream emotional analysis to the retained cues. Experiments show that FAB‑G yields consistent gains in emotion, arousal, and valence prediction, achieves stronger agreement with human‑marked salient attributes under Dice and Tversky metrics, and produces substantially more compact final explanations than prompting‑based baselines. Cross‑dataset evaluation further suggests that attribute‑grounded salience selection transfers beyond the source distribution of EmoArt, while also revealing attribute‑specific boundary cases. The dataset and project page are available at https://zhiliangzhang.github.io/EmoArt‑130k/
Authors:Jan Miksa, Patryk Krukowski, Przemysław Spurek, Dawid Damian Rymarczyk, Marcin Sendera
Abstract:
Machine unlearning has reached a critical bottleneck. As traditional weight‑space interventions focus primarily on erasing targeted concepts, they often fail to prevent the unintended suppression of other significant representations. This leads to substantial collateral damage, with essential knowledge being forgotten, because these methods lack formal mathematical guarantees for the preservation of neutral concepts. To avoid degradation, they are frequently forced into conservative updates. We propose BARRIER (Bounded Activation Regions for Robust Information Erasure), a paradigm‑shifting framework that shifts the locus of intervention from static model weights to the dynamic geometry of hidden‑layer activations. Unlike existing methods, BARRIER employs Interval Arithmetic (IA) on SVD‑based projections of the activation space to encapsulate the specific target region within a bounding hypercube. By driving unlearning updates exclusively within this forget interval and mathematically bounding the model response on the complement, we ensure rigorous protection of the retain distribution. This geometric construction transforms the preservation of knowledge from an empirical heuristic into a formal optimization target with a probabilistic tail bound on functional drift. Crucially, this stability permits highly aggressive unlearning updates within the forget region. Empirical evaluations demonstrate that BARRIER matches state‑of‑the‑art trade‑offs across classifiers and diffusion models, maximizing targeted concept erasure while safeguarding the integrity of all other representations. Our code is available at https://github.com/OneAndZero24/BARRIER.
Authors:Huanyang Tong, Kai Liu, Fangjun Kuang, Huiling Chen
Abstract:
Biomedical Vision‑‑Language Models (VLMs) have shown remarkable promise in few‑shot medical diagnosis but face a critical bottleneck: fragility to prompt variations.Existing adaptation frameworks typically optimize visual and textual prompts as independent streams, relying on ideal ``Golden Prompts''. In clinical reality, where descriptions are often noisy and heterogeneous, this modality isolation leads to unstable cross‑modal alignment.
To address this, we propose BiomedAP, a vision‑informed dual‑anchor framework with gated cross‑modal fusion.BiomedAP enforces synergistic alignment through two mechanisms: (1) Gated Cross‑Modal Fusion, which enables layer‑wise interaction between modalities, acting as a dynamic noise regulator to suppress irrelevant textual cues; and (2) a Dual‑Anchor Constraint that regularizes learnable prompts toward stable semantic centroids derived from both expert templates (High Anchors) and few‑shot visual prototypes (Low Anchors).
Extensive experiments across 11 benchmarks demonstrate that BiomedAP consistently surpasses baselines, achieving competitive few‑shot accuracy and markedly enhanced robustness under prompt perturbations.
Our code is available at: https://github.com/tongdiedie/BiomedAP.
Keywords: Vision‑Language Models; Prompt Learning; Parameter‑Efficient Fine‑Tuning; Few‑shot Learning
Authors:Oswin Gosal, Edwin Arkel Rios, Augusto Christian Surya, Fernando Mikael, Bo-Cheng Lai, Min-Chun Hu
Abstract:
Fine‑grained image recognition classifies subcategories such as bird species or car models. While state‑of‑the‑art (SOTA) models are accurate, they are often too resource‑intensive for deployment on constrained devices. Knowledge distillation addresses this by transferring knowledge from a large teacher model to a smaller student model. A key challenge is selecting the right teacher, as it heavily impacts student performance. This paper introduces a teacher selection metric, Ratio 1‑2, based on teacher prediction ratios. Extensive analysis of over one thousand experiments across 3 students, 8 teachers, and 8 datasets under 4 training strategies demonstrates that our metric improves teacher selection by 18% over previous methods, enabling small student models to achieve up to 17% accuracy gains. Experiment codebase is available at: \hrefhttps://github.com/arkel23/FGIR‑KD‑Teacherhttps://github.com/arkel23/FGIR‑KD‑Teacher.
Authors:Qingji Dong, Hang Dong, Mingqin Chen, Rui Zhang, Yitong Wang
Abstract:
Large‑scale pre‑trained diffusion models have been extensively adopted for real‑world image Super‑Resolution because of their powerful generative priors through textual guidance. However, when super‑resolving high‑resolution images with patch‑wise inference strategy, most existing diffusion‑based SR methods tend to suffer from over‑generation, due to the misalignment between the global prompt from LR image and the incomplete semantic information of local patches during each inference step. On the other hand, most existing methods also failed to generate detailed texture in local patches due to the overemphasis on global generation capabilities in network designs and training strategies. To address this issue, we present DreamSR, a novel SR model that suppresses local over‑generation and improves fine‑detail synthesis, thereby achieving visually faithful results with ultra‑high‑quality details. Specifically, we propose a dual‑branch MM‑ControlNet, where the ControlNet generates local textual feature with patch‑level prompts while the pre‑trained DiT provides global textual feature with global prompts, thereby mitigating over‑generation and ensuring semantic consistency across patches. We also design a comprehensive training strategy with stage‑specific data processing pipelines and a Receptive‑Field Enhancement strategy, enhancing the model's capability to capture patch information and effectively restore local textures. Extensive experiments demonstrate that DreamSR outperforms state‑of‑the‑art methods, providing high‑quality SR results. Code and model are available at https://github.com/jerrydong0219/DreamSR.
Authors:Nisha Huang, Yizhou Lin, Jie Guo, Xiu Li, Tong-Yee Lee, Zitong Yu
Abstract:
Recently, diffusion‑based material transfer methods rely on image fine‑tuning or complex architectures with auxiliary networks but face challenges such as text dependency, additional computational costs, and feature misalignment. To address these limitations, we propose DealMaTe, using \underlinedepth, norm\underlineal, and \underlinelighting images for \underlinematerial \underlinetransf\underlineer. DealMaTe is a simplified diffusion framework that eliminates text guidance and reference networks. We design a lightweight 3D information injection method, Multi‑Dim 3D Shader LoRA, which, without modifying the base model weights, enables compatible control conditions and achieves harmonious and stable results. Additionally, we optimize the attention mechanism with Shader Causal Mutual Attention and key‑value (KV) caching to reduce inference latency caused by multiple conditions, improve computational efficiency, and achieve high‑quality material transfer results with low architectural complexity. Extensive experiments covering a wide variety of objects and lighting conditions consistently demonstrate that DealMaTe achieves remarkable high‑fidelity material transfer under arbitrary input materials. The code is available at https://github.com/haha‑lisa/DealMaTe.
Authors:Yuchun Wang, Xiaosong Li, Gefei Liang, Yang Liu
Abstract:
Multimodal 3D MRI brain tumor segmentation is a pivotal step in radiotherapy target delineation, surgical planning and post‑treatment assessment. Existing methods often assume artifact‑free MRI images. However, inevitable patient motion during scanning introduces artifacts and blur that degrade boundary and texture features, leading to poor segmentation performance. To bridge this gap, we introduce Degradation‑Aware Blur‑Segmentation Net (DABSeg), a synchronous deblurring 3D multimodal MRI segmentation network that unifies blur removal and accurate segmentation. Specifically, we propose a feature‑domain motion‑deblurring stem to compensate for blur and rebalance intensity. Concurrently, the backbone network embeds a blur‑aware cross‑modal cross‑attention module and multi‑scale residual aggregation to yield effective modality complementarity. Notably, we optimize a joint loss that combines weighted Dice with a clear‑reference reconstruction term, where imbalanced weights are applied to small targets to boost learning intensity and predictive stability for small lesions and border regions. Systematic comparisons and ablation experiments on the BraTS2020 dataset under both clear and degenerative conditions consistently demonstrate that DABSeg surpasses state‑of‑the‑art methods in tumor Dice score and boundary precision. These results validate the effectiveness of degenerative‑aware cross‑task collaborative learning in improving the robustness and clinical utility of multi‑modal 3D brain tumor segmentation under realistic degenerative conditions. The source code is available at https://github.com/YuchunWang24/DABSeg_ICPR
Authors:Yan Luo, Ahmadou Aidara, Jingyi Lu, Jeremy Moebel, Kai Han, Mengyu Wang
Abstract:
Classifier‑free guidance (CFG) is the primary control over how strongly text semantics move a flow‑based sampler, yet standard practice holds its scale fixed across the entire ODE trajectory. This is a fundamental mismatch: early steps are noise‑dominated and carry weak semantic signal, while late steps commit image structure and demand stronger directional commitment; more critically, the value of any guidance strength depends on whether the guided velocity is consistent with the model's current dynamics or working against them. We propose Velocity‑Adaptive Guidance Scale (VAGS), a training‑free replacement that multiplies the nominal scale by a bounded factor combining a temporal signal‑level term with the cosine similarity between task‑relevant velocity fields. For inversion‑free editing, VAGS measures the alignment between source‑ and target‑guided velocities, so edit strength at each step reflects local compatibility between preservation and transformation. For generation, VAGS‑Gen uses the alignment between unconditional and conditional velocities as the analogous signal. Neither variant requires fine‑tuning, auxiliary networks, or extra forward passes, and fixed CFG is recovered as a special case. On PIE‑Bench and DIV2K for editing, and COCO17, CUB‑200, and Flickr30K for generation, VAGS consistently improves structural fidelity and generation quality over fixed CFG and recent training‑free guidance variants. The code is publicly available at https://github.com/Harvard‑AI‑and‑Robotics‑Lab/Velocity_Adaptive_Guidance_Scale.
Authors:Xin Zou, Ruimeng Liu, Chang Tang, Zhenglai Li, Xinwang Liu, Kunlun He, Wanqing Li
Abstract:
Multi‑View Clustering (MVC) has gained significant attention for its ability to leverage complementary information across diverse views. However, existing deep MVC methods often struggle with view‑distribution entanglement during cross‑view fusion, which hampers the quality of the shared latent space and leads to suboptimal Figures. To address this issue, we propose the Generalized Multi‑view Auto‑Encoder (GMAE), a framework designed to preserve cross‑view complementarity through disentangled representation learning. Specifically, GMAE employs dual‑path autoencoders to decouple source features into view‑specific and view‑common embeddings, facilitating the discovery of clearer clustering structures. We further construct cross‑view adversarial discriminators to guide view‑specific encoders in capturing more discriminative features. By strategically modulating mutual information, GMAE effectively aligns distributions and prevents representation collapse, ensuring the generation of robust, non‑trivial embeddings. Comprehensive experiments on 13 benchmark datasets demonstrate that GMAE consistently outperforms state‑of‑the‑art methods in both complete and incomplete MVC tasks. Our code implementation is available at the repository: https://github.com/obananas/GMAE.
Authors:Jiale Liu, Jungang Li, Jieming Yu, Xinglin Yu, Zihao Dongfang, Zongjian Ding, Kaifeng Ding, Yi Yang, Lidong Chen, Yang Zou, Shunwen Bai, Jiahuan Zhang, Haoran Huang, Shan Huang, Yudong Gao, Mingjun Cheng
Abstract:
Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructions do not directly provide a sparse, comparable, and geometry‑consistent panoramic training interface. Dense trajectories duplicate nearby views, source‑specific rendering policies yield heterogeneous annotations, and sparse heuristics may miss important regions or introduce depth‑inconsistent observations. We study how to convert 3D assets into sparse panoramic RGB‑D‑pose data that preserves complete scene coverage with low redundancy and auditable provenance. We propose COVER (Coverage‑Oriented Viewpoint curation with ERP Range‑depth warping), a training‑free ERP viewpoint curator that projects geometry observed from selected views into candidate ERP probes, scores incremental coverage, and penalizes depth conflicts. Under bounded proxy error, its greedy coverage proxy preserves the standard coverage‑style approximation behavior up to an additive error term. Using COVER, we build CM‑EVS (Coverage‑curated Metric ERP View Set), a panoramic RGB‑D‑pose dataset with 36,373 curated ERP frames from 1,275 indoor scenes across Blender indoor, HM3D, and ScanNet++, complemented by outdoor panoramas from TartanGround and OB3D re‑encoded into the same schema. Each frame provides full‑sphere RGB, metric range depth, calibrated pose; COVER‑produced indoor frames include per‑step provenance logs. With a median of only 25 frames per indoor scene, CM‑EVS covers all 13 unified room types while maintaining compact scene‑level coverage. Experiments show that COVER improves the coverage‑conflict trade‑off, making CM‑EVS a sparse, compact, and auditable RGB‑D‑pose resource for geometry‑consistent panoramic 3D learning.
Authors:Ryohei Goto, Takuya Fujihashi, Shunsuke Saruwatari, Fumio Okura
Abstract:
We propose a method of estimating a 3D human pose from a single view without 3D supervision. The key to our method is to leverage the 2D diffusion priors of motion diffusion models (MDMs) pre‑trained on large 2D human pose datasets. Specifically, we extend multi‑view ancestral sampling of diffusion models to the task of 2D‑3D lifting of human pose. To this end, we newly propose a conditional multi‑view ancestral sampling (cMAS) that optimizes the 3D pose such that its multi‑view projections follow the manifold in 2D MDM noise space, while conditioning the 3D pose to match the given 2D poses and anatomical constraints of humans. Experiments on the Yoga dataset demonstrate that our method achieves better cross‑domain performance compared to state‑of‑the‑art supervised and unsupervised 3D pose estimation methods, including extreme human poses where 3D supervision is unavailable. Code is available at: https://github.com/asaa0001/c‑MAS.
Authors:Jiaxuan Zhao, Ali Bereyhi
Abstract:
Modern deep learning models for change detection (CD) often struggle to explicitly represent task‑relevant semantic differences. This paper proposes the Latent Difference Guidance (LDGuid) framework that explicitly learns and injects semantic differences into CD models. LDGuid deploys adversarial autoencoding to implement a difference embedding (DE) module. The DE module is pretrained via the information bottleneck method, restricting it to learn only task‑relevant differences between pre‑ and post‑event samples. The learned latent difference is then used as an explicit guidance signal in the CD model. We validate LDGuid by integrating it into U‑Net, BIT, and AERNet baselines for CD and evaluating it on LEVIR‑CD, WHU‑CD, SVCD, and CaBuAr datasets. Experimental results show that LDGuid enhances segmentation performance across all benchmarks, with particularly remarkable gains in challenging settings affected by spectral noise. The results further highlight the ability of LDGuid in incorporating domain knowledge, such as task‑specific spectral indices. Our findings suggest that semantic difference learning can drastically enhance the robustness of CD in remote sensing.
Authors:Xinmin Feng, Li Li, Dong Liu, Feng Wu
Abstract:
To fit diverse display and bandwidth constraints, high‑frame‑rate videos are temporally downscaled to low‑frame‑rate (LFR) and later upscaled, requiring joint optimization for effective frame‑rate rescaling. However, existing methods typically link the two operations via training objectives, without fully exploiting their reciprocal nature, which may cause high‑frequency information loss. Moreover, they overlook the impact of lossy codecs on LFR videos, limiting real‑world applicability. In this work, we propose an end‑to‑end framework for compression‑aware frame‑rate rescaling, named TVRN. To regularize high‑frequency information lost during frame‑rate downscaling, TVRN adopts an invertible architecture that combines a Multi‑Input Multi‑Output Temporal Wavelet Transform with a high‑frequency reconstruction module. To enable end‑to‑end training through non‑differentiable lossy codecs, we design a surrogate network that approximates their gradients. Finally, to improve robustness under various compression levels, we extend TVRN to an asymmetric architecture by incorporating compression‑aware features learned via a learning‑to‑rank strategy. Extensive experiments show that TVRN outperforms existing methods in reconstruction quality under industrial video compression settings. Source code is publicly available at https://github.com/fengxinmin/TVRN_public.
Authors:Sunghwan Steve Cho, Yunseok Han, Jaeyoung Do
Abstract:
Longitudinal chest X‑ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA benchmarks focus on single images or short‑horizon image pairs. We introduce MI‑CXR, a benchmark for standardized evaluation of Multi‑Interval longitudinal reasoning over multi‑visit CXR sequences, without requiring free‑form report generation or additional clinical context. MI‑CXR comprises five‑way multiple‑choice questions over five‑visit patient timelines and instantiates three complementary task families: Temporal Event Localization, Interval‑wise Change Reasoning, and Global Trajectory Summarization, which assess clinically grounded visual reasoning over time. Evaluating 14 state‑of‑the‑art vision‑language models (VLMs) shows low overall performance, with an average accuracy of 29.3%, only modestly above random guessing. Using stage‑wise diagnostic probing, we find that models often produce locally plausible interval descriptions but fail to enforce temporal constraints or compose evidence into globally consistent decisions over the full timeline. These findings reveal key limitations of current VLMs and establish MI‑CXR as a principled benchmark for longitudinal medical reasoning. The benchmark is available at https://github.com/AIDASLab/MI‑CXR
Authors:Hao Yang, Xianping Ma, Peifeng Ma, Man-On Pun
Abstract:
High‑resolution remote sensing imagery is critical for environmental monitoring, urban mapping, and land cover analysis, but its transmission is often hindered by limited bandwidth and high communication costs. Conventional pipelines transmit full‑resolution pixel data, resulting in redundant and inefficient delivery. This paper proposes a text‑guided remote sensing image transmission system that replaces complete high‑resolution data with low‑resolution images accompanied by compact textual descriptions. An onboard text generator produces spatial and semantic summaries, reducing the transmitted data volume to approximately 2% of the original size. For ground‑based reconstruction, a text‑conditioned image restoration model is introduced, which leverages cross‑modal learning to recover fine spatial details and maintain semantic coherence. Experimental results on the Alsat‑2B, UC Merced Land Use, and Aerial Image datasets demonstrate that the proposed framework achieves reconstruction PSNRs of 16.36 dB, 26.87 dB, and 27.41 dB, respectively, enabling efficient and information‑preserving image transfer for remote sensing applications. The implementation will be made publicly available at \hrefhttps://github.com/haoyangofficial/textrssrGitHub.
Authors:Bingwen Qiu, Yuan Liu, Junqi Bai, Tong Jiang, Ben Liang, Fangzhou Chen, Xiubao Sui, Qian Chen
Abstract:
A fundamental challenge in point cloud object detection lies in the conflict between the extreme sparsity of distant points and the need for remote context understanding. The existing methods typically use 1D serialization to expand the receptive field, which inevitably discards already scarce local geometric details and reduces detection of distant and small objects. To address this issue, we propose 3DTMDet, a novel detection network that synergistically combines state space models (Mamba) with Transformers. The core idea is to utilize SSM's linear complexity and advantages in long sequence modeling to effectively capture global interactions between sparse and distant points, while using Transformer modules with local attention to encode fine‑grained geometric structures in local point sets, preserving accurate shape information. We propose the 3D Hybrid Mamba Transformer (3DHMT) block, which uses an SSM‑Attention‑SSM pipeline to balance global context understanding and local detail preservation, effectively alleviating the tension between receptive field enlargement and geometric preservation in remote detection. In addition, we introduced a voxel generation block inspired by LiDAR physics, which diffuses features along the sensor observation direction to reconstruct the complete object structure of occlusion and distant areas. Extensive experiments conducted on the KITTI and ONCE datasets have shown that 3DTMDet outperforms state‑of‑the‑art detectors. The code is available at https://github.com/QiuBingwen/3DTMDet.
Authors:Mingtong Dai, Guanqi Peng, Yongjie Bai, Feng Yan, Chunjie Chen, Lingbo Liu, Liang Lin, Xinyu Wu
Abstract:
Previous imitation learning policies predict future actions at every control step, whether in smooth motion phases or precise, contact‑rich operation phases. This uniform treatment is wasteful: most steps in a manipulation trajectory traverse free space and carry little task‑relevant information, while a small fraction of \emphkey steps around contacts, grasps, and alignment demand dense, high‑resolution prediction. We propose a novel \emphaction relabeling mechanism: at each timestep in a skip segment, we replace the behavior cloning target with the action at the entrance of the next key segment, enabling the policy to leap over redundant steps in a single decision. The resulting Skip Policy (SkiP) dynamically leaps over skip segments and intensively refines actions in key segments, within a single unified network requiring no learned skip planner or hierarchical structure. To automatically partition demonstrations into key and skip segments without manual annotation, we introduce \emphMotion Spectrum Keying (MSK), a fast, task‑agnostic procedure that detects local motion complexity from action signals. Extensive experiments across 72 simulated manipulation tasks and three real‑robot tasks show that SkiP reduces executed steps by 15‑‑40% while matching or improving success rates across various policy backbones. Project page: \texttthttps://pgq18.github.io/SkiP‑page/.
Authors:Dongjae Lee, Wooseong Yang, Yifu Tao, Maurice Fallon, Ayoung Kim
Abstract:
Neural distance fields offer a compact and continuous representation of 3D geometry, making them attractive for incremental LiDAR mapping. However, their online optimization is vulnerable to catastrophic forgetting, where new observations can degrade previously reconstructed geometry. Replay‑based training is commonly used to address this issue, but existing methods typically rely on passive replay buffers and uniform sampling, which can waste memory on redundant observations and under‑train poorly constrained regions. We propose LAPS, a replay management framework for incremental neural mapping that improves both replay retention and replay allocation during online updates. LAPS combines reliability‑based active pooling to retain reliable historical samples under limited memory with uncertainty‑guided active sampling to focus optimization on under‑constrained regions. Experiments on synthetic and real‑world benchmarks show that LAPS consistently improves reconstruction completeness while maintaining competitive geometric accuracy. On Oxford Spires, it improves recall by 4.66 pp and F1‑score by 3.79 pp over PIN‑SLAM on the Blenheim Palace 05 sequence. We release our open source implementation at: https://github.com/dongjae0107/LAPS.
Authors:Libo Sun, Po-wei Harn, Peixiong He, Xiao Qin
Abstract:
Mixture‑of‑Experts (MoE) networks promise favorable accuracy‑compute trade‑offs, yet practical vision deployments are hindered by expert collapse and limited end‑to‑end efficiency gains. We study when sparse top‑k routing with hard capacity constraints helps in vision classification, evaluated under multi‑seed protocols on four benchmarks (CIFAR‑10/100, Tiny‑ImageNet, ImageNet‑1K). We observe a \emphcompute‑leverage pattern: positive accuracy gaps require a substantial fraction ρ of total FLOPs to be routed; at ImageNet scale this is necessary but not sufficient, as multi‑expert routing (k \geq 2) is additionally required. Two controlled experiments isolate these factors. A hidden‑size sweep on CIFAR‑10 yields both predicted sign reversals across standard and depthwise backbones, ruling out backbone family as the active variable. An ImageNet‑1K ablation that varies only top‑k ‑‑ holding architecture, initialization, and ρ fixed ‑‑ reverses the gap from positive to negative across all five seeds. A per‑sample variant of Soft MoE that softmaxes over experts rather than the batch rescues CIFAR‑100 above the dense baseline, identifying batch‑axis dispatch as the dominant failure mode in per‑sample CNN settings. Code and aggregate results: https://github.com/libophd/sparse‑moe‑vision‑rho.
Authors:Tinghui Zhu, Sheng Zhang, James Y. Huang, Selena Song, Xiaofei Wen, Yuankai Li, Hoifung Poon, Muhao Chen
Abstract:
Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially pronounced in tasks where generated videos must satisfy explicit spatial, temporal, or logical constraints. Inspired by the role of reinforcement learning with verifiable rewards (RLVR) in reasoning‑oriented language models, we introduce VideoRLVR, a practical recipe for optimizing video diffusion models with rule‑based feedback. VideoRLVR formulates video reasoning as the generation of verifiable visual trajectories and consists of an SDE‑GRPO optimization backbone, dense decomposed rewards, and an Early‑Step Focus strategy for efficient training. The Early‑Step Focus strategy restricts policy optimization to the early denoising phase, reducing training latency by about 40% while preserving performance. We evaluate VideoRLVR on Maze, FlowFree, and Sokoban, three procedurally generated domains with objective success criteria. Across these tasks, VideoRLVR consistently improves over supervised fine‑tuning baselines, with dense decomposed rewards proving especially important in low‑success‑rate settings. Our RL‑optimized model also outperforms the evaluated proprietary and open‑source video generation models on these verifiable reasoning benchmarks and out‑of‑domain benchmarks. These results suggest that verifiable RL can move video models beyond perceptual imitation toward more reliable rule‑consistent visual reasoning.
Authors:Po-Chien Luan, Wuyang Li, Yang Gao, Alexandre Alahi
Abstract:
Human trajectory forecasting is crucial for safe navigation in crowded environments, requiring models that balance accuracy with computational efficiency. Efficiently modeling social interactions is key to performance in dense crowds. Yet, most recent methods rely on attention mechanisms, which are effective at capturing complex dependencies, but incur quadratic computational costs that scale poorly with the growing number of neighbors. Recently, Selective State‑Space Models have provided a linear‑time alternative; however, their inherently sequential design is misaligned with the unstructured and dynamic nature of social interactions. To address this challenge, we propose Social‑Mamba, a forecasting architecture that reformulates social interactions as structured sequential processes. At its core is the Cycle Mamba block, a novel module that enables continuous bidirectional information flow. Social‑Mamba organizes agents on an egocentric grid and introduces social triplet factorization, which decomposes interactions into temporal, egocentric, and goal‑centric scans. These are dynamically integrated through a learnable social gate and global scan to generate accurate and efficient trajectory predictions. Extensive experiments on five trajectory forecasting benchmarks show that Social‑Mamba achieves state‑of‑the‑art accuracy while offering superior parameter efficiency and computational scalability. Furthermore, embedding Social‑Mamba into a flow‑matching framework further enhances both accuracy and efficiency, establishing it as a flexible and robust foundation for future trajectory forecasting research. The code is publicly available: https://github.com/vita‑epfl/Social‑Mamba
Authors:Luca Bompani, Manuele Rusci, Luca Benini, Daniele Palossi, Francesco Conti
Abstract:
Modern smart vision sensors need on‑device intelligence to process video streams, as cloud computing is often impractical due to bandwidth, latency, and privacy constraints. However, these sensory systems typically rely on ultra‑low‑power microcontrollers (MCUs) with limited memory and compute, making conventional video object detection methods, which require feature storage or multi‑frame buffering, unfeasible. To address this challenge, we introduce Multi‑Resolution Rescored ByteTrack (MR2‑ByteTrack), a Video Object Detection (VOD) method tailored for MCU‑based embedded vision nodes. MR2‑ByteTrack reduces computational cost by alternating between full‑ and low‑resolution inference, while linking detections across frames via ByteTrack and correcting misclassifications through the Rescore algorithm, which applies probability union rules to aggregate detection confidence scores across frames. We apply our approach to both a CNN‑based detector and a Transformer‑based model, demonstrating its generality across architectures with fundamentally different spatial processing. Experiments on ImageNetVID demonstrate that MR2‑ByteTrack maintains accuracy, achieving mAP scores of up to 49.0 for the CNN‑based models and 48.7 for the Transformer, while reducing multiply‑accumulate operations by as much as 53% for the CNNs and 32% for the Transformer. When deployed on GAP9, an ultra‑low‑power RISC‑V multicore MCU, our method yields up to 55% energy savings compared to processing only full‑resolution images, enabling the first real‑time Transformer‑based VOD on an MCU‑class embedded vision node. Code available at https://github.com/Bomps4/Multi_Resolution_Rescored_ByteTrack/tree/IEEE_Access
Authors:Le Jiang, Xiangyu Bai, Bishoy Galoaa, Shayda Moezzi, Caleb James Lee, Tooba Imtiaz, Edmund Yeh, Jennifer Dy, Yanzhi Wang, Sarah Ostadabbas
Abstract:
We present PanoWorld, a panoramic video world model that generates geometry‑consistent 360\degree video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly constrain the underlying 3D scene state, producing outputs that appear plausible yet exhibit inconsistent depth, broken correspondences, and implausible motion across the spherical surface. We address this gap by framing panoramic video generation as a geometry‑ and dynamics‑consistent latent state modeling problem rather than pure visual synthesis. Building on a pre‑trained perspective video world model, we introduce two lightweight regularizers: a depth consistency loss against pseudo ground‑truth panoramic depth, and a trajectory consistency loss that supervises the 3D world‑frame positions of tracked points across time. We further apply spherical‑geometry‑aware adaptation to the conditioning and positional encoding. We additionally introduce PanoGeo, a unified geometry‑aware panoramic video dataset with consistent depth, trajectory, and prompt annotations across diverse real and synthetic sources, used for both training and stratified evaluation. Experiments show that PanoWorld improves geometric consistency over prior panoramic generation methods while maintaining competitive visual realism, establishing that panoramic video generation must be treated as a geometric modeling problem to support the holistic spatial understanding requirements of embodied AI applications. Code is available at https://github.com/ostadabbas/PanoWorld.
Authors:Emre Hayir, Lorin Crawford, Alex X. Lu
Abstract:
Microscopy images contain rich information about how cells respond to perturbations, making them essential to applications like drug screening. To quantify images, researchers often use representation extraction methods, and recent years have seen a proliferation of deep learning methods. While measuring the quality of these representations is essential, evaluation remains fragmented, with each proposed model evaluated on different tasks and datasets, using custom pipelines and metrics, making it difficult to fairly compare models. Here, we introduce MorphoHELM, a comprehensive open benchmark for evaluating feature extraction methods for Cell Painting, the most widely‑used morphological profiling assay. MorphoHELM consolidates evaluation standards in the field, extends and corrects them to be more robust, and evaluates on the widest range of methods to date. A defining feature of the benchmark is that each task is evaluated at different degrees of batch effects (or technical noise), directly quantifying how the ability of methods to detect biological signal degrades as noise increases. Together, these properties enable MorphoHELM to detect trade‑offs between methods, and we demonstrate that models that excel at certain kinds of biological signal are weaker at others. We show that no existing model outperforms classic computer vision analytic strategies across all settings, which remain the strongest general use‑case representations. All datasets, code, and evaluation tools are publicly available at https://github.com/microsoft/MorphoHELM.
Authors:Blaž Rolih, Matic Fučka, Filip Wolf, Luka Čehovin Zajc
Abstract:
Remote sensing change detection (RSCD) aims to localise changes between two images of the same geographic region. In practice, change masks often follow region‑level annotation conventions rather than purely local appearance differences, making them context‑dependent and occasionally ambiguous. Most state‑of‑the‑art methods utilise per‑pixel discriminative classification, which produces a single prediction per input and fails to explicitly model the changed region as a coherent whole. A natural alternative is generative formulation, which can model a distribution of plausible masks, enabling sampling to capture ambiguity and encourage global consistency. However, existing generative RSCD approaches typically lag behind strong discriminative baselines due to the high computational cost of pixel‑space generation and the complexity of their conditioning mechanisms. To address the limitations of prior discriminative and generative methods, we propose ChangeFlow, a generative framework that reformulates change detection as the synthesis of a change mask in latent space via rectified flow. ChangeFlow is guided by a structured yet lightweight conditioning signal, and its stochastic design naturally supports sampling‑based prediction ensembling. Namely, aggregating multiple predicted change masks improves robustness, while sample agreement provides a practical confidence estimation that highlights ambiguous regions. Across four benchmarks, ChangeFlow achieves an average F1 of 80.4%, improving by 1.3 points on average over the previous best method, while maintaining inference speed comparable to recent strong baselines. Project page: https://blaz‑r.github.io/changeflow_cd
Authors:Arsha Nagrani, Jasper Uijilings, Shyamal Buch, Tobias Weyand, Sudheendra Vijayanarasimhan, Bo Hu, Ramin Mehran, David A Ross, Cordelia Schmid
Abstract:
Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate reasoning steps, and most provide answers only in the text domain. We introduce Minerva‑Ego, a benchmark for evaluating complex egocentric visual reasoning. We extend recent high‑quality video data sources recorded from egocentric / embodied settings with a set of challenging, multi‑step multimodal questions and spatiotemporally‑dense human‑annotated reasoning traces. Benchmarking experiments show that state‑of‑the‑art models still have a large gap to human performance. To investigate this gap in detail, we annotate each reasoning trace in the dataset with the objects of interest required to solve the question, as spatiotemporal mask annotations. Through extensive evaluations, we identify that prompting frontier models with hints of 'where' and 'when' to look yields substantial improvements in performance. Minerva‑Ego can be downloaded at https://github.com/google‑deepmind/neptune.
Authors:Darryl Cherian Jacob, Xinyu Liu, Kai Wang, Pan He
Abstract:
Vision‑language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existing VLM‑based VAD methods suffer from a fundamental mismatch between training and inference in both data distribution and model configuration. First, most approaches rely on static post‑training adaptation, limiting generalization under distribution shifts such as unseen environments or anomaly types. Second, they train VLMs on sparse frames from long videos, but perform inference on densely sampled short segments, creating inconsistencies between training and testing. To address these limitations, we propose COPRA, a conditional parameter adaptation framework for VLM‑based VAD. Instead of fixed prompts or shared parameter updates, COPRA generates input‑specific parameter updates to dynamically adapt a frozen VLM for each video segment during both training and inference. Experiments show strong performance on standard VAD benchmarks, consistently outperforming static baselines in both in‑domain and cross‑domain settings. Moreover, COPRA generalizes beyond VAD to unseen tasks such as multiple‑choice Video Question Answering and Dense Captioning. These results highlight COPRA as an effective weight‑space generation framework for scalable, adaptive, and context‑aware video understanding. The code will be released at https://github.com/THE‑MALT‑LAB/COPRA
Authors:AmirHossein Naghi Razlighi, Aryan Mikaeili, Ali Mahdavi-Amiri, Daniel Cohen-Or, Yiorgos Chrysanthou
Abstract:
Motion‑centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produce specific, localized actions or state transitions in an existing clip. We introduce Sound Sparks Motion, a training‑free framework that enables motion editing in an audio‑visual video generation model by tuning its internal multimodal conditioning signals at test time. Rather than modifying model weights, our method tunes only two lightweight variables: an audio latent derived from the source video and a residual perturbation in the text‑conditioning. We find that this combination can encourage motion edits that the underlying model often struggles to realize under prompt‑only control. Since there is no direct way to evaluate temporal alignment between text and motion, we guide the tuning process using a vision‑language model that provides feedback indicating whether the intended motion appears in the generated video. This simple supervision yields an effective semantic objective for motion editing, while regularization and perceptual‑temporal constraints help preserve content and visual quality. Beyond per‑video tuning, we show that the learned latent controls are transferable across videos, suggesting that they capture reusable motion‑edit directions rather than overfitting to a single example.
Our results highlight multimodal conditioning tuning, particularly through the audio pathway, as a promising direction for motion‑aware video editing, and suggest that test‑time tuning can serve as a lightweight probing mechanism that helps reveal latent motion controls embedded in the model's multimodal conditioning. Code and data are available via our project page: https://amirhossein‑razlighi.github.io/Sound_Sparks_Motion/
Authors:Tianyu Yu, Kechen Fang, Zihao Wan, Kaidong Zhang, Yicheng Zhang, Jun Song, Bo Zheng, Yuan Yao
Abstract:
Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis suggests this architecture suffers from an alignment challenge: visual features remain distant from the text space in the initial layers of the LLM, forcing the model to waste critical depth~\citezhang‑etal‑2024‑investigating,artzy‑schwartz‑2024‑attend on superficial modality alignment rather than deep understanding and complex reasoning. In this work, we propose Deep Pre‑Alignment (DPA), a novel architecture that replaces the standard ViT encoder with a small VLM as perceiver, ensuring visual features are deeply aligned with the text space of the target large language model. Comprehensive experiments demonstrate the effectiveness of DPA. On the 4B parameter scale, DPA outperforms baselines by 1.9 points across 8 multimodal benchmarks, with gains widening to 3.0 points at the 32B scale. Moreover, by offloading alignment to the perceiver, DPA achieves a 32.9% reduction in language capability forgetting over 3 text benchmarks. We further demonstrate that these gains are consistent across different LLM families including Qwen3 and LLaMA 3.2, highlighting the generality of our approach. Beyond performance, DPA also offers a seamless upgrade path for current VLM development, requiring only a modular replacement for the visual encoder with marginal computation overhead.
Authors:Zeqing Wang, Danze Chen, Zhaohu Xing, Zizhao Tong, Yinhan Zhang, Xingyi Yang, Yeying Jin
Abstract:
Current game world models simulate environments from a subjective, player‑centric perspective. However, by treating the Non‑Player Character (NPC) merely as background pixels, these models cannot capture interactions between the player and NPC. In that sense, they act as passive video renderers rather than real simulation engines, lacking the physical understanding needed to model action‑induced NPC reactivities. We introduce ReactiveGWM, a reactive game world model that synthesizes dynamic interactions between the player and NPC. Instead of entangling all interaction dynamics, ReactiveGWM explicitly decouples player controls from NPC behaviors. Player actions are injected into the diffusion backbone via a lightweight additive bias, while high‑level NPC responses (e.g., Offense, Control, Defense) are grounded through cross‑attention modules. Crucially, these modules learn a game‑agnostic representation of interactive logic. This enables zero‑shot strategy transfer: our learned modules can be plugged directly into off‑the‑shelf, unannotated world models of different games. This instantly unlocks steerable NPC interactions without any domain‑specific retraining. Evaluated on two Street Fighter games, ReactiveGWM maintains fine‑grain player controllability while achieving robust, prompt‑aligned NPC strategy adherence, paving the way for scalable, strategy‑rich interaction with the NPC.
Authors:Ruozhen He, Meng Wei, Ziyan Yang, Vicente Ordonez
Abstract:
Multi‑shot video generation extends single‑shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a challenge over long sequences. Existing evaluations typically use independently generated prompt sets with limited entity coverage and simple consistency metrics, making standardized comparison difficult. We introduce EntityBench, a benchmark of 140 episodes (2,491 shots) derived from real narrative media, with explicit per‑shot entity schedules tracking characters, objects, and locations simultaneously across easy / medium / hard tiers of up to 50 shots, 13 cross‑shot characters, 8 cross‑shot locations, 22 cross‑shot objects, and recurrence gaps spanning up to 48 shots. It is paired with a three‑pillar evaluation suite that disentangles intra‑shot quality, prompt‑following alignment, and cross‑shot consistency, with a fidelity gate that admits only accurate entity appearances into cross‑shot scoring. As a baseline, we propose EntityMem, a memory‑augmented generation system that stores verified per‑entity visual references in a persistent memory bank before generation begins. Experiments show that cross‑shot entity consistency degrades sharply with recurrence distance in existing methods, and that explicit per‑entity memory yields the highest character fidelity (Cohen's d = +2.33) and presence among methods evaluated. Code and data are available at https://github.com/Catherine‑R‑He/EntityBench/.
Authors:Ziyu Guo, Rain Liu, Xinyan Chen, Pheng-Ann Heng
Abstract:
Visual reasoning, often interleaved with intermediate visual states, has emerged as a promising direction in the field. A straightforward approach is to directly generate images via unified models during reasoning, but this is computationally expensive and architecturally non‑trivial. Recent alternatives include agentic reasoning through code or tool calls, and latent reasoning with learnable hidden embeddings. However, agentic methods incur context‑switching latency from external execution, while latent methods lack task generalization and are difficult to train with autoregressive parallelization. To combine their strengths while mitigating their limitations, we propose ATLAS, a framework in which a single discrete 'word', termed as a functional token, serves both as an agentic operation and a latent visual reasoning unit. Each functional token is associated with an internalized visual operation, yet requires no visual supervision and remains a standard token in the tokenizer vocabulary, which can be generated via next‑token prediction. This design avoids verbose intermediate visual content generation, while preserving compatibility with the vanilla scalable SFT and RL training, without architectural or methodological modifications. To further address the sparsity of functional tokens during RL, we introduce Latent‑Anchored GRPO (LA‑GRPO), which stabilizes the training by anchoring functional tokens with a statically weighted auxiliary objective, providing stronger gradient updates. Extensive experiments and analyses demonstrate that ATLAS achieves superior performance on challenging benchmarks while maintaining clear interpretability. We hope ATLAS offers a new paradigm inspiring future visual reasoning research.
Authors:Kaixin Zhu, Yiwen Tang, Yifan Yang, Renrui Zhang, Bohan Zeng, Ziyu Guo, Ruichuan An, Zhou Liu, Qizhi Chen, Delin Qu, Jaehong Yoon, Wentao Zhang
Abstract:
High‑quality 3D scene reconstruction has recently advanced toward generalizable feed‑forward architectures, enabling the generation of complex environments in a single forward pass. However, despite their strong performance in static scene perception, these models remain limited in responding to dynamic human instructions, which restricts their use in interactive applications. Existing editing methods typically rely on a 2D‑lifting strategy, where individual views are edited independently and then lifted back into 3D space. This indirect pipeline often leads to blurry textures and inconsistent geometry, as 2D editors lack the spatial awareness required to preserve structure across viewpoints. To address these limitations, we propose VGGT‑Edit, a feed‑forward framework for text‑conditioned native 3D scene editing. VGGT‑Edit introduces depth‑synchronized text injection to align semantic guidance with the backbone's spatial poses, ensuring stable instruction grounding. This semantic signal is then processed by a residual transformation head, which directly predicts 3D geometric displacements to deform the scene while preserving background stability. To ensure high‑fidelity results, we supervise the framework with a multi‑term objective function that enforces geometric accuracy and cross‑view consistency. We also construct the DeltaScene Dataset, a large‑scale dataset generated through an automated pipeline with 3D agreement filtering to ensure ground‑truth quality. Experiments show that VGGT‑Edit substantially outperforms 2D‑lifting baselines, producing sharper object details, stronger multi‑view consistency, and near‑instant inference speed. The project page is https://chriszkxxx.github.io/VGGT‑Edit/.
Authors:Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li, Xueyan Zou
Abstract:
Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human judgment or learned graders, which can be subjective and weakly diagnostic for geometric failures. We introduce PDI‑Bench (Perspective Distortion Index), a quantitative framework for auditing geometric coherence in generated videos. Given a generated clip, we obtain object‑centric observations via segmentation and point tracking (e.g., SAM 2, MegaSaM, and CoTracker3), lift them to 3D world‑space coordinates via monocular reconstruction, and compute a set of projective‑geometry residuals capturing three failure dimensions: scale‑depth alignment, 3D motion consistency, and 3D structural rigidity. To support systematic evaluation, we build PDI‑Dataset, covering diverse scenarios designed to stress these geometric constraints. Across state‑of‑the‑art video generators, PDI reveals consistent geometry‑specific failure modes that are not captured by common perceptual metrics, and provides a diagnostic signal for progress toward physically grounded video generation and physical world model. Our code and dataset can be found at https://pdi‑bench.github.io/.
Authors:Yifan Wang, Tong He
Abstract:
Camera‑controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existing methods usually learn camera‑specific conditioning through camera encoders, control branches, or attention and positional‑encoding modifications, which often require post‑training on large‑scale camera‑annotated videos. Training‑free alternatives avoid such post‑training, but often shift the cost to test‑time optimization or extra denoising‑time guidance. We propose Warp‑as‑History, a simple interface that turns camera‑induced warps into camera‑warped pseudo‑history with target‑frame positional alignment and visible‑token selection. Given a target camera trajectory, we construct camera‑warped pseudo‑history from past observations and feed it through the model's visual‑history pathway. Crucially, we align its positional encoding with the target frames being denoised and remove warped‑history tokens without valid source observations. Without any training, architectural modification, or test‑time optimization, this interface reveals a non‑trivial zero‑shot capability of a frozen video generation model to follow camera trajectories. Moreover, lightweight offline LoRA finetuning on only one camera‑annotated video further improves this capability and generalizes to unseen videos, improving camera adherence, visual quality, and motion dynamics without test‑time optimization or target‑video adaptation. Extensive experiments on diverse datasets confirm the effectiveness of our method.
Authors:Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, Enze Xie
Abstract:
We introduce SANA‑WM, an efficient 2.6B‑parameter open‑source world model natively trained for one‑minute generation, synthesizing high‑fidelity, 720p, minute‑scale videos with precise camera control. SANA‑WM achieves visual quality comparable to large‑scale industrial baselines such as LingBot‑World and HY‑WorldPlay, while significantly improving efficiency. Four core designs drive our architecture: (1) Hybrid Linear Attention combines frame‑wise Gated DeltaNet (GDN) with softmax attention for memory‑efficient long‑context modeling. (2) Dual‑Branch Camera Control ensures precise 6‑DoF trajectory adherence. (3) Two‑Stage Generation Pipeline applies a long‑video refiner to stage‑1 outputs, improving quality and consistency across sequences. (4) Robust Annotation Pipeline extracts accurate metric‑scale 6‑DoF camera poses from public videos to yield high‑quality, spatiotemporally consistent action labels. Driven by these designs, SANA‑WMdemonstrates remarkable efficiency across data, training compute, and inference hardware: it uses only ~213K public video clips with metric‑scale pose supervision, completes training in 15 days on 64 H100s, and generates each 60s clip on a single GPU; its distilled variant can be deployed on a single RTX 5090 with NVFP4 quantization to denoise a 60s 720p clip in 34s. On our one‑minute world‑model benchmark, SANA‑WM demonstrates stronger action‑following accuracy than prior open‑source baselines and achieves comparable visual quality at 36× higher throughput for scalable world modeling.
Authors:Chenyu Lian, Hong-Yu Zhou, Jing Qin
Abstract:
Disease screening is critical for early detection and timely intervention in clinical practice. However, most current screening models for medical images suffer from limited interpretability and suboptimal performance. They often lack effective mechanisms to reference historical cases or provide transparent reasoning pathways. To address these challenges, we introduce EviScreen, an evidential reasoning framework for disease screening that leverages region‑level evidence from historical cases. The proposed EviScreen offers retrospection interpretability through regional evidence retrieved from dual knowledge banks. Using this evidential mechanism, the subsequent evidence‑aware reasoning module makes predictions using both the current case and evidence from historical cases, thereby enhancing disease screening performance. Furthermore, rather than relying on post‑hoc saliency maps, EviScreen enhances localization interpretability by leveraging abnormality maps derived from contrastive retrieval. Our method achieves superior performance on our carefully established benchmarks for real‑world disease screening, yielding notably higher specificity at clinical‑level recall. Code is publicly available at https://github.com/DopamineLcy/EviScreen.
Authors:Kam Man Wu, Haolin Yang, Qingyu Chen, Yihu Tang, Jingye Chen, Qifeng Chen
Abstract:
Recent advances in image generation have made it easy to produce high‑quality images. However, these outputs are inherently flattened, entangling foreground elements, background, and text within a fixed canvas. As a result, flexible post‑generation editing remains challenging, revealing a clear last‑mile gap toward practical usability. Existing approaches either rely on scarce proprietary layered assets or construct partially synthetic data from limited structural priors. However, both strategies face fundamental challenges in scalability. In this work, we investigate whether pure synthetic layered data can improve graphic design decomposition. We make the assumption that, in graphic design, effective decomposition does not require modeling inter‑layer dependencies as precisely as in natural‑image composition, since design elements are often intentionally arranged as modular and semantically separable components. Concretely, we conduct a data‑centric study based on CLD baseline, which is a state‑of‑the‑art layer decomposition framework. Based on the baseline, we construct our own synthetic dataset, SynLayers, generate textual supervision using vision language models, and automate inference inputs with VLM‑predicted bounding boxes. Our study reveals three key findings: (1) even training with purely synthetic data can outperform non‑scalable alternatives such as the widely used PrismLayersPro dataset, demonstrating its viability as a scalable and effective substitute; (2) performance consistently improves with increased training data scale, while gains begin to saturate at around 50K samples; and (3) synthetic data enables balanced control over layer‑count distributions, avoiding the layer‑count imbalance commonly observed in real‑world datasets. We hope this data‑centric study encourages broader adoption of synthetic data as a practical foundation for layered design editing systems.
Authors:Sining Ang, Yuguang Yang, Canyu Chen, Yan Wang
Abstract:
End‑to‑end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule‑based planning metrics that measure safety, feasibility, progress, and comfort. This creates a training‑‑evaluation mismatch: trajectories close to the logged path may violate planning rules, while alternatives farther from the demonstration can remain valid and high‑scoring. The mismatch is especially limiting for proposal‑selection planners, whose performance depends on candidate‑set coverage and scorer ranking quality. We propose CLOVER, a Closed‑LOop Value Estimation and Ranking framework for end‑to‑end autonomous driving planning. CLOVER follows a lightweight generator‑‑scorer formulation: a generator produces diverse candidate trajectories, and a scorer predicts planning‑metric sub‑scores to rank them at inference time. To expand proposal support beyond single‑trajectory imitation, CLOVER constructs evaluator‑filtered pseudo‑expert trajectories and trains the generator with set‑level coverage supervision. It then performs conservative closed‑loop self‑distillation: the scorer is fitted to true evaluator sub‑scores on generated proposals, while the generator is refined toward teacher‑selected top‑k and vector‑Pareto targets with stability regularization. We analyze when an imperfect scorer can improve the generator, showing that scorer‑mediated refinement is reliable when scorer‑selected targets are enriched under the true evaluator and updates remain conservative. On NAVSIM, CLOVER achieves 94.5 PDMS and 90.4 EPDMS, establishing a new state of the art. On the more challenging NavHard split, it obtains 48.3 EPDMS, matching the strongest reported result. On supplementary nuScenes open‑loop evaluation, CLOVER achieves the lowest L2 error and collision rate among compared methods. Code data will be released at https://github.com/WilliamXuanYu/CLOVER.
Authors:Mukul Ranjan, Prince Jha, Khushboo Kumari, Zhiqiang Shen
Abstract:
Vision‑Language Models (VLMs) are increasingly applied to cultural heritage materials, from digital archives to educational platforms. This work identifies a fundamental issue in how these models interpret historical artifacts. We define this phenomenon as cultural anachronism, the tendency to misinterpret historical objects using temporally inappropriate concepts, materials, or cultural frameworks. To quantify this phenomenon, we introduce the Temporal Anachronism Benchmark for Vision‑Language Models (TAB‑VLM), a dataset of 600 questions across six categories, designed to evaluate temporal reasoning on 1,600 Indian cultural artifacts spanning prehistoric to modern periods. Systematic evaluations of ten state‑of‑the‑art models reveal significant deficiencies on our benchmark, and even the best model (GPT‑5.2) achieves only 58.7% overall accuracy. The performance gap persists across varying architectures and scales, suggesting that cultural anachronism represents a significant limitation in visual AI systems, regardless of model size. These findings highlight the disparity between current VLM capabilities and the requirements for accurately interpreting cultural heritage materials, particularly for non‑Western visual cultures underrepresented in training data. Our benchmark provides a foundation for enhancing temporal cognition in multimodal AI systems that interact with historical artifacts. The dataset and code are available in our project page.
Authors:Chengshuai Yang, Lei Xing, Gregory Entin, Roopa Vemulapalli, Lisa Casey, Raiyan Tripti Zaman
Abstract:
Background. RGB‑trained capsule‑endoscopy classifiers underperform on small‑vessel vascular findings by
conflating hemoglobin contrast with bile and illumination falloff. Thus, here we test whether a Monte
Carlo‑inspired analytic model can compute hemoglobin from RGB signal built upon extracted classifier.
Methods. On Kvasir‑Capsule (47,238 frames, video‑level 70/15/15 split, 11 evaluable classes) we evaluate two
software‑only configurations against RGB‑only EfficientNet‑B0 across 6 seeds: (i) a prior P_blood =
sigma(alpha (H_norm ‑ 0.5)) Phi(r) fused as 2 zero‑init auxiliary channels; (ii) a distillation head
training a 3‑channel RGB backbone to predict P_blood. Significance: paired DeLong, McNemar, bootstrap CIs
with Bonferroni correction.
Results. Across 6 seeds (n=6,423), the analytic prior provides a small but direction‑consistent macro‑AUC
improvement: RGB‑only 0.760 +/‑ 0.027, input‑fusion 0.783 +/‑ 0.024 (paired Delta = +0.023, sign‑positive on
5/6 seeds), distillation 0.773 +/‑ 0.028. The largest robust per‑class lift is on Lymphangiectasia, where AUC
rises from RGB 0.238 +/‑ 0.057 to input‑fusion 0.337 +/‑ 0.019, sign‑consistent across all 6 seeds. On rare
focal‑vascular classes (Angiectasia, Blood ‑ fresh) the prior's per‑seed effects are bimodal: seed=42 reaches
Angiectasia AUC 0.528 ‑> 0.916, but the cross‑seed mean is 0.646 ‑> 0.608 with sigma_PI = 0.23 ‑ reported as
a high‑variance per‑seed exemplar.
Conclusion. A Monte Carlo‑inspired analytic prior provides a small, direction‑consistent macro‑AUC
improvement on Kvasir‑Capsule across 6 seeds with the largest robust per‑class lift on Lymphangiectasia; the
distillation variant runs on plain 3‑channel RGB and yields a free interpretability heatmap.
Authors:Wuyang Li, Yang Gao, Mariam Hassan, Lan Feng, Wentao Pan, Po-Chien Luan, Alexandre Alahi
Abstract:
We propose EverAnimate, an efficient post‑training method for long‑horizon animated video generation that preserves visual quality and character identity. Long‑form animation remains challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk‑based generation prone to accumulated drift: (i) low‑level quality drift, such as progressive degradation of static backgrounds, and (ii) high‑level semantic drift, such as inconsistent character identity and view‑dependent attributes. To address this issue, EverAnimate restores drifted flow trajectories by anchoring generation to a persistent latent context memory, consisting of two complementary mechanisms. (i) Persistent Latent Propagation maintains a context memory across chunks to propagate identity and motion in latent space while mitigating temporal forgetting. (ii) Restorative Flow Matching introduces an implicit restoration objective during sampling through velocity adjustment, improving within‑chunk fidelity. With only lightweight LoRA tuning, EverAnimate outperforms state‑of‑the‑art long‑animation methods in both short‑ and long‑horizon settings: at 10 seconds, it improves PSNR/SSIM by 8%/7% and reduces LPIPS/FID by 22%/11%; at 90 seconds, the gains increase to 15%/15% and 32%/27%, respectively.
Authors:Man Wang, Chenyang Liu, Wenjun Li, Feng Ni, Bing Jia, Baoqi Huang, Riting Xia, Zhenwei Shi
Abstract:
Remote sensing image change captioning (RSICC) aims to achieve high‑level semantic understanding of genuine changes occurring between bi‑temporal images. Despite notable progress, existing methods are fundamentally limited by a shared modeling assumption: changed and unchanged image pairs, which have intrinsically different semantic granularities, are processed under a unified modeling strategy. This modeling inconsistency leads to semantic entanglement between coarse‑grained change existence judgment and fine‑grained semantic understanding.To address the above limitation, we propose a novel hierarchical semantic disentangling network (HiSem) that explicitly disentangles semantic representations of different granularities. Specifically, we first introduce the Bidirectional Differential Attention Modulation (BDAM) module that leverages discrepancy‑aware attention to enhance cross‑temporal interactions, thereby amplifying true change signals while suppressing irrelevant variations. Building upon this, we design a Hierarchical Adaptive Semantic Disentanglement (HASD) module that performs adaptive routing at two hierarchical levels: a coarse‑grained image‑level routing mechanism distinguishes changed and unchanged image pairs, while a fine‑grained token‑level Mixture‑of‑Experts (MoE) block models diverse and heterogeneous change semantics for changed samples. Extensive experiments on two benchmark datasets demonstrate that HiSem outperfoms previous methods, achieving a significant improvement of +7.52% BLEU‑4 on the WHU‑CDC dataset. More importantly, our approach provides a structured perspective for RSICC by explicitly aligning model design with the intrinsic semantic heterogeneity of bi‑temporal scenes. The code will be available at https://github.com/Man‑Wang‑star/HiSem
Authors:Ming Qian, Zimin Xia, Changkun Liu, Shuailei Ma, Wen Wang, Zeran Ke, Bin Tan, Hang Zhang, Gui-Song Xia
Abstract:
Generating a street‑level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade‑off: geometry‑colorization models achieve high geometric fidelity but are typically building‑focused and lack semantic diversity. In contrast, proxy‑based models use feed‑forward image‑to‑3D frameworks to generate holistic scenes by jointly learning geometry and texture, a process that yields rich content but coarse and unstable geometry. We attribute these geometric failures to the extreme viewpoint gap and sparse, inconsistent supervision inherent in satellite‑to‑street data. We introduce Sat3DGen to address these fundamental challenges, which embodies a geometry‑first methodology. This methodology enhances the feed‑forward paradigm by integrating novel geometric constraints with a perspective‑view training strategy, explicitly countering the primary sources of geometric error. This geometry‑centric strategy yields a dramatic leap in both 3D accuracy and photorealism. For validation, we first constructed a new benchmark by pairing the VIGOR‑OOD test set with high‑resolution DSM data. On this benchmark, our method improves geometric RMSE from 6.76m to 5.20m. Crucially, this geometric leap also boosts photorealism, reducing the Fréchet Inception Distance (FID) from ~40 to 19 against the leading method, Sat2Density++, despite using no extra tailored image‑quality modules. We demonstrate the versatility of our high‑quality 3D assets through diverse downstream applications, including semantic‑map‑to‑3D synthesis, multi‑camera video generation, large‑scale meshing, and unsupervised single‑image Digital Surface Model (DSM) estimation. The code has been released on https://github.com/qianmingduowan/Sat3DGen.
Authors:Chenxing Jiang, Zhe Tong, Pusen Gao, Peize Liu, Yang Xu, Chuan Fang, Ping Tan, Shaojie Shen
Abstract:
Stereo matching on top‑bottom equirectangular images provides an effective framework for full‑surround perception, as vertically aligned epipolar lines enable the use of advanced perspective stereo architectures that are largely driven by large‑scale datasets and monocular priors. However, the performance of such adaptations is severely limited by the scarcity of omnidirectional stereo datasets and the degradation of perspective monocular priors under spherical distortions. To address these challenges, we propose H‑OmniStereo, a zero‑shot omnidirectional stereo matching framework. First, we construct high‑quality synthetic dataset comprising over 2.8 million top‑bottom equirectangular stereo pairs to scale up training. Second, we introduce an equirectangular monocular normal estimator, specifically operating in a heading‑aligned coordinate system. Beyond providing distortion‑robust and cross‑view‑consistent geometric priors for establishing reliable correspondences in stereo matching, this design boosts training efficiency and accommodates train‑test FoV mismatches. Extensive experiments show that our approach achieves higher accuracy than existing methods on out‑of‑domain datasets and successfully generalizes to real‑world consumer camera setups using a single model. The model and dataset will be released at https://github.com/JIANG‑CX/H‑OmniStereo.
Authors:Hanxu Zhang, Chen Jia, Hui Liu, Xu Cheng, Fan Shi, Shengyong Chen
Abstract:
Achieving pixel‑level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing crack topology modeling with computational efficiency, often failing to reconcile high segmentation quality with low resource demands. To address these limitations, we propose the Ultra‑Compact Structure‑Calibrated Vision RWKV (SCRWKV), a network that achieves high‑precision modeling via a novel Structure‑Field Encoder (SFE) backbone while maintaining linear complexity. The SFE integrates the Adaptive Multi‑scale Cascaded Modulator (AMCM) to enhance texture representation and utilizes the Structure‑Calibrated Insight Unit (SCIU) as its core engine. Specifically, the SCIU employs the Geometry‑guided Bidirectional Structure Transformation (GBST) to capture topological correlations and integrates the Dynamic Self‑Calibrating Decay (DSCD) into Dy‑WKV to suppress noise propagation. Furthermore, we introduce a lightweight Cross‑Scale Harmonic Fusion (CSHF) decoder to achieve precise feature aggregation. Systematic evaluations on multiple benchmarks characterized by complex textures and severe interference demonstrate that SCRWKV, with only 1.22M parameters, significantly outperforms SOTA methods. Achieving an F1 score of 0.8428 and mIoU of 0.8512 on the TUT dataset, the model confirms its robust potential for efficient real‑world deployment. The code is available at https://github.com/zhxhzy/SCRWKV.
Authors:Xiyu Ren, Zhaowei Wang, Yiming Du, Zhongwei Xie, Chi Liu, Xinlin Yang, Haoyue Feng, Wenjun Pan, Tianshi Zheng, Baixuan Xu, Zhengnan Li, Yangqiu Song, Ginny Wong, Simon See
Abstract:
Memory is essential for large vision‑language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capability: long‑context LVLMs and memory‑augmented agents. However, no existing benchmark conducts a systematic comparison of the two on questions that genuinely require multimodal evidence. To close this gap, we introduce MEMLENS, a comprehensive benchmark for memory in multimodal multi‑session conversations, comprising 789 questions across five memory abilities (information extraction, multi‑session reasoning, temporal reasoning, knowledge update, and answer refusal) at four standard context lengths (32K‑256K tokens) under a cross‑modal token‑counting scheme. An image‑ablation study confirms that solving MEMLENS requires visual evidence: removing evidence images drops two frontier LVLMs below 2% accuracy on the 80.4% of questions whose evidence includes images. Evaluating 27 LVLMs and 7 memory‑augmented agents, we find that long‑context LVLMs achieve high short‑context accuracy through direct visual grounding but degrade as conversations grow, whereas memory agents are length‑stable but lose visual fidelity under storage‑time compression. Multi‑session reasoning caps most systems below 30%, and neither approach alone solves the task. These results motivate hybrid architectures that combine long‑context attention with structured multimodal retrieval. Our code is available at https://github.com/xrenaf/MEMLENS.
Authors:Zheng Hui, Yunlong Bai
Abstract:
Recent breakthroughs in video diffusion models have significantly accelerated the development of video editing techniques. However, existing methods often rely on inpainting video frames based on masked input, which requires extracting the target video mask in advance, and the precision of the segmentation directly affects the quality of the completion. In this paper, we present SEDiT, a novel one‑stage video Subtitle Erasure approach via One‑step Diffusion Transformer. We introduce a mask‑free inference approach that enables direct erasure of the targeted subtitle. The proposed one‑stage framework mitigates the sub‑optimality inherent in the two‑stage processing of prior models. Since subtitle removal is a localized editing task in which most pixels remain unchanged, the underlying distribution shift is minimal, making it well‑suited to one‑step generation under rectified flow. We empirically validate the reliability of one‑step denoising and further provide a formal theoretical justification. Under the localized‑editing structure of subtitle removal, the conditional optimal transport (OT) map and its induced rectified flow velocity field are Lipschitz continuous with respect to the latent variable, which underpins the theoretical feasibility of one‑step sampling. To address the challenge of long‑term temporal consistency, we adopt a hybrid training strategy by occasionally conditioning the model with a clean first‑frame latent. This facilitates temporal continuity, allowing each segment during inference to leverage the output of its predecessor. To avoid visible seams caused by cropping and reinserting processed targets, particularly in scenarios involving substantial motion, we feed the original video directly into SEDiT. Thanks to one‑step and chunk‑wise streaming inference, our method can efficiently handle native 1440p video with infinite length.
Authors:Sukju Oh, Sukkyu Sun
Abstract:
Online surgical phase recognition (SPR) underpins context‑aware operating‑room systems and requires committing to a prediction at every frame from past context alone. Surgical video poses three demands that natural‑video recognizers do not jointly address: procedures span tens of thousands of frames, time flows non‑uniformly as long routine stretches are punctuated by brief phase‑defining transitions, and the visual domain is narrow so backbone features are strongly correlated across channels. Existing recognizers either let per‑frame cost grow with elapsed length, or hold cost bounded but advance state at a uniform rate with channel‑independent dynamics, leaving the latter two demands unaddressed. We present SurgicalMamba, a causal SPR model built on Mamba2's structured state‑space duality (SSD) that holds per‑frame cost at O(d). It introduces three SSD‑compatible components that jointly address these demands: a dual‑path SSD block that separates long‑ and short‑term regimes at the level of recurrent state; intensity‑modulated stepping, a continuous‑time time‑warp that adapts the slow path's effective rate to phase‑relevant information; and state regramming, a per‑chunk Cayley rotation that opens cross‑channel mixing in the otherwise axis‑aligned SSM recurrence. The learned rotation planes inherit a phase‑aligned structure without any direct supervision, offering an interpretable internal signature of surgical workflow. Across seven public SPR benchmarks, SurgicalMamba reaches state‑of‑the‑art accuracy and phase‑level Jaccard under strict online evaluation: 94.6%/82.7% on Cholec80 (+0.7 pp/+2.2 pp over the strongest prior) and 89.5%/68.9% on AutoLaparo (+1.7 pp/+2.0 pp), at 238.74 fps on a single GPU. Ablations isolate the contribution of each component. The code is publicly available at https://github.com/sukjuoh/Surgical‑Mamba.
Authors:Zhuohao Chen, Zeng Li, Yifei Zhang, Chang Liu, Yu Zhou
Abstract:
Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine‑grained character strokes. Training such models relies on large amounts of annotated data. Recent self‑supervised approaches, such as Masked Image Modeling (MIM), alleviate this dependency by leveraging large‑scale unlabeled data. Yet most existing MIM methods operate at a single spatial scale and fail to capture the hierarchical nature of scene text. In this work, we introduce Masked Next‑Scale Prediction (MNSP), a unified self‑supervised framework designed to explicitly model cross‑scale structural evolution. The framework incorporates Next‑Scale Prediction (NSP), which learns hierarchical representations by predicting higher‑resolution features from lower‑resolution contexts. Naive scale prediction, however, tends to produce spatially diffuse attention, directing the model toward background regions rather than textual structures. MNSP resolves this limitation by jointly learning cross‑scale prediction and masked image reconstruction. NSP captures global layout priors across resolutions, while masked reconstruction imposes strong local constraints that guide attention toward informative text regions. A Multi‑scale Linguistic Alignment module further maintains semantic consistency across different resolutions. Extensive experiments demonstrate that MNSP achieves state‑of‑the‑art performance, reaching 86.2% average accuracy on the challenging Union14M benchmark and 96.7% across six standard datasets. Additional analyses show that our method improves robustness under extreme scale and layout variations. Code is available at https://github.com/CzhczhcHczh/MNSP
Authors:Marcello Ceresini, Federico Pirazzoli, Andrea Bertogalli, Lorenzo Cipelli, Filippo D'Addeo, Anthony Dell'Eva, Alessandro Paolo Capasso, Alberto Broggi
Abstract:
We present a flow‑matching planner for autonomous driving that directly outputs actionable control trajectories defined by acceleration and curvature profiles. The model is conditioned on a bird's‑eye‑view (BEV) raster of the surrounding scene and generates control sequences in a small number of Ordinary Differential Equations (ODE) integration steps, enabling low‑latency inference suitable for real‑time closed‑loop re‑planning. We train exclusively on urban scenarios (real urban city streets, intersections and roundabouts of the city of Parma, Italy) collected from a 2D traffic simulator with reactive agents, and evaluate in closed‑loop on both in‑distribution and markedly out‑of‑distribution environments, including multi‑lane highways and unseen urban scenarios. Our results show that the model generalizes reliably to these unseen conditions, maintaining stable closed‑loop control and successfully completing scenarios that differ substantially from the training distribution. We attribute this to the BEV representation, which provides a geometry‑centric view of the scene that is inherently less sensitive to distributional shifts, and to the flow‑matching formulation, which learns a smooth vector field that degrades gracefully under distribution shift. We provide video demonstrations of closed‑loop behavior at https://marcelloceresini.github.io/DirectControlFlowMatching.
Authors:Lukas Roming, Felix Lehnerer, Jonas V. Funk, Andreas Michel, Georg Maier, Thomas Längle, Jürgen Beyerer
Abstract:
Visual anomaly detection (AD) for industrial inspection is a highly relevant task in modern production environments. The problem becomes particularly challenging when training and deployment data differ due to changes in acquisition conditions during production. In the VAND 4.0 Industrial Track, models must remain robust under distribution shifts such as varying illumination and their performance is assessed on the MVTec AD 2 dataset. To address this setting, we propose a training‑free and class‑agnostic anomaly detection pipeline based on the work of SuperAD. Our approach improves generalization through several modifications designed to enhance robustness under distribution shifts. These adaptations include using a DINOv3 backbone, overlapping patch‑wise processing, intensity‑based augmentations, improved memory‑bank subsampling for better coverage of the data distribution, and iterative morphological closing for cleaner and more spatially consistent anomaly maps. Unlike methods that rely on class‑specific architectures or per‑class hyperparameter tuning, our method uses a single architecture and one shared hyperparameter configuration across all object classes. This makes the approach well suited for industrial deployment, where product variants and appearance changes must be handled with minimal adaptation effort. We achieve segmentation F1 scores of 62.61%, 57.42%, and 54.35% on test public, private, and private mixed of MVTec AD 2 respectively, thereby outperforming SuperAD and other state‑of‑the‑art methods. Code is available at https://github.com/LukasRoom/SuperADD.
Authors:Leon Davies, Qinggang Meng, Mohamad Saada, Baihua Li, Simon Sølvsten
Abstract:
Monocular 3D object detection remains challenging because metric size and depth are underdetermined by single‑view evidence, particularly under occlusion, truncation, and projection‑induced scale‑depth ambiguity. Although recent methods improve depth and geometric reasoning, metric size remains unstable in unified multi‑class settings, where class variability and partial visibility broaden plausible size modes. We propose MonoPRIO, a unified monocular 3D detector that targets this bottleneck through adaptive prior conditioning in the size pathway. MonoPRIO constructs class‑aware size prototypes offline, routes each decoder query to a soft mixture prior, applies uncertainty‑aware log‑space conditioning, and uses Cluster‑Aligned Prior (CAP) regularisation on matched positives during training. On the official KITTI test server, MonoPRIO achieves the strongest fully reported unified multi‑class result among methods reporting complete Car, Pedestrian, and Cyclist metrics. In the car‑only setting, it also achieves the strongest 3D bounding‑box AP across Easy/Moderate/Hard categories among compared methods without extra data, while using substantially less compute than MonoCLUE. Ablations and diagnostics show complementary gains from routed injection and CAP, with the largest benefits in ambiguity‑prone, partially occluded, and low‑data regimes. These findings indicate that adaptive priors are most effective when image evidence underdetermines metric size, while atypical geometry or extreme visibility loss can still cause mismatch between routed priors and true instance geometry. Code, trained models, result logs, and reproducibility material are available at https://github.com/bigggs/MonoPRIO.
Authors:Yuejiao Su, Xinshen Zhang, Zhen Ye, Lei Yao, Lap-Pui Chau, Yi Wang
Abstract:
Understanding human‑‑environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing multimodal large language models (MLLMs) still struggle with accurate interaction reasoning and fine‑grained pixel grounding. To this end, this paper introduces EARL, an Egocentric Analysis‑guided Reinforcement Learning framework that explicitly transfers coarse interaction semantics to query‑oriented answering and grounding. Specifically, EARL adopts a two‑stage parsing framework including coarse‑grained interpretation and fine‑grained response. The first stage holistically interprets egocentric interactions and generates a structured textual description. The second stage produces the textual answer and pixel‑level mask in response to the user query. To bridge the two stages, we extract a global interaction descriptor as a semantic prior, which is integrated via a novel Analysis‑guided Feature Synthesizer (AFS) for query‑oriented reasoning. To optimize heterogeneous outputs, including textual answers, bounding boxes, and grounding masks, we design a multi‑faceted reward function and train the response stage with GRPO. Experiments on Ego‑IRGBench show that EARL achieves 65.48% cIoU for pixel grounding, outperforming previous RL‑based methods by 8.37%, while OOD grounding results on EgoHOS indicate strong transferability to unseen egocentric grounding scenarios.
Authors:Saqib Nazir, Ardhendu Behera
Abstract:
Label‑free single‑cell imaging offers a scalable, non‑invasive alternative to fluorescence‑based cytometry, yet inferring molecular phenotypes directly from bright‑field morphology remains challenging. We present a unified Deep Learning (DL) framework that jointly performs White Blood Cell (WBC) classification and continuous protein‑expression regression from label‑free Differential Phase Contrast (DPC) images. Our model employs a Hybrid architecture that fuses convolutional fine‑grained texture features with transformer‑based global representations through a learnable cross‑branch gating module, enabling robust morpho‑molecular inference from DPC images. To support downstream interpretability, we further incorporate a Large Language Model (LLM) that generates concise, biologically grounded summaries of the predicted cell states. Experiments on the Berkeley Single Cell Computational Microscopy (BSCCM) and Blood Cells Image benchmarks demonstrate strong performance, achieving a 91.3% WBC classification accuracy and a 0.72 Pearson correlation for CD16 expression regression on BSCCM. These results underscore the promise of label‑free single‑cell imaging for cost‑effective hematological profiling, enabling simultaneous phenotype identification and quantitative biomarker estimation without fluorescent staining. The source code is available at https://github.com/saqibnaziir/Single‑Cell‑Phenotyping.
Authors:Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen, Laurence Tianruo Yang, Yurun Jin, Haishan Liu, Changti Wu, Hang Yuan, Cong Huang, Kai Chen
Abstract:
Robot imitation data are often multimodal: similar visual‑language observations may be followed by different action chunks because human demonstrators act with different short‑horizon intents, task phases, or recent context. Existing frame‑conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter‑chunk conflict and unstable execution. We introduce IntentVLA, a history‑conditioned VLA framework that encodes recent visual observations into a compact short‑horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12‑task ambiguity‑aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short‑horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines
Authors:Omkar Oak, Rukmini Nazre, Rujuta Budke, Suraj Sawant
Abstract:
Urban vegetation monitoring plays a vital role in understanding environmental changes, yet comprehensive datasets for this purpose remain limited. To address this gap, we present the Temporal Remote‑sensing Repository for Analyzing Change Detection (TERRA‑CD), a benchmark dataset comprising 5,221 Sentinel‑2 image pairs from 2019 and 2024, covering 232 cities across the USA and Europe. The dataset features three distinct annotation schemes: 4‑class land cover mapping masks, 3‑class vegetation change masks, and 13‑class semantic change masks capturing all possible land cover transitions. Using various deep learning approaches including Siamese networks, STANet variants, Bi‑SRNet, Changemask, Post‑Classification Comparison, and HRSCD strategies, we evaluated the dataset's effectiveness for both vegetation Multi‑class Change Detection as well as Semantic Change Detection. The proposed dataset and methods are available at https://github.com/omkarsoak/TERRA‑CD.
Authors:ZhiXin Sun
Abstract:
With the rapid evolution of computer vision, vision‑based methodologies for water level and river surface velocity estimation have reached significant maturity. Compared to traditional sensing, these techniques offer superior interpretability, automated data archiving, and enhanced system robustness. However, challenges such as environmental sensitivity, limited precision, and complex site calibration persist. This work proposes an integrated framework that synergizes state‑of‑the‑art (SOTA) vision models with statistical modeling. By leveraging physical priors and robust filtering strategies, we improve the accuracy of water level detection and flow estimation. Code will be available at https://github.com/sunzx97/Vision_Based_Water_Level_and_Flow_Estimation.git
Authors:Yu He, Fang Li, Haoyang Tong, Lichen Ma, Xinyuan Shan, Jingling Fu, Dong Chen, Luohang Liu, Junshi Huang, Yan Li
Abstract:
Recent advances in generative models have empowered impressive layered image generation, yet their success is largely confined to graphic design domains. The layering of in‑the‑wild images remains an underexplored problem, limiting fine‑grained editing and applications of images in real‑world scenarios. Specifically, challenges remain in scalable layered data and the modeling of object interaction in natural images, such as illumination effects and structural boundary. To address these bottlenecks, we propose a novel framework for high‑fidelity natural image decomposition. First, we introduce an Agent‑driven Data Decomposition (ADD) pipeline that orchestrates agents and tools to synthesize layered data without manual intervention. Utilizing this pipeline, we construct a large‑scale dataset, named LiWi‑100k, with over 100,000 high‑quality layered in‑the‑wild images. Second, we present a novel framework that jointly improves photometric fidelity and alpha boundary accuracy. Specifically, shadow‑guided learning explicitly models the illumination effects, and degradation‑restoration objective provides boundary‑correction supervision by recovering clean foreground image from degraded one. Extensive experiments demonstrate that our framework achieves state‑of‑the‑art (SoTA) performance in natural image decomposition, outperforming existing models in RGB L1 and Alpha IoU metrics. We will soon release our code and dataset.
Authors:Fuhao Li, Shaofeng You, Jiagao Hu, Yu Liu, Yuxuan Chen, Zepeng Wang, Fei Wang, Daiguo Zhou, Jian Luan
Abstract:
Evaluating object removal in images and videos remains challenging because the task is inherently one‑to‑many, yet existing metrics frequently disagree with human perception. Full‑reference metrics reward copy‑paste behaviors over genuine erasure; no‑reference metrics suffer from systematic biases such as favoring blurry results; and global temporal metrics are insensitive to localized artifacts within edited regions. To address these limitations, we propose RC (Removal Coherence), a pair of perception‑aligned metrics: RC‑S, which measures spatial coherence via sliding‑window feature comparison between masked and background regions, and RC‑T, which measures temporal consistency via distribution tracking within shared restored regions across adjacent frames. To validate RC and support community benchmarking, we further introduce PROVE‑Bench, a two‑tier real‑world benchmark comprising PROVE‑M, an 80‑video paired dataset with motion augmentation, and PROVE‑H, a 100‑video challenging subset without ground truth. Together, RC metrics and PROVE‑Bench form the PROVE (Perceptual RemOVal cohErence) evaluation framework for visual media. Experiments across diverse image and video benchmarks demonstrate that RC achieves substantially stronger alignment with human judgments than existing evaluation protocols. The code for RC metrics and PROVE‑Bench are publicly available at: https://github.com/xiaomi‑research/prove/.
Authors:Ling Li, Changjie Chen, Yuyan Wang, Jiaqing Lyu, Kenglun Chang, Yiyun Chen, Zhidong Deng
Abstract:
In multi‑view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a specific moment. While providing accurate spatial information, this traditional approach often overlooks the rich temporal dependencies between adjacent frames. We propose a novel 3D human pose estimation input method: the sparse interleaved input to address this. This method leverages images captured from different camera views at various time points (e.g., View 1 at time t and View 2 at time t+δ), allowing our model to capture rich spatio‑temporal information and effectively boost performance. More importantly, this approach offers two key advantages: First, it can theoretically increase the output pose frame rate by N times with N cameras, thereby breaking through single‑view frame rate limitations and enhancing the temporal resolution of the production. Second, using a sparse subset of available frames, our method can reduce data redundancy and simultaneously achieve better performance. We introduce the DenseWarper model, which leverages epipolar geometry for efficient spatio‑temporal heatmap exchange. We conducted extensive experiments on the Human3.6M and MPI‑INF‑3DHP datasets. Results demonstrate that our method, utilizing only sparse interleaved images as input, outperforms traditional dense multi‑view input approaches and achieves state‑of‑the‑art performance. The source code for this work is available at: https://github.com/lingli1724/DenseWarper‑ICLR2026
Authors:Jiahao Tian, Yiwei Wang, Gang Yu, Chi Zhang
Abstract:
Autoregressive video diffusion models support real‑time synthesis but suffer from error accumulation and context loss over long horizons. We discover that attention heads in AR video diffusion transformers serve functionally distinct roles as local heads for detail refinement, anchor heads for structural stabilization, and memory heads for long‑range context aggregation, yet existing methods treat them uniformly, leading to suboptimal KV cache allocation. We propose Head Forcing, a training‑free framework that assigns each head type a tailored KV cache strategy: local and anchor heads retain only essential tokens, while memory heads employ a hierarchical memory system with dynamic episodic updates for long‑range consistency. A head‑wise RoPE re‑encoding scheme further ensures positional encodings remain within the pretrained range. Without additional training, Head Forcing extends generation from 5 seconds to minute‑level duration, supports multi‑prompt interactive synthesis, and consistently outperforms existing baselines. Project Page: https://jiahaotian‑sjtu.github.io/headforcing.github.io/.
Authors:Yiheng Li, Yang Yang, Zichang Tan, Gao Li, Zhen Lei, Wenhao Wang
Abstract:
As the misuse of AI‑generated images grows, generalizable image detection techniques are urgently needed. Recent state‑of‑the‑art (SOTA) methods adopt aligned training datasets to reduce content, size, and format biases, empowering models to capture robust forgery cues. A common strategy is to employ reconstruction techniques, e.g., VAE and DDIM, which show remarkable results in diffusion‑based methods. However, such reconstruction‑based approaches typically introduce limited and homogeneous artifacts, which cannot fully capture diverse generative patterns, such as GAN‑based methods. To complement reconstruction‑based fake images with aligned yet diverse artifact patterns, we propose a GAN‑based upsampling approach that mimics GAN‑generated fake patterns while preserving content, size, and format alignment. This naturally results in two aligned but distinct types of fake images. However, due to the domain shift between reconstruction‑based and upsampling‑based fake images, direct mixed training causes suboptimal results, where one domain disrupts feature learning of the other. Accordingly, we propose a Separate Expert Fusion (SEF) framework to extract complementary artifact information and reduce inter‑domain interference. We first train domain‑specific experts via LoRA adaptation on a frozen foundational model, then conduct decoupled fusion with a gating network to adaptively combine expert features while retaining their specialized knowledge. Rather than merely benefiting GAN‑generated image detection, this design introduces diverse and complementary artifact patterns that enable SEF to learn a more robust decision boundary and improve generalization across broader generative methods. Extensive experiments demonstrate that our method yields strong results across 13 diverse benchmarks. Codes are released at: https://github.com/liyih/SEF_AIGC_detection.
Authors:Jiashun Zhu, Ronghao Fu, Jiasen Hu, Nachuan Xing, Xu Na, Xiao Yang, Zhiwen Lin, Weipeng Zhang, Lang Sun, Zhiheng Xue, Haoran Liu, Weijie Zhang, Bo Yang
Abstract:
Interpreting ultra‑high‑resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large‑scale scenes. Existing remote sensing vision‑language models can inspect local regions with zooming and cropping tools, but most exploration strategies follow either a one‑shot focus or a single sequential trajectory. Such single‑path exploration can lose global context, leave scattered regions unvisited, and revisit or count the same evidence multiple times. To this end, we propose GeoVista, a planning‑driven active perception framework for UHR remote sensing interpretation. Instead of committing to one zooming path, GeoVista first builds a global exploration plan, then verifies multiple candidate regions through branch‑wise local inspection, while maintaining an explicit evidence state for cross‑region aggregation and de‑duplication. To enable this behavior, we introduce APEX‑GRO, a cold‑start supervised trajectory corpus that reformulates diverse UHR tasks as Global‑Region‑Object interactive reasoning processes with a unified, scale‑invariant spatial representation. We further design an Observe‑Plan‑Track mechanism for global observation, adaptive region inspection, and evidence tracking, and align the model with a GRPO‑based strategy using step‑wise rewards for planning, localization, and final answer correctness. Experiments on RSHR‑Bench, XLRS‑Bench, and LRS‑VQA show that GeoVista achieves state‑of‑the‑art performance. Code and dataset are available at https://github.com/ryan6073/GeoVista
Authors:Yubo Zhao, Yujin Chai, Yunao Dong, Chengfeng Zhao, Zijiao Zeng, Yuan Liu, Chi-Keung Tang
Abstract:
Recovering 4D human‑object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation‑based learning. Recent methods can reconstruct temporally coherent human and object trajectories, but these trajectories often remain visual artifacts while failing to preserve stable contact, functional manipulation, or physical plausibility when used as reference motions for humanoid‑object simulation. This reveals a fundamental interaction gap: HOI reconstruction should not stop at tracking a human and an object, but should recover the relation that makes their motion a coherent interaction. We introduce HA‑HOI, a framework for reconstructing physically plausible 4D HOI animation from in‑the‑wild monocular videos. Instead of treating the human and object as independent entities in an ambiguous monocular 3D space, we propose a human‑first, object‑follow formulation. The human motion is recovered as the interaction anchor, and the object is reconstructed, aligned, and refined relative to the human action. The resulting kinematic trajectory is then projected into a physics‑based humanoid‑object simulation, where it acts as a teacher trajectory for stable physical rollout. Across benchmark and in‑the‑wild videos, HA‑HOI improves human‑object alignment, contact consistency, temporal stability, and simulation readiness over prior monocular HOI reconstruction methods. By moving beyond visually plausible trajectory recovery toward physically grounded interaction animation, our work takes a step toward turning general monocular HOI videos into scalable demonstrations for humanoid‑object behavior. Project page: https://knoxzhao.github.io/real2sim_in_HOI/
Authors:Ledun Zhang, Yatu Ji, Xufei Zhuang, Xinying Yao
Abstract:
Existing object removal tools often rely on manual masks or text prompts, making precise removal difficult for non‑expert users in complex scenes and often leading to incomplete removal or unnatural background completion. To address this issue, we present ClickRemoval, an open‑source interactive object removal tool built on pretrained Stable Diffusion models and driven solely by user clicks. Without additional training, hand‑drawn masks, or text descriptions, ClickRemoval localizes target objects and restores the background through self‑attention modulation during denoising. Experiments show that ClickRemoval achieves competitive results across quantitative metrics and user studies. We release a complete software package at https://github.com/zld‑make/ClickRemoval under the Apache‑2.0 license.
Authors:Yize Liu, Siyuan Yan, Ming Hu, Lie Ju, Xieji Li, Feilong Tang, Wei Feng, Zongyuan Ge
Abstract:
Dermatological diagnosis requires integrating fine‑grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (MLLMs) facilitate interactive medical image analysis, their application in dermatology is hindered by insufficient domain‑specific grounding and hallucinations. To address these issues, we propose DermAgent, a collaborative multi‑tool agent that orchestrates seven specialized vision and language modules within a Plan‑Execute‑Reflect framework. DermAgent delivers stepwise, traceable diagnostic reasoning through three core components. First, it employs complementary visual perception tools for comprehensive morphological description, dermoscopic concept annotation, and disease diagnosis. Second, to overcome the lack of domain prior, a dual‑modality retrieval module anchors every prediction in external evidence by cross‑referencing 413,210 diagnosed image cases and 3,199 clinical guideline chunks. To further mitigate hallucinations, a deterministic critic module conducts strict post‑hoc auditing via confidence, coverage, and conflict gates, automatically detecting inter‑source disagreements to trigger targeted self‑correction. Extensive experiments on five dermatology benchmarks demonstrate that DermAgent consistently outperforms state‑of‑the‑art MLLMs and medical agent baselines across zero‑shot fine‑grained disease diagnosis, concept annotation, and clinical captioning tasks, exceeding GPT‑4o by 17.6% in skin disease diagnostic accuracy and 3.15% in captioning ROUGE‑L. Our code is available at https://github.com/YizeezLiu/DermAgent.
Authors:Yuanhang Yao, Ping Qian, Zhu Liu, Long Ma, Weimin Wang
Abstract:
Single‑frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel‑wise annotations are expensive. Point supervision with online label evolution reduces annotation cost; however, lightweight CNN detectors often lack sufficient semantics, leading to noisy pseudo‑masks and unstable optimization. To address this, we propose a hierarchical VFM‑driven knowledge distillation framework that uses a frozen Vision Foundation Model (VFM) during training. We formulate point‑supervised learning as a bilevel optimization process: the inner loop adapts a VFM‑embedded teacher on reweighted training samples, while the outer loop transfers validation‑guided knowledge to a lightweight student to mitigate pseudo‑label noise and training‑set bias. We further introduce Semantic‑Conditioned Affine Modulation (SCAM) to inject VFM semantics into CNN features at multiple layers. In addition, a dynamic collaborative learning strategy with cluster‑level sample reweighting enhances robustness to imperfect pseudo‑masks. Experiments on diverse challenging cases across multiple ISTD backbones demonstrate consistent improvements in detection accuracy and training stability. Our code is available at https://github.com/yuanhang‑yao/semantic‑prior.
Authors:Yang Yue, Fangyun Wei, Tianyu He, Jinjing Zhao, Zanlin Ni, Zeyu Liu, Jiayi Guo, Lei Shi, Yue Dong, Li Chen, Ji Li, Gao Huang, Dong Chen
Abstract:
Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine‑grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete‑tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing diverse content uniformly. To address this, we propose InsightTok, a simple yet effective discrete visual tokenization framework that enhances text and face fidelity through localized, content‑aware perceptual losses. With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details. Overall, our results highlight the potential of specialized supervision in tokenizer training for advancing discrete image generation.
Authors:David Huang, Guile Wu, Chengjie Huang, Bingbing Liu, Dongfeng Bai
Abstract:
Recent feed‑forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per‑scene optimization paradigm by enabling effective multi‑view reconstruction in a single forward pass. However, most existing methods struggle to achieve a balance between reconstruction quality and computational efficiency, which limits their scalability and efficiency. Although some efficient visual geometry transformers have recently emerged, they typically use the same sparsity ratio across layers and frames and lack mechanisms to adaptively learn representative tokens to capture global relationships, leading to suboptimal performance. In this work, we propose TurboVGGT, a novel approach that employs an efficient visual geometry transformer with adaptive alternating attention for fast multi‑view 3D reconstruction. Specifically, TurboVGGT employs an end‑to‑end trainable framework with adaptive sparse global attention guided by adaptive sparsity selection to capture global relationships across frames and frame attention to aggregate local details within each frame. In the adaptive sparse global attention, TurboVGGT adaptively learns representative tokens with varying sparsity levels for global geometry modeling, considering that token importance varies across frames, attention layers operate tokens at different levels of abstraction, and global dependencies rely on structurally informative regions. Extensive experiments on multiple 3D reconstruction benchmarks demonstrate that TurboVGGT achieves fast multi‑view reconstruction while maintaining competitive reconstruction quality compared with state‑of‑the‑art methods. Project page: https://turbovggt.github.io/.
Authors:Yang Zheng, Wen Li, Zhaoqiang Liu
Abstract:
Diffusion models (DMs) have exhibited remarkable efficacy in various image restoration tasks. However, existing approaches typically operate within the high‑dimensional pixel space, resulting in high computational overhead. While methods based on latent DMs seek to alleviate this issue by utilizing the compressed latent space of a variational autoencoder, they require repeated encoder‑decoder inference. This introduces significant additional computational burdens, often resulting in runtime performance that is even inferior to that of their pixel‑space counterparts. To mitigate the computational inefficiency, this work proposes projecting data into lower‑dimensional subspaces using dynamic resolution DMs to accelerate the inference process. We first fine‑tune pre‑trained DMs for dynamic resolution priors and adapt DPS and DAPS, which are two widely used pixel‑space methods for general image restoration tasks, into the proposed framework, yielding methods we refer to as SubDPS and SubDAPS, respectively. Given the favorable inference speed and reconstruction fidelity of SubDAPS, we introduce an enhanced variant termed SubDAPS++ to further boost both reconstruction efficiency and quality. Empirical evaluations across diverse image datasets and various restoration tasks demonstrate that the proposed methods outperform recent DM‑based approaches in the majority of experimental scenarios. The code is available at https://github.com/StarNextDay/SubDAPS.git.
Authors:Junchao Zhu, Ruining Deng, Junlin Guo, Tianyuan Yao, Chongyu Qu, Juming Xiong, Zhengyi Lu, Yanfan Zhu, Marilyn Lionts, Yuechen Yang, Yu Wang, Shilin Zhao, Haichun Yang, Yuankai Huo
Abstract:
Inferring spatially resolved gene expression from histology images offers a cost‑effective complement to spatial transcriptomics (ST). However, existing methods reduce this task to a simple morphology‑to‑expression mapping, where visual similarity does not guarantee molecular consistency. Meanwhile, single‑cell data has amassed rich resources far surpassing the scale of ST data, yet it remains underexplored in vision‑omics modeling. Furthermore, current approaches commit to a monolithic paradigm with bottlenecks, unable to balance expressive flexibility with biological fidelity. To bridge these gaps, we propose DUET, a novel dual‑paradigm framework that synergizes parametric prediction and memory‑based retrieval under cellular inductive priors. DUET implements a parallel regression‑retrieval paradigm, adaptively reconciling the outputs of its complementary pathways. To mitigate aleatoric vision ambiguity, we incorporate large‑scale single‑cell references to impose molecular states as biological constraints for faithful learning. Building upon structural refinement, we further design a lightweight adapter to dynamically assign branch preference across spatial contexts to achieve optimal performance. Extensive experiments on three public datasets across varied gene scales demonstrate that DUET achieves SOTA performance, with consistent gains contributed by each proposed component. Code is available at https://github.com/Junchao‑Zhu/DUET
Authors:Wei Dong, Han Zhou, Terry Ji, Guanhua Zhao, Shahab Asoodeh, Yulun Zhang, Guangtao Zhai, Jun Chen, Xiaohong Liu
Abstract:
Adverse weather removal (AWR) in real‑world images remains challenging due to heterogeneous and unseen degradations, while distortion‑driven training often yields overly smooth results. We propose PVRF, a unified framework that integrates zero‑shot soft weather perceptions with velocity‑constrained rectified‑flow refinement. PVRF introduces an AWR‑specific question answering module (AWR‑QA) that uses frozen vision‑‑language models (VLMs) to estimate soft probabilities of weather types and low‑level attribute scores. These perceptions condition restoration networks via attribute‑modulated normalization (AMN) and weather‑weighted adapters (WWA), producing an anchor estimate for refinement. We then learn a terminal‑consistent residual rectified flow with perception‑adaptive source perturbation and a terminal‑consistent velocity parameterization to stabilize learning near the terminal regime. Extensive experiments show that PVRF improves both fidelity and perceptual quality over state‑of‑the‑art baselines, with strong cross‑dataset generalization on single and combined degradations. Code will be released at https://github.com/dongw22/PVRF.
Authors:Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia
Abstract:
Diffusion Transformers (DiTs) and related flow‑based architectures are now among the strongest text‑to‑image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood. In this work, we study massive activations: a small subset of hidden‑state channels whose responses are consistently much larger than the rest. We show that, despite their sparsity, these few channels effectively draw the whole picture, in three complementary senses. First, they are functionally critical: a controlled disruption probe that zeroes the massive channels causes a sharp collapse in generation quality, while disrupting an equally‑sized set of low‑statistic channels has marginal effect. Second, they are spatially organized: restricting image‑stream tokens to massive channels and clustering them yields coherent partitions that closely align with the main subject and salient regions, exposing a structured spatial code hidden inside an apparently outlier‑like subspace. Third, they are transferable: transporting massive activations from one prompt‑conditioned trajectory into another, shifts the final image toward the source prompt while preserving substantial content from the target, producing localized semantic interpolation rather than unstructured pixel blending. We exploit this property in two use cases: text‑conditioned and image‑conditioned semantic transport, where massive activations transport enables prompt interpolation and subject‑driven generation without any additional training. Together, these results recast massive activations not as activation anomalies, but as a sparse prompt‑conditioned carrier subspace that organizes and controls semantic information in modern DiT models.
Authors:Dongxia Liu, Jie Ma, Xiaochen Yang, Jiancheng Zhang, Bin Xia, Zhehan Kan, Nisha Huang, Jun Liang, Wenming Yang, Jin Li
Abstract:
The creation of cinematic‑quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor‑intensive and computationally expensive within traditional production workflows. While generative diffusion models have shown promise in diverse artistic workflows, their capacity for high‑fidelity animal simulation remains largely unexploited. We present MoZoo, a generative dynamics solver that bypasses conventional refinement to synthesize high‑fidelity animal videos from coarse meshes under multimodal guidance. We propose Role‑Aware RoPE (RAR‑RoPE) which employs role‑based index remapping to synchronize motion alignment while decoupling reference information via fixed temporal offsets. Complementing this, Asymmetric Decoupled Attention partitions the latent sequence to enforce a unidirectional information flow, effectively preventing feature interference and improving computational efficiency. To address the scarcity of high‑quality training data, we introduce MoZoo‑Data, a synthetic‑to‑real pipeline that leverages a rendering engine and an inverse mapping approach to construct a large‑scale dataset of paired sequences. Furthermore, we establish MoZooBench, a comprehensive benchmark with 120 mesh‑video pairs. Experimental results demonstrate that MoZoo achieves high‑fidelity fur simulation across diverse animal skeletons and layouts, preserving superior temporal and structural consistency.
Authors:Minghao Sun, Chongyang Xu, Yitao Xie, Buzhen Huang, Kun Li
Abstract:
Multi‑person 3D reconstruction is pivotal for real‑world interaction analysis, yet remains challenging due to severe occlusions and depth ambiguity. Current approaches typically rely on single‑modality inputs, which inherently lack geometric guidance. Furthermore, these methods often reconstruct subjects in isolation, neglecting the collective group context essential for resolving ambiguities in crowded scenes. To address these limitations, we propose Contrastive Multi‑modal Hypergraph Reasoning to synergize semantic, geometric, and pose cues for crowd reconstruction. We first initialize robust node representations by combining RGB features, geometric priors, and occlusion‑aware incomplete poses. Additionally, we introduce a pelvis depth indicator as a global spatial anchor, aligning visual features with a metric‑scale‑agnostic depth ordering. Subsequently, we construct a shared‑topology hypergraph that moves beyond pairwise constraints to model higher‑order crowd dynamics. To improve feature fusion, we design a hypergraph‑based contrastive learning scheme that jointly enhances intra‑modal discriminability and enforces cross‑modal orthogonality. This mechanism enables the network to propagate global context effectively, allowing it to infer missing information even under severe occlusion. Extensive experiments on the Panoptic and GigaCrowd benchmarks confirm that our method achieves new state‑of‑the‑art performance. Code and pre‑trained models are available at https://github.com/SunMH‑try/CoMHR.
Authors:Ido Sobol, Kihyuk Sohn, Yoav Blum, Egor Zakharov, Max Bluvstein, Andrea Vedaldi, Or Litany
Abstract:
We often aim to generate images that are both photorealistic and 3D‑consistent, adhering to precise geometry, material, and viewpoint controls. Typically, this is achieved by fine‑tuning an image generator, pre‑trained on billions of real images, using renders of synthetic 3D assets, where annotations for control signals are available. While this approach can learn the desired controls, it often compromises the realism of the images due to domain gap between photographs and renders. We observe that this issue largely arises from the model learning an unintended association between the presence of control signals and the synthetic appearance of the images. To address this, we introduce Realiz3D, a lightweight framework for training diffusion models, that decouples controls and visual domain. The key idea is to explicitly learn visual domain, real or synthetic, separately from other control signals by introducing a co‑variate that, fed into small residual adapters, shifts the domain. Then, the generator can be trained to gain controllability, without fitting to specific visual domain. In this way, the model can be guided to produce realistic images even when controls are applied. We enhance control transferability to the real domain by leveraging insights about roles of different layers and denoising steps in diffusion‑based generators, informing new training and inference strategies that further mitigate the gap. We demonstrate the advantages of Realiz3D in tasks as text‑to‑multiview generation and texturing from 3D inputs, producing outputs that are 3D‑consistent and photorealistic.
Authors:Zijie Wu, Lixin Xu, Puhua Jiang, Sicong Liu, Chunchao Guo, Xiang Bai
Abstract:
Video‑guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical deployment faces a critical yet frequently overlooked hurdle: the pose misalignment dilemma. In real‑world scenarios, the initial pose of a user‑provided static mesh rarely aligns with the starting frame of a reference video. Naively forcing a mesh to follow a mismatched trajectory inevitably leads to severe geometric distortion or animation failure. To address this, we present Rectified Dynamic Mesh (R‑DMesh), a unified framework designed to generate high‑fidelity 4D meshes that are ``rectified'' to align with video context. Unlike standard motion transfer approaches, our method introduces a novel VAE that explicitly disentangles the input into a conditional base mesh, relative motion trajectories, and a crucial rectification jump offset. This offset is learned to automatically transform the arbitrary pose of the input mesh to match the video's initial state before animation begins. We process these components via a Triflow Attention mechanism, which leverages vertex‑wise geometric features to modulate the three orthogonal flows, ensuring physical consistency and local rigidity during the rectification and animation process. For generation, we employ a Rectified Flow‑based Diffusion Transformer conditioned on pre‑trained video latents, effectively transferring rich spatio‑temporal priors to the 3D domain. To support this task, we construct Video‑RDMesh, a large‑scale dataset of over 500k dynamic mesh sequences specifically curated to simulate pose misalignment. Extensive experiments demonstrate that R‑DMesh not only solves the alignment problem but also enables robust downstream applications, including pose retargeting and holistic 4D generation.
Authors:Lavsen Dahal, Yubraj Bhandari, Geoffrey Rubin, Joseph Y. Lo
Abstract:
Automated CT triage requires models that are simultaneously accurate across diverse pathologies and reliable under institutional shift. While Vision Transformers provide strong visual representations, many clinically significant findings are defined by quantitative imaging biomarkers rather than appearance alone. We introduce JANUS, a physiology‑guided dual‑stream architecture that conditions visual embeddings on macro‑radiomic priors via Anatomically Guided Gating. On the MERLIN test set (N=5082), JANUS attains macro‑AUROC 0.88 and AUPRC 0.74, outperforming all reproduced baselines. It generalizes to an external dataset N=2000; AUROC 0.87), with the largest gains on findings defined by size and attenuation as well as improved calibration on both datasets. We further quantify prediction suppression using the Physiological Veto Rate (PVR), showing that under domain shift JANUS reduces high‑confidence false positives substantially more often than true positives. Together, these results are consistent with physically grounded conditioning that improves both discrimination and reliability in CT triage. Code is made publicly available at github repository https://github.com/lavsendahal/janus and model weights are at https://huggingface.co/lavsendahal/janus.
Authors:Minjoon Jung, Byoung-Tak Zhang, Lorenzo Torresani
Abstract:
Video temporal grounding (VTG) takes an untrimmed video and a natural‑language query as input and localizes the temporal moment that best matches the query. Existing methods rely on large, task‑specific datasets requiring costly manual annotation. We introduce EvoGround, a framework of two coupled self‑evolving agents, a proposer and a solver, that learn temporal grounding from raw videos without any human‑labeled data. The proposer generates query‑‑moment pairs from raw videos, while the solver learns to ground them and feeds back signals that improve the proposer in return. Through this self‑reinforcing reinforcement‑learning loop, the two agents are initialized from the same backbone and mutually improve across iterations. Trained on 2.5K unlabeled videos, EvoGround matches or surpasses fully supervised models across multiple VTG benchmarks, while emerging as a state‑of‑the‑art fine‑grained video captioner without manual labels.
Authors:Guney Tombak, Ertunc Erdil, Ender Konukoglu
Abstract:
Cross‑modal 3D medical image analysis requires voxelwise representations that remain anatomically consistent across imaging contrasts, scanners, and acquisition protocols. Recent work has shown that frozen 2D Vision Transformer (ViT) foundation models can support such representations, but typical pipelines extract features along a single anatomical axis and adapt those features inside a registration solver for one image pair at a time, leaving complementary viewing directions unused and producing representations that do not transfer to new volumes. We introduce VoxCor, a training‑free fit‑‑transform method for reusable volumetric feature representations from frozen 2D ViT foundation models. During an offline fitting phase, VoxCor combines triplanar ViT inference with a compact closed‑form weighted partial least squares (WPLS) projection that uses fitting‑time voxel correspondences to select modality‑stable anatomical directions in the triplanar feature space. At transform time, new volumes are mapped by triplanar ViT inference and linear projection alone, without fine‑tuning or registration. Voxel correspondences can then be queried directly by nearest‑neighbor search. We evaluate VoxCor on intra‑subject Abdomen MR‑‑CT and inter‑subject HCP T2w‑‑T1w tasks using deformable registration, voxelwise k‑nearest‑neighbor segmentation, and segmentation‑center landmark localization. VoxCor improves the hardest cross‑subject, cross‑modality transfer settings, reduces encoder sensitivity for dense correspondence transfer, and yields registration performance competitive with handcrafted descriptors and learned 3D features. This positions VoxCor as a reusable feature layer for downstream multimodal analysis beyond pairwise registration. Code, configuration files, and implementation details are publicly available on GitHub at \hrefhttps://github.com/guneytombak/VoxCorguneytombak/VoxCor.
Authors:Feiyu Tan, Qi Xie, Zongben Xu, Deyu Meng
Abstract:
Image restoration is an inherently ill posed inverse problem. Equivariant networks that embed geometric symmetry priors can mitigate this ill posedness and improve performance. However, current understanding of the relationship between network equivariance and data symmetry remains largely heuristic. Particularly for real world data with imperfect symmetry, existing research lacks a systematic theoretical framework to quantify symmetry, select transformation groups, or evaluate model data alignment. To bridge this gap, we conduct an analysis from an optimization perspective and formalize the intrinsic relationship among data symmetry priors, model equivariance, and generalization capability. Specifically, we propose for the first time a quantifiable definition of non strict symmetry at the dataset level (rather than sample level) and use it as a constraint to formulate the restoration inverse problem. We then show that the equivariance for restoration models can be naturally derived from this inverse problems incorporated the proposed symmetry constraints, and that the equivariance error of the optimal restoration operator is strictly bounded by the data symmetry error and the discretization mesh size. Furthermore, by analyzing the network's empirical risk, we demonstrate that aligning equivariance with data symmetry optimizes the bias variance trade off, minimizing the total expected risk. Guided by these insights, we propose a Sample Adaptive Equivariant Network that uses a hypernetwork and transformation learnable equivariant convolutions to dynamically align with each sample's inherent symmetry. Extensive experiments on super resolution, denoising, and deraining validate our theoretical findings and show significant superiority over standard baselines and traditional equivariant models. Our code and supplementary material are available at https://github.com/tanfy929/SA‑Conv.
Authors:Christina Kassab, Hyeonjae Gil, Matías Mattamala, Ayoung Kim, Maurice Fallon
Abstract:
Scene graphs are becoming a standard representation for robot navigation, providing hierarchical geometric and semantic scene understanding. However, most scene graph mapping methods rely on depth cameras or LiDAR sensors. In this work, we present LEXI‑SG, the first dense monocular visual mapping system for open‑vocabulary 3D scene graphs using only RGB camera input. Our approach exploits the semantic priors of open‑vocabulary foundation models to partition the scene into rooms, deferring feed‑forward reconstruction to when each room is fully observed ‑‑ enabling scalable dense mapping without sliding‑window scale inconsistencies. We propose a room‑based factor graph formulation to globally align room reconstructions while preserving local map consistency and naturally imposing the semantic scene graph hierarchy. Within each room, we further support open‑vocabulary object segmentation and tracking. We validate LEXI‑SG on indoor scenes from the Habitat‑Matterport 3D and self‑collected egocentric office sequences. We evaluate its performance against existing feed‑forward SLAM methods, as well as established scene graphs baselines. We demonstrate improved trajectory estimation and dense reconstruction, as well as, competitive performance in open‑vocabulary segmentation. LEXI‑SG shows that accurate, scalable, open‑vocabulary 3D scene graphs can be achieved from monocular RGB alone. Our project page and office sequences are available here: https://ori‑drs.github.io/lexisg‑web/.
Authors:Yuchao Gu, Guian Fang, Yuxin Jiang, Weijia Mao, Song Han, Han Cai, Mike Zheng Shou
Abstract:
Few‑step video generation has been significantly advanced by consistency distillation. However, the performance of consistency‑distilled models often degrades as more sampling steps are allocated at test time, limiting their effectiveness for any‑step video diffusion. This limitation arises because consistency distillation replaces the original probability‑flow ODE trajectory with a consistency‑sampling trajectory, weakening the desirable test‑time scaling behavior of ODE sampling. To address this limitation, we introduce AnyFlow, the first any‑step video diffusion distillation framework based on flow maps. Instead of distilling a model for only a few fixed sampling steps, AnyFlow optimizes the full ODE sampling trajectory. To this end, we shift the distillation target from endpoint consistency mapping (z_t\rightarrow z_0) to flow‑map transition learning (z_t\rightarrow z_r) over arbitrary time intervals. We further propose Flow Map Backward Simulation, which decomposes a full Euler rollout into shortcut flow‑map transitions, enabling efficient on‑policy distillation that reduces test‑time errors (i.e., discretization error in few‑step sampling and exposure bias in causal generation). Extensive experiments across both bidirectional and causal architectures, at scales ranging from 1.3B to 14B parameters, demonstrate that AnyFlow achieves performance matches or surpasses consistency‑based counterparts in the few‑step regime, while scaling with sampling step budgets.
Authors:Cenwei Zhang, Suncheng Xiang, Lei You
Abstract:
Medical segmentation foundation models such as SAM and MedSAM provide strong prompt‑driven segmentation, but their image encoders are still too large for many clinical settings. Compression is also risky in medicine because a model can keep high Dice while losing boundary fidelity. We propose MedCore, a structured pruning framework for MedSAM. The main idea is to preserve two kinds of structures: structures that became important during SAM‑to‑MedSAM adaptation, and structures that have high boundary leverage. We identify the first type by a dual‑intervention score that compares zeroing a group with resetting it to its original SAM weight. We identify the second type by boundary‑aware Fisher estimation. We also introduce a boundary leverage principle, which shows that compression‑induced boundary displacement is controlled by logit perturbation on the boundary divided by the logit spatial gradient. This principle explains why boundary metrics can degrade even when Dice remains high. On polyp segmentation benchmarks, MedCore reduces parameters by 60.0% and FLOPs by 58.4% while achieving Dice 0.9549, Boundary F1 0.6388, and HD95 5.14 after recovery fine‑tuning. It also reaches 86.6% parameter reduction and 90.4G FLOPs with strong boundary quality. Our analysis further shows that MedSAM lies in a head‑fragile boundary regime: head‑pruning steps have 2.887 times larger 95th‑percentile boundary leverage than MLP‑pruning steps, and this logit‑level effect is consistent with BF1 and HD95 degradation. Our code is available at https://github.com/cenweizhang/MedCore.
Authors:Vladislav Makarov, Mark Gizetdinov, Dmitry Yudin
Abstract:
Scene graph generation provides a compact structured representation for visual perception, but accurate and fast graph prediction from images and videos remains challenging. Recent VLM‑based methods can generate scene graphs end‑to‑end as structured text, yet often produce long outputs with irrelevant objects and relations. We present SceneGraphVLM, a compact method for image and video scene graph generation with small visual language models. SceneGraphVLM serializes graphs in a token‑efficient TOON format and trains the model in two stages: supervised fine‑tuning followed by reinforcement learning with hallucination‑aware rewards that balance relation coverage and precision while penalizing unsupported objects and relations. For videos, the model can optionally condition each frame on the previously generated graph, providing lightweight short‑term context without tracking or post‑processing. We evaluate SceneGraphVLM on PSG, PVSG, and Action Genome. With compact VLMs and vLLM‑accelerated decoding, SceneGraphVLM achieves a strong quality‑speed trade‑off, improves precision‑oriented SGG metrics while preserving reasonable recall, and generates complete scene graphs with approximately one‑second latency. Code and implementation details are available at: https://github.com/markus0440/SceneGraphVLM.git.
Authors:Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
Abstract:
In this paper, we propose GTA‑VLA(Guide, Think, Act), an interactive Vision‑Language‑Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visual cues. Existing VLA models learn a direct "Sense‑to‑Act" mapping from multimodal observations to robot actions. While effective within the training distribution, such tightly coupled policies are brittle under out‑of‑domain (OOD) shifts and difficult to correct when failures occur. Although recent embodied Chain‑of‑Thought (CoT) approaches expose intermediate reasoning, they still lack a mechanism for incorporating human spatial guidance, limiting their ability to resolve visual ambiguities or recover from mistakes. To address this gap, our framework allows users to optionally guide the policy with spatial priors, such as affordance points, boxes, and traces, which the subsequent reasoning process can directly condition on. Based on these inputs, the model generates a unified spatial‑visual Chain‑of‑Thought that integrates external guidance with internal task planning, aligning human visual intent with autonomous decision‑making. For practical deployment, we further couple the reasoning module with a lightweight reactive action head for efficient action execution. Extensive experiments demonstrate the effectiveness of our approach. On the in‑domain SimplerEnv WidowX benchmark, our framework achieves a state‑of‑the‑art 81.2% success rate. Under OOD visual shifts and spatial ambiguities, a single visual interaction substantially improves task success over existing methods, highlighting the value of interactive reasoning for failure recovery in embodied control. Details of the project can be found here: https://signalispupupu.github.io/GTA‑VLA_ProjPage/
Authors:Wudi Chen, Zhiyuan Zha, Xin Yuan, Shigang Wang, Bihan Wen, Jiantao Zhou, Gang Yan, Zipei Fan, Ce Zhu
Abstract:
Recent advances have demonstrated that coded aperture snapshot spectral imaging (CASSI) systems show great potential for capturing 3D hyperspectral images (HSIs) from a single 2D measurement. Despite the inherent spectral continuity of scenes captured by CASSI, most existing reconstruction methods are restricted to fixed, discrete spectral outputs, thereby precluding continuous spectral reconstruction or spectral super‑resolution. To address this challenge, we propose Phy‑CoSF, which synergizes deep unfolding networks with implicit neural representations, establishing a new paradigm for continuous spectral reconstruction and super‑resolution in CASSI. Specifically, we propose a two‑phase architecture that bridges discrete‑wavelength training with continuous spectral rendering, enabling the synthesis of high‑fidelity HSIs at arbitrary target wavelengths. At the core of our framework lies the continuous spectral fields (CoSF) module, embedded within each unfolding stage as a dynamic prior, which comprises a triple‑branch cross‑domain feature mixer for comprehensive spatial‑frequency‑channel feature fusion, alongside a spectral synthesis head that generates spectral intensities by querying continuous wavelength coordinates. Extensive experimental results demonstrate that Phy‑CoSF not only achieves continuous modeling at arbitrary spectral resolutions but also outperforms many state‑of‑the‑art methods in both reconstruction fidelity and spectral detail preservation. Our code and more results are available at: https://github.com/PaiDii/Phy‑CoSF.git.
Authors:Jaeyung Kim, YoungJoon Yoo
Abstract:
Vector Quantized Variational Autoencoder (VQ‑VAE) has become a fundamental framework for learning discrete representations in image modeling. However, VQ‑VAE models must tokenize entire images using a finite set of codebook vectors, and this capacity limitation restricts their ability to capture rich and diverse representations. In this paper, we propose ArcCosine Additive Margin VQ‑VAE (ArcVQ‑VAE), a novel vector quantization framework that introduces a spherical angular‑margin prior (SAMP) for the codebook of a conventional VQ‑VAE. The proposed SAMP consists of Ball‑Bounded Norm Regularization, which constrains all codebook vectors within a time‑dependent Euclidean ball, and ArcCosine Additive Margin Loss, which encourages greater angular separability among latent vectors. This formulation promotes more discriminative and uniformly dispersed latent representations within the constrained space, thereby improving effective latent‑space coverage and leading to improved codebook utilization. Experimental results on standard image reconstruction and generation tasks show that ArcVQ‑VAE achieves competitive performance against baseline models in terms of reconstruction accuracy, representation diversity, and sample quality. The code is available at: https://github.com/goals4292/ArcVQ‑VAE
Authors:Tiange Zhang, Rongqun Lin, Xiandong Meng, Haofeng Wang, Xing Tian, Qi Zhang, Siwei Ma
Abstract:
Content‑adaptive compression has always been a key direction in neural video coding (NVC), aiming to mitigate the domain gap between training and testing data. Such gaps often arise from distributional discrepancies between training and inference data, which may cause noticeable performance degradation when the testing content differs from the training distribution. To tackle this challenge, we propose DCVC‑DT, a domain transfer enhanced neural video compression framework. Specifically, we design a lightweight online domain transfer (DT) mechanism that dynamically adapts the encoded latent representation during inference, effectively bridging the domain gap without modifying the encoder or decoder parameters. In addition, we develop a frame‑level dynamic RD (Rate and Distortion) adjustment scheme that actively regulates the ratio of R and D in the loss function based on quality fluctuation, thereby improving rate‑distortion performance. Extensive experiments demonstrate that DCVC‑DT achieves up to 6.21% bitrate savings over the baseline DCVC‑DC, while significantly enhancing generalization to unseen testing data and alleviating error propagation. Our code is available at https://github.com/SunnyMass/DCVC‑DT.
Authors:Huan Wang, Jun Shen, Haoran Li, Zhenyu Yang, Jun Yan, Ousman Manjang, Yanlong Zhai, Di Wu, Guansong Pang
Abstract:
Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prototype‑based FL is in the spotlight, since shared global prototypes offer semantic anchors for aligning client‑specific local prototypes. However, existing methods update global prototypes at the prototype‑level via averaging local prototypes or refining global anchors, which often leads to semantic drift across clients and subsequently yields a misaligned global signal. To alleviate this issue, we introduce hyper‑prototypes, defined by a set of learnable global class‑wise prototypes to preserve underlying semantic knowledge across clients. The hyper‑prototypes are optimized via gradient matching to align with class‑relevant characteristics distilled directly from clients' real samples, rather than prototype‑level descriptors. We further propose FedHPro, a Federated Hyper‑Prototype Learning framework, to leverage hyper‑prototypes to promote inter‑class separability via mutual‑contrastive learning with client‑specific margin, while encouraging intra‑class uniformity through a consistency penalty. Comprehensive experiments under diverse heterogeneous scenarios confirm that 1) hyper‑prototypes produce a more semantically consistent global signal, and 2) FedHPro achieves state‑of‑the‑art performance on several benchmark datasets. Code is available at \hrefhttps://github.com/mala‑lab/FedHProhttps://github.com/mala‑lab/FedHPro.
Authors:Lilin Zhang, Yimo Guo, Yue Li, Jiancheng Shi, Xianggen Liu
Abstract:
Deep neural networks are highly vulnerable to adversarial examples, i.e.,small perturbations that can significantly degrade model performance. While adversarial training has become the primary defense strategy, most studies focus on balanced datasets, overlooking the challenges posed by real‑world long‑tail data. Motivated by the fact that perturbations in adversarial examples inherently alter the training distribution, we theoretically investigate their impact. We first revisit adversarial training for long‑tail data and identify two key limitations: (i) a skewed training objective caused by class imbalance, and (ii) unstable evolution of adversarial distributions. Furthermore, we show that perturbations can simultaneously address both adversarial vulnerability and class imbalance. Based on these insights, we propose RobustLT, a plug‑and‑play framework that adaptively adjusts perturbations during adversarial training. Extensive experiments demonstrate that RobustLT consistently enhances adversarial robustness and class‑balance on long‑tailed datasets. The code is available at \hrefhttps://github.com/zhang‑lilin/RobustLThttps://github.com/zhang‑lilin/RobustLT.
Authors:Junhyuk Jeon, Seokhyeon Hong, Junyong Noh
Abstract:
Text‑driven motion diffusion models are capable of generating realistic human motions, but text alone often struggles to express fine‑level nuances of motion, commonly referred to as style. Recent approaches have tackled this challenge by attaching a style injection mechanism to a pretrained text‑driven diffusion model. Existing stylization methods, however, either require style‑specific fine‑tuning of existing models or rely on heavy ControlNet‑based architectures, limiting efficiency and generalization to unseen styles. We propose a lightweight style conditioning framework that dynamically modulates a pretrained diffusion model through hypernetwork‑generated LoRA parameters. A style reference motion is encoded into a global style embedding, which is mapped by a hypernetwork to low‑rank updates applied at each denoising step of the diffusion model. By structuring the style latent space with a supervised contrastive loss, our framework reliably captures diverse stylistic attributes, improves generalization to unseen styles, and supports optimization‑based guidance without requiring predefined style categories. Experiments on the HumanML3D and 100STYLE datasets show state‑of‑the‑art stylization results, while achieving improved stylization for unseen styles.
Authors:Yunheng Wang, Yuetong Fang, Taowen Wang, Lusong Li, Kun Liu, Junzhe Xu, Zizhao Yuan, Yixiao Feng, Jiaxi Zhang, Wei Lu, Zecui Zeng, Renjing Xu
Abstract:
Vision‑and‑Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning from simulation to real‑world deployment, primarily due to perceptual instability (e.g., lighting variations and motion blur) and under‑specified instructions. While existing methods attempt to bridge this gap by scaling up model size and training data, we argue that the bottleneck lies in the lack of robust spatial grounding and cross‑domain priors. In this paper, we propose StereoNav, a robust Vision‑Language‑Action framework designed to enhance real‑world navigation consistency. To address the inherent gap between synthetic training and physical execution, we introduce Target‑Location Priors as a persistent bridge. These priors provide stable visual guidance that remains invariant across domains, effectively grounding the agent even when instructions are vague. Furthermore, to mitigate visual disturbances like motion blur and illumination shifts, StereoNav leverages stereo vision to construct a unified representation of semantics and geometry, enabling precise action prediction through enhanced depth awareness. Extensive experiments on R2R‑CE and RxR‑CE demonstrate that StereoNav achieves state‑of‑the‑art egocentric RGB performance, with SR and SPL scores of 81.1% and 68.3%, and 67.5% and 52.0%, respectively, while using significantly fewer parameters and less training data than prior scaling‑based approaches. More importantly, real‑world robotic deployments confirm that StereoNav substantially improves navigation reliability in complex, unstructured environments. Project page: https://yunheng‑wang.github.io/stereonav‑public.github.io.
Authors:Kangye Ji, Yuan Meng, Jianbo Zhou, Ye Li, Chen Tang, Zhi Wang
Abstract:
Action diffusion excels at high‑fidelity action generation but incurs heavy computational costs owing to its iterative denoising nature. Despite current technologies showing promise in accelerating diffusion transformers by reusing the cached features, they struggle to adapt to policy dynamics arising from diverse perceptions and multi‑round rollout iterations in open environments. We propose test‑time sparsity to tackle this challenge, which aims to accelerate action diffusion by dynamically predicting prunable residual computations for each model forward at test time. However, two bottlenecks remain in this paradigm: 1) repetitive conditional encoding and pruning offset most potential speed gains, and 2) the features cached from previous denoising timesteps cannot constrain large pruning errors under aggressive sparsity. To address the first bottleneck, we design a highly parallelized inference pipeline that minimizes the non‑decoder delay to milliseconds. Specifically, we first design a lightweight pruner that shares the encoder with the diffusion transformer. Then, we decouple the encoding and pruning from the autoregressive denoising loop by processing all denoising timesteps in parallel, and overlap the pruner with the decoder forward inference through asynchronism. To overcome the second bottleneck, we introduce an omnidirectional reusing strategy, which achieves 95% sparsity by selectively reusing features cached from the current forward, previous denoising timesteps, and earlier rollout iterations. To learn the rollout‑level reusing strategies, we sample a few action trajectories to supervise the sparsified diffusion step by step. Extensive experiments demonstrate that our method reduces FLOPs by 92% and accelerates action generation by 5x, achieving lossless performance with an inference frequency of 47.5 Hz. Our code is available at https://github.com/ky‑ji/Test‑time‑Sparsity.
Authors:G. Dofri Vidarsson, Liying Lu, Sabine Süsstrunk
Abstract:
Illuminant estimation aims to infer scene illumination from image measurements despite intrinsic ambiguities between surface reflectance and lighting. Most existing methods operate on trichromatic RGB images and are therefore fundamentally limited by the restricted spectral information available. Hyperspectral imaging provides a much richer representation of scene radiance and has the potential to alleviate these ambiguities. However, its high dimensionality poses computational and statistical challenges. In this work, we systematically study the effect of spectral dimensionality and representation choice on illuminant estimation performance using hyperspectral data. We adopt the practical and effective Color‑by‑Correlation (CbC) framework as the estimation backbone and analyze its behavior under different spectral dimensionality reduction strategies. Our results offer practical insights into how hyperspectral information can be efficiently exploited for illuminant estimation and identify conditions under which compact spectral representations outperform conventional RGB‑based approaches. The code is available at https://github.com/IVRL/Reduced‑Spectral‑Color‑Constancy.
Authors:Abdelrahman Eldesokey, Merey Ramazanova, Ahmad Sait, Ansar Khangeldin, Karen Sanchez, Tong Zhang, Bernard Ghanem
Abstract:
Text‑to‑image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing evaluation practices typically apply uniform annotation mechanisms, such as Likert‑scale or binary question answering (BQA), across heterogeneous evaluation skills, despite fundamental differences in their nature. In this work, we revisit T2I evaluation through the lens of skill‑aligned annotation, where annotation strategies reflect the underlying characteristics of each evaluation skill. We systematically compare skill‑aligned annotation against uniform baselines and show that it produces more consistent evaluation signals, with higher inter‑annotator agreement and improved stability across models. Finally, we present an automated pipeline that instantiates the proposed evaluation protocol, enabling scalable and fine‑grained evaluation with spatially grounded feedback. Our work highlights that improving the foundations of image evaluation can increase reliability and efficiency without simply scaling annotation effort. We hope this motivates further research on refining evaluation protocols as a central component of reliable model assessment.
Authors:Hongli Liu, Yu Wang, Shengjie Zhao
Abstract:
Few‑shot action recognition (FSAR) requires models to generalize to novel action categories from only a handful of annotated samples. Despite progress with vision‑language models, existing approaches still suffer from semantic‑temporal misalignment, where static textual prompts fail to capture decisive visual cues that appear sparsely across sequences, and from inadequate modeling of multi‑scale temporal dynamics, as short‑term discriminative cues and long‑range dependencies are often either oversmoothed or fragmented. To address these challenges, we propose Semantic Temporal Adaptive Representation Learning (STAR), a unified framework, consisting of a semantic‑alignment component and a temporal‑aware component, effectively bridging the semantic and temporal gaps and transferring the sequence modeling capability of Mamba into the FSAR. The semantic alignment module introduces a Temporal Semantic Attention (TSA) mechanism, which performs frame‑level cross‑modal alignment with textual cues, ensuring fine‑grained semantic‑temporal consistency. The temporal‑aware module incorporates a Semantic Temporal Prototype Refiner (STPR) that integrates semantic‑guided Mamba blocks with multi‑frequency temporal sampling and bidirectional state‑space refinement, yielding semantically aligned prototypes with enhanced discriminative fidelity and temporal consistency. Furthermore, temporally dependent class descriptors derived from large language models (LLMs) provide long‑range semantic guidance. Extensive experiments on five FSAR benchmarks demonstrate the consistent superiority of STAR over state‑of‑the‑art methods. For instance, STAR achieves up to 8.1% and 6.7% gains on the SSv2‑Full and SSv2‑Small datasets under the 1‑shot setting, and 7.3% on HMDB51, validating its effectiveness under limited supervision. The code is available at https://github.com/HongliLiu1/STAR‑main.
Authors:Geng Li, Yuxin Peng
Abstract:
Fine‑grained recognition in everyday life is often not a closed‑book classification problem: when encountering unfamiliar objects, humans actively search, compare visual details, and verify evidence before deciding. Existing benchmarks primarily evaluate visually recognition, leaving this active external knowledge acquisition ability underexplored. We study fine‑grained knowledge acquisition, where a system must seek, verify, and use external evidence to answer open‑ended fine‑grained recognition questions. We introduce FIKA‑Bench, a leakage‑aware and evidence‑grounded collection of 311 public‑source and real‑life instances. To ensure high quality, every example is filtered against frontier closed‑book models to remove memorized cases and audited to eliminate image‑answer leakage, retaining only samples supported by verified evidence. Our evaluation of latest Large Multimodal Models (LMMs) and agents reveals that the task remains a formidable challenge: the best system reaches only 25.1% accuracy, with no model exceeding 30%. Crucially, we find that merely equipping models with tools is insufficient to bridge this gap; agent failures are predominantly driven by wrong entity retrieval and poor visual judgement. These results show that reliable knowledge acquisition needs better agent designs that focus on fine‑grained recognition.
Authors:Zheng Chen, Ruofan Yang, Jin Han, Dehua Song, Zichen Zou, Chunming He, Yong Guo, Yulun Zhang
Abstract:
Diffusion‑based models have shown strong performance in video super‑resolution (VSR) and video frame interpolation (VFI). However, their role in the coupled space‑time video super‑resolution (STVSR) setting remains limited. Existing diffusion‑based STVSR approaches suffer from two issues: (1) low inference efficiency and (2) insufficient utilization of spatiotemporal information. These limitations impede deployment. To address these issues, we introduce DiffST, an efficient spatiotemporal‑aware video diffusion framework for real‑world STVSR. To improve efficiency, we adapt a pre‑trained diffusion model for one‑step sampling and process the entire video directly rather than operating on individual frames. Furthermore, to enhance spatiotemporal information utilization, we introduce cross‑frame context aggregation (CFCA) and video representation guidance (VRG). The CFCA module aggregates information across multiple keyframes to produce intermediate frames. The VRG module extracts video‑level global features to guide the diffusion process. Extensive experiments show that DiffST obtains leading results on real‑world STVSR tasks. It also maintains high inference efficiency, running about 17× faster than previous diffusion‑based STVSR methods. Code is available at: https://github.com/zhengchen1999/DiffST.
Authors:Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu, Zhen Wang, Donglian Qi, Yunfeng Yan, Xi Chen
Abstract:
Multimodal large laboratory models (MLLMs) still struggle with spatial understanding under the dominant perspective‑image paradigm, which inherits the narrow field of view of human‑like perception. For navigation, robotic search, and 3D scene understanding, 360‑degree panoramic sensing offers a form of supersensing by capturing the entire surrounding environment at once. However, existing MLLM pipelines typically decompose panoramas into multiple perspective views, leaving the spherical structure of equirectangular projection (ERP) largely implicit. In this paper, we study pano‑native understanding, which requires an MLLM to reason over an ERP panorama as a continuous, observer‑centered space. To this end, we first define the key abilities for pano‑native understanding, including semantic anchoring, spherical localization, reference‑frame transformation, and depth‑aware 3D spatial reasoning. We then build a large‑scale metadata construction pipeline that converts mixed‑source ERP panoramas into geometry‑aware, language‑grounded, and depth‑aware supervision, and instantiate these signals as capability‑aligned instruction tuning data. On the model side, we introduce PanoWorld with Spherical Spatial Cross‑Attention, which injects spherical geometry into the visual stream. We further construct PanoSpace‑Bench, a diagnostic benchmark for evaluating ERP‑native spatial reasoning. Experiments show that PanoWorld substantially outperforms both proprietary and open‑source baselines on PanoSpace‑Bench, H Bench, and R2R‑CE Val‑Unseen benchmarks. These results demonstrate that robust panoramic reasoning requires dedicated pano‑native supervision and geometry‑aware model adaptation. All source code and proposed data will be publicly released.
Authors:Jiahao Chen, Zihui Zhang, Yafei Yang, Jinxi Li, Shenxing Wei, Zhixuan Sun, Bo Yang
Abstract:
We introduce EvObj for unsupervised 3D instance segmentation that bridges the geometric domain gap between synthetic pretraining data and real‑world point clouds. Current methods suffer from structural discrepancies when transferring object priors from synthetic datasets (e.g., ShapeNet) to real scans (e.g., ScanNet), particularly due to morphological variations and occlusion artifacts. To address this, EvObj integrates two innovative modules: (1) An object discerning module that dynamically refines object candidates, enabling continuous adaptation of object priors to target domains; and (2) An object completion module that reconstructs partial geometries after discovering objects. We conduct extensive experiments on both real‑world and synthetic datasets, demonstrating superior 3D object segmentation performance over all baselines while achieving state‑of‑the‑art results.
Authors:David Iagaru, Nina M. Gottschling, Anders C. Hansen, Josselin Garnier
Abstract:
Artificial intelligence (AI) has transformed imaging inverse problems, from medical diagnostics to Earth observation. Yet deep neural networks can produce hallucinations, realistic‑looking but incorrect details, undermining their reliability, especially when ground truth data is unavailable. We develop a theoretical framework showing that such hallucinations are not merely artifacts of particular models, but can arise from the ill‑posed nature of the inverse problem itself. We derive necessary and sufficient conditions for hallucinations, together with computable bounds on their magnitude that depend only on the forward model. Building on this theory, we introduce algorithms to: (1) estimate the minimum hallucination magnitude achievable by any reconstruction model for a given input; (2) assess the faithfulness of reconstructed details by a given reconstruction model. Experiments across three imaging tasks demonstrate that our approach applies broadly, including to modern generative models, and provides a principled way to quantify and evaluate AI hallucinations.
Authors:Sangin Lee, Seokjun Kwon, Jeongmin Shin, Namil Kim, Yukyung Choi
Abstract:
General object detection (OD) struggles to detect objects in the target domain that differ from the training distribution. To address this, recent studies demonstrate that training from multiple source domains and explicitly processing them separately for multi‑source domain adaptation (MSDA) outperforms blending them for unsupervised domain adaptation (UDA). However, existing MSDA methods learn domain‑agnostic features from domain‑specific RGB images while preserving domain‑specific information from the domain‑agnostic feature map. To address this, we propose MS‑DePro: Multi‑Source Detector with Depth and Prompt, composed of (1) depth‑guided localization and (2) multi‑modal guided prompt learning. We leverage domain‑agnostic input modalities, namely depth maps and text, to encode domain‑agnostic characteristics. Specifically, we utilize depth maps to generate domain‑agnostic region proposals for localization and integrate multi‑modal features to align learnable text embeddings for classification. MS‑DePro achieves state‑of‑the‑art performance on MSDA benchmarks, and comprehensive ablations demonstrate the effectiveness of our contributions. Our code is available on https://github.com/sejong‑rcv/Multi‑Modal‑Guided‑Multi‑Source‑Domain‑Adaptation‑for‑Object‑Detection.
Authors:Jiayu Chen, Junbei Tang, Wenbiao Zhao, Maoliang Li, Jiayi Luo, Zihao Zheng, Jiawei Yang, Guojie Luo, Xiang Chen
Abstract:
Autoregressive video generation enables streaming and open‑ended long video synthesis, but still suffers from long‑term degradation caused by accumulated errors. Existing KVCache strategies usually apply unified historical‑frame retention, implicitly assuming homogeneous historical dependencies across attention heads. We revisit historical‑frame attention and reveal three distinct head types: Anchor Heads require broad long‑range context, Wave Heads exhibit periodic temporal dependencies, and Veil Heads focus on initial and adjacent frames. Based on this finding, we propose Pyramid Forcing, a head‑aware pyramidal KVCache framework that identifies head types offline, assigns behavior‑specific cache policies, and supports heterogeneous cache lengths via efficient ragged‑cache attention. Experiments on Self Forcing and Causal Forcing show that Pyramid Forcing consistently improves long‑horizon generation quality on VBench‑Long, increasing the 60‑second Self Forcing score from 77.87 to 81.21 while enhancing motion dynamics, visual fidelity, and semantic consistency. Project: https://if‑lab‑pku.github.io/Pyramid‑Forcing/.
Authors:Guangqian Yang, Tong Ding, Wenlong Hou, Yue Xun, Ye Du, Qian Niu, Shujun Wang
Abstract:
Clinical diagnostic workups typically follow a modality escalation pathway: after initial clinical evaluation, clinicians begin with routine structural imaging (e.g., MRI), selectively add sequences such as FLAIR or T2 to refine the differential, and reserve molecular imaging (e.g., amyloid‑PET) for cases that remain uncertain after standard evaluation. Consequently, patients are observed with heterogeneous and often incomplete modality subsets. However, most current AI models assume fixed data modalities as the model inputs. In this paper, we present BrainAnytime, a unified pretraining framework pretrained on 34,899 3D brain scans from five datasets that support brain image analysis under arbitrary modality availability spanning multi‑sequence MRI and amyloid‑PET. A single model accepts whatever imaging is available, from a lone T1 scan to a full multimodal workup. Pretraining learns structural‑molecular correspondences between MRI and PET via cross‑modal distillation (RCMD) and prioritizes disease‑vulnerable anatomy via atlas‑guided curriculum masking (PACM), all within a shared 3D masked autoencoder (Multi‑MAE3D). Across four downstream tasks and five clinically motivated modality settings, BrainAnytime largely outperforms modality‑specific models, missing‑modality baselines, and large‑scale brain MRI pretrained foundation models on most modality settings. Notably, it surpasses the strongest missing‑modality baselines with relative improvements of 6.2% and 7.0% in average accuracy on CN vs. AD and CN vs. MCI classification, respectively. Code is available at https://github.com/SDH‑Lab/BrainAnytime.
Authors:Ziqi Wen, Parsa Madinei, Miguel P. Eckstein
Abstract:
Evaluating whether large vision‑language models (VLMs) align with human perception for high‑level semantic scene comprehension remains a challenge. Traditional white‑box interpretability methods are inapplicable to closed‑source architectures and passive metrics fail to isolate causal features. We introduce Counterfactual Semantic Saliency (CSS). This black‑box, model‑agnostic framework quantifies the importance of objects by measuring the semantic shift induced by their causal ablation from a scene. To evaluate AI‑human semantic alignment, we tested prominent VLMs against a human psychophysics baseline comprising 16,289 valid responses across 307 complex natural scenes and 1,306 high‑fidelity counterfactual variants. Our analysis reveals a pervasive scene comprehension gap: models exhibit an overreliance (relative to humans) on large objects (size bias), objects at the center of the image (center bias), and high saliency objects. In contrast, models rely less on people in the scenes than our human participants to describe the images. A model's size bias is a primary driver explaining variations in model‑human semantic divergence. Code and data will be available at https://github.com/starsky77/Counterfactual‑Semantic‑Saliency.
Authors:Zihang Xu, Xiaoyang Liu, Zheng Chen, Yulun Zhang, Xiaokang Yang
Abstract:
Text image super‑resolution (Text‑SR) requires more than visually plausible detail synthesis: slight errors in stroke topology may alter character identity and break readability. Existing methods improve text fidelity with stronger recognition‑based or generative priors, yet they still face two unresolved challenges under severe degradation: the text condition extracted from low‑quality inputs can itself be unreliable, and a plausible global prior does not fully determine fine‑grained stroke boundaries. We present PRISM, a single‑step diffusion‑based Text‑SR framework that addresses these two challenges through Flow‑Matching Prior Rectification (FMPR) and a Structure‑guided Uncertainty‑aware Residual Encoder (SURE). FMPR constructs a privileged training‑time prior from paired low‑quality/high‑quality latents and learns a flow matching that transports degraded embeddings toward this restoration‑oriented prior space, yielding more accurate and reliable global text guidance. SURE further predicts uncertainty‑aware structural residuals to selectively absorb reliable local boundary evidence while suppressing ambiguous stroke cues. Together, these components enable explicit global prior rectification and local structure refinement within a single diffusion restoration pass. Experiments on both synthetic and real‑world benchmarks show that PRISM achieves state‑of‑the‑art performance with millisecond‑level inference. Our dataset and code will be available at https://github.com/faithxuz/PRISM.
Authors:Hansheng Chen, Jan Ackermann, Minseo Kim, Gordon Wetzstein, Leonidas Guibas
Abstract:
Flow‑based generation in high‑dimensional spaces is difficult because velocity prediction requires modeling high‑dimensional noise, even when data has strong low‑rank structure. We present Asymmetric Flow Modeling (AsymFlow), a rank‑asymmetric velocity parameterization that restricts noise prediction to a low‑rank subspace while keeping data prediction full‑dimensional. From this asymmetric prediction, AsymFlow analytically recovers the full‑dimensional velocity without changing the network architecture or training/sampling procedures. On ImageNet 256×256, AsymFlow achieves a leading 1.57 FID, outperforming prior DiT/JiT‑like pixel diffusion models by a large margin. AsymFlow also provides the first‑ever route for finetuning pretrained latent flow models into pixel‑space models: aligning the low‑rank pixel subspace to the latent space gives a seamless initialization that preserves the latent model's high‑level semantics and structure, so finetuning mainly improves low‑level mismatches rather than relearning pixel generation. We show that the pixel AsymFlow model finetuned from FLUX.2 klein 9B establishes a new state of the art for pixel‑space text‑to‑image generation, beating its latent base on HPSv3, DPG‑Bench, and GenEval while qualitatively showing substantially improved visual realism.
Authors:Feijiang Li, Zhenxiong Li, Jieting Wang, Zizheng Jiu, Saixiong Liu, Liang Du
Abstract:
Image clustering aims to partition unlabeled image datasets into distinct groups. A core aspect of this task is constructing and leveraging prior knowledge to guide the clustering process. Recent approaches introduce semantic descriptions as prior information, most of which typically relying on matching‑based techniques with predefined vocabularies. However, the limited matching space restricts their adaptability to downstream clustering tasks. Moreover, these methods primarily focus on reducing bias to improve performance, frequently overlooking the importance of variance reduction. To address these limitations, we propose GSEC (Image Clustering based on Generative Semantic Guidance and Bi‑Layer Ensemble), a framework designed to reduce bias through generative semantic guidance and mitigate variance via ensemble learning. Our method employs Multimodal Large Language Models to generate semantic descriptions and derive image embeddings via weighted averaging. Additionally, a bi‑layer ensemble strategy integrates cross‑modal information through BatchEnsemble in the inner layer and aligns outputs via an alignment mechanism in the outer layer. Comparative experiments demonstrate that GSEC outperforms 18 state‑of‑the‑art methods across six benchmark datasets, while further analysis confirms its effectiveness in simultaneously reducing both bias and variance. The code is available at https://github.com/2017LI/GSEC.git.
Authors:Hanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo, Tianyu He, Xin Jin, Zhibo Chen
Abstract:
Recent developments in generative models and large‑scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains including spatial intelligence, embodied intelligence, and autonomous driving. While achieving remarkable progress, existing approaches to 3D world generation typically prioritize appearance prediction with limited modeling of the underlying geometry, leading to issues such as unreliable scene structure estimation and degraded cross‑view consistency. To address these limitations, motivated by the coarse‑to‑fine nature of human visual perception, we propose GTA, a novel image‑to‑3D world generation method following a Geometry‑Then‑Appearance paradigm. Specifically, given a single input image, to improve the structural fidelity of synthesized 3D scenes, GTA adopts a two‑stage framework with two dedicated video diffusion models, which first generate coarse geometric structure from novel viewpoints and then synthesize fine‑grained appearance conditioned on the predicted geometry. To further enhance cross‑view appearance consistency, we introduce a random latent shuffle strategy during the training process, along with a test‑time scaling scheme that improves perceptual quality without compromising quantitative performance. Extensive experiments have demonstrated that our proposed method consistently outperforms existing approaches in terms of fidelity, visual quality, and geometric accuracy. Moreover, GTA is shown to be effective as a general enhancement module that further improves the generation quality of existing image‑to‑3D world pipelines, as well as supporting multiple downstream applications and exhibiting favorable data efficiency during model training, highlighting its versatility and broad applicability. Project page: https://hanxinzhu‑lab.github.io/GTA/.
Authors:Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, Jiahao Kong, Weijun Zeng, Jutao Xiao, Jie Yang, Wentao Zhang, Bin Wang, Conghui He
Abstract:
Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc‑VQA evaluations score only the final answer and leave the supporting evidence unchecked. This answer‑only approach masks a critical failure mode: a model can land on the correct answer while grounding it in the wrong passage ‑‑ a critical risk in high‑stakes domains like law, finance, and medicine, where every conclusion must be traceable to a specific source region. To address this, we introduce CiteVQA, a benchmark that requires models to return element‑level bounding‑box citations alongside each answer, evaluating both jointly. CiteVQA comprises 1,897 questions across 711 PDFs spanning seven domains and two languages, averaging 40.6 pages per document. To ensure fidelity and scalability, the ground‑truth citations are generated by an automated pipeline‑which identifies crucial evidence via masking ablation‑and are subsequently validated through expert review. At the core of our evaluation is Strict Attributed Accuracy (SAA), which credits a prediction only when the answer and the cited region are both correct. Auditing 20 MLLMs reveals a pervasive Attribution Hallucination: models frequently produce the right answer while citing the wrong region. The strongest system (Gemini‑3.1‑Pro‑Preview) achieves an SAA of only 76.0, and the strongest open‑source MLLM reaches just 22.5. Ultimately, towards trustworthy document intelligence, CiteVQA exposes a reliability gap that answer‑only evaluations overlook, providing the instrumentation needed to close it. Our repository is available at https://github.com/opendatalab/CiteVQA.
Authors:Kaixiang Zhao, Tianrun Yu, Aoxu Zhang, Junhao Su, Porter Jenkins, Amanda Hughes
Abstract:
The proliferation of sophisticated image editing tools and generative artificial intelligence models has made verifying the authenticity of digital images increasingly challenging, with important implications for journalism, forensic analysis, and public trust. Although numerous forensic algorithms, ranging from handcrafted methods to deep learning‑based detectors, have been developed for manipulation detection, individual methods often suffer from limited robustness, fragmented evidence, or weak generalization across manipulation types and image conditions. To address these limitations, we present FRAME, a method for Forensic Routing and Adaptive Multi‑path Evidence fusion for image manipulation detection. FRAME organizes diverse forensic algorithms into a multi‑path analysis space, adaptively selects informative forensic paths for each input image, and fuses complementary evidence to improve detection and localization performance. By moving beyond single‑method analysis and fixed fusion strategies, FRAME provides a more robust and flexible approach to image forensic reasoning while preserving interpretable forensic cues from multiple evidence sources. Experimental results demonstrate the effectiveness of FRAME across diverse manipulation scenarios. Code is available at \hrefhttps://github.com/kzhao5/FRAMEhttps://github.com/kzhao5/FRAME.
Authors:Andreas Maier, Jeta Sopa, Gozde Gul Sahin, Paula Perez-Toro, Siming Bayer
Abstract:
Wu et al. (2026) showed that most frontier large language models (LLMs) recommend a sponsored, roughly twice‑as‑expensive flight when their system prompt contains a soft sponsorship cue. We reproduce their evaluation on ten open‑weight chat models plus the two of their twenty‑three models that are still reachable today (gpt‑3.5‑turbo, gpt‑4o). All reported rates in this paper are produced under the same judge the original paper used (gpt‑4o); we additionally store every label under an open‑weight (gpt‑oss‑120b) and a smaller proprietary (gpt‑4o‑mini) judge for an ablation. Three findings emerge. First, a prose description of an LLM evaluation pipeline is not, on its own, sufficient for accurate reproduction: we surfaced three silent implementation failures that each shifted a reported rate by tens of percentage points. Second, the central claims do generalise ‑ the gpt‑3.5‑turbo logistic‑regression intercept of alpha = 0.81 is within four points of the original alpha = 0.86, and 200 of 200 trials on gpt‑3.5‑turbo and gpt‑4o promote a payday lender to a financially distressed user. Third, a thirty‑token user prompt that asks the assistant for a neutral comparison table first cuts sponsored recommendation from 46.9% to 1.0% averaged across our ten open‑source models, and from 53.0% to 0% averaged across the two OpenAI models. AI literacy and price‑comparison portals are likely market‑level mitigations; the harmful‑product cell is bounded by neither. Raw data, labels and analysis scripts are at https://github.com/akmaier/Paper‑LLM‑Ads .
Authors:Paul Hoareau, Kuan Yi Wang, Brandon Bujak, Roy Sun, Govind Nair, Irene Cortese, Charidimos Tsagkas, Daniel Reich, Julien Cohen-Adad
Abstract:
INTRODUCTION | Fully supervised 3D segmentation of high‑resolution ex vivo MRI is limited by the prohibitive cost of volumetric annotation, forcing reliance on sparse 2D slices. Weakly supervised Sparse‑to‑Dense frameworks bridge this gap, but guidelines remain ambiguous regarding human‑centric visual enhancements and transferring optimization strategies across dimensions. We analyze divergent regularization needs for multi‑class segmentation of high‑resolution ex vivo spinal cord MRI.
METHODS | We used 9.4T MRI of multiple sclerosis spinal cords (>104,000 slices) with sparse annotations (428 slices). A 2D Teacher trained on sparse slices generated dense pseudo‑labels to train a 3D Student. We systematically evaluated the impact of human‑centric preprocessing, spatial augmentation, and soft‑label regularization on both architectures.
RESULTS | We identified a critical divergence in training dynamics. The 2D Teacher required strong spatial augmentation and soft‑labeling to overcome data scarcity, improving White Matter Lesion Dice scores by >11 points. However, propagating these techniques to the 3D Student degraded its performance. Furthermore, human‑centric preprocessing (e.g., CLAHE) disrupted global statistical cues, dropping Gray Matter Lesion Dice scores by ~25 points.
DISCUSSION | Our study highlights a perception divergence (human‑centric contrast enhancement harms machine models) and a regularization conflict across dimensions. 3D architectures trained on dense pseudo‑labels exhibit fundamentally different optimization landscapes than 2D counterparts and require distinct, conservative regularization. Code and models: https://github.com/ivadomed/model_seg_sc‑gm‑lesion_human_ms_exvivo_t2star.
Authors:Yichen Feng, Yuetai Li, Chunjiang Liu, Yuanyuan Chen, Fengqing Jiang, Yue Huang, Hang Hua, Zhengqing Yuan, Kaiyuan Zheng, Luyao Niu, Bhaskar Ramasubramanian, Basel Alomair, Xiangliang Zhang, Misha Sra, Zichen Chen, Radha Poovendran, Zhangchen Xu
Abstract:
Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing solutions reduce this judgment to predicting a scalar score for a single image. We first ask whether such scores faithfully capture comparative preference: in a controlled study with eight expert annotators, score‑derived rankings align poorly with the same annotators' direct comparisons, while direct ranking yields substantially higher inter‑annotator agreement on best‑ and worst‑image labels. Motivated by this finding, we introduce the Visual Aesthetic Benchmark (VAB), which casts aesthetic evaluation as comparative selection over candidate sets with matched subject matter. VAB contains 400 tasks and 1,195 images across fine art, photography, and illustration, with labels derived from the consensus of 10 independent expert judges per task. Evaluating 20 frontier MLLMs and six dedicated visual‑quality reward models, we find that the strongest system identifies both the best and the worst image correctly across three random permutations of the candidate order in only 26.5% of tasks, far below the 68.9% achieved by human experts. Fine‑tuning a 35B‑parameter model on 2,000 expert examples brings its accuracy close to that of a 397B‑parameter open‑weight model, suggesting that the comparative signal in VAB is transferable. Together, these results expose a clear and measurable gap between current multimodal models and expert aesthetic judgment, and VAB provides the first set‑based, expert‑grounded testbed on which that gap can be tracked and closed.
Authors:Ahmed Heakl, Youssef Mohamed, Abdullah Sohail, Rania Elbadry, Ahmed Nassar, Peter W. J. Staar, Fahad Shahbaz Khan, Imran Razzak, Salman Khan
Abstract:
Multilingual document understanding remains limited for low‑resource languages due to scarce training data and model‑based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs high‑fidelity OCR datasets and benchmarks covering 82 languages and 9 evaluation tasks. Our dual pipelines, differential rendering of native DOCX documents and synthetic LaTeX‑based generation for right‑to‑left scripts produce precise structural annotations in a unified DocTag format encoding layout, text, and component types, without learned models for core annotation. Evaluating 16 state‑of‑the‑art models reveals persistent gaps in low‑resource scripts. We show that Direct Preference Optimization (DPO) using rendering‑derived ground truth as positive signal achieves stable multilingual adaptation, improving both in‑domain (+1.9%) and out‑of‑domain (+1.8%) accuracy without measurable base‑language degradation, where supervised fine‑tuning degrades out‑of‑domain performance by up to 21%. Our best variant, DocAtlas‑DeepSeek, improves +1.7% over the strongest baseline. Code is available at https://github.com/ahmedheakl/DocAtlas .
Authors:Mohamed Ahmed Mohamed, Xiaowei Huang
Abstract:
Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real‑world foggy data remains a significant bottleneck. In this paper, we propose Clear2Fog (C2F), an end‑to‑end, physics‑based pipeline that simulates fog on clear‑weather datasets while ensuring sensor‑level consistency across camera and LiDAR. By using monocular depth estimation and a novel atmospheric light estimation method, C2F overcomes structural artifacts and chromatic biases common in existing techniques. A human perceptual study confirms C2F's physical realism, with the generated images being preferred 92.95% of the time over an established method. Utilising a training set of 270,000 images from the Waymo Open Dataset, we conduct an extensive data efficiency study to investigate how environmental diversity influences model robustness. Our findings reveal that models trained on mixed‑density fog datasets at 75% scale outperform those trained on fixed‑density datasets at 100% scale. Furthermore, we investigate the sim‑to‑real transfer by fine‑tuning pre‑trained models on real‑world foggy data. We demonstrate that a tenfold increase over the default fine‑tuning learning rate successfully overcomes negative transfer from synthetic biases, resulting in a 1.67 mAP improvement over real‑only baselines. The C2F pipeline provides a scalable framework for enhancing the reliability of autonomous systems in adverse weather and demonstrates the potential of diverse synthetic datasets for efficient model training.
Authors:Jisu Nam, Jahyeok Koo, Soowon Son, Jaewoo Jung, Honggyu An, Junhwa Hur, Seungryong Kim
Abstract:
Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per‑frame geometry, recovering object motion in this geometry remains challenging and benefits from strong motion priors learned from real‑world videos. Existing 3D trackers either follow iterative paradigms trained from scratch on synthetic data or fine‑tune 3D reconstruction models learned from static multi‑view images, both lacking real‑world motion priors. Pre‑trained video diffusion transformers (video DiTs) offer rich spatio‑temporal priors from internet‑scale videos, making them a promising foundation for 3D tracking. However, their frame‑anchored formulation, which generates each frame's content, is fundamentally mismatched with reference‑anchored dense 3D tracking, which must follow the same physical points from a reference frame across time. We present TrackCraft3R, the first method to repurpose a video DiT as a feed‑forward dense 3D tracker. Given a monocular video and its frame‑anchored reconstruction pointmap, TrackCraft3R predicts a reference‑anchored tracking pointmap that follows every pixel of the first frame across time in a single forward pass, along with its visibility. We achieve this through two designs: (i) a dual‑latent representation that uses per‑frame geometry latents and reference‑anchored track latents as dense queries, and (ii) temporal RoPE alignment, which specifies the target timestamp of each track latent. Together, these designs convert the per‑frame generative paradigm of video DiTs into a reference‑anchored tracking formulation with LoRA fine‑tuning. TrackCraft3R achieves state‑of‑the‑art performance on standard sparse and dense 3D tracking benchmarks, while running 1.3x faster and using 4.6x less peak memory than the strongest prior method. We further demonstrate robustness to large motions and long videos.
Authors:Chenhao Qiu, Yechao Zhang, Xin Luo, Shien Song, Xusheng Liu
Abstract:
Long video question answering requires locating sparse, time‑scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long videos introduce long‑horizon search and verification, which often necessitates multi‑turn, agentic interaction. We show that existing LVU agents can exhibit "evidence misalignment": they produce correct answers that are not supported by the retrieved or inspected evidence. To characterize this failure, we introduce two diagnostics (temporal groundedness and semantic groundedness) and use them to reveal two pressures that amplify misalignment: prompt pressure from shared‑context saturation at inference time and reward pressure from outcome‑only optimization during training. These findings point to a structural root cause: the coupled agent paradigm conflates long‑horizon planning with answer authority. We therefore propose the decoupled planner‑inspector framework, which separates planning from answer authority and gates final answering on pixel‑level verification. Across four long‑video benchmarks, our framework improves both answer accuracy and evidence alignment, achieving 55.1% on LVBench and 62.0% on LongVideoBench while producing interpretable search trajectories. Moreover, the decoupled architecture scales consistently with increased search budgets and supports plug‑and‑play upgrades of the MLLM backbone without retraining the planner. Code and models are available at https://github.com/Echochef/VideoSEAL.
Authors:Jinyue Li, Yuzhou Yu, Jingjing Yang, Meng Fu, Yani Zhang, Shuyao He, Dianlong Ge, Xin Ning, Yannan Chu, Qiankun Li
Abstract:
The accurate classification of benign and malignant pulmonary nodules in CT scans is critical for early lung cancer screening, yet remains challenging due to the multi‑scale and heterogeneous nature of pulmonary nodules. While deep learning offers potential for auxiliary diagnosis, most existing models act as "black boxes", lacking the transparency and explainability required for trustworthy clinical integration. To address this issue, we propose M3Net, a novel 3D network for pulmonary nodule classification inspired by the hierarchical diagnostic workflow of radiologists, which integrates multi‑scale contextual information from fine‑grained structures to global anatomical relationships. Our framework constructs a progressive multi‑scale input, from fine‑grained nodule structures to local semantics and global spatial relationships. M3Net employs scale‑specific encoders and ensures cross‑scale semantic consistency through latent space projection and mutual information maximization. Extensive experiments on the public LIDC‑IDRI dataset and a self‑collected clinical dataset (USTC‑FHLN) demonstrate that our method achieves state‑of‑the‑art performance, with accuracies of 86.96% and 84.24% respectively, outperforming the best baseline by 3.26% and 2.17%. The results validate that M3Net provides a more robust and clinically relevant solution for pulmonary nodule classification. The code is available at https://github.com/jylEcho/M3‑Net.
Authors:Youssef Aboelwafa, Hicham G. Elmongui, Marwan Torki
Abstract:
Low‑light image enhancement is challenging due to complex degradations, including amplified noise, artifacts, and color distortion. While Retinex‑based deep learning methods have achieved promising results, they primarily rely on single‑modality RGB information. We propose M2Retinexformer (Multi‑Modal Retinexformer), a novel framework that extends Retinexformer by incorporating depth cues, luminance priors, and semantic features within a progressive refinement pipeline. Depth provides geometric context that is invariant to lighting variations, while luminance and semantic features offer explicit guidance on brightness distribution and scene understanding. Modalities are extracted at multiple scales and fused through cross‑attention, with adaptive gating dynamically balancing illumination‑guided self‑attention and cross‑attention based on the reliability of auxiliary cues. Evaluations on the LOL, SID, SMID, and SDSD benchmarks demonstrate overall improvements over Retinexformer and recent state‑of‑the‑art methods. Code and pretrained weights are available at https://github.com/YoussefAboelwafa/M2Retinexformer
Authors:Jiaping Lin, Fei Shen, Junzhe Li, Ping Nie, Fei Yu, Ming Li, Haizhou Li
Abstract:
Existing training‑free approaches for GUI grounding often rely on multiple inference runs, such as iterative cropping or candidate aggregation, to identify target elements. Despite this additional computation, each forward pass still independently interprets the instruction and parses the visual layout, without enabling progressive interaction among visual tokens. In this paper, we study what happens during GUI grounding in Vision‑Language Models (VLMs) and identify a previously overlooked bottleneck. We show that grounding follows a two‑stage paradigm: the prefill stage determines candidate UI elements, while the decoding stage subsequently refines the final coordinates. This asymmetry establishes prefill as the critical step, as errors in candidate selection cannot be effectively corrected during decoding. Based on this observation, we propose Re‑Prefill, a training‑free method that revisits inference by introducing an attention‑guided second prefill stage to refine target selection. Specifically, visual tokens that consistently receive high attention from the query position, i.e., the final token, across layers are extracted as a preliminary target hypothesis and appended to the input, together with the instruction hidden states, enabling the model to deeply re‑think its decision before coordinate generation. Experiments across four VLMs and five benchmarks, including ScreenSpot‑Pro, ScreenSpot‑V2, OSWorld‑G, UI‑Vision, and MMBench‑GUI, demonstrate consistent improvements without additional training, with gains of up to 4.3% on ScreenSpot‑Pro. Code will be available at https://github.com/linjiaping1/Re‑Prefill.
Authors:Miaosen Zhang, Xiaohan Zhao, Zhihong Tan, Zhou Huoshen, Yijia Fan, Yifan Yang, Kai Qiu, Bei Liu, Justin Wagle, Chenzhong Yin, Mingxi Cheng, Ji Li, Qi Dai, Chong Luo, Xu Yang, Xin Geng, Baining Guo
Abstract:
Computer‑use agents (CUAs) automate on‑screen work, as illustrated by GPT‑5.4 and Claude. Yet their reliability on complex, low‑frequency interactions is still poor, limiting user trust. Our analysis of failure cases from advanced models suggests a long‑tail pattern in GUI operations, where a relatively small fraction of complex and diverse interactions accounts for a disproportionate share of task failures. We hypothesize that this issue largely stems from the scarcity of data for complex interactions. To address this problem, we propose a new benchmark CUActSpot for evaluating models' capabilities on complex interactions across five modalities: GUI, text, table, canvas, and natural image, as well as a variety of actions (click, drag, draw, etc.), covering a broader range of interaction types than prior click‑centric benchmarks that focus mainly on GUI widgets. We also design a renderer‑based data‑synthesis pipeline: scenes are automatically generated for each modality, screenshots and element coordinates are recorded, and an LLM produces matching instructions and action traces. After training on this corpus, our Phi‑Ground‑Any‑4B outperforms open‑source models with fewer than 32B parameters. We will release our benchmark, data, code, and models at https://github.com/microsoft/Phi‑Ground.git
Authors:Haiwen Diao, Penghao Wu, Hanming Deng, Jiahao Wang, Shihao Bai, Silei Wu, Weichen Fan, Wenjie Ye, Wenwen Tong, Xiangyu Fan, Yan Li, Yubo Wang, Zhijie Cao, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Yuwei Niu, Yue Zhu, Bo Liu, Chengguang Lv, Haojia Yu, Haozhe Xie, Hongli Wang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jingcheng Ni, Junxiang Xu, Kaihuan Liang, Lianqiang Shi, Linjun Dai, Linyan Wang, Oscar Qian, Peng Gao, Pengfei Liu, Qingping Sun, Rui Shen, Ruisi Wang, Shengnan Ma, Shuang Yang, Siyi Xie, Siying Li, Tianbo Zhong, Xiangli Kong, Xuanke Shi, Yang Gao, Yongqiang Yao, Yves Wang, Zhengqi Bai, Zhengyu Lin, Zixin Yin, Wenxiu Sun, Ruihao Gong, Quan Wang, Lewei Lu, Lei Yang, Ziwei Liu, Dahua Lin
Abstract:
Recent large vision‑language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova‑U1, a native unified multimodal paradigm built upon NEO‑unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova‑U1‑8B‑MoT and SenseNova‑U1‑A3B‑MoT, built on dense (8B) and mixture‑of‑experts (30B‑A3B) understanding baselines, respectively. Designed from first principles, they rival top‑tier understanding‑only VLMs across text understanding, vision‑language perception, knowledge reasoning, agentic decision‑making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge‑intensive any‑to‑image (X2I) synthesis, complex text‑rich infographic generation, and interleaved vision‑language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre‑/post‑training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision‑language‑action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.
Authors:Christen Millerdurai, Shaoxiang Wang, Yaxu Xie, Vladislav Golyanik, Didier Stricker, Alain Pagani
Abstract:
Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head‑mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand‑centric manipulation tasks, where sensing must remain compact and unobtrusive. While monocular RGB methods have made progress, they remain constrained by depth‑scale ambiguity and struggle to generalize across the diverse optical configurations of head‑mounted devices. As a result, models typically require extensive training on device‑specific datasets, which are costly and laborious to acquire. This paper addresses these challenges by introducing EgoForce, a monocular 3D hand reconstruction framework that recovers robust, absolute 3D hand pose and its position from the user's (camera‑space) viewpoint. EgoForce operates across fisheye, perspective, and distorted wide‑FOV camera models using a single unified network. Our approach combines a differentiable forearm representation that stabilizes hand pose, a unified arm‑hand transformer that predicts both hand and forearm geometry from a single egocentric view, mitigating depth‑scale ambiguity, and a ray space closed‑form solver that enables absolute 3D pose recovery across diverse head‑mounted camera models. Experiments on three egocentric benchmarks show that EgoForce achieves state‑of‑the‑art 3D accuracy, reducing camera‑space MPJPE by up to 28% on the HOT3D dataset compared to prior methods and maintaining consistent performance across camera configurations. For more details, visit the project page at https://dfki‑av.github.io/EgoForce.
Authors:Yihao Meng, Zichen Liu, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Yue Yu, Hanlin Wang, Haobo Li, Jiapeng Zhu, Yanhong Zeng, Xing Zhu, Yujun Shen, Qifeng Chen, Huamin Qu
Abstract:
Autoregressive video generation aims at real‑time, open‑ended synthesis. Yet, cinematic storytelling is not merely the endless extension of a single scene; it requires progressing through evolving events, viewpoint shifts, and discrete shot boundaries. Existing autoregressive models often struggle in this setting. Trained primarily for short‑horizon continuation, they treat long sequences as extended single shots, inevitably suffering from motion stagnation and semantic drift during long rollouts. To bridge this gap, we introduce CausalCine, an interactive autoregressive framework that transforms multi‑shot video generation into an online directing process. CausalCine generates causally across shot changes, accepts dynamic prompts on the fly, and reuses context without regenerating previous shots. To achieve this, we first train a causal base model on native multi‑shot sequences to learn complex shot transitions prior to acceleration. We then propose Content‑Aware Memory Routing (CAMR), which dynamically retrieves historical KV entries according to attention‑based relevance scores rather than temporal proximity, preserving cross‑shot coherence under bounded active memory. Finally, we distill the causal base model into a few‑step generator for real‑time interactive generation. Extensive experiments demonstrate that CausalCine significantly outperforms autoregressive baselines and approaches the capability of bidirectional models while unlocking the streaming interactivity of causal generation. Demo available at https://yihao‑meng.github.io/CausalCine/
Authors:Runhui Huang, Jie Wu, Rui Yang, Zhe Liu, Hengshuang Zhao
Abstract:
In this paper, we propose AlphaGRPO, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR‑Diffusion Unified Multimodal Models (UMMs) to enhance multimodal generation capabilities without an additional cold‑start stage. Our approach unlocks the model's intrinsic potential to perform advanced reasoning tasks: Reasoning Text‑to‑Image Generation, where the model actively infers implicit user intents, and Self‑Reflective Refinement, where it autonomously diagnoses and corrects misalignments in generated outputs. To address the challenge of providing stable supervision for real‑world multimodal generation, we introduce the Decompositional Verifiable Reward (DVReward). Unlike holistic scalar rewards, DVReward utilizes an LLM to decompose complex user requests into atomic, verifiable semantic and quality questions, which are then evaluated by a general MLLM to provide reliable and interpretable feedback. Extensive experiments demonstrate that AlphaGRPO yields robust improvements across multimodal generation benchmarks, including GenEval, TIIF‑Bench, DPG‑Bench and WISE, while also achieving significant gains in editing tasks on GEdit without training on editing tasks. These results validate that our self‑reflective reinforcement approach effectively leverages inherent understanding to guide high‑fidelity generation. Project page: https://huangrh99.github.io/AlphaGRPO/
Authors:Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xiaohan Yu, Lin Gu, Gim Hee Lee
Abstract:
Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked existing approaches. This paper presents AmbiSuR, a framework that explores an intrinsic solution upon Gaussian Splatting for the photometric ambiguity‑robust surface 3D reconstruction with high performance. Starting by revisiting the foundation, our investigation uncovers two built‑in primitive‑wise ambiguities in representation, while revealing an intrinsic potential for ambiguity self‑indication in Gaussian Splatting. Stemming from these, a photometric disambiguation is first introduced, constraining ill‑posed geometry solution for definite surface formation. Then, we propose an ambiguity indication module that unleashes the self‑indication potential to identify and further guide correcting underconstrained reconstructions. Extensive experiments demonstrate our superior surface reconstructions compared to existing methods across various challenging scenarios, excelling in broad compatibility. Project: https://fictionarry.github.io/AmbiSuR‑Proj/ .
Authors:Alan Z. Song, Yinjie Chen, Mu Nan, Rui Zhang, Jiahang Cao, Weijian Mai, Muquan Yu, Hossein Adeli, Deva Ramanan, Michael J. Tarr, Andrew F. Luo
Abstract:
Vision Transformers (ViTs) achieve strong data‑driven scaling by leveraging all‑to‑all self‑attention. However, this flexibility incurs a computational cost that scales quadratically with image resolution, limiting ViTs in high‑resolution domains. Underlying this approach is the assumption that pairwise token interactions are necessary for learning rich visual‑semantic representations. In this work, we challenge this assumption, demonstrating that effective visual representations can be learned without any direct patch‑to‑patch interaction. We propose VECA (Visual Elastic Core Attention), a vision transformer architecture that uses efficient linear‑time core‑periphery structured attention enabled by a small set of learned cores. In VECA, these cores act as a communication interface: patch tokens exchange information exclusively through the core tokens, which are initialized from scratch and propagated across layers. Because the N image patches only directly interact with a resolution invariant set of C learned "core" embeddings, this yields linear complexity O(N) for predetermined C, which bypasses quadratic scaling. Compared to prior cross‑attention architectures, VECA maintains and iteratively updates the full set of N input tokens, avoiding a small C‑way bottleneck. Combined with nested training along the core axis, our model can elastically trade off compute and accuracy during inference. Across classification and dense tasks, VECA achieves performance competitive with the latest vision foundation models while reducing computational cost. Our results establish elastic core‑periphery attention as a scalable alternative building block for Vision Transformers.
Authors:Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu, Hu Yu, Siming Fu, Yuming Li, Zeyue Xue, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao
Abstract:
Recent advances in joint audio‑video generation have been remarkable, yet real‑world applications demand strong per‑modality fidelity, cross‑modal alignment, and fine‑grained synchronization. Reinforcement Learning (RL) offers a promising paradigm, but its extension to multi‑objective and multi‑modal joint audio‑video generation remains unexplored. Notably, our in‑depth analysis first reveals that the primary obstacles to applying RL in this stem from: (i) multi‑objective advantages inconsistency, where the advantages of multimodal outputs are not always consistent within a group; (ii) multi‑modal gradients imbalance, where video‑branch gradients leak into shallow audio layers responsible for intra‑modal generation; (iii) uniform credit assignment, where fine‑grained cross‑modal alignment regions fail to get efficient exploration. These shortcomings suggest that vanilla RL fine‑tuning strategy with a single global advantage often leads to suboptimal results. To address these challenges, we propose OmniNFT, a novel modality‑aware online diffusion RL framework with three key innovations: (1) Modality‑wise advantage routing, which routes independent per‑reward advantages to their respective modality generation branches. (2) Layer‑wise gradient surgery, which selectively detaches video‑branch gradients on shallow audio layers while retaining those for cross‑modal interaction layers. (3) Region‑wise loss reweighting, which modulates policy optimization toward critical regions related to audio‑video synchronization and fine‑grained alignment. Extensive experiments on JavisBench and VBench with the LTX‑2 backbone demonstrate that OmniNFT achieves comprehensive improvements in audio and video perceptual quality, cross‑modal alignment, and audio‑video synchronization.
Authors:Hao Zhu, Shuo Jin, Wenbin Liao, Jiayu Xiao, Yan Zhu, Siyue Yu, Feng Dai
Abstract:
Pursuing training‑free open‑vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep‑seated spatial bias in CLIP. To overcome the limitations of existing solutions, this work moves beyond the CLIP‑based paradigm and harnesses the recent spatially‑aware dino.txt framework to facilitate more efficient and high‑quality dense prediction. While dino.txt exhibits robust spatial awareness, we find that the semantic ambiguity of text queries gives rise to severe mismatch within its dense cross‑modal interactions. To address this, we introduce Visual‑guided Prompt evolution (VIP) to rectify the semantic expressiveness of text queries in dino.txt, unleashing its potential for fine‑grained object perception. Towards this end, VIP integrates alias expansion with a visual‑guided distillation mechanism to mine valuable semantic cues, which are robustly aggregated in a saliency‑aware manner to yield a high‑fidelity prediction. Extensive evaluations demonstrate that VIP: 1. surpasses the top‑leading methods by 1.4%‑8.4% average mIoU, 2. generalizes well to diverse challenging domains, and 3. requires marginal inference time and memory overhead.
Authors:Luca Parolari, Pietro Gori, Lamberto Ballan, Carlo Biffi, Loic Le Folgoc
Abstract:
Learning robust representations of polyp tracklets is key to enabling multiple AI‑assisted colonoscopy applications, from polyp characterization to automated reporting and retrieval. Supervised contrastive learning is an effective approach for learning such representations, but it typically relies on correct positive and negative definitions. Collecting these labels requires linking tracklets that depict the same underlying polyp entity throughout the video, which is costly and demands specialized clinical expertise. In this work, we leverage the sequential workflow of colonoscopy procedures to derive self‑supervised associations from temporal structure. Since temporally derived associations are not guaranteed to be correct, we introduce a noise‑aware contrastive loss to account for noisy associations. We demonstrate the effectiveness of the learned representations across multiple downstream tasks, including polyp retrieval and re‑identification, size estimation, and histology classification. Our method outperforms prior self‑supervised and supervised baselines, and matches or exceeds recent foundation models across all tasks, using a lightweight encoder trained on only 27 videos. Code is available at https://github.com/lparolari/ntssl.
Authors:Junxian Li, Kai Liu, Zizhong Ding, Zhixin Wang, Zhikai Chen, Renjing Pei, Yulun Zhang
Abstract:
The development of separate‑encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. In this paper, we focus on understanding‑side visual token reduction for improving the efficiency of separate‑encoder UMMs. While this topic has been widely studied for MLLMs, existing methods typically rely on attention scores, text‑image similarity and so on, implicitly assuming that the final objective is discriminative reasoning. This assumption does not hold for UMMs, where understanding‑side visual tokens must also preserve the model's capabilities for editing images. We propose G^2TR, a generation‑guided visual token reduction framework for separate‑encoder UMMs. Our key insight is that the generation branch provides a task‑agnostic signal for identifying understanding‑side visual tokens that are not only semantically relevant but also important for latent‑space image reconstruction and generation. G^2TR estimates token importance from consistency with VAE latent, performs balanced token selection, and merges redundant tokens into retained representatives to reduce information loss. The method is training‑free, plug‑and‑play, and applied only after the understanding encoding stage, making it compatible with existing UMM inference pipelines. Experiments on image understanding and editing benchmarks show that G^2TR substantially reduces visual tokens and prefill computation by 1.94x while maintaining both reasoning accuracy and editing quality, outperforming baselines on almost all benchmarks. Code is at: https://github.com/lijunxian111/G2TR.
Authors:Moussa Kassem Sbeyti, Joshua Holstein, Philipp Spitzer, Nadja Klein, Gerhard Satzger
Abstract:
High‑quality labeled data is essential for training robust machine learning models, yet obtaining annotations at scale remains expensive. AI‑assisted annotation has therefore become standard in large‑scale labeling workflows. However, in tasks where model predictions carry two independent components, a class label and spatial boundaries, a model may classify an object with high confidence while mislocalizing it. Existing AI‑assisted workflows offer annotators no signal about where spatial errors are most likely. Without such guidance, humans may systematically underinspect subtly misplaced boxes. We address this by studying the effect of visualizing spatial uncertainty via a purpose‑built interface. In a controlled study with 120 participants, those receiving uncertainty cues achieve higher label quality while being faster overall. A box‑level analysis confirms that the cues redirect annotator effort toward high‑uncertainty predictions and away from well‑localized boxes. These findings establish localization uncertainty as a lever to improve human‑in‑the‑loop annotation. Code is available at https://mos‑ks.github.io/MUHA/.
Authors:Luming Wang, Hao Shi, Jiajun Zhai, Kailun Yang, Kaiwei Wang
Abstract:
Egocentric 3D hand pose estimation and gesture recognition are essential for immersive augmented/virtual reality, human‑computer interaction, and robotics. However, conventional frame‑based cameras suffer from motion blur and limited dynamic range, while existing event‑based methods are hindered by ego‑motion interference, monocular depth ambiguity, and the lack of large‑scale real‑world stereo datasets. To overcome these limitations, we propose EgoEV‑HandPose, an end‑to‑end framework for joint 3D bimanual pose estimation and gesture recognition from stereo event streams. Central to our approach is KeypointBEV, a flexible stereo fusion module that lifts features into a canonical bird's‑eye‑view space and employs an iterative reprojection‑guided refinement loop to progressively resolve depth uncertainty and enforce kinematic consistency. In addition, we introduce EgoEVHands, the first large‑scale real‑world stereo event‑camera dataset for egocentric hand perception, containing 5,419 annotated sequences with dense 3D/2D keypoints across 38 gesture classes under varying illumination. Extensive experiments demonstrate that EgoEV‑HandPose achieves state‑of‑the‑art performance with an MPJPE of 30.54mm and 86.87% Top‑1 gesture recognition accuracy, significantly outperforming RGB‑based stereo and prior event‑camera methods, particularly in low‑light and bimanual occlusion scenarios, thereby setting a new benchmark for event‑based egocentric perception. The established dataset and source code will be publicly released at https://github.com/ZJUWang01/EgoEV‑HandPose.
Authors:Xinjia Li, Rui Wang, Qiurong Peng, Lingfei Ye, Dengrong Zhang, Haoyu Zhang
Abstract:
Farmland Semantic Change Detection (SCD) is essential for cultivated land protection, yet existing benchmarks and models remain insufficient for fine‑grained farmland conversion monitoring. Current datasets often lack dedicated "from‑to" annotations, while visual change detection models are easily disturbed by phenology‑induced pseudo‑changes caused by crop rotation, seasonal variation, and illumination differences. To address these challenges, we construct HZNU‑FCD, a large‑scale fine‑grained farmland SCD benchmark with a unified five‑class farmland‑to‑non‑farmland annotation protocol. It contains 4,588 bitemporal image pairs with pixel‑level labels for practical farmland protection. Based on this benchmark, we propose a large‑small collaborative SCD framework that integrates a task‑driven small visual model with a frozen large vision‑language model. The small model, Fine‑grained Difference‑aware Mamba (FD‑Mamba), learns dense change representations for boundary preservation and small‑region localization. The large‑model pathway, Cross‑modal Logical Arbitration (CMLA), introduces CLIP‑based textual priors for prompt‑guided semantic arbitration and pseudo‑change suppression. To enable effective collaboration, we design a hard‑region co‑training strategy that supervises the CMLA semantic score map only on low‑confidence pixels. Experiments show that our method achieves 97.63% F1, 96.32% IoU, and 96.35% SCD_IoU_mean on HZNU‑FCD with only 6.65M trainable parameters. Compared with the multimodal ChangeCLIP‑ViT, which leverages vision‑language information for change detection, our method improves F1 by 10.19 percentage points on HZNU‑FCD. It also achieves 91.43% F1 and 84.21% IoU on LEVIR‑CD, and 93.85% F1 and 88.41% IoU on WHU‑CD, demonstrating strong robustness and generalization. The code is available at https://github.com/Lovelymili/FD‑Mamba.
Authors:Yaofang Liu, Kangning Cui, Meng Chu, Zhaoqing Li, Suiyun Zhang, Jean-Michel Morel, Xiaodong Cun, Haoxuan Che, Rui Liu, Raymond H. Chan
Abstract:
Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to serialize this intent into text, a bottleneck that compresses signals like spatial structure, exact appearance, and glyph shape. We propose \emphvisual‑to‑visual (V2V) generation, in which the user conditions a generative model with a visual specification page rather than a text prompt. The page is not an edit target, but a visual document that specifies the desired output. We introduce V2V‑Zero, a training‑free framework that exposes this interface in existing vision‑language model (VLM) conditioned generators by replacing text‑only conditioning with final‑layer hidden states extracted from visual pages, exploiting the fact that the frozen VLM already maps both text and images into the generator's conditioning space.
On GenEval, V2V‑Zero reaches 0.85 with a frozen Qwen‑Image backbone, closely matching its optimized text‑to‑image performance without fine‑tuning. To evaluate the broader V2V space, we introduce Simple‑V2V Bench, spanning seven visual‑conditioning tasks and seven models, including GPT Image 2, Nano Banana 2, Seedream 5.0 Lite, open‑weight baselines, and a video extension. V2V‑Zero scores 32.7/100, outperforming evaluated open‑weight image baselines and revealing a clear capability hierarchy: attribute binding is strong, content generation is unreliable, and structural control remains hard even for commercial systems. A HunyuanVideo‑1.5 extension scores 20.2/100, showing the interface transfers beyond images. Mechanistic analysis shows the default reasoning path is primarily visually routed, with 95.0% of conditioning‑token attention mass on visual‑page hidden states.
Authors:Shuo Ni, Tong Wang, Jing Zhang, He Chen, Haonan Guo, Ning Zhang, Bo Du
Abstract:
Vision‑Language Models (VLMs) increasingly operate on ultra‑high‑resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scale mismatch between large‑scale scene context and micro‑scale targets. We refer to this empirical gap as a "resolution illusion": higher input resolution provides the appearance of richer visual detail, but does not necessarily yield reliable perception of spatially small, task‑relevant evidence. To benchmark this challenge, we introduce UHR‑Micro, a benchmark comprising 11,253 instructions grounded in 1,212 UHR images, designed to evaluate VLMs at the spatial limits of native Earth observation imagery. UHR‑Micro spans diverse micro‑target scales, context requirements, task families, and visual conditions, and provides diagnostic annotations that support controlled evaluation and fine‑grained error attribution. Experiments with representative high‑resolution VLMs show substantial failures in spatial grounding and evidence parsing, despite access to high‑resolution inputs. Further analysis suggests that these failures are not fully resolved by increasing model capacity, but are closely tied to insufficient guidance in locating and using task‑relevant micro‑evidence. Motivated by this finding, we propose Micro‑evidence Active Perception (MAP), a reference agent that decomposes queries into evidence‑seeking steps, actively inspects candidate regions, and grounds its answers in localized observations. MAP‑Agent improves micro‑level perception by making high‑resolution reasoning evidence‑centered rather than image‑centered. Together, UHR‑Micro and MAP‑Agent provide a diagnostic platform for evaluating, understanding, and advancing high‑resolution reasoning in Earth observation VLMs. Datasets and source code were released at https://github.com/MiliLab/UHR‑Micro.
Authors:Xin Cheng, Xihua Wang, Ying Ba, Yuyue Wang, Kaisi Guan, Yinbo Wang, Wenpu Li, Ruihua Song
Abstract:
Recent advancements in video‑audio joint generation have achieved remarkable success in semantic correspondence. However, achieving precise temporal synchronization, which requires fine‑grained alignment between audio events and their visual triggers, remains a challenging problem. The post‑training method for joint generation is largely dominated by Supervised Fine‑Tuning, but the commonly used Mean Squared Error loss provides insufficient penalties for subtle temporal misalignments. Direct Preference Optimization offers an alternative by introducing explicit misaligned counterparts to better improve temporal sensitivity. In this paper we propose a post‑training framework SyncDPO, leveraging DPO to improve the temporal sensitivity of V‑A joint generation. Conventional DPO pipelines typically depend on costly sampling‑and‑ranking procedures to construct preference pairs, resulting in substantial computational cost. To improve efficiency, we introduce a suite of on‑the‑fly rule‑based negative construction strategies that distort temporal structures without incurring additional annotation or sampling. We demonstrate that the temporal alignment capability can be effectively reinforced by providing explicit negative supervision through temporally distorted V‑A pairs. Accordingly, we implement a curriculum learning strategy that progressively increases the difficulty of negative samples, transitioning from coarse misalignment to subtle inconsistencies. Extensive objective and subjective experiments across four diverse benchmarks, ranging from ambient sound videos to human speech videos, demonstrate that SyncDPO significantly outperforms other methods in improving model's temporal alignment capability. It also demonstrates superior generalization on out‑of‑distribution benchmark by capturing intrinsic motion‑sound dynamics. Demo and code is available in https://syncdpo.github.io/syncdpo/.
Authors:Md Abulkalam Azad, Vegard Holmstrøm, John Nyberg, Lasse Lovstakken, Håvard Dalen, Bjørnar Grenne, Andreas Østvik
Abstract:
Myocardial point tracking (MPT) has recently emerged as a promising direction for motion estimation in echocardiography, driven by advances in general‑purpose point tracking methods. However, myocardial motion fundamentally differs from motion encountered in natural videos, as it arises from physiologically constrained deformation that is spatially and temporally continuous throughout the cardiac cycle. Consequently, motion trajectories typically remain locally confined despite substantial tissue deformation. Motivated by these properties, we revisit the architectural design for MPT and find that coarse initialization in commonly used two‑stage coarse‑to‑fine architectures may be unnecessary in this domain. In this work, we propose a fine‑stage‑only architecture, EchoTracker2, which enriches pixel‑precise features with local spatiotemporal context and integrates them with long‑range joint temporal reasoning for robust tracking. Experimental results across in‑distribution, out‑of‑distribution (OOD), and public synthetic datasets show that our model improves position accuracy by 6.5% and reduces median trajectory error by 12.2% relative to a domain‑specific state‑of‑the‑art (SOTA) model. Compared to the best general‑purpose point tracking method, the improvements are 2.0% and 5.3%, respectively. Moreover, EchoTracker2 shows better agreement with expert‑derived global longitudinal strain (GLS) and enhances test‑rest reproducibility. Source code will be available at: https://github.com/riponazad/ptecho.
Authors:Yexing Xu, Wei Feng, Shen Zhang, Haohan Wang, Yuxin Qin, Yaoyu Li, Ao Ma, Yuhao Luo, Lu Wang, Xudong Ren, Haoran Wang, Run Ling, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Longguang Wang, Yulan Guo
Abstract:
Generating realistic and user‑preferred advertisements is a key challenge in e‑commerce. Existing approaches utilize multiple independent models driven by click‑through‑rate (CTR) to controllably create attractive image or text advertisements. However, their pipelines lack cross‑modal perception and rely on CTR that only reflects average preferences. Therefore, we explore jointly generating personalized image‑text advertisements from historical click behaviors. We first design a Unified Advertisement Generative model (Uni‑AdGen) that employs a single autoregressive framework to produce both advertising images and texts. By incorporating a foreground perception module and instruction tuning, Uni‑AdGen enhances the realism of the generated content. To further personalize advertisements, we equip Uni‑AdGen with a coarse‑to‑fine preference understanding module that effectively captures user interests from noisy multimodal historical behaviors to drive personalized generation. Additionally, we construct the first large‑scale Personalized Advertising image‑text dataset (PAd1M) and introduce a Product Background Similarity (PBS) metric to facilitate training and evaluation. Extensive experiments show that our method outperforms baselines in general and personalized advertisement generation. Our project is available at https://github.com/JD‑GenX/Uni‑AdGen.
Authors:Haofeng Liu, Yang Zhou, Ziheng Wang, Zhengbo Xu, Zhan Peng, Jie Ma, Jun Liang, Shengfeng He, Jing Li
Abstract:
Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors offer visual fidelity but lack geometric correspondence. Existing methods either propagate geometric errors throughout generation or suffer from signal conflicts when fusing both statically. We introduce MoCam, which employs structured denoising dynamics to orchestrate a coordinated progression from geometry to appearance within the diffusion process. MoCam first leverages geometric priors in early stages to anchor coarse structures and tolerate their incompleteness, then switches to appearance priors in later stages to actively correct geometric errors and refine details. This design naturally unifies static and dynamic view synthesis by temporally decoupling geometric alignment and appearance refinement within the diffusion process. Experiments demonstrate that MoCam significantly outperforms prior methods, particularly when point clouds contain severe holes or distortions, achieving robust geometry‑appearance disentanglement.
Authors:Xiaofeng Tan, Jun Liu, Bin-Bin Gao, Yuanting Fan, Xi Jiang, Chengjie Wang, Hongsong Wang, Feng Zheng
Abstract:
RLHF is widely used to align flow‑matching text‑to‑image models with human preferences, but often leads to severe diversity collapse after fine‑tuning. In RL, diversity is often assumed to correlate with policy entropy, motivating entropy regularization. However, we show this intuition breaks in flow models: policy entropy remains constant, even while perceptual diversity collapses. We explain this mismatch both theoretically and empirically: the constant entropy arises from the fixed, pre‑defined noise schedule, while the diversity collapse is driven by the mode‑seeking nature of policy gradients. As a result, policy entropy fails to prevent the model from converging to a narrow high‑reward region in the perceptual space. To this end, we introduce perceptual entropy that captures diversity in a perceptual space and maintains the property of standard entropy. Building upon this insight, we propose two entropy‑regularized strategies, Perceptual Entropy Constraint and Perceptual Constraints on Generation Space, to preserve perceptual diversity and improve the quality. Experiments across two base models, neural and rule‑based rewards, and three perceptual spaces demonstrate consistent gains in the quality‑diversity trade‑off; PEC achieves the best overall score of 0.734 (vs. baseline's 0.366); a complementary setting of PEC further reaches a diversity average of 0.989 (vs. baseline's 0.047). Our project page (https://xiaofeng‑tan.github.io/projects/PEC) is publicly available.
Authors:Muhammad Aqeel, Maham Nazir, Uzair Khan, Marco Cristani, Francesco Setti
Abstract:
Zero‑shot anomaly detection aims to identify defects in unseen categories without target‑specific training. Existing methods usually apply the same feature transformation to all samples, treating normal and anomalous data uniformly despite their fundamentally asymmetric distributions, compact normals versus diverse anomalies. We instead exploit this natural asymmetry by proposing AVA‑DINO, an anomaly‑aware vision‑language adaptation framework with dual specialized branches for normal and anomalous patterns that adapt frozen DINOv3 visual features. During training on auxiliary data, the two branches are learned jointly with a text‑guided routing mechanism and explicit routing regularization that encourages branch specialization. At test time, only the input image and fixed, predefined language descriptions are used to dynamically combine the two branches, enabling an asymmetric activation. This design prevents degenerate uniform routing and allows context‑specific feature transformations. Experiments across nine industrial and medical benchmarks demonstrate state‑of‑the‑art performance, achieving 93.5% image‑AUROC on MVTec‑AD and strong cross‑domain generalization to medical imaging without domain‑specific fine‑tuning. https://github.com/aqeeelmirza/AVA‑DINO
Authors:Che Liu, Lichao Ma, Xiangyu Tony Zhang, Yuxin Zhang, Haoyang Zhang, Xuerui Yang, Fei Tian
Abstract:
Omni‑modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer a query. We study whether current omni‑modal benchmarks separate visual shortcuts from genuine audio‑visual‑language evidence integration, and how post‑training behaves under a visually debiased evaluation setting. We audit nine omni‑modal benchmarks with visual‑only probing, remove visually solvable queries, and retain full subsets when filtering is undefined or would make comparisons unstable. This yields OmniClean, a cleaned evaluation view with 8,551 retained queries from 16,968 audited queries. On OmniClean, we evaluate OmniBoost, a three‑stage post‑training recipe based on Qwen2.5‑Omni‑3B: mixed bi‑modal SFT, mixed‑modality RLVR, and SFT on self‑distilled data. Balanced bi‑modal SFT gives limited and uneven gains, RLVR provides the first broad improvement, and self‑distillation reshapes the benchmark profile. After SFT on self‑distilled data, the 3B model reaches performance comparable to, and in aggregate slightly above, Qwen3‑Omni‑30B‑A3B‑Instruct without using a stronger omni‑modal teacher. These results show that omni‑modal progress is easier to interpret when evaluation controls visual leakage, and that small omni‑modal models can benefit from staged post‑training with self‑distilled omni‑query supervision. Project page: https://cheliu‑computation.github.io/omni/
Authors:Xinyi Zhang, Manuel Günther
Abstract:
Deep Learning has revolutionized machine learning, reaching unprecedented levels of accuracy, but at the cost of reduced interpretability. Especially in image processing systems, deep networks transform local pixel information into more global concepts in a highly obscured manner. Explainable AI methods for image processing try to shed light on this issue by highlighting the regions of the image that are important for the prediction task. Among these, Class Activation Mapping (CAM) and its gradient‑based variants compute attributions based on the feature map and upscale them to the image resolution, assuming that feature map locations are influenced only by underlying regions. Perturbation‑based methods, such as CorrRISE, on the other hand, try to provide pixel‑level attributions by perturbing the input with fixed patches and checking how the output of the network changes. In this work, we propose Feature Activation Map Explanation (FAME), which combines both worlds by using network gradients to compute changes to the input image, manipulating it in a gradient‑driven way rather than using fixed patches. We apply this technique on two common tasks, image classification and face recognition, and show that CAM's above‑mentioned assumption does not hold for deeper networks. We qualitatively and quantitively show that FAME produces attribution maps that are competitive state‑of‑the‑art systems. Our code is available: \footnotesize https://github.com/AIML‑IfI/fame.
Authors:Zhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang, Jiawei Chen, Zhuoqi Zeng, Wei Zhang, Chengjie Wang, Jian Yang, Ying Tai
Abstract:
Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel‑space models from scratch demands prohibitive computational and data resources. To address this, we propose the Latent‑to‑Pixel (L2P) transfer paradigm, an efficient framework that directly harnesses the rich knowledge of pre‑trained LDMs to build powerful pixel‑space models. Specifically, L2P discards the VAE in favor of large‑patch tokenization and freezes the source LDM's intermediate layers, exclusively training shallow layers to learn the latent‑to‑pixel transformation. By utilizing LDM‑generated synthetic images as the sole training corpus, L2P fits an already smooth data manifold, enabling rapid convergence with zero real‑data collection. This strategy allows L2P to seamlessly migrate massive latent priors to the pixel space using only 8 GPUs. Furthermore, eliminating the VAE memory bottleneck unlocks native 4K ultra‑high resolution generation. Extensive experiments across mainstream LDM architectures show that L2P incurs negligible training overhead, yet performs on par with the source LDM on DPG‑Bench and reaches 93% performance on GenEval.
Authors:Sohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler, Christos Sakaridis, Suha Kwak
Abstract:
The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in safety‑critical domains. This paper offers the first comprehensive study on robust PVOS (RobustPVOS). We first construct a new, comprehensive benchmark with two real‑world evaluation datasets of 351 video clips and more than 2,500 object masks under real‑world adverse conditions. At the same time, we generate synthetic training data by applying diverse and temporally varying corruptions to existing VOS datasets. Moreover, we present a new RobustPVOS method, dubbed Memory‑object‑conditioned Gated‑rank Adaptation (MoGA). The key to successfully performing RobustPVOS is two‑fold: effectively handling object‑specific degradation and ensuring temporal consistency in predictions. MoGA leverages object‑specific representations maintained in memory across frames to condition the robustification process, which allows the model to handle each tracked object differently in a temporally consistent way. Extensive experiments on our benchmark validate MoGA's efficacy, showing consistent and significant improvements across diverse corruption types on both synthetic and real‑world datasets, establishing a strong baseline for future RobustPVOS research. Our benchmark is publicly available at https://sohyun‑l.github.io/RobustPVOS_project_page/.
Authors:Gengluo Li, Shangpin Peng, Xingyu Wan, Chengquan Zhang, Hao Feng, Xin Xu, Pian Wu, Bang Li, Zengmao Ding, Yongge Liu, Yipei Ye, Yang Yang, Zhan Shu, Guojun Yan, Zhe Li, Can Ma, Weiping Wang, Yu Zhou, Han Hu
Abstract:
Vision Large Language Models (VLLMs) have achieved remarkable success in modern text‑rich visual understanding. However, their perceptual robustness in the face of the continuous morphological evolution of historical writing systems remains largely unexplored. Existing ancient text datasets typically focus on isolated historical periods, failing to capture the systematic visual distribution shifts spanning thousands of years. To bridge this gap and empower Digital Humanities, we introduce Chronicles‑OCR, the first comprehensive benchmark specifically designed to evaluate the cross‑temporal visual perception capabilities of VLLMs across the complete evolutionary trajectory of Chinese characters, known as the Seven Chinese Scripts. Curated in collaboration with top‑tier institutional domain experts, the dataset comprises 2,800 strictly balanced images encompassing highly diverse physical media, ranging from tortoise shells to paper‑based calligraphy. To accommodate the drastic morphological and topological variations across different historical stages, we propose a novel Stage‑Adaptive Annotation Paradigm. Based on this, Chronicles‑OCR formulates four rigorous quantitative tasks: cross‑period character spotting, fine‑grained archaic character recognition via visual referring, ancient text parsing, and script classification. By isolating visual perception from semantic reasoning, Chronicles‑OCR provides an authoritative platform to expose the limitations of current VLLMs, paving the way for robust, evolution‑aware historical text perception. Chronicles‑OCR is publicly available at https://github.com/VirtualLUOUCAS/Chronicles‑OCR.
Authors:Maham Nazir, Muhammad Aqeel, Richong Zhang, Francesco Setti
Abstract:
Multimodal video summarization requires visual features that align semantically with language generation. Traditional approaches rely on CNN features trained for object classification, which represent visual concepts as discrete categories not aligned with natural language. We propose ClipSum, a framework that leverages frozen CLIP vision‑language features with explicit temporal modeling and dimension‑adaptive fusion for instructional video summarization. CLIP's contrastive pre‑training on 400M image‑text pairs yields visual features semantically aligned with the linguistic concepts that text decoders generate, bridging the vision‑language gap at the representation level. On YouCook2, ClipSum achieves 33.0% ROUGE‑1 versus 30.5% for ResNet‑152 with 4x lower dimensionality (512 vs. 2048), demonstrating that semantic alignment matters more than feature capacity. Frozen CLIP (33.0%) surpasses fine‑tuned CLIP (32.3%), showing that preserving pre‑trained alignment is more valuable than task‑specific adaptation. https://github.com/aqeeelmirza/clipsum
Authors:Qi Zhao, Jun Chen, Ivor Tsang, Guang Dai
Abstract:
While modern diffusion models excel at generating diverse single images, extending this to sequential generation reveals a fundamental challenge: balancing narrative dynamism with multi‑character coherence. Existing methods often falter at this trade‑off, leading to artifacts where characters lose their identity or the story stagnates. To resolve this critical tension, we introduce RealDiffusion, a unified framework designed to reconcile robust coherence with narrative dynamism. Heat diffusion serves as a dissipative prior that averages neighboring features along the sequence and removes high‑frequency noise within the subject region. This suppresses attribute drift and stabilizes identity across frames. A region‑aware stochastic process then introduces small perturbations that explore nearby modes and prevent collapse so the story maintains pose change and scene evolution. We thus introduce a lightweight, training‑free Physics‑informed Attention mechanism that injects controllable physical priors into the self‑attention layers during inference. By modeling feature evolution as a configurable physical system, our method regularizes spatio‑temporal relationships without suppressing intentional, prompt‑driven changes. Extensive experiments demonstrate that RealDiffusion achieves substantial gains in character coherence while preserving narrative dynamism, outperforming state‑of‑the‑art approaches. Code is available at https://github.com/ShmilyQi‑CN/RealDiffusion.
Authors:Huiyu Yi, Zhiming Xu, Dunwei Tu, Zhicheng Wang, Baile Xu, Furao Shen
Abstract:
The Nearest Class Mean (NCM) classifier is widely favored in Class‑Incremental Learning (CIL) for its superior resistance to catastrophic forgetting compared to Fully Connected layers. While Neural Collapse (NC) theory supports NCM's optimality by assuming features collapse into single points, non‑linear feature drift and insufficient training in CIL often prevent this ideal state. Consequently, classes manifest as complex manifolds rather than collapsed points, rendering the single‑point NCM suboptimal. To address this, we propose Hierarchical‑Cluster SOINN (HC‑SOINN), a novel classifier that captures the topological structure of these manifolds via a ``local‑to‑global'' representation. Furthermore, we introduce Structure‑Topology Alignment via Residuals (STAR) method, which employs a fine‑grained pointwise trajectory tracking mechanism to actively deform the learned topology, allowing it to adapt precisely to complex non‑linear feature drift. Theoretical analysis and Procrustes distance experiments validate our framework's resilience to manifold deformations. We integrated HC‑SOINN into seven state‑of‑the‑art methods by replacing their original classifiers, achieving consistent improvements that highlight the effectiveness and robustness of our approach. Code is available at https://github.com/yhyet/HC_SOINN.
Authors:Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh
Abstract:
Listwise reranking is a key yet computationally expensive component in vision‑centric retrieval and multimodal retrieval‑augmented generation (M‑RAG) over long documents. While recent VLM‑based rerankers achieve strong accuracy, their practicality is often limited by long visual‑token sequences and multi‑step autoregressive decoding. We propose ZipRerank, a highly efficient listwise multimodal reranker that directly addresses both bottlenecks. It reduces input length via a lightweight query‑image early interaction mechanism and eliminates autoregressive decoding by scoring all candidates in a single forward pass. To enable effective learning, ZipRerank adopts a two‑stage training strategy: (i) listwise pretraining on large‑scale text data rendered as images, and (ii) multimodal finetuning with VLM‑teacher‑distilled soft‑ranking supervision. Extensive experiments on the MMDocIR benchmark show that ZipRerank matches or surpasses state‑of‑the‑art multimodal rerankers while reducing LLM inference latency by up to an order of magnitude, making it well‑suited for latency‑sensitive real‑world systems. The code is available at https://github.com/dukesun99/ZipRerank.
Authors:Yixu Feng, Zinan Zhao, Yanxiang Ma, Chenghao Xia, Chengbin Du, Yunke Wang, Chang Xu
Abstract:
Vision‑Language‑Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real‑time deployment. Existing token pruning methods suffer from a fundamental trade‑off: aggressive compression using pruning inevitably discards critical geometric details like contact points, leading to severe performance degradation. This forces a compromise, limiting the achievable compression rate and thus the potential speedup. We argue that breaking this trade‑off requires rethinking compression as a geometry‑aware, continuous token resampling in the vision encoder. To this end, we propose the Differentiable Grid Sampler (GridS), a plug‑and‑play module that performs task‑aware, continuous resampling of visual tokens in VLA. By adaptively predicting a minimal set of salient coordinates and extracting features via differentiable interpolation, GridS preserves essential spatial information while achieving drastic compression (with fewer than 10% original visual tokens). Experiments on both LIBERO benchmark and a real robotic platform demonstrate that validating the lowest feasible visual token count reported to date, GridS achieves a 76% reduction in FLOPs with no degradation in the success rate. The code is available at https://github.com/Fediory/Grid‑Sampler.
Authors:Minseok Kang, Minhyeok Lee, Jungho Lee, Minjung Kim, Donghyeong Kim, Dayeon Lee, Heeseung Choi, Ig-jae Kim, Sangyoun Lee
Abstract:
As Video Large Language Models (Video‑LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across frames. Training‑free token compression has emerged as a practical solution to this bottleneck. However, existing temporal compression methods rely primarily on cross‑frame token similarity or segmentation heuristics, overlooking each token's semantic role within its frame and failing to adapt compression strength to the compressibility of each frame pair. In this work, we propose OTT‑Vid, a transport‑derived allocation framework for temporal token compression. Our approach consists of two stages: spatial pruning identifies representative content within each frame, and optimal transport (OT) is then solved between neighboring frames to estimate temporal compressibility. We formulate this OT with non‑uniform token mass, which protects semantically important tokens from aggressive compression, and a locality‑aware cost that captures both feature and spatial disparities. The resulting transport plan jointly balances token importance and matching cost, while its total cost defines the transport difficulty of each frame pair, which we use to allocate compression budgets dynamically. Experiments on six benchmarks spanning video question answering and temporal grounding show that OTT‑Vid preserves 95.8% of VQA and 73.9% of VTG performance while retaining only 10% of tokens, consistently outperforming existing state‑of‑the‑art training‑free compression methods.
Authors:Xianzhe Fan, Yuxiang Lu, Shenyuan Gao, Xiaoyang Wu, Ruihua Han, Manling Li, Hengshuang Zhao
Abstract:
Vision‑Language‑Action (VLA) models are often brittle in fine‑grained manipulation, where minor action errors during the critical phases can rapidly escalate into irrecoverable failures. Since existing VLA models rely predominantly on successful demonstrations for training, they lack an explicit awareness of failure during these critical phases. To address this, we propose DreamAvoid, a critical‑phase test‑time dreaming framework that enables VLA models to anticipate and avoid failures. We also introduce an autonomous boundary learning paradigm to refine the system's understanding of the subtle boundary between success and failure. Specifically, we (1) utilize a Dream Trigger to determine whether the execution has entered a critical phase, (2) sample multiple candidate action chunks from the VLA via an Action Proposer, and (3) employ a Dream Evaluator, jointly trained on mixed data (success, failure, and boundary cases), to "dream" the short‑horizon futures corresponding to the candidate actions, evaluate their values, and select the optimal action. We conduct extensive evaluations on real‑world manipulation tasks and simulation benchmarks. The results demonstrate that DreamAvoid can effectively avoid failures, thereby improving the overall task success rate. Our code is available at https://github.com/XianzheFan/DreamAvoid.
Authors:SeongMin Jin, Doo Seok Jeong
Abstract:
Learning latent representations that capture both semantic and spatial information is central to efficient spatio‑semantic reasoning. However, many existing approaches rely on implicit latent structures combined with dense feature maps or task‑specific heads, limiting computational efficiency and flexibility. We propose WorldComp2D, a novel lightweight representation learning framework that explicitly structures latent space geometry according to object identity and spatial proximity using multiscale local receptive fields. This framework consists of (i) a proximity‑dependent encoder that maps a given observation into a spatio‑semantic latent space and (ii) a localizer that infers the coordinates of objects in the input from the resulting spatio‑semantic representation. Using facial landmark localization as a proof‑of‑concept, we show that, compared to SoTA lightweight models, WorldComp2D reduces the numbers of parameters and FLOPs by up to 4.0X and 2.2X, respectively, while maintaining real‑time performance on CPU. These results demonstrate that explicitly structured latent spaces provide an efficient and general foundation for spatio‑semantic reasoning. This framework is open‑sourced at https://github.com/JinSeongmin/WorldComp2D.
Authors:Lezhong Wang, Mehmet Onurcan Kaya, Siavash Bigdeli, Jeppe Revall Frisvad
Abstract:
Recent single‑image relighting methods, powered by advanced generative models, have achieved impressive photorealism on synthetic benchmarks. However, their effectiveness in the complex visual landscape of the real world remains largely unverified. A critical gap exists, as current datasets are typically designed for multi‑view reconstruction and fail to address the unique challenges of single‑image relighting. To bridge this synthetic‑to‑real gap, we introduce WildRelight, the first in‑the‑wild dataset specifically created for evaluating single‑image relighting models. WildRelight features a diverse collection of high‑resolution outdoor scenes, captured under strictly aligned, temporally varying natural illuminations, each paired with a high‑dynamic‑range environment map. Using this data, we establish a rigorous benchmark revealing that state‑of‑the‑art models trained on synthetic data suffer from severe domain shifts. The strictly aligned temporal structure of WildRelight enables a new paradigm for domain adaptation. We demonstrate this by introducing a physics‑guided inference framework that leverages the captured natural light evolution as a self‑supervised constraint. By integrating Diffusion Posterior Sampling (DPS) with temporal Sampling‑Aware Test‑Time Adaptation (TTA), we show that the dataset allows synthetic models to align with real‑world statistics on‑the‑fly, transforming the intractable sim‑to‑real challenge into a tractable self‑supervised task. The dataset and code will be made publicly available to foster robust, physically‑grounded relighting research.
Authors:Shivam Kumar
Abstract:
We introduce ShapeCodeBench, a synthetic benchmark for perception‑to‑program reconstruction: given a rendered raster image, a model must emit an executable drawing program that a deterministic evaluator re‑renders and compares with the target. The v1 DSL has four primitives on a 512 x 512 black‑on‑white canvas, but every instance is generated from a seeded RNG, so fresh held‑out sets can be created to reduce exact‑instance contamination. We release a frozen eval_v1 split with 150 samples across easy, medium, and hard tiers, scored by exact match, pixel accuracy, foreground IoU, parse success, and execution success. We evaluate an empty‑program floor, a classical computer‑vision heuristic, Claude Opus 4.7 at high and max effort, and GPT‑5.5 at medium and extra_high reasoning effort. The heuristic is competitive on easy scenes but collapses when overlaps fuse components; the strongest multimodal configuration preserves much of the foreground structure but still misses exact match because of small parameter errors. Best overall exact match remains low, so ShapeCodeBench is far from saturated. The benchmark code, frozen dataset, run artifacts, and paper sources are released to support independent replication and extension.
Authors:Yaxuan Song, Jianan Fan, Tianyi Wang, Qiuyue Hu, Hang Chang, Heng Huang, Weidong Cai
Abstract:
Histopathology whole‑slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular architecture and functional programs defining pathological states, whereas RNA sequencing (RNA‑seq) provides genome‑wide transcriptional profiles at substantial cost, thereby motivating WSI‑based genome‑wide transcriptomic prediction. Existing approaches for predicting gene expression from WSIs predominantly rely on deterministic regression with one‑to‑one mapping, limiting their ability to capture biological heterogeneity and predictive uncertainty. We propose RNA‑FM, a flow‑matching generative framework for genome‑wide bulk RNA‑seq prediction from WSIs. RNA‑FM formulates transcriptomic prediction as a continuous‑time conditional transport problem, learning a velocity field that maps a simple prior to the target gene expression distribution conditioned on morphologies. By integrating pathway‑level structure, RNA‑FM enables scalable and biologically interpretable genome‑wide gene expression imputation. Extensive experiments demonstrate that RNA‑FM consistently outperforms state‑of‑the‑art approaches while maintaining biological meaningfulness. Code is available at https://github.com/YXSong000/RNA‑FM.
Authors:Conglang Zhang, Yifan Zhan, Qingjie Wang, Zhanpeng Ouyang, Yu Li, Zihao Yang, Xiaoyang Guo, Weiqiang Ren, Qian Zhang, Zhen Dong, Yinqiang Zheng, Wei Yin, Zhengqing Chen
Abstract:
Closed‑loop driving simulation requires real‑time interaction beyond short offline clips, pushing current driving world models toward autoregressive (AR) rollout. Existing AR distillation approaches typically rely on frame sinks or student‑side degradation training. The former transfers poorly to driving due to fast ego‑motion and rapid scene changes, while the latter remains bounded by the teacher's single‑pass output length and thus provides only a limited supervision horizon. A natural question is: can the teacher itself be extended via AR rollout to provide unbounded‑horizon supervision at bounded memory cost? The key difficulty is that a standard teacher drifts under its own predictions, contaminating the supervision it provides. Our key insight is to make the teacher rollout‑capable, ensuring reliable supervision from its own AR rollouts. This is instantiated as HorizonDrive, an anti‑drifting training‑and‑distillation framework for AR driving simulation. First, scheduled rollout recovery (SRR) trains the base model to reconstruct ground‑truth future clips from prediction‑corrupted histories, yielding a teacher that remains stable across long AR rollouts. Second, the rollout‑capable teacher is extended via AR rollout, providing long‑horizon distribution‑matching supervision under bounded memory, while a short‑window student aligns to it with teacher rollout DMD (TRD) for efficient real‑time deployment. HorizonDrive natively supports minute‑scale AR rollout under bounded memory; on nuScenes, HorizonDrive reduces FID by 52% and FVD by 37%, and lowers ARE and DTW by 21% and 9% relative to the strongest long‑horizon streaming baselines, while remaining competitive with single‑pass driving video generators.
Authors:Mingtao Xian, Yifeng Yang, Qinying Gu, Xinbing Wang, Nanyang Ye
Abstract:
Multimodal Large Language Models (MLLMs) have shown strong performance in multi‑image cross‑modal retrieval, yet suffer from severe position bias, where predictions are dominated by input order rather than semantic relevance. Through empirical analysis, we identify a phenomenon termed Logit‑Attention Divergence, in which output logits are heavily biased while internal attention maps remain well‑aligned with relevant visual evidence. This observation reveals a fundamental limitation of existing logit‑level calibration methods such as PriDe. Based on this insight, we propose a training‑free, attention‑guided debiasing framework that leverages intrinsic attention signals for instance‑level correction at inference time, requiring only a minimal calibration set with negligible computational overhead. Experiments on MS‑COCO‑based benchmarks show that our method substantially improves permutation invariance and achieves state‑of‑the‑art performance, enhancing accuracy by over 40% compared to baselines. Code is available at https://github.com/brightXian/LAD.
Authors:Feng Chen, Xianghui Wang, Yuxuan Chen, Boying Li, Yefei He, Zeyu Zhang, Yicheng Wu
Abstract:
Vision‑Language‑Action (VLA) models predominantly adopt action chunking, i.e., predicting and committing to a short horizon of consecutive low‑level actions in a single forward pass, to amortize the inference cost of large‑scale backbones and reduce per‑step latency. However, committing these multi‑step predictions to real‑world execution requires balancing success rate against inference efficiency, a decision typically governed by fixed execution horizons tuned per task. Such heuristics ignore the state‑dependent nature of predictive reliability, leading to brittle performance in dynamic or out‑of‑distribution settings. In this paper, we introduce A3, an Adaptive Action Acceptance mechanism that reframes dynamic execution commitment as a self‑speculative prefix verification problem. A3 first computes a trajectory‑wise consensus score of actions via group sampling, then selects a representative draft and prioritizes downstream verification. Specifically, it enforces: (1) consensus‑ordered conditional invariance, which validates low‑consensus actions by judging whether they remain consistent when re‑decoded conditioned on high‑consensus actions; and (2) prefix‑closed sequential consistency, which guarantees physical rollout integrity by accepting only the longest continuous sequence of verified actions starting from the beginning. Consequently, the execution horizon emerges as the longest verifiable prefix satisfying both internal model logic and sequential execution constraints. Experiments across diverse VLA models and benchmarks demonstrate that A3 eliminates the need for manual horizon tuning while achieving a superior trade‑off between execution robustness and inference throughput.
Authors:Fanpu Cao, Xin Zou, Xuming Hu, Hui Xiong
Abstract:
Multimodal large language models (MLLMs) have become a key interface for visual reasoning and grounded question answering, yet they remain vulnerable to visual hallucinations, where generated responses contradict image content or mention nonexistent objects. A central challenge is that hallucination is not always caused by a simple lack of visual attention: the model may still assign substantial attention mass to image tokens while internally drifting toward an incorrect answer. In this paper, we show that the high‑frequency structure of visual attention, measured by layer‑wise Laplacian energy, reveals both the layer where hallucinated preferences emerge and the layer where the ground‑truth answer transiently recovers. Building on this finding, we propose LaSCD (Laplacian‑Spectral Contrastive Decoding), a training‑free decoding strategy that selects informative layers via Laplacian energy and remaps next‑token logits in closed form. Experiments on hallucination and general multimodal benchmarks show that LaSCD consistently reduces hallucination while preserving general capabilities, highlighting its potential as a faithful decoding paradigm. The code is available at https://github.com/macovaseas/LaSCD.
Authors:Zhenxi Zhang, Yitao Zhuang, Yao Pu, Peixin Yu, Zirong Li, Yan Xia, Hui Li, Bin Li, Fuchen Zheng, Ge Ren
Abstract:
Anatomical structure masks are widely adopted in radiotherapy dose prediction, as they provide explicit geometric constraints that facilitate structure‑dose coupling. However, conventional manual delineation of these masks requires precise annotation of structure boundaries relevant to radiotherapy, which is time‑consuming and labor‑intensive. To address these limitations, we propose a scribble‑guided dose prediction framework that relies solely on anatomical structures annotated with sparse scribbles. Specifically, we design a Scribble Completion Module (SCM) to generate dense anatomical masks by propagating sparse scribble labels to semantically similar voxels. During the propagation process, a supervoxel‑based regularization is introduced to preserve geometric boundary consistency to ensure anatomical plausibility. Furthermore, we propose a Structure‑Guided Dose Generation Module (SGDGM) to strengthen the correspondence between sparse structural cues and dose distribution. Herein, the completed dense masks derived from scribbles serve as structural guidance to condition the dose prediction network. This scribble‑mask‑dose consistency encourages high‑dose concentration within target volumes while effectively sparing surrounding organs‑at‑risk. Extensive experiments on the open‑source GDP‑HMM dataset demonstrate that the proposed method maintains superior dose prediction performance while substantially reducing annotation cost, providing a practical paradigm for dose prediction under sparse structural annotation. The code and reannotated scribbles are made publicly available at https://github.com/iCherishxixixi/ScribbleDose.
Authors:Hai Jiang, Zhen Liu, Yinjie Lei, Songchen Han, Bing Zeng, Shuaicheng Liu
Abstract:
In this paper, we propose a zero‑reference diffusion‑based framework, named ZeroIDIR, for illumination degradation image restoration, which decouples the restoration process into adaptive illumination correction and diffusion‑based reconstruction while being trained solely on low‑quality degraded images. Specifically, we design an adaptive gamma correction module that performs spatially varying exposure correction to generate illumination‑corrected only representations to mitigate exposure bias and serve as reliable inputs for subsequent diffusion processes, where a histogram‑guided illumination correction loss is introduced to regularize the corrected illumination distribution toward that of natural scenes. Subsequently, the illumination‑corrected image is treated as an intermediate noisy state for the proposed perturbed consistency diffusion model to reconstruct details and suppress noise. Moreover, a perturbed diffusion consistency loss is proposed to constrain the forward diffusion trajectory of the final restored image to remain consistent with the perturbed state, thus improving restoration fidelity and stability in the absence of supervision. Extensive experiments on publicly available benchmarks show that the proposed method outperforms state‑of‑the‑art unsupervised competitors and is comparable to supervised methods while being more generalizable to various scenes. Code is available at https://github.com/JianghaiSCU/ZeroIDIR.
Authors:Jimin Tang, Wenyuan Zhang, Junsheng Zhou, Zian Huang, Kanle Shi, Shenkun Xu, Yu-Shen Liu, Zhizhong Han
Abstract:
Gaussian Splatting has achieved remarkable progress in multi‑view surface reconstruction, yet it exhibits notable degradation when only few views are available. Although recent efforts alleviate this issue by enhancing multi‑view consistency to produce plausible surfaces, they struggle to infer unseen, occluded, or weakly constrained regions beyond the input coverage. To address this limitation, we present VidSplat, a training‑free generative reconstruction framework that leverages powerful video diffusion priors to iteratively synthesize novel views that compensate for missing input coverage, and thereby recover complete 3D scenes from sparse inputs. Specifically, we tackle two key challenges that enable the effective integration of generation and reconstruction. First, for 3D consistent generation, we elaborate a training‑free, stage‑wise denoising strategy that adaptively guides the denoising direction toward the underlying geometry using the rendered RGB and mask images. Second, to enhance the reconstruction, we develop an iterative mechanism that samples camera trajectories, explores unobserved regions, synthesizes novel views, and supplements training through confidence weighted refinement. VidSplat performs robustly to sparse input and even a single image. Extensive experiments on widely used benchmarks demonstrate our superior performance in sparse‑view scene reconstruction.
Authors:Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao, Hao Tang
Abstract:
Presentation generation is moving beyond static slide creation toward end‑to‑end presentation video generation with research grounding, multimodal media, and interactive delivery. We introduce PresentAgent‑2, an agentic framework for generating presentation videos from user queries. Given an open‑ended user query and a selected presentation mode, PresentAgent‑2 first summarizes the query into a focused topic and performs deep research over presentation‑friendly sources to collect multimodal resources, including relevant text, images, GIFs, and videos. It then constructs presentation slides, generates mode‑specific scripts, and composes slides, audio, and dynamic media into a complete presentation video. PresentAgent‑2 supports three independent presentation modes within a unified framework: Single Presentation, which generates a single‑speaker narrated presentation video; Discussion, which creates a multi‑speaker presentation with structured speaker roles, such as for asking guiding questions, explaining concepts, clarifying details, and summarizing key points; and Interaction, which independently supports answering audience questions grounded in the generated slides, scripts, retrieved evidence, and presentation context. To evaluate these capabilities, we build a multimodal presentation benchmark covering single presentation, discussion, and interaction scenarios, with task‑specific evaluation criteria for content quality, media relevance, dynamic media use, dialogue naturalness, and interaction grounding. Overall, PresentAgent‑2 extends presentation generation from document‑dependent slide creation to query‑driven, research‑grounded presentation video generation with multimodal media, dialogue, and interaction. Code: https://github.com/AIGeeksGroup/PresentAgent‑2. Website: https://aigeeksgroup.github.io/PresentAgent‑2.
Authors:Haoyu Zhang, Zeyu Zhang, Zedong Zhou, Yang Zhao, Hao Tang
Abstract:
Transformer‑based 3D reconstruction has emerged as a powerful paradigm for recovering geometry and appearance from multi‑view observations, offering strong performance across challenging visual conditions. As these models scale to larger backbones and higher‑resolution inputs, improving their efficiency becomes increasingly important for practical deployment. However, modern 3D transformer pipelines face two coupled challenges: dense multi‑view attention creates substantial token‑mixing overhead, and low‑precision execution can destabilize geometry‑sensitive representations and degrade depth, pose, and 3D consistency. To address the first challenge, we propose Lite3R, a model‑agnostic teacher‑student framework that replaces dense attention with Sparse Linear Attention to preserve important geometric interactions while reducing attention cost. To address the second challenge, we introduce a parameter‑efficient FP8‑aware quantization‑aware training (FP8‑aware QAT) strategy with partial attention distillation, which freezes the vast majority of pretrained backbone parameters and trains only lightweight linear‑branch projection layers, enabling stable low‑precision deployment while retaining pretrained geometric priors. We further evaluate Lite3R on two representative backbones, VGGT and DA3‑Large, over BlendedMVS and DTU64, showing that it substantially reduces latency (1.7‑2.0x) and memory usage (1.9‑2.4x) while preserving competitive reconstruction quality overall. These results demonstrate that Lite3R provides an effective algorithm‑system co‑design approach for practical transformer‑based 3D reconstruction. Code: https://github.com/AIGeeksGroup/Lite3R. Website: https://aigeeksgroup.github.io/Lite3R.
Authors:Ajay Vikram Periasami, Junlin Wang, Bhuwan Dhingra
Abstract:
Image‑to‑code generation tests whether a vision‑language model (VLM) can recover the structure of an image enough to express it as executable code. Existing benchmarks either focus on narrow visual domains, depend on paired executable reference code, or rely on generic rubrics that miss domain‑specific reconstruction errors. We introduce Vision2Code, a reference‑code‑free benchmark and evaluation framework for multi‑domain image‑to‑code generation. Vision2Code contains 2,169 test examples from 15 source datasets that span charts and plots, geometry, graphs, scientific imagery, documents, and 3D spatial scenes. Models generate executable programs, which we render and score against the source image using a VLM rater with dataset‑specific rubrics and deterministic guardrails for severe semantic failures. We report render‑success diagnostics that separate code execution failures from reconstruction quality. Human validation shows that this evaluation protocol aligns better with human judgments than either a generic visual rubric or embedding‑similarity baselines. Across nine open‑weight and proprietary models, we find that image‑to‑code performance is domain‑dependent: leading models perform well on regular chart‑ and graph‑like visuals but remain weak on spatial scenes, chemistry, documents, and circuit‑style diagrams. Finally, we show that evaluator‑filtered model outputs can serve as training data to improve image‑to‑code capability, with Qwen3.5‑9B improving from 1.60 to 1.86 on the benchmark without paired source programs. Vision2Code provides a reproducible testbed for measuring, diagnosing, and improving image‑to‑code generation. Our code and data are publicly available at https://image2code.github.io/vision2code/.
Authors:Xueqi Cheng, Yushun Dong
Abstract:
Multimodal large language models (MLLMs) have heterogeneous strengths across OCR, chart understanding, spatial reasoning, visual question answering, cost, and latency. Effective MLLM routing therefore requires more than estimating query difficulty: a router must match the multimodal requirements of the current image‑question input with the capabilities of each candidate model. We propose LatentRouter, a router that formulates MLLM routing as counterfactual multimodal utility prediction. Given an image‑question query, LatentRouter extracts learned multimodal routing capsules, represents each candidate MLLM with a model capability token, and performs latent communication between these states to estimate how each model would perform if selected. A distributional outcome head predicts model‑specific counterfactual quality, while a bounded capsule correction refines close decisions without allowing residual signals to dominate the prediction. The resulting utility‑based policy supports performance‑oriented and performance‑cost routing, and handles changing candidate pools through shared per‑model scoring with availability masking. Experiments on MMR‑Bench and VL‑RouterBench show that LatentRouter outperforms fixed‑model, feature‑level, and learned‑router baselines. Additional analyses show that the gains are strongest on multimodal task groups where model choice depends on visual, layout‑sensitive, or reasoning‑oriented requirements, and that latent communication is the main contributor to the improvement. The code is available at: https://github.com/LabRAI/LatentRouter.
Authors:Bulat Maksudov, Vladislav Kurenkov, Kathleen M. Curran, Alessandra Mileo
Abstract:
Existing medical‑agent benchmarks deliver imaging as pre‑selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiology‑agent benchmark in which the agent operates an OHIF viewer and an Orthanc DICOM server through twenty‑one function‑calling tools that span slice navigation, windowing, series selection, pixel‑coordinate annotation, and structured reporting. ABRA contains 655 programmatically generated tasks across three difficulty tiers and eight types (viewer control, metadata QA, vision probe, annotation, longitudinal comparison, BI‑RADS reporting, and oracle variants of annotation and BI‑RADS reporting), drawn from LIDC‑IDRI, Duke Breast Cancer MRI, and NLST New‑Lesion LongCT. Each episode is scored along Planning, Execution, and Outcome (Bluethgen et al., 2025) by task‑type‑specific automatic scorers. Ten current models, five closed‑weight and five open‑weight, reach at least 89% Execution on real annotation but only 0‑25% Outcome; on the paired oracle variant where a simulated detector supplies the finding, Outcome on the same task reaches 69‑100% across the models evaluated, localising the bottleneck to perception rather than tool orchestration. Code, task generators, and scorers are released at https://github.com/Luab/ABRA
Authors:Elias B. Krey, Nils Neukirch, Nils Strodthoff
Abstract:
Intermediate feature representations represent the backbone for the expressivity and adaptability of deep neural networks. However, their geometric structure remains poorly understood. In this submission, we provide indirect insights into this matter by applying a broad selection of manipulations in input space, ranging from geometric and photometric transformations to local masking and semantic manipulations using generative image editing models, and assess the feasibility of learning a mapping in the feature space, mapping from the original to the manipulated feature map. To this end, we devise different types of mappings, from linear to non‑linear and local to global mappings and assess both the reconstruction quality of the mapping as well as the semantic content of the mapped representations. We demonstrate the feasibility of learning such mappings for all considered transformations. While global (transformer) models that operate on the full feature map often achieve best results, we show that the same can be achieved with a shared linear model operating on a single feature vector typically with very little degradation in reconstruction quality, even for highly non‑trivial semantic manipulations. We analyze the corresponding mappings across different feature layers and characterize them according to dominance of weight vs. bias and the effective rank of the linear transformations. These results provide hints for the hypothesis that the feature space is to a first degree of approximation organized in linear structures. From a broader perspective, the study demonstrates that generative image editing models might open the door to a deeper understanding of the feature space through input manipulation.
Authors:Qi Cai, Jingwen Chen, Chengmin Gao, Zijian Gong, Yehao Li, Yingwei Pan, Yi Peng, Zhaofan Qiu, Kai Yu, Yiheng Zhang, Hao Ai, Siying Bai, Yang Chen, Zhihui Chen, Fengbin Gao, Ying Guo, Dong Li, Zhen Shen, Leilei Shi, Jing Wang, Siyu Wang, Yimeng Wang, Rui Zheng, Ting Yao, Tao Mei
Abstract:
The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDream‑O1‑Image, a natively unified generative foundation model via pixel‑space Diffusion Transformer, that pioneers a paradigm shift from modular architectures to an end‑to‑end in‑context visual generation engine. By mapping raw image pixels, text tokens, and task‑specific conditions into a single shared token space, HiDream‑O1‑Image achieves a structural unification of multimodal inputs within an Unified Transformer (UiT) architecture. This native encoding paradigm eliminates the need for separate VAEs or disjoint pre‑trained text encoders, allowing the model to treat diverse generation and editing tasks as a consistent in‑context reasoning process. Extensive experiments show that HiDream‑O1‑Image excels across various generation tasks, including text‑to‑image generation, instruction‑based editing, and subject‑driven personalization. Notably, with only 8B parameters, HiDream‑O1‑Image (8B) achieves performance parity with or even surpasses established state‑of‑the‑art models with significantly larger parameters (e.g., 27B Qwen‑Image). Crucially, to validate the immense scalability of this paradigm, we successfully scale the architecture up to over 200B parameters. Experimental results demonstrate that this massive‑scale version HiDream‑O1‑Image‑Pro (200B+) unlocks unprecedented generative capabilities and superior performance, establishing new state‑of‑the‑art benchmarks. Ultimately, HiDream‑O1‑Image highlights the immense potential of natively unified architectures and charts a highly scalable path toward next‑generation multimodal AI.
Authors:Yaman Kindap, Manfred Opper, Benjamin Dupuis, Umut Simsekli, Tolga Birdal
Abstract:
Modelling extreme events and heavy‑tailed phenomena is central to building reliable predictive systems in domains such as finance, climate science, and safety‑critical AI. While Lévy processes provide a natural mathematical framework for capturing jumps and heavy tails, Bayesian inference for Lévy‑driven stochastic differential equations (SDEs) remains intractable with existing methods: Monte Carlo approaches are rigorous but lack scalability, whereas neural variational inference methods are efficient but rely on Gaussian assumptions that fail to capture discontinuities. We address this tension by introducing a neural exponential tilting framework for variational inference in Lévy‑driven SDEs. Our approach constructs a flexible variational family by exponentially reweighting the Lévy measure using neural networks. This parametrization preserves the jump structure of the underlying process while remaining computationally tractable. To enable efficient inference, we develop a quadratic neural parametrization that yields closed‑form normalization of the tilted measure, a conditional Gaussian representation for stable processes that facilitates simulation, and symmetry‑aware Monte Carlo estimators for scalable optimization. Empirically, we demonstrate that the method accurately captures jump dynamics and yields reliable posterior inference in regimes where Gaussian‑based variational approaches fail, on both synthetic and real‑world datasets.
Authors:Dong-Yang Li, Wang Zhao, Yuxin Chen, Wenbo Hu, Meng-Hao Guo, Fang-Lue Zhang, Ying Shan, Shi-Min Hu
Abstract:
Recent advances in 3D generative models have rapidly improved image‑to‑3D synthesis quality, enabling higher‑resolution geometry and more realistic appearance. Yet fidelity, which measures pixel‑level faithfulness of the generated 3D asset to the input image, still remains a central bottleneck. We argue this stems from an implicit 2D‑3D correspondence issue: most 3D‑native generators synthesize shape in canonical space and inject image cues via attention, leaving pixel‑to‑3D associations ambiguous. To tackle this issue, we draw inspiration from 3D reconstruction and propose Pixal3D, a pixel‑aligned 3D generation paradigm for high‑fidelity 3D asset creation from images. Instead of generating in a canonical pose, Pixal3D directly generates 3D in a pixel‑aligned way, consistent with the input view. To enable this, we introduce a pixel back‑projection conditioning scheme that explicitly lifts multi‑scale image features into a 3D feature volume, establishing direct pixel‑to‑3D correspondence without ambiguity. We show that Pixal3D is not only scalable and capable of producing high‑quality 3D assets, but also substantially improves fidelity, approaching the fidelity level of reconstruction. Furthermore, Pixal3D naturally extends to multi‑view generation by aggregating back‑projected feature volumes across views. Finally, we show pixel‑aligned generation benefits scene synthesis, and present a modular pipeline that produces high‑fidelity, object‑separated 3D scenes from images. Pixal3D for the first time demonstrates 3D‑native pixel‑aligned generation at scale, and provides a new inspiring way towards high‑fidelity 3D generation of object or scene from single or multi‑view images. Project page: https://ldyang694.github.io/projects/pixal3d/
Authors:Chang Liu, Haoning Wu, Weidi Xie
Abstract:
Open‑world object counting remains brittle: despite rapid advances in vision‑language models (VLMs), reliably counting the objects a user intends is far from solved. We argue that a central reason is that counting granularity is left implicit; users may refer to a specific identity, an attribute, an instance type, a category, or an abstract concept, yet most methods treat "what to count" as a single, category‑level matching problem. In this work, we redefine open‑world counting as multi‑grained counting, where visual exemplars specify target appearance and fine‑grained text, with optional negative prompts, specifies the intended semantic granularity across five explicit levels. Making granularity explicit, however, exposes a critical data bottleneck: existing counting datasets lack the multi‑category scenes, controlled distractors, and instance‑level annotations needed to verify fine‑grained prompt semantics. To address this, we propose the first fully automatic data‑scaling pipeline that integrates controllable 3D synthesis with consistent image editing and VLM‑based filtering, and use it to construct KubriCount, the largest and most comprehensively annotated counting dataset to date, supporting both training and multi‑grained evaluation. Systematic benchmarking reveals that both multimodal large language models and specialist counting models exhibit severe prompt‑following failures under fine‑grained distinctions. Motivated by these findings, we train HieraCount, a multi‑grained counting model that jointly leverages text and visual exemplars as complementary target specifications. HieraCount substantially improves multi‑grained counting accuracy and generalizes robustly to challenging real‑world scenarios. The project page is available here: https://verg‑avesta.github.io/KubriCount/.
Authors:Wei Chow, Linfeng Li, Xian Sun, Lingdong Kong, Zefeng Li, Qi Xu, Hang Song, Tian Ye, Xian Wang, Jinbin Bai, Shilin Xu, Xiangtai Li, Junting Pan, Shaoteng Liu, Ran Zhou, Tianshu Yang, Songhua Liu
Abstract:
Diffusion models dominate image editing, yet their global denoising mechanism entangles edited regions with surrounding context, causing modifications to propagate into areas that should remain intact. We propose a fundamentally different approach by leveraging Masked Generative Transformers (MGTs), whose localized token‑prediction paradigm naturally confines changes to intended regions. We present EditMGT, an MGT‑based editing framework that is the first of its kind. Our approach employs multi‑layer attention consolidation to aggregate cross‑attention maps into precise edit localization signals, and region‑hold sampling to explicitly prevent token flipping in non‑target areas. To support training, we construct CrispEdit‑2M, a 2M‑sample high‑resolution (>1024) editing dataset spanning seven categories. With only 960M parameters, EditMGT achieves state‑of‑the‑art image similarity on multiple benchmarks while delivering 6x faster editing, demonstrating that MGTs offer a compelling alternative to diffusion‑based editing.
Authors:Lingdong Kong, Ao Liang, Tianyi Yan, Hongsi Liu, Wesley Yang, Ziqi Huang, Xian Sun, Wei Yin, Jialong Zuo, Yixuan Hu, Dekai Zhu, Dongyue Lu, Youquan Liu, Guangfeng Jiang, Linfeng Li, Xiangtai Li, Long Zhuo, Lai Xing Ng, Benoit R. Cottereau, Changxin Gao, Liang Pan, Wei Tsang Ooi, Ziwei Liu
Abstract:
Today's driving world models can generate remarkably realistic dash‑cam videos, yet no single model excels universally. Some generate photorealistic textures but violate basic physics; others maintain geometric consistency but fail when subjected to closed‑loop planning. This disconnect exposes a critical gap: the field evaluates how real generated worlds appear, but rarely whether they behave realistically. We introduce WorldLens, a unified benchmark that measures world‑model fidelity across the full spectrum, from pixel quality and 4D geometry to closed‑loop driving and human perceptual alignment, through five complementary aspects and 24 standardized dimensions. Our evaluation of six representative models reveals that no existing approach dominates across all axes: texture‑rich models violate geometry, geometry‑aware models lack behavioral fidelity, and even the strongest performers achieve only 2‑3 out of 10 on human realism ratings. To bridge algorithmic metrics with human perception, we further contribute WorldLens‑26K, a 26,808‑entry human‑annotated preference dataset pairing numerical scores with textual rationales, and WorldLens‑Agent, a vision‑language evaluator distilled from these judgments that enables scalable, explainable auto‑assessment. Together, the benchmark, dataset, and agent form a unified ecosystem for assessing generated worlds not merely by visual appeal, but by physical and behavioral fidelity.
Authors:Xiran Zhao, Jing Jin, Yan Bai, Zhongan Wang, Yifeng Sun, Yihang Lou, Xuanyu Zhu, Tao Feng, Yingna Wu
Abstract:
Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which do not fully reflect continuous inspection processes in real industrial scenarios. We introduce MMVIAD (Multi‑view Multi‑task Video Industrial Anomaly Detection), to the best of our knowledge the first continuous multi‑view video dataset for industrial anomaly detection and understanding, together with a benchmark for multi‑task evaluation. MMVIAD contains object‑centric 2‑second inspection clips with approximately 120 degrees of camera motion, covering 48 object categories, 14 environments, and 6 structural anomaly types. It supports anomaly detection, defect classification, object classification, and anomaly visible‑time localization. Systematic evaluations on MMVIAD show that current commercial and open‑source video MLLMs remain far below human performance, especially for fine‑grained defect recognition and temporal grounding. To improve transferable anomaly understanding, we further develop a two‑stage post‑training pipeline where PS‑SFT (Perception‑Structured Supervised Fine‑Tuning) initializes perception‑structured reasoning and VISTA‑GRPO (Visibility‑grounded Industrial Structured Temporal Anomaly Group Relative Policy Optimization) refines the model with semantic‑gated defect reward and visibility‑aware temporal reward, producing the final model VISTA. On MMVIAD‑Unseen, VISTA improves the base model's average score across the four tasks from 45.0 to 57.5, surpassing GPT‑5.4. Source code is available at https://github.com/Georgekeepmoving/MMVIAD.
Authors:Juyi Lin, Arash Akbari, Yumei He, Lin Zhao, Haichao Zhang, Arman Akbari, Xingchen Xu, Zoe Y. Lu, Enfu Nan, Hokin Deng, Edmund Yeh, Sarah Ostadabbas, Yun Fu, Jennifer Dy, Pu Zhao, Yanzhi Wang
Abstract:
Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real‑world dynamics. However, evaluating whether generated videos actually follow these rules remains challenging. Existing physics‑focused video benchmarks have made important progress, but they still face three key challenges, including the coarse evaluation frameworks that hide law‑specific failures, response biases and fatigue that undermine the validity of annotation judgments, and automated evaluators that are insufficiently physics‑aware or difficult to audit. To address those challenges, we introduce PhyGround, a criteria‑grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid‑body mechanics, fluid dynamics, and optics. Each law is operationalized through observable sub‑questions to enable per‑law diagnostics. We evaluate eight modern video generation models through a large‑scale, quality‑controlled human study, grounded on social science lab experiment design. A total of 459 annotators provided 5,796 complete annotations and over 37.4K fine‑grained labels; after quality control, the retained annotations exhibited high split‑half model‑ranking correlations (Spearman's rho > 0.90). To support reproducible automated evaluation, we release PhyJudge‑9B, an open physics‑specialized VLM judge. PhyJudge‑9B achieves substantially lower aggregate relative bias than Gemini‑3.1‑Pro (3.3% vs. 16.6%). We release prompts, human annotations, model checkpoints, and evaluation code on the project page https://phyground.github.io/.
Authors:Yifeng Yang, Jubo Feng, Jing Xu, Xinbing Wang, Qinying Gu, Nanyang Ye
Abstract:
Vision‑language models enable OOD detection by comparing image alignment with ID labels and negative semantics. Existing negative‑label‑based methods mainly rely on static negative labels constructed before inference, limiting their ability to cover diverse and evolving OOD concepts. Although test‑time expansion provides a natural solution, naively learning negative semantics from potential OOD samples may introduce hard ID contamination. To address this issue, we propose a Test‑time ID‑prototype‑separated Negative Semantics learning method, termed TINS. TINS learns sample‑specific negative text embeddings via image‑to‑text modality inversion and introduces ID‑prototype‑separated regularization to keep them separated from ID semantics. To further stabilize negative semantics expansion, TINS employs group‑wise aggregation scoring and a buffer update strategy. Extensive experiments across Four‑OOD, OpenOOD, Temporal‑shift, and Various ID settings show consistent improvements over strong baselines. Notably, on the Four‑OOD benchmark with ImageNet‑1K as ID, TINS reduces the average FPR95 from 14.04% to 6.72%. Our code is available at https://github.com/zxk1212/tins.
Authors:Kaicong Huang, Weiheng Oh, Thomas Guggisberg, Ruimin Ke
Abstract:
Automated transit payment analysis is vital for scalable fare auditing and passenger analytics, yet practice still relies on limited manual inspection. Prior vision‑ and skeleton‑based methods remain brittle under noisy onboard surveillance and often depend on poorly generalizable handcrafted features. Building on the success of graph convolutional networks in human action recognition, we observe that skeleton features excel at modeling global spatiotemporal dependencies but tend to underemphasize the subtle local relative motions that distinguish payment actions. In contrast, RGB features preserve fine‑grained spatial details yet often lack reliable temporal continuity in surveillance footage. To bridge both system‑level deployment needs and model‑level design challenges, we present iPay, an integrated payment action recognition framework for onboard transit surveillance system. iPay adopts a multimodal mixture‑of‑experts architecture with four tightly coupled streams: (1) an RGB expert stream emphasizing local evidence via region‑focused computation; (2) a skeleton expert stream modeling articulated motion with a graph convolutional backbone; (3) a dual‑attention fusion stream enabling skeleton‑to‑RGB temporal transfer and RGB‑to‑skeleton spatial enhancement; and (4) a prior‑driven Spatial Difference Discriminator (SDD) that explicitly models hand‑to‑anchor relative motion to improve task‑specific discriminability. We also collaborate with local transit agencies to collect over 55 hours of real onboard surveillance footage, yielding 500+ payment clips. Experiments show that iPay outperforms prior methods and achieves 83.45% recognition accuracy with competitive computational efficiency, making it suitable for edge deployment. Code is available at https://github.com/ccoopq/iPay.
Authors:Yangneng Chen, Junlin Li, Weijun Yao, Xilai Ma, Guodong Du, Wenya Wang, Jing Li
Abstract:
Large Vision‑Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet their reliability is persistently undermined by hallucinations‑generating text that contradicts visual input. Recent studies often attribute these errors to inadequate visual attention. In this work, we analyze the attention mechanisms via the logit lens, uncovering a distinct anomaly we term Vocabulary Hijacking. We discover that specific visual tokens, defined as Inert Tokens, disproportionately attract attention. Crucially, when their intermediate hidden states are projected into the vocabulary space, they consistently decode to a fixed set of unrelated words (termed Hijacking Anchors) across layers, revealing a rigid semantic collapse. Leveraging this semantic rigidity, we propose Hijacking Anchor‑Based Identification (HABI), a robust strategy to accurately localize these Inert Tokens. To quantify the impact of this phenomenon, we introduce the Non‑Hijacked Visual Attention Ratio (NHAR), a novel metric designed to identify attention heads that remain resilient to hijacking and are critical for factual accuracy. Building on these insights, we propose Hijacking‑Aware Visual Attention Enhancement (HAVAE), a training‑free intervention that selectively strengthens the focus of these identified heads on salient visual content. Extensive experiments across multiple benchmarks demonstrate that HAVAE significantly mitigates hallucinations with no additional computational overhead, while preserving the model's general capabilities. Our code is publicly available at https://github.com/lab‑klc/HAVAE.
Authors:Hongyou Zhou, Marc Toussaint, Ling Shao, Zihan Ye
Abstract:
Despite strong zero‑shot performance, SAM is unreliable under domain shift due to Mask‑level Confidence Confusion (MCC), where a single IoU‑based mask score fails to reflect pixel‑wise reliability near boundaries. Motivated by the contrast between texture‑biased shortcuts in neural networks and shape‑centric processing in human vision, we model out‑of‑domain variation as appearance shifts and non‑rigid deformations that jointly stress calibration. We propose Segment Anything with Robust Uncertainty‑Accuracy Correlation (RUAC) for robust pixel‑wise uncertainty estimation under appearance and deformation shifts. RUAC adds a lightweight uncertainty head, trains it with a collaborative style‑deformation attack that jointly perturbs texture and geometry, and applies Uncertainty‑Accuracy Alignment to ensure uncertainty consistently highlights erroneous pixels even under adversarial perturbations. Across 23 zero‑shot domains, RUAC improves segmentation quality and yields more faithful uncertainty with stronger uncertainty‑accuracy correlation. Project page: https://hongyouzhou.github.io/ruac/.
Authors:Guoquan Wei, Liu Shi, Chong Chen, Qiegen Liu
Abstract:
Despite extensive research on computed tomography (CT) denoising, few studies exploit projection‑domain data characteristics to mitigate noise correlation. To bridge this gap, this work proposes FrequencyCT, the first zero‑shot self‑supervised method for pseudo‑sample generation in the frequency domain for low‑dose CT denoising. Specifically, by exploiting the distinct frequency‑domain distributions of noise and true signal, a regional low‑frequency anchoring technique is proposed. Applying phase‑preserving noise and mask perturbations to the high‑frequency region generates pseudo‑samples for self‑supervision. Driven by the exponential correlation between noise variance of noisy projections and the underlying true signal, consistent data truncation is applied to the generated samples to stabilize optimization gradients. Evaluation results on multiple public and real datasets confirm the clinical application potential of this research, which provides an innovative perspective for the field of denoising. The code is available at: https://github.com/yqx7150/FrequencyCT.
Authors:Chen Zhong, Xiao An, Jiaxing Sun, Zihan Gui, Guangyi Yang, Wei He
Abstract:
Low‑level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpretable scalar scores rather than characterizing physics‑driven RS degradations, deviating markedly from the diagnostic needs of RS experts. While Vision‑Language Models (VLMs) present a compelling alternative by delivering language‑grounded IQA, their visual priors are heavily biased toward ground‑level natural images. Consequently, whether VLMs can overcome this domain gap to perceive and articulate RS artifacts remains insufficiently studied. To bridge this gap, we propose SenseBench, the first dedicated diagnostic benchmark for RS low‑level visual perception and description. Driven by a physics‑based hierarchical taxonomy that unifies both non‑reference and reference‑based paradigms, SenseBench features over 10K meticulously curated instances across 6 major and 22 fine‑grained RS degradation categories. Specifically, two complementary protocols are designed for evaluation: objective low‑level visual perception and subjective diagnostic description. Comprehensive evaluation of 29 state‑of‑the‑art VLMs reveals not only skewed domain priors and multi‑distortion collapse, but also fluency illusion and a perception‑description inversion effect. We hope SenseBench provides a robust evaluation testbed and high‑quality diagnostic data to advance the development of VLMs in RS low‑level perception. Code and datasets are available \hrefhttps://github.com/Zhong‑Chenchen/SenseBench\textcolorbluehere.
Authors:Lingjun Zhang, Changjie Wu, Linzhe Shi, Jiangyang Li, Jiaxin Liu, Lei Yang, Hang Zhang, Mu Xu, Hong Wang
Abstract:
End‑to‑end autonomous driving systems are increasingly integrating Vision‑Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptations from general domains, lacking in‑depth exploration tailored to autonomous driving scenarios, particularly within visual reasoning modules. In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird's‑eye‑view (BEV) space, thereby enabling long‑horizon modeling of future world states. We also introduce an efficient and adaptive text reasoning mechanism that utilizes additional social knowledge and reasoning capabilities to further improve driving performance in challenging long‑tail scenarios. We present a novel, efficient, and effective approach that achieves state‑of‑the‑art (SOTA) results on the closed‑loop Bench2drive benchmark. Codes are available at: https://github.com/hotdogcheesewhite/DeepSight.
Authors:Zhilei Shu, Shangwen Zhu, Zihang Liang, Xiaofan Li, Qianyu Peng, Xinyu Cui, Bo Ye, Yiming Li, Fan Cheng, Jian Zhao, Yang Cao, Zheng-Jun Zha, Ruili Feng
Abstract:
Director‑style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events ‑‑ a regime in which 68% of general clips and over 99% of robotics/gameplay clips contain overlapping events, yet existing multi‑event generators rest on a single‑active‑prompt assumption. However, modern video generators, such as Diffusion Transformers (DiT), represent time as discrete points through point‑wise positional encodings. This formulation creates a fundamental dimension mismatch: temporally extended intervals and overlapping events are mathematically unrepresentable to the attention mechanism. In this paper, we propose Time Interval Encoding (TIE), a principled, plug‑and‑play interval‑aware generalization of rotary embeddings that elevates time intervals to first‑class primitives inside DiT cross‑attention. Rather than introducing another heuristic interval embedding, we show that, within RoPE‑compatible bilinear attention, TIE is characterized by two basic principles: Temporal Integrability, which requires an event to aggregate positional evidence over its full duration, and Duration Invariance, which removes the trivial bias toward longer intervals. Under a uniform kernel, this characterization yields an efficient closed‑form sinc‑based solution that preserves the standard attention interface and naturally attenuates boundary noise through interval integration. Empirically, TIE preserves the visual quality of the base DiT model while substantially improving temporal controllability. In our experiments on the OmniEvents dataset, it improves human‑verified Temporal Constraint Satisfaction Rate from 77.34% to 96.03% and reduces temporal boundary error from 0.261s to 0.073s, while also improving trajectory‑level temporal alignment metrics. The code and dataset are available at https://github.com/MatrixTeam‑AI/TIE.
Authors:Yuecheng Liu, Junda Cheng, Longliang Liu, Wenjing Liao, Hanrui Cheng, Yuzhou Wang, Xin Yang
Abstract:
Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine‑detail regions and temporal inconsistencies. We argue that current approaches, which primarily rely on temporal smoothing via Transformers, struggle to maintain strict 3D geometric consistency‑particularly under rotations or drastic view changes. To address this, we propose GemDepth, a framework built on the insight that an explicit awareness of camera motion and global 3D structure is a prerequisite for 3D consistency. Distinctively, GemDepth introduces a Geometry‑Embedding Module (GEM) that predicts inter‑frame camera poses to generate implicit geometric embeddings. This injection of motion priors equips the network with intrinsic 3D perception and alignment capabilities. Guided by these geometric cues, our Alternating Spatio‑Temporal Transformer (ASTT) captures latent point‑level correspondences to simultaneously enhance spatial precision for sharp details and enforce rigorous temporal consistency. Furthermore, GemDepth employs a data‑efficient training strategy, effectively bridging the gap between high efficiency and robust geometric consistency. As shown in Fig.2, comprehensive evaluations demonstrate that GemDepth achieves state‑of‑the‑art performance across multiple datasets, particularly in complex dynamic scenarios. The code is publicly available at: https://github.com/Yuecheng919/GemDepth.
Authors:Weiqi Yan, Lixin Chen, Xiangrui Hou, Zhipeng Cai, Youbiao Wang, Yangyang Shi, Yu Zang, Cheng Wang
Abstract:
Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion‑on‑motion regime, ego‑motion activates background edges across buildings, vegetation, and horizon structures, while the UAV may appear as a sparse event cluster. Unlike static‑ or ground‑observer event‑based UAV detection, onboard UAV‑view detection breaks the clean‑background assumption because sensor ego‑motion can activate dense background events over the entire field of view. To explore this practical problem, we present M^2E‑UAV, to the best of our knowledge, the first onboard UAV‑view motion‑on‑motion event‑based dataset and benchmark for tiny UAV detection, where both the sensing platform and the target UAV are moving. M^2E‑UAV provides synchronized event streams and IMU measurements collected from an onboard sensing platform, together with event‑level UAV foreground labels derived from temporally propagated 10 Hz bounding‑box annotations. The processed benchmark contains 87,223 training samples and 21,395 validation samples across four scene families: sunny building‑forest, sunny farm‑village, sunset building‑forest, and sunset farm‑village. We define a train/validation split and an evaluation protocol for comparing representative existing baselines across event‑frame, voxel‑grid, and point‑set representations, with optional IMU input. The benchmark results show that existing baselines remain limited under sparse tiny‑target evidence and dense ego‑motion‑induced background events. Code and benchmark files will be released at https://github.com/Wickyan/M2E‑UAV.
Authors:Zijun Shen, Sihan Yang, Ruichuan An, Ziyu Guo, Hao Liang, Ming Lu, Renrui Zhang, Wentao Zhang
Abstract:
Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works largely rely on implicit token‑level alignment via supervised fine‑tuning, which fails to fully capture the potential synergy between comprehension and creation. In this work, we propose Sync‑R1, an end‑to‑end reinforcement learning framework that jointly optimizes personalized understanding and generation within a single, explicit reasoning loop. Through this unified feedback process, Sync‑R1 enables personalized comprehension to guide content creation, while the resulting generation quality reciprocally refines understanding within an integrated reward landscape. To efficiently orchestrate this dual‑task synergy, we introduce Sync‑GRPO, a reinforcement learning method utilizing an ensemble reward system. Furthermore, we propose Dynamic Group Scaling (DGS), which adaptively filters low‑potential trajectories to reduce gradient variance and accelerate convergence. To better reflect real‑world complexity, we introduce UnifyBench++, featuring denser textual descriptions and richer user contexts. Experimental results demonstrate that Sync‑R1 achieves state‑of‑the‑art performance, showcasing superior cross‑task reasoning and robust personalization without requiring complex cold‑start procedures. The code and the UnifyBench++ dataset will be released at: https://github.com/arctanxarc/UniCTokens.
Authors:Keming Wu, Yijing Cui, Wenhan Xue, Qijie Wang, Xuan Luo, Zhiyuan Feng, Zuhao Yang, Sudong Wang, Sicong Jiang, Haowei Zhu, Zihan Wang, Ping Nie, Wenhu Chen, Bin Wang
Abstract:
Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world‑state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground‑truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human‑aligned two‑part methodology: Process‑aware Reasoning Verification uses structured QA and reasoning‑phase diagnostics to detect temporal and causal failures, while Multi‑dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert‑annotated pairs over 1.4K videos, supporting pair‑wise and point‑wise reward‑model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world‑aware video generation at https://github.com/UniX‑AI‑Lab/WorldReasonBench/.
Authors:Minqing Huang, Yujiao Xiang, Zihan Liang, Jiajie Huang, Jingqi Wang, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang, Gong Che
Abstract:
Vision‑Language‑Action (VLA) models have emerged as a promising paradigm for end‑to‑end autonomous driving. However, existing reasoning mechanisms still struggle to provide planning‑oriented intermediate representations: textual Chain‑of‑Thought (CoT) fails to preserve continuous spatiotemporal structure, while latent world reasoning remains difficult to use as a direct condition for action generation. In this paper, we propose CoWorld‑VLA, a multi‑expert world reasoning framework for autonomous driving, where world representations serve as explicit conditions to guide action planning. CoWorld‑VLA extracts complementary world information through multi‑source supervision and encodes it into expert tokens within the VLA, thereby providing planner‑accessible conditioning signals. Specifically, we construct four types of tokens: semantic interaction, geometric structure, dynamic evolution, and ego trajectory tokens, which respectively model interaction intent, spatial structure, future temporal dynamics, and behavioral goals. During action generation, CoWorld‑VLA employs a diffusion‑based hierarchical multi‑expert fusion planner, which is coupled with scene context throughout the joint denoising process to generate continuous ego trajectories. Experiments show that CoWorld‑VLA achieves competitive results in both future scene generation and planning on the NAVSIM v1 benchmark, demonstrating strong performance in collision avoidance and trajectory accuracy. Ablation studies further validate the complementarity of expert tokens and their effectiveness as planning conditions for action generation. Code will be available at https://github.com/AFARI‑Research/CoWorld‑VLA.
Authors:Xi Jiang, Yinjie Zhao, Zesheng Yang, Feng Zheng
Abstract:
Visual anomaly detection (VAD) is crucial in many real‑world fields, such as industrial inspection, medical imaging, infrastructure monitoring, and remote sensing. However, the specific anomaly definitions, data modalities, and annotation standards across different domains make it difficult to transfer single‑domain trained VAD models. Vision‑language models (VLMs), pre‑trained on large‑scale cross‑domain data, can perform visual perception under task instructions, offering a promising solution for cross‑domain VAD. However, single‑inference VLM judgments are unreliable, since they rely more on prior knowledge than on normal‑sample references or fine‑grained feature evidence. We therefore present AnomalyClaw, a training‑free VAD agent that turns anomaly judgment into a multi‑round refutation process. In each round, the agent proposes candidate anomalies and refutes each against normal‑sample references, drawing on a 13‑tool library for visual verification, reference parsing, and frozen expert probing. On the CrossDomainVAD‑12 benchmark (12 datasets), AnomalyClaw achieves consistent macro‑AUROC improvements over single‑step direct inference with +6.23 pp on GPT‑5.5, +7.93 pp on Seed2.0‑lite, and +3.52 pp on Qwen3.5‑VL‑27B. We further introduce an optional verbalized self‑evolution extension. It builds an online rulebook from internal‑branch disagreement without oracle labels. On Qwen3.5‑VL‑27B, it delivers a +2.09 pp mean gain, comparable to a K = 10 oracle‑label supervised baseline (+1.99 pp). These results show that agentic refutation improve anomaly understanding and reasoning of VLMs, rather than merely aggregating tool outputs.
Authors:Mingwei Xing, Xinliang Wang, Yifeng Shi
Abstract:
This work explores a simple yet powerful lightweight adapter design for feed‑forward 3D Gaussian Splatting (3DGS). Existing methods typically apply complex, architecture‑specific designs on top of the generic pipeline of image feature extraction \rightarrow multi‑view interaction \rightarrow feature decoding. However, constrained by the scale bottleneck of 3D training data and the low‑pass filtering effect of deep networks, these methods still fall short in cross‑domain generalization and high‑frequency geometric fidelity. To address these problems, we propose AdaptSplat, which demonstrates that without complex component engineering, introducing a single adapter of only 1.5M parameters into the generic architecture is sufficient to achieve superior performance. Specifically, we design a lightweight Frequency‑Preserving Adapter (FPA) that extracts direction‑aware high‑frequency structural priors from the shallow features of a powerful vision foundation model backbone, and seamlessly integrates them into the generic pipeline via high‑frequency positional encodings and adaptive residual modulation. This effectively compensates for the high‑frequency attenuation caused by over‑smoothing in deep features, improving the fitting accuracy of Gaussian primitives on complex surfaces and sharp boundaries. Extensive experiments demonstrate that AdaptSplat achieves state‑of‑the‑art feed‑forward reconstruction performance on multiple standard benchmarks, with stable generalization across domains. Code available at: https://github.com/xmw666/AdaptSplat.
Authors:Xiaobin Hu, Enpu Zuo, Lanping Hu, Kaiwen Yang, Dianshu Liao, Tianyi Zhang, Bo Yin, Yinsi Zhou, Shidong Pan, Xiaoyu Sun
Abstract:
Privacy protection has become a critical requirement in the era of ubiquitous visual data sharing, imposing higher demands on efficient and robust privacy detection algorithms. However, current robust detection models are severely hindered by the lack of comprehensive datasets. Existing privacy‑oriented datasets often suffer from limited scale, coarse‑grained annotations, and narrow domain coverage, failing to capture the intricate details of sensitive information in realworld environments. To bridge this gap, we present a large‑scale, fine‑grained Visual Privacy Dataset (VPD‑100K), designed to facilitate generalized privacy detection. We establish a holistic taxonomy comprising four primary domains: Human Presence, On‑Screen Personally Identifiable Information (PII), Physical Identifiers, and Location Indicators, containing 100,000 images annotated with 33 fine‑grained classes and over 190,000 object instances. Statistical analysis reveals that our dataset features long‑tailed distributions, small object scales, and high visual complexity. These characteristics make the dataset particularly valuable for demanding, unconstrained applications such as live streaming, where actors frequently face unintentional, realtime information leakage. Furthermore, we design an effective frequency‑enhanced lightweight module consisting of frequency‑domain attention fusion and adaptive spectral gating mechanism that breaks the limitations of spatial pixel intensity to better capture the subtle details of sensitive information. Extensive experiments conducted on both diverse image and streaming videos benchmarks consistently demonstrate the effectiveness of our VPD‑100K dataset and the wellcurated frequency mechanism. The code and dataset are available at https://vpd‑100k.github.io/.
Authors:Federico Pizzolato, Francesco Pasti, Nicola Bellotto
Abstract:
Terrain segmentation is a fundamental capability for autonomous mobile robots operating in unstructured outdoor environments. However, state‑of‑the‑art models are incompatible with the memory and compute constraints typical of microcontrollers, limiting scalable deployment in small robotics platforms. To address this gap, we develop a complete framework for robust binary terrain segmentation on a low‑cost microcontroller. At the core of our approach we design Nano‑U, a highly compact binary segmentation network with a few thousand parameters. To compensate for the network's minimal capacity, we train Nano‑U via Quantization‑Aware Distillation (QAD), combining knowledge distillation and quantization‑aware training. This allows the final quantized model to achieve excellent results on the Botanic Garden dataset and to perform very well on TinyAgri, a custom agricultural field dataset with more challenging scenes. We deploy the quantized Nano‑U on a commodity microcontroller by extending MicroFlow, a compiler‑based inference engine for TinyML implemented in Rust. By eliminating interpreter overhead and dynamic memory allocation, the quantized model executes on an ESP32‑S3 with a minimal memory footprint and low latency. This compiler‑based execution demonstrates a viable and energy‑efficient solution for perception on low‑cost robotic platforms.
Authors:Soichiro Okazaki, Tatsuya Sasaki, Hiroki Ohashi
Abstract:
Open‑vocabulary object detection (OVOD) aims to detect both seen and unseen categories, yet existing methods often struggle to generalize to novel objects due to limited integration of global and local contextual cues. We propose DetRefiner, a simple yet effective plug‑and‑play framework that learns to fuse global and local features to refine open‑vocabulary detection. DetRefiner processes global image features and patch‑level image features from foundational models (e.g., DINOv3) through a lightweight Transformer encoder. The encoder produces a class vector capturing image‑level attributes and patch vectors representing local region attributes, from which attribute reliability is inferred to recalibrate the base model's confidence. Notably, DetRefiner is trained independently of the base OVOD model, requiring neither access to its internal features nor retraining. At inference, it operates solely on the base detector's predictions, producing auxiliary calibration scores that are merged with the base detector's scores to yield the final refined confidence. Despite this simplicity, DetRefiner consistently enhances multiple OVOD models across COCO, LVIS, ODinW13, and Pascal VOC, achieving gains of up to +10.1 AP on novel categories. These results highlight that learning to fuse global and local representations offers a powerful and general mechanism for advancing open‑world object detection. Our codes and models are available at https://github.com/hitachi‑rd‑cv/detrefiner.
Authors:Longteng Guo, Xuanxu Lin, Dongze Hao, Tongtian Yue, Pengkang Huo, Jiatong Ma, Yuchen Liu, Jing Liu
Abstract:
Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi‑step inference across various subjects. Existing benchmarks for multimodal large language models (MLLMs) often fail to capture the complexity and traceability of reasoning processes necessary for rigorous evaluation. To fill this gap, we introduce SciVQR, a multimodal benchmark covering 54 subfields in mathematics, physics, chemistry, geography, astronomy, and biology. SciVQR includes domain‑specific visuals, such as equations, charts, and diagrams, and challenges models to combine visual comprehension with reasoning. The tasks range from basic factual recall to complex, multi‑step inferences, with 46% including expert‑authored solutions. SciVQR not only evaluates final answers but also examines the reasoning process, providing insights into how models reach their conclusions. Our evaluation of leading MLLMs, including both proprietary and open‑source models, reveals significant limitations in handling complex multimodal reasoning tasks, underscoring the need for improved multi‑step reasoning and better integration of interdisciplinary knowledge in advancing MLLMs toward true scientific intelligence. The dataset and evaluation code are publicly available at https://github.com/CASIA‑IVA‑Lab/SciVQR.
Authors:Yeo Keat Ee, Debaditya Roy, Chen Li, Hao Zhang, Basura Fernando
Abstract:
Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challenges such as action variability, ambiguous boundaries, and high annotation costs remain, especially in new or low‑resource domains. Grammar‑based approaches improve segmentation with structural priors but rely on complex parsing limiting scalability. In this work, we propose a lightweight, constraint‑based refinement framework that enhances TAS predictions by integrating statistical structural priors such as transition confidence, action boundary sets, and per‑class duration, that can be directly extracted from annotated data. These constraints are integrated into a modified Viterbi decoding algorithm, allowing inference‑time refinement without retraining or added model complexity. Our approach improves both fully and semi‑supervised TAS models by correcting structural prediction errors while maintaining high efficiency. Code is available at https://github.com/LUNAProject22/CAD
Authors:Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu, Wen-Kai Kuo, Jun-Wei Hsieh
Abstract:
The Vision Transformer (ViT) achieves remarkable accuracy across visual tasks but remains computationally expensive for edge deployment. This paper presents MicroViTv2, a lightweight Vision Transformer optimized for real‑device efficiency. Built upon the original MicroViT, the proposed model is designed based on reparameterized design, specifically Reparameterized Patch Embedding (RepEmbed) and Reparameterized Depth‑Wise convolution mixer (RepDW) for faster inference, and introduces the Single Depth‑Wise Transposed Attention (SDTA) to capture long‑range dependencies with minimal redundancy. Despite slightly higher FLOPs, MicroViTv2 improves accuracy up to 0.5% compared to its predecessor and surpassing MobileViTv2, EdgeNeXt, and EfficientViT while maintaining fast inference and high energy efficiency on Jetson AGX Orin. Experiments on ImageNet‑1K and COCO demonstrate that hardware‑aware design and structural re‑parameterization are key to achieving high accuracy and low energy consumption, validating the need to evaluate efficiency beyond FLOPs. Code is available at https://github.com/novendrastywn/MicroViT.
Authors:Manyu Li, Ruian He, Chenxi Ma, Weimin Tan, Bo Yan
Abstract:
Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as microscopy remains limited by the scarcity of domain‑specific training data and the difficulty of encoding fine‑grained expert knowledge into model parameters. To bridge the gap, we introduce MicroWorld, a framework that constructs a multimodal attributed property graph (MAPG) from large‑scale scientific image‑‑caption corpora and leverages it to augment MLLM reasoning at inference time without any domain‑specific fine‑tuning. MicroWorld extracts biomedical entities and relations via scispaCy or LLM‑based triplet mining, aligns images and entities in a shared embedding space using Qwen3‑VL‑Embedding, and assembles a knowledge graph comprising approximately 111K nodes and 346K typed edges spanning eight relation categories. At inference time, a graph‑augmented retrieval pipeline matches query entities to the MAPG and injects structured knowledge context into the MLLM prompt. On the MicroVQA benchmark, MicroWorld improves the reasoning performance of Qwen3‑VL‑8B‑Instruct by 37.5%, outperforming GPT‑5 by 13.0% to achieve a new state‑of‑the‑art. Furthermore, it yields a 6.0% performance gain on the MicroBench benchmark. Extensive experiments demonstrate the enhanced generalization capability introduced by MicroWorld. A qualitative case study further reveals both the mechanisms through which structured knowledge improves reasoning and the failure modes that point to promising future directions. Code and data are available at https://github.com/ieellee/MicroWorld.
Authors:Xiangkai Wang, Yun Zhao, Dongyi He, Qingling Xia, Gen Li, Xinlai Xing, Yuchi Pan, Bin Jiang
Abstract:
Motor imagery electroencephalography (MI‑EEG) decoding offers a non‑invasive route for post‑stroke rehabilitation, but cross‑patient use remains difficult because pathological neural reorganization changes task‑related EEG dynamics, aperiodic activity, local excitability, cross‑regional coordination, and trial‑level brain‑state context. This makes source‑learned MI representations unreliable for unseen patients. To address this problem, we propose CFSPMNet, a cross‑patient adaptation framework that models post‑stroke MI‑EEG as latent neural‑state organization. CFSPMNet combines a Fourier‑Reorganized State Mamba Network (FRSM) with Shared‑Private Prototype Matching (SPPM). FRSM represents each trial as a latent physiological token sequence, reorganizes token states in the Fourier domain, and uses Fourier‑derived trial context to guide Mamba state‑space propagation. SPPM improves target pseudo‑label updating by combining semantic confidence with shared‑private physiological consistency, filtering confident but physiologically inconsistent target predictions. Leave‑one‑subject‑out experiments on two stroke MI‑EEG datasets show that CFSPMNet outperforms representative CNN‑, Transformer‑, Mamba‑, and adaptation‑based baselines, achieving average accuracies of 68.23% on XW‑Stroke and 73.33% on 2019‑Stroke, with gains of 5.63 and 8.25 percentage points over the strongest competitors. Ablation, sensitivity, feature‑alignment, pseudo‑label selection, and neurophysiological visualization analyses further support the roles of Fourier‑domain token‑state reorganization and calibrated pseudo‑label updating. These results suggest that latent neural‑state modeling can improve rehabilitation‑oriented cross‑patient BCI decoding. Code is available at https://github.com/wxk1224/CFSPMNet.
Authors:Feihong Yan, Shaoyu Liu, Haixuan Wang, Shuai Lu, Linfeng Zhang, Huiqi Li, Xiangyang Ji
Abstract:
Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents direct generation at higher resolutions. Naively transferring training‑free extrapolation methods from LLMs or diffusion models to VAR yields three characteristic failure modes: global repetition, local repetition, and detail degradation. We trace them to a unified band‑stage mismatch: VAR generates images in a coarse‑to‑fine, scale‑wise process where each stage is driven by a distinct dominant RoPE frequency band, and each failure mode emerges when the dominant band of a particular stage is disrupted. Building on this insight, we propose Stage‑Aware RoPE Remapping, a training‑free strategy that assigns each frequency band a stage‑specific remapping rule, jointly suppressing all three failure modes. We further observe that attention becomes systematically dispersed as the image resolution increases. Existing methods typically depend on predefined attention scaling factors, which are neither adaptive to the target resolution nor capable of faithfully capturing the actual extent of attention dispersion. We therefore propose Entropy‑Driven Adaptive Attention Calibration, which quantifies dispersion via a resolution‑invariant normalized entropy and yields a closed‑form per‑head scaling factor that realigns the extrapolated‑resolution attention entropy with its training‑resolution counterpart. Extensive experiments show that our method consistently outperforms prior resolution‑extrapolation methods in both structural coherence and fine‑detail fidelity. Our code is available at https://github.com/feihongyan1/ExtraVAR.
Authors:Yeongtak Oh, Dongwook Lee, Sangkwon Park, Heeseung Kim, Sungroh Yoon
Abstract:
While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision‑language, with unified omnimodal benchmarking that jointly covers text, image, and audio still limited, and lacking the methodological rigor to account for absent‑persona scenarios or systematic grounding studies. We introduce Omni‑Persona, the first comprehensive benchmark for omnimodal personalization. We formalize the task as cross‑modal routing over the \emphPersona Modality Graph, encompassing 4 task groups and 18 fine‑grained tasks across ~750 items. To rigorously diagnose grounding behavior, we propose \emphCalibrated Accuracy (\mathrmCal), which jointly rewards correct grounding and appropriate abstention, incorporating absent‑persona queries within a unified evaluation framework. On our dedicated experiments, three diagnostic findings emerge: (i) open‑source models show a consistent audio‑vs‑visual grounding gap that RLVR partially narrows via dense rule‑based supervision; (ii) answerable recall and parameter scale are incomplete diagnostics, since strong recall can coexist with absent‑persona hallucination and larger models do not always achieve higher \mathrmCal, exposing calibration as a separate evaluation axis; and (iii) SFT is bounded by the difficulty of constructing annotated ground‑truth supervision at scale, while RLVR generalizes more consistently through outcome‑level verifiable feedback yet drifts toward conservative behavior and lower generation quality under our reward design. Omni‑Persona thus serves as a diagnostic framework that surfaces the pitfalls of omnimodal personalization, guiding future post‑training and reward design.
Authors:Yuna Lee, Kyoungho Min, Yulhwa Kim
Abstract:
Recent advancements in Vision‑Language Models (VLMs) enable large language models (LLMs) to process high‑resolution images, significantly improving real‑world multimodal understanding. However, this capability introduces a large number of vision tokens, resulting in substantial computational overhead. To mitigate this issue, various vision token pruning methods have been proposed. Nevertheless, existing approaches predominantly rely on learned semantic features within the model to capture visual redundancy. Moreover, they lack adaptive mechanisms to adjust pruning strategies according to the complexity of the input image. In this paper, we propose ERASE, a two‑stage vision token pruning framework that identifies and retains salient tokens through pruning strategies adaptive to image complexity. Experiment results demonstrate that ERASE significantly reduces vision tokens while preserving accuracy. For Qwen2.5‑VL‑7B, at a token pruning ratio of 85%, ERASE retains 89.46% of the original model accuracy, whereas the best prior method retains only 78.1%. Our code is available at https://github.com/Tuna‑Luna/ERASE.
Authors:Zhongyu Xia, Guanyu Zhu, Guo Tang, Wenhao Chen, Yongtao Wang
Abstract:
End‑to‑end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state‑of‑the‑art models achieving near‑perfect scores on widely used open‑loop and closed‑loop benchmarks. This saturation does not mean that the problem has been solved; instead, it reveals that current benchmarks remain limited in scenario diversity, object variety, and the breadth of driving capabilities they evaluate. In particular, they lack sufficient long‑tail scenarios involving rare but safety‑critical objects and fail to assess advanced decision‑making such as legal compliance, ethical reasoning, and emergency response. To address these gaps, we propose HiDrive, a new closed‑loop benchmark for end‑to‑end autonomous driving that emphasizes long‑tail scenarios and a richer evaluation of driving capabilities. HiDrive introduces a diverse set of rare objects and uncommon traffic situations, and expands evaluation from basic driving skills to more advanced capabilities, including rule compliance, moral reasoning, and context‑dependent emergency maneuvers. Correspondingly, we extend previous collision‑avoidance‑centered metrics into a comprehensive evaluation system that encompasses collision and braking, traffic‑rule compliance, and moral‑reasoning indicators. Built on a more advanced physics engine, HiDrive provides physically realistic lighting and high‑fidelity visual rendering, offering a more challenging and realistic testbed for assessing whether autonomous driving systems can handle the complexity of real‑world deployment. The HiDrive software, source code, digital assets, and documentation are available at https://github.com/VDIGPKU/HiDrive.
Authors:Kuan Zhang, Dongchen Liu, Qiyue Zhao, Tianyu Xin, Yue Su, Haisheng Wang, Han Yin, Hongbo Ma, Peize Li, Tianjun Gu, Xiangnan Wu, Xinran Zhang, Yongxuan Li, Zirong Chen, Yiming Li
Abstract:
The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this singular physical existence into a multiverse of games, each governed by entirely different rules, aesthetics, physics, and objectives. This omni‑reality adaptability is a hallmark of general intelligence. As Artificial Intelligence progresses towards Artificial General Intelligence, the multiverse of games has evolved from mere entertainment into the ultimate ground for training and evaluating AGI. The pursuit of this generality has unfolded across four eras: from environment‑specific symbolic and reinforcement learning agents, to current large foundation models acting as generalist players, and toward a future creator stage where agent both creates new game worlds and continually evolves within them. We trace the full lifecycle of a generalist game player along four interdependent pillars: Dataset, Model, Harness, and Benchmark. Every advance across these pillars can be read as an attempt to break one of five fundamental trade‑offs that currently bound the whole system. Building on this end‑to‑end view, we chart a five‑level roadmap, progressing from single‑game mastery to the ultimate creator stage in which the agent simultaneously creates and evolves within theoretical game multiverse. Taken together, our work offers a unified lens onto a rapidly shifting field,and a principled path toward the omnipotent generalist agent capable of seamlessly mastering any challenge within the multiverse of games, thereby paving the way for AGI.
Authors:Qingchao Jiang, Zhenxuan Hou, Zhiying Zhu, Zhenxing Qian, Xinpeng Zhang, Zaiwang Gu
Abstract:
With the rapid development of deep generative models, forged facial images are massively exploited for illegal activities. Although existing synthetic face detection methods have achieved significant progress, they suffer from the inherent limitation of overconfidence due to their reliance on the Softmax activation function. Thus, these methods often lead to unreliable predictions when encountering unknown Out‑of‑Distribution (OOD) images, and cannot ascertain the model's uncertainty in its prediction. Meanwhile, most existing methods require massive high‑quality annotated data, which greatly limits their practicability across diverse scenarios. To address these limitations, we propose EMSFD (Evidence‑based decision Modeling for Synthetic Face Detection with uncertainty‑driven active learning), an approach designed to enhance detection reliability and generalizability. Specifically, EMSFD models class evidence using the Dirichlet distribution and explicitly incorporates model uncertainty into the prediction process. Furthermore, during training, the estimated uncertainty is exploited to prioritize more informative samples from the unlabeled pool for annotation, thereby reducing labeling cost and improving model generalization. Extensive experimental evaluations demonstrate that our method enhances the interpretability of synthetic face detection. Meanwhile, our method yields a 15% increase in accuracy compared to existing state‑of‑the‑art (SOTA) baselines, which demonstrates the superior detection performance and generalizability of our approach. Our code is available at: https://github.com/hzx111621/EMSFD.
Authors:Junzhe Chen, Siyuan Meng, Yuxi Chen, Man Zhao, Wenyao Gui, Xiaojie Guo
Abstract:
Video large language models (Video‑LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize event recognition, action understanding, or coarse temporal reasoning, while rarely testing whether models can preserve the identity, state, and continuity of the same object across occlusion, disappearance, reappearance, state transitions, and cross‑object interactions. We introduce TOC‑Bench, a diagnostic benchmark for evaluating temporal object consistency in Video‑LLMs. TOC‑Bench is object‑track grounded: each queried subject is linked to a per‑frame trajectory and a structured temporal event timeline. To ensure that questions require temporally ordered visual evidence rather than language priors, single‑frame shortcuts, or unordered frame cues, we design a three‑layer temporal‑necessity filtering protocol, which removes 60.7% of candidate QA pairs and retains 17,900 temporally dependent items across 10 diagnostic dimensions. From this pool, we construct a human‑verified benchmark with 2,323 high‑quality QA pairs over 1,951 videos. Experiments on representative Video‑LLMs show that temporal object consistency remains a major unsolved challenge, with notable weaknesses in event counting, event ordering, identity‑sensitive reasoning, and hallucination‑aware verification, even when models perform well on general video understanding benchmarks. These results suggest that object‑centric temporal coherence is a key bottleneck for current Video‑LLMs, and that TOC‑Bench provides a focused platform for diagnosing and improving object‑aware temporal reasoning. The resource is available at https://github.com/cjzcjz666/toc_bench.git.
Authors:Anushree Berlia
Abstract:
We present Loom, an outfit recommendation system that combines neural embedding retrieval with structured domain scoring to generate complete, coherent outfits from fashion catalogs. Given an anchor clothing item, Loom retrieves complementary pieces via slot‑constrained approximate nearest neighbor search over FashionCLIP embeddings, then scores candidate outfits using a multi‑objective function that integrates six signals: embedding similarity, color harmony, formality consistency, occasion coherence, style direction, and within‑outfit diversity.
We introduce two techniques that address limitations of purely learned or purely rule‑based approaches: (1) semantic material weight, which uses CLIP embedding geometry to infer garment heaviness for layer compatibility without hand‑coded material taxonomies; and (2) vibe/anti‑vibe occasion priors, which embed prose descriptions of occasion contexts as anchor vectors in CLIP space and score items by differential affinity.
Ablation experiments on a catalog of 620 items show that each component contributes measurably to outfit quality: the full system achieves a mean outfit score of 0.179 with a 9.3% hard violation rate, compared to 0.054 score and 16.0% violations for a category‑constrained random baseline, a 3.3x improvement in score and 42% reduction in violations. Direction reranking is the single indispensable component: removing it drops score to 0.052, essentially equal to random. The system generates three stylistically distinct outfits in under 5 seconds on commodity hardware.
Authors:Zhipeng Liu, Chunbo Luo
Abstract:
Vision‑language models (VLMs) enable text‑guided object detection but degrade severely under cross‑view scenarios where ground and aerial viewpoints differ in altitude, scale, and spatial layout. These geometric changes introduce systematic complexity variations between viewpoints, e.g., ground view images contain dense and highly occluded structures, while aerial images are sparse and globally organized. Fixed VLM fusion mechanisms cannot handle this discrepancy. We propose CrossVL, a framework combining Complexity‑Aware Pathway Aggregation (CPA) and Paired Curriculum Learning (PCL) for enhanced cross‑view detection for VLM. CPA estimates scene complexity from multimodal statistics and routes visual features through multiple pathways to obtain view‑specific representations. PCL leverages semantic consistency of synchronized ground‑aerial pairs to provide stable early supervision and then gradually shifts toward randomized sampling. On MAVREC, CrossVL improves Florence‑2's aerial mAP from 58.66% to 61.03% and reduces the ground‑aerial performance gap from 8.63pp to 6.65pp, while also achieving a 3.3x reduction in variance across random seeds. CPA provides stable complexity‑aware feature aggregation, and PCL enhances optimization dynamics. Together, they demonstrate that coordinated architectural and training adaptations are crucial for robust cross‑view VLM detection.
Authors:Ke Zhang, Yunjie Tian, Dongdi Zhao, Yijiang Li, Yuanye Liu, Vishal M Patel, Di Fu
Abstract:
On‑policy distillation (OPD), which supervises a student on its own sampled trajectories, has emerged as a data‑efficient post‑training method for improving reasoning while avoiding the reward dependence of reinforcement learning and the catastrophic forgetting often observed in standard supervised fine‑tuning. However, standard OPD typically computes teacher supervision under noisy student‑generated contexts and often relies on a single stochastic teacher rollout per prompt. As a result, the supervision signal can be high‑variance: the sampled teacher trajectory can be incorrect, uninformative, or poorly matched to the student's current reasoning behavior. To address this limitation, we propose BRTS, a Best‑of‑N Rollout Teacher Selection framework for on‑policy distillation. BRTS augments standard student‑context OPD with a teacher‑context supervision branch constructed from the curated teacher trajectory. Rather than distilling from the first sampled teacher rollout, BRTS samples a small pool of teacher trajectories and selects the auxiliary trajectory using a simple priority rule: correctness first, student alignment second. When multiple correct teacher trajectories are available, BRTS chooses the one most aligned with the student's current behavior; when unconditioned teacher samples fail on harder prompts, it invokes a ground‑truth‑conditioned recovery step to elicit a natural derivation. The selected trajectory is then used to provide reliable teacher‑context supervision inside the OPD loop, augmented with an auxiliary loss on the teacher trajectory. Experiments on AIME 2024, AIME 2025, and AMC 2023 show that BRTS improves over standard OPD on challenging reasoning benchmarks, with the largest gains on harder datasets. Our code is available at https://github.com/BWGZK‑keke/BRTS.
Authors:Yufeng Hong, Xiaotian Zhou, Yingyan Li, Xiangpo Zhou, Lin Liu, Yadan Luo, Shaoqing Xu, Lei Yang, Ziying Song
Abstract:
Existing latent world models for autonomous driving have opened a promising path toward future‑aware driving intelligence. However, they typically treat future latent states as prediction targets or auxiliary signals, rather than directly conditioning trajectory planning. This can entangle current and future features in latent space. In this work, we propose DriveFuture, a future‑aware latent world modeling framework for autonomous driving that explicitly learns planning‑oriented foresight by conditioning the current latent state modeling process on future world states. Specifically, during training, the model first predicts future latent world states from the current latent state and ego action, and then refines the prediction against the ground‑truth future latent state via cross‑attention. The resulting future‑aware latent serves as an explicit condition for a diffusion‑based trajectory planner. During inference, DriveFuture conditions on the predicted future latent state instead of the ground‑truth future state. DriveFuture achieves SOTA performance on the public NAVSIM benchmarks, reaching 55.5 EPDMS on NAVSIM‑v2 \textcolorbluenavhard, 89.9 EPDMS on NAVSIM‑v2 \textcolorbluenavtest, and 90.7 PDMS on NAVSIM‑v1 \textcolorbluenavtest, respectively. These results suggest that the key to latent world modeling lies not merely in simulating future states, but more importantly in conditioning current decision‑making on future states. Notably, as of April 2026, DriveFuture ranks 1st on the \hrefhttps://huggingface.co/spaces/AGC2025/e2e‑driving‑navhardNAVSIM‑v2 \textcolorbluenavhard leaderboard and achieves SOTA performance on \hrefhttps://huggingface.co/spaces/AGC2024‑P/e2e‑driving‑navtestNAVSIM‑v1 \textcolorbluenavtest.
Authors:Yicheng Ji, Zhizhou Zhong, Jun Zhang, Qin Yang, XiTai Jin, Ying Qin, Wenhan Luo, Shuiyang Mao, Wei Liu, Huan Li
Abstract:
Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long‑horizon video generation with real‑time responsiveness, as exemplified by the Self Forcing training paradigm. However, existing AR video diffusion models still suffer from significant attention complexity and severe memory overhead due to the redundant key‑value (KV) caches across historical frames, which limits scalability. In this paper, we tackle this challenge by introducing KV cache compression into autoregressive video diffusion. We observe that attention heads in mainstream AR diffusion models exhibit markedly distinct attention patterns and functional roles that remain stable across samples and denoising steps. Building on our empirical study of head‑wise functional specialization, we divide the attention heads into two categories: static heads, which focus on transitions across autoregressive chunks and intra‑frame fidelity, and dynamic heads, which govern inter‑frame motion and consistency. We then propose Forcing‑KV, a hybrid KV cache compression strategy that performs structured static pruning for static heads and dynamic pruning based on segment‑wise similarity for dynamic heads. While maintaining output quality, our method achieves a generation speed of over 29 frames per second on a single NVIDIA H200 GPU along with 30% cache memory reduction, delivering up to 1.35x and 1.50x speedups on LongLive and Self Forcing at 480P resolution, and further scaling to 2.82x speedup at 1080P resolution. Code and demo videos are provided at https://zju‑jiyicheng.github.io/Forcing‑KV‑Page.
Authors:Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi, Boyan Wang, Liang He, Xinze Zhou, Sezgin Er, Ibrahim Ethem Hamamci, Zongwei Zhou, Alan Yuille
Abstract:
Medical vision‑language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual question answering (VQA) benchmarks collapse model capabilities into a single accuracy score, obscuring where and why models fail. We propose DeepTumorVQA, a hierarchical benchmark that follows the multi‑stage evidence chain in tumor diagnosis and decomposes 3D CT reasoning into four stages: recognition, measurement, visual reasoning, and medical reasoning. Higher‑level questions remain independently scorable, while their ground‑truth evidence chains are defined over lower‑level primitives. The benchmark contains 476K questions across 42 clinical subtypes on 9,262 3D CT volumes. In addition to a direct reasoning mode for VLMs, DeepTumorVQA provides tool‑interaction environments for agent evaluation, where a model can call external tools, including segmentation models, measurement programs, and medical knowledge modules, before answering the question. Evaluating over 30 model configurations, we find that reliable quantitative measurement is the primary bottleneck, making later‑stage visual and medical reasoning harder for VLMs, while tool augmentation substantially mitigates this issue. When tools are available, leveraging medical knowledge and tools to reason about medical images becomes a new challenge. We further show that ground‑truth step‑by‑step tool‑use traces from DeepTumorVQA can supervise agents and reduce tool‑use and reasoning failures. This stage‑wise progression from recognition to measurement to visual and medical reasoning provides a concrete roadmap for future medical VLM and AI agent studies. All data and code are released at https://github.com/Schuture/DeepTumorVQA.
Authors:Aws Khalil, Jaerock Kwon
Abstract:
Teleoperation systems are fundamentally limited by communication latency, which degrades situational awareness and control performance. Predictive display aims to mitigate this limitation by presenting an estimate of the current visual state rather than delayed observations. While recent advances in generative video models enable high‑quality video synthesis, their suitability for latency‑sensitive predictive display remains unclear. This paper presents a zero‑shot benchmark of off‑the‑shelf generative video models for short‑horizon predictive display, without task‑specific fine‑tuning. We formulate the problem as rollout‑based future frame prediction and develop a unified benchmarking pipeline using simulated driving data from the CARLA simulator. Five publicly released video models spanning transformer‑based and diffusion‑based families are evaluated across two resolutions and two conditioning regimes (multi‑frame and single‑frame). Performance is assessed using prediction accuracy (mean absolute difference), per‑rollout latency, peak GPU memory usage, and temporal error evolution across the prediction horizon. On this zero‑shot benchmark, no tested model simultaneously achieves low rollout error, non‑divergent per‑step error behavior, and real‑time inference at the source frame rate. Increasing model scale or resolution yields limited and, in some cases, inverted improvements. These findings highlight a gap between general‑purpose generative video synthesis and the requirements of predictive display in teleoperation, suggesting that practical deployment will require either explicit short‑horizon temporal supervision, in‑domain adaptation, or aggressive inference optimization rather than direct application of off‑the‑shelf models. Code, configurations, and qualitative results are released on the project page: https://bimilab.github.io/paper‑GenPD
Authors:Zichen Zou, Xiaosong Jia, Zuxuan Wu, Yu-Gang Jiang
Abstract:
Visual Geometry Grounded Transformer (VGGT) advances 3D reconstruction via scalable Transformer architecture, but the quadratic complexity of global attention prevents long context application. StreamVGGT enables streaming with causal attention, yet its KV cache grows linearly with frames, causing memory overflow and quality degradation. We present RetrieveVGGT, a training‑free framework, which formulates context construction for VGGT as a retrieval problem. By retrieving a fixed number of relevant frames at each step, VGGT maintains a controllable memory budget, which is close to its training context length. Interestingly, we find that the similarity between current frame queries and cached history frame keys at the first global attention layer of VGGT is already a strong indicator of relevance, eliminating the need for additional learned scoring. To enhance information diversity similar to a recommender system, we propose Segment Sampling so that the retrieval spans distinct relevant segments rather than a single high‑similarity region. We design a pose‑aware spatial memory mechanism that organizes history frames according to their already estimated camera poses, enabling location‑aware retrieval. Extensive experiments demonstrate that RetrieveVGGT achieves state‑of‑the‑art performance, outperforming StreamVGGT, TTT3R, and InfiniteVGGT while maintaining constant memory usage regardless of sequence length. Code is available at https://github.com/zzctmd/RetrieveVGGT.
Authors:Alvin Kimbowa, Moein Heidari, David Liu, Ilker Hacihaliloglu
Abstract:
While U‑Net architectures remain the gold standard for medical image segmentation, their deployment in resource‑constrained environments demands aggressive model compression. However, finding an optimally efficient configuration is computationally prohibitive, typically requiring exhaustive train‑and‑evaluate cycles to find the smallest model that maintains peak performance. In this paper, we introduce a training‑free selection framework to automatically identify ultralightweight, dataset‑specific U‑Net configurations directly at initialization. We observe that systematically scaling down U‑Net channel width induces a sharp transition from a stable performance plateau to representational capacity collapse. To pinpoint this boundary without training, we propose a Jacobian‑based sensitivity metric that scores discrete, width‑capped U‑Net variants using a small set of unlabeled images. By analyzing the total variation of this sensitivity curve, we isolate the smallest stable configuration, which we denote as XTinyU‑Net. Evaluated across six diverse medical datasets within the nnU‑Net framework, XTinyU‑Net achieves segmentation accuracy comparable to the heavy nnU‑Net baseline with 400x‑1600x fewer parameters, and outperforms contemporary lightweight architectures while utilizing 5x‑72x fewer parameters. Code is publicly accessible on https://github.com/alvinkimbowa/nntinyunet.git.
Authors:Zhenxuan Zeng, Lingxuan Wang, Sheng Yang, Yanan He, Mingxia Chen, Wei Suo, Peng Wang
Abstract:
Accurate High‑Definition (HD) map construction is critical for autonomous driving, yet existing methods face a fundamental trade‑off: vectorization‑based approaches preserve topology but struggle with geometric fidelity, while rasterization‑based approaches enable precise geometric supervision but produce unstructured outputs. To bridge this gap, we propose GSMap, a novel framework that unifies both paradigms via a learnable 2D Gaussian representation. Each map element is modeled as an ordered sequence of 2D Gaussians, whose centers correspond to the vertices of the vectorized polyline/polygon. This formulation enables simultaneous optimization through: (1) Differentiable rasterization that enforces pixel‑level geometric constraints, and (2) Topology‑aware vectorization that maintains structural regularity. Experiments on both nuScenes and Argoverse2 demonstrate that our Gaussian‑based representation effectively unifies geometric and topological learning, achieving significant performance improvements and demonstrating strong compatibility with existing HD mapping architectures. Code will be available at https://github.com/peakpang/GSMap
Authors:Can Li, Zhoujian Li, Ren Li, Jie Gu, Lei Lei, Jingmin Chen, Lei Sun
Abstract:
World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challenging because deformable linear, planar, and volumetric objects evolve under high‑dimensional deformation, noisy interactions, and complex material response. The model must therefore infer a physical state from visual observations, roll it forward under new interactions, and render the resulting dynamics with high visual fidelity. We present DeformMaster, a video‑derived interactive physics‑neural world model that turns real interaction videos into an online interactive model of deformable objects within a unified dynamics‑and‑appearance framework. DeformMaster preserves structured physical rollout while using a neural residual to compensate for unmodeled effects, grounds sparse hand motion as distributed compliant actuator for hand‑continuum interaction, represents material response with spatially varying constitutive experts, and drives high‑fidelity 4D appearance from the predicted physical evolution. Experiments on real‑world deformable‑object sequences demonstrate DeformMaster's ability to roll out future dynamics and render dynamic appearance, outperforming state‑of‑the‑art baselines while supporting novel action rollout, material‑parameter variation, and dynamic novel‑view synthesis. Project page: https://can‑lee.github.io/deformmaster‑web/
Authors:Siran Peng, Xiangyu Zhu, Shang-Qi Deng, Liang-Jian Deng, Zhen Lei
Abstract:
Remote sensing image fusion aims to create a high‑resolution multi/hyper‑spectral image from a high‑resolution image with limited spectral information and a low‑resolution image with abundant spectral data. Recently, deep learning (DL) techniques have shown significant effectiveness in this area. Most DL‑based methods approach image fusion as a 2D problem by encoding spectral information into feature map channels. However, our research suggests that this strategy introduces notable spectral distortions. In contrast, some methods consider spectral data as an additional dimension, utilizing standard 3D convolutions to preserve spectral information. Nevertheless, in a standard 3D convolutional layer, the same set of kernels is applied across all input regions, which we have found to be sub‑optimal for image fusion. Furthermore, standard 3D convolutions necessitate substantial computational resources. To address these challenges, we propose a novel convolutional paradigm called Adaptive 3D Convolution (Ada3D) for remote sensing image fusion. Ada3D applies a unique set of 3D kernels to each input voxel, enabling the capture of fine‑grained details. These adaptive kernels are generated through a two‑step process: (i) spatial and spectral kernels are derived from their respective image sources; (ii) these two types of kernels are then combined to form content‑aware 3D kernels that effectively integrate spatial and spectral information. Additionally, adaptive biases are introduced to enhance the convolutional outcome at the voxel level. Furthermore, we incorporate the group convolution technique to reduce computational complexity. As a result, Ada3D offers full adaptivity in an efficient manner. Evaluation results across five datasets demonstrate that our method achieves SOTA performance, underscoring the superiority of Ada3D. The code is available at https://github.com/PSRben/Ada3D.
Authors:Shanwen Tan, Hao Li, Jingtao Zhang, Xiaosong Jia, Xue Yang, Shaofeng Zhang, Yanyong Zhang
Abstract:
Streaming long‑video generation faces a central challenge in continuous semantic switching, requiring adaptive memory to preserve coherent visual evolution. Current approaches rely on cache rebuilding at prompt boundaries or fixed memory budgets, but they introduce redundant computation and limit flexible semantic adaptation. This limitation arises from a mismatch between cached video history and prompt updates, as memory preserves visual continuity while prompt switches demand rapid semantic adaptation. Motivated by this observation, we present SWIFT, Semantic Windowing and Injection for Flexible Transitions, a training‑free framework for multi‑prompt long‑video generation that enables efficient semantic switching while preserving temporal coherence in causal video diffusion models. SWIFT introduces a lightweight Semantic Injection Cache that augments cached video memory rather than reconstructing it from scratch at every prompt boundary. To avoid uniformly perturbing all attention channels, we further perform head‑wise semantic injection, so that each attention head receives a prompt update proportional to its alignment with the current video state. In addition, we introduce an Adaptive Dynamic Window that allocates temporal memory according to prompt phase, using larger local context near switching boundaries and smaller windows during stable segments to reduce average inference cost. To preserve long‑range semantic consistency under compressed local attention, we further maintain segment‑level semantic anchors that summarize prompt‑conditioned video history and reintroduce it as compact memory tokens. Compared with current state‑of‑the‑art methods, SWIFT preserves generation quality while achieving 22.6 FPS on a single H100 GPU, establishing a substantially more efficient solution for multi‑prompt long‑video generation. Our code is available at https://github.com/ShanwenTan/SWIFT.
Authors:Junkang Zhou, Yefei He, Feng Chen, Weijie Wang, Bohan Zhuang
Abstract:
Large‑scale autoregressive models have demonstrated remarkable capabilities in image generation. However, their sequential raster‑scan decoding relies on strictly next‑token prediction, making inference prohibitively expensive. Existing acceleration methods typically either introduce entirely new generation paradigms that necessitate costly pre‑training from scratch, or enable parallel generation at the expense of a training‑inference gap or altered prediction objectives. In this paper, we introduce FlashAR, a lightweight post‑training adaptation framework that efficiently adapts a pre‑trained raster‑scan autoregressive model into a highly parallel generator based on two‑way next‑token prediction. Our key insight is that effective adaptation should minimize modifications to the pre‑trained model's original training objective to preserve its learned prior. Accordingly, we retain the original AR head as a horizontal head for row‑wise prediction and introduce a complementary, lightweight vertical head for column‑wise prediction. To facilitate efficient adaptation, we branch the vertical head from an intermediate layer rather than the final layer, bypassing the inherent horizontal head bias. Moreover, since horizontal and vertical predictions capture complementary dependencies whose relative importance varies across target positions, we employ a learnable fusion gate to dynamically combine the two predictions at each position. To further reduce adaptation cost, we propose a two‑stage adaptation pipeline: the vertical head is first initialized through adaptation from the pre‑trained autoregressive model before jointly fine‑tuned with backbone to adapt to the new decoding paradigm. Extensive experiments on LlamaGen and Emu3.5 show that FlashAR achieves up to a 22.9x speedup for 512x512 image generation through a lightweight post‑training with merely 0.05% of the original training data.
Authors:Shogo Noguchi
Abstract:
Recent conditional image generation methods can improve controllability by generating images that are faithful to conditions such as sketches, human poses, segmentation maps, and depth. By applying these techniques to image augmentation while preserving annotations, generated images can be used as additional training data and can improve recognition performance. However, for high‑level driving tasks such as traffic‑rule extraction and driving‑behavior understanding, simply using annotations as conditions is insufficient. Instead, images must be augmented while preserving the detailed high‑level structure of the original scene. One possible solution is to use multiple conditions so that generated images retain diverse structural cues after generation. However, when multiple conditions are used, conflicts among conditions can prevent reliable structure preservation. In this work, we input semantic segmentation, depth, and edges extracted from the original image into a multi‑condition image generation model, thereby providing rich structural information as conditions. We further propose a modeling approach for handling conflicts among multiple conditions and show that it enables image generation with stronger structural preservation. We also build a generation framework and evaluation protocol for driving tasks, establishing a basis for comparison with prior and future models. As a result, this work contributes to image generation research by addressing condition conflicts in multi‑condition generation and provides an important step toward mitigating data scarcity in high‑level autonomous‑driving tasks.
Authors:Fangzheng Wu, Brian Summa
Abstract:
Attention sinks ‑‑ tokens that receive disproportionate attention mass ‑‑ are assumed to be functionally important in autoregressive language models, but their role in diffusion transformers remains unclear. We present a causal analysis in text‑to‑image diffusion, dynamically identifying dominant attention recipients per timestep and suppressing them via paired, training‑free interventions on the score and value paths. Across 553 GenEval prompts on Stable Diffusion~3 (with SDXL corroboration), removing these sinks does not degrade text‑image alignment (CLIP‑T) or preference proxies (ImageReward, HPS‑v2) at k=1; only under stronger interventions (k\!\geq\!10) does HPS‑v2 exhibit a metric‑dependent boundary, while CLIP‑T remains robust throughout. The perceptual shifts induced by suppression are nonetheless \emphsink‑specific ‑‑ ~\!6× larger than equal‑budget random masking ‑‑ revealing an empirical dissociation between trajectory‑level perturbation and \emphsemantic alignment in diffusion transformers. \footnoteCode available at https://github.com/wfz666/ICML26‑attention‑sink.
Authors:Boxuan Zhang, Jianing Zhu, Qifan Wang, Jiang Liu, Ruixiang Tang
Abstract:
Recent generative models can produce images that appear highly realistic, raising challenges in distinguishing real and AI‑generated images. Yet existing detectors based on pre‑trained feature extractors tend to over‑rely on global semantics, limiting sensitivity to the critical micro‑defects. In this work, we propose Micro‑Defects expose Macro‑Fakes (MDMF), a local distribution‑aware detection framework that amplifies micro‑scale statistical irregularities into macro‑level distributional discrepancies. To avoid localized forensic cues being diluted by plain aggregation, we introduce a learnable Patch Forensic Signature that projects semantic patch embeddings into a compact forensic latent space. We then use Maximum Mean Discrepancy (MMD) to quantify distributional discrepancies between generated and real images. Our theory‑grounded analysis shows that patch‑wise modeling yields provably larger discrepancies when localized forensic signals are present in generated images, enabling more reliable separation from real images. Extensive experiments demonstrate that MDMF consistently outperforms baseline detectors across multiple benchmarks, validating its general effectiveness. Project page: https://zbox1005.github.io/MDMF‑project/
Authors:Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, Fang Dong
Abstract:
Volumetric video (VV) streaming enables real‑time, immersive access to remote 3D environments, powering telepresence, ecological monitoring, and robotic teleoperation. These applications turn VV streaming into a real‑time interface to remote physical environments, imposing new system‑level demands for photorealistic scene representation, low‑latency interaction, and robust performance under heterogeneous networks. 3D Gaussian Splatting (3DGS) has been widely used for real‑time photorealistic rendering, offering superior visual quality and rendering performance, but it faces challenges due to bandwidth consumption. Furthermore, as the foundation of adaptive VV streaming, existing Levels of Detail (LoD) methods based on density are not well‑suited to Gaussian representations, leading to visible gaps and severe quality degradation. Recent studies have also explored attribute compression techniques to reduce bandwidth consumption. Our preliminary studies reveal that aggressive attribute compression primarily causes color distortion, which can be effectively corrected in the rendered image using a reference image. Motivated by these findings, we propose a novel Color‑Adaptive scheme for adaptive VV streaming that uses vector quantization (VQ) to establish LoDs and correct color distortions with low‑resolution reference images. We further present CAGS, an adaptive VV streaming system compatible with diverse Gaussian representations, which integrates the Color‑Adaptive scheme by rendering reference images on the streaming server and performing color restoration on the client. Extensive experiments on our prototype system demonstrate that CAGS outperforms the existing adaptive streaming systems in PSNR by 5~20 dB under fluctuating bandwidth, operates significantly faster than existing scalable Gaussian compression methods, and generalizes across different Gaussian representations.
Authors:Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou, Robert Berke, Mauricio Reyes
Abstract:
Vision‑‑language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and clinical text, thereby leveraging complementary clinical information. However, radiological findings are long‑tailed in practice, leaving some conditions underrepresented and making zero‑shot inference essential. Yet current CLIP‑style medical VLMs are sensitive to prompt variations and often lack trustworthy external knowledge at inference time, which hinders reliable clinical deployment. We present KEPIL, a prompt‑robust framework that integrates curated medical knowledge to stabilize zero‑shot generalization. KEPIL comprises: (i) \emphdynamic prompt enrichment using ontologies with LLM assistance, (ii) a \emphsemantic‑aware contrastive loss aligning embeddings of equivalent prompt variants via a dual‑embedding objective, and (iii) \emphentity‑centric report standardization to yield ontology‑aligned representations. Across seven benchmarks, KEPIL achieves state‑of‑the‑art zero‑shot inference performance; under prompt‑variation tests, it improves AUC by \(6.37%\) on CheXpert and by \(4.11%\) on average. These results suggest that structured knowledge and robust prompt design are key to clinically reliable radiology‑facing VLMs. Code will be released at https://github.com/Roypic/KEPIL.
Authors:Jiankun Peng, Jianyuan Guo, Yiguang Yang, Yue Liu, Jiashuang Yan, Ying Xu
Abstract:
Online topological planning has become an effective paradigm for Vision‑Language Navigation in Continuous Environments (VLN‑CE), but existing methods still suffer from two limitations: redundant local depth information and weakened focus on current frontier candidates as the topological graph grows. To address this, we propose LCGNav, a modular local geometric enhancement framework for topological VLN. LCGNav explicitly converts candidate depth views into 3D point clouds and applies physical truncation based on the agent's reachable range, enabling more compact local geometric modeling. It further introduces a dimension‑preserving local fusion strategy with transient state degradation, so that geometric enhancement is applied only to the currently relevant ghost nodes without changing the original planner interface. Experiments on R2R‑CE and RxR‑CE show that LCGNav serves as an effective cross‑architecture enhancement module, consistently improving multiple key metrics of representative online topological baselines with low additional training cost. When integrated with ETP‑R1, LCGNav achieves the best performance among the compared online topological methods on the val‑unseen splits of the R2R‑CE and RxR‑CE benchmarks. The code is available at https://github.com/shannanshouyin/LCGNav.
Authors:Mikhail G. Mozerov
Abstract:
This paper presents a novel Direct Integration Theorem (DIT), derived as a non‑trivial corollary of the classical Central Slice Theorem (CST). The DIT provides a mathematically consistent transition from the continuous to the discrete domain ‑ a fundamental challenge in computed tomography ‑ thereby eliminating the need for frequency‑domain interpolation without resorting to conventional ramp‑filtering. The proposed approach circumvents two principal limitations inherent in traditional methods: (i) the zero‑frequency singularity and spectral distortions introduced by the mandatory ramp‑filtering step, and (ii) discretization inaccuracies associated with frequency‑domain interpolation. Based on the DIT, we develop a rigorous framework for consistent discrete solutions of the inverse Radon problem. Mathematical modeling demonstrates that this approach achieves quasi‑exact reconstruction, with errors constrained solely by sampling parameters and grid geometry. Furthermore, while Filtered Back Projection (FBP) inherently distorts the variance of the reconstructed image, the DIT‑based algorithm preserves it. Comparative simulations confirm that the proposed method eliminates common artifacts, such as intensity cupping, and consistently outperforms FBP in terms of PSNR, SSIM, and reprojection fidelity, faithfully restoring the original image's statistical characteristics.
Authors:Yixin Tang, Jiawei Guo, Junxian Li, Zhiteng Li, Jixin Zhao, Bingya Zhang, Chenbo Wang, Yulun Zhang, Shangchen Zhou
Abstract:
Recently, diffusion‑based object removal models have achieved impressive results in eliminating objects and their associated visual effects. However, they indiscriminately denoise all tokens across all timesteps, ignoring that removal usually involves small foreground regions. This strategy introduces substantial computational overhead and prolonged inference times. To overcome this computational burden, we propose a latent discriminator to implement Region‑aware Adversarial Distillation (RAD), yielding a highly efficient few‑step model named FlashClear. Furthermore, tailored to few‑step diffusion models, we propose FPAC (Foreground‑Prioritized Asymmetric Attention and Caching), a training‑free acceleration strategy. Extensive experiments demonstrate that our framework provides massive acceleration while maintaining or exceeding the performance of our base model, ObjectClear. Notably, on the OBER benchmark, our FlashClear achieves up to 8.26× and 122× speedup over ObjectClear and OmniPaint, respectively, while maintaining high visual quality and fidelity.
Authors:Tri Cao, Khoi Le, Thong Nguyen, Cong-Duy Nguyen, Quynh Vo, Anh Tuan Luu, Chunyan Miao, See-Kiong Ng, Shuicheng Yan, Bryan Hooi
Abstract:
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio‑temporal monitoring, the ability to persistently track object identities, states, and relations over time. Existing benchmarks obscure this deficit by relying on single final‑answer evaluations for queries that can often be resolved via local visual cues or statistical priors. To rigorously diagnose this, we introduce STEMO‑Bench (Spatio‑TEmporal MOnitoring), a benchmark of human‑verified object‑centric facts that evaluates intermediate reasoning by decomposing queries into sub‑questions, distinguishing genuine temporal understanding from coincidental correctness. To address failure modes exposed by STEMO, we propose STEMO‑Track, a novel object‑centric framework that explicitly constructs and reasons over structured object trajectories via chunk‑wise state extraction and temporal aggregation. Extensive experiments demonstrate that our object‑centric framework significantly reduces hallucinated answers and improves spatio‑temporal reasoning consistency over state‑of‑the‑art MLLMs.
Authors:Yu Li, Volker Schwieger
Abstract:
In LiDAR‑based environment perception systems, ground segmentation is a key preprocessing step supporting various applications such as mapping and navigation. Although extensively studied, problems such as reflection noise and isolated ground remain challenging. To address these issues, we propose FugSeg, a fast uncertainty‑aware ground segmentation method. A polar grid map is adopted as the point cloud representation to ensure generalizability across LiDAR types. Building on that, we develop a within‑ and cross‑segment ground labeling strategy that identifies not only directly visible ground cells but also those that are isolated or occluded. During this process, an adaptive slope is introduced, which incorporates measurement uncertainties to enhance its reliability under complex terrain. Finally, to achieve point‑level ground segmentation, a fine‑grained ground elevation estimation method is introduced. Throughout the complete workflow, reflection noise is explicitly handled via the proposed noisy ground cells. We conduct comprehensive evaluations on four public datasets covering both structured and unstructured environments. Results show that FugSeg outperforms state‑of‑the‑art non‑learning methods, achieving the highest F1, accuracy, and mIoU across all datasets, while maintaining the fastest runtime (135 Hz and 487 Hz for 64‑ and 32‑layer LiDARs) using a single CPU thread, making it suitable for resource‑limited systems. The code will be available at https://github.com/Leo‑YuLi/FugSeg.
Authors:Xiang Feng, Jiawei Zhou, Zhangfeng Huang, Kewei Wang, Shanshan Ye, Jinxin Hu, Zulong Chen, Yong Luo, Jing Zhang
Abstract:
Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end‑to‑end answer accuracy. We introduce DocScope, a benchmark that formulates long‑document QA as a structured reasoning trajectory prediction problem: given a complete PDF document and a question, the model outputs evidence pages, supporting evidence regions, relevant factual statements, and a final answer. We design a four‑stage evaluation protocol ‑‑ Page Localization, Region Grounding, Fact Extraction, and Answer Verification ‑‑ that audits each level of the trajectory independently through inter‑stage decoupling, with all judges selected and calibrated via human alignment studies. DocScope comprises 1,124 questions derived from 273 documents, with all hierarchical evidence annotations completed by human annotators. We benchmark 6 proprietary models, 12 open‑weight models, and several domain‑specific systems. Our experiments reveal that answer accuracy cannot substitute for trajectory‑level evaluation: even among correct answers, the highest observed rate of complete evidence chains is only 29%. Across all models, region grounding remains the weakest trajectory stage. Furthermore, the primary difficulty stems from aggregating evidence dispersed across long distances and multiple document clusters, while an oracle study identifies faithful perception and fact extraction as the dominant capability bottleneck. Cross‑architecture comparisons further suggest that activated parameter count matters more than total scale. The benchmark and code will be publicly released at https://github.com/MiliLab/DocScope.
Authors:Hoang M. Truong, Hai Nguyen-Truong, Dang Huynh
Abstract:
Open‑vocabulary semantic segmentation requires adapting image‑level vision‑language models such as CLIP to dense pixel‑level prediction, which is challenging due to the mismatch between hierarchical structure and semantic alignment in the embedding space. While recent works leverage hyperbolic geometry to model hierarchical relationships, they align embeddings across hierarchical levels but overlook semantic misalignment among embeddings within the same level. In this work, we propose HyRo, a hyperbolic fine‑tuning framework that decouples hierarchical and semantic alignment in the Poincaré ball model. HyRo aligns hierarchical levels by adjusting the hyperbolic radius and refines semantic relationships through angular alignment using an orthogonal transformation that theoretically preserves the hyperbolic radius. Experiments on standard open‑vocabulary semantic segmentation benchmarks demonstrate that HyRo achieves state‑of‑the‑art performance over prior methods.
Authors:Piotr Borycki, Magdalena Trędowicz, Jacek Tabor, Łukasz Struski, Przemysław Spurek
Abstract:
Ante‑hoc interpretability methods based on prototypes provide highly accurate explanations by utilizing the intuitive "this looks like that" reasoning paradigm. On the other hand, post‑hoc models can explain predictions for a single image without relying on an underlying dataset or requiring costly neural network retraining. Recent approaches successfully solve the retraining problem for prototype‑based networks. However, they still face a fundamental limitation: they require access to a subset of data (e.g., a test or validation set) to search for and extract the visual prototypes. In this paper, we address this issue and introduce ProDG: Generative Prototypes for Data‑Free Post‑Hoc Explainability, a novel framework that leverages generative models to synthesize pure, high‑fidelity prototypes directly from the frozen model's weights, completely eliminating the dependency on any external data. By establishing this new frontier in Data‑Free XAI, ProDG unlocks robust visual interpretability for privacy‑sensitive domains, where original data is strictly restricted or fundamentally inaccessible. Project page: https://github.com/piotr310100/ProDG
Authors:Junli Zha, Jiahui Wang, Xinkai Lu, Jinbo Wang
Abstract:
Vision‑Language Models (VLMs) exhibit systematic bias toward visual illusions, recalling memorized facts rather than perceiving actual visual differences. This paper presents a training‑free framework for the 5th DataCV Challenge Task 1 at CVPR 2026, addressing this perception‑versus‑memory conflict through three complementary strategies:(1) illusion‑aware image preprocessing that weakens illusion‑inducing context via type‑specific transformations (edge extraction, color isolation, morphological processing, and reference‑line overlay), (2) anti‑illusion prompt engineering guiding VLMs toward qualitative visual comparison, and (3) multi‑vote ensemble that further improves robustness. Our method achieves 90.48% accuracy on the official 630‑image test set using Claude (claude‑opus‑4‑6) with 5‑vote majority ensemble, and 98.41% on a human‑verified subset. The approach requires no finetuning, relying solely on visual manipulation and prompt design. Our solution secured 2nd place in the challenge, only 0.47% behind the 1st‑place solution. Code is available at https://github.com/jasminezz/sf‑illusion‑aware‑vlm.git.
Authors:Zhen-Hao Xie, Yan Wang, Hao Sun, Han-Jia Ye, De-Chuan Zhan, Da-Wei Zhou
Abstract:
Class‑Incremental Learning (CIL) requires a learning system to learn new classes while retaining previously learned knowledge. However, in real‑world scenarios such as autonomous driving, a system trained on urban roads in sunny weather may later need to operate in rural or highway environments with different traffic patterns and weather conditions. This requires the model not only to overcome catastrophic forgetting, but also to effectively handle domain shifts. In this paper, we propose CrOss‑sample Relational Fusion (CORF), a unified framework to address domain shift and catastrophic forgetting simultaneously. To enhance generalizability, we perform selective refinement of training samples by leveraging spatial contribution maps to highlight semantically informative regions. Furthermore, we incorporate predictive confidence to adaptively weigh samples, thereby facilitating the learning of domain‑agnostic representations. To alleviate forgetting, we propose a cascaded distillation framework that captures cross‑sample relational dependencies across multiple feature hierarchies, enabling multi‑grained knowledge transfer from previous tasks. CORF can be seamlessly integrated into existing CIL algorithms to enhance their generalizability, achieving competitive performance across various benchmark datasets. Code is available at https://github.com/LAMDA‑CL/TMM26‑CORF .
Authors:Haimin Luo, Min Ouyang, Lan Xu, Jingyi Yu
Abstract:
Hair is a rich medium of visual and cultural expression, yet its digital modeling remains challenging due to the duality of fluidity and structure. Many existing generative approaches rely primarily on continuous diffusion fields, which entangle global topology with local texture and obscure the semantic and structural organization of hairstyles. To address this, we propose HairGPT, a strand‑centric framework that treats strands as generative primitives and formulates realistic 3D hairstyle synthesis as a dual‑decoupled autoregressive sequence modeling problem. Our method applies spatial decoupling across semantic scalp regions and structural decoupling along a hierarchical strand representation, progressing from global layout to fine‑grained style. We further introduce a geometric tokenizer and region‑aware semantic annotations to guide strand‑level generation, enabling compositional editing, synthesis of rare and complex hairstyles, and adaptation to stylized domains. By aligning generative modeling with the workflow of digital grooming, HairGPT turns hair generation from opaque texture synthesis into a structured and semantically controllable authoring process, supporting robust semantic conditioning and high‑fidelity results across realistic and stylized domains. Project Page: https://haiminluo.github.io/hairgpt/
Authors:Ziyang Ding, Linjian Meng, Yiming Wu, Yuhan Li, Yuhao Liu, Zhen Zhao
Abstract:
Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead of decoding intermediate steps into discrete tokens. However, existing works typically rely on hard alignment objectives to force latent representations to match predefined visual features, thereby severely limiting the exploratory of latent reasoning process. To address this problem, we propose CoLVR (Contrastive Optimization for Latent Visual Reasoning). To obtain a more exploratory visual reasoning, CoLVR introduces a latent contrastive training framework. Firstly, CoLVR learns diverse and exploratory representations with a latent contrastive objective guided by angle‑based perturbation, which expands the semantic latent space and avoids over‑constrained embedding. Then, CoLVR employs a latent trajectory contrastive reward for RL (Reinforcement Learning) post‑training to enable fine‑grained optimization of latent visual reasoning process and thus fostering diverse reasoning behaviors. Experiments demonstrate that CoLVR significantly enhances the exploratory capability of latent representations, achieving average improvements of 5.83% on VSP and 8.00% on Jigsaw, while also outperforming existing latent models on out of domain benchmarks, with a 3.40% gain on MMStar. The data, codes, and models are released at https://github.com/Oscar‑dzy/CoLVR.
Authors:Benlei Cui, Fangao Zeng, Weitao Jiang, Yuwen Zhai, Haiwen Hong, Longtao Huang, Hui Xue, Wenxiang Shang, Pipei Huang
Abstract:
Product poster generation poses distinct challenges beyond general poster design, requiring both faithful preservation of product appearance and precise control over dense, multi‑line text layouts. Prior methods typically adopt inpainting frameworks augmented with auxiliary modules such as ControlNet and OCR encoders. However, these approaches introduce architectural complexity and computational overhead while still suffering from text errors and subject extension artifacts. We present SimplePoster, a simple yet effective inpainting‑based framework that achieves faithful subject preservation and accurate, position‑controllable text rendering without external controllers. Our approach builds on two observations: (1) full‑parameter fine‑tuning of the base model effectively suppresses subject extension, outperforming ControlNet‑based alternatives; and (2) a zero‑cost character‑level position encoding enables geometry‑aware text generation without dedicated layout modules. Experiments show that SimplePoster achieves a 98.7% subject preservation rate, compared to 55.2% for SeedEdit 3.0 and 85.3% for PosterMaker, while also improving text rendering accuracy. Code, models, benchmark and a part of training data will be available at https://github.com/Alibaba‑YuFeng/SIMPLEPOSTER
Authors:Joowon Kim, Seungho Shin, Joonhyung Park, Eunho Yang
Abstract:
Recent "Thinking with Video" approaches use Video Generation Models (VGMs) for visual reasoning by producing temporally coherent Chain‑of‑Frames as reasoning artifacts. Even strong VGMs, however, exhibit two recurring failure modes on goal‑directed tasks: long‑horizon drift on multi‑step tasks and mid‑clip simulation errors that compound. Both stem from the absence of explicit reasoning built upon the VGM's short‑horizon visual prior, a role naturally filled by Vision‑Language Models (VLMs), but where to place the VLM is non‑trivial: upfront plans commit before any frame is generated and post‑hoc critiques over whole videos intervene too late. We propose VLM‑VGM Collaborative Video Reasoning (CollabVR), a closed‑loop framework that couples the VLM with the VGM at step‑level granularity: the VLM plans the immediate next action, inspects the clip the VGM generates, and folds the verifier's diagnosis directly into the next action prompt to repair detected failures. On Gen‑ViRe and VBVR‑Bench, CollabVR improves both open‑source and closed‑source VGMs over single‑inference, Pass@k, and prior test‑time scaling baselines at matched compute, with the largest gains on the hardest tasks. It also yields further improvements on top of a reasoning‑fine‑tuned VGM, indicating that step‑level VLM supervision is orthogonal to and stackable with reasoning‑oriented fine‑tuning. We provide video samples and additional qualitative results at our project page: https://joow0n‑kim.github.io/collabvr‑project‑page.
Authors:Jiaming Liang, Chi-Man Pun, Weisi Lin, Greta Seng Peng Mok
Abstract:
Learned image compression (LIC) integrates deep neural networks (DNNs) to map high‑dimensional images into compact latent representations, reducing redundancy and achieving superior rate‑distortion (RD) performance in benign settings. Unfortunately, due to inherent vulnerabilities in DNNs, LIC systems are susceptible to adversarial perturbations that lead to downstream deterioration, compression rate degradation, untargeted distortion, and both local semantic manipulation (LSM) and low‑resolution (3×28×28) global semantic manipulation (GSM). However, high‑resolution GSM remains unexplored due to its intractability. Notably, the existing project gradient descent (PGD) method achieves near‑perfect white‑box attacks for classification, segmentation, and other tasks, yet fails to generalize to high‑resolution GSM. Our theoretical and empirical analyses reveal that well‑performing GSM drives adversarial examples from the Identity Region to the Amplification Region through the Lazying‑Oscillating‑Refining stages. General \ell_\infty‑bounded attacks fail on high‑resolution GSM because their step‑size schedules cannot accommodate both the Oscillating and Refining stages. Based on this, we propose the Periodic Geometric Decay schedule that enables \ell_\infty‑bounded high‑resolution GSM. To verify our approach, we integrate it with PGD, yielding a minimal variant, PGD^2‑GSM. Extensive experiments on the Kodak (3×768×512) demonstrate that our PGD^2‑GSM is the first to stably achieve high‑resolution GSM, thereby exposing a novel threat to LIC systems. Code is available at https://github.com/chinaliangjiaming/PGD2‑GSM.
Authors:Weiren Zhao, Yi Dong, Cheng Chen
Abstract:
Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generation as disjoint objectives, lacking a meaningful functional synergy. In this work, we identify and address a critical question in unified medical modeling: what form of understanding truly benefits generation. We present SynerMedGen, a unified framework built on the proposed principle of generation‑aligned understanding, which synergizes understanding objectives with generation tasks via task alignment. SynerMedGen introduces three generation‑aligned understanding tasks and a two‑stage training strategy that transfers generation‑beneficial representations learned during understanding training to medical image synthesis. Remarkably, even with understanding training alone, our SynerMedGen achieves strong zero‑shot performance across 22 medical image synthesis tasks and demonstrates robust generalization to unseen datasets. When combined with generation training, SynerMedGen consistently outperforms state‑of‑the‑art specialized medical image synthesis models as well as recent unified medical models. We also release a large‑scale dataset named SynerMed consisting of 1M paired synthesis samples and 2M generation‑derived understanding instances to support further research on understanding‑generation synergy. Our project can be accessed at https://github.com/Mhilab/SynerMedGen.
Authors:Md. Shakhoyat Rahman Shujon, Sheikh Md. Galib Mahim, Md. Milon Islam, Md Rezwanul Haque, Md Rabiul Islam, Hamdi Altaheri, Fakhri Karray
Abstract:
We propose CAST, a dual‑stream architecture that utilizes channel‑aware spatial transfer learning for isolated sign language recognition addressing the challenges of magnitude‑only 60~GHz radar Range‑Time Maps (RTM). The proposed framework combines three physics‑aware architectures with pretrained vision backbones, which operate under radar‑only constraints across clinical and alphabetical gestures. First, an explicit decibel‑to‑linear inversion is combined with a windowed fast Fourier transform that extracts Cadence Velocity Diagrams (CVD) while avoiding the harmonic artifacts that arise from the spectral analysis of log‑compressed signals. Second, a cross‑antenna spatial attention module applies attention to raw antenna channels before the convolution, preserving inter‑receiver amplitude covariance. Third, an asymmetric cross‑attention mechanism fuses representations from parallel ConvNeXt‑Tiny (CVD) and EfficientNetV2‑S (RTM) backbones. Extensive experiments reveal that the architecture achieves a Top‑1 accuracy of 80.5% under 5‑fold cross‑validation, establishing a 3.3% improvement over the best single‑model baseline (77.2%). The findings suggest that physics‑aware signal representations form a promising direction for radar‑only sign language recognition under constrained sensor modalities. The source code is available at: https://github.com/Shakhoyat/CAST‑at‑SignEval2026.
Authors:Soyeon Na, Seung Young Noh, Ju Yong Chang
Abstract:
Egocentric human mesh recovery (HMR) from monocular head‑mounted cameras is increasingly important for AR/VR applications, but remains challenging due to the lack of reliable ground‑truth (GT) annotations based on parametric human body models such as SMPL and SMPL‑X for real egocentric images. Existing egocentric HMR methods typically rely on pseudo‑GT and focus on body pose estimation, which limits their ability to recover fine‑grained whole‑body details such as hands and face. We study egocentric whole‑body human mesh recovery and propose a prior‑guided learning framework that reconstructs whole‑body meshes from a single egocentric image. We construct more accurate optimization‑based pseudo‑GT aligned with 3D joint supervision, and leverage multiple priors by adapting an exocentric HMR foundation model together with a diffusion‑based pose prior. A deterministic undistortion module is further adopted to handle fisheye distortions in egocentric images. Experiments across multiple egocentric benchmarks demonstrate improved whole‑body reconstruction compared to state‑of‑the‑art methods, and show that our optimization‑based pseudo‑GT is substantially more accurate than existing regression‑based pseudo‑GT. To facilitate reproducibility, the code and dataset annotations are publicly available at https://github.com/naso06/EgoSMPLX.
Authors:Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami
Abstract:
Multimodal Large Language Models (MLLMs) have made rapid progress in single‑video understanding, yet their ability to reason across multiple independent video streams remains poorly understood. Existing multi‑video benchmarks rely largely on human‑annotated real‑world footage, limiting the precision of spatial, temporal, and physical ground truth and making it difficult to diagnose model failures. We introduce SYNCR, a controlled synthetic benchmark for cross‑video reasoning with programmatically verified grounding. Built using Habitat, Kubric, and CLEVRER simulator engines, SYNCR contains 8,163 multi‑video question‑answer pairs grounded in 9,650 unique videos. It evaluates MLLMs across eight tasks spanning four diagnostic pillars: Temporal Alignment, Spatial Tracking, Comparative Reasoning, and Holistic Synthesis.
Our zero‑shot evaluation of leading open‑ and closed‑weight MLLMs reveals a substantial gap between current models and humans: the best model achieves only 52.5% average accuracy, compared to an 89.5% human baseline. Models perform relatively well on temporal ordering but struggle with precise physical and spatial reasoning, with the best model reaching only 26.0% accuracy on Kinematic Comparison. We further find that parameter scaling and reasoning‑specialized post‑training improve temporal alignment capabilities, but do not reliably address fine‑grained physical tracking or global spatial synthesis. Finally, an exploratory sim‑to‑real correlation analysis suggests that several SYNCR tasks track model‑level trends on real‑world multi‑video benchmarks, while also exposing reasoning capabilities underrepresented by existing evaluations. Code available at https://github.com/SaraGhazanfari/SYNCR.
Authors:Weijing Wu, Qihua Liang, Bineng Zhong, Haiying Xia, Zhiyi Mo, Shuxiang Song
Abstract:
Refining visual representations by eliminating their internal feature‑level redundancy is crucial for simultaneously optimizing the performance and computational cost of models in visual tracking. To enhance their performance, many contemporary Transformer‑based trackers leverage a larger number of historical template frames to capture richer spatio‑temporal cues. However, this strategy leads to a massive number of input visual tokens. This creates two critical issues: it imposes a quadratic computational burden and can also degrade the tracker's overall performance. To bridge this gap, we propose a compress‑then‑interact tracking framework, ETCTrack, that learns to efficiently compress template tokens from historical template frames into a robust target representation, moving beyond handcrafted rules. Our method first employs the Adaptive Token Compressor to dynamically construct compact yet highly discriminative template tokens by filtering out redundant visual tokens. These refined template tokens are then processed by our Hierarchical Interaction Encoder to achieve a deep, adaptive interaction with the search features. Refined search features ensure subsequent precise target localization. Experiments on seven benchmarks demonstrate that our method outperforms current state‑of‑the‑art trackers. ETCTrack‑B224 reduces the number of template tokens by 60%, leading to a 21.4% reduction in MACs with only a 0.4% drop in accuracy. The source code are available at https://github.com/PJD‑WJ/ETCTrack.
Authors:Yize Cai, Rui Feng, Anlan Yu, Baoshen Guo, Zhiqing Hong
Abstract:
Human Activity Recognition (HAR) from wearable sensors supports broad healthcare and behavior science applications. However, data heterogeneity and the scarcity of labeled data limit its real‑world generalization. Recent advances in self‑supervised learning (SSL) in vision and language domains have shown strong capability for learning generalizable representations from unlabeled data. Yet, few studies have systematically compared the generalization performance of SSL methods or explored how to adapt them for generalizable HAR. To address these gaps, we present BenchHAR, a unified framework for evaluating the generalization capability of SSL methods for sensor‑based HAR on unseen target distributions. BenchHAR curates a large‑scale dataset (~258K samples) and evaluates eight representative SSL methods across 12 encoder‑classifier architectures. Our results reveal that existing SSL methods struggle to achieve satisfactory generalization performance. We find that: (1) For HAR models, the hybrid paradigm (combining reconstruction and contrastive pretraining) achieves the best overall performance. The CNN encoder exhibits the strongest ability to learn generalizable representations, while more expressive classifier architectures further improve generalization. (2) For data scale, increasing the amount of pretraining data from downstream activity classes consistently improves generalization, while adding more labeled data yields limited gains. Interestingly, incorporating unlabeled data from non‑downstream activity classes does not improve generalization. (3) Sensor data collected from custom‑grade devices generalizes better than that from research‑grade devices, and data from limb transfers more effectively to trunk positions. BenchHAR provides a unified benchmark and actionable insights for generalizable sensor‑based HAR systems. Our code is available at https://github.com/saiketa/HAR‑Bench.
Authors:Lennard M. van Karnenbeek, Hilde G. A. van der Pol, Mark Wijkhuizen, Eva Poelman, Caroline A. Drukker, Theo Ruers, Freija Geldof, Behdad Dashtbozorg
Abstract:
Purpose: We aim to enhance the image quality of point‑of‑care ultrasound (POCUS) devices using deep learning and a novel paired dataset of POCUS and high‑end ultrasound images.
Approach: We collected the first accurately paired dataset using a custom‑built automated gantry system of low‑end POCUS and high‑end ultrasound images. A conditional generative adversarial network (cGAN) was utilized based on the pix2pix architecture, with a U‑Net generator that incorporates both L1 and structural similarity index (SSIM) losses to improve perceptual quality. Pretraining on a simulation dataset further boosts performance. Evaluation was performed on 1064 paired ex vivo tissue and phantom ultrasound image sets.
Results: Our approach improves the SSIM from 0.29 to 0.54 and PSNR from 19.16 dB to 22.41 dB. No‑reference metrics also indicate substantial enhancement, with the Natural Image Quality Evaluator (NIQE) and Perception‑based Image Quality Evaluator (PIQE) scores dropping from 7.95 to 4.44 and 31.12 to 19.99, respectively.
Conclusions: This work presents the first publicly available accurately paired dataset of low‑end POCUS to high end ultrasound images. Additionally, our results demonstrate the potential of the proposed framework to overcome hardware limitations of handheld POCUS, enhancing its diagnostic value in low‑resource and point‑of‑care settings. The POCUS‑IQ Dataset is publicly available at https://github.com/NKI‑MedTech‑AI/POCUS‑IQ.
Authors:Jiazheng Li, Chi-Hao Wu, Yunze Liu, Kaize Ding, Jundong Li, Chuxu Zhang
Abstract:
Understanding ultra‑long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For current multimodal LLMs: even with million‑token context windows, frame budgets cover only tens of minutes of densely sampled video, and most evidence is discarded before inference begins. Memory‑augmented and agentic approaches help with scale, but their retrieval remains fragmented across modalities and lacks long‑range narrative summaries that span days or weeks. We propose MAGIC‑Video, a training‑free framework built around a multimodal memory graph with interleaved narrative chain: the graph unifies episodic, semantic, and visual content through six typed edges and supports cross‑modal retrieval, while the chain distils long‑horizon entity biographies and recurring activity events. At inference time, an agentic loop interleaves graph retrieval with narrative fact injection, covering both the modality and time dimensions of ultra‑long video in a single retrieval pipeline. On EgoLifeQA, Ego‑R1 and MM‑Lifelong, MAGIC‑Video consistently outperforms strong general‑purpose, long‑video, and agentic baselines, with gains of 10.1, 7.4, and 5.9 points over the prior best agentic system on each benchmark. Code is available at https://github.com/lijiazheng0917/MAGIC‑video.
Authors:Logan Mann, Ajit Saravanan, Ishan Dave, Shikhar Shiromani, Saadullah Ismail, Yi Xia, Emily Huang
Abstract:
A pervasive intuition holds that vision‑language models (VLMs) are most trustworthy when their attention maps look sharp: concentrated attention on the queried region should imply a confident, calibrated answer. We test this Attention‑Confidence Assumption directly. We instrument three open‑weight VLM families (LLaVA‑1.5, PaliGemma, Qwen2‑VL; 3‑7B parameters) with a unified mechanistic pipeline ‑‑ the VLM Reliability Probe (VRP) ‑‑ that compares attention structure, generation dynamics, and hidden‑state geometry against a single correctness label. Three results emerge. (i) Attention structure is a near‑zero predictor of correctness (R_pb(C_k,y)=0.001, 95% CI [‑0.034,0.036]; R_pb(H_s,y)=‑0.012, [‑0.047,0.024] on a pooled n=3,090 split), even though attention remains causally necessary for feature extraction (top‑30% patch masking drops accuracy by 8.2‑11.3 pp, p<0.001). (ii) Reliability becomes legible later in the computation: a single hidden‑state linear probe reaches AUROC>0.95 on POPE for two of three families, and self‑consistency at K=10 is the strongest behavioral predictor we measure at 10x inference cost (R_pb=0.43). (iii) Causal neuron‑level ablations expose a sharp architectural split with direct monitor‑design implications: late‑fusion LLaVA concentrates reliability in a fragile late bottleneck (‑8.3 pp object‑identification accuracy after top‑5 probe‑neuron ablation), whereas early‑fusion PaliGemma and Qwen2‑VL distribute it widely and absorb destruction of ~50% of their peak‑layer hidden dimension with <=1 pp degradation. The takeaway is narrow but consequential: in 3‑7B VLMs, reliability is read more reliably off hidden‑state geometry, layer‑wise margin formation, and sparse late‑layer circuits than off attention‑map sharpness.
Authors:Maria Stoica, Abdelrahman Hekal, Alessio Lomuscio
Abstract:
Reliable out‑of‑distribution (OOD) detection is a critical requirement for the safe deployment of machine learning systems. Despite recent progress, state‑of‑the‑art OOD detectors are highly susceptible to adversarial attacks, which undermines their trustworthiness in automated systems. To address this vulnerability, we apply median smoothing to baseline OOD detection scores, balancing clean and adversarial accuracies. Our key insight is that the noisy samples generated for median smoothing can be repurposed to quantify the local instability of the base score. We observe that OOD samples exhibit higher instability under perturbation. Based on this, we propose ROSS, a novel and robust post‑hoc OOD detector that leverages the instability of baseline scores to further distinguish between in‑distribution (ID) and OOD samples. ROSS achieves symmetric robustness, performing strongly against both score‑minimising and score‑maximising attacks, unlike prior work. This symmetric defence leads to state‑of‑the‑art robustness, outperforming prior methods by up to 40 AUROC points. We demonstrate ROSS's effectiveness on extensive experiments across CIFAR‑10, CIFAR‑100, and ImageNet. Code is available at: https://github.com/Abdu‑Hekal/ROSS.
Authors:Bo Ye, Kai Gan, Tong Wei, Min-Ling Zhang
Abstract:
Generalized Category Discovery (GCD) seeks to identify novel categories from unlabeled data while retaining the classification ability of seen categories. Prior GCD methods commonly leverage transferable representations from pre‑trained models, adapting to downstream datasets via partial fine‑tuning (updating only the final ViT block) and visual prompt tuning (appending learnable vectors to inputs). However, conventional partial fine‑tuning offers limited flexibility, as it fails to adapt the entire model; meanwhile, visual prompt tuning is prone to overfitting, due to its sensitivity to initialization and inherently constrained capacity. To address these limitations, we propose LAGCD, a simple yet effective GCD approach that embeds a residual linear adapter into each ViT block. From the perspective of feature sparsity, we systematically show that non‑linearity in conventional adapters impairs performance, whereas our linear adapter enhances it by enabling more flexible model capacity. We further introduce an auxiliary distribution alignment loss to mitigate the negative impact of biased predictions between seen and novel categories. Extensive experiments on both generic and fine‑grained datasets confirm that LAGCD consistently improves performance over many sophisticated baselines. The source code is available at https://github.com/yebo0216best/LAGCD
Authors:Weicai Yan, Xinhua Ma, Wang Lin, Tao Jin
Abstract:
Parameter‑efficient fine‑tuning methods introduce a small number of training parameters, enabling pre‑trained models to adapt rapidly to new data distributions. While these methods have shown promising results, they exhibit notable limitations. First, most existing methods operate in the signal space domain, which results in substantial information redundancy. Second, most existing methods utilize fixed prompts or adaptation layers, failing to fully account for the multi‑scale characteristics of signals. To address these challenges, we propose the Multi‑Scale Frequency Adapter (FreqAdapter), which integrates textual information and performs multi‑scale fine‑tuning of signals in the frequency domain. Additionally, we introduce a multi‑scale adaptation strategy to optimize receptive fields across different frequency ranges, further enhancing the model's representational capacity. Extensive experiments on multimodal models, including CLIP and LLaVA, demonstrate that FreqAdapter significantly improves both performance and efficiency. FreqAdapter improves performance with minimal cost and fast convergence within one epoch. Code is available at https://github.com/Kelvin‑ywc/FreqAdapter.
Authors:Girmaw Abebe Tadesse, Titien Bartette, Andrew Hassanali, Allen Kim, Jonathan Chemla, Andrew Zolli, Yves Ubelmann, Caleb Robinson, Inbal Becker-Reshef, Juan Lavista Ferres
Abstract:
Monitoring archaeological sites at scale is vital for protecting cultural heritage, yet pinpointing when disturbances occur remains difficult because visual cues are subtle and ground‑truth data are sparse. We introduce WATCH, a framework for month‑level change‑event localization over PlanetScope satellite mosaics (2017‑2024, 4.7 m/px) that supports three complementary scoring approaches: (i) Temporal Embedding Distance (TED), a training‑free method that scores month‑to‑month deviations from a local temporal reference; (ii) Self‑Supervised Change Detection (SSCD), an ensemble of reconstruction, forecasting, and latent‑novelty signals; and (iii) a Weakly Supervised (WS) temporal localization model trained with sparse event‑month labels. We benchmark WATCH on 1,943 archaeological sites in Afghanistan using embeddings from six foundation models (CLIP, GeoRSCLIP, SatMAE, Prithvi‑EO‑2.0, DINOv3, and Satlas‑Pretrain) alongside a handcrafted spectral and texture baseline, and assess cross‑regional generalization on sites in Syria, Turkey, Pakistan, and Egypt. The unsupervised approaches (TED, SSCD) consistently outperform the weakly supervised alternative. TED with SatMAE achieves the highest exact‑month recall (55% at m=0), while TED with GeoRSCLIP, CLIP, or Satlas‑Pretrain reaches 92.5% within a three‑month tolerance (m=3). Handcrafted features remain competitive for exact‑month detection under weak supervision. Our directional margin analysis reveals systematic temporal biases: SSCD paired with GeoRSCLIP or Prithvi‑EO‑2.0 exhibits the strongest early‑warning profile, detecting anomalies before the recorded event, while TED favors confirmation‑oriented detection after a change has materialized. These results show that satellite imagery combined with foundation‑model embeddings enables scalable, decision‑relevant heritage monitoring. Code: https://github.com/microsoft/WATCH
Authors:Zi-Yi Jia, Zi-Jian Cheng, Xin-Yue Zhang, Kun-Yang Yu, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo
Abstract:
Multi‑model learning has attracted great attention in visual‑text tasks. However, visual‑tabular data, which plays a pivotal role in high‑stakes domains like healthcare and industry, remains underexplored. In this paper, we introduce VT‑Bench, the first unified benchmark for standardizing vision‑tabular discriminative prediction and generative reasoning tasks. VT‑Bench aggregates 14 datasets across 9 domains (medical‑centric, while covering pets, media, and transportation) with over 756K samples. We evaluate 23 representative models, including unimodal experts, specialized visual‑tabular models, general‑purpose vision‑language models (VLMs), and tool‑augmented methods, highlighting substantial challenges of visual‑tabular learning. We believe VT‑Bench will stimulate the community to build more powerful multi‑modal vision‑tabular foundation models.
Benchmark: https://github.com/Ziyi‑Jia990/VT‑Bench
Authors:Peng Liao, Shangsong Liang, Lin Chen, Peijia Zheng
Abstract:
Inertial Measurement Unit (IMU)‑based Human Activity Recognition (HAR) aims to interpret and classify user behaviors from temporal motion signals. Recently, deep learning frameworks have advanced this task by learning and extracting discriminative spatiotemporal representations, significantly improving recognition performance. However, IMU‑based HAR still faces several critical challenges, particularly limited training samples and static knowledge utilization, both of which severely hinder its large‑scale deployment. In this paper, we introduce MoRA, the first Retrieval‑Augmented Module specifically designed for motion series. It can be flexibly integrated into any existing HAR model, enhancing recognition performance while maintaining inference efficiency. To address issues such as information redundancy in retrieval results and rigid fusion strategies, we propose an uncertainty‑adaptive fusion unit within MoRA. This unit leverages previous physical knowledge from IMU signals to dynamically adjust the fusion strategy between original outputs and retrieved information, enabling more robust recognition. Extensive experiments on ten real‑world datasets demonstrate that MoRA significantly improves the performance of existing IMU‑based HAR models, consistently delivering stable and effective gains. The source code of MoRA is available at: https://github.com/liavonpenn/mora.
Authors:Daniel Dauner, Valentin Charraut, Bastian Berle, Tianyu Li, Long Nguyen, Jiabao Wang, Changhui Jing, Maximilian Igl, Holger Caesar, Boris Ivanovic, Yiyi Liao, Andreas Geiger, Kashyap Chitta
Abstract:
The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain largely untapped. Each dataset adopts different 2D and 3D modalities, such as cameras, lidar, ego states, annotations, traffic lights, and HD maps, with different rates and synchronization schemes. They come in fragmented formats requiring complex dependencies that cannot natively coexist in the same development environment. Further, major inconsistencies in annotation conventions prevent training or measuring generalization across multiple datasets. We present 123D, an open‑source framework that unifies such multi‑modal driving data through a single API. To handle synchronization, we store each modality as an independent timestamped event stream with no prescribed rate, enabling synchronous or asynchronous access across arbitrary datasets. Using 123D, we consolidate eight real‑world driving datasets spanning 3,300 hours and 90,000 kilometers, together with a synthetic dataset with configurable collection scripts, and provide tools for data analysis and visualization. We conduct a systematic study comparing annotation statistics and assessing each dataset's pose and calibration accuracy. Further, we showcase two applications 123D enables: cross‑dataset 3D object detection transfer and reinforcement learning for planning, and offer recommendations for future directions. Code and documentation are available at https://github.com/kesai‑labs/py123d.
Authors:Wei Yu, Yunhang Qian
Abstract:
Recent event‑based image reconstruction methods predominantly rely on Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to process complementary event information. However, these architectures face fundamental limitations: CNNs often fail to capture global feature correlations, whereas ViTs incur quadratic computational complexity (e.g., O(n^2)), hindering their application in high‑resolution scenarios. To address these bottlenecks, we introduce EmambaIR, an Efficient visual State Space Model designed for image reconstruction using spatially sparse and temporally continuous event streams. Our framework introduces two key components: the cross‑modal Top‑k Sparse Attention Module (TSAM) and the Gated State‑Space Module (GSSM). TSAM efficiently performs pixel‑level top‑k sparse attention to guide cross‑modal interactions, yielding rich yet sparse fusion features. Subsequently, GSSM utilizes a nonlinear gated unit to enhance the temporal representation of vanilla linear‑complexity (O(n)) SSMs, effectively capturing global contextual dependencies without the typical computational overhead. Extensive experiments on six datasets across three diverse image reconstruction tasks ‑ motion deblurring, deraining, and High Dynamic Range (HDR) enhancement ‑ demonstrate that EmambaIR significantly outperforms state‑of‑the‑art methods while offering substantial reductions in memory consumption and computational cost. The source code and data are publicly available at: https://github.com/YunhangWickert/EmambaIR
Authors:Zhen Fang, Wenxuan Huang, Yu Zeng, Yiming Zhao, Shuang Chen, Kaituo Feng, Yunlong Lin, Lin Chen, Zehui Chen, Shaosheng Cao, Feng Zhao
Abstract:
Existing Flow Matching (FM) text‑to‑image models suffer from two critical bottlenecks under multi‑task alignment: the reward sparsity induced by scalar‑valued rewards, and the gradient interference arising from jointly optimizing heterogeneous objectives, which together give rise to a 'seesaw effect' of competing metrics and pervasive reward hacking. Inspired by the success of On‑Policy Distillation (OPD) in the large language model community, we propose Flow‑OPD, the first unified post‑training framework that integrates on‑policy distillation into Flow Matching models. Flow‑OPD adopts a two‑stage alignment strategy: it first cultivates domain‑specialized teacher models via single‑reward GRPO fine‑tuning, allowing each expert to reach its performance ceiling in isolation; it then establishes a robust initial policy through a Flow‑based Cold‑Start scheme and seamlessly consolidates heterogeneous expertise into a single student via a three‑step orchestration of on‑policy sampling, task‑routing labeling, and dense trajectory‑level supervision. We further introduce Manifold Anchor Regularization (MAR), which leverages a task‑agnostic teacher to provide full‑data supervision that anchors generation to a high‑quality manifold, effectively mitigating the aesthetic degradation commonly observed in purely RL‑driven alignment. Built upon Stable Diffusion 3.5 Medium, Flow‑OPD raises the GenEval score from 63 to 92 and the OCR accuracy from 59 to 94, yielding an overall improvement of roughly 10 points over vanilla GRPO, while preserving image fidelity and human‑preference alignment and exhibiting an emergent 'teacher‑surpassing' effect. These results establish Flow‑OPD as a scalable alignment paradigm for building generalist text‑to‑image models. The codes and weights will be released in: https://github.com/CostaliyA/Flow‑OPD .
Authors:Ismail Aljosevic, Amir Masoud Almasi, Ana Parovic, Ashkan Shafiei
Abstract:
In this paper, we propose a modular framework for 6D pose estimation based on keypoint heatmap regression. Our approach combines YOLOv10m for object detection with a ResNet18‑based network that predicts 2D heatmaps from RGB images. Keypoints extracted from these heatmaps are used to estimate the 6D object pose via the PnP RANSAC algorithm. We compare different keypoint selection strategies to assess their impact on pose accuracy. Additionally, we extend the baseline by incorporating depth data using a cross‑fusion architecture, which enables interaction between RGB and depth features at multiple stages. We further explore general training improvements, such as experimenting with activation functions and learning rate scheduling strategies to improve model performance. Our best RGB‑only model achieved a mean ADD‑based accuracy of 84.50%, while the RGB‑D fusion model reached 92.41% on the LINEMOD dataset. The code is available at https://github.com/ameermasood/HeatNet.
Authors:Kaidi Jia, Yujie Lin, Chengyi Yang, Jiayao Ma, Jinsong Su
Abstract:
Vision‑language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. However, existing methods primarily fine‑tune the language decoder, leading to superficial forgetting that fails to erase underlying visual representations and often introduces object hallucination. We propose HFRU, a reinforcement unlearning framework that operates on the vision encoder for deep semantic removal. Our two‑stage approach combines alignment disruption with GRPO‑based optimization using a composite reward, including an abstraction reward that encourages semantically valid substitutions and mitigates hallucinations. Experiments on object recognition and face identity tasks show that HFRU achieves over 98% forgetting and retention performance, while introducing negligible object hallucination, significantly outperforming prior methods.Our code and implementation details are available at https://github.com/XMUDeepLIT/HFRU.
Authors:Henry Marichal, Diego Passarella, Gregory Randall
Abstract:
Tree ring marking remains a key step in dendrometry and dendrochronology, but it is often performed manually, making the process time‑consuming, subjective, and difficult to scale to large image datasets.
We present the Tree Ring Analyzer Suite (TRAS), an open‑source graphical software for automatic delineation, manual correction, and measurement of tree rings in wood cross‑sectional images. TRAS integrates three complementary detection algorithms: the classical image‑processing method CS‑TRD and two deep‑learning approaches, DeepCS‑TRD and INBD. The interface allows users to refine automatic detections, remove false positives, and manually add missing rings. It also computes dendrochronological metrics such as earlywood and latewood areas, ring perimeter, equivalent ring width, and custom path‑based ring‑width measurements.
TRAS was evaluated on 18 expertly annotated Pinus taeda L. cross‑section images. DeepCS‑TRD achieved the best automatic detection performance, with an F‑score of 81.0% and precision of 86.4%. Automatic detection reduced the required manual correction effort to approximately 20% of ring boundaries. For one‑dimensional ring‑width measurements, TRAS showed excellent agreement with CooRecorder (r > 0.99). Common detection errors, such as jump propagation or false positives near knots, were easily corrected through the postprocessing interface.
TRAS provides a flexible and reproducible solution for tree‑ring analysis on Windows, macOS, and Linux. Code is available at the https://hmarichal93.github.io/tras.
Authors:Vicent Caselles-Ballester, Eloy Martínez-Heras, Giuseppe Pontillo, Zoe Mendelsohn, Elena M. Marrón, Juan Luis García Fernández, Laia Subirats, Jon Stutters, Jeremy Chataway, Frederik Barkhof, Sara Llufriu, Ferran Prados
Abstract:
Multiple sclerosis (MS) expresses substantial clinical and radiological heterogeneity, which poses significant challenges for automatic lesion segmentation. The current deep learning‑based SOTA is highly susceptible to changes in both distribution, e.g., changes in scanner; as well as the structure of inputs, evident in the current divide between cross‑sectional and longitudinal approaches. We introduce TimeLesSeg, a unified contrast‑agnostic framework designed to segment MS lesions regardless of the presence of a temporal dimension in its inputs, with a single convolutional neural network. Our approach models pathological priors through lesion masks, which are processed together with the current scan. Cross‑sectional processing is enabled by exposing the model to training cases where no prior information is available, which are modeled with an empty mask, allowing it to operate seamlessly in both scenarios. To overcome the scarcity and inconsistency of longitudinal datasets, we propose a novel generative pipeline in which patterns of lesion evolution are simulated by stochastically deforming each individual lesion with morphological operations, producing realistic prior timepoints. In parallel, we achieve contrast agnosticism through Gaussian mixture model‑based domain randomization, enabling the network to experience a wide spectrum of intensity profiles. Results on three publicly available and two in‑house datasets show that TimeLesSeg outperforms the contrast‑agnostic state of the art on single‑modality inputs across overlap‑ and distance‑based metrics. In longitudinal processing, our method outperforms SAMSEG, and captures lesion load dynamics more accurately than both the former and LST‑AI. All source code related to the development of TimeLesSeg is available at https://github.com/NeuroADaS‑Lab/TimeLesSeg.
Authors:Giacomo Spigler
Abstract:
Active vision ‑‑ where a policy controls its own gaze during manipulation ‑‑ has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared benchmark to compare approaches or quantify what active vision contributes, on which task types, and under what conditions. We introduce TAVIS, evaluation infrastructure for active‑vision imitation learning, with two complementary task suites ‑‑ TAVIS‑Head (5 tasks, global search via pan/tilt necks) and TAVIS‑Hands (3 tasks, local occlusion via wrist cameras) ‑‑ on two humanoid torso embodiments (GR1T2, Reachy2), built on IsaacLab. TAVIS provides three evaluation primitives: a paired headcam‑vs‑fixedcam protocol on identical demonstrations; GALT (Gaze‑Action Lead Time), a novel metric grounded in cognitive science and HRI that quantifies anticipatory gaze in learned policies; and procedural ID/OOD splits. Baseline experiments with Diffusion Policy and π_0 reveal that (i) active‑vision generally helps, but benefits are task‑conditional rather than uniform; (ii) multi‑task policies degrade sharply under controlled distribution shifts on both suites; and (iii) imitation alone yields anticipatory gaze, with median lead times comparable to the human teleoperator reference. Code, evaluation scripts, demonstrations (LeRobot v3.0; ~2200 episodes) and trained baselines are released at https://github.com/spiglerg/tavis and https://huggingface.co/tavis‑benchmark.
Authors:Lang Zhang, JinYi Yoon, Matthew Corbett, Abhijit Sarkar, Bo Ji
Abstract:
Driver cognitive distraction is a major cause of road collisions and remains difficult to detect. Unlike manual or visual distraction, cognitive distraction is diverted by thoughts unrelated to driving, even when the driver appears visually attentive and exhibits no explicit physical movements. In this work, we propose EyeCue, a gaze‑empowered egocentric video understanding framework, to detect driver cognitive distraction. A key insight is that cognitive distraction manifests in the interaction between eye gaze and visual context. To capture this interaction, EyeCue integrates eye gaze with egocentric video to enable context‑aware modeling of the driver's attention over time. Furthermore, to tackle the limited scale and diversity of existing datasets, we introduce CogDrive, a comprehensive multi‑scenario dataset that augments four existing driving datasets with cognitive distraction annotations. Through extensive evaluations on CogDrive, we show that EyeCue achieves the highest accuracy of 74.38%, outperforming 11 baselines from 6 model families by over 7%. Notably, EyeCue can achieve an accuracy of over 70% across various driving scenarios (different road types, times of day, and weather conditions) with strong generalizability. These results highlight the importance of modeling gaze‑context interactions and the effectiveness of cross‑modal interaction modeling for multimodal cognitive distraction detection. Our codes and CogDrive dataset resources are available at https://github.com/langzhang2000/EyeCue.
Authors:Boyang Dai, Chaoqi Chen, Yizhou Yu
Abstract:
Out‑of‑distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regular entangled representations to discriminate in‑distribution (ID) and OOD data, neglecting the rich contextual information within images. This issue is particularly challenging for detecting near‑OOD, as models with simplicity bias struggle to learn discriminative features in disentangled representations. The human visual system can use the co‑occurrence of objects in the natural environment to facilitate scene understanding. Inspired by this, we propose an Object‑Centric OOD detection framework that learns to capture Object CO‑occurrence (OCO) patterns within images. The proposed method introduces a new OOD detection paradigm that understands object co‑occurrence within an image by predicting disentangled representations for the test sample, then adaptively divides patterns into three scenarios based on object co‑occurrence patterns observed in ID training data, and finally performs OOD detection in a divide‑and‑conquer manner. By doing so, OCO can distinguish near‑OOD by considering the semantic contextual relationships present in their images, avoiding the tendency to focus solely on simple, easily learnable regions. We evaluate OCO through experiments across challenging and full‑spectrum OOD settings, demonstrating competitive results and confirming its ability to address both semantic and covariate shifts. Code is released at https://github.com/Michael‑McQueen/OCO.
Authors:Hugo Vigna, Samuel Bontemps
Abstract:
The Hessian spectrum of trained deep networks exhibits a characteristic structure: a continuous bulk of near‑zero eigenvalues and a small number of large outlier eigenvalues (spikes), confirming the relevance of Random Matrix Theory in deep learning. The spike count matches the number of classes minus one. While prior work has described this structure, no method has exploited it operationally to improve classification performance. We propose Hessian Surgery, a post‑hoc optimization method that directly perturbs model weights along spike eigenvectors to rebalance per‑class accuracy without retraining. We introduce (i) a spike‑class sensitivity matrix that quantifies the directional derivative of each class's accuracy along each spike eigenvector, (ii) a constrained optimization of perturbation coefficients that targets weak classes while preserving strong ones, and (iii) an adaptive amplitude control that raises or lowers the perturbation budget based on iteration‑level improvement signals. We obtain encouraging results on CIFAR‑10 and ISIC‑2019 on both balanced accuracy and standard deviation.
Authors:Ritul Jangir, Arkya Jyoti Bagchi, Aiman Farooq, Mangalton Okram, Saurabh Seetaram Korgaonkar, Deepak Mishra
Abstract:
High‑fidelity surgical video generation can greatly improve medical training and the development of AI, adapting these generative models for precise video editing remains a formidable challenge. Modifying surgical attributes, such as instrument tissue interactions or procedural phases is challenging due to the strict anatomical and temporal constraints. In this paper, we propose OphEdit, a novel training‑free framework for the text‑guided editing of ophthalmic surgical videos. Our approach leverages a deterministic second‑order ODE inversion pipeline to capture Attention Value (V) tensors from the original video. By selectively injecting these stored tensors into the conditional Classifier‑Free Guidance (CFG) branch during the denoising phase, OphEdit rigorously preserves the intricate anatomical geometry of the eye while seamlessly mapping text‑driven semantic modifications onto the video stream. Clinical evaluations demonstrates that OphEdit effectively handles complex surgical transformations, such as instrument swaps and procedural variations, with superior structural fidelity and temporal consistency compared to natural‑domain video editors. Our work represents the first application of training‑free video editing in the ophthalmic surgical domain, offering a scalable solution for generating diverse, annotated medical datasets without the need for exhaustive manual recording or costly model fine‑tuning. The code and prompts can be accessed at https://github.com/ophedit/OphEdit
Authors:Pei Zhang, Yunkai Liang, Kaiqiang Wang
Abstract:
Underwater environments impose severe constraints on conventional imaging systems and demand solutions that balance high‑quality sensing with strict resource efficiency. While emerging event cameras offer a promising alternative, their potential in aquatic scenarios remains largely unexplored. Through the lens of neuromorphic vision, this work pioneers the investigation of motion fields that serve as key media for agile underwater perception. Built upon spiking neural networks, we introduce a self‑supervised framework to estimate per‑pixel optical flow from asynchronous event streams, elegantly bypassing the long‑standing bottleneck of underwater data scarcity. Extensive evaluations demonstrate that our method achieves competitive visual and quantitative results against leading techniques while operating with superior computational efficiency. By bridging neuromorphic sensing and aquatic intelligence, this work opens new frontiers for lightweight, real‑time, and low‑cost perception on resource‑constrained underwater edge platforms.
Authors:Zijia Fu, Yuanfei Huang, Lizhi Wang, Hua Huang
Abstract:
Lens flares, caused by complex optical aberrations, severely degrade image quality especially in nighttime photography. Although recent restoration methods have made remarkable progress, most still rely on spatially uniform processing. They are failing to handle the region‑dependent restoration demands of flare scenes, where saturated light sources should be preserved, flare artifacts removed, and background details recovered. To address this challenge, we propose DeflareMambav2, a prior‑guided Mamba framework for lens flare removal. Specifically, we introduce a Flare Prior Network (FPN) to estimate flare priors and guide adaptive restoration. Besides, a novel radial serialization strategy breaks spatially homogeneous processing by performing flare‑aware targeted sampling, and better supports long‑range modeling in State Space Models (SSMs). Based on these priors, the backbone adopts a dual‑level adaptive scheme. It explicitly preserves light‑source regions to avoid over‑processing, and applies curriculum‑based restoration to the remaining contaminated areas while calibrating restoration intensity at the pixel level. Extensive experiments demonstrate that DeflareMambav2 achieves state‑of‑the‑art performance with reduced parameter burden. Code is available at https://github.com/BNU‑ERC‑ITEA/DeflareMambav2.
Authors:Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park
Abstract:
Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human‑robot interaction. However, this task remains a highly challenging problem because egocentric hand motion is driven by complex human intent, exhibits highly dexterous articulations, and is observed under drastic viewpoint shifts induced by ego‑motion. In this work, we introduce EggHand, a foundation‑model‑based framework for egocentric hand pose forecasting that unifies multimodal semantic reasoning with dynamic motion modeling. Our approach couples an action decoder from a Vision‑Language‑Action (VLA) model, which captures the structured temporal dynamics of hand motion, with an egocentric video‑text encoder that provides viewpoint‑aware contextual information learned from large‑scale first‑person video. Together, these components overcome the brittleness of generic visual encoders under ego‑motion and enable joint reasoning over motion, context, and high‑level intent‑without relying on body pose or external tracking. Experiments on the EgoExo4D dataset show that EggHand sets a new state of the art in forecasting accuracy, remains robust under severe ego‑motion, and further enables controllable prediction via language‑based task prompts. Project page: https://jyoun9.github.io/EggHand
Authors:Zepeng Yang, Junxuan Bai, Hao Li, Ju Dai, Junjun Pan, Yongfeng Yin, Bin Li
Abstract:
The rapid advances in deep learning have significantly enhanced the accuracy of multimodal 3D human pose estimation (HPE). However, the state‑of‑the‑art (SOTA) HPE pipelines still rely on Transformers, whose quadratic complexity makes real‑time processing for long sequences impractical. Mamba addresses this issue through selective state‑space modeling, enabling efficient sequence processing without sacrificing representational power. Nevertheless, it struggles to capture complex spatial dependencies in multimodal settings. To bridge this gap, we propose VIMCAN, a hybrid architecture that combines the efficient sequence modeling of Mamba with the spatial reasoning of Cross‑Attention, and performs robust visual‑inertial fusion and human pose estimation between RGB keypoints and wearable IMU data. By leveraging Mamba's dynamic parameterization for temporal modeling and Attention for spatial dependency extraction, VIMCAN achieves superior accuracy, with mean per‑joint position errors (MPJPE) of 17.2 mm on TotalCapture and 45.3 mm on 3DPW. VIMCAN outperforms prior Transformer‑based and other SOTA approaches while supporting real‑time inference at over 60 frames per second on consumer‑grade hardware. The source code is available at \hrefhttps://github.com/Eddieyzp/VIMCANthis GitHub repository.
Authors:Grzegorz Wilczynski, Mikołaj Zielinski, Bartosz Świrta, Dominik Belter, Przemysław Spurek
Abstract:
3D vision systems are fundamentally constrained by their reliance on visual overlap: reconstruction methods require it for geometric alignment, while generative models use it to enforce multi‑view consistency. This limitation is particularly acute in real‑world scenarios such as distributed swarm robotics or crowd‑sourced data collection, where capturing overlapping perspectives, both in terms of spatial and appearance overlap, is often impossible. We introduce Generative Reconstruction from Disjoint Views as a new paradigm, establish a comprehensive dataset, and propose specialized evaluation metrics for zero‑overlap scenarios. Our benchmarking demonstrates that existing state‑of‑the‑art methods fail catastrophically on this task, producing disconnected geometries or semantically incoherent reconstructions. To address these limitations, we propose GLADOS, a general, modular framework that operates through three stages: (1) Generative Bridging, where foundation models synthesize intermediate perspectives to connect disjoint inputs; (2) Robust Coarse 3D Reconstruction, that establish coarse geometric scaffold via global alignment which absorbs local contradictions from generative process; and (3) Iterative Context Expansion and Consistency Optimization to fill missing regions and unify the reconstruction. As an architectureagnostic framework, GLADOS enables seamless integration of future advances in generation, reconstruction, and inpainting. The source code is available at: https://github.com/gwilczynski95/GLADOS.
Authors:Christopher Ries, Moussa Kassem Sbeyti, Nicolas Bianco, Nadja Klein
Abstract:
Conformal Prediction (CP) is a distribution‑free method for constructing prediction sets with marginal finite‑sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety‑critical object detection. However, object detection introduces structured multi‑output predictions, complicating the application of classical CP theory developed for single outputs. In addition, standard, unscaled CP produces fixed‑width prediction intervals across inputs, leading to unnecessary width for low‑uncertainty predictions. While scaled CP addresses this by adapting the interval width to an input‑dependent uncertainty estimate, prior work has neither systematically compared unscaled and scaled CP for multi‑class object detection, nor integrated CP with a complementary uncertainty quantification method in this setting. We fill this gap by: (i) applying CP coordinate‑wise to bounding box corners with a Bonferroni correction for box‑level guarantees; (ii) scaling the resulting intervals using per‑prediction aleatoric uncertainty estimates derived from a probabilistic object detector trained with loss attenuation, evaluated in uncalibrated and two calibrated variants; (iii) extending to a two‑step pipeline that constructs prediction sets for the class using RAPS and conditions the conformalized bounding boxes on the predicted class set. Across three autonomous driving datasets (KITTI, BDD, CODA), including a cross‑domain setting under distribution shift, scaled CP consistently improves interval sharpness over unscaled CP, achieving up to 19% higher IoU and 39% lower interval scores, without sacrificing coverage. Class‑wise calibration further improves coverage for both variants with a negligible effect on sharpness. Together, these improvements yield more actionable uncertainty estimates for real‑time, real‑world object detection.
Authors:Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng, Bing Ma, Kai Yu, Tianxiang Zheng, Qinglin Lu, Zhen Cui
Abstract:
Human image animation has witnessed significant advancements, yet generating high‑fidelity hand motions remains a persistent challenge due to their high degrees of freedom and motion complexity. While reinforcement learning from human feedback, particularly direct preference optimization, offers a potential solution, it necessitates the construction of strict preference pairs. However, curating such pairs for dynamic hand regions is prohibitively expensive and often impractical due to frame‑wise inconsistencies. In this paper, we propose Implicit Preference Alignment (IPA), a data‑efficient post‑training framework that eliminates the need for paired preference data. Theoretically grounded in implicit reward maximization, IPA aligns the model by maximizing the likelihood of self‑generated high‑quality samples while penalizing deviations from the pretrained prior. Furthermore, we introduce a Hand‑Aware Local Optimization mechanism to explicitly steer the alignment process toward hand regions. Experiments demonstrate that our method achieves effective preference optimization to enhance hand generation quality, while significantly lowering the barrier for constructing preference data. Codes are released at https://github.com/mdswyz/IPA
Authors:Bohan Hou, Jiuning Gu, Jiayan Guo, Ronghao Dang, Sicong Leng, Xin Li, Xuemeng Song, Jianfei Yang
Abstract:
Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoint rather than part of an interleaved search trajectory. We introduce InterLV‑Search, a benchmark for Interleaved Language‑Vision Agentic Search, in which textual and visual evidence is repeatedly used to condition later search. It contains 2,061 examples across three levels: active visual evidence seeking, controlled offline interleaved multimodal search, and open‑web interleaved multimodal search. Beyond existing benchmarks, it also includes multimodal multi‑branch samples that involve comparison between multiple entities during the evidence search. We construct Level 1 and Level 2 with automated pipelines and Level 3 with a machine‑led, human‑supervised open‑web pipeline. We further provide InterLV‑Agent for standardized tool use, trajectory logging, and evaluation. Experiments on proprietary and open‑source multimodal agents show that current systems remain far from solving interleaved multimodal search, with the best model below 50% overall accuracy, highlighting challenges in visual evidence seeking, search control, and multimodal evidence integration. We release the benchmark data and evaluation code at https://github.com/hbhalpha/InterLV‑Search‑Bench
Authors:Honghua Chen, Zitong Xu, Huiyu Duan, Xinyun Zhang, Xiongkuo Min, Guangtao Zhai
Abstract:
Recent text‑guided image editing (TIE) models have achieved remarkable progress, however, many edited results still suffer from artifacts, unintended modifications, and suboptimal aesthetics. Although several benchmarks and evaluation methods have been proposed, most existing approaches rely on scalar scores and lack interpretability. This limitation largely stems from the absence of high‑quality interpretation datasets for TIE and effective reward models to train interpretable evaluators. To address these challenges, we introduce ReasonEdit‑22K, the first dataset that combines 22K edited images with 113K Chain‑of‑Thought (CoT) samples, along with 1.3M human judgments assessing these interpretations in terms of logicality, accuracy, and usefulness. Building upon this dataset, we propose RE‑Reward, a multimodal large language model (MLLM)‑based reward model designed to provide human‑aligned feedback for evaluating interpretable reasoning in image editing. Furthermore, we develop ReasonEdit, which is trained using reward signals derived from RE‑Reward and the Group Relative Policy Optimization (GRPO) algorithm to learn an interpretable evaluation model. Extensive experiments demonstrate that ReasonEdit achieves superior alignment with human preferences and exhibits strong generalization across public benchmarks. In addition, it is capable of generating high‑quality interpretable evaluation text, enabling more transparent and trustworthy assessment for image editing. The code is available at https://github.com/IntMeGroup/ReasonEdit.
Authors:Zitong Xu, Huiyu Duan, Yifei Nie, Mingda Du, Sijing Wu, Xiongkuo Min, Tianyi Zheng, Jian Zhang, Shusong Xu, Jinwei Chen, Bo Li, Guangtao Zhai
Abstract:
Recent text‑guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine‑grained issues such as unnatural objects, lighting mismatch, and unexpected changes. Existing refinement approaches either rely on costly iterative regeneration or employ vision‑language models (VLMs) with weak spatial grounding, often resulting in semantic drift and unreliable local corrections. To address these limitations, we first construct EditFHF‑15K, a dataset of fine‑grained human feedback for edited images, comprising (1) 15K images from 12 TIE models spanning 43 editing tasks, (2) 60K annotated artifact regions and 80K editing failure regions, each accompanied by textual reasoning, and (3) 45K mean opinion scores (MOSs) assessing perceptual quality, instruction following, and visual consistency. Based on EditFHF‑15K, we propose EditRefiner, a hierarchical, interpretable, and human‑aligned agentic framework that reformulates post‑editing correction as a human‑like perception‑reasoning‑action‑evaluation loop. Specifically, we introduce: (1) a perception agent that detects contextual saliency maps of artifacts and editing failures, (2) a reasoning agent that interprets these perceptual cues to perform human‑aligned diagnostic inference, (3) an action agent that uses the reasoning output to plan and execute localized re‑editing, and (4) an evaluation agent that assesses the re‑edited image and guides the action agent on whether further refinements are required. Extensive experiments demonstrate that EditRefiner consistently outperforms state‑of‑the‑art methods in distortion localization, diagnose accuracy and human perception alignment, establishing a new paradigm for self‑corrective and perceptually reliable image editing. The code is available at https://github.com/IntMeGroup/EditRefiner.
Authors:Linxiao Shi, Siming Zheng, Zerong Wang, Hao Zhang, Jinwei Chen, Bo Li, Shifeng Chen, Peng-Tao Jiang
Abstract:
Existing mobile devices are constrained by compact optical designs, such as small apertures, which make it difficult to produce natural, optically realistic bokeh effects. Although recent learning‑based methods have shown promising results, they still struggle with photos captured under high digital zoom levels, which often suffer from reduced resolution and loss of fine details. A naive solution is to enhance image quality before applying bokeh rendering, yet this two‑stage pipeline reduces efficiency and introduces unnecessary error accumulation. To overcome these limitations, we propose MagicBokeh, a unified diffusion‑based framework designed for high‑quality and efficient bokeh rendering. Through an alternative training strategy and a focus‑aware masked attention mechanism, our method jointly optimizes bokeh rendering and super‑resolution, substantially improving both controllability and visual fidelity. Furthermore, we introduce degradation‑aware depth module to enable more accurate depth estimation from low‑quality inputs. Experimental results demonstrate that MagicBokeh efficiently produces photorealistic bokeh effects, particularly on real‑world low‑resolution images, paving the way for future advancements in bokeh rendering. Our code and models are available at https://github.com/vivoCameraResearch/MagicBokeh.
Authors:Fengqiang Wan, Yipeng Lin, Kan Lv, Yang Yang
Abstract:
Pre‑trained models with parameter‑efficient fine‑tuning (PEFT) have demonstrated promising potential for class‑incremental learning (CIL), yet catastrophic forgetting still persists when adapting models to new tasks. In this paper, we present a novel perspective on catastrophic forgetting through the analysis of inter‑layer relation drift, i.e., the progressive disruption of relationships among layer‑wise representations during the learning of new tasks. We theoretically show that the increase of such drift reduces the classification margins of previously learned tasks, thereby degrading overall model performance. To address this issue, we propose \underlineSelf‑\underlineRectifying inter‑layer \underlineRelation Low‑Rank Adaptation~(SR^2‑LoRA), a simple yet effective method that mitigates catastrophic forgetting by constraining inter‑layer relation drift. Specifically, SR^2‑LoRA constructs the relation matrices induced by the previous and current models on current‑task samples, and aligns the corresponding singular values. We further theoretically show that this alignment exhibits greater robustness to estimation perturbations than direct entry‑wise alignment. Extensive experiments on standard CIL benchmarks demonstrate that SR^2‑LoRA effectively mitigates catastrophic forgetting, with its advantages becoming more pronounced as the number of tasks increases. Code is available in the \hrefhttps://github.com/FqWan24/SR‑2‑LoRArepository.
Authors:Shuai Zhang, Zhecheng Shi, Zhuxiao Li, Jing Ou, Tengxi Wang, Yuan Liu, Wufan Zhao
Abstract:
Semantic segmentation of large‑scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse sampling pattern of LiDAR and the view‑dependent geometric distortion in image observations complicate cross‑modal alignment and hinder stable fusion. Inspired by the fact that 2D images captured by cameras are representations of the 3D world, we recognize that the features learned from 2D and 3D segmentation share some common semantics, while other aspects remain modality‑specific. This insight motivates a unified multimodal framework for joint 2D‑3D semantic segmentation. We combine a SAM‑based vision encoder with a SPTNet‑based geometric encoder to extract complementary semantic and geometric representations. The resulting features from both modalities are explicitly decomposed into shared and private subspaces, where the shared components summarize semantic factors common to both domains, and the private components preserve properties that are unique to each modality. A lightweight attention‑based fusion module aggregates the shared features into a consistent cross‑modal representation, and a regularized training objective ensures both semantic alignment and subspace independence. Experiments on the SemanticKITTI and nuScenes benchmarks demonstrate consistent improvements in segmentation accuracy over representative multimodal baselines, accompanied by competitive computational efficiency. Cross‑domain evaluation on nuScenes USA‑Singapore shows stable performance under distribution shifts, demonstrating strong generalization. The implementation code is publicly available at: https://github.com/shuaizhang69/UniD‑Shift.
Authors:Simin Huo, Ning LI
Abstract:
Video‑language models (VLMs) face rapid inference costs as visual token counts scale with video length. For example, 32 frames at 448×448 resolution already yield >8,000 visual tokens in Qwen3‑VL, making LLM prefill the dominant throughput bottleneck. Existing methods often rely on global similarity or attention‑guided compression, incurring offsets to their gains. We propose Temporal Token Fusion (TTF), a training‑free, plug‑and‑play pre‑LLM token compression framework that exploits structured temporal redundancy in video. TTF automatically selects an anchor frame, then for each subsequent frame, performs a local window similarity search (e.g.,3× 3), fusing tokens that exceed a threshold. The compressed sequence maintains positional consistency across both prefill and decoding through coordinate realignment, enabling seamless integration with existing VLM pipelines. On Qwen3‑VL‑8B with threshold t=0.70, TTF removes about 67% of visual tokens while retaining 99.5% of the baseline accuracy and introducing only \approx0.16\,GFLOPs of matching overhead. Overall, TTF offers a practical, efficient solution for video understanding. The code is available at \hrefhttps://github.com/Cominder/ttfhttps://github.com/Cominder/ttf
Authors:Junwei Wen, Deshui Miao, Guangming Lu, Xin Li, Wenjie Pei
Abstract:
Video Reasoning Segmentation (VRS) aims to segment target objects in videos based on implicit instructions that convey human intent and temporal logic. Existing MLLM‑based methods predict masks with a [SEG] token after selecting frames via simple sampling or an auxiliary MLLM, where limited supervision and frame‑language similarity rules often yield narrow‑scope keyframe choices that weaken holistic temporal understanding and lead to brittle localization in complex multi‑object scenes. To address these issues, we introduce RCoT‑Seg, a video‑of‑thought framework that factorizes VRS into temporal video reasoning (TVR) and keyframe target perception (KTP), explicitly separating temporal reasoning from spatial perception. Specifically, in the TVR stage, an agentic keyframe selection module, initialized with a curated CoT‑start corpus and refined by GRPO under task‑aligned rewards, is proposed to generate and reselect the keyframe through self‑evaluation, strengthening moment localization and temporal reasoning. In the KTP stage, RCoT‑Seg performs high‑resolution segmentation on the selected frame and propagates masks with SAM2‑based methods across the sequence, replacing heuristic sampling and external selectors while improving spatial precision and inter‑frame consistency. Extensive experimental results demonstrate that the proposed RCoT‑Seg achieves favorable performance against the state‑of‑the‑art methods. The code and models will be publicly released at https://github.com/Victor‑wjw/RCoT‑Seg.
Authors:Yang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu, Renliang Weng, Jianjun Qian, Jian Yang, Jin Xie
Abstract:
World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, progress in LiDAR‑based world models has lagged behind those built on camera videos or occupancy data, primarily due to two core challenges: the inherent disorder of LiDAR point clouds and the difficulty of distinguishing dynamic objects from static structures. To address these issues, we propose GEM: a Generative LiDAR world model that leverages deformable mamba architecture, significantly improving fidelity and imaginative capability. Specifically, leveraging the structural similarity between sequential laser scanning and Mamba's processing mechanism, we first tokenize LiDAR sweeps into compact representations via a custom LiDAR scene tokenizer. After unsupervised disentanglement of tokenized features via a dynamic‑static separator, a tri‑path deformable Mamba is introduced to perform selective scanning and adaptive gating fusion over the disentangled features, leading to enhanced spatial‑temporal understanding of the world evolution. Optionally, a planner and a BEV layout controller can be integrated to explore the model's capability for autonomous rollout and its potential to generate ``what‑if" scenarios. Extensive experiments show that GEM achieves state‑of‑the‑art performances across diverse benchmarks and evaluation settings, demonstrating its superiority and effectiveness. Project page: https://github.com/wuyang98/GEM.
Authors:Yecong Wan, Fan Li, Mingwen Shao, Wangmeng Zuo
Abstract:
Generalizable novel view synthesis aims to render unseen views from uncalibrated input images without requiring per‑scene optimization. Recent feed‑forward approaches based on 3D Gaussian Splatting have achieved promising efficiency and rendering quality. However, most of them assign a fixed number of Gaussians to each pixel or voxel, ignoring the spatially varying complexity of real‑world scenes. Such uniform allocation often wastes Gaussian primitives in smooth regions while providing insufficient capacity for fine structures, complex geometry, and high‑frequency details. This motivates us to predict region‑dependent primitive cardinalities rather than impose a fixed primitive budget everywhere, enabling a more expressive 3D scene representation. Therefore, we propose SplatWeaver, a generalizable novel view synthesis framework that is able to dynamically allocate Gaussian primitives over different regions in a feed‑forward manner. Specifically, SplatWeaver introduces cardinality Gaussian experts and a pixel‑level routing scheme, wherein each expert specializes in producing a specific number of primitives from 0 to M, and the routing scheme coordinates these experts to adaptively determine how many Gaussian primitives should be allocated to each spatial location. Moreover, SplatWeaver incorporates a high‑frequency prior with attendant guidance module and routing regularization to stabilize expert selection and promote complexity‑aware allocation. By leveraging high‑frequency cues, the routing process is encouraged to assign more Gaussian primitives to fine structures and textured regions, while suppressing redundancy in smooth areas. Extensive experiments across diverse scenarios show that SplatWeaver consistently outperforms state‑of‑the‑art methods, delivering more faithful novel‑view renderings with fewer Gaussian primitives. Project Page: https://yecongwan.github.io/SplatWeaver/
Authors:Wenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama
Abstract:
A latent world model may achieve accurate short‑horizon prediction while still inducing a latent space that is poorly aligned with planning. A key issue is spatiotemporal mismatch: these models are often trained with local predictive supervision, but deployed for long‑horizon goal‑directed search in latent spaces where Euclidean distance may not reflect what is reachable within a finite action budget. We present the Reachability‑Correction auxiliary objective (RC‑aux), a lightweight correction for this mismatch in reconstruction‑free latent world models. RC‑aux keeps the world‑model backbone unchanged and adds planning‑aligned supervision along two axes. Along the time axis, multi‑horizon open‑loop prediction trains the model beyond one‑step consistency. Along the space axis, budget‑conditioned reachability supervision, together with temporal hard negatives, encourages the latent space to distinguish states that are eventually reachable from those reachable within the current planning horizon. At test time, the learned reachability signal can also be used by a reachability‑aware planner to favor trajectories that are both goal‑directed and attainable under the available budget. We instantiate RC‑aux on LeWorldModel and evaluate it under both continuation‑training and matched‑from‑scratch settings. Across goal‑conditioned pixel‑control tasks and a LIBERO‑Goal extension, RC‑aux improves LeWM‑style planning with modest additional cost. These results suggest that planning with latent world models depends not only on predictive accuracy, but also on whether the learned representation encodes the temporal and geometric structure required by downstream search. The code is available at https://github.com/Guang000/RC‑aux.
Authors:Junchuan Zhao, Qifan Liang, Ye Wang
Abstract:
Co‑speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user‑specified gestural style. Existing VQ‑VAE based co‑speech gesture generation methods improve generation quality but fail to encode semantic structure into the motion representation or explicitly disentangle content from style, limiting both semantic coherence and personalization fidelity. We present PersonaGest, a two‑stage framework addressing both limitations. In the first stage, a semantic‑guided RVQ‑VAE disentangles motion content and gestural style within the residual quantization structure, where a Semantic‑Aware Motion Codebook (SMoC) organizes the content codebook by gesture semantics and contrastive learning further enforces content‑style separation. In the second stage, a Masked Generative Transformer generates content tokens via a semantic‑aware re‑masking strategy, followed by a cascade of Style Residual Transformers conditioned on a reference motion prompt for style control. Extensive experiments demonstrate state‑of‑the‑art performance on objective metrics and perceptual user studies, with strong style consistency to the reference prompt. Our project page with demo videos is available at https://danny‑nus.github.io/PersonaGest/
Authors:Qianwen Ma, Yang Xu, Shangwei Deng, Xiaobo Li, Haofeng Hu
Abstract:
Infrared small target detection (IRSTD) remains challenging due to the scarcity of useful target cues and the presence of severe background clutter. Most current methods rely on conventional feature learning and local interaction modeling, where features are represented in Euclidean space. However, such designs may still be limited in describing the subtle differences of weak targets and the contextual relations between targets and backgrounds. To address these limitations, we propose LoHGNet, an IRSTD network that integrates Lorentz geometric encoding with high‑order relation learning. By introducing Lorentz manifold based feature learning, LoHGNet offers a different feature representation from conventional IRSTD methods and provides new discriminative cues for IRSTD. Specifically, a Lorentz encoding branch is constructed with the Geometric Attention Guided Lorentz Residual Convolution Module (GA‑LRCM) to perform feature modeling under hyperbolic geometric constraints and enhance the hierarchical geometric representation capability of weak targets. Subsequently, the hyperbolic features are mapped into the Euclidean tangent space through logarithmic mapping, and a High‑Order Relation Learning Module (HORL) is designed to model the high‑order contextual dependencies between targets and backgrounds via hypergraph construction, thereby improving target discrimination in complex backgrounds. Experimental results on three datasets demonstrate that the proposed LoHGNet achieves competitive performance in both detection accuracy and adaptability to complex scenes. The code will be available at https://github.com/Kingwin97.
Authors:Chamuditha Jayanga Galappaththige, Jason Lai, Timothy Patten, Donald Dansereau, Niko Suenderhauf, Dimity Miller
Abstract:
Scene change detection methods built on Gaussian splatting universally follow a render‑then‑compare paradigm: the pre‑change scene is rendered into 2D and compared against post‑change images via pixel or feature residuals. This change detection problem with Gaussian Splatting has been treated as a question about pixels; we treat it as a question about primitives. We provide direct evidence that native primitive attributes alone ‑‑ position, anisotropic covariance, and color ‑‑ carry sufficient signal for scene change detection. What makes primitive‑space comparison hard is the under‑constrained nature of Gaussian splatting representation: independent optimizations yield primitive solutions whose count, positions, shapes, and colors differ even where nothing has changed. We address this challenge with anisotropic models of geometric and photometric drift, complemented by a per‑primitive observability term that reflects the extent to which each Gaussian is constrained by the camera geometry. Operating directly on primitives gives our method, GD‑DIFF, two properties that distinguish it from render‑then‑compare methods. First, change maps are multi‑view consistent by construction, where prior work had to learn this through an additional optimization objective. Second, geometric and appearance changes are scored separately, identifying not just where but what kind of change occurred, distinguishing structural changes (e.g., an added object) from surface‑level ones (e.g., a color change) without supervision or external model dependencies. On real‑world benchmarks, GS‑DIFF surpasses the prior state‑of‑the‑art approach by ~17% in mean Intersection over Union.
Authors:Jun Dai, Renbiao Jin, Bo Xu, Yutian Chen, Linning Xu, Mulin Yu, Tianfan Xue, Shi Guo
Abstract:
3D reconstruction methods such as 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) achieve impressive photorealism but fail when input images suffer from severe motion blur. While event cameras provide high‑temporal‑resolution motion cues, existing event‑assisted approaches rely on low‑resolution sensors and strict synchronization, limiting their practicality for handheld 3D capture on common devices, such as smartphones. We introduce a flexible, high‑resolution asynchronous RGB‑Event dual‑camera system and a corresponding reconstruction framework. Our approach first reconstructs sharp images from the event data and then employs a cross‑domain pose estimation module based on the Visual Geometry Transformer (VGGT) to obtain robust initialization for 3DGS. During optimization, we employ a structure‑driven event loss and view‑specific consistency regularizers to mitigate the ill‑posed behavior of traditional event losses and deblurring losses, ensuring both stable and high‑fidelity reconstruction. We further contribute AsyncEv‑Deblur, a new high‑resolution RGB‑Event dataset captured with our asynchronous system. Experiments demonstrate that our method achieves state‑of‑the‑art performance on both our challenging dataset and existing benchmarks, substantially improving reconstruction robustness under severe motion blur. Project page: https://openimaginglab.github.io/AsyncEvGS/
Authors:Han Jang, Junhyeok Lee, Heeseong Eum, Joon Jang, Yoseob Han, Seung Hong Choi, Kyu Sung Choi
Abstract:
Precise molecular subtyping of gliomas, including isocitrate dehydrogenase (IDH) mutation and 1p/19q codeletion, directly guides surgical and therapeutic decisions, yet currently relies on invasive tissue sampling. Deep learning on structural MRI has emerged as a non‑invasive alternative, but anatomy‑only approaches cannot capture the hemodynamic signatures that distinguish molecular subtypes. Radiogenomics based on dynamic susceptibility contrast (DSC) MRI holds immense potential for non‑invasively characterizing glioma molecular subtypes, yet clinical deployment has been hindered by inter‑site variability and the limitations of voxel‑wise analysis. We introduce HiPerfGNN, a framework that first learns discrete hemodynamic representations from raw time‑intensity curves using a vector‑quantized variational autoencoder (VQ‑VAE). These quantized perfusion codes define coarse‑level graph nodes representing functional tumor habitats, each of which is hierarchically subdivided into fine‑level subregions guided by structural MRI. A hierarchical graph neural network then propagates information across scales for molecular prediction. On an internal cohort (n=475), the model achieved AUCs of 0.96 (IDH), 0.89 (1p/19q), and 0.84 (WHO grade), and maintained robust IDH performance (AUC 0.89) on an independent external cohort (n=397) without recalibration. Gradient‑based saliency analysis confirms biologically grounded attention patterns aligned with known glioma pathophysiology. Our results demonstrate the added value of integrating perfusion dynamics into radiogenomic pipelines for glioma molecular subtyping. Code is available at https://github.com/janghana/HiPerfGNN.
Authors:Haoming Wang, Wei Gao
Abstract:
Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology‑preserving representations of 3D space. While modern Vision‑Language Models (VLMs) demonstrate emergent spatial reasoning from 2D egocentric inputs, it remains unclear whether they construct an analogous 3D internal representation. In this paper, we demonstrate that current VLMs do possess a latent topological map of 3D scenes, but it is heavily overshadowed by non‑geometric visual semantics, such as color and shape. By isolating this spatial subspace through cross‑scene linear feature extraction, we extract a clean spatial subspace that causally controls the model's spatial outputs. We mathematically shape this latent representation and prove its correspondence to the Laplacian eigenmaps of the scene's 3D Gaussian‑kernel graph, converging to the physical 3D space in the continuous limit. Motivated by this geometric identification, we further introduce a mathematically principled latent regularization method for VLMs, based on Dirichlet energy. Applying this single‑term regularizer to a minimal 500‑step supervised VLM fine‑tuning (SFT) on simple synthetic data yields significant improvements on real‑world spatial benchmarks, outperforming standard SFT and competitive baselines by up to 12.1% in spatial tasks involving scene topology understanding. Source code is available at https://github.com/pittisl/vlm‑latent‑shaping
Authors:Talha Ilyas, Deval Mehta, Zongyuan Ge
Abstract:
Skeleton‑based human activity recognition has achieved strong empirical performance, yet most existing models remain black boxes and difficult to interpret. In this work, we introduce a neurosymbolic formulation of skeleton‑based HAR that reframes action recognition as concept‑driven first‑order logical reasoning over motion primitives. Our framework bridges representation learning and symbolic inference by grounding first‑order logic predicates in learnable spatial and temporal motion concepts. Specifically, we employ a standard spatio‑temporal skeleton encoder to extract latent motion representations, which are then mapped to interpretable concept predicates via a spatio‑temporal concept decoder that explicitly separates pose‑centric and dynamics‑centric abstractions. These concept predicates are composed through differentiable first‑order logic layers, enabling the model to learn human‑readable logical rules that govern action semantics. To impose semantic structure on the learned concepts, we align skeleton representations with LLM‑derived descriptions of atomic motion primitives, establishing a shared conceptual space for perception and reasoning. Extensive experiments on NTU RGB+D 60/120 and NW‑UCLA demonstrate that our approach achieves competitive recognition performance while providing explicit, interpretable explanations grounded in logical structure. Our results highlight neurosymbolic reasoning as an effective paradigm for interpretable spatio‑temporal action understanding. Code: https://github.com/Mr‑TalhaIlyas/REASON
Authors:Xinyu Zhang, Zhengtong Xu, Yutian Tao, Yeping Wang, Yu She, Abdeslam Boularias
Abstract:
World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature‑based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to hallucination. However, current feature‑based approaches rely on direct regression, which leads to blurry or collapsed predictions in complex interactions, while generative modeling in high‑dimensional feature spaces still remains challenging. In this work, we discover that a new type of latent action representation, which we refer to as Residual Latent Action (RLA), can be easily learned from DINO residuals. We also show that RLA is predictive, generalizable, and encodes temporal progression. Building on RLA, we propose RLA World Model (RLA‑WM), which predicts RLA values via flow matching. RLA‑WM outperforms both state‑of‑the‑art feature‑based and video‑diffusion world models on simulation and real‑world datasets, while being orders of magnitude faster than video diffusion. Furthermore, we develop two robot learning techniques that use RLA‑WM to improve policy learning. The first one is a minimalist world action model with RLA that learns from actionless demonstration videos. The second one is the first visual RL framework trained entirely inside a world model learned from offline videos only, using a video‑aligned reward and no online interactions or handcrafted rewards. Project page: https://mlzxy.github.io/rla‑wm
Authors:Daniil Lisus, Cedric Le Gentil, Timothy D. Barfoot
Abstract:
This paper introduces Dr‑BA, a first‑of‑its‑kind radar bundle adjustment (BA) framework that operates directly on 2D spinning radar intensity images. Unlike camera or lidar sensors, radar is largely unaffected by precipitation, making it a critical modality for autonomous systems that require all‑weather robustness. Existing state estimation approaches using spinning radar typically extract sparse point clouds from range‑azimuth‑intensity measurements and apply point cloud alignment techniques to estimate vehicle motion, scene structure, or to localize within an existing map. In contrast, Dr‑BA uses the full radar returns from multiple scans to jointly estimate dense maps and sensor poses. By formulating the problem as a separable optimization, we derive an efficient and general solution that decouples pose estimation from mapping. In addition to solving the BA problem, this formulation naturally extends to direct radar‑only localization (DRL) within a previously built map. Dr‑BA achieves state‑of‑the‑art radar‑based BA and cross‑session localization performance, demonstrated on more than 200 km of on‑road data across five distinct routes. Our implementation is publicly available at https://github.com/utiasASRL/dr_ba.
Authors:Do Xuan Long, Yale Song, Min-Yen Kan, Tomas Pfister, Long T. Le
Abstract:
Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse over long horizons. We present A^2RD, an Agentic Auto‑Regressive Diffusion architecture that decouples creative synthesis from consistency enforcement. A^2RD formulates long video synthesis as a closed‑loop process that synthesizes and self‑improves video segment‑by‑segment through a Retrieve‑‑Synthesize‑‑Refine‑‑Update cycle. It comprises three core components: (i) Multimodal Video Memory that tracks video progression across modalities; (ii) Adaptive Segment Generation that switches among generation modes for natural progression and visual consistency; and (iii) Hierarchical Test‑Time Self‑Improvement that self‑improves each segment at frame and video levels to prevent error propagation. We further introduce LVBench‑C, a challenging benchmark with non‑linear entity and environment transitions to stress‑test long‑horizon consistency. Across public and LVBench‑C benchmarks spanning one‑ to ten‑minute videos, A^2RD outperforms state‑of‑the‑art baselines by up to 30% in consistency and 20% in narrative coherence. Human evaluations corroborate these gains while also highlighting notable improvements in motion and transition smoothness.
Authors:Kirill Trapeznikov, Gabriel Mancino-Ball, Jonathan Li, Paul Cummer, Jai Aslam, Danial Samadi Vahdati, Tai Nguyen, Matthew C. Stamm, Peter Bautista, Michael Davinroy, Laura Cassani, Jill Crisman
Abstract:
The proliferation of generative video technologies has intensified the need for reliable methods to detect and characterize synthetic media. To address this challenge, we organized the \hrefhttps://safe‑video‑2025.dsri.orgSAFE: Synthetic Video Detection Challenge, co‑located with the Authenticity and Provenance in the Age of Generative AI (APAI) Workshop at ICCV 2025. The competition invited participants to develop and evaluate algorithms capable of distinguishing real from synthetic videos under fully blind evaluation conditions with over 600 submissions from 12 teams over a 90 day span. Hosted on the Hugging Face platform, the challenge comprised two primary tasks: (1) detection of synthetic video content generated by diverse state‑of‑the‑art models, and (2) detection of synthetic content following common post‑processing operations such as resizing, re‑compression, motion blur and others. The challenge data consisted of 13 modern high quality synthetic video models with generated content matched to real videos from 21 diverse and challenge sources, all adding up to 20 hours of 6,000 video samples. This paper describes the challenge design, dataset construction, evaluation methodology, and outcomes, offering insights into the generalization and robustness of contemporary synthetic video detection methods. Our findings highlight measurable progress in cross‑generator generalization but also persistent vulnerabilities to post‑processing artifacts. https://safe‑video‑2025.dsri.org
Authors:Ernie Chu, Vishal M. Patel
Abstract:
Diffusion Transformers (DiTs) have achieved state‑of‑the‑art video generation quality, but they incur immense computational cost because standard inference applies the same number of denoising steps uniformly to every token in the sequence. It is well known that human vision ignores vast amounts of redundant motion. Why, then, do our densest models treat every spatiotemporal token with equal priority? In this paper, we introduce Heterogeneous Step Allocation (HSA), a training‑free inference algorithm that assigns varying step budgets to different spatiotemporal tokens based on their velocity dynamics. To resolve the resulting sequence‑length mismatch without sacrificing global context, HSA introduces a KV‑cache synchronization mechanism that allows active tokens to attend to the full sequence while entirely bypassing inactive tokens. Furthermore, we derive a cached Euler update that advances the latent states of skipped tokens in a single operation without additional model evaluations. We evaluate HSA on the Wan‑2 and LTX‑2 models for both text‑to‑video (T2V) and image‑to‑video (I2V) generation. Our results demonstrate that HSA significantly outperforms previous state‑of‑the‑art caching methods and the vanilla Flow Matching baseline, especially at aggressive acceleration regimes (e.g., 50% and 25% runtimes). Crucially, HSA achieves a superior quality‑runtime Pareto frontier without the need for expensive offline profiling, robustly preserving structural integrity and generation quality even under tight computational budgets.
Project page: https://ernestchu.github.io/hsa
Authors:Zhifeng Gu, Yuqi Wang, Bing Wang
Abstract:
Relative spatial relations provide a compact representation of spatial structure and are fundamental to relative spatial reasoning in 3D layout generation. Recent works leverage Multimodal Large Language Models (MLLMs) to infer such relations, but the inferred relations are often unreliable and are typically handled with post‑hoc heuristics. In this paper, we propose R^3L, a general framework that improves the reliability and consistency of relative spatial reasoning for 3D layout generation. Our key motivation is that multi‑hop reasoning requires repeated reference‑frame transformations, which accumulate errors in inferred relations and lead to semantic and metric drift. To mitigate this, we propose invariant spatial decomposition to break coupled relation chains, and consistent spatial imagination to promote self‑consistency through an imagine‑and‑revise loop. We further introduce supportive spatial optimization to ease pose optimization via global‑to‑local coordinate re‑parameterization. Extensive experiments across diverse scene types and instructions demonstrate that R^3L produces more physically feasible and semantically consistent layouts. Notably, our analysis shows that resolving frame‑induced inconsistencies is crucial for reliable multi‑hop relative spatial reasoning. The code is available at https://github.com/Neal2020GitHub/R3L.
Authors:Yufan Deng, Daquan Zhou
Abstract:
Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly annotated human activity data. We present HumanNet, a one‑million‑hour human‑centric video corpus that captures how humans interact with the physical world at scale. HumanNet spans both first‑person and third‑person perspectives and covers fine‑grained activities, human‑object interactions, tool use, and long‑horizon behaviors across diverse real‑world environments. Beyond raw video, the dataset provides interaction‑centric annotations, including captions, motion descriptions, and hand and body‑related signals, enabling motion‑aware and interaction‑aware learning. Beyond scale, HumanNet introduces a systematic data curation paradigm for embodied learning, where human‑centric filtering, temporal structuring, viewpoint diversity, and annotation enrichment are treated as first‑class design principles. This design transforms unstructured internet video into a scalable substrate for representation learning, activity understanding, motion generation, and human‑to‑robot transfer. We conduct a first‑step validation on the value of this design through controlled vision‑language‑action ablation: under a fixed set of validation data, continued training from the Qwen VLM model with 1000 hours of egocentric video drawn from HumanNet surpasses the continued training with 100 hours of real‑robot data from Magic Cobot, indicating that egocentric human video could be a scalable and cost‑effective substitute for robot data. By building this project, we aim to explore the opportunity to scale embodied foundation models using human‑centric videos, rather than relying solely on robot‑specific data.
Authors:Omar El Khalifi, Thomas Rossi, Oscar Fossey, Thibault Fouque, Ulysse Mizrahi, Philip Torr, Ivan Laptev, Fabio Pizzati, Baptiste Bellot-Gurlet
Abstract:
For artistic applications, video generation requires fine‑grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero‑shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per‑frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image‑to‑video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two‑phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose‑only guidance refines high‑frequency details without over‑constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose‑only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera‑consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.
Authors:Borui Zhang, Bo Zhang, Bo Wang, Wenzhao Zheng, Yuhao Cheng, Liang Tang, Yiqiang Yan, Jie Zhou, Jiwen Lu
Abstract:
GUI grounding is a critical capability for enabling GUI agents to execute tasks such as clicking and dragging. However, in complex scenarios like the ScreenSpot‑Pro benchmark, existing models often suffer from suboptimal performance. Utilizing the proposed Masked Prediction Distribution (MPD) attribution method, we identify that the primary sources of errors are twofold: high image resolution (leading to precision bias) and intricate interface elements (resulting in ambiguity bias). To address these challenges, we introduce Bias‑Aware Manipulation Inference (BAMI), which incorporates two key manipulations, coarse‑to‑fine focus and candidate selection, to effectively mitigate these biases. Our extensive experimental results demonstrate that BAMI significantly enhances the accuracy of various GUI grounding models in a training‑free setting. For instance, applying our method to the TianXi‑Action‑7B model boosts its accuracy on the ScreenSpot‑Pro benchmark from 51.9% to 57.8%. Furthermore, ablation studies confirm the robustness of the BAMI approach across diverse parameter configurations, highlighting its stability and effectiveness. Code is available at https://github.com/Neur‑IO/BAMI.
Authors:Weiqing Xiao, Hong Li, Xiuyu Yang, Houyuan Chen, Wenyi Li, Tianqi Liu, Shaocong Xu, Chongjie Ye, Hao Zhao, Beibei Wang
Abstract:
Recent advances have shown that large‑scale video diffusion models can be repurposed as neural renderers by first decomposing videos into intrinsic scene representations and then performing forward rendering under novel illumination. While promising, this paradigm fundamentally relies on accurate intrinsic decomposition, which remains highly unreliable for real‑world videos and often leads to distorted appearances, broken materials, and accumulated temporal artifacts during relighting. In this work, we present Relit‑LiVE, a novel video relighting framework that produces physically consistent, temporally stable results without requiring prior knowledge of camera pose. Our key insight is to explicitly introduce raw reference images into the rendering process, enabling the model to recover critical scene cues that are inevitably lost or corrupted in intrinsic representations. Furthermore, we propose a novel environment video prediction formulation that simultaneously generates relit videos and per‑frame environment maps aligned with each camera viewpoint in a single diffusion process. This joint prediction enforces strong geometric‑illumination alignment and naturally supports dynamic lighting and camera motion, significantly improving physical consistency in video relighting while easing the requirement of known per‑frame camera pose. Extensive experiments demonstrate that Relit‑LiVE consistently outperforms state‑of‑the‑art video relighting and neural rendering methods across synthetic and real‑world benchmarks. Beyond relighting, our framework naturally supports a wide range of downstream applications, including scene‑level rendering, material editing, object insertion, and streaming video relighting. The Project is available at https://github.com/zhuxing0/Relit‑LiVE.
Authors:Hao Dong, Hongzhao Li, Shupan Li, Muhammad Haris Khan, Eleni Chatzi, Olga Fink
Abstract:
Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation protocols. Current research is fragmented, with studies varying significantly across datasets, modality configurations, and experimental settings. Furthermore, existing benchmarks focus predominantly on action recognition, often neglecting critical real‑world challenges such as input corruptions, missing modalities, and model trustworthiness. This lack of standardization obscures a reliable assessment of the field's advancement. To address this issue, we introduce MMDG‑Bench, the first unified and comprehensive benchmark for MMDG, which standardizes evaluation across six datasets spanning three diverse tasks: action recognition, mechanical fault diagnosis, and sentiment analysis. MMDG‑Bench encompasses six modality combinations, nine representative methods, and multiple evaluation settings. Beyond standard accuracy, it systematically assesses corruption robustness, missing‑modality generalization, misclassification detection, and out‑of‑distribution detection. With 7, 402 neural networks trained in total across 95 unique cross‑domain tasks, MMDG‑Bench yields five key findings: (1) under fair comparisons, recent specialized MMDG methods offer only marginal improvements over ERM baseline; (2) no single method consistently outperforms others across datasets or modality combinations; (3) a substantial gap to upper‑bound performance persists, indicating that MMDG remains far from solved; (4) trimodal fusion does not consistently outperform the strongest bimodal configurations; and (5) all evaluated methods exhibit significant degradation under corruption and missing‑modality scenarios, with some methods further compromising model trustworthiness.
Authors:Jakub Stępień, Marcin Mazur, Jacek Tabor, Przemysław Spurek
Abstract:
Sparse Autoencoders (SAEs) have become an important tool in mechanistic interpretability, helping to analyze internal representations in both Large Language Models (LLMs) and Vision Transformers (ViTs). By decomposing polysemantic activations into sparse sets of monosemantic features, SAEs aim to translate neural network computations into human‑understandable concepts. However, common architectures such as TopK SAEs rely on a fixed sparsity level. They enforce the same number of active features (K) across all inputs, ignoring the varying complexity of real‑world data. Natural data often lies on manifolds with varying local intrinsic dimensionality, meaning the number of relevant factors can change significantly across samples. This suggests that a fixed sparsity level is not optimal. Simple inputs may require only a few features, while more complex ones need more expressive representations. Using a constant K can therefore introduce noise in simple cases or miss important structure in more complex ones. To address this issue, we propose SoftSAE, a sparse autoencoder with a Dynamic Top‑K selection mechanism. Our method uses a differentiable Soft Top‑K operator to learn an input‑dependent sparsity level k. This allows the model to adjust the number of active features based on the complexity of each input. As a result, the representation better matches the structure of the data, and the explanation length reflects the amount of information in the input. Experimental results confirm that SoftSAE not only finds meaningful features, but also selects the right number of features for each concept. The source code is available at: https://github.com/St0pien/SoftSAE.
Authors:Hongcan Guo, Qinyu Zhao, Yian Zhao, Shen Nie, Rui Zhu, Qiushan Guo, Feng Wang, Tao Yang, Hengshuang Zhao, Guoqiang Wei, Yan Zeng
Abstract:
Large language models have achieved remarkable success under the autoregressive paradigm, yet high‑quality text generation need not be tied to a fixed left‑to‑right order. Existing alternatives still struggle to jointly achieve generation efficiency, scalable representation learning, and effective global semantic modeling. We propose Cola DLM, a hierarchical latent diffusion language model that frames text generation through hierarchical information decomposition. Cola DLM first learns a stable text‑to‑latent mapping with a Text VAE, then models a global semantic prior in continuous latent space with a block‑causal DiT, and finally generates text through conditional decoding. From a unified Markov‑path perspective, its diffusion process performs latent prior transport rather than token‑level observation recovery, thereby separating global semantic organization from local textual realization. This design yields a more flexible non‑autoregressive inductive bias, supports semantic compression and prior fitting in continuous space, and naturally extends to other continuous modalities. Through experiments spanning 4 research questions, 8 benchmarks, strictly matched ~2B‑parameter autoregressive and LLaDA baselines, and scaling curves up to about 2000 EFLOPs, we identify an effective overall configuration of Cola DLM and verify its strong scaling behavior for text generation. Taken together, the results establish hierarchical continuous latent prior modeling as a principled alternative to strictly token‑level language modeling, where generation quality and scaling behavior may better reflect model capability than likelihood, while also suggesting a concrete path toward unified modeling across discrete text and continuous modalities.
Authors:Ziyun Zeng, Yiqi Lin, Guoqiang Liang, Mike Zheng Shou
Abstract:
In recent years, open‑source efforts like Senorita‑2M have propelled video editing toward natural language instruction. However, current publicly available datasets predominantly focus on local editing or style transfer, which largely preserve the original scene structure and are easier to scale. In contrast, Background Replacement, a task central to creative applications such as film production and advertising, requires synthesizing entirely new, temporally consistent scenes while maintaining accurate foreground‑background interactions, making large‑scale data generation significantly more challenging. Consequently, this complex task remains largely underexplored due to a scarcity of high‑quality training data. This gap is evident in poorly performing state‑of‑the‑art models, e.g., Kiwi‑Edit, because the primary open‑source dataset that contains this task, i.e., OpenVE‑3M, frequently produces static, unnatural backgrounds. In this paper, we trace this quality degradation to a lack of precise background guidance during data synthesis. Accordingly, we design a scalable pipeline that generates foreground and background guidance in a decoupled manner with strict quality filtering. Building on this pipeline, we introduce Sparkle, a dataset of ~140K video pairs spanning five common background‑change themes, alongside Sparkle‑Bench, the largest evaluation benchmark tailored for background replacement to date. Experiments demonstrate that our dataset and the model trained on it achieve substantially better performance than all existing baselines on both OpenVE‑Bench and Sparkle‑Bench. Our proposed dataset, benchmark, and model are fully open‑sourced at https://showlab.github.io/Sparkle/.
Authors:Fangda Chen, Shanshan Zhao, Longrong Yang, Chuanfu Xu, Zhigang Luo, Long Lan
Abstract:
Video diffusion models perform well in short‑video synthesis, but their training‑free extension to long videos often suffers from content drift, temporal inconsistency, and over‑smoothed dynamics. Existing methods improve temporal consistency by combining a global branch with a local branch, but they often further decompose appearance consistency and temporal dynamics within each branch using predefined criteria. This assignment is unreliable when appearance and action progression are tightly coupled, such as in camera motion and sequential motion. We analyze the video temporal extension issue from a singular‑spectrum perspective and show that enlarged self‑attention windows induce spectral concentration: spectral energy becomes dominated by a few low‑rank singular directions, preserving coarse structure but suppressing high‑rank spatial details and motion‑rich temporal variations. To mitigate this problem, we propose FreeSpec, a training‑free spectral reconstruction framework for long‑video generation. FreeSpec decomposes global and local features with singular value decomposition, and uses the global branch as low‑rank spectral guidance and the local branch as a high‑rank reconstruction basis. This spectrum‑level fusion avoids the rigid feature partitioning of previous decomposition rules, preserving long‑range consistency while better retaining spatial details and temporal dynamics. Experiments on Wan2.1 and LTX‑Video demonstrate that FreeSpec improves long‑video generation, especially for temporal dynamics, while maintaining strong visual quality and temporal consistency. Project demo: https://fdchen24.github.io/FreeSpec‑Website/.
Authors:Canyu Zhao, Hao Chen, Yunze Tong, Yu Qiao, Jiacheng Li, Chunhua Shen
Abstract:
Reinforcement learning fine‑tuning has become the dominant approach for aligning diffusion models with human preferences. However, assessing images is intrinsically a multi‑dimensional task, and multiple evaluation criteria need to be optimized simultaneously. Existing practice deal with multiple rewards by training one specialist model per reward, optimizing a weighted‑sum reward R(x)=\sum_k w_k R_k(x), or sequentially fine‑tuning with a hand‑crafted stage schedule. These approaches either fail to produce a unified model that can be jointly trained on all rewards or necessitates heavy manually tuned sequential training. We find that the failure stems from using a naive weighted‑sum reward aggregation. This approach suffers from a sample‑level mismatch because most rollouts are specialist samples, highly informative for certain reward dimensions but irrelevant for others; consequently, weighted summation dilutes their supervision. To address this issue, we propose MARBLE (Multi‑Aspect Reward BaLancE), a gradient‑space optimization framework that maintains independent advantage estimators for each reward, computes per‑reward policy gradients, and harmonizes them into a single update direction without manually‑tuned reward weighting, by solving a Quadratic Programming problem. We further propose an amortized formulation that exploits the affine structure of the loss used in DiffusionNFT, to reduce the per‑step cost from K+1 backward passes to near single‑reward baseline cost, together with EMA smoothing on the balancing coefficients to stabilize updates against transient single‑batch fluctuations. On SD3.5 Medium with five rewards, MARBLE improves all five reward dimensions simultaneously, turns the worst‑aligned reward's gradient cosine from negative under weighted summation in 80% of mini‑batches to consistently positive, and runs at 0.97X the training speed of baseline training.
Authors:Pranav Mantini, Shishir K. Shah
Abstract:
We address the challenge of knowledge composition in Vision‑Language Models (VLMs), where accumulating expertise across multiple domains or tasks typically leads to catastrophic forgetting. We introduce GeoStack (Geometric Stacking), a modular framework that allows independently trained domain experts to be composed into a unified model. By imposing geometric and structural constraints on the adapter manifold, GeoStack ensures the foundational knowledge of the base model is preserved. Furthermore, we mathematically demonstrate a weight‑folding property that achieves constant‑time inference complexity (O(1)), regardless of the number of integrated experts. Experimental results across multi‑domain adaptation and class‑incremental learning show that GeoStack provides an efficient mechanism for long‑term knowledge composition while significantly mitigating catastrophic forgetting. Code is available at https://github.com/QuantitativeImagingLaboratory/GeoStack.
Authors:Tao Liu, Hao Yan, Mengting Chen, Taihang Hu, Zhengrong Yue, Zihao Pan, Jinsong Lan, Xiaoyong Zhu, Ming-Ming Cheng, Bo Zheng, Yaxing Wang
Abstract:
Step distillation has become a leading technique for accelerating diffusion models, among which Distribution Matching Distillation (DMD) and Consistency Distillation are two representative paradigms. While consistency methods enforce self‑consistency along the full PF‑ODE trajectory to steer it toward the clean data manifold, vanilla DMD relies on sparse supervision at a few predefined discrete timesteps. This restricted discrete‑time formulation and mode‑seeking nature of the reverse KL divergence tends to exhibit visual artifacts and over‑smoothed outputs, often necessitating complex auxiliary modules ‑‑ such as GANs or reward models ‑‑ to restore visual fidelity. In this work, we introduce Continuous‑Time Distribution Matching (CDM), migrating the DMD framework from discrete anchoring to continuous optimization for the first time. CDM achieves this through two continuous‑time designs. First, we replace the fixed discrete schedule with a dynamic continuous schedule of random length, so that distribution matching is enforced at arbitrary points along sampling trajectories rather than only at a few fixed anchors. Second, we propose a continuous‑time alignment objective that performs active off‑trajectory matching on latents extrapolated via the student's velocity field, improving generalization and preserving fine visual details. Extensive experiments on different architectures, including SD3‑Medium and Longcat‑Image, demonstrate that CDM provides highly competitive visual fidelity for few‑step image generation without relying on complex auxiliary objectives. Code is available at https://github.com/byliutao/cdm.
Authors:Shouvik Sardar, Sourish Das
Abstract:
Cocoa (Theobroma cacao) is a critical cash crop for millions of smallholder farmers in West Africa, where Cocoa Swollen Shoot Virus Disease (CSSVD) and anthracnose cause devastating yield losses. Automated disease detection from leaf images is essential for early intervention, yet deploying such systems in resource‑constrained settings demands models that are small, fast, and require no internet connectivity. Existing edge‑deployable plant disease systems rely on end‑to‑end deep learning without uncertainty quantification, while Bayesian methods for edge devices focus on hardware‑level inference architectures rather than agricultural applications. We bridge this gap with TinyBayes, the first framework to combine a closed‑form Bayesian classifier with a mobile‑grade computer vision pipeline for crop disease detection. Our pipeline uses YOLOv8‑Nano (5.9 MB) for lesion localisation, MobileNetV3‑Small (3.5 MB) for feature extraction, and the Jacobi prior; a Bayesian method that provides a closed form non‑iterative estimators via projection, for the classification. The Jacobi‑DMR (Distributed Multinomial Regression) classifier adds only 13.5 KB to the pipeline, bringing the total model size within 9.5 MB, while achieving 78.7% accuracy on the Amini Cocoa Contamination Challenge dataset and enabling end‑to‑end CPU inference under 150 ms per image. We benchmark against seven classifiers including Random Forest, SVM, Ridge, Lasso, Elastic Net, XGBoost, and Jacobi‑GP, and demonstrate that the Jacobi‑DMR offers the best trade‑off between accuracy, model size, and inference speed for edge deployment. We have proved the asymptotic equivalence and consistency, asymptotic normality and the bias correction of Jacobi‑DMR. All data and codes are available here: https://github.com/shouvik‑sardar/TinyBayes
Authors:Ke Zhang, Bomin Wang, Hangqi Zhou, Xiahai Zhuang
Abstract:
Curating fully annotated datasets for medical image segmentation is labour‑intensive and expertise‑demanding. To alleviate this problem, prior studies have explored scribble annotations for weakly supervised segmentation. Existing solutions mainly compute losses on annotated areas and generate pseudo labels by propagating annotations to adjacent regions. However, these methods often suffer from inaccurate and unrealistic segmentations due to insufficient supervision and incomplete shape information. In contrast, we first investigate the principle of good scribble annotations, which leads to efficient scribble forms via supervision maximization and randomness simulation. We further introduce regularization terms to encode the spatial relationship and the shape constraints, where the EM algorithm is utilized to estimate the mixture ratios of label classes. These ratios are critical in identifying the unlabeled pixels for each class and correcting erroneous predictions, thus the accurate estimation lays the foundation for the incorporation of spatial prior. Finally, we integrate the efficient scribble supervision with the prior into a framework, referred to as ZScribbleSeg, and apply it to multiple scenarios. Leveraging only scribble annotations, ZScribbleSeg achieves competitive performance on six segmentation tasks including ACDC, MSCMRseg, BTCV, MyoPS, Decathlon‑BrainTumor and Decathlon‑Prostate. Our code will be released via https://github.com/DLwbm123/ZScribbleSeg.
Authors:Shiao Wang, Xiao Wang, Duoqing Yang, Wenhao Zhang, Bo Jiang, Lin Zhu, Yonghong Tian, Bin Luo
Abstract:
Despite significant progress, RGB‑based trackers remain vulnerable to challenging imaging conditions, such as low illumination and fast motion. Event cameras offer a promising alternative by asynchronously capturing pixel‑wise brightness changes, providing high dynamic range and high temporal resolution. However, existing event‑based trackers often neglect the intrinsic spatial sparsity and temporal density of event data, while relying on a single fixed temporal‑window sampling strategy that is suboptimal under varying motion dynamics. In this paper, we propose an event sparsity‑aware tracking framework that explicitly models event‑density variations across multiple temporal scales. Specifically, the proposed framework progressively injects sparse, medium‑density, and dense event search regions into a three‑stage Vision Transformer backbone, enabling hierarchical multi‑density feature learning. Furthermore, we introduce a sparsity‑aware Mixture‑of‑Experts module to encourage expert specialization under different sparsity patterns, and design a dynamic pondering strategy to adaptively adjust the inference depth according to tracking difficulty. Extensive experiments on FE240hz, COESOT, and EventVOT demonstrate that the proposed approach achieves a favorable trade‑off between tracking accuracy and computational efficiency. The source code will be released on https://github.com/Event‑AHU/OpenEvTracking.
Authors:Xiaochen Huang, Honggang Chen, Weicheng Zhang, Xiaobo Dai, Yongyi Li, Linbo Qing, Xiaohai He
Abstract:
In multimedia application scenarios, images captured under low‑illumination conditions often lead to lower accuracy in visual perception tasks compared to those taken in well‑lit environments.
To tackle this challenge, we propose AMIEOD, an image enhancement‑enabled object detection framework for low‑illumination scenes, where the two tasks are jointly optimized in a detection performance‑oriented manner.
Specifically, to fully exploit the information in poorly lit images, a Multi‑Experts Image Enhancement Module (MEIEM) is proposed, which leverages diverse enhancement strategies.
On this basis, aiming to better align the MEIEM with the detection task, we propose a Detection‑Guided Regression Loss (DGRL) that utilizes the detection result to decide the regression target.
Moreover, to dynamically select the most suitable enhancement strategy from MEIEM during inference, we construct an Expert Selection Module (ESM) guided by the proposed Detection‑Guided Cross‑Entropy (DGCE) loss, which formulates the optimization of ESM as a classification task.
The improved method is well‑matched with current detection algorithms to improve their performance in dim scenes.
Extensive experiments on multiple datasets demonstrate that the proposed method significantly improves object detection accuracy in low‑illumination conditions. Our code has been released at https://github.com/scujayfantasy/AMIEOD
Authors:Jun Li, Peifeng Lai, Xuhang Lou, Jinpeng Wang, Yuting Wang, Ke Chen, Yaowei Wang, Shu-Tao Xia
Abstract:
Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries and rich video content inevitably introduces uncertainty into the retrieval process. In this setting, vague queries often induce semantic ambiguity across videos, a challenge that is further exacerbated by the sparse temporal supervision within videos, which fails to provide sufficient matching evidence. To address this, we propose Holmes, a hierarchical evidential learning framework that aggregates multi‑granular cross‑modal evidence to quantify and model uncertainty explicitly. At the inter‑video level, similarity scores are interpreted as evidential support and modeled via a Dirichlet distribution. Based on the proposed three‑fold principle, we perform fine‑grained query identification, which then guides query‑adaptive calibrated learning. At the intra‑video level, to accumulate denser evidence, we formulate a soft query‑clip alignment via flexible optimal transport with an adaptive dustbin, which alleviates sparse temporal supervision while suppressing spurious local responses. Extensive experiments demonstrate that Holmes outperforms state‑of‑the‑art methods. Code is released at https://github.com/lijun2005/ICML26‑Holmes.
Authors:Shichao Kan, Xuyang Zhang, Haojie Zhang, Zhe Zhu, Yigang Cen, Yixiong Liang, Lianlei Shan, Linna Zhang, Zhe Qu, Jiazhi Xia
Abstract:
Evaluating image captions without references remains challenging because global embedding similarity often misses fine‑grained mismatches such as hallucinated objects, missing attributes, or incorrect relations. We propose MSD‑Score, a reference‑free metric that models image patch and text token embeddings as von Mises‑Fisher mixtures on the unit hypersphere. Instead of treating each modality as a single point, MSD‑Score formulates image‑text matching as a multi‑scale distributional scoring problem. Semantic discrepancies are quantified via a weighted bi‑directional KL divergence and combined with global similarity in a multi‑scale framework for both single‑ and multi‑candidate evaluations. Extensive experiments show that MSD‑Score achieves state‑of‑the‑art correlation with human judgments among reference‑free metrics. Beyond accuracy, its probabilistic formulation yields transparent and decomposable diagnostics of local grounding errors, providing a deterministic complementary signal to holistic similarity metrics and judge‑based evaluators.
Authors:Xiangyue Zhang, Yiyi Cai, Kunhang Li, Kaixing Yang, You Zhou, Zhengqing Li, Xuangeng Chu, Jiaxu Zhang, Haiyang Liu
Abstract:
We propose PersonaGesture, a diffusion‑based pipeline for single‑reference co‑speech gesture personalization of unseen speakers. Given target speech and one motion clip from a new speaker, the model must synthesize gestures that follow the new utterance while retaining speaker‑specific pose choices, without per‑speaker optimization. This setting is useful for avatars and virtual agents, but it is hard because the reference mixes stable speaker habits with utterance‑specific trajectories. PersonaGesture consists of two key components, Adaptive Style Infusion (ASI) and Implicit Distribution Rectification (IDR), to separate temporal identity evidence from residual statistic correction. A Style Perceiver first encodes the variable‑length reference into compact speaker‑memory tokens. ASI injects these tokens into denoising through zero‑initialized residual cross‑attention, enabling style evidence to affect motion formation without replacing the pretrained speech‑to‑motion prior. Building on this, IDR applies a length‑aware diagonal affine map in latent space to correct residual channel‑wise moments estimated from the same reference. Across BEAT2 and ZeroEGGS, we evaluate quantitative metrics, reference‑identity controls, same‑audio diagnostics, qualitative comparisons, and human preference. Experiments show that separating denoising‑time speaker memory from conservative post‑generation moment correction improves unseen‑speaker personalization over collapsed style codes, full‑reference attention, and one‑clip finetuning. Project: https://xiangyue‑zhang.github.io/PersonaGesture.
Authors:Youcan Xu, Jiaxin Shi, Zhen Wang, Wensong Song, Feifei Shao, Chen Liang, Jun Xiao, Long Chen
Abstract:
Camera‑controlled video‑to‑video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcasting. However, existing implicit synthesis methods fundamentally rely on non‑causal, full‑sequence processing and rigid prefix‑style temporal concatenation. This architectural paradigm mandates bidirectional attention, resulting in prohibitive computational latency, quadratic complexity scaling, and inherent incompatibility with real‑time streaming or variable‑length inputs. To overcome these limitations, we introduce \textttRealCam, a novel autoregressive framework for interactive, real‑time camera‑controlled V2V generation. We first design a high‑fidelity teacher model grounded in a Cross‑frame In‑context Learning paradigm. By interleaving source and target frames into synchronized contextual pairs, our design inherently enables length‑agnostic generalization and naturally facilitates causal adaptation, breaking the rigid prefix bottleneck. We then distill this teacher into a few‑step causal student via Self‑Forcing with Distribution Matching Distillation, enabling efficient, on‑the‑fly streaming synthesis. Furthermore, to mitigate severe loop inconsistency in closed‑loop trajectories, we propose Loop‑Closed Data Augmentation (LoopAug), a novel paradigm that synthesizes globally consistent loop sequences from existing multiview datasets. Extensive experiments demonstrate that \textttRealCam achieves state‑of‑the‑art visual fidelity and temporal consistency while enabling truly interactive camera control with orders‑of‑magnitude faster inference than existing paradigms. Our project page is at https://xyc‑fly.github.io/RealCam/.
Authors:Tommy Carstensen
Abstract:
Systematic reviews and meta‑analyses frequently require numerical data that authors report only as figures, yet manual digitisation is slow and does not scale. We present PlotPick, an open‑source tool that uses vision‑language models (VLMs) to batch‑extract structured tabular data from scientific figures. We evaluate six VLMs from three providers on two established chart‑to‑table benchmarks (ChartX and PlotQA) and compare against the dedicated chart‑to‑table model DePlot. All six VLMs outperform DePlot on both benchmarks. On ChartX (restricted to bar charts, line charts, box plots, and histograms; n=300), VLMs achieve 88‑96% recall versus 71% for DePlot. On PlotQA (n=529), VLMs achieve 86‑99% RMSF1 versus 94% for DePlot. The gap is largest on chart types absent from the dedicated models' training data: on box plots, DePlot achieves 24% RMSF1 while VLMs achieve 83‑97%. PlotPick is available at https://plotpick.streamlit.app.
Authors:Xiao Wang, Ziwen Wang, Weizhe Kong, Wentao Wu, Yuehang Li, Aihua Zheng, Chenglong Li, Jin Tang
Abstract:
Vehicle Re‑identification (Re‑ID) aims to retrieve the most similar image to a given query from images captured by non‑overlapping cameras. Extending vehicle Re‑ID from image‑only queries to text‑based queries enables retrieval in real‑world scenarios where only a witness description of the target vehicle is available. In this paper, we propose PFCVR, a Part‑level Fine‑grained Cross‑modal Vehicle Retrieval model for text‑to‑image vehicle re‑identification. PFCVR constructs locally paired images and texts at the part level and introduces learnable part‑query tokens that aggregate both part‑specific and full‑sentence context before aligning with visual part features. On top of this explicit local alignment, a bi‑directional mask recovery module lets each modality reconstruct its masked content under the guidance of the other, implicitly bridging local correspondences into global feature alignment. Furthermore, we construct a new large‑scale dataset called T2I‑VeRW, which contains 14,668 images covering 1,796 vehicle identities with fine‑grained part‑level annotations. Experimental results on the T2I‑VeRI dataset show that PFCVR achieves 29.2% Rank‑1 accuracy, improving over the best competing method by +3.7% percentage points. On the newly proposed T2I‑VeRW benchmark, PFCVR achieves 55.2% Rank‑1 accuracy, outperforming a comprehensive set of recent state‑of‑the‑art methods. Source code will be released on https://github.com/Event‑AHU/Neuromorphic_ReID
Authors:Zhangquan Chen, Manyuan Zhang, Xinlei Yu, Xiang An, Bo Li, Xin Xie, ZiDong Wang, Mingze Sun, Shuang Chen, Hongyu Li, Xiaobin Hu, Ruqi Huang
Abstract:
Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vision‑language models (VLMs). Prior approaches either verbalize spatial‑temporal reasoning entirely as text, which is inherently verbose and imprecise for complex dynamics, or rely on external geometric modules that increase inference complexity without fostering intrinsic model capability. In this paper, we present 4DThinker, the first framework that enables VLMs to "think with 4D" through dynamic latent mental imagery, i.e., internally simulating how scenes evolve within the continuous hidden space. Specifically, we first introduce a scalable, annotation‑free data generation pipeline that synthesizes 4D reasoning data from raw videos. We then propose Dynamic‑Imagery Fine‑Tuning (DIFT), which jointly supervises textual tokens and 4D latents to ground the model in dynamic visual semantics. Building on this, 4D Reinforcement Learning (4DRL) further tackles complex reasoning tasks via outcome‑based rewards, restricting policy gradients to text tokens to ensure stable optimization. Extensive experiments across multiple dynamic spatial reasoning benchmarks demonstrate that 4DThinker consistently outperforms strong baselines and offers a new perspective toward 4D reasoning in VLMs. Our code is available at https://github.com/zhangquanchen/4DThinker.
Authors:Christian Wachinger, Bernhard Renger, Christopher Späth, Jan Kirschke, Marcus Makowski
Abstract:
Interpreting quantitative CT biomarkers, such as organ volume and tissue attenuation, requires large‑scale healthy reference distributions. However, creating these is challenging because clinical datasets are often heavily enriched with pathology. Here, we develop an evidence‑grounded, cross‑verified large language model (LLM) ensemble to filter pathological findings from radiology reports, enabling the construction of pathology‑reduced cohorts from over 350,000 CT examinations. Five LLMs, first, flag structure‑level abnormality candidates grounded in verbatim report evidence and, second, resolve disagreements via cross‑verification. Using distribution‑aware generalized additive models for location, scale, and shape, we establish comprehensive whole‑body reference charts for 106 anatomical structures (volumes and attenuation) across adulthood, accounting for age, sex, contrast enhancement, and acquisition parameters. Longitudinal analyses reveal structure‑ and contrast‑dependent changes distinct from cross‑sectional trends. These resources facilitate covariate‑adjusted centile scoring from routine CT, supporting standardized quantitative phenotyping, multi‑site imaging studies, and scalable opportunistic screening research.
Authors:Junhui Yin, Nan Pu, Xinyu Zhang, Lingfeng Yang, Lin Wu, Xiaojie Wang, Zhun Zhong
Abstract:
Prompt learning has become an effective and widely used technique in enhancing vision‑language models (VLMs) such as CLIP for various downstream tasks, particularly in zero‑shot classification within specific domains. Existing methods typically focus on either learning class‑shared prompts for a given domain or generating instance‑specific prompts through conditional prompt learning. While these methods have achieved promising performance, they often overlook class‑specific knowledge in prompt design, leading to suboptimal outcomes. The underlying reasons are: 1) class‑specific prompts offer more fine‑grained supervision compared to coarse class‑shared prompts, which helps prevent misclassification of data from different classes into a single class; 2) compared to class‑specific prompts, instance‑specific prompts neglect the richer class‑level information across multiple instances, potentially causing data from the same class to be divided into multiple classes. To effectively supplement the class‑specific knowledge into existing methods, we propose a plug‑and‑play Class‑Aware Knowledge Injection (CAKI) framework. CAKI comprises two key components, i.e., class‑specific prompt generation and query‑key prompt matching. The former encodes class‑specific knowledge into prompts from few‑shot samples that belong to the same class and stores the learned prompts in a class‑level knowledge bank. The latter provides a plug‑and‑play mechanism for each test instance to retrieve relevant class‑level knowledge from the knowledge bank and inject such knowledge to refine model predictions. Extensive experiments demonstrate that our CAKI effectively improves the performance of existing methods on base and novel classes. Code is publicly available at \hrefhttps://github.com/yjh576/CAKIthis https URL.
Authors:Sankarshana Venugopal, Mohammad Mostafavi, Jonghyun Choi
Abstract:
Diffusion‑based image‑to‑image (I2I) translation excels in high‑fidelity generation but suffers from slow sampling in state‑of‑the‑art Diffusion Bridge Models (DBMs), often requiring dozens of function evaluations (NFEs). We introduce DBMSolver, a training‑free sampler that exploits the semi‑linear structure of DBM's underlying SDE and ODE via exponential integrators, yielding highly‑efficient 1st‑ and 2nd‑order solutions. This reduces NFEs by up to 5x while boosting quality (e.g., FID drops 53% on DIODE at 20 NFEs vs. 2nd‑order baseline). Experiments on inpainting, stylization, and semantics‑to‑image tasks across resolutions up to 256x256 show DBMSolver sets new SOTA efficiency‑quality tradeoffs, enabling real‑world applicability. Our code is publicly available at https://github.com/snumprlab/dbmsolver.
Authors:Kunchong Shi, Jing Zhang
Abstract:
Current Chinese calligraphy generation methods suffer from poor stroke rendering and unrealistic ink morphology, resulting in outputs with limited visual fidelity and artistic fluidity. To address this problem, we propose InkDiffuser, a diffusion‑based generative framework for one‑shot Chinese calligraphy synthesis. To guarantee high‑fidelity rendering, we introduce two core contributions: a high‑frequency enhancement mechanism and a Differentiable Ink Structure (DIS) loss that explicitly regularizes ink morphology. Inspired by the observation that high‑frequency information in individual samples typically carries contour details, we enhance content extraction by explicitly fusing high‑frequency representations for more accurate font structure. Furthermore, we propose a differentiable ink structure loss that integrates differentiable morphological operations into the diffusion process. By allowing the model to learn an explicit decomposition of ink‑trace structures, DIS facilitates fine‑grained refinement of stroke contours and delivers significantly improved visual realism in the generated calligraphy. Extensive experiments on various calligraphic styles and complex characters demonstrate that InkDiffuser can generate superior calligraphy fonts with realistic ink rendering effects from only a single reference glyph and outperform existing few‑shot font generation approaches in structural consistency, detail fidelity, and visual authenticity. The code is available at the following address: https://github.com/JingVIPLab/InkDiffuser.
Authors:Megha Mariam K. M, Vineeth N. Balasubramanian, C. V. Jawahar
Abstract:
The communication of scientific knowledge has become increasingly multimodal, spanning text, visuals, and speech through materials such as research papers, slides, and recorded presentations. These different representations collectively convey a study's reasoning, results, and insights, offering complementary perspectives that enrich understanding. However, despite their shared purpose, such materials are rarely connected in a structured way. The absence of explicit links across formats makes it difficult to trace how concepts, visuals, and explanations correspond, limiting unified exploration and analysis of research content. To address this gap, we introduce the Multimodal Conference Dataset (MCD), the first benchmark that integrates research papers, presentation videos, explanatory videos, and slides from the same works. We evaluate a range of embedding‑based and vision‑language models to assess their ability to discover fine‑grained cross‑format correspondences, establishing the first systematic benchmark for this task. Our results show that vision‑language models are robust but struggle with fine‑grained alignment, while embedding‑based models capture text‑visual correspondences well but equations and symbolic content form distinct clusters in the embedding space. These findings highlight both the strengths and limitations of current approaches and point to key directions for future research in multimodal scientific understanding. To ensure reproducibility, we release the resources for MCD at https://github.com/meghamariamkm2002/MCD
Authors:Zhengru Fang, Yanan Ma, Yu Guo, Senkang Hu, Yixian Zhang, Hangcheng Cao, Wenbo Ding, Yuguang Fang
Abstract:
When a chest X‑ray shows consolidation but the question asks which finding is present, a medical vision‑language model may answer "No consolidation." This is more than an incorrect choice: it is a polarity reversal that emits a clinical statement contradicting the image. We study this failure as negated‑option attraction, where a model is drawn to a negated answer option even when it conflicts with both the visual evidence and the question. We introduce CXR‑ContraBench (Chest X‑Ray Contradiction Benchmark), a diagnostic benchmark spanning internal ReXVQA slices and external OpenI and CheXpert protocols. The benchmark centers on present‑finding questions, where selecting "No X" despite visible X creates the main clinical risk, and uses absent‑finding questions as secondary tests of whether models copy negated wording. Across CheXpert protocols, the failure is substantial and persistent. On a strict direct presence probe, MedGemma and Qwen2.5‑VL reach only 31.49% and 30.21% accuracy, respectively; on a matched 135,754‑record CheXpert training‑split protocol, both models select negated options on over 62% of presence questions. Chain‑of‑thought prompting reduces some presence‑side reversals but does not eliminate them and can amplify absence‑side contradictions. Finally, QCCV‑Neg (Question‑Conditioned Consistency Verifier for Negation) deterministically repairs the measured polarity‑confused subset without retraining, raising MedGemma and Qwen2.5‑VL to 96.60% and 95.32% accuracy on the direct presence probe. These results show that standard accuracy can hide a clinically meaningful inference‑time polarity failure. Source code and benchmark construction scripts are available at https://github.com/fangzr/cxr‑contrabench‑code.
Authors:Hao Wang, Shiqi Wang, Qi Liu
Abstract:
Generating realistic 3D Human‑Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high‑level semantic intent with strict low‑level physical constraints. Existing methods excel at semantic alignment, however, they struggle to maintain precise object contact. We reveal a key finding termed Geometric Forgetting: as diffusion model depth increases, semantic feature tend to overshadow object geometry feature, causing the model to lose its perception to object geometry. To address this, we propose MaMi‑HOI, a hierarchical framework reconciling Macro‑level kinematic fluidity with Micro‑level spatial precision. First, to counteract geometric forgetting, we introduce the Geometry‑Aware Proximity Adapter (GAPA), which explicitly re‑injects dense object details to perform residual snapping corrections for precise contact. Nevertheless, such aggressive local enforcement can disrupt global dynamics, leading to robotic stiffness. In response, we introduce the Kinematic Harmony Adapter (KHA), which proactively aligns whole‑body posture with spatial objectives, ensuring the skeleton actively accommodates constraints without compromising naturalness. Extensive experiments validate that MaMi‑HOI simultaneously achieves natural motion and precise contact. Crucially, it extends generation capabilities to long‑term tasks with complex trajectories, effectively bridging the gap between global navigation and high‑fidelity manipulation in 3D scenes. Code is available at https://github.com/DON738110198/MaMi‑HOI.git
Authors:Anh H. Vo, Sungyo Lee, Phil-Joong Kim, Soo-Mi Choi, Yong-Guk Kim
Abstract:
Recent advances in large language models (LLMs) have significantly improved language‑driven 3D content generation, but most existing approaches still treat scene generation and user interaction as separate processes, limiting the adaptability and immersive potential of interactive multimedia systems. This paper presents a unified framework that closes the loop between language‑driven 3D scene generation and immersive user interaction. Given natural language instructions, the system first constructs structured scene representations using LLMs, and then optimizes spatial layouts via reinforcement learning under geometric and semantic constraints. The generated environments are deployed in a virtual reality setting to facilitate HRI‑in‑the‑loop, where user interactions provide continuous feedback to align generated content with human perception and usability. By tightly coupling generation and interaction, the proposed framework enables more responsive, adaptive, and realistic multimedia experiences. Experiments on the ALFRED benchmark demonstrate state‑of‑the‑art performance in task‑based scene generation. Furthermore, qualitative results and user studies show consistent improvements in immersion, interaction quality, and task efficiency, highlighting the importance of closed‑loop integration of generation and interaction for next‑generation multimedia systems. Our project page can be found at https://proj‑showcase.github.io/h3ds/.
Authors:Xiwen Luo, Jia Li, Rencheng Song, Yu Liu, Juan Cheng
Abstract:
Emotion recognition from facial videos enables non‑contact inference of human emotional states. Although facial expressions are widely used cues, they cannot fully reflect intrinsic affective states. Remote photoplethysmography (rPPG) provides complementary physiological information, but it is highly susceptible to noise and inter‑subject variability, limiting generalization to unseen individuals. Existing multimodal methods combine facial and rPPG features, yet their fusion strategies often disrupt pretrained facial representations and lack explicit mechanisms to suppress subject‑specific variations. To address these issues, we propose a subject‑invariant cross‑modal prompt‑tuning framework for video‑based emotion recognition. Specifically, rPPG waveforms are transformed into noise‑robust time‑frequency representations (TFRs), from which modality‑complementary prompts are generated to modulate facial tokens within a frozen Vision Transformer (ViT). This design enables effective cross‑modal interaction while preserving the generalizable facial representations learned by the pretrained backbone. In addition, we introduce a decoupled shared‑specific adapter (DSSA) into each ViT layer to explicitly separate subject‑shared and subject‑specific components, thereby improving cross‑subject generalization. Experiments on the MAHNOB‑HCI and DEAP benchmarks demonstrate that the proposed method consistently outperforms strong baselines in both recognition accuracy and generalization ability, highlighting its effectiveness for video‑based emotion recognition.
Authors:Yiyang Shen, Yin Yang, Kun Zhou, Tianjia Shao
Abstract:
We introduce S2C‑3D, a novel sparse‑view 3D reconstruction framework for high‑fidelity and complete scene reconstruction from as few as six to eight images. Our framework features three components: a specialized diffusion model for scene‑specific image restoration, a training‑free view‑consistency conditioned sampling process in the diffusion model for refined Gaussian optimization, and a camera trajectory planning scheme to ensure comprehensive scene coverage. The specialized diffusion model is developed by finetuning a pretrained architecture on the input views and their corresponding degraded counterparts. The adaptation to the scene distribution allows the model to repair Gaussian renderings while effectively eliminating domain gaps. Meanwhile, the trajectory planning scheme optimizes scene coverage by connecting each newly sampled camera to its two nearest neighbors. By iteratively constructing paths and retaining only those that significantly enhance visibility, the scheme establishes a trajectory that covers the entire scene. To address multi‑view conflicts, the view‑consistency conditioned sampling process quantifies the consistency between neighboring repaired images. This information is injected as a condition into the sampling process of the frozen diffusion model, facilitating the generation of view‑consistent images without additional training. Consequently, our approach produces high‑fidelity 3D Gaussians that are robust to artifacts. Experimental results demonstrate that S2C‑3D outperforms state‑of‑the‑art methods, constructing high‑quality scenes that are free from missing regions, blurring, or other artifacts with very sparse inputs. The source code and data are available at https://gapszju.github.io/S2C‑3D.
Authors:Panqi Yang, Haodong Jing, Jiahao Chao, Tingyan Xiang, Li Lin, Yao Hu, Yang Luo, Yongqiang Ma
Abstract:
Unified visual tokenization faces a fundamental trade‑off between high‑fidelity pixel reconstruction (spatial equivariance) and semantic abstraction (conceptual invariance). We attribute this conflict to Manifold Misalignment: naive joint optimization induces opposing gradients, creating a zero‑sum game between reconstruction and perception. To address this, we propose MUSE, a framework based on Topological Orthogonality. By treating Structure as an orthogonal bridge, MUSE decouples optimization within Transformers: structural gradients refine attention topology, while semantic gradients update feature values. This turns destructive interference into Mutual Reinforcement. Experiments show that MUSE breaks the trade‑off, achieving state‑of‑the‑art generation quality (gFID 3.08) and surpassing its teacher InternViT‑300M in linear probing (85.2% vs. 82.5%), demonstrating that structurally aligned reconstruction can enhance semantic perception. Code is available at https://github.com/PanqiYang1/MUSE.
Authors:Yuxuan Han, Xin Ming, Tianxiao Li, Zhuofan Shen, Qixuan Zhang, Lan Xu, Feng Xu
Abstract:
High‑quality facial appearance capture has traditionally required costly studio recording. Recent works consider an in‑the‑wild smartphone‑based setup; however, their model‑based inverse rendering paradigm struggles with the complex disentanglement of reflectance from unknown illumination. To bridge this gap, we propose to shift the paradigm into training a powerful delighting network as a prior to constrain the optimization. We leverage the OLAT dataset and the rendered Light Stage scans for training, and propose Dataset Latent Modulation (DLM) to seamlessly integrate these heterogeneous data sources. Specifically, by conditioning the core network on learnable source‑aware tokens, we decouple dataset‑specific styles from physical delighting principles, enabling the emergence of a delighting prior that outperforms existing proprietary models. This powerful delighting prior enables a simple and automatic appearance capture pipeline that achieves high‑quality reflectance estimation from casual video inputs, outperforming prior arts by a large margin. Furthermore, we leverage our appearance capture method to transform the multi‑view NeRSemble dataset into NeRSemble‑Scan, a large‑scale collection of 4K‑resolution relightable scans. By open‑sourcing our model and the NeRSemble‑Scan dataset, we democratize high‑end facial capture and provide a new foundation for the research community to build photorealistic digital humans.
Authors:Gabriel Jeanson, David-Alexandre Duclos, William Larrivée-Hardy, Noé Cochet, Matěj Boxan, Anthony Deschênes, François Pomerleau, Philippe Giguère
Abstract:
Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour‑intensive and geographically constrained. While Uncrewed Aerial Vehicles (UAVs) offer scalable data collection, the transition to deep learning‑based interpretation is bottlenecked by the severe scarcity of expert‑annotated imagery, particularly in complex, visually heterogeneous regeneration zones. This paper addresses the dual challenges of data scarcity and extreme class imbalance in the semantic segmentation of fine‑grained forest regeneration species by providing a scalable framework that reduces reliance on manual photo‑interpretation for high‑resolution, millimetre‑level aerial imagery. Importantly, we leverage the large‑scale vision‑language Nano Banana Pro model to simultaneously generate high‑fidelity images and their corresponding pixel‑aligned semantic masks from prompts. We introduce WilDReF‑Q‑V2, an expansion of a natural forest dataset with 13 977 new unlabelled and 50 labelled real images, as well as the Gen4Regen dataset, featuring 2101 pairs of synthetic images and semantic masks. Our methodology integrates real‑world data with AI‑generated images, highlighting that AI‑generated data is highly complementary to real‑world data, with unified training yielding an F1 score improvement of over 15 %pt compared to purely supervised baselines. Furthermore, we demonstrate that even small quantities of prompt‑generated data significantly improve performance for underrepresented species, some of which saw per‑species F1 score gains of up to 30 %pt. We conclude that vision‑language models can serve as agile data generators, effectively bootstrapping perception tasks for niche AI domains where expert labels are scarce or unavailable. Our datasets, source code, and models will be available at https://norlab‑ulaval.github.io/gen4regen.
Authors:Anh Vu Nguyen, Dino Sejdinovic, Tat-Jun Chin
Abstract:
Edge learning refers to training machine learning models deployed on edge platforms, typically using new data accumulated onboard. The computational limitations on edge devices affect not only model optimisation, but also calculation of the predictive uncertainty of the current model on the unlabelled data, which is vital for informing model updating. In this paper, we investigate edge learning in the context of performing deep image regression on a remote sensing satellite, where a deep network is executed by an onboard computer to regress a scalar y from an input image, e.g., y is the percentage of pixels indicating cloud coverage or land use. We propose an uncertainty‑guided edge learning (UGEL) algorithm that can accurately prioritise the data to speed up training convergence of the on‑board regression model. Underpinning UGEL is the calculation of predictive uncertainty based on deep beta regression, where a deep network is used to estimate the parameters of a beta distribution for which the target y for an input image has a high likelihood. Compared to established methods for uncertainty estimation that are either too costly on edge devices (e.g., require many forward passes per sample) or make strict assumptions on the predictive distribution (e.g., Gaussian), deep beta regression is computable in a single forward pass and allows more general predictive distributions. Results show that UGEL delivers faster‑converging edge learning than active or semi‑supervised learning. Code and models are publicly available at https://github.com/anh‑vunguyen/UGEL.
Authors:Nan Yang, Julian Straub, Fan Zhang, Richard Newcombe, Jakob Engel, Lingni Ma
Abstract:
Tracking 3D human motion from egocentric multi‑camera headset is challenged by severe egomotion, partial visibility or occlusions and lack of training data. Existing methods designed for monocular video often require static or slowly‑moving cameras and cannot efficiently leverage multi‑view, calibrated and localized input. This makes them brittle and prone to fail on dynamic egocentric captures. We propose LAMP (Localization Aware Multi‑camera People Tracking): a novel, simple framework to solve this via early disentanglement of observer and target motion. LAMP introduces a two‑step process. First, we leverage the known device 6 DoF motion and calibration to convert detected 2D body keypoints from all cameras over a temporal window into a unified 3D world reference frame. Second, an end‑to‑end‑trained spatio‑temporal transformer fits 3D human motion directly to this 3D ray cloud. This "lift‑then‑fit" approach allows LAMP to learn and leverage a natural human motion prior in the world‑space, as well as providing an elegant framework to flexibly incorporate information from multiple temporally asynchronous, partially observing and moving cameras. LAMP achieves state‑of‑the‑art results on monocular benchmarks, while significantly outperforming baselines for our targeted egocentric setting.
Authors:Phillipp Fanta-Jende, Francesco Vultaggio, Alexander Kern, Yasmin Loeper, Markus Gerke
Abstract:
We present egenioussBench, a visual localisation benchmark built on geospatial reference data: a city‑scale airborne 3D mesh and a CityGML LoD2 model. This pairing reflects deployable mapping assets and supports true scalability beyond traditional SfM‑based approaches. The query data comprise smartphone images with centimetre‑accurate, map‑independent ground truth obtained via PPK and GCP/CP‑aided adjustment. From 2,709 images, we derive a non‑co‑visible subset by estimating the full co‑visibility matrix from rendered depth and selecting a maximum independent set; the released data include a test split of 42 non‑co‑visible images with withheld ground truth and a validation split of 412 sequential images with poses, e.g. for training of pose regressors and self‑validation. The benchmark features a public leaderboard evaluated with binning metrics at multiple pose‑error thresholds alongside global statistics (median, RMSE, outlier ratio), ensuring fair, like‑for‑like comparison across mesh‑ and LoD2‑based methods. Together, these design choices expose realistic cross‑view and cross‑domain challenges while providing a rigorous, scalable path for advancing large‑scale visual localisation. We make the evaluation code and data availeable at https://github.com/fratopa/egenioussBench and https://www.egeniouss.eu/
Authors:Till Beemelmanns, Alexey Nekrasov, Stefan Vilceanu, Jonas Steinhaus, Timo Woopen, Bastian Leibe, Lutz Eckstein
Abstract:
Reliable uncertainty estimation for 3D object detection is critical for deploying safe autonomous systems, yet modern detectors remain poorly calibrated, especially under distribution shifts. Although post‑hoc calibration methods address this issue and provide improved calibration for in‑distribution tests, they fail to adapt in distribution‑shifted scenarios. In this work, we address this issue and introduce a density‑aware calibration method that couples post‑hoc calibrators with the feature density of latent object queries from DETR‑style 3D object detectors. These queries form a compact, location and class‑aware feature, ideal for density estimation, allowing our approach to adjust model confidences in distribution‑shift scenarios. By fitting a density estimator on these query features, our approach jointly recalibrates both classification and bounding box regression uncertainties. On both a multi‑view camera and LiDAR‑based detector, our approach consistently outperforms standard post‑hoc methods in both in‑distribution and distribution‑shifted scenarios. Code available https://tillbeemelmanns.github.io/query2uncertainty/ .
Authors:Zeren Jiang, Yushi Lan, Yihang Luo, Yufan Deng, Zihang Lai, Edgar Sucar, Christian Rupprecht, Iro Laina, Diane Larlus, Chuanxia Zheng, Andrea Vedaldi
Abstract:
Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high‑quality datasets with dense, complete, and accurate geometric annotations. To address this limitation, we introduce Syn4D, a multiview synthetic dataset of dynamic scenes that includes ground‑truth camera motion, depth maps, dense tracking, and parametric human pose annotations. A key feature of Syn4D is the ability to unproject any pixel into 3D to any time and to any camera. We conduct extensive evaluations across multiple downstream tasks to demonstrate the utility and effectiveness of the proposed dataset, including 4D scene reconstruction, 3D point tracking, geometry‑aware camera retargeting, and human pose estimation. The experimental results highlight Syn4D's potential to facilitate research in dynamic scene understanding and spatiotemporal modeling.
Authors:Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du, Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang, Steven Hoi
Abstract:
The landscape of high‑performance image generation models is currently shifting from the inefficient multi‑step ones to the efficient few‑step counterparts (e.g, Z‑Image‑Turbo and FLUX.2‑klein). However, these models present significant challenges for direct continuous supervised fine‑tuning. For example, applying the commonly used fine‑tuning technique would compromise their inherent few‑step inference capability. To address this, we propose D‑OPSD, a novel training paradigm for step‑distilled diffusion models that enables on‑policy learning during supervised fine‑tuning. We first find that the modern diffusion models, where the LLM/VLM serves as the encoder, can inherit its encoder's in‑context capabilities. This enables us to formulate the training as an on‑policy self‑distillation process. Specifically, during training, we make the model act as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll‑outs. By optimizing on the model's own trajectory and under its own supervision, D‑OPSD enables the model to learn new concepts, styles, etc., without sacrificing the original few‑step capacity.
Authors:Shuang Chen, Kaituo Feng, Hangting Chen, Wenxuan Huang, Dasen Dai, Quanxin Shou, Yunlong Lin, Xiangyu Yue, Shenghua Gao, Tianyu Pang
Abstract:
Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence verification, and multi‑step reasoning. Despite rapid progress, top‑tier multimodal search agents remain difficult to reproduce, largely due to the absence of open high‑quality training data, transparent trajectory synthesis pipelines, or detailed training recipes. To this end, we introduce OpenSearch‑VL, a fully open‑source recipe for training frontier multimodal deep search agents with agentic reinforcement learning. First, we curated a dedicated pipeline to construct high‑quality training data through Wikipedia path sampling, fuzzy entity rewriting, and source‑anchor visual grounding, which jointly reduce shortcuts and one‑step retrieval collapse. Based on this pipeline, we curate two training datasets, SearchVL‑SFT‑36k for SFT and SearchVL‑RL‑8k for RL. Besides, we design a diverse tool environment that unifies text search, image search, OCR, cropping, sharpening, super‑resolution, and perspective correction, enabling agents to combine active perception with external knowledge acquisition. Finally, we propose a multi‑turn fatal‑aware GRPO training algorithm that handles cascading tool failures by masking post‑failure tokens while preserving useful pre‑failure reasoning through one‑sided advantage clamping. Built on this recipe, OpenSearch‑VL delivers substantial performance gains, with over 10‑point average improvements across seven benchmarks, and achieves results comparable to proprietary commercial models on several tasks. We will release all data, code, and models to support open research on multimodal deep search agents.
Authors:Yunhan Yang, Chunshi Wang, Junliang Ye, Yang Li, Zanxin Chen, Zehuan Huang, Yao Mu, Zhuo Chen, Chunchao Guo, Xihui Liu
Abstract:
Synthesizing physics‑grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on static geometry, overlooking the functional properties essential for interaction. We propose that interactive asset generation must be rooted in functional logic and hierarchical physics. To bridge this gap, we introduce PhysForge, a decoupled two‑stage framework supported by PhysDB, a large‑scale dataset of 150,000 assets with four‑tier physical annotations. First, a VLM acts as a "physical architect" to plan a "Hierarchical Physical Blueprint" defining material, functional, and kinematic constraints. Second, a physics‑grounded diffusion model realizes this blueprint by synthesizing high‑fidelity geometry alongside precise kinematic parameters via a novel KineVoxel Injection (KVI) mechanism. Experiments demonstrate that PhysForge produces functionally plausible, simulation‑ready assets, providing a robust data engine for interactive 3D content and embodied agents.
Authors:Bernhard Kainz, Johanna P Mueller, Matthew Baugh, Cosmin Bercea
Abstract:
Zero‑shot anomaly localisation via vision‑language models (VLMs) offers a compelling approach for rare pathology detection, yet its performance is fundamentally limited by the absence of healthy anatomical context. We reformulate zero‑shot localisation as a comparative inference problem in which anomalies are identified through structured comparison against reference distributions of normal anatomy. We introduce WALDO, a training‑free framework grounded in optimal transport theory that enables comparative reasoning through: (i) entropy‑weighted Sliced Wasserstein distances for anatomically‑aware reference selection from DINOv2 patch distributions, (ii) Goldilocks zone sampling exploiting the non‑monotonic relationship between reference similarity and localisation accuracy, and (iii) self‑consistency aggregation via weighted non‑maximum suppression. We theoretically analyse the Goldilocks effect through distributional divergence, and show that references with moderate similarity minimize a bias‑variance trade‑off in comparative visual reasoning. On the NOVA brain MRI benchmark, WALDO with Qwen2.5‑VL‑72B achieves 43.5_\pm1.6% mAP@30 (95% CI: [40.4, 46.7]), representing a 19% relative improvement over zero‑shot baselines. Cross‑model evaluation shows consistent gains: GPT‑4o achieves 32.0_\pm6.5% and Qwen3‑VL‑32B achieves 32.0_\pm6.6% mAP@30. Paired McNemar tests confirm statistical significance (p<0.01). Source code is available at https://github.com/bkainz/WALDO_MICCAI26_demo .
Authors:Yu-Hsi Chen, Abd-Krim Seghouane
Abstract:
Domain Generalization (DG) aims to learn representations that remain robust under out‑of‑distribution (OOD) shifts and generalize effectively to unseen target domains. While recent invariant learning strategies and architectural advances have achieved strong performance, explicitly discovering a structured domain‑invariant subspace through second‑order statistics remains underexplored. In this work, we propose CPCANet, a novel framework grounded in Common Principal Component Analysis (CPCA), which unrolls the iterative Flury‑Gautschi (FG) algorithm into fully differentiable neural layers. This approach integrates the statistical properties of CPCA into an end‑to‑end trainable framework, enforcing the discovery of a shared subspace across diverse domains while preserving interpretability. Experiments on four standard DG benchmarks demonstrate that CPCANet achieves state‑of‑the‑art (SOTA) performance in zero‑shot transfer. Moreover, CPCANet is architecture‑agnostic and requires no dataset‑specific tuning, providing a simple and efficient approach to learning robust representations under distribution shift. Code is available at https://github.com/wish44165/CPCANet.
Authors:Jingsen Zhu, Silvia Sellán, Alexander Terenin
Abstract:
We develop a framework for task‑specific active next‑best‑view selection in 3D reconstruction from point clouds, by casting the problem in the language of Bayesian decision theory. Our framework works by (a) placing a prior distribution over the space of implicit surfaces, (b) using recently‑developed stochastic surface reconstruction methods to calculate the resulting posterior distribution, then (c) using the posterior distribution to carefully reason about which view to scan next. This enables us to perform camera selection in a manner that is directly optimized for the intended use of the reconstructed data ‑ meaning, we reduce uncertainty only in those regions that make a difference in the task at hand, as opposed to prior approaches that reduce it uniformly across space. We evaluate our method across three distinct downstream tasks: semantic classification, segmentation, and PDE‑guided physics simulation. Experimental results demonstrate that our framework achieves superior task performance with fewer views compared to commonly used baselines and prior general uncertainty‑reduction techniques.
Authors:Maxim V. Shugaev, Md Reshad Ul Hoque, Bridget Kennedy, Joseph T. Riley, Fiona Hwang, Justin Hagen, Harvir Ghuman, Ethan Garcia-O'Donnell, Syed Noor Qadri, Freddie Santiago, Mun Wai Lee
Abstract:
Video sequence capturing through refractive dynamic media, such as a turbulent air or water surface, often suffer from severe geometric distortions and temporal instability. While recent advances address mild atmospheric turbulence, no existing benchmarks systematically evaluate restoration methods under strong and highly nonuniform refractive conditions. We present a comprehensive benchmark for geometric distortion removal in video, covering a range from turbulence‑like mild warping to strong discontinuous refractive deformations. The benchmark includes both laboratory‑captured real data and synthetic sequences generated for static scenes via physics‑based light refraction modeling across four distortion levels and multiple surface wave types. We evaluate a spectrum of methods from simple baselines and classical registration algorithms to advanced learning‑based approaches including DATUM and our proposed diffusion based V‑cache for high and extreme distortions regimes. Evaluation uses both pixel‑level (PSNR, SSIM), and perceptual (LPIPS, DINO, CLIP) metrics providing the first large scale analysis of geometric distortion removal. Our benchmark establishes a new foundation for developing and evaluating algorithms capable of reconstructing video from highly distorted optical environments. Our code and datasets are available at https://github.com/iafoss/refractive‑mfir‑benchmark.
Authors:Andranik Sargsyan, Shant Navasardyan
Abstract:
Accurate image segmentation is essential for modern computer vision applications such as image editing, autonomous driving, and medical image analysis. In recent years, Dichotomous Image Segmentation (DIS) has become a standard task for training and evaluating highly accurate segmentation models. Existing DIS approaches often fail to preserve fine‑grained details or fully capture the semantic structure of the foreground. To address these challenges, we present FlowDIS, a novel dichotomous image segmentation method built on the flow matching framework, which learns a time‑dependent vector field to transport the image distribution to the corresponding mask distribution, optionally conditioned on a text prompt. Moreover, with our Position‑Aware Instance Pairing (PAIP) training strategy, FlowDIS offers strong controllability through text prompts, enabling precise, pixel‑level object segmentation. Extensive experiments demonstrate that our method significantly outperforms state‑of‑the‑art approaches both with and without language guidance. Compared with the best prior DIS method, FlowDIS achieves a 5.5% higher F_β^ω measure and 43% lower MAE (\mathcalM) on the DIS‑TE test set. The code is available at: https://github.com/Picsart‑AI‑Research/FlowDIS
Authors:Yuan Wu, Zhiqiang Yan, Jiawei Lian, Zhengxue Wang, Jian Yang
Abstract:
3D occupancy prediction aims to infer dense, voxel‑wise scene semantics from sensor observations, where the 2D‑to‑3D view transformation serves as a crucial step in bridging image features and volumetric representations. Most previous methods rely on a fixed projection space, where 3D reference points are uniformly sampled along pillars. However, such sampling struggles to capture the sparsity and height variations of real‑world scenes, leading to ambiguous correspondences and unreliable feature aggregation. To address these challenges, we propose HiPR, a camera‑LiDAR occupancy framework with Height‑Guided Projection Reparameterization. HiPR first encodes LiDAR into a BEV height map to capture the maximum height of the point cloud. HiPR then adjusts the sampling range of each pillar using the height prior, enabling adaptive reparameterization of the projection space. As a result, the projected points are redistributed into geometrically meaningful regions rather than fixed ranges. Meanwhile, we mask out the invalid parts of the height map to avoid misleading the feature aggregation. In addition, to alleviate the training instability caused by noisy LiDAR‑derived heights, we introduce a training‑time Progressive Height Conditioning strategy, which gradually transitions the conditioning signal from ground‑truth heights to LiDAR heights. Extensive experiments demonstrate that HiPR consistently outperforms existing state‑of‑the‑art methods while maintaining real‑time inference. The code and pretrained models can be found at https://github.com/yanzq95/HiPR.
Authors:Avhishek Biswas, Apala Pramanik, Eylem Ekici, Mehmet C. Vuran
Abstract:
Millimeter‑wave (mmWave) frequencies promise multi‑gigabit connectivity for vehicle‑to‑everything (V2X) networks, but face challenges in terms of severe path loss and mobility‑related beam misalignment. Reliable V2X connectivity requires fast, double‑directional beam alignment. However, existing methods suffer from high training overhead and limited generalization to unseen scenarios. This paper presents VIsion‑based BEamforming(VIBE), a hybrid model‑based, closed‑loop, learning architecture for real‑time double‑directional mmWave beam management primed by camera sensing. VIBE fuses machine learning, model‑based reasoning, and closed‑loop RF feedback to balance beam‑pair establishment latency with link quality. VIBE bypasses exhaustive training overhead and accelerates link establishment by leveraging camera observations to reduce the beam‑search space. Lightweight beam refinement and offset tracking mechanisms adaptively refine beams in response to dynamic application requirements. VIBE is implemented and evaluated across online indoor/outdoor testbeds, public datasets, and real‑time vehicular experiments, demonstrating strong generalization capabilities, making it suitable for real‑time V2X communication. Comparisons with 5G NR hierarchical beamforming show that VIBE consistently maintains lower outage rates. Furthermore, VIBE outperforms state‑of‑the‑art end‑to‑end ML models for beam selection when evaluated on public datasets and achieves outage rates as low as 1.1‑1.4 %. The results show that a hybrid model‑based, closed‑loop learning architecture is better suited for real‑world mmWave vehicular connectivity than end‑to‑end trained ML models. For reproducibility, we publish our code to https://github.com/UNL‑CPN‑Lab/Look‑Once‑Beam‑Twice.
Authors:Wen Wen, Hao Chen, Shiliang Zhang
Abstract:
Lifelong person re‑identification (LReID) aims to train a generalizable model with sequentially collected data. However, such models often suffer from semantic drift, limited adaptability, and catastrophic forgetting as new domains emerge. Existing exemplar‑free approaches largely rely on visual‑only distillation or parameter regularization, while overlooking the potential of auxiliary modalities, such as text, to preserve semantic stability and enable incremental plasticity. We observe that the frozen text encoder in pretrained vision‑language models can serve as a stable semantic anchor across domains. To decouple the roles of vision and text, we propose Prompt‑Anchored vision‑text Distillation (PAD), an asymmetric vision‑text framework for semantic alignment and cross‑domain generalization. On the textual side, we distill prompts to preserve vision‑text alignment under a fixed semantic space, acting as a global semantic reference rather than a dominant learning signal. On the visual side, an EMA‑based teacher with an adaptive prompt pool enables domain‑wise adaptation by allocating new slots while freezing past ones. Extensive experiments show that PAD substantially outperforms state‑of‑the‑art methods across seen and unseen domains, achieving a strong balance between stability and plasticity. Project page is available at https://github.com/zu‑zi/PAD.
Authors:Ali Shibli, Andrea Nascetti, Yifang Ban
Abstract:
Wildfire burned‑area mapping is essential for damage assessment, emissions modeling, and understanding fire‑climate interactions across diverse ecological regions. Recent geospatial foundation models provide strong general‑purpose representations for satellite imagery, yet there is still no clear understanding of how to efficiently adapt these models for downstream Earth observation tasks, particularly under geographic and temporal domain shift. This study evaluates three state‑of‑the‑art Geospatial Foundation Models (GFMs) ‑ Terramind, DINOv3, and Prithvi‑v2 ‑ for burned‑area mapping across the United States and Canada using Sentinel‑2 data. Leveraging 3,820 wildfire events from 2017‑2023, we conduct spatial and temporal generalization tests across diverse biomes. We systematically compare full fine‑tuning, decoder‑only fine‑tuning, and Low‑Rank Adaptation (LoRA) for adapting each model. Across all experiments, LoRA provides the strongest cross‑domain generalization while updating less than 1% of parameters, demonstrating a favorable trade‑off between accuracy and efficiency. Prithvi‑v2 with LoRA achieves the highest overall accuracy and the largest improvement compared to full fine‑tuning. These findings indicate that geospatial foundation models, when adapted using lightweight parameter‑efficient methods such as LoRA, offer a robust and scalable solution for large‑scale burned‑area mapping. Code is available at https://github.com/alishibli97/wildfire‑lora‑gfm.
Authors:Raphaël Delécluse, Hazem Wannous, Laurent Guimas
Abstract:
This companion paper reports the ICPR 2026 TVRID competition on privacy‑aware top‑view person re‑identification. We present the competition setting, the released RGB‑Depth dataset, and a summary of final results with descriptions of the top entries. TVRID contains 86 identities captured by four synchronized overhead Intel RealSense D455 cameras, with paired RGB/Depth streams and structured geometric variation across flat, ascent, descent, and oblique viewpoints. The evaluation protocol includes three tracks: RGB Re‑ID, Depth Re‑ID, and RGB\leftrightarrowDepth cross‑modal retrieval. Submissions are ranked using mAP and CMC‑1 under a unified server‑side evaluation. The final results show a clear difficulty ordering (RGB > Depth > Cross‑Modal), highlighting both the challenge of modality‑constrained retrieval and the feasibility of strong performance with modality‑invariant learning. By releasing the dataset at https://zenodo.org/records/17909410, the evaluation scripts at https://github.com/RaphaelDel/ICPR‑TVRID, and the accompanying documentation, TVRID establishes a reproducible benchmark for top‑view, depth‑based, and cross‑modal person re‑id.
Authors:Mohamed Elhabebe, Ayman El-Baz, Qing Liu
Abstract:
Automated glaucoma detection is critical for preventing irreversible vision loss and reducing the burden on healthcare systems. However, ensuring fairness across diverse patient populations remains a significant challenge. In this paper, we propose FairEnc, a fair pretraining method for vision‑language models (VLMs) that enables simultaneous debiasing across multiple sensitive attributes. FairEnc jointly mitigates biases in both textual and visual modalities with respect to multiple sensitive attributes, including race, gender, ethnicity, and language. Specifically, for the textual encoder, we leverage a large language model to generate synthetic clinical descriptions with varied sensitive attributes while preserving disease semantics, and employ a contrastive alignment objective to encourage demographic‑invariant representations. For the visual encoder, we propose a dual‑level fairness strategy that combines mutual information regularization to reduce statistical dependence between learned features and demographic groups, with multi‑discriminator adversarial debiasing. Comprehensive experiments on the publicly available Harvard‑FairVLMed dataset demonstrate that FairEnc effectively reduces demographic disparity as measured by DPD and DEOdds while achieving strong diagnostic performance under both zero‑shot and linear probing evaluations. Additional experiments on the private FairFundus dataset show that FairEnc consistently preserves fairness advantages under cross‑domain and cross‑modality settings and maintains diagnostic performance within a competitive range. These results highlight FairEnc's ability to generalize fairness under distribution shifts, supporting its potential for more equitable deployment in real‑world clinical settings. Our codebase and synthetic clinical notes are available at https://github.com/Mohamed‑Elhabebe/FairEnc
Authors:Xinze Li, Bohan Yang, Pengxu Chen, Yiyuan Wang, Hongcheng Luo, Wentao Cheng, Weifeng Su
Abstract:
3D Gaussian Splatting (3DGS) has emerged as an advanced technique for real‑time novel view synthesis by representing scene geometry and appearance using differentiable Gaussian primitives. However, efficiently computing precise Gaussian‑tile intersections remains a critical task in the rasterization pipeline. To this end, we propose QuadBox, a method that leverages four axis‑aligned bounding boxes to tightly encapsulate projected Gaussians in a discrete manner. First, we derive a geometry‑aware stretching factor that enables the construction of a tile‑aligned QuadBox, which covers the elliptical projection and largely excludes irrelevant tiles. Second, we introduce QPass, a single‑pass tile traversal algorithm that exhaustively exploits the discrete nature of QuadBox, ensuring that the tile intersection check is performed with simple interval tests. Experiments on public datasets show that our method accelerates the rendering speed of 3DGS by 1.85×. Code is available at \hrefhttps://github.com/Powertony102/QuadBoxhttps://github.com/Powertony102/QuadBox.
Authors:Laura Bravo-Sánchez, Matthieu Armando, Romain Brégier, Grégory Rogez, Serena Yeung-Levy, Fabien Baradel
Abstract:
Recovering 3D human pose and shape from a single image remains a cornerstone of human‑centric vision, yet most methods assume adult subjects and optimize each person independently. These assumptions fail in real‑world, all‑age scenes, where body proportions and depth must be resolved jointly. We introduce Anny‑Fit, a multi‑person, camera‑space optimization framework for all‑age 3D human mesh recovery (HMR). Unlike existing per‑person fitting methods, Anny‑Fit jointly optimizes all individuals directly in the camera coordinate system, enforcing global spatial consistency. At the core of our approach is the use of multiple forms of expert knowledge ‑‑ including metric depth maps, instance segmentation, 2D keypoints, and, VLM‑derived semantic attributes such as age and gender ‑‑ each obtained from dedicated off‑the‑shelf networks. These complementary signals jointly guide the optimization, constraining the depth‑scale ambiguity characteristic of all‑age scenes. Across diverse datasets, Anny‑Fit consistently improves 2D reprojection accuracy (+13 to 16), relative depth ordering (+6 to 7), 3D estimation error (‑9 to ‑29) and shape estimation (+25 to +82), producing more coherent scenes. Finally, we show that VLM‑based semantic knowledge can be distilled into an HMR model via the pseudo‑ground‑truth annotations produced by Anny‑Fit on training data, enabling it to learn semantically meaningful shape parameters while improving HMR performance. Our approach bridges adult‑only and all‑age modeling by enabling zero‑shot adaptation of adult‑trained HMR pipelines to the full age spectrum without retraining. Code is publicly available at https://github.com/naver/anny‑fit.
Authors:Yihan Lin, Haoyang Li, Yang Li, Haitao Shen, Yihan Zhao, Chao Shao, Jing Zhang
Abstract:
Latent actions serve as an intermediate representation that enables consistent modeling of vision‑language‑action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of latent action supervision from two perspectives: (i) regularizing the trajectory via image‑based latent actions, and (ii) unifying the target space with action‑based latent actions. Under a unified VLA baseline, we instantiate and compare four representative integration strategies. Our results reveal a formulation‑task correspondence: image‑based latent actions benefit long‑horizon reasoning and scene‑level generalization, whereas action‑based latent actions excel at complex motor coordination. Furthermore, we find that directly supervising the VLM with discrete latent action tokens yields the most effective performance. Finally, our experiments offer initial insights into the benefits of latent action supervision in mixed‑data, suggesting a promising direction for VLA training. Code is available at https://github.com/RUCKBReasoning/From_Pixels_to_Tokens.
Authors:Zhiwei Yang, Pengfei Song, Yucong Meng, Kexue Fu, Shuo Wang, Zhijian Song
Abstract:
Weakly Supervised Semantic Segmentation (WSSS) with image‑level labels typically leverages Class Activation Maps (CAMs) to achieve pixel‑level predictions. Recently, Contrastive Language‑Image Pre‑training (CLIP) has been introduced to generate CAMs in WSSS. However, previous WSSS methods solely adopt CLIP's vision‑language paired property for dense localization, neglecting its inherently limited dense knowledge across both visual and text modalities, which renders CAM generation suboptimal. In this work, we propose DiCLIP, a novel WSSS framework that leverages the generative diffusion model to enhance CLIP's dense knowledge across two modalities. Specifically, Visual Correlation Enhancement (VCE) and Text Semantic Augmentation (TSA) modules are proposed for dense prediction enhancement. To improve the spatial awareness of visual features, our VCE module utilizes diffusion's reliable spatial consistency to mitigate the over‑smoothing issue in CLIP's attention. It designs the Attention Clustering Refinement (ACR) module to reliably extract diverse correlation maps from the diffusion model. The correlation maps act as a diversity bias for CLIP's self‑attention, recursively pushing its visual features towards a more discriminative dense distribution. To augment the semantics of text embeddings, our TSA module argues that a single text modality is insufficient to encompass the variability of visual categories. Thus, we leverage diffusion's generative power to maintain a dynamic key‑value cache model, shifting CAM generation from a patch‑text matching mechanism to a novel visual knowledge retrieval paradigm. With these enhancements, DiCLIP not only outperforms state‑of‑the‑art methods on PASCAL VOC and MS COCO but also significantly reduces training costs. Code is publicly available at https://github.com/zwyang6/DiCLIP.
Authors:Boyue Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu
Abstract:
UAV‑ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two‑stream methods suffer from isolated feature extraction and rely heavily on implicit appearance matching, which struggles to establish reliable correspondence under drastic view differences, leading to tracking unreliability. To address these limitations, we propose VL‑UniTrack, a fully unified framework enhanced by visual‑language prompts. By encoding features from both views within a single shared encoder, our method breaks the barrier of feature isolation to facilitate sufficient cross‑view interaction. To overcome the ambiguity caused by relying solely on appearance matching, we design visual‑language geometric prompting module, which fuses language descriptions with visual features to generate learnable prompts. These prompts are then fed into our prompt‑guided cross‑view adapter module to enable sufficient cross‑view feature interaction and to guide the learning of view‑specific feature representations. Furthermore, a confidence‑modulated mutual distillation loss is proposed to regularize the training by mitigating noise propagation. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance on the latest benchmark. The code can be downloaded in https://github.com/xuboyue1999/VL‑UniTrack.git
Authors:Jiaqian Zhang, Hao Wei, Chenyang Ge, Yanhui Zhou
Abstract:
Perceptual image compression focuses on preserving high visual quality under low‑bitrate constraints. Most existing approaches to perceptual compression leverage the strong generative capabilities of generative adversarial networks or diffusion models, at the cost of substantial model complexity. To this end, we present an efficient perceptual image compression method that exploits the long‑range modeling capability and linear computational complexity of state space models, with a particular focus on Mamba. Unlike existing methods that rely on an inherently fixed scanning order and consequently impair semantic continuity and spatial correlation, we develop a semantic‑aware Mamba block (SAMB) to enable scanning guided by dynamically clustered semantic features, thereby alleviating the strict causality constraints and long‑range information decay inherent to Mamba. Inspired by singular value decomposition, we design an SVD‑inspired redundancy reduction module (SVD‑RRM) that performs a low‑rank approximation on the latent features by introducing a learnable soft threshold, leading to channel‑wise redundancy information reduction. The proposed SAMB is integrated into both the encoder and decoder of the compression framework, whereas the SVD‑RRM is incorporated only in the encoder. Extensive experiments demonstrate that our method performs favorably against state‑of‑the‑art approaches in terms of rate‑distortion‑perception tradeoff and model complexity. The source code and pretrained models will be available at https://github.com/Jasmine‑aiq/SAMIC.
Authors:Kaili Zheng, Kaiwen Wang, Xun Zhu, Chenyi Guo, Ji Wu
Abstract:
Humans constantly interact with their surroundings. Existing end‑to‑end multi‑person human mesh recovery methods, typically based on the DETR framework, capture inter‑human relationships through self‑attention across all human queries. However, these approaches model interactions only implicitly and lack explicit reasoning about how humans interact with objects and with each other. In this paper, we propose InterMesh, a simple yet effective framework that explicitly incorporates human‑environment interaction information into human mesh recovery pipeline. By leveraging a human‑object interaction detector, InterMesh enriches query representations with structured interaction semantics, enabling more accurate pose and shape estimation. We design lightweight modules, Contextual Interaction Encoder and Interaction‑Guided Refiner, to integrate these features into existing HMR architectures with minimal overhead. We validate our approach through extensive experiments on 3DPW, MuPoTS, CMU Panoptic, Hi4D, and CHI3D datasets, demonstrating remarkable improvements over state‑of‑the‑art methods. Notably, InterMesh reduces MPJPE by 9.9% on CMU Panoptic and 8.2% on Hi4D, highlighting its effectiveness in scenarios with complex human‑object and inter‑human interactions. Code and models are released at https://github.com/Kelly510/InterMesh.
Authors:Anagh Malik, Dorian Chan, Xiaoming Zhao, David B. Lindell, Oncel Tuzel, Jen-Hao Rick Chang
Abstract:
We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud, to construct. Specifically, Velox trains an encoder to compress spatiotemporal color point clouds into a set of dynamic shape tokens. These tokens are supervised using two complementary decoders: a 4D surface decoder, which models the time‑varying surface distribution capturing the geometry; and a Gaussian decoder, which maps the tokens to 3D Gaussians, helping learn appearance. To demonstrate the utility of our representation, we evaluate it across three downstream tasks ‑‑ video‑to‑4D generation, 3D tracking, and cloth simulation via image‑to‑4D generation ‑‑ and observe strong performances in all settings.
Authors:Binh Long Nguyen, Kien Nguyen, Sridha Sridharan, Clinton Fookes, Peyman Moghadam
Abstract:
We introduce Ilov3Splat, a novel framework for instance‑level open‑vocabulary 3D scene understanding built on 3D Gaussian Splatting (3D‑GS). Most prior work depends on 2D rendering‑based matching or point‑level semantic association, which undermines cross‑view consistency, lacks coherent instance‑level reasoning, and limits precision in downstream 3D tasks. To address these limitations, our method jointly optimizes scene geometry and semantic representations by augmenting Gaussian splats with view‑consistent feature fields. Specifically, we leverage multi‑resolution hash embedding to efficiently encode language‑aligned CLIP features, enabling dense and coherent language grounding in 3D space. We further train an instance feature field using contrastive loss over SAM masks, supporting fine‑grained object distinction across views. At inference time, CLIP‑encoded queries are matched against the learned features, followed by two‑stage 3D clustering to retrieve relevant Gaussian groups. This enables our framework to identify arbitrary objects in 3D scenes based on natural language descriptions, without requiring category supervision or manual annotations. Experiments on standard benchmarks demonstrate that Ilov3Splat outperforms prior open‑vocabulary 3D‑GS methods in both object selection and instance segmentation, offering a flexible and accurate solution for language‑driven 3D scene understanding. Project page: https://csiro‑robotics.github.io/Ilov3Splat.
Authors:Jingtao Zhou, Xirui Kang, Feiyang Huang, Lai-Man Po
Abstract:
Existing prompt learning for VLMs exhibits a modality asymmetry, predominantly optimizing text tokens while still relying on frozen visual encoder as holistic extractor and neglecting the spectral granularity essential for fine‑grained discrimination. To bridge this, we introduce Disentangling Spectral Granularity for Prompt Learning (SpecPL), which approaches prompt learning from a novel spectral perspective via Counterfactual Granule Supervision. Specifically, we leverage a frozen VAE to decompose visual signals into semantic low‑frequency bands and granular high‑frequency details. A frozen Visual Semantic Bank anchors text representations to universal low‑frequency invariants, mitigating overfitting. Crucially, fine‑grained discrimination is driven by counterfactual granule training: by permuting high‑frequency signals, we compel the model to explicitly distinguish visual granularity from semantic invariance. Uniquely, SpecPL serves as a universal plug‑and‑play booster, revitalizing text‑oriented baselines like CoOp and MaPLe via visual‑side guidance. Experiments on 11 benchmarks demonstrate competitive state‑of‑the‑art performance, achieving a new performance ceiling of 81.51% harmonic‑mean accuracy. These results validate that spectral disentanglement with counterfactual supervision effectively bridges the gap in the stability‑generalization trade‑off. Code is released at https://github.com/Mlrac1e/SpecPL‑Prompt‑Learning.
Authors:ZhiXin Sun
Abstract:
In recent years, object detection has achieved significant progress, especially in the field of open‑vocabulary object detection. Unlike traditional methods that rely on predefined categories, open‑vocabulary approaches can detect arbitrary objects based on human‑provided prompts. With the advancement of prompt‑based detection techniques, models such as SAM3 can even outperform some category‑specific detectors trained on particular datasets without requiring additional training on those datasets. However, despite these advancements, false positives and false negatives still occur. In practical engineering applications, persistent misdetections or missed detections of the same object are unacceptable. Yet retraining the model every time such errors occur incurs substantial costs in terms of human effort, computational resources, and time. Therefore, how to leverage existing false positive and false negative samples to prevent such errors from recurring remains a highly challenging and urgent problem. To address this issue, we propose EBOD (Example‑Based Object Detection), which integrates a prompt‑based detector (SAM3) with robust feature matching modules (DINOv3 and LightGlue). The proposed framework effectively suppresses the repeated occurrence of false positives and false negatives by leveraging previous error examples, without requiring additional model retraining. Code is available at https://github.com/sunzx97/examples_based_object_detection.
Authors:Chunwei Tian, Jingyuan Xie, Qi Zhang, Chao Li, Wangmeng Zuo, Shichao Zhang
Abstract:
Deep neural networks enriched with structural information have been widely employed for facial expression recognition tasks. However, these methods often depend on hierarchical information rather than face property to finish expression recognition. In this paper, we propose a cross‑modal network with strong biological and structural information for facial expression recognition (CMNet). CMNet can respectively learn expression information via face symmetry on a whole face, left and right half faces to extract complementary facial features. To prevent negative effect of biological and structural information fusion, a salient facial information refinement module can obtain salient facial expression information to improve stability of an obtained facial expression classifier. To reduce reliance on unilateral facial features, a half‑face alignment optimization mechanism is designed to align obtained expression information of learned left and right half faces. Our experimental results demonstrate that CMNet outperforms several novel methods, i.e., SCN and LAENet‑SA for facial expression recognition. Codes can be obtained at https://github.com/hellloxiaotian/CMNet.
Authors:Shuo Wang, Jilin Mei, Fuyang Liu, Wenfei Guan, Fanjie Kong, Zhihua Zhao, Shuai Wang, Chen Min, Yu Hu
Abstract:
Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off‑road scenes, its performance degrades due to high‑frequency geometry, ego‑motion jitter, and increased non‑rigid dynamics. These factors introduce conflicting Gaussian observations across timestamps, leading to either over‑smoothed renderings or structural artifacts. To address this issue, we propose Ground4D, a spatially‑grounded 4D feedforward framework for pose‑free off‑road reconstruction. The key idea is to resolve temporal conflicts through spatially localized conditioning. Specifically, we introduce voxel‑grounded temporal Gaussian aggregation, which partitions the canonical Gaussian space into spatial voxels and performs query‑conditioned temporal attention within each voxel. Intra‑voxel softmax normalization ensures that temporal selectivity and spatial occupancy become mutually reinforcing rather than conflicting. We furthermore introduce surface normal cues as auxiliary geometric guidance to regularize the geometry of Gaussian primitives. Extensive experiments on ORAD‑3D and RELLIS‑3D demonstrate that Ground4D consistently outperforms existing feedforward methods in reconstruction quality and generalizes zero‑shot to unseen off‑road domains. Project page and code:https://github.com/wsnbws/Ground4D.
Authors:Yupeng Gao, Tianyu Li, Guoqing Wang, Yang Yang
Abstract:
Remote Sensing Image Change Captioning (RSICC) aims to generate spatially grounded natural language descriptions of scene evolution from bi‑temporal imagery, moving beyond binary change masks toward semantic‑level understanding. However, existing methods rely on implicit feature differencing without explicitly modeling structured change semantics, and struggle to reconcile the conflicting representation demands of change detection and caption generation. In addition, current benchmarks provide limited coverage of high‑resolution urban construction scenarios. To address these challenges, we propose PTNet, a prototype‑guided task‑adaptive framework for joint change captioning and detection. PTNet explicitly models structured change semantics through a learnable prototype bank that guides cross‑temporal interaction, disentangles task‑specific representations via multi‑head gating, and injects detection‑derived spatial priors into caption generation, enabling coherent semantic correspondence while preserving fine‑grained spatial sensitivity. Furthermore, we construct UCCD, a large‑scale UAV‑based benchmark comprising 9,000 high‑resolution image pairs and 45,000 annotated sentences for urban construction monitoring. Extensive experiments on UCCD and WHU‑CDC demonstrate that PTNet consistently outperforms existing methods. The dataset and source code are publicly available at https://github.com/G124556/ptnet.
Authors:Lin Song, Wenbo Li, Guoqing Ma, Wei Tang, Bo Wang, Yuan Zhang, Yijun Yang, Yicheng Xiao, Jianhui Liu, Yanbing Zhang, Guohui Zhang, Wenhu Zhang, Hang Xu, Nan Jiang, Xin Han, Haoze Sun, Maoquan Zhang, Haoyang Huang, Nan Duan
Abstract:
We present JoyAI‑Image, a unified multimodal foundation model for visual understanding, text‑to‑image generation, and instruction‑guided image editing. JoyAI‑Image couples a spatially enhanced Multimodal Large Language Model (MLLM) with a Multimodal Diffusion Transformer (MMDiT), allowing perception and generation to interact through a shared multimodal interface. Around this architecture, we build a scalable training recipe that combines unified instruction tuning, long‑text rendering supervision, spatially grounded data, and both general and spatial editing signals. This design gives the model broad multimodal capability while strengthening geometry‑aware reasoning and controllable visual synthesis. Experiments across understanding, generation, long‑text rendering, and editing benchmarks show that JoyAI‑Image achieves state‑of‑the‑art or highly competitive performance. More importantly, the bidirectional loop between enhanced understanding, controllable spatial editing, and novel‑view‑assisted reasoning enables the model to move beyond general visual competence toward stronger spatial intelligence. These results suggest a promising path for unified visual models in downstream applications such as vision‑language‑action systems and world models.
Authors:Chamani Shiranthika, Hadi Hadizadeh, Parvaneh Saeedi
Abstract:
Federated Learning enables decentralized training by aggregating model updates across clients without sharing raw data, while Split Federated Learning further partitions the model between clients and a server to reduce computation and communication at the client side. However, decentralized medical institutions rarely operate on a single shared task, making standard Federated and SplitFed collaborations poorly aligned with real clinical workflows. Multi‑task FL extends these frameworks by allowing clients to handle different tasks, but often introduces instability and privacy vulnerabilities. This study proposes MuCALD‑SplitFed, a multi‑task SplitFed framework that integrates causal representation learning and latent diffusion. Experiments show MuCALD‑SplitFed consistently improves segmentation, while baseline SplitFed fails to converge. The proposed approach further reduces information leakage at split points, mitigating reconstruction‑based and membership inference attacks. Additionally, MuCALD SplitFed outperforms state‑of‑the‑art personalized FL and multi‑task FL approaches. The code repository is: https://github.com/ChamaniS/MuCALD_SplitFed.
Authors:Nicolas Michel, Maorong Wang, Jiangpeng He, Toshihiko Yamasaki
Abstract:
Deep learning models continue to scale, with some requiring more storage than many large‑scale datasets. Thus, we introduce a new paradigm: Continual Distillation (CD), where a student learns sequentially from a stream of teacher models without retaining access to earlier teachers. CD faces two challenges: teacher training data is unavailable, and teachers have varying expertise. We show that external unlabeled data enables Unseen Knowledge Transfer (UKT), allowing the student to acquire information from domains not present in the training data, while known to the teacher. We also show that sequential distillation causes Unseen Knowledge Forgetting (UKF) when transferred knowledge is lost after training on later teachers. To better trade off between UKT and UKF, we propose Self External Data Distillation (SE2D), a method that preserves logits on external data to stabilize learning across heterogeneous teachers. Experiments on multiple benchmarks show that SE2D reduces UKF and improves cross‑domain generalization. The code and implementation for this work are publicly available at: https://github.com/Nicolas1203/continual_distillation.
Authors:You Qin, Kai Liu, Shengqiong Wu, Kai Wang, Shijian Deng, Yapeng Tian, Junbin Xiao, Yazhou Xing, Yinghao Ma, Bobo Li, Roger Zimmermann, Lei Cui, Furu Wei, Jiebo Luo, Hao Fei
Abstract:
Audio‑Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of large foundation models, joint modeling of audio and vision has become increasingly crucial, i.e., not only for understanding but also for controllable generation and reasoning across dynamic, temporally grounded signals. Recent advances, such as Meta MovieGen and Google Veo‑3, highlight the growing industrial and academic focus on unified audio‑vision architectures that learn from massive multimodal data. However, despite rapid progress, the literature remains fragmented, spanning diverse tasks, inconsistent taxonomies, and heterogeneous evaluation practices that impede systematic comparison and knowledge integration. This survey provides the first comprehensive review of AVI through the lens of large foundation models. We establish a unified taxonomy covering the broad landscape of AVI tasks, ranging from understanding (e.g., speech recognition, sound localization) to generation (e.g., audio‑driven video synthesis, video‑to‑audio) and interaction (e.g., dialogue, embodied, or agentic interfaces). We synthesize methodological foundations, including modality tokenization, cross‑modal fusion, autoregressive and diffusion‑based generation, large‑scale pretraining, instruction alignment, and preference optimization. Furthermore, we curate representative datasets, benchmarks, and evaluation metrics, offering a structured comparison across task families and identifying open challenges in synchronization, spatial reasoning, controllability, and safety. By consolidating this rapidly expanding field into a coherent framework, this survey aims to serve as a foundational reference for future research on large‑scale AVI.
Authors:Prajnan Goswami, Tianye Ding, Feng Liu, Huaizu Jiang
Abstract:
Visual correspondence across image‑to‑image (2D‑2D), image‑to‑point cloud (2D‑3D), and point cloud‑to‑point cloud (3D‑3D) geometric matching forms the foundation for numerous 3D vision tasks. Despite sharing a similar problem structure, current methods use task‑specific designs with separate models for each modality combination. We present UniCorrn, the first correspondence model with shared weights that unifies geometric matching across all three tasks. Our key insight is that Transformer attention naturally captures cross‑modal feature similarity. We propose a dual‑stream decoder that maintains separate appearance and positional feature streams. This design enables end‑to‑end learning through stack‑able layers while supporting flexible query‑based correspondence estimation across heterogeneous modalities. Our architecture employs modality‑specific backbones followed by shared encoder and decoder components, trained jointly on diverse data combining pseudo point clouds from depth maps with real 3D correspondence annotations. UniCorrn achieves competitive performance on 2D‑2D matching and surpasses prior state‑of‑the‑art by 8% on 7Scenes (2D‑3D) and 10% on 3DLoMatch (3D‑3D) in registration recall. Project website: https://neu‑vi.github.io/UniCorrn
Authors:Evangelos Ntavelis, Sean Wu, Mohamad Shahbazi, Fabio Maninchedda, Dmitry Kostiaev, Artem Sevastopolsky, Vittorio Megaro, Trevor Phillips, Alejandro Blumentals, Shridhar Ravikumar, Mehak Gupta, Reinhard Knothe, Jeronimo Bayer, Matthias Vestner, Simon Schaefer, Thomas Etterlin, Christian Zimmermann, Mathias Deschler, Peter Kaufmann, Stefan Brugger, Sebastian Martin, Brian Amberg, Tom Runia
Abstract:
We propose HeadsUp, a scalable feed‑forward method for reconstructing high‑quality 3D Gaussian heads from large‑scale multi‑camera setups. Our method employs an efficient encoder‑decoder architecture that compresses input views into a compact latent representation. This latent representation is then decoded into a set of UV‑parameterized 3D Gaussians anchored to a neutral head template. This UV representation decouples the number of 3D Gaussians from the number and resolution of input images, enabling training with many high‑resolution input views. We train and evaluate our model on an internal dataset with more than 10,000 subjects, which is an order of magnitude larger than existing multi‑view human head datasets. HeadsUp achieves state‑of‑the‑art reconstruction quality and generalizes to novel identities without test‑time optimization. We extensively analyze the scaling behavior of our model across identities, views, and model capacity, revealing practical insights for quality‑compute trade‑offs. Finally, we highlight the strength of our latent space by showcasing two downstream applications: generating novel 3D identities and animating the 3D heads with expression blendshapes.
Authors:Zhangnan Jiang, Zichen Yang
Abstract:
Nowadays as convolution neural networks demonstrate its powerful problem‑solving ability in the area of image processing, efforts have been made to reconstruct detailed face shapes from 2D face images or videos. However, to make the full use of CNN, a large number of labeled data is required to train the network. Coarse morphable face model has been used to synthesize labeled data. However, it is hard for coarse morphable face models to generate photo‑realistic data with detail such as wrinkles. In this project, we present a pipeline that reconstructs a human face 3D model from a single RGB image. The pipeline includes face detection, landmark detection, regression of 3DMM model parameters, and soft rendering. Mentor: Zhipeng Fan (Email: zf606@nyu.edu) Code Repository: https://github.com/SeVEnMY/3d‑face‑ reconstruction Code Reference: https://github.com/sicxu/Deep3DFaceRecon pytorch
Authors:Zhiyu Pan, Xiongjun Guan, Jianjiang Feng, Jie Zhou
Abstract:
Contactless fingerprint recognition has gained increasing attention due to its advantages in hygiene and acquisition flexibility. However, the absence of physical contact constraints introduces severe nonlinear geometric distortions caused by free finger poses in 3D space, resulting in a substantial cross‑modal domain gap between contactless and conventional contact‑based fingerprints. Existing solutions largely rely on explicit geometric correction or image enhancement, which are fragile under extreme pose variations. In this paper, we propose Identity‑Consistent Multi‑Pose Generation of Contactless Fingerprints (IMPOSE), a physics‑inspired framework that synthesizes identity‑preserving, multi‑pose contactless fingerprint samples to empower recognition models. IMPOSE consists of three stages: (1) rolled fingerprint identity generation via latent diffusion with discrete codebook representations, (2) cross‑modal translation from rolled to contactless modality guided by Sauvola‑based local adaptive binarization as an identity anchor, and (3) physics‑based multi‑pose simulation through 3D finger model texture mapping and projection. The generated samples maintain strict identity consistency at the ridge topology level and spatial alignment with standard fingerprint coordinate space. Extensive experiments on the UWA and PolyU CL2CB databases demonstrate that fine‑tuning fixed‑length dense descriptors (FDD) with IMPOSE‑synthesized data achieves state‑of‑the‑art cross‑modal matching, reducing EER to 8.74% on UWA and 2.26% on PolyU CL2CB. Synthetic data also yields consistent gains across mainstream representations including DeepPrint and AFRNet, and the hybrid strategy combining synthetic and real data achieves the best overall results. The code and generated samples are available at https://github.com/Yu‑Yy/IMPOSE.
Authors:Xun Jiang, Yufan Gu, Disen Hu, Yuqing Hou, Yazhou Yao, Fumin Shen, Heng Tao Shen, Xing Xu
Abstract:
Multimodal learning often grapples with the challenge of low‑quality data, which predominantly manifests as two facets: modality imbalance and noisy corruption. While these issues are often studied in isolation, we argue that they share a common root in the predictive uncertainty towards the reliability of individual modalities and instances during learning. In this paper, we propose a unified framework, termed Conformal Predictive Self‑Calibration (CPSC), which leverages conformal prediction to equip the model with the ability to perform self‑guided calibration on‑the‑fly. The core of our proposed CPSC lies in a novel self‑calibrating training loop that seamlessly integrates two key modules: (1) Representation Self‑Calibration, which decomposes unimodal features into components, and selectively fuses the most robust ones identified by a conformal predictor to enhance feature resilience. (2) Gradient Self‑Calibration, which recalibrates the gradient flow during backpropagation based on instance‑wise reliability scores, steering the optimization towards more trustworthy directions. Furthermore, we also devise a self‑update strategy for the conformal predictor to ensure the entire system co‑evolves consistently throughout the training process. Extensive experiments on six benchmark datasets under both imbalanced and noisy settings demonstrate that our CPSC framework consistently outperforms existing state‑of‑the‑art methods. Our code is available at https://github.com/XunCHN/CPSC.
Authors:Robert Martinko, Daniel Steininger, Julia Simon, Andreas Trondl, Matthias Blaickner
Abstract:
Rising global food demand and growing climate pressure increase the need for sustainable, precise agricultural practices. Automated, individualized plant treatment relies on fine‑grained visual analysis, yet leaf‑level segmentation remains underexplored despite its value for assessing crop health, growth dynamics, yield potential and localized stress symptoms. Progress is limited by a lack of dedicated datasets, especially regarding species coverage, and by the absence of systematic evaluations of modern instance‑segmentation architectures for this task. We address these gaps by surveying current data and identifying four suitable, publicly available leaf‑segmentation datasets. Using them, we compare one‑stage, two‑stage and Transformer‑based detectors and identify a YOLO26 model configuration to provide the best trade‑off for real‑world precision‑agriculture tasks. Extensive cross‑domain generalization experiments reveal substantial performance drops across plant species and recording setups, especially for models trained solely on laboratory data. To strengthen data availability, we introduce a new benchmark dataset with leaf‑level masks for 23 plant species, created via semi‑automatic annotation of selected CropAndWeed images. A model trained on all four existing datasets achieves a mean mAP50‑95 of 83.9% across their corresponding test sets and 40.2% on our new benchmark, demonstrating improved generalization and highlighting the need for diverse leaf‑segmentation datasets in robust precision agriculture.
Authors:Faraz Kayani, Sarmad Kayani, Asad Ahmed, Radu Timofte, Dmitry Ignatov
Abstract:
While deep‑learning‑based image restoration has achieved unprecedented fidelity, deployment on mobile Neural Processing Units (NPUs) remains bottlenecked by operator incompatibility and memory‑access overhead. We propose an NPU‑aware hardware‑algorithm co‑design approach for real‑world image denoising on mobile NPUs. Our approach employs a high‑capacity teacher to supervise a lightweight student network specifically designed to leverage the tiled‑memory architectures of modern mobile SoCs. By prioritizing NPU‑native primitives ‑‑ standard 3x3 convolutions, ReLU activations, and nearest‑neighbor upsampling ‑‑ and employing a progressive context expansion strategy (up to 1024x1024 crops), the model achieves 37.66 dB PSNR / 0.9278 SSIM on the validation benchmark and 37.58 dB PSNR / 0.9098 SSIM on the held‑out test benchmark at full resolution (2432x3200) in the Mobile AI 2026 challenge. Following the official challenge rules, the inference runtime is measured under a standardized Full HD (1088x1920) protocol, where it runs in 34.0 ms on the MediaTek Dimensity 9500 and 46.1 ms on the Qualcomm Snapdragon 8 Elite NPU. We further reveal an "Inference Inversion" effect, where strict adherence to NPU‑compatible operations enables dedicated NPU execution up to 3.88x faster than the integrated mobile GPU. The 1.96M‑parameter student recovers 99.8% of the teacher's restoration quality via high‑alpha knowledge distillation (alpha = 0.9), achieving a 21.2x parameter reduction while closing the PSNR gap from 1.63 dB to only 0.05 dB. These results establish hardware‑aware distillation as an effective strategy for unifying high‑fidelity denoising with practical deployment across diverse mobile NPU architectures. The proposed lightweight student model (LiteDenoiseNet) and its training statistics are provided in the NN Dataset, available at https://github.com/ABrain‑One/NN‑Dataset.
Authors:Yazhe Wan, Changjae Oh
Abstract:
Open‑vocabulary object detection aims to recognize objects from an open set of categories, which leverages vision‑language models (VLMs) pre‑trained on large‑scale image‑text data. The cooperative paradigm combines an object detector with a VLM to achieve zero‑shot recognition of novel objects. However, VLMs pre‑trained on full images often struggle to capture local object details, limiting their effectiveness when applied to region‑level detection. We present Decoupled Adaptivity Training (DAT), a self‑supervised fine‑tuning approach to improve VLMs for cooperative model‑based object detection. Given a cooperative model consists of a closed‑set detector and a VLM, we first construct a region‑aware pseudo‑labeled dataset using a pre‑trained closed‑set object detector, in which regions corresponding to novel objects may be present but remain unlabeled or mislabeled. We then fine‑tune the visual backbone of the VLM in a decoupled manner, which enhances local feature alignment while preserving global semantic knowledge via weight interpolation. DAT is a plug‑and‑play module that requires no inference overhead and fine‑tunes less than 0.8M parameters. Experiments on the COCO and LVIS datasets show that DAT consistently improves detection performance on both novel and known categories, establishing a new state of the art in cooperative open‑vocabulary detection.
Authors:Zhuoyue Zhang, Jihua Zhu, Chaowei Fang, Jian Liu, Ajmal Saeed Mian
Abstract:
Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground‑truth tube centers as decoder positional embeddings, causing spatio‑temporal positional leakage. Moreover, they supervise inter‑frame motion with deterministic proxy targets that systematically discard distributional structure by collapsing multimodal trajectory uncertainty into conditional means. To address these limitations, we propose Diffusion Masked Pretraining (DiMP), a unified self‑supervised framework for dynamic point clouds. DiMP introduces diffusion modeling into both positional inference and motion learning. It first applies forward diffusion noise only to masked tube centers, then predicts clean centers from visible spatio‑temporal context. This removes positional leakage while preserving visible coordinates as clean temporal anchors. DiMP also reformulates point‑wise inter‑frame displacement supervision as a DDPM noise‑prediction objective conditioned on decoded representations. This design drives the encoder to target the full conditional distribution of plausible motions under a variational surrogate, rather than collapsing to a single deterministic estimate. Extensive experiments demonstrate that DiMP consistently improves downstream accuracy over the backbone alone, with absolute gains of 11.21% on offline action segmentation and 13.65% under causally constrained online inference.Codes are available at https://github.com/InitalZ/DiMP.git.
Authors:Lorenzo Beltrame, Jules Salzinger, Filip Svoboda, Phillipp Fanta-Jende, Jasmin Lampert, Radu Timofte, Marco Körner
Abstract:
Shadows cast by terrain and tall structures remain a major obstacle for high‑resolution satellite image analysis, degrading classification, detection, and 3D reconstruction performance. Public resources offering geometry‑consistent paired shadow/shadow‑free satellite imagery are essentially missing, and most Earth‑observation datasets are designed for shadow detection or 3D modelling rather than removal. Existing deep shadow‑removal datasets either target ground‑level or aerial scenes or rely on unpaired and weakly supervised formulations rather than explicit satellite pairs. We address this gap with deSEO, a geometry‑aware and physics‑informed methodology that, to the best of our knowledge, is the first to derive paired supervision for satellite shadow removal from the S‑EO shadow detection dataset through a fully replicable pipeline. For each tile, deSEO selects a minimally shadowed acquisition as a weak reference and pairs it with shadowed counterparts using temporal and geometric filtering, Jacobian‑based orientation normalisation, and LoFTR‑RANSAC registration. A per‑pixel validity mask restricts learning to reliably aligned regions, enabling supervision despite residual off‑nadir parallax. In addition to this paired dataset, we develop a DSM‑aware deshadowing model that combines residual translation, perceptual objectives, and mask‑constrained adversarial learning. In contrast, a direct adaptation of a UAV‑based SRNet/pix2pix architecture fails to converge under satellite viewpoint variability. Our model consistently reduces the visual impact of cast shadows across diverse illumination and viewing conditions, achieving improved structural and perceptual fidelity on held‑out scenes. deSEO therefore provides the first reproducible, geometry‑aware paired dataset and baseline for shadow removal in satellite Earth observation.
Authors:Carlijn Lems, Sander Moonemans, Natálie Klubíčková, Biagio Brattoli, Taebum Lee, Seokhwi Kim, Veronica Vilaplana, Laura Pons, Sapir Hochman, Mauricio Eduardo Suárez-Franck, Pedro Luis Fernandez, Julius Drachneris, Donatas Petroska, Renaldas Augulis, Arvydas Laurinavicius, Domingos Oliveira, Diana Montezuma, Anouk B. Bouwmeester, Dominique van Midden, Anne-Marie Vos, Shoko Vos, Jolique van Ipenburg, Maschenka Balkenhol, Koen Winkler, Iris Nagtegaal, Konnie Hebeda, Uta Flucke, Katrien Grünberg, Josef Skopal, Brinder S. Chohan, Jordi Temprana-Salvador, Enrico Munari, Luca Cima, Giulia Querzoli, Yosamin Gonzalez Belisario, Jaeike W. Faber, Geert J. L. H. van Leenders, Jan H. von der Thüsen, Lodewijk A. A. Brosens, Ronald R. de Krijger, Pieter Wesseling, Sandrine Florquin, Mateusz Maniewski, Adam Kowalewski, Robert Barna, Dina Tiniakos, Joan Lop Gros, Rogier Donders, Jake S. F. Maurits, Ming Yang Lu, Chengkuan Chen, Faisal Mahmood, Jeroen van der Laak, Nadieh Khalili, Frédérique Meeuwsen, Francesco Ciompi
Abstract:
Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent benchmarking to assess its potential in assisting pathologists in routine diagnostics. We created DALPHIN, the first multicentric open benchmark for pathology AI copilots, comprising 1236 images from 300 cases, spanning 130 rare to common diagnoses, 6 countries, and 14 subspecialties. The DALPHIN design and dataset are introduced alongside a human performance benchmark of 31 pathologists from 10 countries with varying expertise. We report results for two general‑purpose (GPT‑5, Gemini 2.5 Pro) and one pathology‑specific copilot (PathChat+) for sequential and independent answer generation. We observed no statistically significant difference from expert‑level performance in four of six tasks for PathChat, 2/6 tasks for Gemini, and 1/6 tasks for GPT. DALPHIN is publicly released with sequestered, indirectly accessible ground truth to foster robust and enduring benchmarking. Data, methods, and the evaluation platform are accessible through dalphin.grand‑challenge.org.
Authors:Remi Chierchia, Léo Lebrat, David Ahmedt-Aristizabal, Olivier Salvado, Clinton Fookes, Rodrigo Santa Cruz
Abstract:
Neural Surface Reconstruction has become a standard methodology for indoor 3D reconstruction, with Signed Distance Functions (SDFs) proving particularly effective for representing scene geometry. A variety of applications require a detailed understanding of the scene context, driving the need for object‑level semantic signals. While recent methods successfully integrate semantic labels, they often inherit the slow training time and limited scalability of multi‑SDF learning. In this paper, we introduce FSTM, a unified approach for learning geometry and semantics through a two‑step process: a geometry warm‑up using RGB inputs and geometric cues, followed by semantic field estimation. By first optimising geometry without semantic supervision, we observe substantial improvements compared to the standard joint optimisation. Rather than relying on specialised modules or complex multi‑SDF designs, FSTM shows that a streamlined formulation is sufficient to achieve strong geometric and semantic reconstructions. Experiments on both synthetic and real‑world indoor datasets show that our method outperforms multi‑SDF approaches. It trains 2.3x faster on Replica, improves robustness to real‑world imperfections on ScanNet++, and achieves higher recall by recovering the surfaces of more objects in the scene. The code will be made available at https://remichierchia.github.io/FSTM.
Authors:Zihao Guo, Jihua Zhu, Jian Liu, Ajmal Saeed Mian
Abstract:
Pre‑trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine‑tuning these models is computationally expensive and storage‑intensive. Parameter‑efficient fine‑tuning (PEFT) offers a promising alternative, but existing PEFT approaches are primarily designed for Transformer‑based backbones and rely on token‑level prompting or feature transformation. Mamba‑based backbones introduce a granularity mismatch between token‑level adaptation and state‑level sequence dynamics. Consequently, straightforward transfer of existing PEFT approaches to frozen Mamba backbones leads to substantial accuracy degradation and unstable optimization. To address this issue, we propose Mantis, the first Mamba‑native PEFT framework for 3D PFMs. Specifically, a State‑Aware Adapter (SAA) is introduced to inject lightweight task‑conditioned control signals into selective state‑space updates, enabling state‑level adaptation while keeping the pre‑trained backbone frozen. Moreover, different valid point cloud serializations are regularized by Dual‑Serialization Consistency Distillation (DSCD), thereby reducing serialization‑induced instability. Extensive experiments across multiple benchmarks demonstrate that our Mantis achieves competitive performance with only about 5% trainable parameters. Our code is available at https://github.com/gzhhhhhhh/Mantis.
Authors:Alexander Matyasko, Xin Lou, Indriyati Atmosukarto, Wei Zhang
Abstract:
Attacking semantic segmentation models is significantly harder than image classification models because an attacker must flip thousands of pixel predictions simultaneously. Standard pixel‑wise cross‑entropy (CE) is ill‑suited to this setting: it tends to overemphasize already‑misclassified pixels, which slows optimization and overstates model robustness. To address these issues, we introduce TsallisPGD, an adversarial attack built on the Tsallis cross‑entropy, a generalization of CE parameterized by q, which adaptively reshapes the gradient landscape by controlling gradient concentration across pixels. By varying q, we steer the attack toward pixels at different confidence levels. We first show that no single fixed‑q is universally optimal, as its effectiveness depends on the dataset, model architecture, and perturbation budget. Motivated by this, we propose a dynamic q‑schedule that sweeps q during optimization. Extensive experiments on Cityscapes, Pascal VOC, and ADE20K show that TsallisPGD, using a single validation‑selected schedule, achieves the best average attack rank across all evaluated settings and improves over CEPGD, SegPGD, CosPGD, JSPGD, and MaskedPGD in reducing accuracy and mIoU on both standard and robust models.
Authors:Siyou Lin, Zhou Xue, Hongwen Zhang, Liang An, Dongping Li, Shaohui Jiao, Yebin Liu
Abstract:
Recent trends in sparse‑view 3D reconstruction have taken two different paths: feed‑forward reconstruction that predicts pixel‑aligned point maps without a complete geometry, and generative 3D reconstruction that generates complete geometry but often with poor input‑alignment. We present Mix3R, a novel generative 3D reconstruction method which mixes feed‑forward reconstruction and 3D generation into a single framework in an aligned manner. Mix3R generates a 3D shape in two stages: a sparse voxel generation stage and a textured geometry generation stage. Unlike pure generative methods, our first‑stage generation jointly produces a coarse 3D structure (sparse voxels), per‑view point maps and camera parameters aligned to that 3D structure. This is made possible by introducing a Mixture‑of‑Transformers architecture that inserts global self‑attentions to a feed‑forward reconstruction model and a 3D generative model, both pretrained on large‑scale data. This design effectively retains the pretrained priors but enables better 2D‑3D alignment. Based on the initial aligned generations of sparse 3D voxels and point maps, we compute an overlap‑based attention bias that is directly added to another pretrained textured geometry generation model, enabling it to correctly place input textures onto generated shapes in a training‑free manner. Our design brings mutual benefits to both feed‑forward reconstruction and 3D generation: The feed‑forward branch learns to ground its predictions to a generative 3D prior, and conversely, the 3D generation branch is conditioned on geometrically informative features from the feed‑forward branch. As a result, our method produces 3D shapes with better input alignment compared with pure 3D generative methods, together with camera pose estimations more accurate than previous feed‑forward reconstruction methods. Our project page is at https://jsnln.github.io/mix3r/
Authors:Lina Zhang, Tonmoy Monsoor, Mehmet Efe Lorasdagi, Prateik Sinha, Chong Han, Peizheng Li, Yuan Wang, Jessica Pasqua, Colin McCrimmon, Rajarshi Mazumder, Vwani Roychowdhury
Abstract:
Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzing clinically significant involuntary movements in neurological disorders remains largely unexplored. This pilot study evaluates the capability of MLLMs for automated recognition of pathological movements in seizure videos. We assessed the zero‑shot performance of state‑of‑the‑art MLLMs on 20 ILAE‑defined semiological features across 90 clinical seizure recordings. MLLMs outperformed fine‑tuned Convolutional Neural Network (CNN) and Vision Transformer (ViT) baseline models on 13 of 18 features without task‑specific training, demonstrating particular strength in recognizing salient postural and contextual features while struggling with subtle, high‑frequency movements. Feature‑targeted signal enhancement (facial cropping, pose estimation, audio denoising) improved performance on 10 of 20 features. Expert evaluation showed that 94.3 percent of MLLM‑generated explanations for correctly predicted cases achieved at least 60 percent faithfulness scores, aligning with epileptologist reasoning. These findings demonstrate the potential of adapting general‑purpose MLLMs for specialized clinical video analysis through targeted preprocessing strategies, offering a path toward interpretable, efficient diagnostic assistance. Our code is publicly available at https://github.com/LinaZhangUCLA/PathMotionMLLM.
Authors:JF Bastien, Sam D'Amico
Abstract:
Video vision‑language models (VLMs) keep paying for visual state the stream already told us was stable. The factory wall did not move, but most VLM pipelines still hand the model dense RGB frames or a fresh prefix again. We study that waste as training‑free anti‑recomputation: reuse state when validation says it survives, and buy fresh evidence when the scene, query, or cache topology requires it.
The largest measured win is after ingest. On frozen Qwen2.5‑VL‑7B‑Instruct‑4bit, adaptive same‑video follow‑up reuse preserves paired choices and correctness on a 93‑query VideoMME breadth setting while reducing follow‑up latency by 14.90‑35.92x. The first query is still cold; the win starts when later questions reuse the same video state. Stress tests bound the result: repeated‑question schedules hold through 50 turns, while dense‑answer‑anchored prompt variation separates conservative fixed K=1 repair from faster aggressive policies that drift.
Fresh‑video pruning is smaller but real. C‑VISION skips timed vision‑tower work before the first answer is generated. On Gemma 4‑E4B‑4bit, the clean 32f short cell reaches 1.316x first‑query speedup with no paired drift or parse failures on 20 items; Qwen shows the fidelity/speed boundary.
Stage‑share ceiling (C‑CEILING) is the accounting guardrail: a component speedup becomes an end‑to‑end speedup only in proportion to the wall‑clock share it accelerates, so C‑VISION and after‑ingest follow‑up reuse do not multiply. Candidate C‑STREAM remains a native‑rate target, not a headline result here. The broader direction is VLM‑native media that expose change, motion, uncertainty, object state, sensor time, and active tiles directly, so models do not have to rediscover the world from dense RGB every frame.
Authors:Abderrahmene Boudiaf, Sajd Javed
Abstract:
High‑throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a "phenotyping bottleneck," where manual data collection is labor‑intensive and prone to observer bias. Conventional closed‑set computer vision systems fail to address this challenge, as they require extensive species‑specific annotation and lack the flexibility to handle diverse breeding populations. To bridge this gap, we present CropVLM, a Vision‑Language Model (VLM) adapted for the agricultural domain via Domain‑Specific Semantic Alignment (DSSA). Trained on 52,987 manually selected image‑caption pairs covering 37 species in natural field conditions, CropVLM effectively maps agronomic terminology to fine‑grained visual features. We further introduce the Hybrid Open‑Set Localization Network (HOS‑Net), an architecture that integrates CropVLM to enable the detection of novel crops solely from natural language descriptions without retraining. By eliminating the reliance on species‑specific training data, CropVLM provides a scalable solution for high‑throughput phenotyping, accelerating genetic gain and facilitating large‑scale biodiversity research essential for sustainable agriculture. The trained model weights and complete pipeline implementation are publicly available at: [https://github.com/boudiafA/CropVLM](https://github.com/boudiafA/CropVLM). In comprehensive evaluations, CropVLM achieves 72.51% zero‑shot classification accuracy, outperforming seven CLIP‑style baselines. Our detection pipeline demonstrates superior zero‑shot generalization to novel species, achieving 49.17 AP50 on our CVTCropDet benchmark and 50.73 AP50 on tropical fruit species, compared to 34.89 and 48.58 for the next‑best method, respectively.
Authors:Seunghyun Ji
Abstract:
LoRA fine‑tuning of diffusion transformers (DiT) on multi‑style data suffers from \emphstyle bleed: a single low‑rank residual cannot represent several distinct artist fingerprints, and the optimizer converges to their average. Mixture‑of‑experts LoRA in the HydraLoRA style replaces the up‑projection with E heads under a router, but when every expert is zero‑initialized the router receives identical gradient from each head and remains at the uniform prior. The experts then evolve permutation‑symmetrically, and the network trains as a single rank‑r LoRA at E× the cost. We present Ortho‑Hydra, a re‑parameterisation that combines an OFT‑style Cayley‑orthogonal shared basis with per‑expert \emphdisjoint output subspaces carved from the top‑(Er) left singular vectors of the pretrained weight. Disjointness makes the router's per‑expert score non‑degenerate at step~0, so specialization receives gradient signal before any expert has trained. We test the predicted deadlock on a DiT pipeline by comparing two HydraLoRA baselines, a zero‑initialized shared‑basis variant and the original σ=0.1 Gaussian‑jitter mitigation, against Ortho‑Hydra under a matched optimiser, dataset, and step budget. Neither baseline leaves the uniform prior within the first 1\textk steps; Ortho‑Hydra begins de‑uniformising within the first few hundred. End‑task generation quality on multi‑style data is out of scope; we report the construction, the cold‑start mechanism, and the routing dynamics it changes. Code: https://github.com/sorryhyun/anima_lora.
Authors:Ryan Faulkenberry, Saurabh Prasad
Abstract:
The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery well without supervised fine‑tuning are valuable, but existing solutions fall behind supervised methods. Recently, DINOv3 surpassed SOTA RS foundation models on the GEO‑bench segmentation benchmark without pre‑training on RS data. Additionally, DINO.txt has enabled open vocabulary semantic segmentation (OVSS) with the DINOv3 backbone. We leverage these developments to form an OVSS model for RS imagery, free of RS‑domain fine‑tuning. Our model, CAFe‑DINO (Cost Aggregation + Feature Upsampling with DINO) exploits the strong OVSS performance of DINOv3 for RS imagery via cost aggregation and training‑free upsampling of text‑image similarity scores. The robust latent of the DINOv3 backbone eliminates the need for fine‑tuning on RS imagery; we instead fine‑tune our model on a RS‑targeted subset of COCO‑Stuff. CAFe‑DINO achieves state‑of‑the‑art performance on key RS segmentation datasets, outperforming OVSS methods fine‑tuned on RS data. Our code and data are publicly available at https://github.com/rfaulk/DINO_Soars.
Authors:Jonas V. Funk
Abstract:
Reliable wildfire spread prediction is vital for risk‑aware emergency planning, yet most deep learning models lack principled uncertainty quantification (UQ). Further, for boundary‑sensitive cases like wildfire spread, evaluating models with global metrics alone is often insufficient. To shift the focus of UQ evaluation toward a more operationally relevant approach, the Fire‑Centered Evaluation Region (FCER) framework is introduced as a spatially conditioned protocol to characterize UQ within critical fire zones. Using FCER, an Ensemble is compared against an distilled single‑pass student model on the WildfireSpreadTS dataset. The student model demonstrates comparable calibration and complementary uncertainty ranking in boundary‑relevant regimes. Code is available at https://github.com/jonasvilhofunk/WildfireUQ‑FCER
Authors:Amirreza Mahbod, Ramona Woitek, Jeanne Shen
Abstract:
In computational pathology, nuclear instance segmentation is a fundamental task with many downstream clinical applications. With the advent of deep learning, many approaches, including convolutional neural networks (CNNs) and vision transformers (ViTs), have been proposed for this task, along with both machine learning‑based and non‑machine learning‑based pre‑ and post‑processing techniques to further boost performance. However, one fundamental aspect that has received less attention is the evaluation pipeline. In this study, we identify four key issues associated with nuclear instance segmentation evaluation and propose corresponding solutions. Our proposed modifications, namely handling vague regions, score normalization, overlapping instances, and border uncertainty, are integrated into a unified framework called NucEval, which enables robust evaluation of nuclear instance segmentation. We evaluate this pipeline using the NuInsSeg dataset, which provides unique characteristics that make it particularly suitable for this study, as well as two additional external datasets, with three CNN‑ and ViT‑based nuclear instance segmentation models, to demonstrate the impact of these modifications on instance segmentation metrics. The code, along with complete guidelines and illustrative examples, is publicly available at: https://github.com/masih4/nuc_eval.
Authors:Bumjun Kim, Albert No
Abstract:
Understanding how textual embeddings contribute to memorization in text‑to‑image diffusion models is crucial for both interpretability and safety. This paper investigates an unexpected behavior of CLIP embeddings in Stable Diffusion, revealing that the model disproportionately relies on specific embeddings. We categorize input tokens as <startoftext>, <prompt>, <endoftext> and <pad> with corresponding embeddings \mathbfv^\mathbfsot, \mathbfv^\mathbfpr, \mathbfv^\mathbfeot, \mathbfv^\mathbfpad. We discover that \mathbfv^\mathbfpr contribute minimally to generation in memorized cases. In contrast, \mathbfv^\mathbfpad strongly affect memorization due to their structural duplication of \mathbfv^\mathbfeot, the only embedding explicitly optimized during CLIP training. This duplication unintentionally amplifies the influence of \mathbfv^\mathbfeot, causing the model to over‑rely on it, thereby driving memorization. Based on these observations, we propose two simple yet effective inference‑time mitigation strategies: (1) Replacing the tokenizer's default <pad> from <eot> to the ! token before embedding, and masking the \mathbfv^\mathbfeot; (2) Partial masking of \mathbfv^\mathbfpad. Both suppress memorization without degrading quality, and are readily deployable without prior detection.
Authors:Xiao Li, Xiang Zheng, Yifeng Gao, Xinyu Xia, Yixu Wang, Xin Wang, Ye Sun, Yunhan Zhao, Ming Wen, Jiayu Li, Zixing Chen, Xun Gong, Yi Liu, Yige Li, Yutao Wu, Cong Wang, Jun Sun, Yixin Cao, Zhineng Chen, Jingjing Chen, Tao Gui, Qi Zhang, Zuxuan Wu, Xipeng Qiu, Xuanjing Huang, Tiehua Zhang, Zhipeng Wei, Kun Wang, Xinfeng Li, Hanxun Huang, Sarah Erfani, James Bailey, Jianping Wang, Chaowei Xiao, Ran He, Bo Li, Xingjun Ma, Yu-Gang Jiang
Abstract:
Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open‑world, safety‑critical environments. As these systems gain autonomy and enter domains such as transportation, healthcare, and industrial or assistive robotics, ensuring their safety becomes both technically challenging and socially indispensable. Unlike digital AI systems, embodied agents must act under uncertain sensing, incomplete knowledge, and dynamic human‑robot interactions, where failures can directly lead to physical harm. This survey provides a comprehensive and structured review of safety research in embodied AI, examining attacks and defenses across the full embodied pipeline, from perception and cognition to planning, action and interaction, and agentic system. We introduce a multi‑level taxonomy that unifies fragmented lines of work and connects embodied‑specific safety findings with broader advances in vision, language, and multimodal foundation models. Our review synthesizes insights from over 500 papers spanning adversarial, backdoor, jailbreak, and hardware‑level attacks; attack detection, safe training and robust inference; and risk‑aware human‑agent interaction. This analysis reveals several overlooked challenges, including the fragility of multimodal perception fusion, the instability of planning under jailbreak attacks, and the trustworthiness of human‑agent interaction in open‑ended scenarios. By organizing the field into a coherent framework and identifying critical research gaps, this survey provides a roadmap for building embodied agents that are not only capable and autonomous but also safe, robust, and reliable in real‑world deployment.
Authors:Yu-Ju Tsai, Brian Price, Qing Liu, Luis Figueroa, Daniil Pakhomov, Zhihong Ding, Scott Cohen, Ming-Hsuan Yang
Abstract:
Personalized image completion aims to restore occluded regions in personal photos while preserving identity and appearance. Existing methods either rely on generic inpainting models that often fail to maintain identity consistency, or assume that suitable reference images are explicitly provided. In practice, suitable references are often not explicitly provided, requiring the system to search for identity‑consistent images within personal photo collections. We present AlbumFill, a training‑free framework that retrieves identity‑consistent references from personal albums for personalized completion. Given an occluded image and a personal album, a vision‑language model infers missing semantic cues to guide composed image retrieval, and the retrieved references are used by reference‑based completion models. To facilitate this task, we introduce a dataset containing 54K human‑centric samples with associated album images. Experiments across multiple baselines demonstrate the difficulty of personalized completion and highlight the importance of identity‑consistent reference retrieval. Project Page: https://liagm.github.io/AlbumFill/
Authors:Yeheng Zong, Pou-Chun Kung, Yike Pan, Seth Isaacson, Yizhou Chen, Ram Vasudevan, Katherine A. Skinner
Abstract:
Accurately recovering human pose and appearance from video is an essential component of scene reconstruction, with applications to motion capture, motion prediction, virtual reality, and digital twinning. Despite significant interest in building realistic human avatars from video, this paper demonstrates that existing methods do not accurately recover the 3D geometry of humans. ViT‑based approaches are not consistently reliable and can overfit to 2D views, while NeRF‑ and Gaussian Splatting‑based avatars treat pose and appearance separately, limiting rendering generalization to new poses. To resolve these shortcomings, this paper proposes HumanSplatHMR, a joint optimization framework that refines 3D human poses while simultaneously learning a high‑fidelity avatar for novel‑view and novel‑pose synthesis. Our key insight is to close the loop between geometric pose estimation and differentiable rendering. Unlike prior human avatar methods that rely on accurate human pose obtained through motion capture systems or offline refinement, which are impractical in in‑the‑wild scenarios, our approach uses only human mesh estimates from a state‑of‑the‑art human pose estimator to better reflect real‑world conditions. Therefore, instead of using the human pose only as a deformation prior, HumanSplatHMR backpropagates photometric, segmentation, and depth losses through a differentiable renderer to the pose parameters and global position. This coupling refines the global 3D pose over time, improving accuracy and alignment while producing better renderings from novel views. Experiments show consistent improvements over pose recovery baselines that omit image‑level refinement and avatar baselines that decouple pose estimation from avatar reconstruction.
Authors:Yining Li, Dongchen Han, Zeyu Liu, Hanyi Wang, Yulin Wang, Gao Huang
Abstract:
While linear‑complexity attention mechanisms offer a promising alternative to Softmax attention for overcoming the quadratic bottleneck, training such models from scratch remains prohibitively expensive. Inheriting weights from pretrained Transformers provides an appealing shortcut, yet the fundamental representational gap between Softmax and linear attention prevents effective weight transfer. In this work, we address this conversion challenge from two perspectives: architectural alignment and representational alignment. We identify Test‑Time Training (TTT) as a linear‑complexity architecture whose two‑layer dynamic formulation is structurally aligned with Softmax attention, enabling direct inheritance of pretrained attention weights. To further align representational properties, including key shift‑invariance and locality, we introduce key instance normalization and a lightweight locality enhancement module. We validate our approach by linearizing Stable Diffusion 3.5 and introduce SD3.5‑T^5 (Transformer To Test Time Training). With only 1 hour of fine‑tuning on 4×H20 GPUs, SD3.5‑T^5 achieves comparable text‑to‑image quality to the fine‑tuned Softmax model, while accelerating inference by 1.32× and 1.47× at 1K and 2K resolutions. Code is available at https://github.com/LeapLabTHU/Transformer‑to‑TTT.
Authors:Danil Tokhchukov, Veronika Morozova, Gonzalo Ferrer
Abstract:
Traditional Simultaneous Localization and Mapping (SLAM) algorithms rely heavily on the static environment assumption, which severely limits their applicability in real‑world spaces populated by moving entities, such as pedestrians. In this work, we propose DynoSLAM, a tightly‑coupled Dynamic GraphSLAM architecture that integrates socially‑aware Graph Neural Networks (GNNs) directly into the factor graph optimization. Unlike conventional approaches that use rigid constant‑velocity heuristics or deterministic single‑agent neural priors, our framework formulates pedestrian motion forecasting as a stochastic World Model. By utilizing Monte Carlo rollouts from a trained GNN, we capture the multimodal epistemic uncertainty of human interactions and embed it into the SLAM graph via a dynamic Mahalanobis distance factor. We demonstrate through extensive simulated experiments that this stochastic formulation not only maintains highly accurate retrospective tracking but also prevents the optimization failures caused by the deterministic "argmax problem". Ultimately, extracting the empirical mean and covariance matrices of future pedestrian states provides a mathematically rigorous, probabilistic safety envelope for downstream local planners, enabling anticipatory and collision‑free robot navigation in densely crowded environments.
Authors:Chenyu Hui, Xiaodi Huang, Siyu Xu, Yunke Wang, Shan You, Fei Wang, Tao Huang, Chang Xu
Abstract:
Vision‑language‑action (VLA) models typically rely on large‑scale real‑world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain gap and limited environmental diversity, resulting in weak real‑world generalization. We present an efficient video augmentation framework that converts simulated VLA videos into realistic training videos while preserving task semantics and action trajectories. Our pipeline extracts structured conditions from simulation via video semantic segmentation and video captioning, rewrites captions to diversify environments, and uses a conditional video transfer model to synthesize realistic videos. To make augmentation practical at scale, we introduce a diffusion feature‑reuse mechanism that reuses video tokens across adjacent timesteps to accelerate generation, and a coreset sampling strategy that identifies a compact, non‑redundant subset for augmentation under limited computation. Extensive experiments on Robotwin 2.0, LIBERO, LIBERO‑Plus, and a real robotic platform demonstrate consistent improvements. For example, our method improves RDT‑1B by 8% on Robotwin 2.0, and boosts π_0 by 5.1% on the more challenging LIBERO‑Plus benchmark. Code is available at: https://github.com/nanfangxiansheng/Seeing‑Realism‑from‑Simulation.
Authors:Giacomo Pacini, Luca Ciampi, Nicola Messina, Nicola Tonellotto, Giuseppe Amato, Fabrizio Falchi
Abstract:
Open‑world text‑guided class‑agnostic counting (CAC) has emerged as a flexible paradigm for counting arbitrary object classes by using natural language prompts. However, current evaluation protocols primarily focus on standard counting errors within single‑category images, overlooking a fundamental requirement: the ability to correctly ground the textual prompt in the visual scene. In this paper, we show that several state‑of‑the‑art CAC models often struggle to determine which object class should be counted based on the given prompt, revealing a misalignment between textual semantics and visual object representations. This limitation leads to spurious counting responses and reduced reliability in real‑world scenarios. To systematically address these limitations, we propose a new evaluation framework focused on model robustness and trustworthiness. Our contribution is two‑fold: (i) we introduce PrACo++ (Prompt‑Aware Counting++), a novel test suite featuring two dedicated evaluation protocols ‑‑ the negative‑label test and the distractor test ‑‑ paired with new specialized metrics; and (ii) we present the MUCCA (MUlti‑Category Class‑Agnostic counting) evaluation dataset, a new collection of real‑world images featuring multiple annotated object categories per scene, unlike existing CAC benchmarks that typically include a single category per image. Our extensive experimental evaluation of 10 state‑of‑the‑art methods shows that, despite strong performance under standard counting metrics, current models exhibit significant weaknesses in understanding and grounding object class descriptions. Finally, we provide a quantitative analysis of how semantic similarity between prompts influences these failures. Overall, our results underscore the need for more semantically grounded architectures and offer a reliable framework for future assessment in open‑world text‑guided CAC methods.
Authors:Ye Zhang, Longguang Wang, Qing Gao, Chaocan Xiang, Mohammed Bennamoun, Yulan Guo
Abstract:
The field of sensor‑based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identify daily activities. Despite the advancements in learning‑based methods, it is challenging to perform information fusion from the temporal perspective due to the complexities in fusing heterogeneous sensor data and establishing long‑term context correlations. This paper proposes a novel triple spectral fusion framework tailored for HAR. First, we develop an adaptive complementary filtering technique for noise suppression and organize each IMU's sensors into posture and motion modality nodes. Given that IMU nodes form a dynamic heterogeneous graph, we then apply adaptive filtering within the graph Fourier domain to merge both homogeneous and heterogeneous node information. Furthermore, an adaptive wavelet frequency selection approach is implemented to suppress context redundancy and shorten the length of features. This approach enhances both timestamp‑based graph aggregation and the correlation of long‑term contexts. Our framework uses adaptive filtering in the Fourier, graph Fourier, and wavelet domains, enabling effective multi‑sensor fusion and context correlation. Extensive experiments on ten benchmark datasets demonstrate the superior performance of our framework. Project page: https://github.com/crocodilegogogo/TSF‑TPAMI2026.
Authors:Romain Valabregue, Ines Khemir, Eric Badinet, François Rousseau, Guillaume Auzias, Reuben Dorent
Abstract:
Synthetic training has recently advanced brain MRI segmentation by enabling contrast‑agnostic models trained entirely on generated data. However, most existing approaches rely on hundreds of automatically labeled templates, introducing systematic biases and limiting their flexibility to incorporate new anatomical structures. We present the Segment It All Model (SIAM), a 3D whole‑head segmentation framework for 16 anatomical structures, trained using only six high‑quality, manually annotated templates. SIAM extends domain randomization to both intensity and shape domains: synthetic image generation ensures contrast variability, while high‑resolution spatial transformations model anatomical differences in cortical thickness and deep nuclei morphology. Unlike prior synthetic models, SIAM simultaneously segments brain as well as extra‑cerebral tissues, including cerebrospinal fluid, vessels, dura mater, skull, and skin, enabling fully automated, preprocessing‑free analysis. Evaluation across eight heterogeneous datasets (N=301), that include multiple contrasts (T1‑weighted, T2‑weighted, CT) and span a wide range of ages, demonstrates that SIAM matches or outperforms state‑of‑the‑art methods for brain structures, in addition to extending automated segmentation to non‑brain structures. The model also exhibits superior consistency across contrasts and repeated acquisitions, together with improved sensitivity to subtle gray matter atrophy. We openly release the model and the label templates at https://github.com/romainVala/SIAM.
Authors:Verena Jasmin Hallitschke, Carsten Eickhoff, Philipp Berens
Abstract:
Vision‑language models hold considerable promise for ophthalmology, but their development depends on large‑scale, high‑quality image‑text datasets that remain scarce. We present PubMed‑Ophtha, a hierarchical dataset of 102,023 ophthalmological image‑caption pairs extracted from 15,842 open‑access articles in PubMed Central. Unlike existing datasets, figures are extracted directly from article PDFs at full resolution and decomposed into their constituent panels, panel identifiers, and individual images. Each image is annotated with its imaging modality ‑‑ color fundus photography, optical coherence tomography, retinal imaging, or other ‑‑ and a mark status indicating the presence of annotation marks such as arrows. Figure captions are split into panel‑level subcaptions using a two‑step LLM approach, achieving a mean average sentence BLEU score of 0.913 on human‑annotated data. Panel and image detection models reach a mAP@0.50 of 0.909 and 0.892, respectively, and figure extraction achieves a median IoU of 0.997. To support reproducibility, we additionally release the human‑annotated ground‑truth data, all trained models, and the full dataset generation pipeline.
Authors:Guangrui Bai, Yifan Mei, Yahui Deng, Yuhan Chen, Yuze Qiu, Wenhai Liu, Erbao Dong
Abstract:
Explicit reconstruction constraints derived from the decoupled representation are further imposed to suppress abnormal channel amplification and chromatic noise. Experiments on LOLv2‑Real, MIT‑Adobe FiveK, and LSRW show that the proposed method achieves competitive or superior quantitative and visual performance, reaching 29.71 dB PSNR and 0.89 SSIM on LOLv2‑Real. DarkFace experiments further indicate improved downstream face detection under low‑light conditions. Code and pretrained models are available at: https://github.com/mubaisam/ICD.
Authors:Yiming Ding, Siyu Cao, Luyuan Jiao, Yixuan Li, Zitong Wang, Zhiyong Liu, Lu Zhang
Abstract:
Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always hold in real‑world scenarios, where queries may correspond to multiple or no moments. Thus, we formulate Generalized Moment Retrieval (GMR), a unified setting that requires retrieving the complete set of relevant moments or predicting an empty set. To enable systematic study of GMR, we introduce Soccer‑GMR, a large‑scale benchmark built on challenging soccer videos that reflect general GMR scenarios, with realistic negative and positive queries. The benchmark is constructed via a duration‑flexible semi‑automated pipeline with human verification, enabling scalable data generation while maintaining high annotation quality. We further design a unified evaluation protocol with complementary metrics tailored for null‑set rejection, positive‑query localization, and end‑to‑end GMR performance. Finally, we establish strong baselines across two modeling paradigms: a lightweight plug‑and‑play GMR adapter for discriminative VMR models, and a GMR‑tailored GRPO reward for fine‑tuning multimodal large language models (MLLMs). Extensive experiments show consistent gains across all metrics and expose key limitations of current methods, positioning GMR as a more realistic and challenging benchmark for video‑language understanding.
Authors:Jintao Guo, Lin Wang, Shumeng Li, Jian Zhang, Yulin Zhou, Luyang Cao, Hairong Zheng, Yinghuan Shi
Abstract:
Existing cross‑subject fMRI decoding methods typically train a model on multiple scanned subjects and then adapt it to a new subject using substantial paired fMRI‑image data. However, in realistic scenarios, new‑subject fMRI data are often limited due to costly data acquisition, and raw data from previous subjects may be inaccessible, leading existing methods to suffer performance degradation during new‑subject adaptation. In this paper, we identify that this degradation stems from two key issues: brain‑side instability caused by large subject differences in fMRI responses, and image‑side supervision unreliability caused by fine‑grained visual details that are not reliably supported by limited fMRI signals. To address these challenges, we propose StableMind, a regularized adaptation framework designed to improve brain‑side representation stability and image‑side supervision reliability. (1) To stabilize brain representations, StableMind reuses ridge projections from the pretrained model as adaptation priors to constrain limited‑data new‑subject adaptation, and applies Fourier‑based feature‑level brain augmentation to improve robustness to individual variability. (2) To improve image supervision reliability, StableMind introduces difficulty‑aware image blur for brain‑image alignment, reducing the influence of fine‑grained visual details that are weakly supported by limited fMRI signals while preserving stable visual structure. Experiments on the Natural Scenes Dataset under a unified 1‑hour adaptation protocol demonstrate that StableMind achieves 84.02% image retrieval accuracy and 81.66% brain retrieval accuracy averaged over four subjects, surpassing the state‑of‑the‑art method by 5.71% brain retrieval accuracy with fewer trainable adaptation parameters. Our code is available at https://github.com/lingeringlight/StableMind.
Authors:Desong Yang, Mang Ye
Abstract:
With recent advancements in large‑scale pre‑trained text‑to‑image (T2I) models, training‑free image editing methods have demonstrated remarkable success. Typically, these methods involve adding noise to a clean image via an inversion process, followed by separate denoising steps for the reconstruction and editing paths during the forward process. However, since the reconstruction path is approximated using noisy latents from mismatched timesteps, existing methods inevitably suffer from accumulated drift, which fundamentally limits reconstruction fidelity. To address this challenge, we systematically analyze the inversion process within the flow transformer and propose DirectEdit, a simple yet effective editing method that eliminates the inherent reconstruction error without introducing additional neural function evaluations (NFEs). Unlike most prior works that attempt to rectify the inversion path, DirectEdit focuses on directly aligning the forward paths, enabling precise reconstruction and reliable feature sharing. Furthermore, we introduce a preservation mechanism based on attention feature injection and multi‑branch mask‑guided noise blending, which effectively balances fidelity and editability. Extensive experiments across diverse scenarios demonstrate that DirectEdit achieves efficient and accurate image editing, delivering superior performance that outperforms state‑of‑the‑art methods. Code and examples are available at https://desongyang.github.io/Directedit.
Authors:Jiaqi Shi, Jin Xiao, Xiaoguang Hu, Wenxuan Ji, Zichong Jia, Zifan Long, Tianyou Chen, Baochang Zhang
Abstract:
In 3D point cloud understanding, the core challenge lies in accurately capturing discriminative features within complex neighborhoods, which directly affects the execution precision of downstream tasks such as embodied AI and autonomous driving. Existing methods explore feature correlation discrimination but are limited to point‑level spatial distribution or channel responses, enabling only coarse‑grained level evaluation. For modern multi‑scale point cloud networks, such coarse‑grained metrics inevitably incur significant information loss in deeper layers. To address this, we propose PointCRA, a novel network with a channel‑level metric‑based enhancement mechanism. Our core idea is to introduce temporal trend variation as a new evaluation dimension to avoid the information loss caused by weight dimension collapse in existing spatial and channel attention mechanisms. On this basis, we construct a multi‑level calibration framework guided by neighborhood homogeneity for weight calibration, and design a dedicated loss function to enhance channel discriminability.PointCRA leverages intrinsic feature priors to adaptively correct feature aggregation, offering interpretability with low parameter overhead. Our method is transferable, interpretable, and efficient. We validate the proposed method on diverse datasets and benchmark models, and further demonstrate its rationality through extensive analytical experiments. Our PointCRA achieves 77.5% mIoU on the S3DIS dataset, 90.4% OA on the ScanObjectNN dataset, and 87.4% instance mIoU on the ShapeNetPart dataset. The code and pretrained weights are publicly available on GitHub: https://github.com/AGENT9717/PointCRA
Authors:Duy Nguyen Huu, Duy Hoang Khuong, Ngu Huynh Cong Viet
Abstract:
Chest radiography is a widely used imaging modality for thoracic disease diagnosis, yet its conventional interpretation remains time‑consuming and heavily dependent on expert knowledge. While deep learning has improved diagnostic efficiency through automated feature extraction, challenges such as class imbalance and the localization of multiple co‑existing pathologies remain unsolved. In this paper, inspired by the strength of Convolutional Block Attention Module (CBAM) in feature refinement and the capability of CNN blocks in feature extraction, we propose a strategy to integrate CBAM into traditional CNN blocks to enhance performance in multi‑label classification tasks. Our method achieves a mean AUC of 0.8695 on ChestXray14 dataset, outperforming several state‑of‑the‑art baselines.Our source code is available at: https://github.com/NNNguyenDuyyy/FETC_CBAM_Enhanced_CNN.git
Authors:Stefanos Pasios
Abstract:
Video game engines have been an important source for generating large volumes of visual synthetic datasets for training and evaluating computer vision algorithms that are to be deployed in the real world. While the visual fidelity of modern game engines has been significantly improved with technologies such as ray‑tracing, a notable sim2real appearance gap between the synthetic and the real‑world images still remains, which limits the utilization of synthetic datasets in real‑world applications. In this letter, we investigate the ability of a state‑of‑the‑art image generation and editing diffusion model (FLUX.2‑4B Klein) to enhance the photorealism of synthetic datasets and compare its performance against a traditional image‑to‑image translation model (REGEN). Furthermore, we propose a hybrid approach that combines the strong geometry and material transformations of diffusion‑based methods with the distribution‑matching capabilities of image‑to‑image translation techniques. Through experiments, it is demonstrated that REGEN outperforms FLUX.2‑4B Klein and that by combining both FLUX.2‑4B Klein and REGEN models, better visual realism can be achieved compared to using each model individually, while maintaining semantic consistency. The code is available at: https://github.com/stefanos50/Hybrid‑Sim2Real
Authors:Yagiz Nalcakan, Hyeongjin Ju, Incheol Park, Sanghyeop Yeo, Youngwan Jin, Shiho Kim
Abstract:
Vision Foundation Models (VFMs) pretrained on large‑scale RGB data have demonstrated remarkable representation quality, yet their applicability to multispectral imaging spanning Near‑Infrared (NIR), Short‑Wave Infrared (SWIR), and Long‑Wave Infrared (LWIR) remains largely unexplored. These spectral modalities offer complementary sensing capabilities critical for robust perception in adverse conditions, but present a fundamental domain gap relative to RGB‑centric pretrained models. We present SpectraDINO, a multispectral VFM that bridges this spectral gap by extending DINOv2 ViT backbones to beyond‑visible modalities through lightweight, per‑modality bottleneck adapters, while preserving the rich representations of the frozen RGB backbone. We introduce a multi‑stage teacher‑student training protocol in which a frozen DINOv2 teacher guides a spectral student via cosine distillation, symmetric contrastive loss, patch‑level alignment, and a novel neighborhood‑structure‑preservation loss. This staged curriculum enables strong cross‑modal alignment without catastrophic forgetting of RGB priors. We evaluate SpectraDINO on multispectral object detection and semantic segmentation across challenging NIR, SWIR, and LWIR benchmarks using widely adopted fusion strategies. SpectraDINO achieves state‑of‑the‑art performance across most benchmarks, validating its effectiveness as a general‑purpose backbone for spectral generalization. The code and weights for model variants are available at https://github.com/Yonsei‑STL/SpectraDINO.
Authors:Shihao Hou, Chikai Shang, Zhiheng Yang, Jiacheng Yang, Xinyi Shang, Junlong Gao, Yiqun Zhang, Yang Lu
Abstract:
Personalized federated learning (PFL) with foundation models has emerged as a promising paradigm enabling clients to adapt to heterogeneous data distributions. However, real‑world scenarios often face the co‑occurrence of non‑IID data and long‑tailed class distributions, presenting unique challenges that remain underexplored in PFL. In this paper, we investigate this long‑tailed personalized federated learning and observe that current methods suffer from two limitations: (i) fine‑tuning degrades performance below zero‑shot baselines due to the erosion of inherent class balance in foundation models; (ii) conventional personalization techniques further transfer this bias to local models through parameter or feature‑level fusion. To address these challenges, we propose Federated Learning via Gradient Purification and Residual Learning (FedPuReL), which preserves balanced knowledge in the global model while enabling unbiased personalization. Specifically, we purify local gradients using zero‑shot predictions to maintain a class‑balanced global model, and model personalization as residual correction atop the frozen global model. Extensive experiments demonstrate that FedPuReL consistently outperforms state‑of‑the‑art methods, achieving superior performance on both global and personalized models across diverse long‑tailed scenarios. The code is available at https://github.com/shihaohou/FedPuReL.
Authors:Abdullah Ahmad Khan, Hamid Laga, Ferdous Sohel
Abstract:
Machine unlearning in Vision‑Language Models (VLMs) is required for compliance with the General Data Protection Regulation (GDPR), yet current evaluation practices are inconsistent. We present the first systematic study of metric reliability in multimodal unlearning. Five standard metrics, Forget Accuracy (FA), Retain Accuracy (RA), Membership Inference Attack (MIA), Activation Distance (AD), and JS divergence (JS), yield conflicting method rankings across three VQA benchmarks (MLLMU‑Bench, UnLOK‑VQA, MMUBench). Kendall tau analysis over 36 unlearned LLaVA‑1.5‑7B models reveals two opposing clusters, FA, RA, MIA and AD, JS, with tau_FA_AD = ‑0.26, reproduced on BLIP‑2 OPT‑2.7B. Agreement is lower in multimodal VQA (average tau = 0.086) than in unimodal classification (average tau = 0.158; difference = 0.072), indicating that dual image‑and‑text pathways amplify inconsistency. We introduce the Unified Quality Score (UQS), a composite metric with weights derived from each metric's Spearman correlation with the oracle distance d(M_hat, M_star), where M_star is the oracle model retrained only on the retain set. RA shows the strongest reliability (rho = 0.484, p = 0.003), while FA is negatively correlated (rho = ‑0.418, p = 0.011). UQS yields stable rankings under 100 random weight perturbations (tau = 0.647 +‑ 0.262). We release the benchmark, 36 checkpoints, and an interactive leaderboard. Code and pre‑computed results are available at https://github.com/neurips26/UnifiedUnl.
Authors:Ce Wang, Zhenyu Hu, Wanjie Sun
Abstract:
Diffusion models have recently achieved remarkable performance in image super‑resolution (SR), but their high computational cost limits practical deployment in remote sensing applications. To address this issue, we propose SlimDiffSR, a lightweight and efficient diffusion‑based framework for real‑world remote sensing image super‑resolution. Unlike existing single‑step diffusion methods that rely on fixed timesteps, we first introduce an uncertainty‑guided timestep assignment strategy to construct a stronger single‑step teacher model, where reconstruction difficulty is explicitly linked to diffusion timesteps, enabling adaptive generative strength. Building upon this teacher, we further present a structured pruning strategy tailored to remote sensing imagery, which systematically removes redundant semantic modules and replaces standard operations with lightweight designs, including frequency‑separable convolution, direction‑separable convolution, and a query‑driven global aggregation module. These components explicitly exploit the unique characteristics of remote sensing data, such as sparse high‑frequency details, strong directional patterns, and long‑range spatial dependencies. To enhance knowledge transfer, we incorporate Maximum Mean Discrepancy (MMD) into the distillation process to align feature distributions between the teacher and student models. Extensive experiments on multiple remote sensing benchmarks demonstrate that SlimDiffSR achieves a favorable balance between efficiency and reconstruction quality. In particular, it attains up to 200× inference acceleration and a 20× reduction in model parameters compared with multi‑step diffusion models, while achieving competitive perceptual quality and clearly outperforming existing lightweight diffusion baselines in efficiency. The code is available at: https://github.com/wwangcece/SlimDiffSR.
Authors:Jianing Zhang, Zijian Zhou, Kai Sun
Abstract:
Pansharpening aims to generate high‑resolution multispectral (HRMS) images by fusing low‑resolution multispectral (LRMS) and high‑resolution panchromatic (PAN) images. Although deep learning has advanced this field, mainstream frequency‑based methods relying on standard scaled dot‑product attention suffer from quadratic computational complexity and fail to exploit the inherent regional sparsity of remote sensing imagery. Furthermore, existing spatial enhancement strategies typically employ static convolution kernels, which struggle to adapt to the complex frequency and regional variations of PAN and MS images. To address these bottlenecks, we propose a Region‑Aware Fusion (RAFNet) Network that synergistically models spatial and frequency information. Specifically, we design a Spatial Adaptive Refinement (SAR) module that leverages the discrete wavelet transform (DWT) for directional frequency separation and K‑means clustering for regional partitioning, which enables the dynamic construction of region‑specific adaptive convolution kernels, achieving spatially and frequency‑adaptive feature enhancement. Moreover, we introduce a Clustered Frequency Aggregation (CFA) module based on a sparse attention mechanism guided by the semantic clusters, which executes a region‑aware sparse attention strategy that drastically reduces computational redundancy while ensuring high‑quality frequency feature extraction. In addition we integrated these modules into a progressive, multi‑level spatial‑frequency network architecture to facilitate robust interaction and accurate image reconstruction. Extensive experiments on multiple benchmark datasets demonstrate that the proposed RAFNet significantly outperforms state‑of‑the‑art pansharpening methods in both reduced‑ and full‑resolution assessments. The code is available at https://github.com/PatrickNod/RAFNet.
Authors:Peggy Joy Lu, Wei-Yu Chen, Yao-Tsung Huang, Vincent Shin-Mu Tseng
Abstract:
We propose HeroCrystal, a novel privacy‑preserving framework for multi‑camera domain‑adaptive object detection, addressing challenges such as data privacy, class imbalance, and heterogeneous architectures. Our framework consists of three key stages. In the Generated Stage, we introduce a one‑shot, target‑aware diffusion‑based generation module that learns visual style from a single target‑domain image while leveraging prompt‑based control to synthesize specific object instances. Unlike conventional style transfer‑based methods that require large target datasets and ignore semantic‑level discrepancies, our approach enables privacy‑preserving augmentation to reduce ethical concerns, and introduces controllable rare object generation to mitigate long‑tailed category degradation. In the Federated Stage, we employ probabilistic Faster R‑CNN on the client side to improve localization accuracy, and a dynamic model contrastive strategy to suppress domain‑specific bias. The server side performs model fusion across heterogeneous architectures without accessing raw data. Finally, in the Distilled Stage, we propose an inconsistent categories integration algorithm to resolve label inconsistency and architecture heterogeneity across clients. Extensive experiments on multiple cross‑domain detection benchmarks demonstrate that our method outperforms existing multi‑source domain adaptation and federated learning baselines under multi‑class, privacy‑preserving settings. Our method improves mAP by +2.1% over prior privacy‑preserving approaches and achieves a new state‑of‑the‑art mAP of 33.4%, highlighting the effectiveness of HeroCrystal in enabling practical multi‑camera AI surveillance systems.
The source code is publicly available at https://github.com/ccuvislab/HeroCrystal.
Authors:Soyeon Kim, Seongwoo Lim, Kyowoon Lee, Jaesik Choi
Abstract:
Feature attribution is central to diagnosing and trusting deep neural networks, and Integrated Gradients (IG) is widely used due to its axiomatic properties. However, IG can yield unreliable explanations when the integration path between a baseline and the input passes through regions with noisy gradients. While Guided Integrated Gradients reduces this sensitivity by adaptively updating low‑gradient‑magnitude features, input‑space guidance still produces intermediate inputs that deviate from the data manifold. To address this limitation, we propose \emphManifold‑Aligned Guided Integrated Gradients (MA‑GIG), which constructs attribution paths in the latent space of a pre‑trained variational autoencoder. By decoding intermediate latent states, MA‑GIG biases the path toward the learned generative manifold and reduces exposure to implausible input‑space regions. Through qualitative and quantitative evaluations, we demonstrate that MA‑GIG produces faithful explanations by aggregating gradients on path features proximal to the input. Consequently, our method reduces off‑manifold noise and outperforms prior path‑based attribution methods across multiple datasets and classifiers. Our code is available at https://github.com/leekwoon/ma‑gig/.
Authors:Maximilian Kellner, Dominik Merkle, Michael Brunklaus, Alexander Reiterer
Abstract:
Large‑scale 3D point clouds can consist of hundreds of millions of points. Even after downsampling, these point clouds are too large for modern 3D neural networks. In order to develop a semantic understanding of the scene, the point clouds are divided into smaller subclouds that can be processed. Typically, this division is done using spherical crops, resulting in a loss of surrounding geometric context. To address this issue, we propose alternative methods that produce subclouds with larger crop sizes while maintaining a similar number of points. Specifically, we compare exponential, Gaussian, and linear cropping methods with the spherical method. We evaluated three 3D deep learning model architectures using multiple indoor and outdoor environment datasets. Our results demonstrate that altering the cropping strategy can enhance model performance, especially for large‑scale outdoor scenes, yielding new state‑of‑the‑art results. Code is available at https://github.com/mvg‑inatech/point_cloud_cropping
Authors:Daniel da Silva Costa, Pedro Nuno de Souza Moura, Adriana C. F. Alvim
Abstract:
In recent years, several advances have been observed in Deep Learning with surprising results. Models in this area have been increasingly used in numerous applications, including those sensitive to human life, which require clear explanations and justifications. Various explainability methods have been proposed, but not many metrics to evaluate these methods. The most commonly used metric is the Intersection over Union (IoU). However, due to the characteristics of the results of the explainability methods, called saliency maps, which do not have a known shape, we hypothesise that there must be a better metric that allows one to find an explainability method that produces results that best resemble the human perception. We propose using different metrics to assess the similarity between human perception and the explanation saliency maps to find a better metric. An investigation was conducted employing a subset of the Chihuahuas images from ImageNet dataset. Several CAM‑based explainability methods were used to generate saliency maps for each chihuahua image. Alignment was measured by applying distance metrics between the bounding box of human annotations and the saliency maps produced by each explainability method. Rankings of the best saliency maps were created using the results of the distance metrics and compared to the ranking obtained using people's choice, collected through crowdsourcing, of the best explanation saliency maps for each selected image. Comparison between rankings was performed using the Rank‑Biased Overlap (RBO) metric. The results indicate the feasibility of our method to find the explainability method that best resembles human perception. In our experiments, the two metrics that best resemble human perception corresponded to Manhattan and Correlation. Besides, the best explainability methods regarding human perception were LayerCAM, Score‑CAM, and IS‑CAM.
Authors:Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung, Weitao Xu
Abstract:
Vision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs), yet efficiently deploying them on FPGAs remains challenging. Linear layers struggle with dynamic activation outliers that render static quantization ineffective, while uniform quantization fails to capture the weight distribution at low bit‑widths. Furthermore, while associative scan accelerates SSMs on GPUs, its memory access patterns are misaligned with the streaming dataflow required by FPGAs. To address these challenges, we present ViM‑Q, a scalable algorithm‑hardware co‑design for end‑to‑end ViM inference on the edge. We introduce a hardware‑aware quantization scheme combining dynamic per‑token activation quantization and per‑channel smoothing to mitigate outliers, alongside a custom 4‑bit per‑block Additive Power‑of‑Two (APoT) weight quantization. The models are deployed on a runtime‑parameterizable FPGA accelerator featuring a linear engine employing a Lookup‑Table (LUT) unit to replace multiplications with shift‑add operations, and a fine‑grained pipelined SSM engine that parallelizes the state dimension while preserving sequential recurrence. Crucially, the hardware supports runtime configuration, adapting to diverse dimensions and input resolutions across the ViM family. Implemented on an AMD ZCU102 FPGA, ViM‑Q achieves an average 4.96x speedup and 59.8x energy efficiency gain over a quantized NVIDIA RTX 3090 GPU baseline for low‑batch inference on ViM‑tiny. This co‑design shows a viable path for deploying ViM models on resource‑constrained edge devices.
Authors:Yuchen Wang, Wenliang Zhong, Lichen Bai, Zikai Zhou, Shitong Shao, Bojun Cheng, Shuo Chen, Shuo Yang, Zeke Xie
Abstract:
Video diffusion models leveraging step distillation or causal distillation have achieved remarkable performance. However, adapting existing LoRAs to these variants remains a critical challenge due to weight space mismatches. We observe that direct application leads to style degradation and structural collapse, yet the underlying mechanisms remain poorly understood. To fill this gap, we delve into the weight space and identify that the incompatibility stems from spectral interference within shared functional clusters defined over singular subspaces. Specifically, our analysis reveals that while both paradigms respect spectral rigidity, they establish conflicting routing pathways that clash through constructive overload or destructive cancellation. To address this issue, we propose Cluster‑Aware Spectral Arbitration (CASA), a data‑free framework that dynamically arbitrates between safeguarding the target's manifold and restoring LoRA alignment based on spectral density. Extensive experiments demonstrate that CASA effectively mitigates artifacts and revives LoRA functionality. Our code is available at https://github.com/Noahwangyuchen/CASA
Authors:Vladislav Pyatov, Gleb Bobrovskikh, Saveliy Galochkin, Nikita Boldyrev, Oleg Voynov, Alexander Filippov, Gonzalo Ferrer, Peter Wonka, Evgeny Burnaev
Abstract:
We introduce CADFS, a data‑centric framework that enables large vision‑language models to generate complex CAD design histories. Existing generative CAD systems are restricted to sketch‑extrude operations due to simplified representations and limited datasets. We address this by introducing a FeatureScript‑based representation and constructing a dataset of 450k real‑world CAD models spanning 15 modeling operations. We obtain the dataset via a new pipeline that reconstructs clean, executable FeatureScript programs and provides multimodal annotations. Fine‑tuning a VLM on this representation yields state‑of‑the‑art results in text‑conditioned CAD generation and image‑based reconstruction, producing more accurate, diverse, and feature‑rich designs than prior frameworks. Ablations show that each individual component of our framework, i.e., the FeatureScript representation, the extended operation set, and representation‑aligned textual descriptions, significantly improves performance. Our framework substantially broadens the complexity and realism achievable in generative CAD. The CADFS framework and the new dataset are available at https://voyleg.github.io/cadfs/.
Authors:Rei Tamaru, Pei Li, Bin Ran
Abstract:
Traffic digital twins are powerful tools for advanced traffic management, and most systems are built on static geometric representations. However, these representations fail to capture the dynamic functional semantics required for behavior‑aware reasoning, such as how a lane operates under complex traffic conditions. To address this gap, we introduce GeoLaneRep, a behavior‑grounded lane representation learning framework for traffic digital twins. GeoLaneRep jointly encodes static lane geometry, observed vehicle trajectories, and operational descriptors into a shared, cross‑camera semantic embedding. The encoder is trained with a joint objective combining contrastive cross‑camera alignment, auxiliary role supervision, and temporal anomaly detection. Across 16 roadside cameras and 132 lanes, the learned embeddings achieve a 0.004 lateral‑rank error and an edge‑role F1 of 1.000 in zero‑shot cross‑camera matching, and an AUROC of 0.991 for window‑level anomaly detection. We further show that the same behavioral embeddings can condition a diffusion‑based generator to synthesize lane geometries that satisfy targeted operational specifications, with 87.9% overall specification accuracy across 38 lane groups. GeoLaneRep thus provides a semantic interface between roadside observations and downstream digital twin tasks, supporting cross‑camera transfer, behavior‑aware monitoring, and goal‑directed lane synthesis. The framework is openly available at https://github.com/raynbowy23/GeoLaneRep.
Authors:Hongkun Pan, Yuwei Wu, Wanyi Hong, Shenghui Hu, Qitong Yan, Yi Yang, Rufei Han, Changju Zhou, Minfeng Zhu, Dongming Han, Wei Chen
Abstract:
Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with high information density (HID) charts characterized by multiple subplots, legends, and dense annotations due to three major challenges: (1) limited fine‑grained perception results in the omission of critical visual cues; (2) redundant or noisy visual information undermines the performance of multimodal reasoning; (3) lack of adaptive deep reasoning relative to the amount of visual information. To tackle these challenges, we present a novel focus‑driven fine‑grained chart reasoning model, Chart‑FR1, to improve perception, focusing efficiency, and adaptive deep reasoning on HID charts. Specifically, we propose Focus‑CoT, a visual focusing chain‑of‑thought that enhances fine‑grained perception by explicitly linking reasoning steps to key visual cues, such as local image regions and OCR signals. Building on this, we introduce Focus‑GRPO, a focus‑driven reinforcement learning algorithm with an information‑efficiency reward that compresses redundant visual information for efficient focusing, and an adaptive KL penalty mechanism that enables flexible control over reasoning depth as more visual cues are discovered. Furthermore, to fill the gap in benchmarks for HID charts, we build HID‑Chart, a challenging benchmark with an information‑density metric designed to evaluate fine‑grained chart reasoning capabilities. Extensive experiments on multiple chart benchmarks demonstrate that Chart‑FR1 outperforms state‑of‑the‑art MLLMs in chart understanding and reasoning. Code is available at https://github.com/phkhub/Chart‑FR1.
Authors:Youyi Zhan, He Wang, Tianjia Shao, Kun Zhou
Abstract:
We propose a method to reconstruct high‑fidelity human avatars from multi‑view video that can run on mobile devices. Many works can model high‑quality Gaussian‑based full‑body avatars from multi‑view video. However, these methods require heavy computation to obtain pose‑dependent appearance, making deployment on mobile devices very difficult. Recent methods distill from pretrained models and model pose‑dependent nonlinear Gaussian attributes by linearly combining global pose features with blendshapes. Although they can run on mobile devices, they suffer some loss of detail. We observe that nearby Gaussians are often highly correlated within a local region of the body, and can be linearly modeled with less error. Therefore, we use local linear blendshapes in small body parts to capture global nonlinear changes of Gaussian attributes. To further reduce computation and model size, we propose to remove blendshapes for Gaussians whose attributes change little, yielding a minimal blendshape representation. Our method is an end‑to‑end training method without a pretrained model. To make it run on multiple devices, we implement our method using WebGPU. Experiments show that our method can render high‑quality human avatars with better details, and can reach 120 FPS at 2K resolution on mobile devices.
Authors:Lilika Makabe, Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita
Abstract:
Multi‑view 3D reconstruction, namely, structure‑from‑motion followed by multi‑view stereo, is a fundamental component of 3D computer vision. In general, multi‑view 3D reconstruction suffers from an unknown scale ambiguity unless a reference object of known size is present in the scene. In this article, we show that multi‑view images captured using a dual‑pixel (DP) sensor can automatically resolve the scale ambiguity, without requiring a reference object or prior calibration. Specifically, the defocus blur observed in DP images provides sufficient information to determine the absolute scale when paired with depth maps (up to scale) recovered from multi‑view 3D reconstruction. Based on this observation, we develop a simple yet effective linear method to estimate the absolute scale, followed by the intensity‑based optimization stage that aligns the left and right DP images by shifting them back toward each other using cross‑view blur kernels. Experiments demonstrate the effectiveness of the proposed approach across diverse scenes captured with different cameras and lenses. Code and data are available at https://github.com/lilika‑makabe/dp‑sfm‑tpami.git
Authors:Favour Nerrise, Lucy Yin, Mohammad H. Abbasi, Kilian M. Pohl, Ehsan Adeli
Abstract:
Brain MRI foundation models learn rich representations of anatomy, but interpreting what clinical information they encode remains an open problem. Standard sparse autoencoders (SAEs) suffer from severe feature collapse in deep transformer layers, and in Alzheimer's disease (AD) research, aging confounds nearly every clinical variable, making naive annotation unreliable. We propose GeoSAE, a geometry‑guided SAE framework that uses the foundation model's learned manifold structure to prevent feature collapse and annotates each surviving feature via age‑deconfounded partial correlations. Applied to ~14k T1‑weighted MRI scans from the Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Australian Imaging biomarkers and Lifestyle (AIBL) datasets, GeoSAE identifies a compact, fully interpretable feature set that predicts mild cognitive impairment (MCI)‑to‑AD conversion (AUC 0.746) using only 2% of the embedding dimensions, while comorbidity‑annotated features achieve only chance‑level performance. The identified features replicate across cohorts without retraining (r=0.97) and localize to neuroanatomically distinct regions consistent with Braak staging. This shows that geometry‑guided SAEs can extract interpretable, biomarkers from frozen brain MRI foundation models.
Authors:Yun Xing, Hanyuan Liu, Jiahao Nie, Shijian Lu
Abstract:
Large Multimodal Models (LMMs) have recently demonstrated their proficiency in holistic visual comprehension. However, most of them struggle to tackle region‑level perception guided by visual prompts, especially for cases where multiple regions are referred simultaneously, or scenarios where global contexts are necessary for precise visual referring. We introduce Contextual Latent Steering (CSteer), a training‑free approach for guiding general LMMs to refer multiple regions contextually, without expensive fine‑tuning or architectural modifications. CSteer starts with pre‑computing contextual vectors that implicitly represent visual referring behaviors, such as differentiation among regions and attention to global contexts, followed by representation editing during inference time. Experimental results on multiple datasets indicate that general LMMs with CSteer outperform tailored referring LMMs in most cases, suggesting a promising solution in training‑free, and setting new state‑of‑the‑art for this field. Code is available at https://github.com/xing0047/csteer.git.
Authors:Taiki Kanaya, Hideo Saito
Abstract:
Single‑image 3D face reconstruction is a core problem in computer vision, with important clinical applications such as cephalometric landmark analysis in orthodontics. Traditionally, this analysis relies on lateral X‑ray imaging; however, frequent X‑ray exposure is impractical due to radiation concerns. While recent research has explored detecting landmarks from lateral RGB images as an alternative, existing methods typically rely on 2D features such as the eyes, mouth, ears, and boundary silhouettes, failing to fully exploit the underlying 3D facial geometry spanning the facial profile and jawline, which is essential for accurate diagnosis. Meanwhile, although 3D face reconstruction from frontal views has seen significant progress, most learning‑based 3D morphable model (3DMM) regressors are developed and benchmarked on near‑frontal images, where appearance cues are abundant. In extreme profile views (yaw \approx 90^\circ), much of the face is occluded, and the available signal is dominated by boundary cues, making accurate 3D reconstruction challenging. In this paper, we bridge this gap with geometry‑conditioned synthetic data and a simple profile‑specific FLAME regression baseline for single lateral images. We introduce ProfileSynth, a dataset created by sampling FLAME shape and pose parameters in extreme yaw ranges and generating photorealistic profile images using a diffusion model conditioned on depth and normal maps. We further study a profile‑specific baseline with visibility‑aware jawline regularization. Our framework provides a practical baseline for "profile × 3DMM" reconstruction and a promising foundation for more accurate, non‑invasive cephalometric analysis from lateral RGB images.
Authors:Sixian Zhang, Yiyao Wang, Xinhang Song, Keming Zhang, Zijian Xu, Shuqiang Jiang
Abstract:
Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi‑scale semantics, and lack a native interface for large models, thus requiring additional training of feature projection for semantic alignment. To this end, we propose the multi‑scale Gaussian‑Language Map (GLMap), which introduces three key designs: (1) explicit geometry, (2) multi‑scale semantics covering both instance and region concepts, and (3) a dual‑modality interface where each semantic unit jointly stores a natural language description and a 3D Gaussian representation. The 3D Gaussians enable compact storage and fast rendering of task‑relevant images via Gaussian splatting. To enable efficient incremental construction, we further propose a Gaussian Estimator that analytically derives Gaussian parameters from dense point clouds without gradient‑based optimization. Experiments on ObjectNav, InstNav, and SQA tasks show that GLMap effectively enhances target navigation and contextual reasoning, while remaining compatible with large‑model‑based methods in a zero‑shot manner. The code is available at https://github.com/sx‑zhang/GLMap.
Authors:Jing Xu, Yuexiao Ma, Xuzhe Zheng, Xing Wang, Shiwei Liu, Chenqian Yan, Xiawu Zheng, Rongrong Ji, Fei Chao, Songwei Liu
Abstract:
Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential iterative denoising. While cache reuse strategies can accelerate generation by skipping redundant denoising steps, existing methods rely on coarse‑grained chunk‑level skipping that fails to capture fine‑grained pixel dynamics. This oversight is critical: pixels with high motion require more denoising steps to prevent error accumulation, while static pixels tolerate aggressive skipping. We formalize this insight theoretically by linking cache errors to residual instability, and propose MotionCache, a motion‑aware cache framework that exploits inter‑frame differences as a lightweight proxy for pixel‑level motion characteristics. MotionCache employs a coarse‑to‑fine strategy: an initial warm‑up phase establishes semantic coherence, followed by motion‑weighted cache reuse that dynamically adjusts update frequencies per token. Extensive experiments on state‑of‑the‑art models like SkyReels‑V2 and MAGI‑1 demonstrate that MotionCache achieves significant speedups of 6.28× and 1.64× respectively, while effectively preserving generation quality (VBench: 1%\downarrow and 0.01%\downarrow respectively). The code is available at https://github.com/ywlq/MotionCache.
Authors:Sen Fang, Hongbin Zhong, Yanxin Zhang, Dimitris N. Metaxas
Abstract:
Existing large‑scale sign language resources typically provide supervision only at the level of raw video‑text alignment and are often produced in laboratory settings. While such resources are important for semantic understanding, they do not directly provide a unified interface for open‑world recognition and translation, or for modern pose‑driven sign language video generation frameworks: 1. RGB‑based pretrained recognition models depend heavily on fixed backgrounds or clothing conditions during recording, and are less robust in open‑world settings than style‑agnostic pose‑processing models. 2. Recent pose‑guided image/video generation models mostly use a unified keypoint representation such as DWPose as their control interface. At present, the sign language field still lacks a data resource that can directly interface with this modern pose‑native paradigm while also targeting real‑world open scenarios. We present SignVerse‑2M, a large‑scale multilingual pose‑native dataset for sign language pose modeling and evaluation. Built from publicly available multilingual sign language video resources, it applies DWPose in a unified preprocessing pipeline to convert raw videos into 2D pose sequences that can be used directly for modeling, resulting in a consolidated corpus of about two million clips covering more than 55 sign languages. Unlike many laboratory datasets, this resource preserves the recording conditions and speaker diversity of real‑world videos while reducing appearance variation through a unified pose representation. Toward this goal, we further provide the data construction pipeline, task definitions, and a simple SignDW Transformer baseline, demonstrating the feasibility of this resource for multilingual pose‑space modeling and its compatibility with modern pose‑driven pipelines, while discussing the evaluation claims it can support as well as its current limitations.
Authors:Ruize He, Dongchen Han, Gao Huang
Abstract:
Existing research largely attributes the global sequence modeling capability of Transformers to the explicit computation of attention weights, a process that inherently incurs quadratic computational complexity. In this work, we offer a novel perspective: we demonstrate that attention can be mathematically reframed as a Multi‑Layer Perceptron (MLP) equipped with dynamically predicted parameters. Through this lens, we explain attention's global modeling power not as explicit token‑wise aggregation, but as an implicit process where dynamically generated parameters act as a compressed representation of the global context. Inspired by this insight, we investigate a fundamental question: can we achieve Transformer‑level sequence global modeling entirely through dynamic parameterization while maintaining linear complexity, effectively replacing explicit attention? To explore this, we design various dynamic parameter prediction strategies and integrate them into standard network layers. Extensive empirical studies on vision models demonstrate that dynamic parameterization can indeed serve as a highly effective, linear‑complexity alternative to explicit attention, opening new pathways for efficient sequence modeling. Code is available at https://github.com/LeapLabTHU/WeightFormer.
Authors:Qian Yin, Di Wen, Kunyu Peng, David Schneider, Zeyun Zhong, Alexander Jaus, Zdravko Marinov, Jiale Wei, Ruiping Liu, Junwei Zheng, Yufan Chen, Chen Zhang, Lei Qi, Rainer Stiefelhagen
Abstract:
Dense temporal annotation of procedural activity videos is vital for action understanding and embodied intelligence but remains labor‑intensive due to reactive tools. Each correction is treated as an isolated edit, limiting reuse of information on annotator uncertainty and model reliability. We introduce IMPACT‑Scribe, a correction‑driven framework for dense labeling that uses each correction to improve future human‑machine collaboration. IMPACT‑Scribe combines uncertainty‑aware boundary scribble supervision, local proposal modeling, cost‑aware query planning, structured propagation, and correction‑driven adaptation. Experiments and a human study show that this closed‑loop design improves labeling quality per effort, enhances boundary accuracy, and fosters better human‑machine interaction over time. The code will be made publicly available at https://github.com/BanzQians/IMPACT_AS.
Authors:Haoshen Zhang, Di Wen, Kunyu Peng, David Schneider, Zeyun Zhong, Alexander Jaus, Zdravko Marinov, Jiale Wei, Ruiping Liu, Junwei Zheng, Yufan Chen, Yufeng Zhang, Yuanhao Luo, Lei Qi, Rainer Stiefelhagen
Abstract:
We present IMPACT‑HOI, a mixed‑initiative framework for annotating egocentric procedural video by constructing structured event graphs for Human‑Object Interactions (HOI), motivated by the need for high‑quality structured supervision for learning robot manipulation from human demonstration. IMPACT‑HOI frames this task as the incremental resolution of a partially specified, onset‑anchored event state. A trust‑calibrated controller selects among direct queries, human‑confirmed suggestions, and conservative completions based on empirical annotator behavior and evidence quality. A risk‑bounded execution protocol, utilizing atomic rollback, ensures that human‑confirmed decisions are preserved against conflicting automated updates. A user study with 9 participants shows a 13.5% reduction in manual annotation actions, a 46.67% event match rate, and zero confirmed‑field violations under the studied protocol. The code will be made publicly available at https://github.com/541741106/IMPACT_HOI.
Authors:Tianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He, Xinze Li, Bingyu Zhu, Wuhui Duan, Congang Chen, Zeyu Fu, Yi Dong, Baoyuan Wu, Jason Li, Guangliang Cheng
Abstract:
Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single‑modality scope, simplified manipulations, or unrealistic distributions, which limit their ability to assess real‑world robustness. To address these limitations, we present Omni‑Fake, a unified omni‑dataset for comprehensive multimodal deepfake detection in social‑media settings. It comprises Omni‑Fake‑Set, a large‑scale, high‑quality dataset with 1M+ samples, and Omni‑Fake‑OOD, an out‑of‑distribution benchmark with 200k+ samples intentionally excluded from training to evaluate generalization. Omni‑Fake spans four modalities (image, audio, video, and audio‑video talking head) and supports a joint detection‑localization‑explanation protocol. On top of Omni‑Fake, we further propose Omni‑Fake‑R1, a reinforcement‑learning‑driven multimodal detector that adaptively integrates visual and auditory cues and outputs structured decisions, localization, and natural‑language explanations. Extensive experiments show significant gains in detection accuracy, cross‑modal generalization, and explainability over state‑of‑the‑art baselines. Project page: https://tianxiao1201.github.io/omni‑fake‑project‑page/
Authors:Joy Dhar, Song Xia, Manish Kumar Pandey, Maryam Haghighat, Azadeh Alavi, Ferdous Sohel, Wenyu Zhang, Nayyar Zaidi
Abstract:
We introduce Hybrid Convolutions with Attention Stochasticity (HyCAS), an adversarial defense that narrows the long‑standing gap between provable robustness under L2 certificates and empirical robustness against strong L attacks, while preserving strong generalization across diverse imaging benchmarks. HyCAS unifies deterministic and randomized principles by coupling 1‑Lipschitz, spectrally normalized convolutions with two stochastic components, spectral normalized random, projection filters and a randomized attention‑noise mechanism, to realize a randomized defense. Injecting smoothing randomness inside the architecture yields an overall <= 2‑Lipschitz network with formal certificates. Exten‑sive experiments on diverse imaging benchmarks, including CIFAR‑10/100, ImageNet‑1k, NIH Chest X‑ray, HAM10000, show that HyCAS surpasses prior leading certified and empirical defenses, boosting certified accuracy by up to 7.3% (on NIH Chest X‑ray) and empirical robustness by up to 3.1% (on HAM10000), without sacrificing clean accuracy. These results show that a randomized Lipschitz constrained architecture can simultaneously improve both certified L2 and empirical L adversarial robustness, thereby supporting safer deployment of deep models in high‑stakes applications. Code: https://github.com/misti1203/HyCAS
Authors:Guotao Liang, Zhangcheng Wang, Chuang Wang, Juncheng Hu, Haitao Zhou, Junhua Liu, Jing Zhang, Dong Xu, Qian Yu
Abstract:
Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence. However, this task remains challenging as it requires bridging discrete code representations with continuous visual dynamics. Existing optimization‑based methods often destroy topological consistency, while general‑purpose LLMs rely on rigid CSS/SMIL transformations, failing to model geometry‑level non‑rigid deformations. To address these limitations, we present VAnim, the first LLM‑based framework for open‑domain text‑to‑SVG animation. We reconceptualize animation not as sequence generation, but as Sparse State Updates (SSU) on a persistent SVG DOM tree. This paradigm compresses sequence length by over 9.8x while preserving the SVG DOM structure and non‑participating elements by construction. To enable precise control, we propose an Identification‑First Motion Planning mechanism that grounds textual instructions in explicit visual entities. Furthermore, to overcome the non‑differentiable nature of SVG rendering, we employ Rendering‑Aware Reinforcement Learning via Group Relative Policy Optimization (GRPO). By leveraging a hybrid reward from a state‑of‑the‑art video perception encoder, we align discrete code updates with high‑fidelity visual feedback. We also introduce SVGAnim‑134k, the first benchmark for vector animation. Extensive experiments demonstrate that VAnim significantly outperforms state‑of‑the‑art baselines in semantic alignment and structural validity, with additional appendix metrics further validating motion quality and identity preservation.
Authors:Liang Peng, Bohan Tan, Zhipeng Zhang, Haobo Li, Yifan Jiao, Xingping Dong, Libo Zhang
Abstract:
Visual query localization (VQL) aims to predict the spatio‑temporal response of the most recent occurrence in a sequence given a query. Currently, most research focuses on visual query localization in 2D videos, while its counterpart in 3D space has received little attention. In this paper, we make the first attempt to address visual query localization in the 3D world by introducing a novel benchmark, dubbed 3DVQL. Specifically, 3DVQL contains 2,002 sequences with around 170,000 frames and 6.4K response track segments from 38 object categories. Each sequence in 3DVQL is provided with multiple modalities, including point clouds, RGB images, and depth images, to support flexible research. To ensure high‑quality annotations, each sequence is manually annotated with multiple rounds of verification and refinement. To the best of our knowledge, 3DVQL is the first benchmark for 3D multimodal visual query localization. To facilitate comparison in subsequent research, we implement a series of representative 3D multimodal VQL baselines using point clouds and RGB images. The experimental results show that existing methods exhibit significant performance variations across different fusion modules. To encourage future research, we propose a lift‑and‑attention fusion algorithm named LaF, which significantly outperforms existing baseline models. Our benchmark and model will be publicly released at https://github.com/wuhengliangliang/3DVQL.
Authors:Kanak Mazumder, Fabian B. Flohr
Abstract:
Online High‑Definition (HD) map construction is a key component of autonomous driving. Recent methods rely on multi‑view camera images for cost‑effective HD map segmentation, but cameras lack depth information for accurate scene geometry. In contrast, LiDAR provides precise 3D measurements but lacks dense semantic cues. In this work, we propose LIE, LiDAR‑only semantic map construction method that employ Knowledge Distillation (KD) to handle the lack of dense semantic and texture cues. Specifically, the teacher branch fuses student LiDAR features and the corresponding 2D intensity map tile to provide dense supervision for segmenting map elements using online distillation scheme. Experimental results show that our method outperforms all single‑modality approaches, achieving 8.2% higher mIoU than the state‑of‑the‑art camera‑based model on nuScenes. LIE is robust over long ranges and under challenging weather and lighting, and efficiently adapts to Argoverse2 with only 10% fine‑tuning, surpassing camera‑based models trained on the full dataset. Source code will be available \hrefhttps://iv.ee.hm.edu/lie/here.
Authors:Jiacheng Yang, Ruichi Zhang, Chikai Shang, Mengke Li, Xinyi Shang, Junlong Gao, Yonggang Zhang, Yang Lu
Abstract:
Long‑tailed data bias decision boundaries toward head classes and degrade tail class accuracy. Diffusion‑based generative augmentation address this problem by generating additional data, while head‑to‑tail transfer further mitigate the generator bias inherit from long‑tailed dataset. However, we show that while head‑to‑tail transfer helps balance the decision space of the classifier, it also induces latent non‑local feature mixing that entangles inter‑class features, causing decision boundary overlap and tail class distribution shift. To address this, we first identify the problem of boundary ambiguity and then propose Decision Boundary‑aware Generation (DBG) framework, which promotes near‑boundary representation learning by generating informative near‑boundary samples. Overall, DBG rebalances the long‑tailed dataset while yielding more separable decision space for long‑tailed learning. Across standard long‑tailed benchmarks, DBG consistently improves tail class and overall accuracy with less inter‑class overlap. The code of DBG is available at https://github.com/keepdigitalabc‑svg/DBG.
Authors:Zhaoyang Li, Zhichao You, Tianrui Li
Abstract:
Although multi‑modal learning has advanced point cloud completion, the theoretical mechanisms remain unclear. Recent works attribute success to the connection between modalities, yet we identify that standard hard projection severs this connection: projecting a sparse point cloud onto the image plane yields an extremely sparse support, which hinders visual prior propagation, a failure mode we term Cross‑Modal Entropy Collapse. To address this practical limitation, we propose SplAttN, which replaces hard projection with Differentiable Gaussian Splatting to produce a dense, continuous image‑plane representation. By reformulating projection as continuous density estimation, SplAttN avoids collapsed sparse support, facilitates gradient flow, and improves cross‑modal connection learnability. Extensive experiments show that SplAttN achieves state‑of‑the‑art performance on PCN and ShapeNet‑55/34. Crucially, we utilize the real‑world KITTI benchmark as a stress test for multi‑modal reliance. Counter‑factual evaluation reveals that while baselines degenerate into unimodal template retrievers insensitive to visual removal, SplAttN maintains a robust dependency on visual cues, validating that our method establishes an effective cross‑modal connection. Code is available at https://github.com/zay002/SplAttN.
Authors:Panagiotis P. Filntisis, George Retsinas, Radek Daněček, Vanessa Sklyarova, Petros Maragos, Timo Bolkart
Abstract:
Recent frameworks like ToFu and TEMPEH provide an automated alternative to classical registration pipelines by predicting 3D meshes in dense semantic correspondence directly from calibrated multi‑view images. However, these learning‑based methods rely on the slow, manual registration pipelines they aim to replace for their training supervision. We overcome this limitation with MOCHI (Multi‑view Optimizable Correspondence of Heads from Images), a multi‑view 3D face prediction framework trained without requiring registered training data. MOCHI eliminates the registration data dependency by enforcing topological consistency through a pseudo‑linear inverse kinematic solver. Semantic alignment is guided by dense keypoints from a 2D landmark predictor trained exclusively on synthetic data. Our analysis further reveals that standard point‑to‑surface distances induce training instabilities and visual artifacts in registration‑free settings. We propose pointmap‑ and normal‑based losses instead, which provide smoother gradients and superior reconstruction fidelity. Finally, we introduce a test‑time optimization scheme that refines network weights over a few dozen iterations. This approach bridges the gap between feed‑forward efficiency and iterative optimization precision, allowing MOCHI to outperform traditional labor‑intensive pipelines in both reconstruction accuracy and visual quality. Code and model are public at: https://filby89.github.io/mochi.
Authors:Abhishek Vivekanandan, Ahmed Abouelazm, J. Marius Zöllner
Abstract:
Motion forecasting often requires trading interpretability for predictive accuracy. Standard anchor‑based architectures rely on opaque latent queries that are highly prone to latent collapse, or naive trajectory sampling that limits multi‑modal diversity. We propose an end‑to‑end differentiable framework that grounds predictions in a comprehensive "motion bank", a structured embedding space of physically realizable trajectories constructed via contrastive learning. Rather than regressing paths from a blank slate, our architecture dynamically retrieves explicit motion priors using a novel Anchor Retrieval Layer. This module adapts orthogonally initialized queries via a Dual‑Level Gated Cross‑Attention mechanism and executes discrete trajectory selection using a Straight‑Through Gumbel‑Softmax estimator to preserve continuous gradient flow. The retrieved semantically grounded anchors are then geometrically refined by a DETR‑style decoder, optimized jointly with a Winner‑Takes‑All (WTA) kinematic Gaussian Mixture Model (GMM), a latent diversity penalty, and a soft‑min weighted endpoint loss. By strictly conditioning the decoding phase on diverse, interpretable motion primitives, our approach eliminates the "black box" of standard latent queries while achieving competitive multi‑modal accuracy on the Argoverse 2 and Waymo Open Motion datasets. Code is available at: https://github.com/abviv/recall2predict
Authors:Jingze Wu, Quan Zhang, Hongfei Suo, Zeqiang Cai, Hongbo Chen
Abstract:
Although reinforcement learning (RL) has significantly advanced reasoning capabilities in large multimodal language models (MLLMs), its efficacy remains limited for lightweight models essential for edge deployments. To address this issue, we leverage causal analysis and experiment to reveal the underlying phenomenon of perceptual bias, demonstrating that RL‑based fine‑tuning compels lightweight models to preferentially adopt perceptual shortcuts induced by data biases, rather than developing genuine reasoning abilities. Motivated by this insight, we propose VideoThinker, a causal‑inspired framework that cultivates robust reasoning in lightweight models through a two‑stage debiasing process. First, the Bias Aware Training stage forges a dedicated "bias model" to embody these shortcut behaviors. Then, the Causal Debiasing Policy Optimization (CDPO) algorithm fine‑tunes the primary model, employing an innovative repulsive objective to actively push it away from the bias model's flawed logic while simultaneously pulling it toward correct, generalizable solutions. Our model, VideoThinker‑R1, establishes a new state‑of‑the‑art in video reasoning efficiency. For same‑scale comparison, requiring no Supervised Fine‑Tuning (SFT) and using only 1 of the training data for RL, it surpasses VideoRFT‑3B with a 3.2% average gain on widely‑used benchmarks and a 7% lead on VideoMME. For cross‑scale comparison, it outperforms the larger Video‑UTR‑7B model on multiple benchmarks, including a 2.1% gain on MVBench and a 3.8% gain on TempCompass. Code is available at https://github.com/falonss703/VideoThinker.
Authors:Ruichi Zhang, Chikai Shang, Jiacheng Yang, Mengke Li, Yang Zhou, Junlong Gao, Yang Lu
Abstract:
Long‑tailed distributions are common in real‑world recognition tasks, where a few head classes have many samples while most tail classes have very few. Recently, fine‑tuning foundation models for long‑tailed learning has gained attention due to their excellent performance. However, most existing methods focus solely on mitigating long‑tailed distribution bias while overlooking concept confusion caused by the long‑tailed distribution. In this paper, we study this problem and attribute it to the mutual exclusivity of single‑label supervision under long‑tailed distributions, which suppresses feature sharing among related classes and amplifies the dominance of head classes, leading to disrupted inter‑class discriminability. To address this, we propose CUE, Concept‑aware mUlti‑label Expansion, which introduces multi‑label concept signals to preserve disrupted inter‑class relationships. Specifically, CUE constructs concept sets by (i) extracting instance‑level visual cues from zero‑shot CLIP and (ii) generating class‑level semantic cues with LLM; the two cues are incorporated via separately weighted Binary Logit‑Adjustment (BLA) auxiliary losses and jointly optimized with the baseline Logit‑Adjustment (LA) loss. Experiments on several long‑tailed benchmarks, CUE achieves balanced and strong performance, surpassing recent state‑of‑the‑art methods. Code is available at: https://github.com/zhangruichi/CUE.
Authors:Kosuke Takemoto, Takafumi Koshinaka
Abstract:
Diffusion‑based virtual try‑on methods achieve photorealistic synthesis through cross‑attention mechanisms that transfer garment features to target body regions. However, these approaches rely on implicit learning of spatial correspondences, struggling to preserve fine details such as text and illustrations. We propose a novel approach, which we call SIFT‑VTON, that utilizes SIFT keypoint matching to provide explicit geometric guidance for diffusion‑based virtual try‑on. Our method applies domain‑specific filtering to SIFT keypoint matches between garment and person images, then converts these correspondences into spatial probability distributions that supervise cross‑attention layers during training. This explicit supervision guides the model to learn precise spatial alignment, concentrating attention on geometrically consistent garment regions. Experiments on the VITON‑HD dataset demonstrate significant improvements on unpaired metrics while maintaining competitive paired reconstruction metrics. Qualitative comparisons show superior preservation of text clarity and pattern alignment. Attention visualizations confirm that our method produces sharply focused attention on relevant garment details. This work demonstrates that classical geometric correspondence methods can effectively enhance modern diffusion models for conditional synthesis tasks. The source code will be available at https://github.com/takesukeDS/SIFT‑VTON.
Authors:Peiyang Liu, Ziqiang Cui, Xi Wang, Di Liang, Wei Ye
Abstract:
Iterative Retrieval‑Augmented Generation (iRAG) has emerged as a powerful paradigm for answering complex multi‑hop questions by progressively retrieving and reasoning over external documents. However, current systems predominantly operate on parsed text, which creates two critical bottlenecks: (1) Coarse‑grained attribution, where users are burdened with manually locating evidence within lengthy documents based on vague text‑level citations; and (2) Visual semantic loss, where the conversion of visually rich documents (e.g., slides, PDFs with charts) into text discards spatial logic and layout cues essential for reasoning. To bridge this gap, we present Chain of Evidence (CoE), a retriever‑agnostic visual attribution framework that leverages Vision‑Language Models to reason directly over screenshots of retrieved document candidates. CoE eliminates format‑specific parsing and outputs precise bounding boxes, visualizing the complete reasoning chain within the retrieved candidate set. We evaluate CoE on two distinct benchmarks: Wiki‑CoE, a large‑scale dataset of structured web pages derived from 2WikiMultiHopQA, and SlideVQA, a challenging dataset of presentation slides featuring complex diagrams and free‑form layouts. Experiments demonstrate that fine‑tuned Qwen3‑VL‑8B‑Instruct achieves robust performance, significantly outperforming text‑based baselines in scenarios requiring visual layout understanding, while establishing a retriever‑agnostic solution for pixel‑level interpretable iRAG. Our code is available at https://github.com/PeiYangLiu/CoE.git.
Authors:Rajesh Sureddi, Shreshth Saini, Avinab Saha, Alan C. Bovik
Abstract:
The development of video game streaming has grown rapidly, with major platforms such as YouTube and Twitch using different codecs. To support quality assessment models that work consistently across any codec, it is necessary to have access to large, diverse subjective gaming quality datasets. Currently, there are only a few available, each having limitations. To address this gap, we present the largest gaming video quality dataset to date, incorporating both user‑generated content (UGC) and professional‑generated content (PGC) with extensive visual diversity. Our dataset covers the most widely used codecs ‑ H.264, H.265, and AV1 ‑ and consists of 4,048 video samples, each annotated by an average of 37 mean opinion score (MOS) ratings. In addition to overall quality scores, we collect coarse‑grained quality attributes, enabling a better understanding of perceptual factors. We study the performance of leading video quality assessment methods on this dataset, including a vision language model that outperforms all the benchmarks. To the best of our knowledge, this is the first dataset that comprehensively addresses gaming video quality assessment across multiple codecs and content types with quality attributes. Our dataset is publicly available at https://rajeshsureddi.github.io/GameScope/.
Authors:Lei He, Jielei Chu, Fengmao Lv, Weide Liu, Tianrui Li, Jun Cheng, Yuming Fang
Abstract:
Unified image restoration using a single model often faces task interference due to diverse degradations. To address this, we propose DACG‑IR (Degradation‑Aware Adaptive Context Gating), which enables explicit perception of degradation characteristics to dynamically modulate feature representations. Our method constructs degradation‑aware contextual representations from the input to modulate attention distribution, frequency‑domain features, and feature aggregation. Specifically, a lightweight multi‑scale degradation‑aware module extracts coarse degradation information and generates layer‑wise prompts. These prompts guide attention temperature and output gating in encoder and decoder blocks for adaptive feature extraction. Additionally, a spatial‑channel dual‑gated adaptive fusion mechanism refines encoder features, suppressing noise propagation from shallow to deep layers. This design effectively suppresses degradation‑induced noise while preserving informative structures. Experiments show DACG‑IR outperforms state‑of‑the‑art methods in single‑task, all‑in‑one, adverse weather removal, and composite degradation settings. Code: https://github.com/HlHomes/DACG‑IR‑code
Authors:Valter Estevam, Rayson Laroca, Helio Pedrini, David Menotti
Abstract:
This paper proposes a novel Zero‑Shot Action Recognition~(ZSAR) method based on contrastive learning. In ZSAR, we aim to classify examples from classes that were missing during training. Two well‑known problems remain in ZSAR: the semantic gap and the domain shift. A semantic gap occurs because label representations come from the textual domain (i.e., language models) and must be associated with visual representations (i.e., CNNs, RNNs, transformer‑based). This multimodal nature implies that the semantic properties of the two spaces are not identical. On the other hand, the domain shift arises from differences between the training and test sets and is inherent to ZSAR once the test set is unknown. One of the most promising methods to address both issues is learning joint embedding spaces. Therefore, we propose a new model that encodes videos and sentences in a joint embedding space, trained by aligning videos with their natural‑language descriptions. We design an automatic negative sampling procedure to augment the training dataset and generate unpaired data, i.e., visual appearance and unrelated descriptions. Our results are state‑of‑the‑art on the UCF‑101 and Kinetics‑400 datasets under several split configurations. Our code is available at https://github.com/valterlej/cezsar.
Authors:Hamidreza Aftabi, John E. Lloyd, Amanda Ding, Benedikt Sagl, Eitan Prisman, Antony Hodgson, Sidney Fels
Abstract:
Mandibular reconstruction with vascularized bone grafts is complicated by donor‑host nonunion, and current virtual surgical planning produces a geometric plan rather than a configuration that explicitly promotes bone union. We present OsteoOpt++, an image‑to‑decision planning loop for patient‑specific mandibular reconstruction. A pre‑operative computed tomography (CT) is converted into a personalized digital twin through template‑to‑patient registration and CT‑derived updates of the muscle and temporomandibular‑joint parameters. Bayesian optimization with an expected‑improvement‑plus acquisition rule then searches six clinically controllable cut‑plane and donor‑positioning variables under an apposition‑driven objective and a safety‑factor‑regularized variant. The workflow was evaluated on three generic defects (body, symphysis, and ramus‑body) and a total of 3+1 patient‑specific cases, with 3 used for optimization and 1 for validation. In the generic cases, against a common surgical approach, cycle‑averaged donor‑mandible apposition increased by up to 29 percentage points (329% relative); in the patient‑specific cases, against the surgeon‑implemented day‑5 post‑operative configuration, by up to 26 percentage points. A 10% sensitivity analysis over eleven modeling parameters capped the change in the apposition‑driven objective at 3% for generic cases and 4% for patient‑specific cases, and the longitudinal case showed Dice overlap of 0.70 and 0.76 between predicted apposition and year‑1 bone formation. Clinically, this provides surgeons with a pre‑operative, image‑driven recommendation for cut‑plane orientation and donor placement that is predicted to improve union conditions over the configurations currently delivered in the operating room. The optimization and patient‑specific modeling code is open source at https://github.com/hamidreza‑aftabi/OsteoOpt.
Authors:Hamed Khatounabadi, Xiaohu Lu, Hayder Radha
Abstract:
The performance of state‑of‑the‑art object detectors degrades significantly under adverse weather, causing a safety‑critical domain shift problem for autonomous vehicles. Recent efforts address this problem by relying on synthetic data to train the object detectors, which limits their real‑world applicability. Meanwhile, pseudo‑labeling is widely used for cross‑dataset domain adaptation problems. However, these methods have not been exploited by weather‑based domain adaptation approaches due to the noisy nature of such labels generated under harsh weather conditions. In this paper, we propose two new approaches to mitigate this weather‑induced domain shift. First, we propose a Weather‑Induced pseudo Label Denoising (WILD) framework that filters noisy pseudo labels generated by real data captured under adverse weather conditions. Second, we develop a novel hybrid training methodology, WILD SAM, that exploits both pseudo‑label denoising and simulation‑based training solutions while using real‑data from the target harsh‑weather domain. We validate both proposed approaches, WILD and WILD SAM, on the recently released Four Seasons dataset across rainy and snowy scenarios. Experiments show that the proposed frameworks improve Average Precision (AP) up to 13% and significantly reduce the weather‑induced performance gap relative to the baseline. The code is available at: https://github.com/Kh‑Hamed/WILD‑SAM
Authors:Johannes B. Thalhammer, Lorenzo D'Amico, Lucy Costello, Sebastian Peterhansl, Daniel Frey, Tina Dorosti, Florian Schaff, Jannis Ahlers, Ronan Smith, Marcus Kitchen, Franz Pfeiffer, Martin Donnelley, Daniela Pfeiffer, Kaye S. Morgan
Abstract:
Propagation‑based X‑ray phase‑contrast imaging (PBI) enables high‑contrast visualization of lung structures and holds strong medical potential. However, safe translation to the clinic will require a substantial radiation dose reduction, which inevitably increases image noise. Supervised convolutional‑neural‑network‑based denoising can restore image quality but depends on paired low‑ and high‑dose datasets, which are rarely available in practice. Self‑supervised methods avoid this limitation, yet most are not well adapted to the inverse problem of PBI computed tomography (CT). We introduce Neighbor2Inverse, a self‑supervised denoising framework designed for low‑dose PBI‑CT that generalizes to clinical CT. Building on the Neighbor2Neighbor principle, each noisy projection is subsampled into two variants that preserve structural information but contain independent noise realizations. These are reconstructed separately, and the resulting pairs are used to train a denoising network directly in the image domain. We benchmark the proposed method against established analytical and self‑supervised denoising approaches. In region‑of‑interest PBI CT experiments, Neighbor2Inverse achieves superior noise suppression while preserving fine structural details, as demonstrated by improved contrast‑to‑noise ratio, spatial resolution, and composite image quality metrics. Competitive performance is also observed on clinical CT data under simulated low‑dose conditions.
This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.
Code, data, and interactive figures are available at https://github.com/J‑3TO/Neighbor2Inverse.
Authors:Junzhe Huang, Xiaoxiao Sun, Yan Yang, Yuxuan Hou, Ruotian Zhang, Sirui Li, Hehe Fan, Serena Yeung-Levy, Xin Yu
Abstract:
Using multimodal foundation models to analyze table images is a high‑value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluations rely largely on structured‑text tables or clean rendered images, leaving the visual complexity of in‑the‑wild table images underexplored. Such images feature varied layouts and diverse domains that demand sophisticated structural perception and numerical reasoning. To bridge this gap, we introduce WildTableBench, the first question‑answering benchmark for naturally occurring table images from real‑world settings. WildTableBench comprises 402 high‑information‑density table images collected from online forums and websites across diverse domains, together with 928 manually annotated and verified questions spanning 17 subtypes across five categories. We evaluate 21 frontier proprietary and open‑source multimodal foundation models on this benchmark. Only one model exceeds 50% accuracy, while all remaining models range from 4.1% to 49.9%. We further conduct diagnostic analyses to characterize model failures and reveal persistent weaknesses in structural perception and reasoning. These results and analyses provide useful insights into current model capabilities and establish WildTableBench as a valuable diagnostic benchmark for table image understanding. Dataset: https://huggingface.co/datasets/jzhuang/WildTableBench Code: https://github.com/hjzhe/WildTableBench Leaderboard: https://hjzhe.github.io/WildTableBench
Authors:Chirag Shinde
Abstract:
We introduce energy‑based constraint networks ‑‑ a modality‑agnostic architecture that learns structural coherence from contrastive pairs. The system processes frozen encoder embeddings through a state‑space model with dual‑head attention, producing a scalar energy measuring structural consistency alongside per‑position energy scores that localize violations. Multiple independently trained branches detect different violation types and compose at inference without interference.
We demonstrate the framework in two domains. In text, the system achieves 93.4% accuracy on trained corruption types and 87.2% on 9 unseen types, using frozen BERT and 7.4M trainable parameters. In vision, the same architecture achieves competitive deepfake detection: 0.959 AUC on FaceForensics++ Deepfakes and 0.870 on Celeb‑DF without any Celeb‑DF training data, using frozen DINOv2 and 3.6M parameters per branch.
The framework supports flexible training: branches learn from designer‑specified corruptions, real‑world paired data, or both. Composable branches require representation compatibility ‑‑ a finding validated through extensive experimentation where five incompatible approaches failed before the compatible one succeeded. The architecture is encoder‑agnostic and domain‑agnostic: changing the domain requires only new corruption strategies; changing the encoder requires only a new input projection layer. To our knowledge, this is the first architecture to learn within‑modality structural coherence as an explicit energy landscape with per‑position decomposition, and to demonstrate that the same architecture transfers across modalities via corruption respecification alone.
Authors:Rui Zhang, Xianzhi Song, Linqi Zhu, Branko Bijeljic, Gensheng Li, Martin J. Blunt
Abstract:
Reliable segmentation of multiphase pore‑scale X‑ray images of rocks is necessary to quantify fluid saturation, connectivity, and interfacial geometry. However, current 3D segmentation methods are typically dataset‑specific, requiring retraining or extensive fine‑tuning whenever rock type, fluid pattern, scanner, or acquisition conditions change. Foundation models such as the Segment Anything Model (SAM) provide strong 2D boundary priors, but they are not directly applicable to 3D data.
We present SAMamba3D, a parameter‑efficient framework that adapts a largely frozen SAM encoder to generalizable 3D pore‑scale segmentation by coupling it with Mamba‑based volumetric context modeling and progressive cross‑scale feature interaction. For sandstone and carbonate datasets, with different fluids, wettability, and scanning conditions, SAMamba3D matches or outperforms current 3D baselines while reducing the need for case‑specific retraining. The resulting segmented images preserve physically meaningful descriptors, including fluid saturation, connectivity, and interface morphology, enabling more reliable and rapid analysis of large 3D multiphase images.
Authors:Lin Sun, Wang Dexian, Jingang Huang, Linglin Zhang, Change Jia, Zhengwei Cheng, Xiangzheng Zhang
Abstract:
Industrial Retrieval‑Augmented Generation (RAG) systems depend on optical character recognition (OCR) to transform visual documents into text. Existing OCR benchmarks rely on character‑level metrics, which inadequately measure downstream RAG effectiveness under real‑world conditions. We introduce an OCR benchmark for industrial RAG systems covering 11 challenging document types, including extreme layouts, high‑resolution pages, complex or watermarked backgrounds, historical documents with non‑standard reading orders, visually decorated text, and documents containing tables and mathematical formulas. Evaluating recent SOTA OCR models under a controlled OCR‑first RAG pipeline shows clear performance degradation on realistic industrial documents despite strong conventional benchmark scores. We find that high OCR accuracy does not necessarily translate into strong downstream RAG performance: structural and semantic errors can cause substantial retrieval failures even when WER/CER remains low. Further analysis shows that this mismatch is category‑dependent, arises through both retrieval‑side and downstream generation‑side failures, and remains stable across representative OCR‑first pipeline choices. The benchmark is publicly available at https://github.com/Qihoo360/InduOCRBench.
Authors:Hongjun Wang, Po Hu, Kai Han
Abstract:
Generalized Category Discovery (GCD) aims to categorize unlabelled instances from both known and unknown classes by transferring knowledge from labelled data of known classes. Existing methods assume all data comes from a single domain, yet real‑world unlabelled data often exhibits domain shifts alongside semantic shifts. We study GCD under domain shifts and propose three frameworks that adapt foundation models, ranging from self‑supervised vision models to vision‑language models. (i) HiLo disentangles domain and semantic features through multi‑level feature extraction and mutual information minimization, combined with PatchMix augmentation and curriculum sampling. (ii) HLPrompt extends HiLo with semantic‑aware spatial prompt tuning to suppress background and domain noise. (iii) VLPrompt leverages vision‑language models via factorized textual prompts and cross‑modal consistency regularization. The three methods share core design principles while operating on different foundation backbones, making them suitable for different deployment scenarios. Extensive experiments on synthetic corruptions and real‑world multi‑domain shifts demonstrate consistent improvements over strong baselines. Project page: https://visual‑ai.github.io/hilo/
Authors:Prabhjot Singh, Manmeet Singh
Abstract:
Operational phase unwrapping is the primary computational bottleneck in InSAR‑based volcanic and seismic monitoring. We challenge the industry trend of adopting high‑complexity computer vision architectures, such as attention mechanisms, without validating their suitability for physics‑constrained geophysical regression. We present the first large‑scale architectural ablation study on a global LiCSAR benchmark (20 frames, 39,724 patches, 651M pixels). Our results reveal a significant "complexity penalty": a vanilla U‑Net (7.76M parameters) achieves R^2=0.834 and RMSE = 1.01 cm, outperforming 11.37M‑parameter attention‑based models by 34% in R^2 and 51% in RMSE. Power Spectral Density (PSD) analysis provides the physical justification: while attention excels at capturing sharp semantic edges in natural images, it injects unphysical high‑frequency artifacts (>0.3 cycles/pixel) into geophysical fields, violating the fundamental smoothness constraints of elastic surface deformation. With a 2.92ms inference latency (a 2.5× speedup), the vanilla U‑Net is the only candidate to comfortably meet the sub‑100ms requirement for operational early‑warning systems. This work bridges the "publication‑to‑practice" gap by proving that convolutional locality outperforms modern complexity for smooth‑field regression, advocating for physics‑informed simplicity in ML4RS. Code available at https://github.com/prabhjotschugh/When‑Less‑is‑More‑InSAR‑Phase‑Unwrapping
Authors:Chamani Shiranthika, Parvaneh Saeedi
Abstract:
Federated learning enables collaborative model training across medical institutions without sharing raw data, but its performance is often limited by domain heterogeneity across clients. Existing approaches to address this challenge fall into two main paradigms: model‑side personalization, which adapts model parameters to each client, and data‑side harmonization, which reduces inter‑client variation at the input level. Despite their widespread use, these strategies have not been systematically compared. In this work, we conduct a comprehensive study across six medical imaging settings‑colon polyp, skin lesion, and breast tumor segmentation, and tuberculosis CXR, brain tumor, and breast tumor classification‑covering diverse types of domain shift. We evaluate a broad set of state‑of‑the‑art harmonization and personalization methods under a unified framework. Our results reveal a conditional trade‑off driven by the nature of heterogeneity: harmonization is more effective when variation is primarily appearance‑based (e.g., CXR classification), while personalization performs better when differences are structural (e.g., colon polyp segmentation). When inter‑client variation is limited, both strategies perform similarly. These findings demonstrate that the effectiveness of adaptation in federated medical imaging depends on the type and magnitude of domain shift rather than the strategy alone. We provide practical guidelines for selecting between harmonization and personalization and highlight directions for future hybrid approaches that combine both paradigms. Code is available at https://github.com/ChamaniS/WhenToAdapt.
Authors:Qi Li, Weining Wang, Shuangjun Du, Bo Peng, Jing Dong, Kun Wang, Zhenan Sun, Ming-Hsuan Yang
Abstract:
Face swapping has witnessed significant progress in recent years, largely driven by advances in deep generative models such as GANs and diffusion models.Despite these advances, existing methods remain fragmented across different paradigms, and their evaluation is highly inconsistent due to the lack of standardized datasets and protocols. Moreover, prior surveys primarily focus on broader deepfake generation or detection, leaving face swapping insufficiently studied as a standalone problem. In this paper, we present a comprehensive survey and benchmark for face swapping. We provide a structured review of existing methods, organizing them into five major paradigms and systematically analyzing their design principles, strengths, and limitations. To enable fair and controlled evaluation, we introduce CASIA FaceSwapping, a high‑quality benchmark with balanced demographic distributions and explicit attribute variations, and establish standardized protocols to assess the robustness of different face swapping methods. Extensive experiments on representative approaches yield new insights into the performance characteristics and limitations of current techniques. Overall, our work provides a unified perspective and a principled evaluation framework to facilitate the development of more robust and controllable face swapping methods. More results can be found at https://github.com/CASIA‑NLPRAI/face‑swapping‑survey.
Authors:George Stoica, Sayak Paul, Matthew Wallingford, Vivek Ramanujan, Abhay Nori, Winson Han, Ali Farhadi, Ranjay Krishna, Judy Hoffman
Abstract:
Flow matching (FM) trains a time‑dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high‑dimensional images, each training sample supervises only a single trajectory and intermediate point, yielding an extremely sparse and high‑variance training signal. This under‑constrained supervision can cause flow collapse, where the learned dynamics memorize specific source‑target pairings, mapping diverse inputs to overly similar outputs, failing to generalize. We introduce Posterior‑Augmented Flow Matching (PAFM), a theoretically grounded generalization of FM that replaces single‑target supervision with an expectation over an approximate posterior of valid target completions for a given intermediate state and condition. PAFM factorizes this intractable posterior into (i) the likelihood of the intermediate under a hypothesized endpoint and (ii) the prior probability of that endpoint under the condition, and uses an importance sampling scheme to construct a mixture over multiple candidate targets. We prove that PAFM yields an unbiased estimator of the original FM objective while substantially reducing gradient variance during training by aggregating information from many plausible continuation trajectories per intermediate. Finally, we show that PAFM improves over FM by up to 3.4 FID50K across different model scales (SiT‑B/2 and SiT‑XL/2), different architectures (SiT and MMDiT), and in both class and text conditioned benchmarks (ImageNet and CC12M), with a negligible increase in the compute overhead. Code: https://github.com/gstoica27/PAFM.git.
Authors:Jaeyoung Chung, Suyoung Lee, Jianfeng Xiang, Jiaolong Yang, Kyoung Mu Lee
Abstract:
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user‑defined segment maps of arbitrary shapes and scales, ensuring global‑scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine‑grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user‑controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.
Authors:Zhanjie Hu, Bolin Zhang, Jianhua Wang, Jianbo Zheng, Chenchen Yan, Takahiro Komamizu, Ichiro Ide, Jiangbo Qian
Abstract:
Temporal Video Grounding (TVG) aims to localize temporal moments in an untrimmed video that semantically correspond to given natural language queries. Recently, Graph Convolutional Networks (GCN) have been widely adopted in TVG to model temporal relations among video clips and enhance contextual reasoning by constructing clip‑level graphs. Despite their effectiveness, existing GCN‑based TVG methods encounter three critical bottlenecks: 1) Most methods construct graph nodes using either static or dynamic features alone, resulting in incomplete visual representation and overlooking complementary semantics, 2) Most methods construct temporal graphs in a query‑agnostic manner, leading to inefficient feature interaction within the temporal graph representation, and 3) Most methods often suffer from a single‑granularity semantic matching, while direct training on complex temporal localization task may lead to slow convergence and suboptimal precision. To address these challenges, we propose Static and Dynamic Graph Alignment Network (SDGAN). First, SDGAN jointly exploits static and dynamic visual features to construct two complementary temporal graphs and performs Position‑wise Nodes Alignment, enabling more expressive and robust visual representation. Second, SDGAN introduces Query‑Clip Contrastive Learning and Adaptive Graph Modeling to explicitly align visual clips with their corresponding textual queries, yielding query‑aware visual representations. Third, SDGAN incorporates multi‑granularity temporal proposals within Progressive Easy‑to‑Hard Training Strategy, effectively bridging coarse‑grained semantic localization and fine‑grained temporal boundary refinement. Extensive experiments on three benchmark datasets demonstrate that SDGAN achieves superior performance across complex TVG scenarios. Codes and datasets are available at https://github.com/ZhanJieHu/SDGAN.
Authors:Jaeyoung Chung, Suyoung Lee, Kyoung Mu Lee
Abstract:
We present a training‑free approach for controllable 3D inpainting based on initial noise optimization. In the structured 3D latent diffusion framework, we observe that the underlying geometric structure is established during the early stages of the diffusion process and exhibits high sensitivity to the initial noise. Such characteristics compromise stability in tasks like inpainting and editing, where the model must ensure strict alignment with the existing context while synthesizing a new structure. In this paper, we introduce a strategy to optimize the initial noise within the structured 3D latent diffusion framework, ensuring high‑fidelity 3D inpainting. Specifically, we update the initial noise by leveraging a backpropagation approximation grounded in the rectified flow model, with the spectral parameterization specially designed for robust and efficient structured 3D latent optimization. Experiments demonstrate consistent improvements in contextual consistency and prompt alignment over representative training‑free inpainting baselines, establishing initial noise control as an independent dimension for 3D inpainting, orthogonal to conventional sampling trajectory manipulation.
Authors:Haojian Huang, Jiahao Shi, Yinchuan Li, Yingcong Chen
Abstract:
Affordance grounding requires identifying where and how an agent should interact in open‑world scenes, where actionable regions are often small, occluded, reflective, and visually ambiguous. Recent systems therefore combine multiple skills (e.g., detection, segmentation, interaction‑imagination), yet most orchestrate them with fixed pipelines that are poorly matched to per‑instance difficulty, offer limited targeted recovery from intermediate errors, and fail to reuse experience from recurring objects. These failures expose a systems problem: test‑time grounding must acquire the right evidence, decide whether that evidence is reliable enough to commit, and do so under bounded inference cost without access to labels. We propose Affordance Agent Harness, a closed‑loop runtime that unifies heterogeneous skills with an evidence store and cost control, retrieves episodic memories to provide priors for recurring categories, and employs a Router to adaptively select and parameterize skills. An affordance‑specific Verifier then gates commitments using self‑consistency, cross‑scale stability, and evidence sufficiency, triggering targeted retries before a final judge fuses accumulated evidence and trajectories into the prediction. Experiments on multiple affordance benchmarks and difficulty‑controlled subsets show a stronger accuracy‑cost Pareto frontier than fixed‑pipeline baselines, improving grounding quality while reducing average skill calls and latency. Project page: https://tenplusgood.github.io/a‑harness‑page/.
Authors:Houyuan Chen, Hong Li, Xianghao Kong, Tianrui Zhu, Shaocong Xu, Weiqing Xiao, Yuwei Guo, Chongjie Ye, Lvmin Zhang, Hao Zhao, Anyi Rao
Abstract:
Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often train separate models for each problem setting, which fixes the input‑output mapping and limits the modeling of correlations across modalities. We present UniVidX, a unified multimodal framework that leverages VDM priors for versatile video generation. UniVidX formulates pixel‑aligned tasks as conditional generation in a shared multimodal space, adapts to modality‑specific distributions while preserving the backbone's native priors, and promotes cross‑modal consistency during synthesis. It is built on three key designs. Stochastic Condition Masking (SCM) randomly partitions modalities into clean conditions and noisy targets during training, enabling omni‑directional conditional generation instead of fixed mappings. Decoupled Gated LoRA (DGL) introduces per‑modality LoRAs that are activated when a modality serves as the generation target, preserving the strong priors of the VDM. Cross‑Modal Self‑Attention (CMSA) shares keys and values across modalities while keeping modality‑specific queries, facilitating information exchange and inter‑modal alignment. We instantiate UniVidX in two domains: UniVid‑Intrinsic, for RGB videos and intrinsic maps including albedo, irradiance, and normal; and UniVid‑Alpha, for blended RGB videos and their constituent RGBA layers. Experiments show that both models achieve performance competitive with state‑of‑the‑art methods across distinct tasks and generalize robustly to in‑the‑wild scenarios, even when trained on fewer than 1,000 videos. Project page: https://houyuanchen111.github.io/UniVidX.github.io/
Authors:Yan Zhang, Daiqing Wu, Huawen Shen, Can Ma, Yu Zhou
Abstract:
Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capability for autonomous GUI agents. Recent reinforcement learning methods (e.g., GRPO) have achieved strong performance, but they rely on expensive multiple rollouts and suffer from sparse signals on hard samples. These limitations make on‑policy self‑distillation (OPSD), which provides dense token‑level supervision from a single rollout, a promising alternative. However, its applicability to GUI grounding remains unexplored. In this paper, we present GUI‑SD, the first OPSD framework tailored for GUI grounding. First, it constructs a visually enriched privileged context for the teacher using a target bounding box and a Gaussian soft mask, providing informative guidance without leaking exact coordinates. Second, it employs entropy‑guided distillation, which adaptively weights tokens based on digit significance and teacher confidence, concentrating optimization on the most impactful and reliable positions. Extensive experiments on six representative GUI grounding benchmarks show that GUI‑SD consistently outperforms GRPO‑based methods and naive OPSD in both accuracy and training efficiency. Code and training data are available at https://zhangyan‑ucas.github.io/GUI‑SD/.
Authors:Massimo Rondelli, Francesco Pivi, Maurizio Gabbrielli
Abstract:
Automatic generation of executable Blender code from natural language remains challenging, with state‑of‑the‑art LLMs producing frequent syntactic errors and geometrically inconsistent objects. We present BlenderRAG, a retrieval‑augmented generation system that operates on a curated multimodal dataset of 500 expert‑validated examples (text, code, image) across 50 object categories. By retrieving semantically similar examples during generation, BlenderRAG improves compilation success rates from 40.8% to 70.0% and semantic normalized alignment from 0.41 to 0.77 (CLIP similarity) across four state‑of‑the‑art LLMs, without requiring fine‑tuning or specialized hardware, making it immediately accessible for deployment. The dataset and code will be available at https://github.com/MaxRondelli/BlenderRAG.
Authors:Hang Wang, Chao Shen, Chenhao Lin, Minghui Yang, Lei Zhang, Cong Wang
Abstract:
The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI‑generated video (AIGV) detection methods primarily focus on uni‑modal or spatiotemporal artifacts, but they overlook the rich cues within the visual‑textual cross‑modal space, especially the temporal stability of semantic alignment. In this work, we identify a distinctive fingerprint in AIGVs, termed cross‑modal temporal artifact (CMTA). Unlike real videos that exhibit natural temporal fluctuations in cross‑modal alignment due to semantic variations, AIGVs display unnaturally stable semantic trajectories governed by given input prompts. To bridge this gap, we propose the CMTA framework, a cross‑modal detection approach that captures these unique temporal artifacts through joint cross‑modal embedding and multi‑grained temporal modeling. Specifically, CMTA leverages BLIP to generate frame‑level image captions and utilizes CLIP to extract corresponding visual‑textual representations. A coarse‑grained temporal modeling branch is then designed to characterize temporal fluctuations in cross‑modal alignment with a GRU. In parallel, a fine‑grained branch is constructed to capture intricate inter‑frame variations from integrated visual‑textual features with a Transformer encoder. Extensive experiments on 40 subsets across four large‑scale datasets, including GenVideo, EvalCrafter, VideoPhy, and VidProM, validate that our approach sets a new state‑of‑the‑art while exhibiting superior cross‑generator generalization. Code and models of CMTA will be released at https://github.com/hwang‑cs‑ime/CMTA
Authors:Hao Wei, Yanhui Zhou, Chenyang Ge, Saeed Anwar, Ajmal Mian
Abstract:
Most recent extreme rescaling methods struggle to preserve semantically consistent structures and produce realistic details, due to the severely ill‑posed nature of low‑ to high‑resolution mapping under scaling factors of 16× or higher. To alleviate the above problems, we propose FaithEIR, a diffusion‑based framework for extreme image rescaling. Inspired by singular value decomposition, we develop learnable reversible transformation that enables invertible downscaling and upscaling in the latent space. To compensate for information loss due to quantization, we propose an adaptive detail prior, a high‑frequency dictionary that captures the empirical average of commonly occurring structures in the training data. Finally, we design a lightweight pixel semantic embedder to provide semantic conditioning for the pretrained diffusion model. We present extensive experimental results demonstrating that our FaithEIR consistently outperforms state‑of‑the‑art methods, achieving superior reconstruction fidelity and perceptual quality. Our code, model weights, and detailed results are released at https://github.com/cshw2021/FaithEIR.
Authors:Nadav Z. Cohen, Ofir Abramovich, Ariel Shamir
Abstract:
Text‑to‑image diffusion models generate images by gradually converting white Gaussian noise into a natural image. White Gaussian noise is well suited for producing diverse outputs from a single text prompt due to its absence of structure. However, this very property limits control over, and predictability of, specific visual attributes, as the noise is not human‑interpretable. In this work, we investigate the characteristics of the input noise in diffusion models. We show that, although all frequencies in white Gaussian noise have comparable statistical energy, low‑frequency components primarily determine the images global structure and color composition, while high‑frequency components control finer details. Building on this observation, we demonstrate that simple manipulations of the low‑frequency noise using low‑frequency image priors can effectively condition the generation process to reconstruct these low‑frequency visual cues. This allows us to define a simple, training‑free method with minimal overhead that steers overall image structure and color, while letting high‑frequency components freely emerge as fine details, enabling variability across generated outputs.
Authors:Yonghao Zhao, Yupeng Gao, Jian Yang, Jin Xie, Beibei Wang
Abstract:
Recent advances in Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have made it standard practice to reconstruct 3D scenes from multi‑view images. Removing objects from such 3D representations is a fundamental editing task that requires complete and seamless inpainting of occluded regions, ensuring consistency in geometry and appearance. Although existing methods have made notable progress in improving inpainting consistency, they often neglect global lighting effects, leading to physically implausible results. Moreover, these methods struggle with view‑dependent non‑Lambertian surfaces, where appearance varies across viewpoints, leading to unreliable inpainting. In this paper, we present 3D Gaussian Object Removal in the Intrinsic Space (GOR‑IS), a novel framework for physically consistent and visually coherent 3D object removal. Our approach decomposes the scene into intrinsic components and explicitly models light transport to maintain global lighting effects consistency. Furthermore, we introduce an intrinsic‑space inpainting module that operates directly in the material and lighting domains, effectively addressing the challenges posed by non‑Lambertian surfaces. Extensive experiments on both synthetic and real‑world datasets demonstrate that our framework substantially improves the physical consistency and visual coherence of object removal, outperforming existing methods by 13% in perceptual similarity (LPIPS) and 2dB in peak signal‑to‑noise ratio (PSNR). Code is publicly available at https://applezyh.github.io/GOR‑IS‑project‑page/
Authors:Huangbiao Xu, Huanqi Wu, Xiao Ke, Yuxin Peng
Abstract:
Real‑world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods typically rely on the unrealistic assumption of full‑modal availability during training to provide reconstruction supervision or cross‑modal priors. This paper tackles the more challenging setting of IML under training‑time incomplete observations, which precludes reliance on a ``God's eye view'' of complete data. We propose LIMSSR (LLM‑Driven Incomplete Multimodal Sequence‑to‑Score Reasoning), a framework that reformulates this challenge as a conditional sequence reasoning task. LIMSSR leverages the semantic reasoning capabilities of Large Language Models via Prompt‑Guided Context‑Aware Modality Imputation and Multidimensional Representation Fusion to infer latent semantics from available contexts without direct reconstruction. To mitigate hallucinations, we introduce a Mask‑Aware Dual‑Path Aggregation to dynamically calibrate inference uncertainty. Extensive experiments on three Action Quality Assessment datasets demonstrate that LIMSSR significantly outperforms state‑of‑the‑art baselines without relying on complete training data, establishing a new paradigm for data‑efficient multimodal learning. Code is available at https://github.com/XuHuangbiao/LIMSSR.
Authors:Zhenhua Ning, Xin Li, Jun Yu, Guangming Lu, Yaowei Wang, Wenjie Pei
Abstract:
While 3D Gaussian Splatting (3DGS) has demonstrated impressive real‑time rendering performance, its efficacy remains constrained by a reliance on heuristic density control. Despite numerous refinements to these handcrafted rules, such methods inherently lack the flexibility to adapt to diverse scenes with complex geometries.
In this paper, we propose a paradigm shift for density control from rigid heuristics to fully learnable policies. Specifically, we introduce LeGS, a framework that reformulates density control as a parameterized policy network optimized via Reinforcement Learning (RL). Central to our approach is the tailored effective reward function grounded in sensitivity analysis, which precisely quantifies the marginal contribution of individual Gaussians to reconstruction quality. To maintain computational tractability, we derive a closed‑form solution that reduces the complexity of reward calculation from O(N^2) to O(N). Extensive experiments on the Mip‑NeRF 360, Tanks \& Temples, and Deep Blending datasets demonstrate that LeGS significantly outperforms state‑of‑the‑art methods, striking a superior balance between reconstruction quality and efficiency. The code will be released at https://github.com/AaronNZH/LeGS
Authors:Kang Yang, Tianci Bu, Peng Wang, Deying Li, Yongcai Wang
Abstract:
Most existing heterogeneous cooperative perception methods depend on prior preparation like offline joint training or tailored collaborator‑model adaptation. Such preprocessing is, however, generally impractical in real scenarios, as agents are usually independently trained by different developers and meet occasionally online. This work investigates \emphpreparation‑free heterogeneous cooperative perception, where agents use independently trained single‑agent detectors without any pre‑deployment coordination. We find direct cross‑agent fusion under this setting greatly underperforms ego‑only perception. We present BOLT, a lightweight plug‑and‑play module that adapts neighboring features online via ego‑as‑teacher distillation, requiring only ego predictions without ground‑truth labels. BOLT leverages high‑confidence ego perception features to guide cross‑agent feature‑domain alignment, while enabling neighbors to contribute features in the ego's low‑confidence regions. With only 0.9M trainable parameters, BOLT improves AP@50 by up to 32.3 points over vanilla unadapted fusion in the preparation‑free setting. It consistently outperforms ego‑only results on DAIR‑V2X and OPV2V, across different encoder pairs and fusion strategies. Code: https://github.com/sidiangongyuan/BOLT.
Authors:YuSheng Lin, Ji-Hwa Tsai, Chun-Shu Wei
Abstract:
Recent EEG‑to‑image retrieval methods leverage pretrained vision encoders and foveation‑inspired priors, but typically assume a fixed, center‑focused view. This center bias conflicts with content‑driven human attention, creating a geometric‑semantic dissociation between visual features and EEG responses. We propose SIMON, a saliency‑aware multi‑view framework for zero‑shot EEG‑to‑image retrieval. SIMON combines foreground segmentation and saliency prediction to select fixation centers via Saliency‑Aware Sampling (SAS), then generates foveated views that emphasize informative object regions while suppressing background clutter. On THINGS‑EEG, SIMON achieves state‑of‑the‑art performance in both intra‑subject and inter‑subject settings, reaching an average Top‑1 accuracy of 69.7% and 19.6%, respectively, consistently outperforming recent competitive baselines. Analyses across sampling granularity, EEG channel topology, and visual/brain encoder backbones further support the robustness of saliency‑aware multi‑view integration. Our code and models are publicly available at https://github.com/simonlink666/SIMON.
Authors:Wei Liu, Hongkai Liu, Zhiying Deng, Yee Whye Teh, Wee Sun Lee
Abstract:
LLM parameter editing methods commonly rely on computing an ideal target hidden‑state at a target layer (referred as anchor point) and distributing the target vector to multiple preceding layers (commonly known as backward spreading) for cooperative editing. Although widely used for a long time, its underlying basis have not been systematically investigated. In this paper, we first conduct a systematic study of its foundations, which helps clarify its capability boundaries, practical considerations, and potential failure modes. Then, we propose a simple and elegant alternative that replaces backward spreading with forward‑propagation. Instead of optimizing the target at the last editing layer, we optimize the anchor point at the first editing layer, and then propagate it forward to obtain accurate and mutually compatible target hidden‑states for all subsequent editing layers. This approach achieves the same computational complexity as existing methods while producing more accurate layer‑wise targets. Our method is simple, without interfering with either the computation of the initial target hidden state or any other components of the subsequent editing pipeline, and thus constituting a benefit for a wide range of LLM parameter editing methods.
Authors:Yuhui Lu, Wenjing Liu, Kun Zhan
Abstract:
Standard diffusion models for graph generation typically rely on uniform time‑stepping, an approach that overlooks the non‑homogeneous dynamics of distributional evolution on complex manifolds. In this paper, we present an information‑geometric framework that reinterprets the diffusion sampling trajectory as a parametric curve on a Riemannian manifold. Our key observation is that the Fisher‑Rao metric provides a principled measure of the intrinsic distance. By analyzing this metric, we derive the Drift Variation Score (DVS), a geometry‑aware indicator that quantifies the instantaneous rate of distributional change. Unlike prior heuristic‑based adaptive samplers, our DVS solver enforces a constant informational speed on the statistical manifold, automatically maintaining a uniform rate of distributional change along the sampling trajectory. This equal arc‑length strategy ensures that each discretization step contributes equally to the information speed. Theoretical analysis verifies that DVS characterizes the local stiffness of the sampling dynamics in the Fisher‑Rao sense. Experimental results on molecule and social network generation show that DVS significantly improves structural fidelity and sampling efficiency. Code is at https://github.com/kunzhan/DVS
Authors:Jingxiang Chen, Mohamed Ibrahim, Yang Liu
Abstract:
We present VkSplat, a high‑performance, cross‑vendor 3D Gaussian Splatting (3DGS) training pipeline implemented fully in Vulkan compute, addressing performance and compatibility limitation of existing training pipelines. With various optimizations, we achieve 3.3× speed and 33% VRAM reduction over CUDA+PyTorch baseline, maintaining quality, and demonstrating compatibility across GPU vendors. To the best of our knowledge, this is the first fully‑Vulkan‑based 3DGS training pipeline that achieves state‑of‑the‑art performance. Code: \hrefhttps://github.com/harry7557558/vksplathttps://github.com/harry7557558/vksplat
Authors:Qianfan Shen, Ningxiao Tao, Qiyu Dai, Tianle Chen, Minghan Qin, Yongjie Zhang, Mengyu Chu, Wenzheng Chen, Baoquan Chen
Abstract:
We consider the problem of synthesizing photorealistic, physically plausible combustion effects in in‑the‑wild 3D scenes. Traditional CFD and graphics pipelines can produce realistic fire effects but rely on handcrafted geometry, expert‑tuned parameters, and labor‑intensive workflows, limiting their scalability to the real world. Recent scene modeling advances like 3D Gaussian Splatting (3DGS) enable high‑fidelity real‑world scene reconstruction, yet lack physical grounding for combustion. To bridge this gap, we propose FieryGS, a physically‑based framework that integrates physically‑accurate and user‑controllable combustion simulation and rendering within the 3DGS pipeline, enabling realistic fire synthesis for real scenes. Our approach tightly couples three key modules: (1) multimodal large‑language‑model‑based physical material reasoning, (2) efficient volumetric combustion simulation, and (3) a unified renderer for fire and 3DGS. By unifying reconstruction, physical reasoning, simulation, and rendering, FieryGS removes manual tuning and automatically generates realistic, controllable fire dynamics consistent with scene geometry and materials. Our framework supports complex combustion phenomena ‑‑ including flame propagation, smoke dispersion, and surface carbonization ‑‑ with precise user control over fire intensity, airflow, ignition location and other combustion parameters. Evaluated on diverse indoor and outdoor scenes, FieryGS outperforms all comparative baselines in visual realism, physical fidelity, and controllability. Project page can be found at https://pku‑vcl‑geometry.github.io/FieryGS/.
Authors:YiFeng Wang, Zhun Sun, Keisuke Sakaguchi
Abstract:
We present Activation Residual Hessian Quantization (ARHQ), a post‑training weight splitting method designed to mitigate error propagation in low‑bit activation‑weight quantization. By constructing an input‑side residual Hessian from activation quantization residuals (G_x), ARHQ analytically identifies and isolates error‑sensitive weight directions into a high‑precision low‑rank branch. This is achieved via a closed‑form truncated SVD on the scaled weight matrix W G^1/2_x . Experimental results on Qwen3‑4B‑Thinking‑2507 demonstrate that ARHQ significantly improves layer‑wise SNR and preserves downstream reasoning performance on ZebraLogic even under aggressive quantization. The code is available at https://github.com/BeautMoonQ/ARHQ.
Authors:Junyoung Lee, Sookwan Han, Jeonghwan Kim, Inhee Lee, Mingi Choi, Jisoo Kim, Wonjung Woo, Hanbyul Joo
Abstract:
Human‑robot collaboration has been studied primarily in dyadic or sequential settings. However, real homes require multiadic collaboration, where multiple humans and robots share a workspace, acting concurrently on interleaved subtasks with tight spatial and temporal coupling. This regime remains underexplored because close‑proximity interaction between humans, robots, and objects creates persistent occlusion and rapid state changes, making reliable real‑time 3D tracking the central bottleneck. No existing platform provides the real‑time, occlusion‑robust, room‑scale perception needed to make this regime experimentally tractable. We present OmniRobotHome, the first room‑scale residential platform that unifies wide‑area real‑time 3D human and object perception with coordinated multi‑robot actuation in a shared world frame. The system instruments a natural home environment with 48 hardware‑synchronized RGB cameras for markerless, occlusion‑robust tracking of multiple humans and objects, temporally aligned with two Franka arms that act on live scene state. Continuous capture within this consistent frame further supports long‑horizon human behavior modeling from accumulated trajectories. The platform makes the multiadic collaboration regime experimentally tractable. We focus on two central problems: safety in shared human‑robot environments and human‑anticipatory robotic assistance, and show that real‑time perception and accumulated behavior memory each yield measurable gains in both.
Authors:Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai
Abstract:
Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene generation, often overlooking comprehensive 3D scene understanding. Conversely, while Large Language Models (LLMs) demonstrate impressive reasoning capabilities, they lack the capacity to predict future geometric evolution, creating a significant disparity between semantic interpretation and physical simulation. To bridge this gap, we propose HERMES++, a unified driving world model that integrates 3D scene understanding and future geometry prediction within a single framework. Our approach addresses the distinct requirements of these tasks through synergistic designs. First, a BEV representation consolidates multi‑view spatial information into a structure compatible with LLMs. Second, we introduce LLM‑enhanced world queries to facilitate knowledge transfer from the understanding branch. Third, a Current‑to‑Future Link is designed to bridge the temporal gap, conditioning geometric evolution on semantic context. Finally, to enforce structural integrity, we employ a Joint Geometric Optimization strategy that integrates explicit geometric constraints with implicit latent regularization to align internal representations with geometry‑aware priors. Extensive evaluations on multiple benchmarks validate the effectiveness of our method. HERMES++ achieves strong performance, outperforming specialist approaches in both future point cloud prediction and 3D scene understanding tasks. The model and code will be publicly released at https://github.com/H‑EmbodVis/HERMESV2.
Authors:Jiawei Yang, Zhengyang Geng, Xuan Ju, Yonglong Tian, Yue Wang
Abstract:
We show that Fréchet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation space. Our idea is simple: decouple the population size for FD estimation (e.g., 50k) from the batch size for gradient computation (e.g., 1024). We term this approach FD‑loss. Optimizing FD‑loss reveals several surprising findings. First, post‑training a base generator with FD‑loss in different representation spaces consistently improves visual quality. Under the Inception feature space, a one‑step generator achieves0.72 FID on ImageNet 256x256. Second, the same FD‑loss repurposes multi‑step generators into strong one‑step generators without teacher distillation, adversarial training or per‑sample targets. Third, FID can misrank visual quality: modern representations can yield better samples despite worse Inception FID. This motivates FDr^k, a multi‑representation metric. We hope this work will encourage further exploration of distributional distances in diverse representation spaces as both training objectives and evaluation metrics for generative models.
Authors:Keming Wu, Zuhao Yang, Kaichen Zhang, Shizun Wang, Haowei Zhu, Sicong Leng, Zhongyu Yang, Qijie Wang, Sudong Wang, Ziting Wang, Zili Wang, Hui Zhang, Haonan Wang, Hang Zhou, Yifan Pu, Xingxuan Li, Fangneng Zhan, Bo Li, Lidong Bing, Yuxin Song, Ziwei Liu, Wenhu Chen, Jingdong Wang, Xinchao Wang, Xiaojuan Qi, Shijian Lu, Bin Wang
Abstract:
Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still struggle with spatial reasoning, persistent state, long‑horizon consistency, and causal understanding. We argue that the field should move beyond appearance synthesis toward intelligent visual generation: plausible visuals grounded in structure, dynamics, domain knowledge, and causal relations. To frame this shift, we introduce a five‑level taxonomy: Atomic Generation, Conditional Generation, In‑Context Generation, Agentic Generation, and World‑Modeling Generation, progressing from passive renderers to interactive, agentic, world‑aware generators. We analyze key technical drivers, including flow matching, unified understanding‑and‑generation models, improved visual representations, post‑training, reward modeling, data curation, synthetic data distillation, and sampling acceleration. We further show that current evaluations often overestimate progress by emphasizing perceptual quality while missing structural, temporal, and causal failures. By combining benchmark review, in‑the‑wild stress tests, and expert‑constrained case studies, this roadmap offers a capability‑centered lens for understanding, evaluating, and advancing the next generation of intelligent visual generation systems.
Authors:Andrea Dunn Beltran, Daniel Rho, Aarav Mehta, Xinqi Xiong, Raúl San José Estépar, Ron Alterovitz, Marc Niethammer, Roni Sengupta
Abstract:
Bronchoscopic navigation relies on registering endoscopic video to a preoperative CT scan, but respiratory motion deforms the airway by 5‑20 mm, creating CT‑to‑body divergence that limits localization accuracy. In practice, this is mitigated through breath‑hold protocols, which attempt to match the intraoperative anatomy to a static CT, but are difficult to reproduce and disrupt clinical workflow. We propose to eliminate the need for breath‑hold protocols by leveraging patient‑specific respiratory modeling. Paired inhale‑exhale CT scans, already acquired for planning, implicitly define the patient‑specific deformation space of the breathing airway. By registering these scans, we reduce respiratory motion to a single scalar breathing phase per frame, constraining all reconstructions to anatomically observed configurations. We embed this representation within a mesh‑anchored Gaussian splatting framework, where a lightweight estimator infers breathing phase directly from endoscopic RGB, enabling continuous, deformation‑aware reconstruction throughout the respiratory cycle without breath‑holds or external sensing. To enable quantitative evaluation, we introduce RESPIRE, a physically grounded bronchoscopy simulation pipeline with per‑frame ground truth for geometry, pose, breathing phase, and deformation. Experiments on RESPIRE show that our approach achieves geometrically faithful reconstruction, over 20x faster training, and 1.22 mm target localization accuracy (within the 3mm clinically relevant tolerances) outperforming unconstrained single‑CT baselines. Please check out our website for additional visuals: https://asdunnbe.github.io/RESPIRE/
Authors:Wenxiao Li, Faqiang Wang, Yuping Duan, Li Cui, Liqiang Zhang, Jun Liu
Abstract:
Topological features play an essential role in ensuring geometric plausibility and structural consistency in image analysis tasks such as segmentation and skeletonization. However, integrating topology‑preserving learning based on simple points into deep learning tasks remains challenging, as existing simple point detection methods are confined to binary images and are non‑differentiable, rendering them incompatible with gradient‑based optimization in modern deep learning. Moreover, morphological and purely data‑driven approaches often fail to guaranty topological consistency. To address these limitations, we propose a novel method that directly computes simple points on continuous‑valued images, enabling differentiable topological inference. Building on this theory, we develop an efficient skeleton extraction algorithm that preserves topological structures in binary and continuous‑valued images. Furthermore, we design a variational model that enforces topological constraints by preserving topologically non‑removable (i.e., non‑simple) points, which can be seamlessly integrated into any deep neural network segmentation with softmax or sigmoid outputs. Experimental results demonstrate that the proposed approach effectively improves topological integrity and structural accuracy across multiple benchmarks. The codes are available in https://github.com/levnsio/CSP.
Authors:Geon Yeong Park, Roman Shapovalov, Rakesh Ranjan, Jong Chul Ye, Andrea Vedaldi, Thu Nguyen-Phuoc
Abstract:
We consider the problem of regenerating 3D objects from 2D images and initial 3D shapes. Most 3D generators operate in a one‑shot fashion, converting text or images to a 3D object with limited controllability. We introduce instead MeshReGen, a 3D regenerator that is conditioned on an initial 3D shape. This conceptually simple formulation allows us to support numerous useful tasks, including 3D enhancement, reconstruction, and editing. MeshReGen uses a new conditioning mechanism based on VecSet, which allows the regenerator to update or improve the input geometry with consistent fine‑grained details. MeshReGen learns a widely applicable regeneration prior from off‑the‑shelf 3D datasets via self‑supervised pretext tasks and augmentations, without additional annotations. We evaluate both the geometric consistency and fine‑grained quality of MeshReGen, achieving state‑of‑the‑art performance in controllable 3D generation across several tasks.
Authors:Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
Abstract:
Recent methods for arbitrary‑skeleton motion capture from monocular video follow a factorized pipeline, where a Video‑to‑Pose network predicts joint positions and an analytical inverse‑kinematics (IK) stage recovers joint rotations. While effective, this design is inherently limited, since joint positions do not fully determine rotations and leave degrees of freedom such as bone‑axis twist ambiguous, and the non‑differentiable IK stage prevents the system from adapting to noisy predictions or optimizing for the final animation objective. In this work, we present the first fully end‑to‑end framework in which both Video‑to‑Pose and Pose‑to‑Rotation are learnable and jointly optimized. We observe that the ambiguity in pose‑to‑rotation mapping arises from missing coordinate system information: the same joint positions can correspond to different rotations under different rest poses and local axis conventions. To resolve this, we introduce a reference pose‑rotation pair from the target asset, which, together with the rest pose, not only anchors the mapping but also defines the underlying rotation coordinate system. This formulation turns rotation prediction into a well‑constrained conditional problem and enables effective learning. In addition, our model predicts joint positions directly from video without relying on mesh intermediates, improving both robustness and efficiency. Both stages share a skeleton‑aware Global‑Local Graph‑guided Multi‑Head Attention (GL‑GMHA) module for joint‑level local reasoning and global coordination. Experiments on Truebones Zoo and Objaverse show that our method reduces rotation error from ~17 degrees to ~10 degrees, and to 6.54 degrees on unseen skeletons, while achieving ~20x faster inference than mesh‑based pipelines. Project page: https://animotionlab.github.io/MoCapAnythingV2/
Authors:Sudong Wang, Weiquan Huang, Xiaomin Yu, Zuhao Yang, Hehai Lin, Keming Wu, Chaojun Xiao, Chen Chen, Wenxuan Wang, Beier Zhu, Yunjian Zhang, Chengwei Qin
Abstract:
The standard post‑training recipe for large multimodal models (LMMs) applies supervised fine‑tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR). However, SFT introduces distributional drift that neither preserves the model's original capabilities nor faithfully matches the supervision distribution. This problem is further amplified in multimodal reasoning, where perception errors and reasoning failures follow distinct drift patterns that compound during subsequent RL. We introduce PRISM, a three‑stage pipeline that mitigates this drift by inserting an explicit distribution‑alignment stage between SFT and RLVR. Building on the principle of on‑policy distillation (OPD), PRISM casts alignment as a black‑box, response‑level adversarial game between the policy and a Mixture‑of‑Experts (MoE) discriminator with dedicated perception and reasoning experts, providing disentangled corrective signals that steer the policy toward the supervision distribution without requiring access to teacher logits. While 1.26M public demonstrations suffice for broad SFT initialization, distribution alignment demands higher‑fidelity supervision; we therefore curate 113K additional demonstrations from Gemini 3 Flash, featuring dense visual grounding and step‑by‑step reasoning on the hardest unsolved problems. Experiments on Qwen3‑VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the SFT‑to‑RLVR baseline on 4B and 8B, respectively. Our code, data, and model checkpoints are publicly available at https://github.com/XIAO4579/PRISM.
Authors:Zeyu Jiang, Changqing Zhou, Xingxing Zuo, Changhao Chen
Abstract:
Existing learning‑based occupancy prediction methods rely on large‑scale 3D annotations and generalize poorly across environments. We present FreeOcc, a training‑free framework for open‑vocabulary occupancy prediction from monocular or RGB‑D sequences. Unlike prior approaches that require voxel‑level supervision and ground‑truth camera poses, FreeOcc operates without 3D annotations, pose ground truth, or any learning stage. FreeOcc incrementally builds a globally consistent occupancy map via a four‑layer pipeline: a SLAM backbone estimates poses and sparse geometry; a geometrically consistent Gaussian update constructs dense 3D Gaussian maps; open‑vocabulary semantics from off‑the‑shelf vision‑language models are associated with Gaussian primitives; and a probabilistic Gaussian‑to‑occupancy projection produces dense voxel occupancy. Despite being entirely training‑free and pose‑agnostic, FreeOcc achieves over 2× improvements in IoU and mIoU on EmbodiedOcc‑ScanNet compared to prior self‑supervised methods. We further introduce ReplicaOcc, a benchmark for indoor open‑vocabulary occupancy prediction, and show that FreeOcc transfers zero‑shot to novel environments, substantially outperforming both supervised and self‑supervised baselines. Project page: https://the‑masses.github.io/freeocc‑web/.
Authors:Shuokun Cheng, Jinghao Shi, Kun Sun
Abstract:
Accurate lesion segmentation is crucial for clinical diagnosis and treatment planning. However, lesions often resemble surrounding tissues and exhibit ill‑defined boundaries, leading to unstable predictions in boundary/transition regions. Moreover, small‑lesion cues can be diluted by multi‑scale feature extraction, causing under‑ or over‑segmentation. To address these challenges, we propose an Uncertainty‑Aware Hypergraph Refinement Network (UHR‑Net). First, we introduce an Uncertainty‑Oriented Instance Contrastive (UO‑IC) pretraining strategy that couples geometry‑aware copy‑paste augmentation with hard‑negative mining of lesion‑like background regions to improve instance‑level discrimination for small and visually ambiguous lesions. Second, we design an Uncertainty‑Guided Hypergraph Refinement (UGHR) block, which derives an entropy‑based uncertainty map from a coarse probability map to guide hypergraph refinement. By splitting hyperedge prototypes into foreground and background groups, UGHR decouples higher‑order interactions and improves refinement in ambiguous regions. Experiments on five public benchmarks demonstrate consistent gains over strong baselines. Code is available at: https://github.com/CUGfreshman/UHR‑Net.
Authors:Jiaying Ying, Heming Du, Kaihao Zhang, Sean M. Tweedy, Xin Yu
Abstract:
Single‑image human mesh recovery provides a compact 3D, person‑centric representation that supports analysis, animation, AR and VR, rehabilitation, and human‑computer interaction. However, prevailing systems impose an intact‑limb prior and degrade on people with limb loss, because fixed‑topology models cannot represent residual limbs. In this work, we present ResiHMR, a residual‑limb aware framework for single‑image 3D human modeling. ResiHMR adopts residual‑limb keypoints and introduces two components: (i) a topology‑adaptive Residual Anchor‑Factor Optimization module that constrains estimation to the observed kinematic subgraph of anatomically valid structures, and (ii) a geometry‑based Residual‑Limb Reconstruction module that estimates residual‑limb boundaries and convex limb‑termination geometry. These components introduce topology‑aware optimization and explicit termination geometry as tools for human mesh recovery under non‑standard limb anatomy. Unlike joint‑removal methods in a fixed topology, ResiHMR explicitly reconstructs residual‑limb surfaces and aligns optimization with limb‑loss topology, which better matches prosthetic biomechanics and real‑world use. To the best of our knowledge, this is the first single‑image HMR system that explicitly reconstructs residual‑limb surfaces and performs topology‑adaptive optimization for individuals with limb loss. On a curated dataset of real‑world images with limb loss, ResiHMR improves reconstruction quality under both SMPLify‑X and HSMR backbones, reducing intact‑joint 2D MPJPE from 41.32 to 37.40 with SMPLify‑X and residual‑limb 2D MPJPE from 73.61 to 23.19 with HSMR.
Authors:Sharayu Nilesh Deshmukh, Kailash A. Hambarde, Joana C. Costa, Hugo Proença, Tiago Roxo
Abstract:
Current DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured in binary settings. Four‑class audio‑visual formulations address this by discriminating manipulation type, but introduce a unresolved problem: models may rely solely on data source integrity to detect DeepFakes without evaluating their semantic consistency. If the DeepFake origin is not in the data source but in its content, can semantic mismatch be assessed by the state‑of‑the‑art? This paper proposes a new evaluation setup, extending the four‑class formulation by explicitly modeling semantic‑level inconsistency between authentic modalities with the introduction a new class: Real Audio‑Real Video with Semantic Mismatch (RARV‑SMM). We assess the robustness of state‑of‑the‑art models in this new realistic DeepFake setting, using the FakeAVCeleb dataset, highlighting the limitations of existing approaches when faced with semantic mismatch data. We further introduce three RARV‑SMM variants that expose distinct architectural vulnerabilities as audio‑visual divergence increases. We also propose a semantic reinforcement strategy that incorporates the semantic mismatch class and ImageBind embeddings to improve DeepFake detection in both our proposed and state‑of‑the‑art settings, on FakeAVCeleb and LAV‑DF, paving the way to more realistic DeepFake detectors. The source code and data are available at https://github.com/.
Authors:Jing Zhang, Wentao Jiang, Tao Huang, Zhiwei Wang, Jianxin Liu, Jian Chen, Ping Ye, Gang Wang, Zengmao Wang, Bo Du, Dacheng Tao
Abstract:
Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: specialized detectors offer strong localization but limited reasoning, whereas multimodal large language models (MLLMs) provide flexible reasoning but weak grounding in specialized medical domains. We present Echo‑α, an agentic multimodal reasoning model for ultrasound interpretation that unifies these strengths within an invoke‑and‑reason framework. Echo‑α is trained to coordinate organ‑specific detector outputs, integrate them with global visual context, and convert the resulting evidence into grounded diagnostic decisions beyond detector‑only inference. This behavior is established through a nine‑task supervised curriculum and then refined by sequential reinforcement learning under different reward trade‑offs, yielding Echo‑α‑Grounding for lesion anchoring and Echo‑α‑Diagnosis for final diagnosis. On multi‑center renal and breast ultrasound benchmarks, Echo‑α outperforms competitive baselines on both grounding and diagnosis. In particular, on cross‑center test sets, Echo‑α‑Grounding attains 56.73%/43.78% F1@0.5 and Echo‑ α‑Diagnosis reaches 74.90%/49.20% overall accuracy on renal/breast ultrasound. These results suggest that agentic multimodal reasoning can turn specialized detectors into verifiable clinical evidence, offering a practical route toward ultrasound AI systems that are more accurate, interpretable, and transferable. The repository is at https://github.com/MiliLab/Echo‑Alpha.
Authors:Ce Chen, Yi Ren, Yuanming Li, Viktor Goriachko, Zhenhui Ye, Zujin Guo, Zhibin Hong, Mingming Gong
Abstract:
Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points, frequently yielding corrupted video shots. We address this fundamental limitation by formalizing the Shot Transition Detection (STD) task. Rather than searching for ambiguous points, STD explicitly detects the continuous temporal segments of transitions. To tackle this, we propose TransVLM, a Vision‑Language Model (VLM) framework for STD. Unlike regular VLMs that predominantly rely on spatial semantics and struggle with fine‑grained inter‑shot dynamics, our method explicitly injects optical flow as a critical motion prior at the input stage. Through a simple yet effective feature‑fusion strategy, TransVLM directly processes concatenated color and motion representations, significantly enhancing its temporal awareness without incurring any additional visual token overhead on the language backbone. To overcome the severe class imbalance in public data, we design a scalable data engine to synthesize diverse transition videos for robust training, alongside a comprehensive benchmark for STD. Extensive experiments demonstrate that TransVLM achieves superior overall performance, outperforming traditional heuristic methods, specialized spatiotemporal networks, and top‑tier VLMs. This work has been deployed to production. For more related research, please visit HeyGen Research (https://www.heygen.com/research) and HeyGen Avatar‑V (https://www.heygen.com/research/avatar‑v‑model). Project page: https://chence17.github.io/TransVLM/
Authors:Fengxian Ji, Jingpu Yang, Zirui Song, Yuanxi Wang, Zhexuan Cui, Yuke Li, Qian Jiang, Xiuying Chen
Abstract:
Despite the rapid progress of large vision‑language models (LVLMs), fine‑grained, state‑conditioned GUI interaction remains challenging. Current evaluations offer limited coverage, imprecise target‑state definitions, and an overreliance on final‑task success, obscuring where and why agents fail. To address this gap, we introduce FineState‑Bench, a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. FineState‑Bench comprises 2,209 instances across desktop, web, and mobile platforms, spanning four interaction families and 23 UI component types, with each instance explicitly specifying an exact target state for fine‑grained state setting. We further propose FineState‑Metrics, a four‑stage diagnostic pipeline with stage‑wise success rates: Localization Success Rate (SR@Loc), Interaction Success Rate (SR@Int), Exact State Success Rate at Locate (ES‑SR@Loc), and Exact State Success Rate at Interact (ES‑SR@Int), and a plug‑and‑play Visual Diagnostic Assistant (VDA) that generates a Description and a bounding‑box Localization Hint to diagnose visual grounding reason via controlled w/ vs.\ w/o comparisons. On FineState‑Bench, exact goal‑state success remains low: ES‑SR@Int peaks at 32.8% on Web and 22.8% on average across platforms. With VDA localization hints, Gemini‑2.5‑Flash gains +14.9 ES‑SR@Int points, suggesting substantial headroom from improved visual grounding, yet overall accuracy is still insufficient for reliable fine‑grained state‑conditioned interaction \hrefhttps://github.com/FengxianJi/FineState‑BenchGithub.
Authors:Shiqi Xu, Moritz Burmester, Katharina Prasse, Isaac Bravo, Stefanie Walter, Margret Keuper
Abstract:
The pervasive growth of digital content, specifically short videos on social media platforms, has significantly altered how topics are discussed and understood in public discourse. In this work, we advance automated visual theme detection by assessing zero‑shot and clustering capabilities on social media data. (1) We evaluated the capabilities of notable VLMs such as VideoChatGPT, PandaGPT, and VideoLLava using zero‑shot image classification and compared their performance to the baseline provided by frame‑wise CLIP image classification. (2) By treating clustering as a minimum cost multicut problem, we aim to uncover insightful patterns in an unsupervised manner. For both analysis strategies, we provide extensive evaluations and practical guidance to practitioners. While VLMs are currently not able to detect climate change specific classes, the clustering results are distinct visual frames. %Given that VLMs are not currently capable to grasp the climate change discourse, we focus the clustering evaluation of image embedding models. We find that both ConvNeXt V2 and DINOv2 produce meaningful clusters, with DINOv2 focusing more on style differences and abstract categories, while ConvNeXt V2 clusters differ in more fine‑grained ways. Code available at https://github.com/KathPra/ClimateVID.git.
Authors:Junan Hu, Jian Liu, Jingxiang Lai, Jiarui Hu, Yiwei Sheng, Shuang Chen, Jian Li, Dazhao Du, Song Guo
Abstract:
Graphical User Interface (GUI) agents have emerged as a promising paradigm for intelligent systems that perceive and interact with graphical interfaces visually. Yet supervised fine‑tuning alone cannot handle long‑horizon credit assignment, distribution shifts, and safe exploration in irreversible environments, making Reinforcement Learning (RL) a central methodology for advancing automation. In this work, we present the first comprehensive overview of the intersection between RL and GUI agents, and examine how this research direction may evolve toward digital inhabitants. We propose a principled taxonomy that organizes existing methods into Offline RL, Online RL, and Hybrid Strategies, and complement it with analyses of reward engineering, data efficiency, and key technical innovations. Our analysis reveals several emerging trends: the tension between reliability and scalability is motivating the adoption of composite, multi‑tier reward architectures; GUI I/O latency bottlenecks are accelerating the shift toward world‑model‑based training, which can yield substantial performance gains; and the spontaneous emergence of System‑2‑style deliberation suggests that explicit reasoning supervision may not be necessary when sufficiently rich reward signals are available. We distill these findings into a roadmap covering process rewards, continual RL, cognitive architectures, and safe deployment, aiming to guide the next generation of robust GUI automation and its agent‑native infrastructure.
Authors:Zujin Guo, Zhenhui Ye, Yi Ren, Yuanming Li, Ce Chen, Zhibin Hong, Chen Change Loy
Abstract:
Existing talking avatar methods typically adopt an image‑to‑video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted, single‑view perspective lacks sufficient temporal and expression cues, limiting the ability to synthesize high‑fidelity talking avatars in customized backgrounds. To this end, we introduce Talking Avatar generation from Video Reference (TAVR), a novel framework that shifts the paradigm by leveraging cross‑scene video inputs. To effectively process these extended temporal contexts and bridge cross‑scene domain gaps, TAVR integrates a token selection module alongside a comprehensive three‑stage training scheme. Specifically, same‑scene video pretraining establishes foundational appearance copying, which is subsequently expanded by cross‑scene reference fine‑tuning for robust cross‑scene adaptation. Finally, task‑specific reinforcement learning aligns the generated outputs with identity‑based rewards to maximize identity similarity. To systematically evaluate cross‑scene robustness, we construct a new benchmark comprising 158 carefully curated cross‑scene video pairs. Extensive experiments show that TAVR benefits from flexible inference‑time video referencing and consistently surpasses existing baselines both quantitatively and qualitatively. This work has been deployed to production. For more related research, please visit \hrefhttps://www.heygen.com/researchHeyGen Research and \hrefhttps://www.heygen.com/research/avatar‑v‑modelHeyGen Avatar‑V.
Authors:Ishrak Hamim Mahi, Siam Ferdous, Md Sakib Sadman Badhon, Nabid Hasan Omi, Md Habibun Nabi Hemel, Farig Yousuf Sadeque, Md. Tanzim Reza
Abstract:
The rapid proliferation of image generation models and other artificial intelligence (AI) systems has intensified concerns regarding data privacy and user consent. As the availability of public datasets declines, major technology companies increasingly rely on proprietary or private user data for model training, raising ethical and legal challenges when users request the deletion of their data after it has influenced a trained model. Machine unlearning seeks to address this issue by enabling the removal of specific data from models without complete retraining. This study investigates a modified SISA (Sharded, Isolated, Sliced, and Aggregated) framework designed to achieve class‑level unlearning in Convolutional Neural Network (CNN) architectures. The proposed framework incorporates a reinforced replay mechanism and a gating network to enhance selective forgetting efficiency. Experimental evaluations across multiple image datasets and CNN configurations demonstrate that the modified SISA approach enables effective class unlearning while preserving model performance and reducing retraining overhead. The findings highlight the potential of SISA‑based unlearning for deployment in privacy‑sensitive AI applications. The implementation is publicly available at https://github.com/SiamFS/ sisa‑class‑unlearning.
Authors:Ekram Alam, Jaydip Sanyal, Akhil Kumar Das, Arijit Bhattacharya, Farhana Sultana
Abstract:
Mango cultivation is crucial in the agricultural sector, significantly contributing to economic development and food security. However, diseases affecting mango leaves can significantly reduce both the production and overall fruit grade. Detecting leaf diseases at an early stage with precision is key to effective disease prevention and sustaining crop productivity. In this paper, we introduce a "deep learning" model named "GourNet", which leverages "Convolutional Neural Networks" to identify infections in mango leaves. We utilize the "MangoLeafBD" (MBD) dataset to train and assess the effectiveness of the presented model. The MBD dataset contains seven disease classes and a Healthy class, making a total of eight classes. To enhance model performance, the images are preprocessed through steps like resizing, rescaling, and data augmentation prior to training. To properly evaluate the model, the dataset is separated into 80% for training, with the remaining 20% equally split between validation and testing. Our model uses only 683,656 total parameters and achieves a classification accuracy of 97%. This research's source code can be found at: https://github.com/ekramalam/GourNet‑Repo.
Authors:Hyeonseo Jang, Jaebyeong Jeon, Joong-Won Hwang, Kibok Lee
Abstract:
Test‑time prompt tuning (TPT) has emerged as a promising technique for enhancing the adaptability of vision‑language models by optimizing textual prompts using unlabeled test data. However, prior studies have observed that TPT often produces poorly calibrated models, raising concerns about the reliability of their predictions. Recent works address this issue by incorporating additional regularization terms that constrain model outputs, which improve calibration but often degrade performance. In this work, we reveal that these regularization strategies implicitly encourage optimization toward flatter minima, and that the sharpness of the loss landscape around adapted prompts is a key factor governing calibration quality. Motivated by this observation, we introduce Flatness‑aware Prompt Pretraining (FPP), a simple yet effective pretraining framework for TPT that initializes prompts within flatter regions of the loss landscape prior to adaptation. We show that simply replacing the initialization in existing TPT pipelines‑‑without modifying any other components‑‑is sufficient to improve both calibration and performance. Notably, FPP requires no labeled data and incurs no additional computational costs during test‑time tuning, making it highly practical for real‑world deployment. The code is available at: https://github.com/YonseiML/fpp.
Authors:Yuyang Li, Yime He, Zeyu Zhang, Dong Gong
Abstract:
Long‑term conversational memory requires retrieving evidence scattered across multiple sessions, yet single‑pass retrieval fails on temporal and multi‑hop questions. Existing iterative methods refine queries via generated content or document‑level signals, but none explicitly diagnoses the evidence gap, namely what is missing from the accumulated retrieval set, leaving query refinement untargeted. We present EviMem, combining IRIS (Iterative Retrieval via Insufficiency Signals), a closed‑loop framework that detects evidence gaps through sufficiency evaluation, diagnoses what is missing, and drives targeted query refinement, with LaceMem (Layered Architecture for Conversational Evidence Memory), a coarse‑to‑fine memory hierarchy supporting fine‑grained gap diagnosis. On LoCoMo, EviMem improves Judge Accuracy over MIRIX on temporal (73.3% to 81.6%) and multi‑hop (65.9% to 85.2%) questions at 4.5x lower latency. Code: https://github.com/AIGeeksGroup/EviMem.
Authors:Bohai Zhang, Wenjie Chen, Mu Li, Kaixing Long, Xing Shen, Xinqiang Yao, Jincheng Yang, Jianting Chen, Wei Yang, Qianjin Feng, Lei Cao
Abstract:
Accurate CT‑MRI registration of the cervical spine is essential for preoperative planning because this region is anatomically complex,highly variable,and vulnerable to injury of the vertebral arteries and spinal cord. However,cervical CT‑MRI registration remains underexplored,particularly for rigid‑deformable hybrid modeling,and the lack of high‑quality annotated multimodal data further limits progress. To address these challenges, we construct and release a comprehensively annotated CT‑MRI dataset, R‑D‑Reg, and propose MSR, a rigid‑deformable hybrid registration framework for complex joint structures. Specifically, MSR includes a rigid registration module for independent local rigid alignment of individual vertebrae and a deformable registration module with an MSL block that combines Mamba‑based global modeling and Swin Transformer‑based local modeling through adaptive gating. The rigid and deformable deformation fields are then fused to generate a hybrid field that better preserves local anatomical consistency. The code and dataset are publicly available at https://github.com/ssc1230609‑spec/MSR‑registration.
Authors:Dahua Gao, Yubo Dong, Anqi Li, Zhenyuan Lin, Ang Gao, Danhua Liu, Guangming Shi
Abstract:
Conventional push‑broom hyperspectral imaging suffers from slow acquisition speeds, precluding real‑time object detection; in contrast, snapshot spectral imaging enables instantaneous hyperspectral images (HSIs) capture, making real‑time object detection feasible, yet its potential is often compromised by time‑consuming post‑capture reconstruction. To address this issue, we propose the Focal U‑shaped Network (FUN), a novel end‑to‑end framework that jointly performs HSI reconstruction and object detection via multi‑task learning. FUN employs a shared U‑shaped backbone, where reconstruction provides underlying spectral information while detection guides semantic‑aware priors learning, facilitating mutually beneficial task interaction. Crucially, we introduce focal modulation, an efficient alternative to self‑attention that modulates spatial and spectral features while reducing quadratic computational complexity, enabling a self‑attention‑free architecture for joint reconstruction and detection. Furthermore, we contribute a new HSI object detection dataset with 8712 annotated objects across 363 HSIs to facilitate evaluation of the proposed method. Experiments demonstrate that FUN achieves state‑of‑the‑art performance on both tasks, using 40% fewer parameters and 30% less computation than recent alternatives, making it promising for future real‑time edge deployment. The code and datasets are available: https://github.com/ShawnDong98/FUN.
Authors:Junyi Ma, Erhang Zhang, Haoran Yang, Ditao Li, Chenyang Xu, Guangming Wang, Hesheng Wang
Abstract:
A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation skills from human video data has attracted rapidly growing attention in recent years, driven by the abundance of human activity videos and advances in computer vision. This line of research promises to enable robots to acquire skills passively from the vast and readily available resource of human demonstrations, substantially favoring scalable learning for generalist robotic systems. Therefore, we present this survey to provide a comprehensive and up‑to‑date review of human‑video‑based learning techniques in robotics, focusing on both human‑robot skill transfer and data foundations. We first review the policy learning foundations in robotics, and then describe the fundamental interfaces to incorporate human videos. Subsequently, we introduce a hierarchical taxonomy of transferring human videos to robot skills, covering task‑, observation‑, and action‑oriented pathways, along with a cross‑family analysis of their couplings with different data configurations and learning paradigms. In addition, we investigate the data foundations including widely‑used human video datasets and video generation schemes, and provide large‑scale statistical trends in dataset development and utilization. Ultimately, we emphasize the challenges and limitations intrinsic to this field, and delineate potential avenues for future research. The paper list of our survey is available at https://github.com/IRMVLab/awesome‑robot‑learning‑from‑human‑videos.
Authors:Al Zadid Sultan Bin Habib, Tanpia Tasnim, Md. Ekramul Islam, Muntasir Tabasum
Abstract:
Learning informative representations from tabular data in remote sensing and environmental science is challenging due to heterogeneity, scarce labels, and redundancy among features. We present ZAYAN (Zero‑Anchor dYnamic feAture eNcoding), a self‑supervised, feature‑centric contrastive framework for tabular data. ZAYAN performs contrastive learning at the feature rather than sample level, removing the need for explicit anchor selection and any reliance on class labels, while encouraging a redundancy‑minimized, disentangled embedding space. The framework has two modules: ZAYAN‑CL, which pretrains feature embeddings via a zero‑anchor contrastive objective with dynamic perturbations and masking, and ZAYAN‑T, a Transformer that conditions on these embeddings for downstream classification. Across eight datasets, including six remote‑sensing tabular benchmarks and two remote‑sensing‑driven flood‑prediction tables from satellite and GIS products, ZAYAN achieves superior accuracy, robustness, and generalization over tabular deep learning baselines, with consistent gains under label scarcity and distribution shift. These results indicate that feature‑level contrastive learning and dynamic feature encoding provide an effective recipe for learning from tabular sensing data.
Authors:Hezhao Liu, Jiacheng Yang, Junlong Gao, Mengke Li, Yiqun Zhang, Shreyank N Gowda, Yang Lu
Abstract:
In open‑world semi‑supervised learning (OWSSL), a model learns from labeled data and unlabeled data containing both known and novel classes. In practical OWSSL applications, models are expected to perform rigorous classification by directly selecting the most semantically relevant label from a candidate set for each sample. Existing OWSSL methods fail to achieve this because novel samples are trained without explicit supervision, and these methods lack mechanisms to extract latent semantic information, resulting in predicted labels that have no semantic correspondence to candidate textual labels. To address this, we introduce SEmantic Capture for Open‑world Semi‑supervised learning (SECOS), which directly predicts textual labels from the candidate set without post‑processing, meeting the requirements of practical OWSSL applications. SECOS leverages external knowledge to extract and align semantic representations across modalities for both known and novel classes, providing explicit supervisory signals for training novel classes. Extensive experiments demonstrate that even when existing OWSSL methods are evaluated under the more lenient post‑hoc matching setting, SECOS still surpasses them by up to 5.4% without such assistance, highlighting its superior effectiveness. Code is available at https://github.com/ganchi‑huanggua/OSSL‑Classification.
Authors:Davide Di Nucci, Riccardo Catalini, Guido Borghi, Roberto Vezzani
Abstract:
Recent advances in 3D reconstruction and neural rendering,particularly 3D Gaussian Splatting, make it feasible and simple to edit 3D scenes and re‑render them as highly realistic images. Therefore, security concerns arise regarding the authenticity of 3D content. Despite this threat, 3D fake detection remains largely unexplored in the literature, and most existing work is limited to 2D space. Therefore, in this paper, we formalize the concept of 3D fake detection and introduce Fake3DGS, a dataset of 3D Gaussian splatting scenes and corresponding rendered views, where fake images are produced by controlled manipulations of geometry, appearance, and spatial layout, while preserving high visual realism. Using this benchmark, we demonstrate that current state‑of‑the‑art 2D detectors struggle to distinguish between original and 3D manipulated images. To bridge this gap, we introduce a 3D‑aware detection method that leverages multi‑view coherence and features derived from the Gaussian splatting representation. Experimental results demonstrate a substantial improvement in recognizing modified 3D content, underscoring the validity of the new dataset and the necessity for authenticity assessment techniques that extend beyond 2D evidence. Code and data are publicly released for future investigations.
Authors:Lechao Zhang, Haoran Xu, Jingyu Gong, Xuhong Wang, Yuan Xie, Xin Tan
Abstract:
Embodied intelligence requires high‑fidelity simulation environments to support perception and decision‑making, yet existing platforms often suffer from data contamination and limited flexibility. To mitigate this, we propose World2Minecraft to convert real‑world scenes into structured Minecraft environments based on 3D semantic occupancy prediction. In the reconstructed scenes, we can effortlessly perform downstream tasks such as Vision‑Language Navigation(VLN). However, we observe that reconstruction quality heavily depends on accurate occupancy prediction, which remains limited by data scarcity and poor generalization in existing models. We introduce a low‑cost, automated, and scalable data acquisition pipeline for creating customized occupancy datasets, and demonstrate its effectiveness through MinecraftOcc, a large‑scale dataset featuring 100,165 images from 156 richly detailed indoor scenes. Extensive experiments show that our dataset provides a critical complement to existing datasets and poses a significant challenge to current SOTA methods. These findings contribute to improving occupancy prediction and highlight the value of World2Minecraft in providing a customizable and editable platform for personalized embodied AI research. Project page:https://world2minecraft.github.io/.
Authors:Xiaomeng Wang, Martha Larson, Zhengyu Zhao
Abstract:
When the visual style of text is considered, a wide variety can be observed in font, color, and size. However, when a word is read, its meaning is independent of the style in which it has been written or rendered. In this paper, we investigate whether, and how, the style in which a word is visualized in an image impacts the description that a Large Visual Language Model (LVLM) provides for the concept to which that word refers. Specifically, we investigate how functional text styles (readability‑oriented, e.g., black sans‑serif) versus decorative styles (display‑oriented, e.g., colored cursive/script) affect LVLMs' descriptions of a concept in terms of the attributes of that concept. Our experiments study the situation in which the LVLM is able to correctly identify the concept referred to by a visual text, i.e., by a word or words rendered as an image, and in which the visual text style should not influence the attribute‑based description that the LVLM produces. Our experimental results reveal that even when the concept is correctly identified, text style influences the model's attribute‑based descriptions of the concept. Our findings demonstrate non‑trivial style leakage from text style into semantic inference and motivate style‑aware evaluation and mitigation for LVLM‑based multimedia systems.
Authors:Shuo Wang, Jilin Mei, Wenfei Guan, Shuai Wang, Yan Xing, Chen Min, Yu Hu
Abstract:
Off‑road nighttime autonomous driving suffers from unreliable visible‑light perception, making infrared modality crucial for accurate freespace detection. However, progress remains limited due to the scarcity of annotated infrared off‑road datasets and the inter‑frame inconsistencies inherent to current single‑frame methods. To address these gaps, we present the IRON dataset, which, to our knowledge, is the first large‑scale infrared dataset for off‑road temporal freespace detection under all‑day conditions, with strong support for nighttime perception. The dataset comprises 24,314 densely annotated infrared images with synchronized RGB images in diverse scenes and different light conditions. Building upon this dataset, we propose IRONet, a novel flow‑free framework for temporal freespace detection that addresses inter‑frame inconsistencies by aggregating historical context via a memory‑attention mechanism and a carefully designed mask decoder. On our IRON dataset, IRONet achieves state‑of‑the‑art performance, reaching 82.93%(+1.19%) IoU and 90.66%(+0.71%) F1 score at real‑time inference. Remarkably, IRONet also exhibits robust generalization to RGB modalities on ORFD and Rellis datasets. Overall, our work establishes a foundation for reliable all‑day off‑road autonomous driving and future research in infrared temporal perception. The code and IRON dataset are available at https://github.com/wsnbws/IRON.
Authors:Wongi Park, Jordan A. James, Myeongseok Nam, Minjae Lee, Soomok Lee, Sang-Hyun Lee, William J. Beksi
Abstract:
We propose a 3D novel sparse‑view synthesis framework for unconstrained real‑world scenarios that contain distractors. Unlike existing methods that primarily perform novel‑view synthesis from a sparse set of constrained images without transient elements or leverage unconstrained dense image collections to enhance 3D representation in real‑world scenarios, our method not only effectively tackles sparse unconstrained image collections, but also shows high‑quality 3D rendering results. To do this, we introduce reference‑guided view refinement with a diffusion model using a transient mask and a reference image to enhance the 3D representation and mitigate artifacts in rendered views. Furthermore, we address sparse regions in the Gaussian field via pseudo‑view generation along with a sparsity‑aware Gaussian replication strategy to amplify Gaussians in the sparse regions. Extensive experiments on publicly available datasets demonstrate that our methodology consistently outperforms existing methods (e.g., PSNR ‑ 17.2%, SSIM ‑ 10.8%, LPIPS ‑ 4.0%) and provides high‑fidelity 3D rendering results. This advancement paves the way for realizing unconstrained real‑world scenarios without labor‑intensive data acquisition. Our project page is available at \hrefhttps://robotic‑vision‑lab.github.io/SaveWildGS/here
Authors:Yang Zhou, Chaoyong Zhang, Ruoyi Hao, Huilin Pan, Yang Zhang, Hongliang Ren
Abstract:
Nasotracheal intubation (NTI) is a critical clinical procedure for establishing and maintaining patient airway patency. Machine‑assisted NTI has emerged as a pivotal approach for optimizing procedural efficiency and minimizing manual intervention. However, visual detection algorithms employed for NTI navigation encounter significant challenges, including complex anatomical environments and suboptimal illumination conditions surrounding the glottis. Additionally, the glottis presents considerable scale variability throughout the procedure, initially appearing as a small, difficult‑to‑capture structure before expanding to occupy nearly the entire field of view. Moreover, traditional visual detection methods often have high computational costs, making real‑time, high‑precision detection on portable devices challenging. To enhance NTI efficacy and address these challenges, this paper proposes a novel glottis segmentation framework optimized for vision‑assisted NTI applications. First, we designed a lightweight, multi‑receptive field feature extraction module to reduce intra‑class differences, achieving robustness to scale variations of the glottis. This module was then stacked to form the backbone and neck of our network. Subsequently, we developed an advanced label assignment method and redefined the number of samples to further reduce intra‑class differences and enhance accuracy in the complex NTI environment. Experiments on three distinct datasets demonstrate that our network surpasses state‑of‑the‑art algorithms, achieving a segmentation mDice of 92.9% with a compact model size of 19 MB and an inference speed exceeding 170 frames per second. % Our code and datasets will be open‑sourced on GitHub after the manuscript is accepted. Our code and datasets are available at https://github.com/HBUT‑CV/GlottisNet.
Authors:Yihong Guo, Youwei Lyu, Jiajun Tang, Yizhuo Zhou, Hongliang Wang, Jinwei Chen, Changqing Zou, Qingnan Fan
Abstract:
Reasoning photo retouching has gained significant traction, requiring models to analyze image defects, give reasoning processes, and execute precise retouching enhancements. However, existing approaches often rely on non‑differentiable external software, creating optimization barriers and suffering from high parameter redundancy and limited generalization. To address these challenges, we propose VeraRetouch, a lightweight and fully differentiable framework for multi‑task photo retouching. We employ a 0.5B Vision‑Language Model (VLM) as the central intelligence to formulate retouching plans based on instructions and scene semantics. Furthermore, we develop a fully differentiable Retouch Renderer that replaces external tools, enabling direct end‑to‑end pixel‑level training through decoupled control latents for lighting, global color, and specific color adjustments. To overcome data scarcity, we introduce AetherRetouch‑1M+, the first million‑scale dataset for professional retouching, constructed via a new inverse degradation workflow. Furthermore, we propose DAPO‑AE, a reinforcement learning post‑training strategy that enhances autonomous aesthetic cognition. Extensive experiments demonstrate that VeraRetouch achieves state‑of‑the‑art performance across multiple benchmarks while maintaining a significantly smaller footprint, enabling mobile deployment. Our code and models are publicly available at https://github.com/OpenVeraTeam/VeraRetouch.
Authors:Peifu Liu, Tingfa Xu, Jie Wang, Huan Chen, Huiyan Bai, Jianan Li
Abstract:
Hyperspectral image classification demands spatially coherent predictions and precise boundary delineation. Yet prevailing superpixel‑based methods face an inherent contradiction: clustering aggregates similar pixels into regions, but the subsequent classifier operates pixel‑wise, undermining regional consistency. Consequently, existing approaches do not guarantee region‑level, boundary‑aligned classification. To address this limitation, we propose the Dual‑stage Spectrum‑Constrained Clustering‑based Classifier (DSCC), an end‑to‑end framework that explicitly decouples clustering from classification by first grouping spectral similar and spatially proximate pixels into spectral supertokens and then performing token‑level prediction. At its core, DSCC computes an image‑level multi‑criteria feature distance between pixels and centers, followed by a locality‑aware assignment regularization, enabling the generation of boundary‑preserving spectral supertokens. A density‑isolation based center selection further yields representative, well‑separated centers, reducing redundancy and improving robustness to scale variation. To accommodate mixed land‑cover compositions within each token, we introduce a soft‑label scheme that encodes class proportions and improves robustness for mixed‑class tokens. DSCC attains a CF1 of 0.728 at 197.75 FPS on the WHU‑OHS dataset, offering a superior accuracy‑efficiency trade‑off compared with state‑of‑the‑art methods. Extensive experiments further validate the effectiveness and generality of the proposed dual‑stage paradigm for hyperspectral image classification. The source code is available at https://github.com/laprf/DSCC.
Authors:Yingrui Wu, Youkang Kong, Mingyang Zhao, Weize Quan, Dong-Ming Yan, Yang Liu
Abstract:
Synthesizing realistic 3D indoor scenes remains challenging due to data scarcity and the difficulty of simultaneously enforcing global architectural constraints and local semantic consistency. Existing approaches often overlook structural boundaries or rely on fully connected relation graphs that introduce redundant generation errors. Inspired by human design cognition, we present CasLayout, a cascaded diffusion framework that decomposes the joint scene generation task into four conditional sub‑stages with explicit physical and semantic roles: (1) predicting furniture quantity and categories, (2) refining object sizes and feature embeddings, (3) modeling spatial relationships in a latent space, and (4) generating Oriented Bounding Boxes (OBBs). This decoupled architecture reduces data requirements and enables flexible integration of Large Language Models (LLMs) and Vision Language Models (VLMs) for zero‑shot tasks such as image‑to‑scene generation. To maintain physical validity within complex floor plans, we explicitly model building elements (e.g., walls, doors, and windows) as conditional constraints. Furthermore, to address the high entropy of dense relation graphs, we introduce a sparse relation graph formulation aligned with human spatial descriptions. By encoding these sparse graphs into a compact latent space using a bidirectional Variational Autoencoder (VAE), the proposed framework provides enhanced relational controllability, allowing generated layouts to better respect functional organization. Experiments demonstrate that CasLayout achieves state‑of‑the‑art performance in fidelity and diversity while enabling improved controllability in practical applications.
Authors:Naeem Rehmat, Muhammad Saad Saeed, Ijaz Ul Haq, Khalid Malik
Abstract:
Web filtering systems rely on accurate web content classification to block cyber threats, prevent data exfiltration, and ensure compliance. However, classification is increasingly difficult due to the dynamic and rapidly evolving nature of the modern web. Embedding‑based zero‑shot approaches map content and category descriptions into a shared semantic space, enabling label assignment without labeled training data, but remain highly sensitive to definition quality. Poorly specified or ambiguous definitions create semantic overlap in the embedding space, leading to systematic misclassification.
In this paper, we propose a training‑free, adaptive iterative definition refinement framework that improves zero‑shot web content classification by progressively optimizing category definitions rather than updating model parameters. Using LLMs as feedback‑driven definition optimizers, we investigate three refinement strategies namely example‑guided, confusion‑aware, and history‑aware, each refining class descriptions using structured signals from misclassified instances. Furthermore, we introduce a human‑labeled benchmark of 10 URL categories with 1,000 samples per class and evaluate across 13 state‑of‑the‑art embedding foundation models. Results demonstrate that iterative definition refinement consistently improves classification performance across diverse architectures, establishing definition quality as a critical and underexplored factor in embedding‑based systems. The dataset is available at https://github.com/naeemrehmat/B2MWT‑10C.
Authors:Youkang Kong, Yang Liu, Yue Dong, Xin Tong, Heung-Yeung Shum
Abstract:
3D shapes from scanning, reconstruction, or AI‑generated content often lack simple quad mesh layouts ‑‑ critical for efficient editing and modeling. Existing quad‑remeshing techniques typically produce complex layouts with irregular loops, leading to tedious manual cleanup and extensive algorithm tuning. We introduce SQuadGen, a diffusion‑based generative framework that leverages Chart Distance Fields (CDF) to synthesize simple quad layouts on 3D shapes. Our approach addresses two key challenges: (1) the discrete nature of mesh connectivity, which hinders learning, and (2) the scarcity of large‑scale datasets with simple quad meshes. To overcome the first, we propose CDF, a continuous surface‑based representation enabling effective learning and synthesis of quad layouts. To address the second, we define loop‑aware simplicity metrics and construct a large‑scale dataset of high‑quality quad layouts recovered from public 3D repositories through a robust quad‑recovery pipeline. Extensive evaluations across diverse 3D inputs show that SQuadGen consistently outperforms existing methods, producing robust, artist‑friendly simple quad layouts.
Authors:Tengya Zhang, Feng Gao, Lin Qi, Junyu Dong, Qian Du
Abstract:
Hyperspectral image super‑resolution is essential for enhancing the spatial fidelity of HSI data, yet existing deep learning methods often struggle with substantial spectral redundancy and the limited non‑linear modeling capacity of standard feed‑forward networks (FFNs). To address these challenges, we propose Spectral Dynamic Attention Network (SDANet), a framework designed to adaptively suppress redundant spectral interactions. SDANet integrates two key components: 1) Dynamic Channel Sparse Attention (DCSA) module that computes channel‑wise correlations and selectively preserves the most informative attention responses through dynamic and data‑dependent sparsification. 2) Frequency‑Enhanced Feed‑Forward Network (FE‑FFN) that jointly models spatial and frequency‑domain representations to enhance non‑linear expressiveness. Extensive experiments on two benchmark datasets demonstrate that SDANet achieves state‑of‑the‑art HISR performance while maintaining competitive efficiency. The code will be made publicly available at https://github.com/oucailab/SDANet.
Authors:Chuanzheng Gong, Feng Gao, Junyan Lin, Junyu Dong, Qian Du
Abstract:
Hyperspectral image (HSI) and SAR/LiDAR data offer complementary spectral and structural information for land‑cover classification. However, their effective fusion remains challenging due to two major limitations: The spectral redundancy in high‑dimensional HSI and the heterogeneous characteristics between multi‑source data. To this end, we propose Representative Spectral Correlation Network (RSCNet), a novel multi‑source image classification framework specifically designed to address the above challenges through spectral selection and adaptive interaction. The network incorporates two key components: (1) Key Band Selection Module (KBSM) that adaptively selects task‑relevant spectral bands from the original HSI under cross‑source guidance, thereby alleviating redundancy and mitigating information loss from conventional PCA‑based spectral reduction. Moreover, the learned band subset exhibits highly discriminative spectral structures that align with discriminative semantic cues, promoting compact yet expressive representations. (2) Cross‑source Adaptive Fusion Module (CAFM) that performs cross‑source attention weighting and local‑global contextual refinement to enhance cross‑source feature interaction. Experiments on three public benchmark datasets demonstrate that our RSCNet achieves superior performance compared with state‑of‑the‑art methods, while maintaining substantially lower computational complexity. Our codes are publicly available at https://github.com/oucailab/RSCNet.
Authors:Chenyang Wu, Lina Lei, Fan Li, Chun-Le Guo, Dehong Kong, Xinran Qin, Zhixin Wang, Ming-Ming Cheng, Chongyi Li
Abstract:
Recent advances in Diffusion Transformer (DiT)‑based video generation technologies have shown impressive results for video object removal. However, these methods still suffer from substantial inference latency. For instance, although MiniMax Remover achieves state‑of‑the‑art visual quality, it operates at only around 10FPS, primarily due to dense computations over the entire spatiotemporal token space, even when only a small masked region actually requires processing. In this paper, we present YOSE, You Only Select Essential Tokens, an efficient fine‑tuning framework. YOSE introduces two key components: Batch Variable‑length Indexing (BVI) and Diffusion Process Simulator (DiffSim) Module. BVI is a differentiable dynamic indexing operator that adaptively selects essential tokens based on mask information, enabling variable‑length token processing across samples. DiffSim provides a diffusion process approximation mechanism for unmasked tokens, which simulates the influence of unmasked regions within DiT self‑attention to maintain semantic consistency for masked tokens. With these designs, YOSE achieves mask‑aware acceleration, where the inference time scales approximately linearly with the masked regions, in contrast to full‑token diffusion methods whose computation remains constant regardless of the mask size. Extensive experiments demonstrate that YOSE achieves up to 2.5X speedup in 70% of cases while maintaining visual quality comparable to the baseline. Code is available at: https://github.com/Wucy0519/YOSE‑CVPR26.
Authors:Yizhou Wu, Shansong Wang, Yuheng Li, Mojtaba Safari, Mingzhe Hu, Chih-Wei Chang, Harini Veeraraghavan, Xiaofeng Yang
Abstract:
Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning‑based methods remain task‑specific and require substantial labeled data. Here we show that a single self‑supervised representation can generalize across heterogeneous brain MRI endpoints. We trained BrainDINO, a self‑distilled foundation model, on approximately 6.6 million unlabeled axial slices from 20 datasets encompassing broad variation in population, disease, and acquisition setting. Using a frozen encoder with lightweight task heads, BrainDINO supported transfer across tumor segmentation, neurodegenerative and neurodevelopmental conditions classification, brain age estimation, post‑stroke temporal prediction, molecular status prediction, MRI sequence classification, and survival modeling. Across tasks and supervision regimes, BrainDINO consistently equaled or exceeded natural‑image and MRI‑specific self‑supervised baselines, with particularly strong advantages under label scarcity. Representation analyses further showed anatomically organized and pathology‑sensitive feature structure in the absence of task‑specific supervision. Our findings indicate that large‑scale slice‑wise self‑supervised learning can yield a unified brain MRI representation that supports diverse neuroimaging tasks without volumetric pretraining or full‑network fine‑tuning, establishing a scalable foundation for robust and data‑efficient brain imaging analysis. Code is available at https://github.com/mclwu22/BrainDINO
Authors:Wanrong Zheng, Yunhao Ge, Laurent Itti
Abstract:
Breakthrough progress in vision‑based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These models can plan a sequence of motions by evaluating the current view at each time step against the task and goal given to the agent. However, current zero‑shot Vision‑and‑Language Navigation (VLN) agents powered by MLLMs still tend to drift off course, halt prematurely, and achieve low overall success rates. We propose Three‑Step Nav to counteract these failures with a three‑view protocol: First, "look forward" to extract global landmarks and sketch a coarse plan. Then, "look now" to align the current visual observation with the next sub‑goal for fine‑grained guidance. Finally, "look backward" audits the entire trajectory to correct accumulated drift before stopping. Requiring no gradient updates or task‑specific fine‑tuning, our planner drops into existing VLN pipelines with minimal overhead. Three‑Step Nav achieves state‑of‑the‑art zero‑shot performance on the R2R‑CE and RxR‑CE dataset. Our code is available at https://github.com/ZoeyZheng0/3‑step‑Nav.
Authors:Alexander Raistrick, Karhan Kayan, Jack Nugent, David Yan, Lingjie Mei, Meenal Parakh, Hongyu Wen, Dylan Li, Yiming Zuo, Erich Liang, Jia Deng
Abstract:
We introduce ProcFunc, a library for Blender‑based procedural 3D generation in Python. ProcFunc provides a library of easy‑to‑use Python functions, which streamline creating, combining, analyzing, and executing procedural generation code. ProcFunc makes it easy to create large‑scale diverse training data, by combinatorial compositions of semantic components. VLMs can use ProcFunc to edit procedural material and geometry code and can create new procedural code with significantly fewer coding errors. Finally, as an example use case, we use ProcFunc to develop a new procedural generator of indoor rooms, which includes a collection of new compositional procedural materials. We demonstrate the detail, runtime efficiency, and diversity of this room generator, as well as its use for 3D synthetic data generation. Please visit https://github.com/princeton‑vl/procfunc for source code.
Authors:Wanyue Zhang, Wenxiang Wu, Wang Xu, Jiaxin Luo, Helu Zhi, Yibin Huang, Shuo Ren, Zitao Liu, Jiajun Zhang
Abstract:
Vision‑language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address this limitation either by scaling spatial supervision with synthetic data or by coupling VLMs with world models at inference time. However, the former often lacks explicit modeling of motion‑conditioned state transitions, while the latter incurs substantial computational overhead. In this work, we propose World2VLM, a training framework that distills spatial imagination from a generative world model into a vision‑language model. Given an initial observation and a parameterized camera trajectory, we use a view‑consistent world model to synthesize geometrically aligned future views and derive structured supervision for both forward (action‑to‑outcome) and inverse (outcome‑to‑action) spatial reasoning. We post‑train the VLM with a two‑stage recipe on a compact dataset generated by this pipeline and evaluate it on multiple spatial reasoning benchmarks. World2VLM delivers consistent improvements over the base model across diverse benchmarks, including SAT‑Real, SAT‑Synthesized, VSI‑Bench, and MindCube. It also outperforms the test‑time world‑model‑coupled methods while eliminating the need for expensive inference‑time generation. Our results suggest that world models can serve not only as inference‑time tools, but also as effective training‑time teachers, enabling VLMs to internalize spatial imagination in a scalable and efficient manner.
Authors:Zijie Wu, Chaohui Yu, Fan Wang, Xiang Bai
Abstract:
Recent advances in 4D content generation have attracted increasing attention, yet creating high‑quality animated 3D models remains challenging due to the complexity of modeling spatio‑temporal distributions and the scarcity of 4D training data. We present AnimateAnyMesh++, a feed‑forward framework for text‑driven animation of arbitrary 3D meshes with substantial upgrades in data, architecture, and generative capability. First, we expand the DyMesh‑XL dataset by mining dynamic content from Objaverse‑XL, increasing the number of unique identities from 60K to 300K and substantially broadening category and motion diversity. Second, we redesign DyMeshVAE‑Flex with power‑law topology‑aware attention and vertex‑normal enhanced features, which significantly improves trajectory reconstruction, local geometry preservation, and mitigates trajectory‑sticking artifacts. Third, we introduce architectural changes to both DyMeshVAE‑Flex and the rectified‑flow (RF) generator to support variable‑length sequence training and generation, enabling longer animations while preserving reconstruction fidelity. Extensive experiments demonstrate that AnimateAnyMesh++ generates semantically accurate and temporally coherent mesh animations within seconds, surpassing prior approaches in quality and efficiency. The enlarged DyMesh‑XL, the upgraded DyMeshVAE‑Flex, and variable‑length RF together deliver consistent gains across benchmarks and in‑the‑wild meshes. We will release code, models, and the expanded DyMesh‑XL upon acceptance of this manuscript to facilitate research in 4D content creation.
Authors:Fangqiang Fan, Zhicheng Zhao, Xiaoliang Ma, Chenglong Li, Jin Tang
Abstract:
Fine‑grained RGBT image semantic segmentation is crucial for all‑weather unmanned aerial vehicle (UAV) scene understanding. However, UAV RGBT image semantic segmentation faces two coupled challenges: cross‑modal spatial misalignment caused by sensor parallax and platform vibration, and severe semantic confusion among fine‑grained ground objects under top‑down aerial views. To address these issues, we propose a Graph‑based Semantic Calibration Network (GSCNet) for unaligned UAV RGBT image semantic segmentation. Specifically, we design a Feature Decoupling and Alignment Module (FDAM) that decouples each modality into shared structural and private perceptual components and performs deformable alignment in the shared subspace, enabling robust spatial correction with reduced modality appearance interference. Moreover, we propose a Semantic Graph Calibration Module (SGCM) that explicitly encodes the hierarchical taxonomy and co‑occurrence regularities among ground‑object categories in UAV scenes into a structured category graph, and incorporates these priors into graph‑attention reasoning to calibrate predictions of visually similar and rare categories. In addition, we construct the Unaligned RGB‑Thermal Fine‑grained (URTF) benchmark, to the best of our knowledge, the largest and most fine‑grained benchmark for unaligned UAV RGBT image semantic segmentation, containing over 25,000 image pairs across 61 semantic categories with realistic cross‑modal misalignment. Extensive experiments on URTF demonstrate that GSCNet significantly outperforms state‑of‑the‑art methods, with notable gains on fine‑grained categories. The dataset is available at https://github.com/mmic‑lcl/Datasets‑and‑benchmark‑code.
Authors:Changhyun Roh, Yonghyun Jeong, Jonghyun Lee, Chanho Eom, Jihyong Oh
Abstract:
Synthesizing a target concept from a single reference image is challenging in diffusion‑based personalized text‑to‑image generation, particularly for sticker personalization where prompts often require explicit attribute edits. With only one reference, test‑time fine‑tuning (TTF) methods tend to overfit, producing visual entanglement, where background artifacts are absorbed into the learned concept, and structural rigidity, where the model memorizes reference‑specific spatial configurations and loses contextual controllability. To address these issues, we introduce SEmantic‑aware single‑image sticker personALization (SEAL), a plug‑and‑play, architecture‑agnostic adaptation module that integrates into existing personalization pipelines without modifying their U‑Net‑based diffusion backbones. SEAL applies three components during embedding adaptation: (1) a Semantic‑guided Spatial Attention Loss, (2) a Split‑merge Token Strategy, and (3) Structure‑aware Layer Restriction. To support sticker‑domain personalization with attribute‑level control, we present StickerBench, a large‑scale sticker image dataset with structured tags under a six‑attribute schema (Appearance, Emotion, Action, Camera Composition, Style, Background). These annotations provide a consistent interface for varying context while keeping target identity fixed, enabling systematic evaluation of identity disentanglement and contextual controllability. Experiments show that SEAL consistently improves identity preservation while maintaining contextual controllability, highlighting the importance of explicit spatial and structural constraints during test‑time adaptation. The code, StickerBench, and project page will be publicly released.
Authors:Mingbo Hong, Feng Liu, Caroline Gevaert, George Vosselman, Hao Cheng
Abstract:
Detectors often suffer from degraded performance, primarily due to the distributional gap between the source and target domains. This issue is especially evident in single‑source domains with limited data, as models tend to rely on confounders (e.g., illumination, co‑occurrence, and style) from the source domain, leading to spurious correlations that hinder generalization. To this end, this paper proposes a novel Basis‑driven framework for domain generalization, namely Bridge, that incorporates causal inference into object detection. By learning the low‑rank bases for front‑door adjustment, Bridge blocks confounders' effects to mitigate spurious correlations, while simultaneously refining representations by filtering redundant and task‑irrelevant components. Bridge can be seamlessly integrated with both discriminative (e.g., DINOv2/3, SAM) and generative (e.g., Stable Diffusion) Vision Foundation Models (VFMs). Extensive experiments across multiple domain generalization object detection datasets, i.e., Cross‑Camera, Adverse Weather, Real‑to‑Artistic, Diverse Weather Datasets, and Diverse Weather DroneVehicle (our newly augmented real‑world UAV‑based benchmark), underscore the superiority of our proposed method over previous state‑of‑the‑art approaches. The project page is available at: https://mingbohong.github.io/Bridge/.
Authors:Shuzhao Xie, Junchen Ge, Weixiang Zhang, Jiahang Liu, Chen Tang, Yunpeng Bai, Shijia Ge, Jingyan Jiang, Yuzhi Huang, Fengnian Yang, Cong Zhang, Xiaoyi Fan, Zhi Wang
Abstract:
3D Gaussian Splatting (3DGS) achieves high‑quality novel view synthesis with real‑time rendering, but its storage cost remains prohibitive for practical deployment. Existing post‑training compression methods still rely on many coupled hyperparameters across pruning, transformation, quantization, and entropy coding, making it difficult to control the final compressed size and fully exploit the rate‑distortion trade‑off. We propose MesonGS++, a size‑aware post‑training codec for 3D Gaussian compression. On the codec side, MesonGS++ combines joint importance‑based pruning, octree geometry coding, attribute transformation, selective vector quantization for higher‑degree spherical harmonics, and group‑wise mixed‑precision quantization with entropy coding. On the configuration side, it treats the reserve ratio and bit‑width allocation as the dominant rate‑distortion knobs and jointly optimizes them under a target storage budget via discrete sampling and 0‑‑1 integer linear programming. We further propose a linear size estimator and a CUDA parallel quantization operator to accelerate the hyperparameter searching process. Extensive experiments show that MesonGS++ achieves over 34× compression while preserving rendering fidelity, outperforming state‑of‑the‑art post‑training methods and accurately meeting target size budgets. Remarkably, without any training, MesonGS++ can even surpass the PSNR of vanilla 3DGS at a 20× compression rate on the Stump scene. Our code is available at https://github.com/mmlab‑sigs/mesongs_plus
Authors:Jun Guo, Qiwei Li, Peiyan Li, Zilong Chen, Nan Sun, Yifei Su, Heyun Wang, Yuan Zhang, Xinghang Li, Huaping Liu
Abstract:
We propose X‑WAM, a Unified 4D World Model that unifies real‑time robotic action execution and high‑fidelity 4D world synthesis (video + 3D reconstruction) in a single framework, addressing the critical limitations of prior unified world models (e.g., UWM) that only model 2D pixel‑space and fail to balance action efficiency and world modeling quality. To leverage the strong visual priors of pretrained video diffusion models, X‑WAM imagines the future world by predicting multi‑view RGB‑D videos, and obtains spatial information efficiently through a lightweight structural adaptation: replicating the final few blocks of the pretrained Diffusion Transformer into a dedicated depth prediction branch for the reconstruction of future spatial information. Moreover, we propose Asynchronous Noise Sampling (ANS) to jointly optimize generation quality and action decoding efficiency. ANS applies a specialized asynchronous denoising schedule during inference, which rapidly decodes actions with fewer steps to enable efficient real‑time execution, while dedicating the full sequence of steps to generate high‑fidelity video. Rather than entirely decoupling the timesteps during training, ANS samples from their joint distribution to align with the inference distribution. Pretrained on over 5,800 hours of robotic data, X‑WAM achieves 79.2% and 90.7% average success rate on RoboCasa and RoboTwin 2.0 benchmarks, while producing high‑fidelity 4D reconstruction and generation surpassing existing methods in both visual and geometric metrics.
Authors:William Grolleau, Astrid Sabourin, Guillaume Lapouge, Catherine Achard
Abstract:
Aerial‑Ground Re‑Identification (AG‑ReID) is constrained by the viewpoint‑domain gap, as drastic viewpoint disparities occlude or distort discriminative features, making cross‑viewpoint image retrieval challenging. While existing methods rely on paired cross‑view annotations, real‑world deployments, such as wilderness search‑and‑rescue (SAR), often lack target‑domain data, requiring retrieval from ground‑level references alone. To our knowledge, we are the first to address this challenge by formalizing the Single‑View AG‑ReID (SV AG‑ReID) setting, where models trained on a single real viewpoint must generalize to an unseen viewpoint. We propose 3D Lifting‑based Elevated Novel‑view Synthesis (3D‑LENS), a unified framework combining geometrically‑consistent novel view synthesis that leverages large‑scale 3D mesh reconstruction, with a robust representation learning scheme to mitigate synthetic‑to‑real bias. Unlike 2D generative baselines that suffer from geometric inconsistencies or prior 3D methods that are restricted to class‑specific templates, our approach ensures view‑consistent synthesis across diverse categories without predefined templates that fail to capture fine‑grained details, such as carried objects. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance on SV AG‑ReID scenarios. Code and data will be released at https://github.com/TurtleSmoke/3D‑LENS.
Authors:Cyril Shih-Huan Hsu, Wig Yuan-Cheng Cheng, Chrysa Papagianni
Abstract:
Deploying Vision‑Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource‑constrained embedded platforms. Conversely, fully offloading inference to the cloud is often impractical in bandwidth‑limited environments, where transmitting raw visual data introduces substantial latency overhead. While recent edge‑cloud collaborative architectures attempt to partition VLM workloads across devices, they typically rely on transmitting fixed‑size representations, lacking adaptability to dynamic network conditions and failing to fully exploit semantic redundancy. In this paper, we propose a progressive semantic communication framework for edge‑cloud VLM inference, using a Meta AutoEncoder that compresses visual tokens into adaptive, progressively refinable representations, enabling plug‑and‑play deployment with off‑the‑shelf VLMs without additional fine‑tuning. This design allows flexible transmission at different information levels, providing a controllable trade‑off between communication cost and semantic fidelity. We implement a full end‑to‑end edge‑cloud system comprising an embedded NXP i.MX95 platform and a GPU server, communicating over bandwidth‑constrained networks. Experimental results show that, at 1 Mbps uplink, the proposed progressive scheme significantly reduces network latency compared to full‑edge and full‑cloud solutions, while maintaining high semantic consistency even under high compression. The implementation code will be released upon publication at https://github.com/open‑ep/ProSemComVLM.
Authors:Zhirong Shen, Rui Huang, Jiacheng Liu, Chang Zou, Peiliang Cai, Shikang Zheng, Zhengyi Shi, Liang Feng, Linfeng Zhang
Abstract:
To address the high sampling cost of Diffusion Transformers (DiTs), feature caching offers a training‑free acceleration method. However, existing methods rely on hand‑crafted forecasting formulas that fail under aggressive skipping. We propose L2P (Learnable Linear Predictor), a simple data‑driven caching framework that replaces fixed coefficients with learnable per‑timestep weights. Rapidly trained in ~20 seconds on a single GPU, L2P accurately reconstructs current features from past trajectories. L2P significantly outperforms existing baselines: it achieves a 4.55x FLOPs reduction and 4.15x latency speedup on FLUX.1‑dev, and maintains high visual fidelity under up to 7.18x acceleration on Qwen‑Image models, where prior methods show noticeable quality degradation. Our results show learning linear predictors is highly effective for efficient DiT inference. Code is available at https://github.com/Aredstone/L2P‑Cache.
Authors:Fengchun Zhang, Qiang Ma, Liuyu Xiang, Jinshan Lai, Tingxuan Huang, Jianwei Hu
Abstract:
Federated domain generalization for person re‑identification (FedDG‑ReID) aims to collaboratively train a pedestrian retrieval model across multiple decentralized source domains such that it can generalize to unseen target environments without compromising raw data privacy. However, this task is significantly challenged by the inherent stylistic gaps across decentralized clients. Without global supervision, models easily succumb to shortcut learning where representations overfit to domain specific camera biases rather than universal identity features. We propose CO‑EVO, a novel federated framework that resolves this semantic‑style conflict through a co‑evolutionary mechanism. On the semantic side, Camera‑Invariant Semantic Anchoring (CSA) learns identity prompts with cross‑camera consistency to establish purified and domain‑agnostic anchors that filter out local imaging noise. On the visual side, Global Style Diversification (GSD), powered by a Global Camera‑Style Bank (GCSB), synthesizes realistic perturbations to expand the visual boundaries of training data. The core of CO‑EVO is its co‑evolutionary loop where purified anchors act as gravitational centers to guide the image encoder toward robust anatomical attributes amidst diverse style variations. Extensive experiments demonstrate that CO‑EVO achieves state‑of‑the‑art (SOTA) performance, proving that the synergy between semantic purification and style expansion is essential for robust cross‑domain generalization. Our code is available at: https://github.com/NanYiyuzurn/ACL‑LGPS‑2026.
Authors:Jianing You, Han Wang, Kang Liu, Jiale Ding, Fengjie Chu, Zihan Guo, Shengyang Li
Abstract:
Automated animal behavior analysis relies on long‑term, interpretable individual trajectories; however, multi‑animal tracking in space science experimental videos remains highly challenging due to weak appearance cues, low‑quality imaging, complex maneuvering behaviors, and frequent interactions. To address this problem, we first construct the SpaceAnimal‑MOT dataset to characterize the motion complexity and long‑term identity preservation challenges in biological videos acquired under microgravity conditions. We then propose ART‑Track (Adaptive Robust Tracking), a motion‑driven tracking framework tailored to this setting. Specifically, multi‑model motion estimation is introduced to handle abrupt maneuvers and nonlinear motion, motion‑state‑driven association is designed to reduce identity switches under dense interactions and temporary mismatch, and uncertainty‑adaptive fusion is used to dynamically balance spatial and motion cues when prediction reliability varies. Experimental results show that ART‑Track significantly reduces identity switches on zebrafish and fruitfly sequences, while maintaining more stable association under occlusion, deformation, and high‑density interactions, thereby providing a more reliable tracking foundation for downstream quantitative behavior analysis. The code is publicly available at https://github.com/yyy7777777/ART_TRACK/tree/main.
Authors:Kuo-Liang Chung, Yu-Cheng Lin, Wu-Chi Chen
Abstract:
Point cloud registration (PCR) is a fundamental task for integrating 3D observations in remote sensing applications. This paper proposes a fast and effective PCR algorithm utilizing probabilistic self‑updating local correspondence and line vector sets. Our dual RANSAC interaction model comprises a global RANSAC evaluating the global correspondence set and a local RANSAC operating on dynamically updated local sets. Initially, these local sets are constructed using angle histogram statistics and line vector length preservation techniques. To improve accuracy, a probabilistic self‑updating strategy refines the local sets after each interaction round. To reduce runtime, we introduce a global early termination condition that optimally balances accuracy and efficiency. Finally, a weighted singular value decomposition estimates the registration solution. Evaluations on public datasets demonstrate our algorithm achieves superior time efficiency and at least a 10% root mean square error improvement over state‑of‑the‑art methods. The C++ source code is publicly available at https://github.com/ivpml84079/Probabilistic‑Self‑Update‑Line‑Vector‑Set‑Based‑Point‑Cloud‑Registration.
Authors:Boxiang Yang, Ning Chen, Xia Yue, Yichang Luo, Yingbo Fan, Haoyuan Zhang, Haoyu Ma, Jun Yue, Shanjun Mao
Abstract:
Recently, Hyperspectral Image (HSI) classification has attracted increasing attention in remote sensing. However, HSI data are inherently high‑dimensional but low‑rank, with discriminative information concentrated on a low‑dimensional latent manifold. In real‑world remote sensing scenarios, the superposition of multiple degradation factors disrupts this intrinsic manifold structure, driving samples away from their original low‑dimensional distribution and introducing substantial redundant and non‑discriminative variations. To better handle this challenge, this paper proposes a manifold‑space diffusion framework (MSDiff) for robust hyperspectral classification under complex degradation conditions. Specifically, the proposed method first maps high‑dimensional, degradation‑affected HSI data into a compact low‑dimensional manifold through a discriminative spectral‑spatial reconstruction task, preserving class semantics and reducing redundant variations. A diffusion‑based generative model is then applied to regularize the spectral‑spatial distribution within the manifold, enabling progressive refinement and stabilization of latent features against residual degradations. The key advantage of the proposed framework lies in performing diffusion‑based distribution modeling directly on the low‑dimensional manifold, effectively decoupling degradation‑induced disturbances from intrinsic discriminative structures and enhancing representation stability under complex degradations. Experimental results on multiple hyperspectral benchmarks demonstrate consistent performance improvements over state‑of‑the‑art methods under diverse composite degradation settings. The code will be available at https://github.com/yangboxiang1207/MSDiff
Authors:Yuqi Li, Qian Zhou, Huiran Duan, Jingjie Wang, Shunli Zhang, Chuanguang Yang, Guoying Zhao, Yingli Tian
Abstract:
Gait recognition is an attractive biometric modality for long‑range and contact‑free identification, but high‑performing gait models often rely on deep and computationally expensive architectures that are difficult to deploy in practice. Knowledge distillation (KD) offers a natural way to transfer knowledge from a powerful teacher to an efficient student; however, standard KD is often less effective for part‑structured gait models, where supervision is formed from both part‑wise classification logits and part‑wise retrieval embeddings. In this paper, we propose GaitKD, a distillation framework that decouples gait knowledge transfer into two complementary components: decision‑level distillation and boundary‑level distillation. Specifically, GaitKD aligns the teacher and student through part‑calibrated logit distillation to transfer inter‑class decision relations, while preserving the teacher‑induced partitioning of the embedding space through an activation‑boundary objective instead of direct feature regression. With a simple aligned part‑wise design, GaitKD supports heterogeneous teacher‑student gait models without introducing additional inference cost. Experimental results across multiple gait recognition benchmarks and teacher‑student configurations show consistent improvements over strong gait baselines. Our study demonstrates that the two transfer components are complementary, and boundary‑preserving distillation provides more stable performance than direct feature regression. Source code is available at https://github.com/liyiersan/GaitKD/
Authors:Mohammed Q. Alkhatib, Ali Jamali
Abstract:
Over the past decade, hyperspectral image (HSI) classification has drawn considerable interest due to HSIs' ability to effectively distinguish terrestrial objects by capturing detailed, continuous spectral information. The strong performance of recent deep learning techniques in tasks like image classification and semantic segmentation has led to their growing use in HSI classification, due to their ability to capture complex spatial and spectral features more effectively than traditional methods. This paper presents MixerCA, a novel lightweight model for HSI classification that leverages depthwise convolution and a self‑attention mechanism. MixerCA integrates depth‑wise convolutions, token and channel mixing, and coordinate attention into a unified structure to decouple spatial and channel interactions, maintain consistent resolution throughout the network, and directly process HSI patches. Extensive experiments on four hyperspectral benchmark datasets reveal MixerCA's clear advantages over several competing algorithms, including 2D‑CNN, 3D‑CNN, Tri‑CNN, HybridSN, ViT, and Swin Transformer. The source code is publicly available at https://github.com/mqalkhatib/MixerCA.
Authors:Emre Ardıç, Yakup Genç
Abstract:
Federated learning is a machine learning paradigm in which multiple devices collaboratively train a model under the supervision of a central server while ensuring data privacy. However, its performance is often hindered by redundant, malicious, or abnormal samples, leading to model degradation and inefficiency. To overcome these issues, we propose novel sample selection methods for image classification, employing a multitask autoencoder to estimate sample contributions through loss and feature analysis. Our approach incorporates unsupervised outlier detection, using one‑class support vector machine (OCSVM), isolation forest (IF), and adaptive loss threshold (AT) methods managed by a central server to filter noisy samples on clients. We also propose a multi‑class deep support vector data description (SVDD) loss controlled by a central server to enhance feature‑based sample selection. We validate our methods on CIFAR10 and MNIST datasets across varying numbers of clients, non‑IID distributions, and noise levels up to 40%. The results show significant accuracy improvements with loss‑based sample selection, achieving gains of up to 7.02% on CIFAR10 with OCSVM and 1.83% on MNIST with AT. Additionally, our federated SVDD loss further improves feature‑based sample selection, yielding accuracy gains of up to 0.99% on CIFAR10 with OCSVM. These results show the effectiveness of our methods in improving model accuracy across various client counts and noise conditions.
Authors:Zaid Nasser, Mikhail Iumanov, Tianhao Li, Maxim Popov, Jaafar Mahmoud, Sergey Kolyubin
Abstract:
We present RADIO‑ViPE (Reduce All Domains Into One ‑‑ Video Pose Engine), an online semantic SLAM system that enables geometry‑aware open‑vocabulary grounding, associating arbitrary natural language queries with localized 3D regions and objects in dynamic environments. Unlike existing approaches that require calibrated, posed RGB‑D input, RADIO‑ViPE operates directly on raw monocular RGB video streams, requiring no prior camera intrinsics, depth sensors, or pose initialization. The system tightly couples multi‑modal embeddings ‑‑ spanning vision and language ‑‑ derived from agglomerative foundation models (e.g., RADIO) with geometric scene information. This coupling takes place in initialization, optimization and factor graph connections to improve the consistency of the map from multiple modalities. The optimization is wrapped within adaptive robust kernels, designed to handle both actively moving objects and agent‑displaced scene elements (e.g., furniture rearranged during ego‑centric session). Experiments demonstrate that RADIO‑ViPE achieves state‑of‑the‑art results on the dynamic TUM‑RGBD benchmark while maintaining competitive performance against offline open‑vocabulary methods that rely on calibrated data and static scene assumptions. RADIO‑ViPE bridges a critical gap in real‑world deployment, enabling robust open‑vocabulary semantic grounding for autonomous robotics and unconstrained in‑the‑wild video streams. Project page: https://be2rlab.github.io/radio_vipe
Authors:Minh-Khoa Le-Phan, Minh-Hoang Le, Trong-Le Do, Minh-Triet Tran
Abstract:
Current deepfake detection models achieve state‑of‑the‑art performance on pristine academic datasets but suffer severe spatial attention drift under real‑world compound degradations, such as blurring and severe lossy compression. To address this vulnerability, we propose a foundation‑driven forensic framework that integrates an extreme compound degradation engine with a structurally constrained, multi‑stream architecture. During training, our degradation pipeline systematically destroys high‑frequency artifacts, optimizing the DINOv2‑Giant backbone to extract invariant geometric and semantic priors. We then process images through three specialized pathways: a Global Texture stream, a Localized Facial stream, and a Hybrid Semantic Fusion stream incorporating CLIP. Through analyzing spatial attribution via Score‑CAM and feature stability using Cosine Similarity, we quantitatively demonstrate that these streams extract non‑redundant, complementary feature representations and stabilize attention entropy. By aggregating these predictions via a calibrated, discretized voting mechanism, our ensemble successfully suppresses background attention drift while acting as a robust geometric anchor. Our approach yields highly stable zero‑shot generalization, achieving Fourth Place in the NTIRE 2026 Robust Deepfake Detection Challenge at CVPR. Code is available at https://github.com/khoalephanminh/ntire26‑deepfake‑challenge.
Authors:Hector G. Rodriguez, Marcus Rohrbach
Abstract:
Multimodal large language models (MLLMs) achieve ever‑stronger performance on visual‑language tasks. Even as traditional visual question answering (VQA) benchmarks approach saturation, reliable deployment requires satisfying low error tolerances in real‑world, out‑of‑distribution (OOD) scenarios. Precisely, selective prediction aims to improve coverage, i.e. the share of inputs the system answers, while adhering to a user‑defined risk level. This is typically achieved by assigning a confidence score to each answer and abstaining on those that fall below a certain threshold. Existing selective prediction methods estimate implicit confidence scores, relying on model internal signals like logits or hidden representations, which are not available for frontier closed‑sourced models. To enable reliable generalization in VQA, we require reasoner models to produce localized visual evidence while answering, and design a selector that explicitly learns to estimate the quality of the localization provided by the reasoner using only model inputs and outputs. We show that SIEVES (Selective Prediction through Visual Evidence Scoring) improves coverage by up to three times on challenging OOD benchmarks (V Bench, HR‑Bench‑8k, MME‑RealWorld‑Lite, VizWiz, and AdVQA), compared to non‑grounding baselines. Beyond better generalization to OOD tasks, the design of the SIEVES selector enables transfer to proprietary reasoners without access to their weights or logits, such as o3 and Gemini‑3‑Pro, providing coverage boosts beyond those attributable to accuracy alone. We highlight that SIEVES generalizes across all tested OOD benchmarks and reasoner models (Pixel‑Reasoner, o3, and Gemini‑3‑Pro), without benchmark‑ or reasoner‑specific training or adaptation. Code is publicly available at https://github.com/hector‑gr/SIEVES .
Authors:Tri-Nhan Vo, Dang Nguyen, Kien Do, Sunil Gupta
Abstract:
Knowledge distillation (KD) is a well‑known technique to effectively compress a large network (teacher) to a smaller network (student) with little sacrifice in performance. However, most KD methods require a large training set and internal access to the teacher, which are rarely available due to various restrictions. These challenges have originated a more practical setting known as black‑box few‑shot KD, where the student is trained with few images and a black‑box teacher. Recent approaches typically generate additional synthetic images but lack an active strategy to promote their diversity, a crucial factor for student learning. To address these problems, we propose a novel training scheme for generative adversarial networks, where we adaptively select high‑confidence images under the teacher's supervision and introduce them to the adversarial learning on‑the‑fly. Our approach helps expand and improve the diversity of the distillation set, significantly boosting student accuracy. Through extensive experiments, we achieve state‑of‑the‑art results among other few‑shot KD methods on seven image datasets. The code is available at https://github.com/votrinhan88/divbfkd.
Authors:Yi Yang, Hao Pan, Yijing Cui, Alla Sheffer, Changjian Li
Abstract:
Articulation modeling aims to infer movable parts and their motion parameters for a 3D object, enabling interactive animation, simulation, and shape editing. In this paper, we present Sketch2Arti, the first sketch‑based articulation modeling system for CAD objects. Our key observation is that designers naturally communicate articulation intent through lightweight sketches (e.g., arrows and strokes) that indicate how parts should move, yet translating such sketches into articulated 3D models remains largely manual. Sketch2Arti bridges this gap by enabling users to specify articulation through simple 2D sketches drawn from a chosen viewpoint. Given a CAD model and user sketches, our approach automatically discovers the corresponding movable parts and predicts their motion parameters, allowing iterative modeling of multiple articulations on complex objects with fine‑grained control. Importantly, Sketch2Arti is trained in a category‑agnostic manner without requiring object category information, leading to strong generalization to diverse objects beyond existing articulation datasets. Moreover, for shell models lacking interior structures, Sketch2Arti supports controllable internal completion guided by user sketches, generating plausible internal components consistent with the existing geometry and predicted motion constraints. Comprehensive experiments and user evaluations demonstrate the effectiveness, controllability, and generalization of Sketch2Arti. The code, dataset, and the prototype system are at https://arlo‑yang.github.io/Sketch2Arti.
Authors:Jing Zhang, Duojie Chen, Wentao Jiang, Zihan Lou, Jianxin Liu, Xinwu Cui, Qinghong Zhao, Bo Du, Christoph F. Dietrich, Dacheng Tao
Abstract:
Robotic ultrasound has advanced local image‑driven control, contact regulation, and view optimization, yet current systems lack the anatomical understanding needed to determine what to scan, where to begin, and how to adapt to individual patient anatomy. These gaps make systems still reliant on expert intervention to initiate scanning. Here we present SAMe, a semantic anatomy mapping engine that provides robotic ultrasound with an explicit anatomical prior layer. SAMe addresses scan initiation as a target‑to‑anatomy‑to‑action process: it grounds under‑specified clinical complaints into structured target organs, instantiates a patient‑specific anatomical representation for the grounded targets from a single external body image, and translates this representation into control‑facing 6‑DoF probe initialization states without any additional registration using preoperative CT or MRI. The anatomical representation maintained by SAMe is explicit, lightweight (single‑organ inference in 0.08s), and compatible with downstream control by design. Across semantic grounding, anatomical instantiation, and real‑robot evaluation, SAMe shows strong performance across the full initialization pipeline. In real‑robot experiments, centroid‑based SAMe initialization outperformed the body‑keypoint‑based heuristic baseline under a budget‑matched single‑target setting for both liver (86.7% versus 46.7%) and kidney (80.0% versus 73.3%) initialization. Furthermore, The trial‑level organ‑hit rate reached 97.3% for liver and 83.3% for kidney when multiple candidate targets were available. These results establish an explicit anatomical prior layer that addresses scan initialization and is designed to support broader downstream autonomous scanning pipelines, providing the anatomical foundation for complaint‑driven, anatomically informed robotic ultrasonography.
Authors:Chengsheng Zhang, Chenghao Sun, Xinyan Jiang, Wei Li, Xinmei Tian
Abstract:
Large Vision‑Language Models (LVLMs) have achieved remarkable progress in visual‑textual understanding, yet their reliability is critically undermined by hallucinations, i.e., the generation of factually incorrect or inconsistent responses. While recent studies using steering vectors demonstrated promise in reducing hallucinations, a notable challenge remains: they inadvertently amplify the severity of residual hallucinations. We attribute this to their exclusive focus on the decoding stage, where errors accumulate autoregressively and progressively worsen subsequent hallucinatory outputs. To address this, we propose Prefill‑Time Intervention (PTI), a novel steering paradigm that intervenes only once during the prefill stage, enhancing the initial Key‑Value (KV) cache before error accumulation occurs. Specifically, PTI is modality‑aware, deriving distinct directions for visual and textual representations. This intervention is decoupled to steer keys toward visually‑grounded objects and values to filter background noise, correcting hallucination‑prone representations at their source. Extensive experiments demonstrate PTI's significant performance in mitigating hallucinations and its generalizability across diverse decoding strategies, LVLMs, and benchmarks. Moreover, PTI is orthogonal to existing decoding‑stage methods, enabling plug‑and‑play integration and further boosting performance. Code is available at: https://github.com/huaiyi66/PTI.
Authors:Jiayi Guo, Linqing Wang, Jiangshan Wang, Yang Yue, Zeyu Liu, Zhiyuan Zhao, Qinglin Lu, Gao Huang, Chunyu Wang
Abstract:
Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text‑to‑image (T2I) tasks, this unified capability allows UMMs to refine outputs after their initial generation, potentially extending the performance upper bound. Current UMM‑based refinement methods primarily follow a refinement‑via‑editing (RvE) paradigm, where UMMs produce editing instructions to modify misaligned regions while preserving aligned content. However, editing instructions often describe prompt‑image misalignment only coarsely, leading to incomplete refinement. Moreover, pixel‑level preservation, though necessary for editing, unnecessarily restricts the effective modification space for refinement. To address these limitations, we propose Refinement via Regeneration (RvR), a novel framework that reformulates refinement as conditional image regeneration rather than editing. Instead of relying on editing instructions and enforcing strict content preservation, RvR regenerates images conditioned on the target prompt and the semantic tokens of the initial image, enabling more complete semantic alignment with a larger modification space. Extensive experiments demonstrate the effectiveness of RvR, improving Geneval from 0.78 to 0.91, DPGBench from 84.02 to 87.21, and UniGenBench++ from 61.53 to 77.41.
Authors:Junchao Cui, Wenqi Shi, Shaoyong Du, Hang He, Xuanzi Ma, Hao Tang, Xiangyang Luo
Abstract:
Worldwide image geo‑localization aims to infer the geographic location of an image captured anywhere on Earth, spanning street, city, regional, national, and continental scales. Existing methods rely on visual features that are sensitive to environmental variations (e.g., lighting, season, and weather) and lack effective post‑processing to filter outlier candidates, limiting localization accuracy. To address these limitations, we propose DualGeo, a two‑stage framework for worldwide image geo‑localization. First, it establishes a geo‑representational foundation by fusing image and semantic segmentation features via bidirectional cross‑attention. The fused features are then aligned with GPS coordinates through dual‑view contrastive learning to build a global retrieval database. Second, it performs geo‑cognitive refinement by re‑ranking retrieved candidates using geographic clustering. It then feeds them into large multimodal models (LMMs) for final coordinate prediction. Experiments on IM2GPS, IM2GPS3k, and YFCC4k show that DualGeo outperforms state‑of‑the‑art methods, improving street‑level (<1 km) and city‑level (<25 km) localization accuracy by 3.6%‑16.58% and 1.29%‑8.77%, respectively. Our code and datasets are available : https://github.com/CJ310177/DualGeo.
Authors:Fabio D'Oronzio, Federico Putamorsi, Leonardo Zini, Marcella Cornia, Lorenzo Baraldi
Abstract:
Despite recent advances, single‑image super‑resolution (SR) remains challenging, especially in real‑world scenarios with complex degradations. Diffusion‑based SR methods, particularly those built on Stable Diffusion, leverage strong generative priors but commonly rely on text conditioning derived from semantic captioning. Such textual descriptions provide only high‑level semantics and lack the spatially aligned visual information required for faithful restoration, leading to a representation gap between abstract semantics and spatially aligned visual details. To address this limitation, we propose GramSR, a one‑step diffusion‑based SR framework that replaces text conditioning with dense visual features extracted from the low‑resolution input using a pre‑trained DINOv3 encoder. GramSR adopts a three‑stage LoRA architecture, where pixel‑level, semantic‑level, and texture‑level LoRA modules are trained sequentially. The pixel‑level module focuses on degradation removal using \ell_2 loss, the semantic‑level module enhances perceptual details via LPIPS and CSD losses, and the texture‑level module enforces feature correlation consistency through a Gram matrix loss computed from DINOv3 features. At inference, independent guidance scales enable flexible control over degradation removal, semantic enhancement, and texture preservation. Extensive experiments on standard SR benchmarks demonstrate that GramSR consistently outperforms existing one‑step diffusion‑based methods, achieving superior structural fidelity and texture realism. The code for this work is available at: https://github.com/aimagelab/GramSR.
Authors:Zi-Yang Bo, Wei Lu, Hongruixuan Chen, Si-Bao Chen, Bin Luo
Abstract:
Shadows are a prevalent problem in remote sensing imagery (RSI), degrading visual quality and severely limiting the performance of downstream tasks like object detection and semantic segmentation. Most prior works treat shadow detection and removal as separate, cascaded tasks, which can lead to cumbersome process and error accumulation. Furthermore, many deep learning methods rely on paired shadow and non‑shadow images for training, which are often unavailable in practice. To address these challenges, we propose Shadow‑Aware and Removal Unified (SARU) Framework , a cohesive two‑stage framework. First, its dual‑branch detection module (DBCSF‑Net) fuses multi‑color space and semantic features to generate high‑fidelity shadow masks, effectively distinguishing shadows from dark objects. Then, leveraging these masks, a novel, training‑free physical algorithm (N^2SGSR) restores illumination by transferring properties from adjacent non‑shadow regions within the single input image. To facilitate rigorous evaluation and foster future work, we also introduce two new benchmark datasets: the RSI Shadow Detection (RSISD) dataset and the Single‑image Shadow Removal Benchmark (SiSRB). Extensive experiments on the AISD and RSISD datasets demonstrate that SARU achieves SOTA shadow detection performance. For shadow removal, our training‑free N^2SGSR algorithm attains an average processing speed of approximately 1.3s, which is over 10 times faster than the SOTA MAOSD while maintains an SRI value close to 0.9 on both the AISD and SiSRB datasets, a level comparable to the advanced RS‑GSSR method. By holistically integrating shadow detection and removal to mitigate error propagation and eliminating the dependency on paired training data, SARU establishes a robust, practical framework for real‑world RSI analysis. The code and datasets are publicly available at: https://github.com/AeroVILab‑AHU/SARU
Authors:Jianyu Wen, Jun Xie, Feng Chen, Zhepeng Wang, Chenhao Wu, Tong Zhang, Yixuan Yu, Piotr Swierczynski
Abstract:
In this paper, we present Self‑DACE++, an improved unsupervised and lightweight framework for Low‑Light Image Enhancement (LLIE), building upon our previous Self‑Reference Deep Adaptive Curve Estimation (Self‑DACE). To better address the trade‑off between computational efficiency and restoration quality, Self‑DACE++ introduces enhanced Adaptive Adjustment Curves (AACs). These curves, governed by minimal trainable parameters, flexibly adjust the dynamic range while preserving the color fidelity, structural integrity, and naturalness of the enhanced images. To achieve an extremely lightweight architecture without sacrificing performance, we propose a randomized order training strategy coupled with a network fusion mechanism, which compresses the model into an efficient iterative inference structure. Furthermore, we formulate a physics‑grounded objective function based on Retinex theory and incorporate a dedicated denoising module to effectively estimate and suppress latent noise in dark regions. Extensive qualitative and quantitative evaluations on multiple real‑world benchmark datasets demonstrate that Self‑DACE++ outperforms existing state‑of‑the‑art methods, delivering superior enhancement quality with real‑time inference capability. The code is available at https://github.com/John‑Wendell/Self‑DACE.
Authors:Luca Parolari, Nicla Faccioli, Lamberto Ballan
Abstract:
Evaluating layout‑guided text‑to‑image generative models requires assessing both semantic alignment with textual prompts and spatial fidelity to prescribed layouts. Assessing layout alignment requires collecting fine‑grained annotations, which is costly and labor‑intensive. Consequently, current benchmarks rarely provide comprehensive layout evaluation and often remain limited in scale or coverage, making model comparison, ranking, and interpretation difficult. In this work, we introduce a closed‑set benchmark (C‑Bench) designed to isolate key generative capabilities while providing varying levels of complexity in both prompt structure and layout. To complement this controlled setting, we propose an open‑set benchmark (O‑Bench) that evaluates models using real‑world prompts and layouts, offering a measure of semantic and spatial alignment in the wild. We further develop a unified evaluation protocol that combines semantic and spatial accuracy into a single score, ensuring consistent model ranking. Using our benchmarks, we conduct a large‑scale evaluation of six state‑of‑the‑art layout‑guided diffusion models, totaling 319,086 generated and evaluated images. We establish a model ranking based on their overall performance and provide detailed breakdowns for text and layout alignment to enhance interpretability. Fine‑grained analyses across scenarios and prompt complexities highlight the strengths and limitations of current models. Code is available at https://github.com/lparolari/cobench.
Authors:Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng, Yang Liu
Abstract:
Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open‑world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and rare concepts. To overcome these limitations, we introduce OmniVTG, a new large‑scale dataset for open‑world VTG, coupled with a Self‑Correction Chain‑of‑Thought (CoT) training paradigm designed to enhance the grounding capabilities of Multimodal Large Language Models (MLLMs). Our OmniVTG is constructed via a novel Semantic Coverage Iterative Expansion pipeline, which first identifies gaps in the vocabulary of existing datasets and collects videos that are highly likely to contain these target concepts. For high‑quality annotation, we leverage the insight that modern MLLMs excel at dense captioning more than direct grounding and design a caption‑centric data engine to prompt MLLMs to generate dense, timestamped descriptions. Beyond the dataset, we observe that simple supervised finetuning (SFT) is insufficient, as a performance gap between rare and common concepts still persists. We find that MLLMs' video understanding ability significantly surpasses their direct grounding ability. Based on this, we propose a Self‑Correction Chain‑of‑Thought (CoT) training paradigm. We train the MLLM to first predict, then use its understanding capabilities to reflect on and refine its own predictions. This capability is instilled via a three‑stage pipeline of SFT, CoT finetuning, and reinforcement learning. Extensive experiments show our approach not only excels at open‑world grounding in our OmniVTG dataset but also achieves state‑of‑the‑art zero‑shot performance on four existing VTG benchmarks. Code is available at https://github.com/oceanflowlab/OmniVTG.
Authors:Divake Kumar, Sina Tayebati, Devashri Naik, Ranganath Krishnan, Amit Ranjan Trivedi
Abstract:
Vision‑language models (VLMs) are increasingly used as automated judges for multimodal systems, yet their scores provide no indication of reliability. We study this problem through conformal prediction, a distribution‑free framework that converts a judge's point score into a calibrated prediction interval using only score‑token log‑probabilities, with no retraining. We present the first systematic analysis of conformal prediction for VLM‑as‑a‑Judge across 3 judges and 14 visual task categories. Our results show that evaluation uncertainty is strongly task‑dependent: intervals cover ~40% of the score range for aesthetics and natural images but expand to ~70% for chart and mathematical reasoning, yielding a quantitative reliability map for multimodal evaluation. We further identify a failure mode not captured by standard evaluation metrics, ranking‑scoring decoupling, where judges achieve high ranking correlation while producing wide, uninformative intervals, correctly ordering responses but failing to assign reliable absolute scores. Finally, we show that interval width is driven primarily by task difficulty and annotation quality, i.e., the same judge and method yield 4.5x narrower intervals on a clean, multi‑annotator captioning benchmark. Code: https://github.com/divake/VLM‑Judge‑Uncertainty
Authors:Wenqi Jia, Zekun Li, Abhay Mittal, Chengcheng Tang, Chuan Guo, Lezi Wang, James Matthew Rehg, Lingling Tao, Size An
Abstract:
Recent advances in text‑driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity‑neutral motion and generate movements using a canonical body representation, ignoring the strong influence of body morphology on motion dynamics. In practice, attributes such as body proportions, mass distribution, and age significantly affect how actions are performed, and neglecting this coupling often leads to physically inconsistent motions. We propose an identity‑aware motion generation framework that explicitly models the relationship between body morphology and motion dynamics. Instead of relying on explicit geometric measurements, identity is represented using multimodal signals, including natural language descriptions and visual cues. We further introduce a joint motion‑shape generation paradigm that simultaneously synthesizes motion sequences and body shape parameters, allowing identity cues to directly modulate motion dynamics. Extensive experiments on motion capture datasets and large‑scale in‑the‑wild videos demonstrate improved motion realism and motion‑identity consistency while maintaining high motion quality. Project page: https://vjwq.github.io/IAM
Authors:Jiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao, Dongze Hao, Xuanxu Lin, Jing Liu
Abstract:
We present M^3‑VQA, a novel knowledge‑based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine‑grained multimodal entity understanding and complex multi‑hop reasoning. Unlike existing VQA datasets that focus on coarse‑grained categories and simple reasoning over single entities, M^3‑VQA introduces diverse multi‑entity questions involving multiple distinct entities from both visual and textual sources. It requires models to perform both sequential and parallel multi‑hop reasoning across multiple documents, supported by traceable, detailed evidence and a curated multimodal knowledge base. We evaluate 16 leading MLLMs under three settings: without external knowledge, with gold evidence, and with retrieval‑augmented input. The poor results reveal significant challenges for MLLMs in knowledge acquisition and reasoning. Models perform poorly without external information but improve markedly when provided with precise evidence. Furthermore, reasoning‑aware agentic retrieval surpasses heuristic methods, highlighting the importance of structured reasoning for complex multimodal understanding. M^3‑VQA presents a more challenging evaluation for advancing the multimodal reasoning capabilities of MLLMs. Our code and dataset are available at https://github.com/CASIA‑IVA‑Lab/M3VQA.
Authors:Ming Li, Jie Wu, Justin Cui, Xiaojie Li, Rui Wang, Chen Chen
Abstract:
While preference optimization is crucial for improving visual generative models, how to effectively scale this paradigm remains largely unexplored. Current open‑source preference datasets contain conflicting preference patterns, where winners excel in some dimensions but underperform in others. Naively optimizing on such noisy datasets fails to learn preferences, hindering effective scaling. To enhance robustness against noise, we propose Poly‑DPO, which extends the DPO objective with an additional polynomial term that dynamically adjusts model confidence based on dataset characteristics, enabling effective learning across diverse data distributions. Beyond biased patterns, existing datasets suffer from low resolution, limited prompt diversity, and imbalanced distributions. To facilitate large‑scale visual preference optimization by tackling data bottlenecks, we construct ViPO, a massive‑scale preference dataset with 1M image pairs at 1024px across five categories and 300K video pairs at 720p+ across three categories. State‑of‑the‑art generative models and diverse prompts ensure reliable preference signals with balanced distributions. Remarkably, when applying Poly‑DPO to our high‑quality dataset, the optimal configuration converges to standard DPO. This convergence validates dataset quality and Poly‑DPO's adaptive nature: sophisticated optimization becomes unnecessary with sufficient data quality, yet remains valuable for imperfect datasets. We validate our approach across visual generation models. On noisy datasets like Pick‑a‑Pic V2, Poly‑DPO achieves 6.87 and 2.32 gains over Diffusion‑DPO on GenEval for SD1.5 and SDXL, respectively. For ViPO, models achieve performance far exceeding those trained on existing open‑source preference datasets. These results confirm that addressing both algorithmic adaptability and data quality is essential for scaling visual preference optimization.
Authors:Xinxin Liu, Ming Li, Zonglin Lyu, Yuzhang Shang, Chen Chen
Abstract:
Human visual preferences are inherently multi‑dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets provide only single, holistic annotations, resulting in severe label noise: images that excel in some dimensions but are deficient in others are simply marked as winner or loser. We theoretically demonstrate that compressing multi‑dimensional preferences into binary labels generates conflicting gradient signals that misguide Diffusion Direct Preference Optimization (DPO). To address this, we propose Semi‑DPO, a semi‑supervised approach that treats consistent pairs as clean labeled data and conflicting ones as noisy unlabeled data. Our method starts by training on a consensus‑filtered clean subset, then uses this model as an implicit classifier to generate pseudo‑labels for the noisy set for iterative refinement. Experimental results demonstrate that Semi‑DPO achieves state‑of‑the‑art performance and significantly improves alignment with complex human preferences, without requiring additional human annotation or explicit reward models during training. We will release our code and models at: https://github.com/L‑CodingSpace/semi‑dpo
Authors:Cheng-Han Lee, Maniratnam Mandal, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
Abstract:
With the rise of mobile video consumption on diverse handheld display resolutions and orientation modes, altering videos to aspect ratios poses challenges. Static cropping and border padding often compromises visual quality, while warping may distort a video's intended meaning. Here we advocate for a more effective approach: cropping significant regions within video frames in a temporal manner, while minimizing distortion and preserving essential content. One barrier to solving this problem is the lack of sufficiently large‑scale database devoted to informing these tasks. Towards filling this gap, we introduce the LIVE‑YouTube Video Cropping (LIVE‑YT VC) database, featuring 1800 videos, annotated by 90 human subjects. Using videos sourced from the YouTube‑UGC and LSVQ Databases, this new resource is the largest publicly‑available subjective video portrait region cropping database. We also introduce a post‑processed version of the database, called LIVE‑YT VC++, whereby a novel intra‑frame temporal filter was deployed to smooth subjective annotations within each video. We demonstrate the usefulness of this new data resource using the SmartVidCrop algorithm and state‑of‑the‑art video grounding models, in hopes of establishing our subjective dataset as a benchmark for future research. Our contributions offer a resource for advancing video aspect ratio transformation models towards ensuring that reshaped mobile‑friendly video content retains its quality and meaning. Since our labels bear resemblances to video saliency annotations, we also conducted an additional analysis to explore the similarity between our labels and video saliency predictions. Finally, we repurposed state‑of‑the‑art video grounding models for aspect ratio change tasks, and fine‑tuned them on our dataset. As a service to the research community, we plan to open source the project.
Authors:Antoine P. Leeman, Shuyu Zhan, Melanie N. Zeilinger, Glen Chou
Abstract:
We propose VISION‑SLS, a method for nonlinear output‑feedback control from high‑resolution RGB images which provides robust constraint satisfaction guarantees under calibrated uncertainty bounds despite partial observability, sensor noise, and nonlinear dynamics. To enable scalability while retaining guarantees, we propose: (i) a learned low‑dimensional observation map from pretrained visual features with state‑dependent error bounds, and (ii) a causal affine time‑varying output‑feedback policy optimized via System Level Synthesis (SLS). We develop a scalable, novel solver for the resulting nonconvex program that leverages sequential convex programming coupled with efficient Riccati recursions. On two simulated visuomotor tasks (a 4D car and a 10D quadrotor) with >= 512 x 512 pixels and a 59D humanoid task with partial observability, our method enables safe, information‑gathering behavior that reduces uncertainty while guaranteeing constraint satisfaction with empirically‑calibrated error bounds. We also validate our method on hardware, safely controlling a ground vehicle from onboard images, outperforming baselines in safety rate and solve times. Together, these results show that learned visual abstractions coupled with an efficient solver make SLS‑based safe visuomotor output‑feedback practical at scale. The code implementation of our method is available at https://github.com/trustworthyrobotics/VISION‑SLS.
Authors:Nikesh Subedi, Loris Bazzani, Ziad Al-Halah
Abstract:
In episodic memory with natural language queries (EM‑NLQ), a user may ask a question (e.g., "Where did I place the mug?") that requires searching a long egocentric video, captured from the user's perspective, to find the moment that answers it. However, queries can be ambiguous or incomplete, leading to incorrect responses. Current methods ignore this key aspect and address EM‑NLQ in a one‑shot setup, limiting their applicability in real‑world scenarios. In this work, we address this gap and introduce the Episodic Memory with Questions and Feedback task (EM‑QnF). Here, the user can provide feedback on the model's initial prediction or add more information (e.g., "Before this. I'm looking for the big blue mug not the white one"), helping the model refine its predictions interactively. To this end, we collect datasets for feedback‑based interaction and propose a lightweight training scheme that avoids expensive sequential optimization. We also introduce a plug‑and‑play Feedback ALignment Module (FALM) that enables existing EM‑NLQ models to incorporate user feedback effectively. Our approach significantly improves over the state of the art on three challenging benchmarks and is better than or competitive with commercial large vision‑language models while remaining efficient. Evaluation with human‑generated feedback shows that it generalizes well to real‑world scenarios.
Authors:Maitreya Patel, Jingtao Li, Weiming Zhuang, Yezhou Yang, Lingjuan Lv
Abstract:
We introduce an efficient, resolution‑agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios, narrowing the gap to diffusion models at scale. At its core is VibeToken, a novel resolution‑agnostic 1D Transformer‑based image tokenizer that encodes images into a dynamic, user‑controllable sequence of 32‑256 tokens, achieving a state‑of‑the‑art efficiency and performance trade‑off. Building on VibeToken, we present VibeToken‑Gen, a class‑conditioned AR generator with out‑of‑the‑box support for arbitrary resolutions while requiring significantly fewer compute resources. Notably, VibeToken‑Gen synthesizes 1024x1024 images using only 64 tokens and achieves 3.94 gFID; by comparison, a diffusion‑based state‑of‑the‑art alternative requires 1,024 tokens and attains 5.87 gFID. In contrast to fixed‑resolution AR models such as LlamaGen ‑‑ whose inference FLOPs grow quadratically with resolution (11T FLOPs at 1024x1024) ‑‑ VibeToken‑Gen maintains a constant 179G FLOPs (63.4x efficient) independent of resolution. We hope VibeToken can help unlock the wide adoption of AR visual generative models in production use cases.
Authors:Nishit Anand, Manan Suri, Christopher Metzler, Dinesh Manocha, Ramani Duraiswami
Abstract:
Controlling illumination in images is essential for photography and visual content creation. While closed‑source models have demonstrated impressive illumination control, open‑source alternatives either require heavy control inputs like depth maps or do not release their data and code. We present a fully open‑source and reproducible pipeline for learning illumination control in diffusion models. Our approach builds a data engine that transforms well‑lit images into supervised training triplets consisting of a poorly‑illuminated input image, a natural language lighting instruction, and a well‑illuminated output image. We finetune a diffusion model on this data and demonstrate significant improvements over baseline SD 1.5, SDXL, and FLUX.1‑dev models in perceptual similarity, structural similarity, and identity preservation. Our work provides a reproducible solution built entirely with open‑source tools and publicly available data. We release all our code, data, and model weights publicly.
Authors:Yu Xin, Gorkem Can Ates, Jun Ma, Sumin Kim, Ying Zhang, Kaleb E Smith, Kuang Gong, Wei Shao
Abstract:
Text guided 3D medical image segmentation offers a flexible alternative to class based and spatial prompt based models by allowing users to specify regions of interest directly in natural language. This paradigm avoids reliance on predefined label sets, reduces ambiguous outputs, and aligns more naturally with clinical workflows. However, existing text guided frameworks are often computationally expensive, exhibit weak text volume feature alignment, and fail to capture fine anatomical details. We propose ESICA, a lightweight and scalable framework that addresses these challenges through three innovations: (1) a similarity matrix based mask prediction formulation that enhances semantic alignment, (2) an efficient decomposed decoder with adapter modules for accurate volumetric decoding, and (3) a two pass refinement strategy that sharpens boundaries and resolves uncertain regions. To improve training stability and generalization, ESICA adopts a two stage scheme consisting of positive only pretraining followed by balanced fine tuning. On the CVPR BiomedSegFM benchmark spanning five imaging modalities (CT, MRI, PET, ultrasound, and microscopy), ESICA achieves state of the art segmentation accuracy, while the compact ESICA4 Lite variant attains similar segmentation performance with substantially fewer parameters, yielding a superior efficiency accuracy trade off. Our framework advances text guided segmentation toward efficient, scalable, and clinically deployable systems. Code will be made publicly available at https://github.com/mirthAI/ESICA.
Authors:Weijie Wang, Xiaoxuan He, Youping Gu, Yifan Yang, Zeyu Zhang, Yefei He, Yanbo Ding, Xirui Hu, Donny Y. Chen, Zhiyuan He, Yuqing Yang, Bohan Zhuang
Abstract:
Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high computational costs and limit scalability. We propose World‑R1, a framework that aligns video generation with 3D constraints through reinforcement learning. To facilitate this alignment, we introduce a specialized pure text dataset tailored for world simulation. Utilizing Flow‑GRPO, we optimize the model using feedback from pre‑trained 3D foundation models and vision‑language models to enforce structural coherence without altering the underlying architecture. We further employ a periodic decoupled training strategy to balance rigid geometric consistency with dynamic scene fluidity. Extensive evaluations reveal that our approach significantly enhances 3D consistency while preserving the original visual quality of the foundation model, effectively bridging the gap between video generation and scalable world simulation.
Authors:Cheng Wang, Zhibin He, Zhihao Peng, Shengyuan Liu, Yufan Hu, Yang Carl, He Lifang, Lichao Sun, Xiang Li, Yixuan Yuan
Abstract:
Agentic artificial intelligence systems promise to accelerate scientific workflows, but neuroimaging poses unique challenges: heterogeneous modalities (sMRI, fMRI, dMRI, EEG), long multi‑stage pipelines, and persistent reproducibility risks. To address this gap, we present NeuroClaw, a domain‑specialized multi‑agent research assistant for executable and reproducible neuroimaging research. NeuroClaw operates directly on raw neuroimaging data across formats and modalities, grounding decisions in dataset semantics and BIDS metadata so users need not prepare curated inputs or bespoke model code. The platform combines harness engineering with end‑to‑end environment management, including pinned Python environments, Docker support, automated installers for common neuroimaging tools, and GPU configuration. In practice, this layer emphasizes checkpointing, post‑execution verification, structured audit traces, and controlled runtime setup, making toolchains more transparent while improving reproducibility and auditability. A three‑tier skill/agent hierarchy separates user‑facing interaction, high‑level orchestration, and low‑level tool skills to decompose complex workflows into safe, reusable units. Alongside the NeuroClaw framework, we introduce NeuroBench, a system‑level benchmark for executability, artifact validity, and reproducibility readiness. Across multiple multimodal LLMs, NeuroClaw‑enabled runs yield consistent and substantial score improvements compared with direct agent invocation. Project homepage: https://cuhk‑aim‑group.github.io/NeuroClaw/index.html
Authors:Hai Wang, Xiaochen Yang, Mingzhi Dong, Jing-Hao Xue
Abstract:
The dream of instantly creating rich 360‑degree panoramic worlds from text is rapidly becoming a reality, yet a crucial gap exists in our ability to reliably evaluate their semantic alignment. Contrastive Language‑Image Pre‑training (CLIP) models, standard AI evaluators, predominantly trained on perspective image‑text pairs, face an open question regarding their understanding of the unique characteristics of 360‑degree panoramic image‑text pairs. This paper addresses this gap by first introducing two concepts: \emph360‑degree textual semantics, semantic information conveyed by explicit format identifiers, and \emph360‑degree visual semantics, invariant semantics under horizontal circular shifts. To probe CLIP's comprehension of these semantics, we then propose novel evaluation methodologies using keyword manipulation and horizontal circular shifts of varying magnitudes. Rigorous statistical analyses across popular CLIP configurations reveal that: (1) CLIP models effectively leverage explicit textual identifiers, demonstrating an understanding of 360‑degree textual semantics; and (2) CLIP models fail to robustly preserve semantic alignment under horizontal circular shifts, indicating limited comprehension of 360‑degree visual semantics. To address this limitation, we propose a LoRA‑based fine‑tuning framework that explicitly instills invariance to circular shifts. Our fine‑tuned models exhibit improved comprehension of 360‑degree visual semantics, though with a slight degradation in original semantic evaluation performance, highlighting a fundamental trade‑off in adapting CLIP to 360‑degree panoramic images. Code is available at https://github.com/littlewhitesea/360Semantics.
Authors:Shiyi Zhang, Yiji Cheng, Tiankai Hang, Zijin Yin, Runze He, Yu Xu, Wenxun Dai, Yunlong Lin, Chunyu Wang, Qinglin Lu, Yansong Tang
Abstract:
Unified multi‑modal understanding/generative models have shown improved image editing performance by incorporating fine‑grained understanding into their Chain‑of‑Thought (CoT) process. However, a critical question remains underexplored: what forms of CoT and training strategy can jointly enhance both the understanding granularity and generalization? To address this, we propose Meta‑CoT, a paradigm that performs a two‑level decomposition of any single‑image editing operation with two key properties: (1) Decomposability. We observe that any editing intention can be represented as a triplet ‑ (task, target, required understanding ability). Inspired by this, Meta‑CoT decomposes both the editing task and the target, generating task‑specific CoT and traversing editing operations on all targets. This decomposition enhances the model's understanding granularity of editing operations and guides it to learn each element of the triplet during training, substantially improving the editing capability. (2) Generalizability. In the second decomposition level, we further break down editing tasks into five fundamental meta‑tasks. We find that training on these five meta‑tasks, together with the other two elements of the triplet, is sufficient to achieve strong generalization across diverse, unseen editing tasks. To further align the model's editing behavior with its CoT reasoning, we introduce the CoT‑Editing Consistency Reward, which encourages more accurate and effective utilization of CoT information during editing. Experiments demonstrate that our method achieves an overall 15.8% improvement across 21 editing tasks, and generalizes effectively to unseen editing tasks when trained on only a small set of meta‑tasks. Our code, benchmark, and model are released at https://shiyi‑zh0408.github.io/projectpages/Meta‑CoT/
Authors:Fan Du, Feng Yan, Jianxiong Wu, Xinrun Xu, Weiye Zhang, Weinong Wang, Yu Guo, Bin Qian, Zhihai He, Fei Wang, Heng Yang
Abstract:
Flow‑based vision‑language‑action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi‑step inference is required to recover action structure from uninformative Gaussian noise, leading to a poor efficiency‑quality trade‑off under real‑time constraints. We address this issue by rethinking the role of the starting point in generative action modeling. Instead of shortening the sampling trajectory, we propose CF‑VLA, a coarse‑to‑fine two‑stage formulation that restructures action generation into a coarse initialization step that constructs an action‑aware starting point, followed by a single‑step local refinement that corrects residual errors. Concretely, the coarse stage learns a conditional posterior over endpoint velocity to transform Gaussian noise into a structured initialization, while the fine stage performs a fixed‑time refinement from this initialization. To stabilize training, we introduce a stepwise strategy that first learns a controlled coarse predictor and then performs joint optimization. Experiments on CALVIN and LIBERO show that our method establishes a strong efficiency‑performance frontier under low‑NFE (Number of Function Evaluations) regimes: it consistently outperforms existing NFE=2 methods, matches or surpasses the NFE=10 π_0.5 baseline on several metrics, reduces action sampling latency by 75.4%, and achieves the best average real‑robot success rate of 83.0%, outperforming MIP by 19.5 points and π_0.5 by 4.0 points. These results suggest that structured, coarse‑to‑fine generation enables both strong performance and efficient inference. Our code is available at https://github.com/EmbodiedAI‑RoboTron/CF‑VLA.
Authors:Yingqian Min, Kun Zhou, Yifan Li, Yuhuan Wu, Han Peng, Yifan Du, Wayne Xin Zhao, Min Yang, Ji-Rong Wen
Abstract:
Recent advancements in reinforcement learning with verifiable rewards (RLVR) have significantly improved the complex reasoning ability of vision‑language models (VLMs). However, its outcome‑level supervision is too coarse to diagnose and correct errors within the reasoning chain. To this end, we propose Perceval, a process reward model (PRM) that enables token‑level error grounding, which can extract image‑related claims from the response and compare them one by one with the visual evidence in the image, ultimately returning claims that contain perceptual errors. Perceval is trained with perception‑intensive supervised training data. We then integrate Perceval into the RL training process to train the policy models. Specifically, compared to traditional GRPO, which applies sequence‑level advantages, we apply token‑level advantages by targeting penalties on hallucinated spans identified by Perceval, thus enabling fine‑grained supervision signals. In addition to augmenting the training process, Perceval can also assist VLMs during the inference stage. Using Perceval, we can truncate the erroneous portions of the model's response, and then either have the model regenerate the response directly or induce the model to reflect on its previous output. This process can be repeated multiple times to achieve test‑time scaling. Experiments show significant improvements on benchmarks from various domains across multiple reasoning VLMs trained with RL, highlighting the promise of perception‑centric supervision as a general‑purpose strategy. For test‑time scaling, it also demonstrates consistent performance gains over other strategies, such as major voting. Our code and data will be publicly released at https://github.com/RUCAIBox/Perceval.
Authors:Dongxing Mao, Yilin Wang, Linjie Li, Zhengyuan Yang, Alex Jinpeng Wang
Abstract:
Despite recent advances in text‑to‑image generation, models still struggle to accurately render prompt‑specified text with correct spatial layout ‑‑ especially in multi‑span, structured settings. This challenge is driven not only by the lack of datasets that align prompts with the exact text and layout expected in the image, but also by the absence of effective metrics for evaluating layout quality. To address these issues, we introduce TextGround4M, a large‑scale dataset of over 4 million prompt‑image pairs, each annotated with span‑level text grounded in the prompt and corresponding bounding boxes. This enables fine‑grained supervision for layout‑aware, prompt‑grounded text rendering. Building on this, we propose a lightweight training strategy for autoregressive T2I models that appends layout‑aware span tokens during training, without altering model architecture or inference behavior. We further construct a benchmark with stratified layout complexity to evaluate both open‑source and proprietary models in a zero‑shot setting. In addition, we introduce two layout‑aware metrics to address the long‑standing lack of spatial evaluation in text rendering. Our results show that models trained on TextGround4M outperform strong baselines in text fidelity, spatial accuracy, and prompt consistency, highlighting the importance of fine‑grained layout supervision for grounded T2I generation.
Authors:Yangping Li, Thomas Pinetz, Michael Hölzel, Marieta Toma, Alexander Effland
Abstract:
In pathology, the spatial distribution and proportions of tissue types are key indicators of disease progression, and are more readily available than fine‑grained annotations. However, these assessments are rarely mapped to pixel‑wise segmentation. The task is fundamentally underdetermined, as many spatially distinct segmentations can satisfy the same global proportions in the absence of pixel‑wise constraints. To address this, we introduce Variational Segmentation from Label Proportions (VSLP), a two‑stage framework that infers dense segmentations from global label proportions, without any pixel‑level annotations. This framework first leverages a pre‑trained transformer model with test‑time augmentation to produce a pixel‑wise confidence estimate. In the second stage, these estimates are fused by solving a variational optimization problem that incorporates a Wasserstein data fidelity term alongside a learned regularizer. Unlike end‑to‑end networks, our variational method can visualize the fidelity‑regularization energy, resulting in more interpretable segmentation. We validate our approach on two public datasets, achieving superior performance over existing weakly supervised and unsupervised methods. For one of these datasets, proportions have been estimated by an experienced pathologist to provide a realistic benchmark to the community. Furthermore, the method scales to an in‑house dataset with noisy pathologist labels, severely outperforming state‑of‑the‑art methods, thereby demonstrating practical applicability. The code and data will be made publicly available upon acceptance at https://github.com/xiaoliangpi/VSLP.
Authors:Dibyadip Chatterjee, Zhanzhong Pang, Fadime Sener, Yale Song, Angela Yao
Abstract:
Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause the video at fixed timestamps, pose questions about current or past events, and score models only at those moments. This protocol leaves streaming predictions untested. To close this gap, we introduce SPOT‑Bench, featuring multi‑turn proactive queries that evaluate general streaming perception and assistive capabilities required by an always‑on, real‑time assistant. SPOT‑Bench comes with Timeliness‑F1, a consolidated metric that measures streaming predictions by their temporal precision and balanced coverage across the entire video. Our benchmark reveals: (i) offline models detect events reliably but spam predictions unprompted; (ii) post‑training for silence reduces spamming but induces unresponsiveness; (iii) half of the streaming video expects no response, which we term dead‑time ‑ compute spent here does not affect response latency. These findings motivate AsynKV, a training‑free streaming adaptation of offline models, that retains their event perception while improving their streaming behavior. AsynKV features a long‑short term memory, utilized efficiently by scaling compute during dead‑time. It serves as a strong baseline on SPOT‑Bench, outperforming existing streaming models, and achieves state‑of‑the‑art on retrospective benchmarks.
Authors:Yiming Zhang, Jiacheng Chen, Jiaqi Tan, Yongsen Mao, Wenhu Chen, Angel X. Chang
Abstract:
Current evaluations of spatial intelligence can be systematically invalid under modern vision‑language model (VLM) settings. First, many benchmarks derive question‑answer (QA) pairs from point‑cloud‑based 3D annotations originally curated for traditional 3D perception. When such annotations are treated as ground truth for video‑based evaluation, reconstruction and annotation artifacts can miss objects that are clearly visible in the video, mislabel object identities, or corrupt geometry‑dependent answers (e.g., size), yielding incorrect or ambiguous QA pairs. Second, evaluations often assume full‑scene access, while many VLMs operate on sparsely sampled frames (e.g., 16‑64), making many questions effectively unanswerable under the actual model inputs. We improve evaluation validity by introducing ReVSI, a benchmark and protocol that ensures each QA pair is answerable and correct under the model's actual inputs. To this end, we re‑annotate objects and geometry across 381 scenes from 5 datasets to improve data quality, and regenerate all QA pairs with rigorous bias mitigation and human verification using professional 3D annotation tools. We further enhance evaluation controllability by providing variants across multiple frame budgets (16/32/64/all) and fine‑grained object visibility metadata, enabling controlled diagnostic analyses. Evaluations of general and domain‑specific VLMs on ReVSI reveal systematic failure modes that are obscured by prior benchmarks, yielding a more reliable and diagnostic assessment of spatial intelligence.
Authors:Jinkun Dai, Yuanxin Ye, Peng Tang, Tengfeng Tang, Xianping Ma, Jing Xiao, Mi Wang
Abstract:
Semantic segmentation of multi‑modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. Current multi‑modal approaches mainly focus on integrating complementary visual modalities, yet neglect the incorporating of non‑visual textual data ‑ a rich source of knowledge that can bridge semantic gaps between visual patterns and real‑world concepts. To address this limitation, we propose TSMNet, a text supervised multi‑modal open vocabulary semantic segmentation network that synergistically integrates textual supervision with visual representation for open‑vocabulary semantic segmentation. Unlike conventional multi‑modal segmentation frameworks, TSMNet introduces a dual‑branch text encoder to extract both scene‑level semantic and object‑level label information from various textual data, enabling dynamic cross‑modal fusion. These text‑derived features dynamically interact with visual embeddings through the proposed text‑guided visual semantic fusion module, enabling domain‑aware feature refinement and human‑interpretable decision‑making. To verify our method, we innovatively construct two new multi‑modal datasets, and carry out extensive experiments to make a comprehensive comparison between the proposed method and other state‑of‑the‑art (SOTA) semantic segmentation models. Results demonstrate that TSMNet achieves superior segmentation accuracy while exhibiting robust generalization capabilities across diverse geographical and sensor‑specific scenarios. This work establishes a new paradigm for explainable remote sensing analysis, demonstrating that textual knowledge integration significantly enhances model generalizability. The source code will be available at https://github.com/yeyuanxin110/TSMNet
Authors:Jiayi Wang, Lichun Zhang, Xiaoqi Zhuang, Jiaqi Zhang, Lu Yu, Yin Zhao
Abstract:
Video technology is advancing toward Ultra High Definition (UHD) and High Dynamic Range (HDR), which intensifies the need for higher compression efficiency for these high‑specification videos. Beyond advances in traditional codecs, neural video codecs (NVCs) have attracted significant research attention and have evolved rapidly over the past few years. The coding artifacts of NVCs often exhibit content‑varying and generative characteristics, which differ from those of conventional codecs and are challenging for traditional video quality assessment (VQA) methods to capture. Therefore, VQA metrics are required to generalize across different codecs, content types, and dynamic ranges to better support video codec research and evaluation. In this paper, we propose FDIM, a feature‑distance‑based generic video quality metric for both traditional and neural video codecs across SDR and HDR formats. FDIM employs a hybrid architecture that integrates deep and hand‑crafted features. The deep feature component learns multi‑scale representations to capture distortions ranging from structural and textural fidelity degradation to high‑level semantic deviations, while the hand‑crafted feature component provides stable complementary cues to improve overall generalization. We trained FDIM on a large‑scale subjective quality assessment dataset (DCVQA) consisting of over 16k video sequences encoded by traditional block‑based hybrid video codecs and end‑to‑end perceptually optimized neural video codecs. Extensive experiments on ten SDR/HDR VQA datasets containing diverse, previously unseen codecs demonstrate that FDIM achieves strong generalization and high correlation with subjective assessment. The source code for FDIM and the DCVQA validation set will be released at https://github.com/MCL‑ZJU/FDIM.
Authors:Yifeng Bai, Zhirong Chen, Erkang Cheng, Haibin Ling
Abstract:
Topology reasoning is crucial for autonomous driving. Current methods primarily focus on instance‑level learning for centerline detection, followed by a sequential module for topology reasoning that relies on simplified MLP layers. Moreover, they often neglect the importance of point‑to‑instance (P2I) relationships in topology reasoning. To address these limitations, we present TopoHR (Topological Hierarchical Representation), a novel end‑to‑end framework that establishes cyclic interaction between centerline detection and topology reasoning, allowing them to iteratively enhance each other. Specifically, we introduce a hierarchical centerline representation including point queries, instance queries, and semantic representations. These multi‑level features are seamlessly integrated and fused within a hierarchical centerline decoder. Furthermore, we design a hierarchical topology reasoning module that captures both fine‑grained P2I relationships and global instance‑to‑instance (I2I) connections within a unified architecture. With these novel components, TopoHR ensures accurate and robust topology reasoning. On the OpenLane‑V2 benchmark, TopoHR refreshes state‑of‑the‑art performance with significant improvements. Notably, compared with previous best results, TopoHR achieves +3.8 in \mathrmDET_\textl, +5.4 in \mathrmTOP_\textll on \textsubset_A and +11.0 in \mathrmDET_\textl, +7.9 in \mathrmTOP_\textll on \textsubset_B, validating the effectiveness of the proposed components. The code will be shared publicly at https://github.com/Yifeng‑Bai/TopoHR.git.
Authors:YuHao Yin, Zongji Wang, Yuanben Zhang, Biqing Li, Jiesong Bai, Junyi Liu
Abstract:
Full 360^\circ novel view synthesis under low‑light conditions remains challenging. Insufficient illumination, noise amplification, and view‑dependent photometric inconsistencies prevent existing methods from jointly preserving geometric consistency and photorealism. Unsupervised approaches often exhibit color drift under large viewpoint variations, while supervised low‑light enhancement models, though effective for 2D tasks, struggle to generalize to new scenes and typically require retraining. To address this issue, we propose MERID‑GS, a Multi‑Scale Explicit Retinex Illumination‑Decoupled Gaussian framework for low‑light 360^\circ synthesis. Based on Retinex theory, the method explicitly separates illumination and reflectance, and suppresses noise propagation while enhancing dark‑region structures via a learnable gain and Illumination‑State‑Guided Frequency Gating. Combined with lightweight Reflection Head and 3D Gaussian Splatting, MERID‑GS adapts to new scenes with only a few shots and enables stable low‑light novel view synthesis from sparse‑view observations. In addition, we construct a low‑light multi‑view dataset covering full 360^\circ scenes for joint evaluation. Thorough experiments across multiple datasets in this area demonstrate that MERID‑GS achieves SOTA performance, exhibiting superior cross‑scene generalization and view consistency. The source code and pre‑trained models are available at https://github.com/YhuoyuH/MERID‑GS..
Authors:Fengxian Ji, Jingpu Yang, Zirui Song, Lang Gao, Junhong Liang, Zhenhao Chen, Jinghui Zhang, Xiuying Chen
Abstract:
Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their performance on paid, real‑world design projects remains uncertain. We introduce ServImage, a benchmark that explicitly correlates model outputs with economic value in commercial design projects. ServImage consists of (i) ServImageBench: a dataset of 1.07k paid commercial design tasks and 2.05k designer deliverables totaling over \295k, covering portrait, product, and digital content, along with 33k candidate images and 33k human annotations. (ii) ServImageScore: an integrated scoring system that combines three quality dimensions: baseline requirements fulfilment, visual execution quality, and commercial necessity satisfaction. These three dimensions are designed to characterize the factors that drive human payment decisions and indicate whether an image is commercially acceptable. (iii) ServImageModel: under this scoring system, we propose a payment prediction model trained on the human‑annotated candidate images, achieving 82.00% accuracy in predicting human payment decisions and producing calibrated payment probabilities. ServImage provides a comprehensive foundation for assessing the commercial viability of image generation models and offers a scalable resource for future research on economically grounded vision systems \hrefhttps://github.com/FengxianJi/ServImageGithub.
Authors:Zichun Guo, Yuling Shi, Wenhao Zeng, Chao Hu, Haotian Lin, Terry Yue Zhuo, Jiawei Chen, Xiaodong Gu, Wenping Ma
Abstract:
Multimodal Large Language Models (MLLMs) have achieved remarkable performance in Visually Rich Document Understanding (VRDU) tasks, but their capabilities are mainly evaluated on pristine, well‑structured document images. We consider content restoration from shredded fragments, a challenging VRDU setting that requires integrating visual pattern recognition with semantic reasoning under significant content discontinuities. To facilitate systematic evaluation of complex VRDU tasks, we introduce ShredBench, a benchmark supported by an automated generation pipeline that renders fragmented documents directly from Markdown. The proposed pipeline ensures evaluation validity by allowing the flexible integration of latest or unseen textual sources to prevent training data contamination. ShredBench assesses four scenarios (English, Chinese, Code, Table) with three fragmentation granularities (8, 12, 16 pieces). Empirical evaluations on state‑of‑the‑art MLLMs reveal a significant performance gap: The method is effective on intact documents; however, once the document is shredded, restoration becomes a significant challenge, with NED dropping sharply as fragmentation increases. Our findings highlight that current MLLMs lack the fine‑grained cross‑modal reasoning required to bridge visual discontinuities, identifying a critical gap in robust VRDU research.
Authors:Chih-Chung Hsu, Xin-Di Ma, Wo-Ting Liao, Chia-Ming Lee
Abstract:
Existing attention accelerators often trade exact softmax semantics, depend on fused Tensor Core kernels, or incur sequential depth that limits FP32 throughput on long sequences. We present ELSA, an algorithmic reformulation of online softmax attention that (i)~preserves exact softmax semantics in real arithmetic with a \emphprovable \mathcalO(u\log n) FP32 relative error bound; (ii)~casts the online softmax update as a prefix scan over an associative monoid (m,S,W), yielding O(n) extra memory and O(\log n) parallel depth; and (iii)~is Tensor‑Core independent, implemented in Triton and CUDA C++, and deployable as a \emphdrop‑in replacement requiring no retraining or weight modification. Unlike FlashAttention‑2/3, which rely on HMMA/GMMA Tensor Core instructions and provide no compatible FP32 path, ELSA operates identically on A100s and resource‑constrained edge devices such as Jetson TX2 ‑‑ making it the only hardware‑agnostic exact‑attention kernel that reduces parallel depth to O(\log n) at full precision. On A100 FP32 benchmarks (1K‑‑16K tokens), ELSA delivers 1.3‑‑3.5× speedup over memory‑efficient SDPA and 1.97‑‑2.27× on BERT; on Jetson TX2, ELSA achieves 1.5‑‑1.6× over Math (64‑‑900 tokens), with 17.8‑‑20.2% throughput gains under LLaMA‑13B offloading at \ge32K. In FP16, ELSA approaches hardware‑fused baselines at long sequences while retaining full FP32 capability, offering a unified kernel for high‑precision inference across platforms. Our code and implementation are available at https://github.com/ming053l/ELSA.
Authors:Fanqing Meng, Lingxiao Du, Zijian Wu, Guanzheng Chen, Xiangyan Liu, Jiaqi Liao, Chonghe Jiang, Zhenglin Wan, Jiawei Gu, Pengfei Zhou, Rui Huang, Ziqi Zhao, Shengyuan Ding, Ailing Yu, Bo Peng, Bowei Xia, Hao Sun, Haotian Liang, Ji Xie, Jiajun Chen, Jiajun Song, Liu Yang, Ming Xu, Qionglin Qiu, Runhao Fu, Shengfang Zhai, Shijian Wang, Tengfei Ma, Tianyi Wu, Weiyang Jin, Yan Wang, Yang Dai, Yao Lai, Youwei Shu, Yue Liu, Yunzhuo Hao, Yuwei Niu, Jinkai Huang, Jiayuan Zhuo, Zhennan Shen, Linyu Wu, Hannah Yao, Charles Chen, Cihang Xie, Yuyin Zhou, Jiaheng Zhang, Zeyu Zheng, Mengkang Hu, Michael Qizhe Shieh
Abstract:
Language‑model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change independently of the agent: new emails arrive, calendar entries shift, knowledge‑base records are updated, and evidence appears across images, scanned PDFs, audio, video, and spreadsheets. Existing benchmarks do not adequately evaluate this setting because they typically run within a single static episode and remain largely text‑centric. We introduce \bench, a benchmark for coworker agents built around multi‑turn multi‑day tasks, a stateful sandboxed service environment whose state evolves between turns, and rule‑based verification. The current release contains 100 tasks across 13 professional scenarios, executed against five stateful sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) and scored by 1537 deterministic Python checkers over post‑execution service state; no LLM‑as‑judge is invoked during scoring. We benchmark seven frontier agent systems. The strongest model reaches 75.8 weighted score, but the best strict Task Success is only 20.0%, indicating that partial progress is common while complete end‑to‑end workflow completion remains rare. Turn‑level analysis shows that performance drops after the first exogenous environment update, highlighting adaptation to changing state as a key open challenge. We release the benchmark, evaluation harness, and construction pipeline to support reproducible coworker‑agent evaluation.
Authors:Xuefen Liu, Xinquan Yang, Mianjie Zheng, Kun Tang, Xuguang Li, Xiaoqi Guo, Linlin Shen, He Meng
Abstract:
As dental caries appear as subtle, low‑contrast lesions in intraoral imaging, existing deep learning models face significant challenges in the early detection of caries. While recent Transformer‑based detectors have shown promising results in natural images, they often fail to capture the domain‑specific anatomical priors crucial for dental caries detection. In this paper, we propose Caries‑DETR, a specialized Transformer framework for caries detection in intraoral images. A Tooth Structure‑aware Query Initialization (TSQI) is designed, leveraging large‑scale intraoral photograph pre‑training and a structure perception branch (SPB) to integrate high‑frequency structural priors, guiding the model to focus on anatomically significant lesion areas. Furthermore, we design a Lesion‑aware Dynamic Loss Refinement (LDLR) to implement quality‑driven hard mining through adaptive loss reweighting based on lesion size, anatomical relevance, and prediction quality, optimizing detection for subtle lesions. Extensive experiments on two public datasets (i.e., AlphaDent and DentalAI) demonstrate that Caries‑DETR achieves a state‑of‑the‑art performance compared to existing methods and exhibits good generalization and robustness. Code and data at https://github.com/XuefenLiu‑SZU/Caries‑DETRhttps://github.com/XuefenLiu‑SZU/Caries‑DETR.
Authors:Xinheng Li, Minghao Chen, Mengqing Wu, Yan Liu, Guanying Huo
Abstract:
Single image dehazing is often constrained by a trade‑off between restoration quality and computational efficiency. While efficient, CNN networks struggle to learn robust priors for dense and non‑homogeneous haze. Conversely, diffusion models provide strong generative priors but suffer from severe inference latency and sampling instability. To address these limitations, we propose ZID‑Net, a novel framework that explicitly decouples diffusion supervision from feed‑forward inference. For efficient inference, we design a frequency‑spatial decoupled feed‑forward backbone. Within this backbone, a Channel‑Spatial Laplacian Mask (CSLM) filters haze‑amplified noise to extract purified structural details, while Lightweight Global Context Blocks (LGCBs) establish long‑range spatial dependencies to capture the global variations of haze. A Dynamic Feature Arbitration Block (DFAB) then adaptively fuses these semantic and structural features for robust reconstruction. To provide this backbone with physical priors without the inference cost, we introduce a Zero‑Inference Prior Propagation Head (ZI‑PPH) during training. ZI‑PPH leverages a conditional diffusion process to predict residual noise, providing degradation‑aware structural supervision to the backbone. By discarding the diffusion branch at test time, ZID‑Net integrates diffusion priors into a pure feed‑forward architecture for accurate and efficient restoration. ZID‑Net achieves 40.75 dB PSNR on the synthetic RESIDE dataset and outperforms existing methods with a 1.13 dB gain on real‑world datasets. Additionally, it yields a 3.06 dB PSNR gain on the StateHaze1k remote sensing dataset with an inference time of just 19.35 ms. The project code is available at: https://github.com/XoomitLXH/ZID‑Net.
Authors:Wentao Zhang, Qi Zhang, Mingkun Xu, Mu You, Henghua Shen, Zhongzhi He, Keyan Jin, Derek F. Wong, Tao Fang
Abstract:
Crop disease diagnosis from field photographs faces two recurring problems: models that score well on benchmarks frequently hallucinate species names, and when predictions are correct, the reasoning behind them is typically inaccessible to the practitioner. This paper describes Agri‑CPJ (Caption‑Prompt‑Judge), a training‑free few‑shot framework in which a large vision‑language model first generates a structured morphological caption, iteratively refined through multi‑dimensional quality gating, before any diagnostic question is answered. Two candidate responses are then generated from complementary viewpoints, and an LLM judge selects the stronger one based on domain‑specific criteria. Caption refinement is the component with the largest individual impact: ablations confirm that skipping it consistently degrades downstream accuracy across both models tested. On CDDMBench, pairing GPT‑5‑Nano with GPT‑5‑mini‑generated captions yields +22.7 pp in disease classification and +19.5 points in QA score over no‑caption baselines. Evaluated without modification on AgMMU‑MCQs, GPT‑5‑Nano reached 77.84% and Qwen‑VL‑Chat reached 64.54%, placing them at or above most open‑source models of comparable scale despite the format shift from open‑ended to multiple‑choice. The structured caption and judge rationale together constitute a readable audit trail: a practitioner who disagrees with a diagnosis can identify the specific caption observation that was incorrect. Code and data are publicly available https://github.com/CPJ‑Agricultural/CPJ‑Agricultural‑Diagnosis
Authors:Lei Kang, Giuseppe De Gregorio, Raphaela Heil, Alicia Fornés, Beáta Megyesi
Abstract:
Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing computational workflows rely on a transcription‑first paradigm, in which handwritten symbols are transcribed prior to decipherment. This intermediate step is labor‑intensive, error‑prone, and not always aligned with the goal of direct plaintext recovery. We propose an end‑to‑end, transcription‑free approach that directly maps handwritten cipher images to plaintext. Using the Copiale cipher as a case study, we introduce the first text‑line‑level dataset pairing cipher images with German plaintext. We show that pretraining on generic handwriting data followed by cipher‑specific fine‑tuning substantially improves decipherment accuracy. Our results demonstrate that transcription‑free image‑to‑plaintext decipherment is both feasible and effective for historical substitution ciphers, offering a simplified and scalable alternative to traditional pipelines. https://github.com/leitro/Decipher‑from‑Pixels‑Copiale
Authors:Francesco Dibitonto, Cigdem Beyan, Vittorio Murino
Abstract:
Recent advances in representation learning have shown that hyperbolic geometry can offer a more expressive alternative to the Euclidean embeddings used in CLIP models, capturing hierarchical structures and leading to better‑organized representations. However, current hyperbolic CLIP variants are trained entirely from scratch, which is computationally expensive and resource‑intensive. In this work, we propose HAC (Hyperbolic Adaptation of CLIP), a parameter‑efficient framework that enables pretrained CLIP models to transition into hyperbolic space via lightweight fine‑tuning. We apply HAC to Visual Question Answering (VQA), where models must interpret visual elements and align them with textual queries. Notably, HAC's training is performed on a dataset with no overlap with any VQA benchmark, resulting in a strict zero‑shot evaluation paradigm that underscores HAC's task‑agnostic adaptability. We evaluate HAC across a diverse suite of VQA benchmarks spanning General, Reasoning, and OCR categories. Both HAC‑S (small) and HAC‑B (medium) consistently surpass Euclidean baselines and prior hyperbolic approaches, with HAC‑B delivering up to a +1.9 point average improvement over CLIP‑B on reasoning‑intensive tasks. Our code is available at https://github.com/fdibiton/HAC
Authors:Guoxi Huang, Ruirui Lin, Yini Li, David R. Bull, Nantheera Anantrasirichai
Abstract:
Videos captured in low‑light and underwater conditions often suffer from distortions such as noise, low contrast, color imbalance, and blur. These issues not only limit visibility but also degrade automatic tasks like detection. Post‑processing is typically required but can be time‑consuming. AI‑based tools for video enhancement also demand significantly more computational resources compared to image‑based methods. This paper introduces a novel framework, Visual Mamba, designed to reduce memory usage and computational time by leveraging the Visual State Space (VSS) model. The framework consists of two modules: (i) a feature alignment module, where spatio‑temporal displacement between input frames is registered in the feature space, and (ii) an enhancement module, where noise removal and brightness adjustment are performed using a UNet‑like architecture, with all convolutional layers replaced by VSS blocks. Experimental results show that the Visual Mamba technique outperforms Transformer and convolution‑based models in both low‑light and underwater video enhancement tasks. Code is available on line at https://github.com/russellllaputa/BVI‑Mamba.
Authors:Pritesh Jha
Abstract:
Intelligent document processing pipelines extract structured entities (tables, images, and text) from documents for use in downstream systems such as knowledge bases, retrieval‑augmented generation, and analytics. A persistent limitation of existing pipelines is that extraction output is produced without any intrinsic mechanism to verify whether it faithfully represents the source. Model‑internal confidence scores measure inference certainty, not correspondence to the document, and extraction errors pass silently into downstream consumers.
We present Reconstruction as Validation (RaV‑IDP), a document processing pipeline that introduces reconstruction as a first‑class architectural component. After each entity is extracted, a dedicated reconstructor renders the extracted representation back into a form comparable to the original document region, and a comparator scores fidelity between the reconstruction and the unmodified source crop. This fidelity score is a grounded, label‑free quality signal. When fidelity falls below a per‑entity‑type threshold, a structured GPT‑4.1 vision fallback is triggered and the validation loop repeats. We enforce a bootstrap constraint: the comparator always anchors against the original document region, never against the extraction, preventing the validation from becoming circular.
We further propose a per‑stage evaluation framework pairing each pipeline component with an appropriate benchmark. The code pipeline is publicly available at https://github.com/pritesh‑2711/RaV‑IDP for experimentation and use.
Authors:Francesco Olivato, Cigdem Beyan, Vittorio Murino
Abstract:
In this work, we study Source‑Free Unsupervised Domain Adaptation under corruption‑induced domain shifts, where performance degradation is caused by natural image corruptions that go beyond additive noise, including blur, weather effects, and digital artifacts. We propose a diffusion‑based, input‑level adaptation framework that operates entirely at test time and keeps all source‑trained models frozen, explicitly targeting robustness to corrupted target inputs. Our method leverages a source‑trained diffusion model as a generative prior and introduces a discriminator‑guided adaptive diffusion strategy that dynamically controls the amount of perturbation applied to each test sample. Rather than relying on a fixed diffusion depth, the discriminator determines, on a per‑image basis, when sufficient forward diffusion has been applied to suppress corruption‑specific artifacts, with each corruption type effectively defining a distinct target domain. This adaptive stopping mechanism applies only the necessary amount of noise to remove domainspecific corruption while preserving class‑discriminative structure. The reverse diffusion process then reconstructs a source‑aligned image, optionally stabilized through structural guidance, which is classified using a frozen source‑trained classifier. We evaluate the proposed approach across a broad spectrum of corruption‑induced target domains, covering 15 diverse corruption types, and demonstrate more balanced robustness with competitive or improved performance across non‑noise corruptions. Additional analyses reveal how the adaptive diffusion schedule responds to different corruption characteristics, highlighting the practicality, generality, and robustness of the proposed framework. The code is publicly available at https://github.com/fmolivato/dgadiffusion/.
Authors:Peng Chen, Wenxuan He, Feng Qian, Guangyao Shi, Jingwen Yan
Abstract:
In the hyperspectral image (HSI) classification task, each pixel is categorized into a specific land‑cover category or material. Convolutional neural networks (CNNs) and transformers have been widely used to extract local and non‑local features in HSI classification. Recent works have utilized a multi‑scale vision transformer (ViT) to enhance spectral feature capture and yield promising results. However, most existing methods still face challenges in the effective joint use of spatial‑spectral information and in preserving information across layers during the propagation process. To address these issues, we propose a synergistic CNN‑Transformer network with pooling attention fusion for HSI classification, which collaboratively utilizes CNNs and ViT to process spatial and spectral features separately. Specifically, we propose a Twin‑Branch Feature Extraction (TBFE) module, which employs 3D and 2D convolution in parallel to comprehensively extract spectral and spatial features from HSI. A hybrid pooling attention (HPA) module is designed to aggregate spatial attention. Moreover, a cascade transformer encoder is employed for global spectral feature extraction, and a simple yet efficient cross‑layer feature fusion (CFF) module is designed to reduce the loss of crucial information in the previous network layers. Extensive experiments are conducted on several representative datasets to demonstrate the superior performance of our proposed method compared to the state‑of‑the‑art works. Code is available at https://github.com/chenpeng052/SCT‑Net.git.
Authors:Simone Mosco, Daniel Fusaro, Alberto Pretto
Abstract:
Understanding the surrounding environment is fundamental in autonomous driving and robotic perception. Distinguishing between known classes and previously unseen objects is crucial in real‑world environments, as done in Anomaly Segmentation. However, research in the 3D field remains limited, with most existing approaches applying post‑processing techniques from 2D vision. To cover this lack, we propose a new efficient approach that directly operates in the feature space, modeling the feature distribution of inlier classes to constrain anomalous samples. Moreover, the only publicly available 3D LiDAR anomaly segmentation dataset contains simple scenarios, with few anomaly instances, and exhibits a severe domain gap due to its sensor resolution. To bridge this gap, we introduce a set of mixed real‑synthetic datasets for 3D LiDAR anomaly segmentation, built upon established semantic segmentation benchmarks, with multiple out‑of‑distribution objects and diverse, complex environments. Extensive experiments demonstrate that our approach achieves state‑of‑the‑art and competitive results on the existing real‑world dataset and the newly introduced mixed datasets, respectively, validating the effectiveness of our method and the utility of the proposed datasets. Code and datasets are available at https://simom0.github.io/lido‑page/.
Authors:Weihao Li, Hongjin Zhao, Gao Zhu, Ge-Peng Ji, Nicholas Wilson, Marta Yebra, Nick Barnes
Abstract:
Wildfires are an escalating global concern due to the devastating impacts on the environment, economy, and human health, with notable incidents such as the 2019‑2020 Australian bushfires and the 2025 California wildfires underscoring the severity of these events. AI‑enabled camera‑based smoke detection has emerged as a promising approach for the rapid detection of wildfires. However, existing wildfire smoke segmentation datasets that are used for training detection and segmentation models are limited in scale, geographically constrained, and often rely on synthetic imagery, which hinders effective training and generalization. To overcome these limitations, we present AusSmoke, a new smoke segmentation dataset collected from Australia to address the data scarcity in this region. Furthermore, we introduce a MultiNational geographically diverse and substantially larger fully‑labelled benchmark, called MultiNatSmoke, that consolidates publicly available international datasets with the newly collected Australian imagery, expanding the scale by an order of magnitude over previous collections. Finally, we benchmark smoke segmentation models, demonstrating improved performance and enhanced generalization across diverse geographical contexts. The project is available at \hrefhttps://github.com/henryzhao0615/MultiNatSmokeGithub.
Authors:Jainum Sanghavi
Abstract:
Vision Transformers trained only on image classification routinely transfer to tasks that demand spatial understanding, yet they receive no spatial supervision during pretraining. We ask where and how robustly such structure is encoded. Probing a frozen ViT‑B/16 layerwise for two complementary properties, local patch boundaries (BSDS500) and per‑patch depth (NYU Depth V2), reveals a clear hierarchy: boundary structure becomes linearly decodable at layers 5‑6 (AP = 0.833), while depth, which requires integrating global cues, peaks two to three layers later at layer 8 (MAE = 0.0875). Both signals collapse at the final classification layer, and random‑weight controls confirm the encodings are learned rather than architectural. Causal interventions add specificity: ablating the single direction a linear depth probe reads degrades depth decoding by up to 165%, while ablating any other direction changes it by less than 1%. Targeted activation patching along that direction shows the depth signal is partially re‑derived at each layer rather than passively carried in the residual stream, with mid‑layer interventions persisting most strongly downstream. The result is that a classification‑trained ViT develops an actively maintained spatial hierarchy that mirrors the early‑to‑late progression observed in the primate visual cortex.
Authors:Soulayma Gazzeh, Giuseppe Mazzola, Liliana Lo Presti, Marco La Cascia
Abstract:
Reliable depth estimation from spherical images is crucial for 360° vision in robotic navigation and immersive scene understanding. However, the onboard spherical camera can experience unintentional pose variations in real‑world robotic platforms that, along with the geometric distortions inherent in equirectangular projections, significantly impact the effectiveness of depth estimation. To study this issue, a novel public benchmark, called Sphere‑Depth, is introduced to systematically evaluate the robustness of monocular depth estimation models from equirectangular images in a reproducible way. Camera pose perturbations are simulated and used to assess the performance of a popular perspective‑based model, Depth Anything, and of spherical‑aware models such as Depth Anywhere, ACDNet, Bifuse++, and SliceNet. Furthermore, to ensure meaningful evaluation across models, a depth calibration‑based error protocol is proposed to convert predicted relative depth values into metric depth values using supervised learned scaling factors for each model. Experiments show that even models explicitly designed to process spherical images exhibit substantial performance degradation when variations in the camera pose are observed with respect to the canonical pose. The full benchmark, evaluation protocol, and dataset splits are made publicly available at: https://github.com/sgazzeh/Sphere_depth
Authors:Emre Ardıç, Yakup Genç
Abstract:
Federated learning (FL) is a distributed machine learning method where multiple devices collaboratively train a model under the management of a central server without sharing underlying data. One of the key challenges of FL is the communication bottleneck caused by variations in connection speed and bandwidth across devices. Therefore, it is essential to reduce the size of transmitted data during training. Additionally, there is a potential risk of exposing sensitive information through the model or gradient analysis during training. To address both privacy and communication efficiency, we combine differential privacy (DP) and adaptive quantization methods. We use Laplacian‑based DP to preserve privacy, which is relatively underexplored in FL and offers tighter privacy guarantees than Gaussian‑based DP. We propose a simple and efficient global bit‑length scheduler using round‑based cosine annealing, along with a client‑based scheduler that dynamically adapts based on client contribution estimated through dataset entropy analysis. We evaluate our approach through extensive experiments on CIFAR10, MNIST, and medical imaging datasets, using non‑IID data distributions across varying client counts, bit‑length schedulers, and privacy budgets. The results show that our adaptive quantization methods reduce total communicated data by up to 52.64% for MNIST, 45.06% for CIFAR10, and 31% to 37% for medical imaging datasets compared to 32‑bit float training while maintaining competitive model accuracy and ensuring robust privacy through differential privacy.
Authors:Shengzhi Li, Jiarun Chen, Karun Sharma, Jiaqi Su, Shichao Pei
Abstract:
Large vision‑language models (VLMs) can recognize what happens in video but fail to count how many times. We introduce PushupBench, 446 long‑form clips (avg. 36.7s) for evaluating repetition counting. The best frontier model achieves 42.1% exact accuracy; open‑source 4B models score ~6%, matching supervised baselines. We show that accuracy alone misleads ‑‑ weaker models exploit the modal count rather than reason temporally. Fine‑tuning on counting with 1k samples transfers to general video understanding: MVBench (+2.15), PerceptionTest (+1.88), TVBench (+4.54), suggesting counting is a proxy for broader temporal reasoning.PushupBench incorporated in \textttlmms‑eval (https://github.com/EvolvingLMMs‑Lab/lmms‑eval/pull/1262) and hosted on (pushupbench.com/)
Authors:Sheng-Wei Chan, Hsin-Jui Pan, Chun-Po Shen, Chia-Min Lin, Yung-Che Wang, Jen-Shiun Chiang
Abstract:
High‑performance semantic segmentation has achieved significant progress in recent years, often driven by increasingly large backbones and higher computational budgets. While effective, such approaches introduce substantial computational overhead and limit accessibility under constrained hardware settings. In this paper, we propose DGM‑Net (Directional Geometric Mamba Network), an efficient architecture that improves modeling capability through structural design rather than increasing model capacity. We introduce Directional Geometric Mamba (G‑Mamba), a linear‑complexity O(N) operator as an alternative to conventional context modeling modules such as ASPP and PPM. To further enhance structural awareness in state space model (SSM)‑based modeling, we design the DGM‑Module, which extracts centripetal flow fields and topological skeletons to guide the scanning process and improve boundary preservation. Without relying on large‑scale pretraining or heavy backbone scaling, DGM‑Net achieves 80.8% mIoU within 28k iterations, 82.3% mIoU on Cityscapes test set, and 45.24% mIoU on ADE20K. In addition, the model maintains stable performance under constrained hardware settings (e.g., batch size of 2 on 8GB VRAM), highlighting its efficiency and practicality. These results demonstrate that incorporating geometric guidance into SSM‑based architectures provides an effective and resource‑efficient direction for semantic segmentation.
Authors:He Hu, Tengjin Weng, Zebang Cheng, Yu Wang, Jiachen Luo, Björn Schuller, Zheng Lian, Laizhong Cui
Abstract:
Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in applications such as social robots and human‑computer interaction, where understanding human emotions is essential. However, existing benchmarks mainly formulate emotion understanding as a static recognition problem, leaving it largely unclear whether current MLLMs can understand emotion as a dynamic process that evolves, shifts between states, and unfolds across diverse social contexts. To bridge this gap, we present EmoTrans, a benchmark for evaluating emotion dynamics understanding in multimodal videos. EmoTrans contains 1,000 carefully collected and manually annotated video clips, covering 12 real‑world scenarios, and further provides over 3,000 task‑specific question‑answer (QA) pairs for fine‑grained evaluation. The benchmark introduces four tasks, namely Emotion Change Detection (ECD), Emotion State Identification (ESI), Emotion Transition Reasoning (ETR), and Next Emotion Prediction (NEP), forming a progressive evaluation framework from coarse‑grained detection to deeper reasoning and prediction. We conduct a comprehensive evaluation of 18 state‑of‑the‑art MLLMs on EmoTrans and obtain two main findings. First, although current MLLMs show relatively stronger performance on coarse‑grained emotion change detection, they still struggle with fine‑grained emotion dynamics modeling. Second, socially complex settings, especially multi‑person scenarios, remain substantially challenging, while reasoning‑oriented variants do not consistently yield clear improvements. To facilitate future research, we publicly release the benchmark, evaluation protocol, and code at https://github.com/Emo‑gml/EmoTrans.
Authors:Chandravardhan Singh Raghaw, Anushka Parwal, Shahid Shafi Dar, Prajakta Darade, Nagendra Kumar
Abstract:
Knee osteoarthritis (KOA) is a degenerative joint disease that can lead to chronic pain, reduced mobility, and long‑term disability. Automated severity grading from knee radiographs can support early assessment, but current methods heavily depend on large labeled datasets and remain sensitive to class imbalance, noisy samples, and variability in clinical annotations. To alleviate these limitations, we propose a Hierarchical fusion of Semi‑Supervised framework with Self‑Supervision (H‑SemiS) for KOA severity grading in knee X‑ray samples using limited annotated data. Rather than treating severity grading as a flat multi‑class problem, H‑SemiS decomposes the task into a sequence of binary sub‑tasks within a semi‑supervised teacher‑student architecture, directly mitigating the impact of class imbalance. To further enhance feature learning from unlabeled data, the framework integrates an adversarial self‑supervised reconstruction module that encourages the network to capture robust anatomical structures. In parallel, a teacher‑student design with quantum‑inspired feature mixing improves discrimination boundaries between adjacent grades when pseudo‑labels are noisy. We comprehensively evaluate H‑SemiS on two challenging multi‑class datasets and assess its generalizability on two binary‑class datasets. Our experimental results demonstrate the superiority of the proposed H‑SemiS framework across multiple evaluation metrics, consistently outperforming several competing baselines and state‑of‑the‑art methods. The code is publicly available at https://github.com/chandravardhan‑singh‑raghaw/H‑SemiS.
Authors:Zhaoxiang Liu, Zhicheng Ma, Kaikai Zhao, Kai Wang, Shiguo Lian
Abstract:
The Convolutional Neural Networks (CNNs) have been the dominant and effective approach for general computer vision tasks. Recently, Kolmogorov‑Arnold neural networks (KANs), based on the Kolmogorov‑Arnold representation theorem, have shown potential to replace Multi‑Layer Perceptrons (MLPs) in deep learning. KANs, which use learnable nonlinear activations on edges and simple summation on nodes, offer fewer parameters and greater explainability compared to MLPs. However, there has been limited exploration of integrating the Kolmogorov‑Arnold representation theorem with convolutional methods for computer vision tasks. Existing attempts have merely replaced learnable activation functions with weights, undermining KANs' theoretical foundation and limiting their potential effectiveness. Additionally, the B‑spline curves used in KANs suffer from computational inefficiency and a tendency to overfit. In this paper, we propose a novel Kolmogorov‑Arnold Convolutional Layer that deeply integrates the Kolmogorov‑Arnold representation theorem with convolution. This layer provides stronger method interpretability because it is based on established mathematical theorems and its design has theoretical alignment. Building on the Kolmogorov‑Arnold Convolutional Layer, we design an efficient network architecture called KAConvNet, which outperforms existing methods combining KAN and convolution, and achieves competitive performance compared to mainstream ViTs and CNNs. We believe that our work offers valuable insight into the field of artificial intelligence and will inspire the development of more innovative CNNs in the 2020s. The code is publicly available at https://github.com/UnicomAI/KAConvNet.
Authors:Kaiwen Huang, Yi Zhou, Yizhe Zhang, Jingxiong Li, Tao Zhou
Abstract:
Semi‑supervised learning addresses label scarcity and high annotation costs in medical image segmentation by exploiting the latent information in unlabeled data to enhance model performance. Traditional discriminative segmentation relies on segmentation masks, neglecting feature‑level distribution constraints. This limits robust semantic representation learning and adaptive modeling of unlabeled data in scenarios with few labels. To address these limitations, we propose SemiGDA, a novel Generative Dual‑distribution Alignment framework for semi‑supervised medical image segmentation. Our SemiGDA overcomes the reliance of discriminative methods on large labeled datasets by aligning feature and semantic distributions to boost semantic learning and scene adaptability. Specifically, we propose a Dual‑distribution Alignment Module (DAM), which employs two structurally distinct encoders to model image and mask feature distributions. It enforces their alignment in the latent space via distributional constraints, establishing structured feature consistency. Moreover, we design a Consistency‑Driven Skip Adapter (CDSA) strategy, which introduces dual skip adapters (Image and Mask) to fuse multi‑scale features via skip connections. Using a consistency loss, CDSA enhances cross‑branch semantic alignment and reinforces fine‑grained semantic consistency. Experimental results on diverse medical datasets show that our method outperforms other state‑of‑the‑art semi‑supervised segmentation methods. Code is released at: https://github.com/taozh2017/SemiGDA.
Authors:Heng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang, Shuai Li, Jian-Fang Hu
Abstract:
Text‑to‑motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine‑grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which limits both semantic alignment and temporal coherence. Inspired by the fact that complex motions are conceptualized hierarchically rather than at a single temporal scale in the human cognitive system, we propose MotionHiFlow, a hierarchical flow matching framework to generate motion progressively by constructing flow path from low to high temporal scales. The flows at lower scales capture high‑level semantics and coarse motion structures, while flows at higher scales refine temporal details. To link the flows across scales, we introduce a novel cross‑scale transition process, ensuring continuity and preserving noise consistency. Furthermore, by integrating a Text‑Motion Diffusion Transformer and a topology‑aware Motion VAE, MotionHiFlow explicitly models structural dependencies among joints via joint‑aware positional encoding and skeletal topology, enabling precise semantic alignment alongside fine‑grained motion details. Extensive experiments on HumanML3D and KIT‑ML benchmarks demonstrate state‑of‑the‑art performance, with ablation studies confirming the effectiveness of the hierarchical design and key components. Code is available at https://github.com/ai‑lh/MotionHiFlow.
Authors:Balaji Darur, Amanmeet Garg, Makarand Tapaswi
Abstract:
Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by requiring identification of salient actions and associated short descriptions for event roles across multiple events. Grounding with VidSitu requires spatio‑temporal localization of key entities across shots and varied appearances.
We posit that coherent video understanding requires consistent identification of entities that play different roles. We propose Multimodal Entity Coreference (MEC) to unite entity descriptions in text with grounding across the video. Towards this, we introduce CineMEC, a multi‑stage approach that unites event role mention groups with visual clusters of entities, without explicit grounding supervision during training. Our approach is designed to exploit the synergy between visual grounding and captioning, where improving one influences the other and vice versa. For evaluation, we extend the VidSitu dataset with grounding annotations. While previous work focuses primarily on descriptions, CineMEC improves consistency across both: captioning (+2.5% CIDEr, +7% LEA) and visual grounding (+18% HOTA).
Authors:Jeremy Ellis
Abstract:
This paper presents a complete, end‑to‑end on‑device vision machine learning pipeline, comprising data acquisition, two‑layer CNN training with Adam optimization, and real‑time inference, executing entirely on a microcontroller‑class device costing 15‑40 USD. Unlike cloud‑based workflows that require external infrastructure and conceal the computational pipeline from the practitioner, this system implements every step of the core ML lifecycle in approximately 1,750 lines of readable C++ that compiles in under one minute using the Arduino IDE, with no external ML dependencies. Running on the Seeed Studio ESP32‑S3 XIAO ML Kit (8 MB PSRAM), the firmware achieves three‑class 64x64 image classification in approximately 9 minutes per training run, with real‑time inference at 6.3 FPS. Key contributions include: correct batch‑level gradient accumulation; pre‑computed resize lookup tables for inference; dual‑format weight export for SD‑free baked‑in deployment; a three‑tier weight priority system (SD binary > baked‑in header > He‑initialization) resolved automatically at boot; a single‑constant network reconfiguration interface; and PSRAM‑aware memory management suited to microcontroller constraints. All source code and reference datasets are released under the MIT License at https://github.com/webmcu‑ai/on‑device‑vision‑ai
Authors:Vitalii Tutevych, Raphael Memmesheimer, Luca Eichler, Dmytro Pavlichenko, Fynn Schilke, Rodja Krudewig, Sven Behnke
Abstract:
Reliable object perception is necessary for general‑purpose service robots. Open‑vocabulary detectors struggle to generalize beyond a few classes and fully supervised training of object detectors requires time‑intensive annotations. We present a semi‑supervised label propagation approach for household object segmentation. A segment proposer generates class‑agnostic masks, and an ensemble of Hopfield networks assigns labels by learning representative embeddings in complementary foundation model embedding spaces (CLIP, ViT, Theia). Our approach scales to 50 object classes with limited annotation overhead and can automatically label 60% of the data in a RoboCup@Home setting, where preparation time is severely constrained. Dataset and code are publicly available at https://github.com/ais‑bonn/label_propagation.
Authors:Ashwin Kumar, Robbie Holland, Corey Barrett, Jangwon Kim, Maya Varma, Zhihong Chen, Yunhe Gao, Greg Zaharchuk, Tara Taghavi, Krishnaram Kenthapadi, Akshay Chaudhari
Abstract:
Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP‑pretrained vision encoder to an LLM using LLaVA‑style finetuning. This two‑stage, decoupled approach introduces a projection layer that can distort visual features. This is especially concerning in medical imaging where subtle cues are essential for accurate diagnoses. In contrast, early‑fusion generative approaches such as Chameleon eliminate the projection bottleneck by processing image and text tokens within a single unified sequence, enabling joint representation learning that leverages the inductive priors of language models. We present CheXmix, a unified early‑fusion generative model trained on a large corpus of chest X‑rays paired with radiology reports. We expand on Chameleon's autoregressive framework by introducing a two‑stage multimodal generative pretraining strategy that combines the representational strengths of masked autoencoders with MLLMs. The resulting models are highly flexible, supporting both discriminative and generative tasks at both coarse and fine‑grained scales. Our approach outperforms well‑established generative models across all masking ratios by 6.0% and surpasses CheXagent by 8.6% on AUROC at high image masking ratios on the CheXpert classification task. We further inpaint images over 51.0% better than text‑only generative models and outperform CheXagent by 45% on the GREEN metric for radiology report generation. These results demonstrate that CheXmix captures fine‑grained information across a broad spectrum of chest X‑ray tasks. Our code is at: https://github.com/StanfordMIMI/CheXmix.
Authors:Peter Kulits, Cordelia Schmid
Abstract:
We train a language model to generate LEGO‑brick build sequences. While prior work has been restricted to discrete, voxel‑like towers, we consider a much broader set of pieces, encompassing thousands of part types with diverse connection semantics. To enable this, we first collect a large‑scale dataset of over 100,000 human‑designed LDraw brick objects and scenes. The complexity of our setting makes it challenging to autoregressively assemble structures that satisfy physical constraints. When predicting block pose directly, build sequences quickly become invalid after a small number of steps. Although pieces are placed in 3D space, it is the spatial relationships of the parts which define the whole. With this in mind, we design a graph‑based program representation that parametrizes structure through connectivity, improving the physical grounding of generated sequences. To enable future applications, we make our dataset and models available for research purposes. https://kulits.github.io/BrickNet
Authors:Rahul Patel
Abstract:
Anemia affects over one billion people globally and remains severely under‑diagnosed in low‑resource regions where laboratory blood tests are inaccessible. This paper presents AnemiaVision, an end‑to‑end web‑based system for non‑invasive anemia screening from smartphone photographs of the palpebral conjunctiva and fingernail beds. The proposed pipeline fine‑tunes a pre‑trained EfficientNet‑B3 backbone with a redesigned three‑layer classifier head incorporating BatchNorm, GELU activations, and high‑rate Dropout (0.45/0.35). Training employs four orthogonal accuracy‑boosting techniques: TrivialAugmentWide for policy‑free image augmentation, RandomErasing for spatial regularisation, Mixup (alpha=0.2) for inter‑class smoothing, and cosine‑annealing scheduling with linear warmup. Early stopping is governed by peak validation accuracy rather than validation loss to prevent premature termination on high‑variance epochs. The deployed Flask application integrates persistent patient‑history management backed by PostgreSQL on Render, with an automated database‑migration entrypoint ensuring zero data loss across redeploys. Ablation experiments demonstrate that accuracy‑first early stopping contributes +1.6% and Mixup contributes +2.8% to final validation accuracy. Overall, the proposed system achieves a validation accuracy of 96.2% and AUC‑ROC of 0.98, compared with 44.9% validation accuracy and AUC‑ROC of 0.58 from the three‑epoch CPU‑only baseline. Sensitivity for the anemic class reaches 0.96, making the system suitable as a first‑line screening tool for community health workers in rural settings. The system is publicly accessible and source code is openly available.
Authors:Nikoo Moradi, Gijs Luijten, Behrus Hinrichs-Puladi, Jens Kleesiek, Victor Alves, Jan Egger, André Ferreira
Abstract:
Diffusion models produce high‑quality synthetic data but suffer from slow inference. We propose 3D Variable‑Step Denoising Diffusion Probabilistic Model (VS‑DDPM) a framework engineered to maintain generative quality while accelerating inference by several factors. We tested our approach on four tasks (missing MRI, tumor removal, MRI‑to‑sCT, and CBCT‑to‑sCT) within the BraTS2025 and SynthRAD2025 challenges. Designed for high efficiency under hardware and time constrains imposed by both challenges. VS‑DDPM achieved state‑of‑the‑art (SOTA) performance in missing MRI synthesis, yielding Dice scores of 0.80, 0.83, and 0.88 for the enhancing tumor, tumor core, and whole tumor regions, respectively, alongside a structural similarity index (SSIM) of 0.95. For MRI tumor removal, the model attained a root mean squared error (RMSE) of 0.053, a peak signal‑to‑noise ratio (PSNR) of 26.77, and an SSIM of 0.918. While the framework demonstrated competitive performance in MRI‑to‑sCT and CBCT‑to‑sCT tasks, it did not reach SOTA benchmarks, potentially due to sensitivities in data pre and post‑processing pipelines or specific loss function configurations. These results demonstrate that VS‑DDPM provides a robust and tunable solution for high‑fidelity 3D medical image synthesis. The code is available in https://github.com/andre‑fs‑ferreira/SynthRAD_by_Faking_it.
Authors:Jia-Mian Wu, Jun Liu, Siqi Li, Xiaoya Wang, Shibai Yin, Huanyu Luo, Lingling Zheng, Qiang Gao, Jigang Yang, Tai-Xiang Jiang
Abstract:
Computed tomography (CT)‑based attenuation and scatter correction improves quantitative PET but adds radiation exposure that is particularly undesirable in pediatric imaging. Existing CT‑free methods are commonly trained in homogeneous settings and often degrade under scanner or radiotracer shifts, which limits their clinical utility. We propose the Generalizable PET Correction Network (GPCN), a dual‑domain network for domain‑robust CT‑free PET attenuation and scatter correction. GPCN combines a multi‑band contextual refinement module, which models pediatric anatomical variability through wavelet‑based multiscale decomposition and long‑range spatial context modeling, with a frequency‑aware spectral decoupling module, which performs coordinate‑conditioned amplitude/phase refinement in the Fourier domain. By synergizing multi‑band spatial contextual modeling with asymmetric frequency‑spectrum decoupling, the network explicitly separates invariant topological structures from domain‑specific noise, thereby achieving precise quantitative recovery of both anatomical organs and focal lesions. This design aims to separate anatomy‑dominant structures from domain‑sensitive spectral residuals and to improve robustness across heterogeneous imaging conditions. We train and evaluate the method on 1085 pediatric whole‑body PET scans acquired with two scanners and five radiotracers. In both joint training and zero‑shot cross‑domain evaluation, GPCN outperforms representative baselines and maintains stable quantitative accuracy on unseen scanner‑tracer combinations. The method is further supported by ablation, region‑wise quantitative analysis, and downstream segmentation experiments. In our cohort, the CT component of the conventional protocol corresponded to an average effective dose of 10.8 mSv, indicating the potential clinical value of reliable CT‑free correction for pediatric PET.
Authors:Hefeng Zhou, Xuan Liu, Sicheng Chen, Wutong Zhang, Wu Yan, Jiong Lou, Chentao Wu, Guangtao Xue, Wei Zhao, Jie Li
Abstract:
Federated cross‑modal retrieval faces severe challenges from heterogeneous client data, particularly non‑IID semantic distributions and missing modalities. Under such heterogeneity, a single global model is often insufficient to capture both shared cross‑modal knowledge and client‑specific characteristics. We propose RCSR, a personalization‑friendly federated framework that integrates prototype anchoring, retrieval‑centric semantic routing, and optional client‑specific adapters. Built on a frozen CLIP backbone, RCSR leverages lightweight shared adapters for global knowledge transfer while supporting efficient local personalization. Prototype anchoring helps unimodal clients align with global cross‑modal semantics, and a server‑side semantic router adaptively assigns aggregation weights based on retrieval consistency to mitigate alignment drift during heterogeneous updates. Extensive experiments on MS‑COCO, Flickr30K, and other benchmarks show that RCSR consistently improves global retrieval accuracy and training stability, while further enhancing client‑level retrieval performance, especially for clients with incomplete modalities. Code is available at https://github.com/RezinChow/RCSR‑Retrieval‑Centric‑Semantic‑Routing.
Authors:Fujun Han, Junan Chen, Xintong Zhu, Jingqi Ye, Xuanjie Mao, Tao Chen, Peng Ye
Abstract:
Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physics olympiads. However, they remain blank and unexplored for Small Object Understanding (SOU) tasks. To fill this gap, we introduce SOUBench, the first and comprehensive benchmark for exploring the small objects understanding capability of existing MLLMs. Specifically, we first design an effective and automatic visual question‑answer generation strategy, constructing a new SOU‑VQA evaluation dataset, with 18,204 VQA pairs, six relevant sub‑tasks, and three dominant scenarios (i.e., Driving, Aerial, and Underwater). Then, we conduct a comprehensive evaluation on 15 state‑of‑the‑art MLLMs and reveal their weak capabilities in small object understanding. Furthermore, we develop SOU‑Train, a multimodal training dataset with 11,226 VQA pairs, to improve the SOU capabilities of MLLMs. Through supervising fine‑tuning of the latest MLLM, we demonstrate that SOU‑Train can effectively enhance the latest MLLM's ability to understand small objects. Comprehensive experimental results demonstrate that, the proposed SOUBench, along with the SOU‑VQA and SOU‑Train datasets, provides a crucial empirical foundation to the community for further developing models with enhanced small object understanding capabilities. Datasets and Code: https://github.com/Hanfj‑X/SOU.
Authors:Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri, Trung Bui, Anh Totti Nguyen
Abstract:
When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision‑language models (VLMs) such as Gemini‑3‑Pro and GPT‑5 only respond with text, which can be difficult for users to verify. We present SketchVLM, a training‑free, model‑agnostic framework that enables VLMs to produce non‑destructive, editable SVG overlays on the input image to visually explain their answers. Across seven benchmarks spanning visual reasoning (maze navigation, ball‑drop trajectory prediction, and object counting) and drawing (part labeling, connecting‑the‑dots, and drawing shapes around objects), SketchVLM improves visual reasoning task accuracy by up to +28.5 percentage points and annotation quality by up to 1.48x relative to image‑editing and fine‑tuned sketching baselines, while also producing annotations that are more faithful to the model's stated answer. We find that single‑turn generation already achieves strong accuracy and annotation quality, and multi‑turn generation opens up further opportunities for human‑AI collaboration. An interactive demo and code are at https://sketchvlm.github.io/.
Authors:Zhimu Zhou, Yanpeng Zhao, Qiuyu Liao, Bo Zhao, Xiaojian Ma
Abstract:
Visual planning represents a crucial facet of human intelligence, especially in tasks that require complex spatial reasoning and navigation. Yet, in machine learning, this inherently visual problem is often tackled through a verbal‑centric lens. While recent research demonstrates the promise of fully visual approaches, they suffer from significant computational inefficiency due to the step‑by‑step planning‑by‑generation paradigm. In this work, we present EAR, an editing‑as‑reasoning paradigm that reformulates visual planning as a single‑step image transformation. To isolate intrinsic reasoning from visual recognition, we employ abstract puzzles as probing tasks and introduce AMAZE, a procedurally generated dataset that features the classical Maze and Queen problems, covering distinct, complementary forms of visual planning. The abstract nature of AMAZE also facilitates automatic evaluation of autoregressive and diffusion‑based models in terms of both pixel‑wise fidelity and logical validity. We assess leading proprietary and open‑source editing models. The results show that they all struggle in the zero‑shot setting, finetuning on basic scales enables remarkable generalization to larger in‑domain scales and out‑of‑domain scales and geometries. However, our best model that runs on high‑end hardware fails to match the zero‑shot efficiency of human solvers, highlighting a persistent gap in neural visual reasoning.
Authors:Ziyun Chen, Fan Liu, Liang Yao, Chuanyi Zhang, Yuye Ma, Wei Zhou
Abstract:
The core objective of image captioning is to achieve lossless semantic compression from visual signals into textual modalities. However, the reliance on manually curated reference texts for evaluation essentially forces models to mimic specific human annotation styles, thereby masking the true descriptive capabilities of advanced foundation models. This systemic misalignment prompts a critical question: Is task‑specific fine‑tuning truly necessary for Remote Sensing Image Captioning, or is the perceived performance gap merely an artifact of flawed evaluation criteria? To investigate this discrepancy, we propose ReconScore, a novel reference‑free evaluation metric. Rather than computing textual similarities, we assess caption quality by its capability to reconstruct the original visual elements solely from the generated text, effectively neutralizing human annotation biases. Applying this metric, we uncover a profound, counterintuitive truth: inherently powerful, unfine‑tuned MLLMs surpass their fine‑tuned counterparts in authentic zero‑shot RSIC tasks. Driven by this structural discovery, we introduce RemoteDescriber, a completely training‑free generation methodology. By employing ReconScore as a self‑correction mechanism, we iteratively refine the semantic precision of MLLM outputs without any computational fine‑tuning overhead. Comprehensive experiments demonstrate that RemoteDescriber achieves state‑of‑the‑art performance on three datasets. Furthermore, we validate ReconScore's reliability and analyze the flaws of traditional metrics. Our code is available at https://github.com/hhu‑czy/RemoteDescriber.
Authors:Chao Pan, Xin Yao
Abstract:
Fast Adversarial Training (FastAT) seeks to achieve adversarial robustness at a fraction of the computational cost incurred by standard multi‑step methods such as PGD‑AT. Although numerous FastAT techniques have been proposed in recent years, fair comparison among them remains elusive. Existing benchmarks and public leaderboards typically permit diverse model architectures, varying training configurations, and external data sources, making it unclear whether reported improvements reflect genuine algorithmic advances or merely more favorable experimental conditions. To address this problem, we introduce the FastAT Benchmark, a controlled evaluation framework built on three core design principles: unified architecture requirements, standardized training settings, and strict prohibition of external or synthetic data. The benchmark implements over twenty representative FastAT methods within a single codebase, enabling direct and reproducible comparison. Each method is assessed through a dual‑metric evaluation framework that measures both adversarial robustness (accuracy under PGD, AutoAttack, and CR Attack) and computational cost (GPU training time and peak memory footprint). Comprehensive experiments on CIFAR‑10, CIFAR‑100, and Tiny‑ImageNet provide reliable baseline measurements and reveal that well‑designed single‑step methods can match or surpass PGD‑AT robustness at substantially lower cost, while no single method dominates across all evaluation dimensions. The complete benchmark, including source code, configuration files, and experimental results, is publicly available to support transparent and fair evaluation of future FastAT research.
Authors:Yiming Pan, Chengwei Hu, Xuancheng Huang, Can Huang, Mingming Zhao, Yuean Bi, Xiaohan Zhang, Aohan Zeng, Linmei Hu
Abstract:
Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a fundamental challenge: the generation process is text‑centric, whereas its quality is governed by visual aesthetics. This modality gap leads current models to frequently produce slides with aesthetically suboptimal layouts. Existing solutions typically rely either on heavy visual reflection, which incurs high inference cost yet yields limited gains; or on fine‑tuning with large‑scale datasets, which still provides weak and indirect aesthetic supervision. In contrast, the explicit use of aesthetic principles as supervision remains unexplored. In this work, we present AeSlides, a reinforcement learning framework with verifiable rewards for Aesthetic layout supervision in Slide generation. We introduce a suite of meticulously designed verifiable metrics to quantify slide layout quality, capturing key layout issues in an accurate, efficient, and low‑cost manner. Leveraging these verifiable metrics, we develop a GRPO‑based reinforcement learning method that directly optimizes slide generation models for aesthetically coherent layouts. With only 5K training prompts on GLM‑4.7‑Flash, AeSlides improves aspect ratio compliance from 36% to 85%, while reducing whitespace by 44%, element collisions by 43%, and visual imbalance by 28%. Human evaluation further shows a substantial improvement in overall quality, increasing scores from 3.31 to 3.56 (+7.6%), outperforming both model‑based reward optimization and reflection‑based agentic approaches, and even edging out Claude‑Sonnet‑4.5. These results demonstrate that such a verifiable aesthetic paradigm provides an efficient and scalable approach to aligning slide generation with human aesthetic preferences. Our repository is available at https://github.com/ympan0508/aeslides.
Authors:Xin Ning, Qiankun Li, Xiaolong Huang, Qiupu Chen, Feng He, Weijun Li, Prayag Tiwari, Xinwang Liu
Abstract:
With the accumulation of resources in the era of big data and the rise of pre‑trained models in deep learning, optimizing neural networks for various tasks often involves different strategies for fine‑tuning pre‑trained models versus training from scratch. However, existing optimizers primarily focus on reducing the loss function by updating model parameters, without fully addressing the unique demands of these two major paradigms. In this paper, we propose DualOpt, a novel approach that decouples optimization techniques specifically tailored for these distinct training scenarios. For training from scratch, we introduce real‑time layer‑wise weight decay, designed to enhance both convergence and generalization by aligning with the characteristics of weight updates and network architecture. For more importantly fine‑tuning, we integrate weight rollback with the optimizer, incorporating a rollback term into each weight update step. This ensures consistency in the weight distribution between upstream and downstream models, effectively mitigating knowledge forgetting and improving fine‑tuning performance. Additionally, we extend the layer‑wise weight decay to dynamically adjust the rollback levels across layers, adapting to the varying demands of different downstream tasks. Extensive experiments across diverse tasks, including image classification, object detection, semantic segmentation, and instance segmentation, demonstrate the broad applicability and state‑of‑the‑art performance of DualOpt. Code is available at https://github.com/qklee‑lz/OLOR‑AAAI‑2024.
Authors:Haonan Chen, Kaiwen Xiao, Bin Tian, Jun Fu
Abstract:
Autonomous parking remains a critical yet challenging task in intelligent driving systems, particularly within constrained urban environments where maneuvering space is limited and precise control is essential. While recent advances in end‑to‑end learning have shown great promise, the lack of high‑quality, structured datasets tailored for parking scenarios remains a significant bottleneck.To address this gap, we present ParkingScenes, a comprehensive multimodal dataset specifically designed for end‑to‑end autonomous parking in simulated scenes. Built on the CARLA simulator, ParkingScenes features structured parking trajectories generated by a Hybrid A planner and a Model Predictive Controller (MPC), providing accurate and reproducible supervision signals. The dataset includes 16 reverse‑in and 6 parallel parking scenarios, each executed under two pedestrian conditions (present and absent), resulting in 704 structured episodes and approximately 105000 frames. Each scenario is repeated 16 times to ensure consistent coverage. Each frame contains synchronized data from four RGB cameras, four depth sensors, vehicle motion states, and Bird's‑Eye View (BEV) representations, enabling rich multimodal fusion and context‑aware learning. To demonstrate the utility of our dataset, we compare models trained on ParkingScenes with those trained on unstructured, manually collected simulation data under identical conditions. Results show significant improvements in performance, underscoring the effectiveness of structured supervision for robust and accurate parking policy learning. By releasing both the dataset and the collection framework, ParkingScenes establishes a scalable and reproducible benchmark for advancing learning‑based autonomous parking systems. The dataset and collection framework will be released at: https://github.com/haonan‑ai/ParkingScenes
Authors:Jeremy Ellis
Abstract:
This paper presents webmcu‑vision‑web, a single‑file, zero‑install browser application for end‑to‑end TinyML vision model training and deployment on the Seeed Studio XIAO ESP32‑S3 Sense (XIAO ML Kit, 15‑‑40 USD). Acting as a browser‑based companion to the on‑device Arduino firmware of Paper 1 [1], it provides a private, fully local machine learning pipeline, from firmware flashing through image collection, CNN training, weight export, and live activation visualization, without any software installation beyond a Chromium‑based browser. The system targets educators, small businesses, and researchers who need to train task‑specific visual classifiers under their exact deployment conditions. Key capabilities include: in‑browser firmware flashing via esptool‑js; an SD card file browser with image preview and inline editing; config.json live‑sync for zero‑recompile hyperparameter adjustment; webcam and ESP32 OV2640 camera image capture; TensorFlow.js CNN training completing a three‑class run (~30 images per class, 20 epochs) in approximately 1 minute browser‑side versus 9 minutes on‑device, enabling a complete collect‑train‑deploy cycle in under 10 minutes; weight export as myWeights.bin and myWeights.h; confusion matrix; and a live Conv2 activation heatmap streamed from the ESP32 during inference. No data leaves the local machine at any stage. A five‑run consistency evaluation on the three‑class reference problem (0Blank, 1Cup, 2Pen) demonstrates stable convergence with mean accuracy and standard deviation reported; all artefacts are released at the repository link below. The repository is a living template for LLM‑assisted adaptation to new hardware and tasks. All source code is MIT‑licensed at https://github.com/webmcu‑ai/webmcu‑vision‑web.
Authors:Liyao Jiang, Ruichen Chen, Keith G. Mills
Abstract:
Pre‑training is a general method that is used in a range of deep learning tasks. By first training a model on one task, and then further training on the downstream task used for final evaluation, the model is forced to learn a more general understanding of the input data. While pre‑training has been applied to 3D Human Pose Estimation (HPE) previously, the scope of datasets used is typically very limited to some strong benchmarks, like Human3.6M. Therefore, in this project, we expand the scope of an existing 3D HPE scheme to be compatible with additional 2D and 3D HPE datasets, like Occlusion Person. We perform an extensive study on how aspects of 2D pre‑training, such as model size, affect downstream performance, and to what extent pre‑training can help the model generalize to different datasets. Experimental results show that 2D pre‑training consistently outperforms training on 3D data alone, particularly in terms of computational efficiency. Finally, using MPII and Human3.6M, we are able to obtain an MPJPE score of under 64.5mm.
Authors:Jinqi Cao, Zhiping Yu, Baihong Lin, Chenyang Liu, Zhenwei Shi, Zhengxia Zou
Abstract:
Recent generative AI models have achieved remarkable breakthroughs in language and visual understanding. However, although these models can generate realistic visual content, their spatial scale remains confined to bounded environments, preventing them from capturing how geographic environments evolve across thousands of kilometers or from modeling the spatial structure of the large‑scale physical world. This limitation poses a critical challenge for ultra‑wide‑area spatial intelligence in Earth observation and simulation, revealing a deeper gap in generative AI: progress has relied primarily on scaling model parameters and training data, while overlooking spatial scale as a core dimension of intelligence. Here, motivated by this missing dimension, we investigate spatial scale as a new scaling axis in foundation models and present MetaEarth3D, the first generative foundation model capable of spatially consistent generation at the planetary scale. Taking optical Earth observation simulation as a testbed, MetaEarth3D enables the generation of multi‑level, unbounded, and diverse 3D scenes spanning large‑scale terrains, medium‑scale cities, and fine‑grained street blocks. Built upon 10 million globally distributed real‑world training images, MetaEarth3D demonstrates both strong visual realism and geospatial statistical realism. Beyond generation, MetaEarth3D serves as a generative data engine for diverse virtual environments in ultra‑wide spatial intelligence. We argue that this study may help empower next‑generation spatial intelligence for Earth observation.
Authors:Rongxiao Guo, Qingchao Chen
Abstract:
Millimeter‑wave (mmWave) radar has shown great potential for contactless, privacy‑preserving, and robust human sensing, yet existing mmWave‑based human mesh reconstruction (HMR) studies are still limited by the lack of benchmarks for generalization analysis under configuration shifts and fair comparison of different algorithms. To address the limitation, we present DGHMesh, a large‑scale dual‑radar mmWave dataset and generalization‑focused benchmark for HMR. It contains data from 15 subjects performing 8 actions, with 360,000 synchronized frames collected from FMCW radar, SFCW radar, RGB images, and high‑precision 3D HMR annotations. In addition, the dataset provides synchronized raw I/Q data from both radar modalities and accurately calibrated radar spatial positions. The benchmark is designed to evaluate HMR methods under diverse measurement configurations, including human position shifts, human orientation shifts, subarray size variations, and cross‑subject settings. Based on DGHMesh, we also propose mmPTM, a query‑based multi‑radar fusion framework that jointly exploits point clouds and imaging tubes for HMR. Extensive experiments are conducted against representative baselines under different settings. The results demonstrate that mmPTM consistently achieves outstanding accuracy and competitive generalization capability across multiple sub‑benchmarks, validating the effectiveness of multi‑radar fusion and the practical value of the proposed dataset and benchmark for mmWave‑based HMR research. DGHMesh and mmPTM are publicly available at https://github.com/SPIresearch/DGHMesh.(The complete benchmark and code will be released after paper publication)
Authors:Bayangmbe Mounmo, Sam Chien, Mile Mitrovic
Abstract:
Industrial CAD workflows require robust, generalizable 3D geometric representations supporting accuracy and explainability. We introduce Shape, a self‑supervised foundation model converting surface meshes into dense per‑token embeddings. Shape combines a structured 3D latent grid, a multi‑scale geometry‑aware tokenizer (MAGNO) with cross‑attention, and a transformer processor using grouped‑query attention and RMSNorm. A learned reconstruction prior enables per‑region attribution for explainable predictions. Pretraining uses masked‑token reconstruction of normalized geometry statistics and multi‑resolution contrastive consistency. The 10.9M‑parameter backbone is pretrained on 61,052 CAD meshes from Thingi10K, MFCAD, and Fusion360. On a held‑out split of 2,983 meshes, Shape achieves reconstruction R2 = 0.729 and 98.1% top‑1 retrieval under the Wang‑Isola protocol, with near‑zero reconstruction train/val gap (contrastive scores use a larger evaluation pool). A 2x2 ablation on loss type and target‑space normalization shows per‑dimension normalization is critical: without it, performance collapses (R2 < 0.14, top‑1 < 88%); with it, both losses succeed (R2 > 0.70, top‑1 > 96%). Smooth‑L1 offers secondary stability. Code, embeddings, and an interactive demo are released at https://github.com/simd‑ai/shape.
Authors:Hyo Jin Jon, Longbin Jin, Eun Yi Kim
Abstract:
CLIP has demonstrated strong generalization in visual domains through natural language supervision, even for video action recognition. However, most existing approaches that adapt CLIP for action recognition have primarily focused on temporal modeling, often overlooking spatial perception. In real‑world scenarios, visual challenges such as low‑light environments or egocentric viewpoints can severely impair spatial understanding, an essential precursor for effective temporal reasoning. To address this limitation, we propose Efficient Visual Prompting for CLIP (EV‑CLIP), an efficient adaptation framework designed for few‑shot video action recognition across diverse scenes and viewpoints. EV‑CLIP introduces two visual prompts: mask prompts, which guide the model's attention to action‑relevant regions by reweighting pixels, and context prompts, which perform lightweight temporal modeling by compressing frame‑wise features into a compact representation. For a comprehensive evaluation, we curate five benchmark datasets and analyze domain shifts to quantify the influence of diverse visual and semantic factors on action recognition. Experimental results demonstrate that EV‑CLIP outperforms existing parameter‑efficient methods in overall performance. Moreover, its efficiency remains independent of the backbone scale, making it well‑suited for deployment in real‑world, resource‑constrained scenarios. The code is available at https://github.com/AI‑CV‑Lab/EV‑CLIP.
Authors:Ze Chen, Lan Chen, Yuanhang Li, Qi Mao
Abstract:
We propose FlowAnchor, a training‑free framework for stable and efficient inversion‑free, flow‑based video editing. Inversion‑free editing methods have recently shown impressive efficiency and structure preservation in images by directly steering the sampling trajectory with an editing signal. However, extending this paradigm to videos remains challenging, often failing in multi‑object scenes or with increased frame counts. We identify the root cause as the instability of the editing signal in high‑dimensional video latent spaces, which arises from imprecise spatial localization and length‑induced magnitude attenuation. To overcome this challenge, FlowAnchor explicitly anchors both where to edit and how strongly to edit. It introduces Spatial‑aware Attention Refinement, which enforces consistent alignment between textual guidance and spatial regions, and Adaptive Magnitude Modulation, which adaptively preserves sufficient editing strength. Together, these mechanisms stabilize the editing signal and guide the flow‑based evolution toward the desired target distribution. Extensive experiments demonstrate that FlowAnchor achieves more faithful, temporally coherent, and computationally efficient video editing across challenging multi‑object and fast‑motion scenarios. The project page is available at https://cuc‑mipg.github.io/FlowAnchor.github.io/.
Authors:Shunpeng Chen, Yukun Song, Changwei Wang, Rongtao Xu, Kexue Fu, Longxiang Gao, Li Guo, Ruisheng Wang, Shibiao Xu
Abstract:
Visual Place Recognition (VPR) determines a query image's geographic location by matching it against geotagged databases. However, existing methods struggle with perceptual aliasing caused by irrelevant regions and inefficient re‑ranking due to rigid candidate scheduling. To address these issues, we introduce FoL++, a method combining robust discriminative region modeling with adaptive re‑ranking. Specifically, we propose a Reliability Estimation Branch to generate spatial reliability maps that explicitly model occlusion resistance. This representation is further optimized by two spatial alignment losses (SAL and SCEL) to effectively align features and highlight salient regions. For weakly supervised learning without manual annotations, a pseudo‑correspondence strategy generates dense local feature supervision directly from aggregation clusters. Our Adaptive Candidate Scheduler dynamically resizes candidate pools based on global similarity. By weighting local matches by reliability and adaptively fusing global and local evidence, FoL++ surpasses traditional independent matching systems. Extensive experiments across seven benchmarks demonstrate that FoL++ achieves state‑of‑the‑art performance with a lightweight memory footprint, improving inference speed by 40% over FoL. Code and models will be released (and merged with FoL) at https://github.com/chenshunpeng/FoL.
Authors:Dominik Kuczkowski, Laura Ruotsalainen
Abstract:
Monocular visual odometry (VO) is a fundamental computer vision problem with applications in autonomous navigation, augmented reality and more. While deep learning‑based methods have recently shown superior accuracy compared to traditional geometric pipelines, particularly in environments where handcrafted features struggle due to poor structure or lighting conditions, most rely on deterministic regression, which lacks the uncertainty awareness required for robust applications. We propose PoseFM, the first framework to reformulate monocular frame‑to‑frame VO as a generative task using Flow Matching (FM). By leveraging FM, we model camera motion as a distribution rather than a point estimate, learning to transform noise into realistic pose predictions via continuous‑time ODEs. This approach provides a principled mechanism for uncertainty estimation and enables robust motion inference under challenging visual conditions. In our evaluations, PoseFM achieves strong performance on TartanAir, KITTI and TUM‑RGBD benchmarks, achieving the lowest absolute trajectory error (ATE) on some of the trajectories and overall being competitive with the best frame‑to‑frame monocular VO methods. Code and model checkpoints will be made available at https://github.com/helsinki‑sda‑group/posefm.
Authors:Dongwei Sun, Jing Yao, Kan Wei, Xiangyong Cao, Chen Wu, Zhenghui Zhao, Pedram Ghamisi, Jun Zhou, Jón Atli Benediktsson
Abstract:
Rapid situational awareness is critical in post‑disaster response. While remote sensing damage assessment is evolving from pixel‑level change detection to high‑level semantic analysis, existing vision‑language methodologies still struggle to provide actionable intelligence for complex strategic queries. They remain severely constrained by unimodal optical dependence, a prevailing bias towards natural disasters, and a fundamental lack of grounded interactivity. To address these limitations, we present ChangeQuery, a unified multimodal framework designed for comprehensive, all‑weather disaster situation awareness. To overcome modality constraints and scenario biases, we construct the Disaster‑Induced Change Query (DICQ) dataset, a large‑scale benchmark coupling pre‑event optical semantics with post‑event SAR structural features across a balanced distribution of natural catastrophes and armed conflicts. Furthermore, to provide the high‑quality supervision required for interactive reasoning, we propose a novel Automated Semantic Annotation Pipeline. Adhering to a ``statistics‑first, generation‑later'' paradigm, this engine automatically transforms raw segmentation masks into grounded, hierarchical instruction sets, effectively equipping the model with fine‑grained spatial and quantitative awareness. Trained on this structured data, the ChangeQuery architecture operates as an interactive disaster analyst. It supports multi‑task reasoning driven by diverse user queries, delivering precise damage quantification, region‑specific descriptions, and holistic post‑disaster summaries. Extensive experiments demonstrate that ChangeQuery establishes a new state‑of‑the‑art, providing a robust and interpretable solution for complex disaster monitoring. The code is available at \hrefhttps://sundongwei.github.io/changequery/https://sundongwei.github.io/changequery/.
Authors:Ran Zhao, Sheng Jin, Size Wu, Kang Liao, Zerui Gong, Zujin Guo, Yang Xiao, Wei Li
Abstract:
Recent text‑to‑image (T2I) models have demonstrated impressive capabilities in photorealistic synthesis and instruction following. However, their reliability in knowledge‑intensive settings remains largely unexplored. Unlike natural image generation, knowledge visualization requires not only semantic alignment but also strict adherence to domain knowledge, structural constraints, and symbolic conventions, exposing a critical gap between visual plausibility and scientific correctness. To systematically study this problem, we introduce KVBench, a curriculum‑grounded benchmark for evaluating knowledge‑intensive T2I generation. KVBench covers six senior high‑school subjects: Biology, Chemistry, Geography, History, Mathematics, and Physics. The benchmark consists of 1,800 expert‑curated prompts derived from over 30 authoritative textbooks. Using this benchmark, we evaluate 14 state‑of‑the‑art open‑ and closed‑source models, revealing substantial deficiencies in logical reasoning, symbolic precision, and multilingual robustness, with open‑source models consistently underperforming proprietary systems. To address these limitations, we further propose KE‑Check, a two‑stage framework that improves scientific fidelity via (1) Knowledge Elaboration for structured prompt enrichment, and (2) Checklist‑Guided Refinement for explicit constraint enforcement through violation identification and constraint‑guided editing. KE‑Check effectively mitigates scientific hallucinations, narrowing the performance gap between open‑source and leading closed‑source models. Data and codes are publicly available at https://github.com/zhaoran66/KVBench.
Authors:Aotian Zheng, Winston Sun, Bahaa Alattar, Vitaly Ablavsky, Jenq-Neng Hwang
Abstract:
CLIP‑based person re‑identification (ReID) methods aggregate spatial features into a single global \texttt[CLS] token optimized for image‑text alignment rather than spatial selectivity, making representations fragile under occlusion and cross‑camera variation. We propose SAGA‑ReID, which reconstructs identity representations by aligning intermediate patch tokens with anchor vectors parameterized in CLIP's text embedding space ‑‑ emphasizing spatially stable evidence while suppressing corrupted or absent regions, without requiring textual descriptions of individual images. Controlled experiments isolate the aggregation mechanism under two qualitatively distinct conditions ‑‑ synthetic masking, where identity signal is absent, and realistic human distractors, where an overlapping person introduces semantically confusing signal ‑‑ with SAGA's advantage over global pooling growing substantially as occlusion increases across both conditions. Benchmark evaluations confirm consistent gains over CLIP‑ReID across standard and occluded settings, with the largest improvements where global pooling is most unreliable: up to +10.6 Rank‑1 on occluded benchmarks. SAGA's aggregation outperforms dedicated sequential patch aggregation on a stronger backbone, confirming that structured reconstruction addresses a bottleneck that backbone quality and architectural complexity alone cannot resolve. Code available at https://github.com/ipl‑uw/Structured‑Anchor‑Guided‑Aggregation‑for‑ReID.
Authors:Peibo Song, Xiaotian Xue, Jinshuo Zhang, Zihao Wang, Jinhua Liu, Shujun Fu, Fangxun Bao, Si Yong Yeo
Abstract:
Multimodal MRI offers complementary information for brain tumor segmentation, but clinical scans often lack one or more modalities, which degrades segmentation performance. In this paper, we propose UniME (Uni‑Encoder Meets Multi‑Encoders), a two‑stage heterogeneous method for brain tumor segmentation with missing modalities that reconciles the trade‑offs among fine‑grained structure capture, cross‑modal complementarity modeling, and exploitation of available modalities. The idea is to decouple representation learning from segmentation via a two‑stage heterogeneous architecture. Stage 1 pretrains a single ViT Uni‑Encoder with masked image modeling to establish a unified representation robust to missing modalities. Stage 2 adds modality‑specific CNN Multi‑Encoders to extract high‑resolution, multi‑scale, fine‑grained features. We fuse these features with the global representation to produce precise segmentations. Experiments on BraTS 2023 and BraTS 2024 show that UniME outperforms previous methods under incomplete multi‑modal scenarios. The code is available at https://github.com/Hooorace‑S/UniME
Authors:Shozaburo Hirano, Norimichi Ukita
Abstract:
Automated sports analysis demands robust multi‑object tracking (MOT), yet segmentation‑based methods often struggle with mask errors and ID switches in dense scenes. We propose SAMIDARE, a framework that enhances SAM2MOT for crowded scenes through three key components: (1) density‑aware mask re‑generation and (2) selective memory updates, both for adaptive mask control to preserve target feature integrity, and (3) state‑aware association and new track initialization, which improves robustness under mutual occlusions and frequent frame‑out events. Evaluated on the SportsMOT dataset, SAMIDARE achieves state‑of‑the‑art performance, outperforming the baseline by 2.5 HOTA and 4.2 IDF1 points on the validation set. These results demonstrate that adaptive feature management using mask control and state‑aware association provide a robust and efficient solution for dense sports tracking. Code is available at https://github.com/ZabuZabuZabu/SAMIDARE
Authors:Weiqiu You, Cassandra Goldberg, Amin Madani, Daniel A. Hashimoto, Eric Wong
Abstract:
Purpose: Accurate assessment of the Critical View of Safety (CVS) during laparoscopic cholecystectomy is essential to prevent bile duct injury, a complication associated with significant morbidity and mortality. While large vision‑language models (LVLMs) offer flexible reasoning, their predictions remain difficult to audit and unreliable on safety‑critical surgical tasks.
Methods: We introduce Sum‑of‑Checks, a framework that decomposes each CVS criterion into expert‑defined reasoning checks reflecting clinically relevant visual evidence. Given a laparoscopic frame, an LVLM evaluates each check, producing a binary judgment and justification. Criterion‑level scores are computed via fixed, weighted aggregation of check outcomes. We evaluate on the Endoscapes2023 benchmark using three frontier LVLMs, comparing against direct prompting, chain‑of‑thought, and sub‑question decomposition, each with and without few‑shot examples.
Results: Sum‑of‑Checks improves average frame‑level mean average precision by 12‑‑14% relative to the best baseline across all three models and criteria. Analysis of individual checks reveals that LVLMs are reliable on observational checks (e.g., visibility, tool obstruction) but show substantial variability on decision‑critical anatomical evidence.
Conclusion: Structuring surgical reasoning into expert‑aligned verification checks improves both accuracy and transparency of LVLM‑based CVS assessment, demonstrating that explicitly separating evidence elicitation from decision‑making is critical for reliable and auditable surgical AI systems.
Code is available at https://github.com/BrachioLab/SumOfChecks.
Authors:Hao-Yu Hsu, Tianhang Cheng, Jing Wen, Alexander G. Schwing, Shenlong Wang
Abstract:
Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in privacy, safety, energy efficiency, and scalability. We explore an alternative: 4D perception without vision. Its goal is to reconstruct human motion and 3D scene layouts purely from everyday wearable sensors. For this we introduce IMU‑to‑4D, a framework that repurposes large language models for non‑visual spatiotemporal understanding of human‑scene dynamics. IMU‑to‑4D uses data from a few inertial sensors from earbuds, watches, or smartphones and predicts detailed 4D human motion together with coarse scene structure. Experiments across diverse human‑scene datasets show that IMU‑to‑4D yields more coherent and temporally stable results than SoTA cascaded pipelines, suggesting wearable motion sensors alone can support rich 4D understanding.
Authors:Kuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant, Ryan Burgert, Yuancheng Xu, Koichi Namekata, Yiwei Zhao, Bolei Zhou, Micah Goldblum, Paul Debevec, Ning Yu
Abstract:
We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re‑synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint. Existing video reshooting methods often struggle with depth estimation artifacts of real‑world dynamic videos, while also failing to preserve content appearance and failing to maintain precise camera control for challenging new trajectories. We build a 4D‑grounded point cloud representation with static pixel segmentation and 4D reconstruction to explicitly preserve seen content and provide rich camera signals, and we train with reconstructed multiview dynamic data for robustness against point cloud artifacts during real‑world inference. Our results demonstrate improved 4D consistency, camera control, and visual quality compared to state‑of‑the‑art baselines under a variety of videos and camera paths. Moreover, our method generalizes to real‑world applications such as dynamic scene expansion and 4D scene recomposition. See our project page for results, code, and models: https://eyeline‑labs.github.io/Vista4D
Authors:Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny, Mustafa Shukor, Alasdair Newson, Matthieu Cord
Abstract:
Despite impressive progress in capabilities of large vision‑language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the visual input. Prior work has attributed hallucinations in LVLMs to factors such as limitations of the vision backbone or the dominance of the language component, yet the relative importance of these factors remains unclear. To resolve this ambiguity, We propose HalluScope, a benchmark to better understand the extent to which different factors induce hallucinations. Our analysis indicates that hallucinations largely stem from excessive reliance on textual priors and background knowledge, especially information introduced through textual instructions. To mitigate hallucinations induced by textual instruction priors, we propose HalluVL‑DPO, a framework for fine‑tuning off‑the‑shelf LVLMs towards more visually grounded responses. HalluVL‑DPO leverages preference optimization using a curated training dataset that we construct, guiding the model to prefer grounded responses over hallucinated ones. We demonstrate that our optimized model effectively mitigates the targeted hallucination failure mode, while preserving or improving performance on other hallucination benchmarks and visual capability evaluations. To support reproducibility and further research, we will publicly release our evaluation benchmark, preference training dataset, and code at https://pegah‑kh.github.io/projects/prompts‑override‑vision/ .
Authors:Yanran Zhang, Wenzhao Zheng, Yifei Li, Bingyao Yu, Yu Zheng, Lei Chen, Jiwen Lu, Jie Zhou
Abstract:
In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fields have evolved distinct architectural paradigms: the former predominantly relies on generative networks, while the latter favors discriminative frameworks. A recent trend in both domains is the use of adversarial information to enhance performance, revealing potential for synergy. However, the significant architectural divergence between them presents considerable challenges. Departing from previous approaches, we propose UniGenDet: a Unified generative‑discriminative framework for co‑evolutionary image Generation and generated image Detection. To bridge the task gap, we design a symbiotic multimodal self‑attention mechanism and a unified fine‑tuning algorithm. This synergy allows the generation task to improve the interpretability of authenticity identification, while authenticity criteria guide the creation of higher‑fidelity images. Furthermore, we introduce a detector‑informed generative alignment mechanism to facilitate seamless information exchange. Extensive experiments on multiple datasets demonstrate that our method achieves state‑of‑the‑art performance. Code: \hrefhttps://github.com/Zhangyr2022/UniGenDethttps://github.com/Zhangyr2022/UniGenDet.
Authors:Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Yongqi Li, Liqiang Nie
Abstract:
Composed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that consists of a reference image and modification text. Although research on CIR has made significant progress, prevailing setups still rely simple modification texts that typically cover only a limited range of salient changes, which induces two limitations highly relevant to practical applications, namely Insufficient Entity Coverage and Clause‑Entity Misalignment. In order to address these issues and bring CIR closer to real‑world use, we construct two instruction‑rich multi‑modification datasets, M‑FashionIQ and M‑CIRR. In addition, we propose TEMA, the Text‑oriented Entity Mapping Architecture, which is the first CIR framework designed for multi‑modification while also accommodating simple modifications. Extensive experiments on four benchmark datasets demonstrate that TEMA's superiority in both original and multi‑modification scenarios, while maintaining an optimal balance between retrieval accuracy and computational efficiency. Our codes and constructed multi‑modification dataset (M‑FashionIQ and M‑CIRR) are available at https://github.com/lee‑zixu/ACL26‑TEMA/.
Authors:Safouane El Ghazouali, Nicola Venturi, Michael Rueegsegger, Umberto Michelucci
Abstract:
Recent advances in deep learning for remote sensing rely heavily on large annotated datasets, yet acquiring high‑quality ground truth for geometric, radiometric, and multi‑domain tasks remains costly and often infeasible. In particular, the lack of accurate depth annotations, controlled illumination variations, and multi‑scale paired imagery limits progress in monocular depth estimation, domain adaptation, and super‑resolution for aerial scenes. We present SyMTRS, a large‑scale synthetic dataset generated using a high‑fidelity urban simulation pipeline. The dataset provides high‑resolution RGB aerial imagery (2048 x 2048), pixel‑perfect depth maps, night‑time counterparts for domain adaptation, and aligned low‑resolution variants for super‑resolution at x2, x4, and x8 scales. Unlike existing remote sensing datasets that focus on a single task or modality, SyMTRS is designed as a unified multi‑task benchmark enabling joint research in geometric understanding, cross‑domain robustness, and resolution enhancement. We describe the dataset generation process, its statistical properties, and its positioning relative to existing benchmarks. SyMTRS aims to bridge critical gaps in remote sensing research by enabling controlled experiments with perfect geometric ground truth and consistent multi‑domain supervision. The results obtained in this work can be reproduced from this Github repository: https://github.com/safouaneelg/SyMTRS.
Authors:Katharina Prasse, Steffen Jung, Isaac Bravo, Stefanie Walter, Patrick Knab, Christian Bartelt, Margret Keuper
Abstract:
Social media platforms have become primary arenas for climate communication, generating millions of images and posts that ‑ if systematically analysed ‑ can reveal which communication strategies mobilise public concern and which fall flat. We aim to facilitate such research by analysing how computer vision methods can be used for social media discourse analysis. This analysis includes application‑based taxonomy design, model selection, prompt engineering, and validation. We benchmark six promptable vision‑language models and 15 zero‑shot CLIP‑like models on two datasets from X (formerly Twitter) ‑ a 1,038‑image expert‑annotated set and a larger corpus of over 1.2 million images, with 50,000 labels manually validated ‑ spanning five annotation dimensions: animal content, climate change consequences, climate action, image setting, and image type. Among the models benchmarked, Gemini‑3.1‑flash‑lite outperforms all others across all super‑categories and both datasets, while the gap to open‑weight models of moderate size remains relatively small. Beyond instance‑level metrics, we advocate for distributional evaluation: VLM predictions can reliably recover population level trends even when per‑image accuracy is moderate, making them a viable starting point for discourse analysis at scale. We find that chain‑of‑thought reasoning reduces rather than improves performance, and that annotation dimension specific prompt design improves performance. We release tweet IDs and labels along with our code at https://github.com/KathPra/Codebooks2VLMs.git.
Authors:Avinash Paliwal, Adithya Iyer, Shivin Yadav, Muhammad Ali Afridi, Midhun Harikumar
Abstract:
Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi‑view data for non‑rigid scenes. We overcome this limitation with a highly scalable self‑supervised framework capable of leveraging internet‑scale monocular videos. Our core contribution is the generation of pseudo multi‑view training triplets, consisting of a source video, a geometric anchor, and a target video. We achieve this by extracting distinct smooth random‑walk crop trajectories from a single input video to serve as the source and target views. The anchor is synthetically generated by forward‑warping the first frame of the source with a dense tracking field, which effectively simulates the distorted point‑cloud inputs expected at inference. Because our independent cropping strategy introduces spatial misalignment and artificial occlusions, the model cannot simply copy information from the current source frame. Instead, it is forced to implicitly learn 4D spatiotemporal structures by actively routing and re‑projecting missing high‑fidelity textures across distinct times and viewpoints from the source video to reconstruct the target. At inference, our minimally adapted diffusion transformer utilizes a 4D point‑cloud derived anchor to achieve state‑of‑the‑art temporal consistency, robust camera control, and high‑fidelity novel view synthesis on complex dynamic scenes.
Authors:Dat To-Thanh, Nghia Nguyen-Trong, Hoang Vo, Hieu Bui-Minh, Tinh-Anh Nguyen-Nhu
Abstract:
Image enhancement models for mobile devices often struggle to balance high output quality with the fast processing speeds required by mobile hardware. While recent deep learning models can enhance low‑quality mobile photos into high‑quality images, their performance is often degraded when converted to lower‑precision formats for actual use on mobile phones. To address this training‑deployment mismatch, we propose an efficient image enhancement model designed specifically for mobile deployment. Our approach uses a hierarchical network architecture with gated encoder blocks and multiscale refinement to preserve fine‑grained visual features. Moreover, we incorporate Quantization‑Aware Training (QAT) to simulate the effects of low‑precision representation during the training process. This allows the network to adapt and prevents the typical drop in quality seen with standard post‑training quantization (PTQ). Experimental results demonstrate that the proposed method produces high‑fidelity visual output while maintaining the low computational overhead needed for practical use on standard mobile devices. The code will be available at https://github.com/GenAI4E/QATIE.git.
Authors:Wenxuan Bao, Yanjun Zhao, Xiyuan Yang, Jingrui He
Abstract:
Pretrained vision‑language models such as CLIP exhibit strong zero‑shot generalization but remain sensitive to distribution shifts. Test‑time adaptation adapts models during inference without access to source data or target labels, offering a practical way to handle such shifts. However, existing methods typically assume that test samples come from a single, consistent domain, while in practice, test data often include samples from mixed domains with distinct characteristics. Consequently, their performance degrades under mixed‑domain settings. To address this, we present Ramen, a framework for robust test‑time adaptation through active sample selection. For each incoming test sample, Ramen retrieves a customized batch of relevant samples from previously seen data based on two criteria: domain consistency, which ensures that adaptation focuses on data from similar domains, and prediction balance, which mitigates adaptation bias caused by skewed predictions. To improve efficiency, Ramen employs an embedding‑gradient cache that stores the embeddings and sample‑level gradients of past test images. The stored embeddings are used to retrieve relevant samples, and the corresponding gradients are aggregated for model updates, eliminating the need for any additional forward or backward passes. Our theoretical analysis provides insight into why the proposed adaptation mechanism is effective under mixed‑domain shifts. Experiments on multiple image corruption and domain‑shift benchmarks demonstrate that Ramen achieves strong and consistent performance, offering robust and efficient adaptation in complex mixed‑domain scenarios. Our code is available at https://github.com/baowenxuan/Ramen .
Authors:Zhiqiu Lin, Chancharik Mitra, Siyuan Cen, Isaac Li, Yuhan Huang, Yu Tong Tiffany Ling, Hewei Wang, Irene Pi, Shihang Zhu, Ryan Rao, George Liu, Jiaxi Li, Ruojin Li, Yili Han, Yilun Du, Deva Ramanan
Abstract:
Video‑language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we define a structured specification for describing subjects, scenes, motion, spatial, and camera dynamics, grounded by hundreds of carefully defined visual primitives developed with professional video creators such as filmmakers. Next, to curate high‑quality captions, we introduce CHAI (Critique‑based Human‑AI Oversight), a framework where trained experts critique and revise model‑generated pre‑captions into improved post‑captions. This division of labor improves annotation accuracy and efficiency by offloading text generation to models, allowing humans to better focus on verification. Additionally, these critiques and preferences between pre‑ and post‑captions provide rich supervision for improving open‑source models (Qwen3‑VL) on caption generation, reward modeling, and critique generation through SFT, DPO, and inference‑time scaling. Our ablations show that critique quality in precision, recall, and constructiveness, ensured by our oversight framework, directly governs downstream performance. With modest expert supervision, the resulting model outperforms closed‑source models such as Gemini‑3.1‑Pro. Finally, we apply our approach to re‑caption large‑scale professional videos (e.g., films, commercials, games) and fine‑tune video generation models such as Wan to better follow detailed prompts of up to 400 words, achieving finer control over cinematography including camera motion, angle, lens, focus, point of view, and framing. Our results show that precise specification and human‑AI oversight are key to professional‑level video understanding and generation. Data and code are available on our project page: https://linzhiqiu.github.io/papers/chai/
Authors:Guangkai Xu, Hua Geng, Huanyi Zheng, Songyi Yin, Yanlong Sun, Hao Chen, Chunhua Shen
Abstract:
Feed‑forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi‑frame models usually produce better cross‑frame consistency, yet they often underperform strong per‑frame methods on single‑frame accuracy. This observation motivates our systematic investigation into the critical factors driving model performance through rigorous ablation studies, which reveals several key insights: 1) Scaling up data diversity and quality unlocks further performance gains even in state‑of‑the‑art visual geometry estimation methods; 2) Commonly adopted confidence‑aware loss and gradient‑based loss mechanisms may unintentionally hinder performance; 3) Joint supervision through both per‑sequence and per‑frame alignment improves results, while local region alignment surprisingly degrades performance. Furthermore, we introduce two enhancements to integrate the advantages of optimization‑based methods and high‑resolution inputs: a consistency loss function that enforces alignment between depth maps, camera parameters, and point maps, and an efficient architectural design that leverages high‑resolution information. We integrate these designs into CARVE, a resolution‑enhanced model for feed‑forward visual geometry estimation. Experiments on point cloud reconstruction, video depth estimation, and camera pose/intrinsic estimation show that CARVE achieves strong and robust performance across diverse benchmarks.
Authors:Kwan Yun, Changmin Lee, Ayeong Jeong, Youngseo Kim, Seungmi Lee, Junyong Noh
Abstract:
Creative face stylization aims to render portraits in diverse visual idioms such as cartoons, sketches, and paintings while retaining recognizable identity. However, current identity encoders, which are typically trained and calibrated on natural photographs, exhibit severe brittleness under stylization. They often mistake changes in texture or color palette for identity drift or fail to detect geometric exaggerations. This reveals the lack of a style‑agnostic framework to evaluate and supervise identity consistency across varying styles and strengths. To address this gap, we introduce StyleID, a human perception‑aware dataset and evaluation framework for facial identity under stylization. StyleID comprises two datasets: (i) StyleBench‑H, a benchmark that captures human same‑different verification judgments across diffusion‑ and flow‑matching‑based stylization at multiple style strengths, and (ii) StyleBench‑S, a supervision set derived from psychometric recognition‑strength curves obtained through controlled two‑alternative forced‑choice (2AFC) experiments. Leveraging StyleBench‑S, we fine‑tune existing semantic encoders to align their similarity orderings with human perception across styles and strengths. Experiments demonstrate that our calibrated models yield significantly higher correlation with human judgments and enhanced robustness for out‑of‑domain, artist drawn portraits. All of our datasets, code, and pretrained models are publicly available at https://kwanyun.github.io/StyleID_page/
Authors:Rawal Khirodkar, He Wen, Julieta Martinez, Yuan Dong, Su Zhaoen, Shunsuke Saito
Abstract:
We present Sapiens2, a model family of high‑resolution transformers for human‑centric vision focused on generalization, versatility, and high‑fidelity outputs. Our model sizes range from 0.4 to 5 billion parameters, with native 1K resolution and hierarchical variants that support 4K. Sapiens2 substantially improves over its predecessor in both pretraining and post‑training. First, to learn features that capture low‑level details (for dense prediction) and high‑level semantics (for zero‑shot or few‑label settings), we combine masked image reconstruction with self‑distilled contrastive objectives. Our evaluations show that this unified pretraining objective is better suited for a wider range of downstream tasks. Second, along the data axis, we pretrain on a curated dataset of 1 billion high‑quality human images and improve the quality and quantity of task annotations. Third, architecturally, we incorporate advances from frontier models that enable longer training schedules with improved stability. Our 4K models adopt windowed attention to reason over longer spatial context and are pretrained with 2K output resolution. Sapiens2 sets a new state‑of‑the‑art and improves over the first generation on pose (+4 mAP), body‑part segmentation (+24.3 mIoU), normal estimation (45.6% lower angular error) and extends to new tasks such as pointmap and albedo estimation. Code: https://github.com/facebookresearch/sapiens2
Authors:Yao Zhang, Zhuchenyang Liu, Thomas Ploetz, Yu Xiao
Abstract:
The world knowledge and reasoning capabilities of text‑based large language models (LLMs) are advancing rapidly, yet current approaches to human motion understanding, including motion question answering and captioning, have not fully exploited these capabilities. Existing LLM‑based methods typically learn motion‑language alignment through dedicated encoders that project motion features into the LLM's embedding space, remaining constrained by cross‑modal representation and alignment. Inspired by biomechanical analysis, where joint angles and body‑part kinematics have long served as a precise descriptive language for human movement, we propose Structured Motion Description (SMD), a rule‑based, deterministic approach that converts joint position sequences into structured natural language descriptions of joint angles, body part movements, and global trajectory. By representing motion as text, SMD enables LLMs to apply their pretrained knowledge of body parts, spatial directions, and movement semantics directly to motion reasoning, without requiring learned encoders or alignment modules. We show that this approach goes beyond state‑of‑the‑art results on both motion question answering (66.7% on BABEL‑QA, 90.1% on HuMMan‑QA) and motion captioning (R@1 of 0.584, CIDEr of 53.16 on HumanML3D), surpassing all prior methods. SMD additionally offers practical benefits: the same text input works across different LLMs with only lightweight LoRA adaptation (validated on 8 LLMs from 6 model families), and its human‑readable representation enables interpretable attention analysis over motion descriptions. Code, data, and pretrained LoRA adapters are available at https://yaozhang182.github.io/motion‑smd/.
Authors:Zeyu Cai, Yuliang Xiu, Renke Wang, Zhijing Shao, Xiaoben Li, Siyuan Yu, Chao Xu, Yang Liu, Baigui Sun, Jian Yang, Zhenyu Zhang
Abstract:
Fitting an underlying body model to 3D clothed human assets has been extensively studied, yet most approaches focus on either single‑modal inputs such as point clouds or multi‑view images alone, often requiring a known metric scale. This constraint is frequently impractical, especially for AI‑generated assets where scale distortion is common. We propose OmniFit, a method that can seamlessly handle diverse multi‑modal inputs, including full scans, partial depth observations, and image captures, while remaining scale‑agnostic for both real and synthetic assets. Our key innovation is a simple yet effective conditional transformer decoder that directly maps surface points to dense body landmarks, which are then used for SMPL‑X parameter fitting. In addition, an optional plug‑and‑play image adapter incorporates visual cues to compensate for missing geometric information. We further introduce a dedicated scale predictor that rescales subjects to canonical body proportions. OmniFit substantially outperforms state‑of‑the‑art methods by 57.1 to 80.9 percent across daily and loose clothing scenarios. To the best of our knowledge, it is the first body fitting method to surpass multi‑view optimization baselines and the first to achieve millimeter‑level accuracy on the CAPE and 4D‑DRESS benchmarks.
Authors:Shiyan Su, Ruyi Zha, Danli Shi, Hongdong Li, Xuelian Cheng
Abstract:
Neural representations (NRs), such as neural fields and 3D Gaussians, effectively model volumetric data in computed tomography (CT) but suffer from severe artifacts under sparse‑view settings. To address this, we propose DiffNR, a novel framework that enhances NR optimization with diffusion priors. At its core is SliceFixer, a single‑step diffusion model designed to correct artifacts in degraded slices. We integrate specialized conditioning layers into the network and develop tailored data curation strategies to support model finetuning. During reconstruction, SliceFixer periodically generates pseudo‑reference volumes, providing auxiliary 3D perceptual supervision to fix underconstrained regions. Compared to prior methods that embed CT solvers into time‑consuming iterative denoising, our repair‑and‑augment strategy avoids frequent diffusion model queries, leading to better runtime performance. Extensive experiments show that DiffNR improves PSNR by 3.99 dB on average, generalizes well across domains, and maintains efficient optimization.
Authors:Yuhan Luo, Tao Chen, Decheng Liu
Abstract:
Nowadays, visual data forgery detection plays an increasingly important role in social and economic security with the rapid development of generative models. Existing face forgery detectors still can't achieve satisfactory performance because of poor generalization ability across datasets. The key factor that led to this phenomenon is the lack of suitable metrics: the commonly used cross‑dataset AUC metric fails to reveal an important issue where detection scores may shift significantly across data domains. To explicitly evaluate cross‑domain score comparability, we propose Cross‑AUC, an evaluation metric that can compute AUC across dataset pairs by contrasting real samples from one dataset with fake samples from another (and vice versa). It is interesting to find that evaluating representative detectors under the Cross‑AUC metric reveals substantial performance drops, exposing an overlooked robustness problem. Besides, we also propose the novel framework Semantic Fine‑grained Alignment and Mixture‑of‑Experts (SFAM), consisting of a patch‑level image‑text alignment module that enhances CLIP's sensitivity to manipulation artifacts, and the facial region mixture‑of‑experts module, which routes features from different facial regions to specialized experts for region‑aware forgery analysis. Extensive qualitative and quantitative experiments on the public datasets prove that the proposed method achieves superior performance compared with the state‑of‑the‑art methods with various suitable metrics.
Authors:Chentao Li, Zirui Gao, Mingze Gao, Yinglian Ren, Jianjiang Feng, Jie Zhou
Abstract:
Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite advancements in Multimodal Large Language Models (MLLMs), current systems often fail to precisely ground the spatial semantics of pointing. Instead, they rely on spurious correlations with visual proximity or object saliency, a phenomenon we term "Referential Hallucination." To address this gap, we introduce EgoPoint‑Bench, a comprehensive question‑answering benchmark designed to evaluate and enhance multimodal pointing reasoning in egocentric views. Comprising over 11k high‑fidelity simulated and real‑world samples, the benchmark spans five evaluation dimensions and three levels of referential complexity. Extensive experiments demonstrate that while state‑of‑the‑art proprietary and open‑source models struggle with egocentric pointing, models fine‑tuned on our synthetic data achieve significant performance gains and robust sim‑to‑real generalization. This work highlights the importance of spatially aware supervision and offers a scalable path toward precise egocentric AI assistants. Project page: https://guyyyug.github.io/EgoPoint‑Bench/
Authors:Yixuan Zhu, Shilin Ma, Haolin Wang, Ao Li, Yanzhe Jing, Yansong Tang, Lei Chen, Jiwen Lu, Jie Zhou
Abstract:
Recent advancements in visual autoregressive models (VAR) have demonstrated their effectiveness in image generation, highlighting their potential for real‑world image super‑resolution (Real‑ISR). However, adapting VAR for ISR presents critical challenges. The next‑scale prediction mechanism, constrained by causal attention, fails to fully exploit global low‑quality (LQ) context, resulting in blurry and inconsistent high‑quality (HQ) outputs. Additionally, error accumulation in the iterative prediction severely degrades coherence in ISR task. To address these issues, we propose VARestorer, a simple yet effective distillation framework that transforms a pre‑trained text‑to‑image VAR model into a one‑step ISR model. By leveraging distribution matching, our method eliminates the need for iterative refinement, significantly reducing error propagation and inference time. Furthermore, we introduce pyramid image conditioning with cross‑scale attention, which enables bidirectional scale‑wise interactions and fully utilizes the input image information while adapting to the autoregressive mechanism. This prevents later LQ tokens from being overlooked in the transformer. By fine‑tuning only 1.2% of the model parameters through parameter‑efficient adapters, our method maintains the expressive power of the original VAR model while significantly enhancing efficiency. Extensive experiments show that VARestorer achieves state‑of‑the‑art performance with 72.32 MUSIQ and 0.7669 CLIPIQA on DIV2K dataset, while accelerating inference by 10 times compared to conventional VAR inference.
Authors:Jingfang Li, Haoran Zhu, Wen Yang, Jinrui Zhang, Fang Xu, Haijian Zhang, Gui-Song Xia
Abstract:
Ultra‑High‑Resolution (UHR) imagery has become essential for modern remote sensing, offering unprecedented spatial coverage. However, detecting small objects in such vast scenes presents a critical dilemma: retaining the original resolution for small objects causes prohibitive memory bottlenecks. Conversely, conventional compromises like image downsampling or patch cropping either erase small objects or destroy context. To break this dilemma, we propose UHR‑DETR, an efficient end‑to‑end transformer‑based detector designed for UHR imagery. First, we introduce a Coverage‑Maximizing Sparse Encoder that dynamically allocates finite computational resources to informative high‑resolution regions, ensuring maximum object coverage with minimal spatial redundancy. Second, we design a Global‑Local Decoupled Decoder. By integrating macroscopic scene awareness with microscopic object details, this module resolves semantic ambiguities and prevents scene fragmentation. Extensive experiments on the UHR imagery datasets (e.g., STAR and SODA‑A) demonstrate the superiority of UHR‑DETR under strict hardware constraints (e.g., a single 24GB RTX 3090). It achieves a 2.8% mAP improvement while delivering a 10× inference speedup compared to standard sliding‑window baselines on the STAR dataset. Our codes and models will be available at https://github.com/Li‑JingFang/UHR‑DETR.
Authors:Javier Sanguino, Carlos Platero, Olga Velasco
Abstract:
This paper deals with the case of using nonlinear diffusion filters to obtain piecewise constant images as a previous process for segmentation techniques.
We first show an intrinsic formulation for the nonlinear diffusion equation to provide some design conditions on the diffusion filters. According to this theoretical framework, we propose a new family of diffusivities; they are obtained from nonlinear diffusion techniques and are related with backward diffusion. Their goal is to split the image in closed contours with a homogenized grey intensity inside and with no blurred edges.
We also prove that our filters satisfy the well‑posedness semi‑discrete and full discrete scale‑space requirements. This shows that by using semi‑implicit schemes, a forward nonlinear diffusion equation is solved, instead of a backward nonlinear diffusion equation, connecting with an edge‑preserving process. Under the conditions established for the diffusivity and using a stopping criterion for the diffusion time, we get piecewise constant images with a low computational effort.
Finally, we test our filter with real images and we illustrate the effects of our diffusivity function as a method to get piecewise constant images.
The code is available at https://github.com/cplatero/NonlinearDiffusion.
Authors:Jinrang Jia, Zhenjia Li, Yifeng Shi
Abstract:
3D Gaussian Splatting (3DGS) has revolutionized neural rendering, yet existing methods remain predominantly research prototypes ill‑suited for production‑level deployment. We identify a critical "Industry‑Academia Gap" hindering real‑world application: unpredictable resource consumption from heuristic Gaussian growth, the "sparsity shield" of current benchmarks that rewards hallucination over physical fidelity, and severe multi‑sensor data pollution. To bridge this gap, we propose YOGO (You Only Gaussian Once), a system‑level framework that reformulates the stochastic growth process into a deterministic, budget‑aware equilibrium. YOGO integrates a novel budget controller for hardware‑constrained resource allocation and an availability‑registration protocol for robust multi‑sensor fusion. To push the boundaries of reconstruction fidelity, we introduce Immersion v1.0, the first ultra‑dense indoor dataset specifically designed to break the "sparsity shield." By providing saturated viewpoint coverage, Immersion v1.0 forces algorithms to focus on extreme physical fidelity rather than viewpoint interpolation, and enables the community to focus on the upper limits of high‑fidelity reconstruction. Extensive experiments demonstrate that YOGO achieves state‑of‑the‑art visual quality while maintaining a strictly deterministic profile, establishing a new standard for production‑grade 3DGS. To facilitate reproducibility, part scenes of Immersion v1.0 dataset and source code of YOGO has been publicly released. The project link is https://jjrcn.github.io/yogo‑project‑home/
Authors:Vishal Rajput
Abstract:
PGD adversarial training, the standard robustness method, can reduce Jacobian Frobenius norm yet worsen clean‑input geometry (e.g., TDI 1.336 vs. ERM 1.093). We show this is not an implementation artifact but a theorem‑level consequence of supervised learning.
We prove that any encoder minimizing supervised loss must retain non‑zero sensitivity along directions correlated with training labels, including directions that are nuisance at test time. This holds across proper scoring rules, architectures, and dataset sizes. We call this the geometric blind spot of supervised learning.
This theorem unifies four empirical phenomena often treated separately: non‑robust features, texture bias, corruption fragility, and the robustness‑accuracy tradeoff. It also explains why suppressing sensitivity in one adversarial direction can redistribute sensitivity elsewhere.
We introduce Trajectory Deviation Index (TDI), a diagnostic of geometric isotropy. Unlike CKA, intrinsic dimension, or Jacobian Frobenius norm alone, TDI captures the failure mode above. In our experiments, PGD attains low Frobenius norm but high TDI, while PMH attains the lowest TDI with one additional training term and no architectural changes.
Across seven tasks, BERT/SST‑2, and ImageNet ViT‑B/16 (backbone family underlying CLIP/DINO/SAM), the blind spot is measurable and repairable. It appears at foundation‑model scale, worsens with model scale and task‑specific fine‑tuning, and is substantially reduced by PMH. PMH also leads on non‑Gaussian corruption types (blur/brightness/contrast) without corruption‑specific training.
Authors:Linkai Liu, Wei Feng, Xi Zhao, Shen Zhang, Xingye Chen, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Yuchen Zhou, Zipeng Guo, Chao Gou
Abstract:
Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a significant focus of recent research. However, while CG has advanced considerably, most efforts have concentrated on generating advertising text and images, leaving Creative Video Generation (CVG) relatively underexplored. This gap is largely due to two major challenges faced by Text‑to‑Video (T2V) models: (a) ambiguous semantic alignment, where models struggle to accurately correlate product selling points with creative video content, and (b) inadequate motion adaptability, resulting in unrealistic movements and distortions. To address these challenges, we develop a comprehensive Advertising Creative Knowledge Base (ACKB) as a foundational resource and propose a knowledge‑driven approach (KD‑CVG) to overcome the knowledge limitations of existing models. KD‑CVG consists of two primary modules: Semantic‑Aware Retrieval (SAR) and Multimodal Knowledge Reference (MKR). SAR utilizes the semantic awareness of graph attention networks and reinforcement learning feedback to enhance the model's comprehension of the connections between selling points and creative videos. Building on this, MKR incorporates semantic and motion priors into the T2V model to address existing knowledge gaps. Extensive experiments have demonstrated KD‑CVG's superior performance in achieving semantic alignment and motion adaptability, validating its effectiveness over other state‑of‑the‑art methods. The code and dataset will be open source at https://kdcvg.github.io/KDCVG/.
Authors:Wadii Boulila, Adel Ammar, Bilel Benjdira, Maha Driss
Abstract:
Self‑supervised learning (SSL) is a standard approach for representation learning in aerial imagery. Existing methods enforce invariance between augmented views, which works well when augmentations preserve semantic content. However, aerial images are frequently degraded by haze, motion blur, rain, and occlusion that remove critical evidence. Enforcing alignment between a clean and a severely degraded view can introduce spurious structure into the latent space. This study proposes a training strategy and architectural modification to enhance SSL robustness to such corruptions. It introduces a per‑sample, per‑factor trust weight into the alignment objective, combined with the base contrastive loss as an additive residual. A stop‑gradient is applied to the trust weight instead of a multiplicative gate. While a multiplicative gate is a natural choice, experiments show it impairs the backbone, whereas our additive‑residual approach improves it. Using a 200‑epoch protocol on a 210,000‑image corpus, the method achieves the highest mean linear‑probe accuracy among six backbones on EuroSAT, AID, and NWPU‑RESISC45 (90.20% compared to 88.46% for SimCLR and 89.82% for VICReg). It yields the largest improvements under severe information‑erasing corruptions on EuroSAT (+19.9 points on haze at s=5 over SimCLR). The method also demonstrates consistent gains of +1 to +3 points in Mahalanobis AUROC on a zero‑shot cross‑domain stress test using BDD100K weather splits. Two ablations (scalar uncertainty and cosine gate) indicate the additive‑residual formulation is the primary source of these improvements. An evidential variant using Dempster‑Shafer fusion introduces interpretable signals of conflict and ignorance. These findings offer a concrete design principle for uncertainty‑aware SSL. Code is publicly available at https://github.com/WadiiBoulila/trust‑ssl.
Authors:Dhruv Parikh, Jacob Fein-Ashley, Rajgopal Kannan, Viktor Prasanna
Abstract:
Large Multimodal Models (LMMs) such as LLaVA are typically trained with an autoregressive language modeling objective, providing only indirect supervision to visual tokens. This often yields weak internal visual representations and brittle behavior under distribution shift. Inspired by recent progress on latent denoising for learning high‑quality visual tokenizers, we show that the same principle provides an effective form of visual supervision for improving internal visual feature alignment and multimodal understanding in LMMs. We propose a latent denoising framework that corrupts projected visual tokens using a saliency‑aware mixture of masking and Gaussian noising. The LMM is trained to denoise these corrupted tokens by recovering clean teacher patch features from hidden states at a selected intermediate LLM layer using a decoder. To prevent representation collapse, our framework also preserves the teacher's intra‑image similarity structure and applies intra‑image contrastive patch distillation. During inference, corruption and auxiliary heads are disabled, introducing no additional inference‑time overhead. Across a broad suite of standard multimodal benchmarks, our method consistently improves visual understanding and reasoning over strong baselines, and yields clear gains on compositional robustness benchmarks (e.g., NaturalBench). Moreover, under ImageNet‑C‑style non‑adversarial common corruptions applied to benchmark images, our method maintains higher accuracy and exhibits reduced degradation at both moderate and severe corruption levels. Our code is available at https://github.com/dhruvashp/latent‑denoising‑for‑lmms.
Authors:Kai Liu, Haoyang Yue, Zeli Lin, Zheng Chen, Jingkai Wang, Jue Gong, Jiatong Li, Xianglong Yan, Libo Zhu, Jianze Li, Ziqing Zhang, Zihan Zhou, Xiaoyang Liu, Radu Timofte, Yulun Zhang, Junye Chen, Zhenming Yan, Yucong Hong, Ruize Han, Song Wang, Li Pang, Heng Zhao, Xinqiao Wu, Deyu Meng, Xiangyong Cao, Weijun Yuan, Zhan Li, Zhanglu Chen, Boyang Yao, Yihang Chen, Yifan Deng, Zengyuan Zuo, Junjun Jiang, Saiprasad Meesiyawar, Sulocha Yatageri, Nikhil Akalwadi, Ramesh Ashok Tabib, Uma Mudenagudi, Jiachen Tu, Yaokun Shi, Guoyi Xu, Yaoxin Jiang, Cici Liu, Tongyao Mu, Qiong Cao, Yifan Wang, Kosuke Shigematsu, Hiroto Shirono, Asuka Shin, Wei Zhou, Linfeng Li, Lingdong Kong, Ce Wang, Xingwei Zhong, Wanjie Sun, Dafeng Zhang, Hongxin Lan, Qisheng Xu, Mingyue He, Hui Geng, Tianjiao Wan, Kele Xu, Changjian Wang, Antoine Carreaud, Nicola Santacroce, Shanci Li, Jan Skaloud, Adrien Gressin
Abstract:
This paper presents the NTIRE 2026 Remote Sensing Infrared Image Super‑Resolution (x4) Challenge, one of the associated challenges of NTIRE 2026. The challenge aims to recover high‑resolution (HR) infrared images from low‑resolution (LR) inputs generated through bicubic downsampling with a x4 scaling factor. The objective is to develop effective models or solutions that achieve state‑of‑the‑art performance for infrared image SR in remote sensing scenarios. To reflect the characteristics of infrared data and practical application needs, the challenge adopts a single‑track setting. A total of 115 participants registered for the competition, with 13 teams submitting valid entries. This report summarizes the challenge design, dataset, evaluation protocol, main results, and the representative methods of each team. The challenge serves as a benchmark to advance research in infrared image super‑resolution and promote the development of effective solutions for real‑world remote sensing applications.
Authors:Yuki Fujimura, Takahiro Kushida, Kazuya Kitano, Takuya Funatomi, Yasuhiro Mukaigawa
Abstract:
We propose WildSplatter, a feed‑forward 3D Gaussian Splatting (3DGS) model for unconstrained images with unknown camera parameters and varying lighting conditions. 3DGS is an effective scene representation that enables high‑quality, real‑time rendering; however, it typically requires iterative optimization and multi‑view images captured under consistent lighting with known camera parameters. WildSplatter is trained on unconstrained photo collections and jointly learns 3D Gaussians and appearance embeddings conditioned on input images. This design enables flexible modulation of Gaussian colors to represent significant variations in lighting and appearance. Our method reconstructs 3D Gaussians from sparse input views in under one second, while also enabling appearance control under diverse lighting conditions. Experimental results demonstrate that our approach outperforms existing pose‑free 3DGS methods on challenging real‑world datasets with varying illumination.
Authors:Yalcin Tur, Mihajlo Stojkovic, Ulas Bagci
Abstract:
Diffusion models have achieved remarkable quality in multi‑modal MRI synthesis, but their computational cost (hundreds of sampling steps and separate models per modality) limits clinical deployment. We observe that this inefficiency stems from an unnecessary starting point: diffusion begins from pure noise, discarding the structural information already present in available MRI sequences. We propose WFM (Wavelet Flow Matching), which instead learns a direct flow from an informed prior, the mean of conditioning modalities in wavelet space, to the target distribution. Because the source and target share underlying anatomy and differ primarily in contrast, this formulation enables accurate synthesis in just 1‑2 integration steps. A single 82M‑parameter model with class conditioning synthesizes all four BraTS modalities (T1, T1c, T2, FLAIR), replacing four separate diffusion models totaling 326M parameters. On BraTS 2024, WFM achieves 26.8 dB PSNR and 0.94 SSIM, within 1‑2 dB of diffusion baselines, while running 250‑1000x faster (0.16‑0.64s vs. 160s per volume). This speed‑quality trade‑off makes real‑time MRI synthesis practical for clinical workflows. Code is available at https://github.com/yalcintur/WFM.
Authors:Zahid Hassan Tushar, Sanjay Purushotham
Abstract:
The NASA PACE mission provides unprecedented hyperspectral observations of ocean color, aerosols, and clouds, offering new insights into how these components interact and influence Earth's climate and air quality. Its Ocean Color Instrument measures light across hundreds of finely spaced wavelength bands, enabling detailed characterization of features such as phytoplankton composition, aerosol properties, and cloud microphysics. However, hyperspectral data of this scale is large, complex, and difficult to label, requiring specialized processing and analysis techniques. Existing foundation models, which have transformed computer vision and natural language processing, are generally trained on standard RGB imagery and therefore struggle to interpret the continuous spectral signatures captured by PACE. While recent advances have introduced hyperspectral foundation models, they are typically trained on cloud‑free observations and often remain limited to single‑sensor datasets due to spectral inconsistencies across instruments. Moreover, existing models tend to be parameter‑heavy and computationally expensive, limiting scalability and adoption in operational settings. To address these challenges, we introduce HyperFM, a parameter‑efficient hyperspectral foundation model that leverages intra‑group and inter‑group spectral attention along with hybrid parameter decomposition to better capture spectral spatial relationships while reducing computational cost. HyperFM demonstrates consistent performance improvements over existing hyperspectral foundation models and task‑specific state‑of‑the‑art methods across four benchmark downstream atmospheric cloud property retrieval tasks. To support further research, we additionally release HyperFM250K, a large‑scale hyperspectral dataset from the PACE mission that includes both clear and cloudy scenes.
Authors:Mahnoor Fatima Saad, Sagnik Majumder, Kristen Grauman, Ziad Al-Halah
Abstract:
Rings like gold, thuds like wood! The sound we hear in a scene is shaped not only by the spatial layout of the environment but also by the materials of the objects and surfaces within it. For instance, a room with wooden walls will produce a different acoustic experience from a room with the same spatial layout but concrete walls. Accurately modeling these effects is essential for applications such as virtual reality, robotics, architectural design, and audio engineering. Yet, existing methods for acoustic modeling often entangle spatial and material influences in correlated representations, which limits user control and reduces the realism of the generated acoustics. In this work, we present a novel approach for material‑controlled Room Impulse Response (RIR) generation that explicitly disentangles the effects of spatial and material cues in a scene. Our approach models the RIR using two modules: a spatial module that captures the influence of the spatial layout of the scene, and a material module that modulates this spatial RIR according to a user‑specified material configuration. This explicitly disentangled design allows users to easily modify the material configuration of a scene and observe its impact on acoustics without altering the spatial structure or scene content. Our model provides significant improvements over prior approaches on both acoustic‑based metrics (up to +16% on RTE) and material‑based metrics (up to +70%). Furthermore, through a human perceptual study, we demonstrate the improved realism and material sensitivity of our model compared to the strongest baselines.
Authors:Amandeep Kaur, Mirali Purohit, Gedeon Muhawenayo, Esther Rolf, Hannah Kerner
Abstract:
New geospatial foundation models introduce a new model architecture and pretraining dataset, often sampled using different notions of data diversity. Performance differences are largely attributed to the model architecture or input modalities, while the role of the pretraining dataset is rarely studied. To address this research gap, we conducted a systematic study on how the geographic composition of pretraining data affects a model's downstream performance. We created global and per‑continent pretraining datasets and evaluated them on global and per‑continent downstream datasets. We found that the pretraining dataset from Europe outperformed global and continent‑specific pretraining datasets on both global and local downstream evaluations. To investigate the factors influencing a pretraining dataset's downstream performance, we analysed 10 pretraining datasets using diversity across continents, biomes, landcover and spectral values. We found that only spectral diversity was strongly correlated with performance, while others were weakly correlated. This finding establishes a new dimension of diversity to be accounted for when creating a high‑performing pretraining dataset. We open‑sourced 7 new pretraining datasets, pretrained models, and our experimental framework at https://github.com/kerner‑lab/pretrain‑where.
Authors:Hyeonwoo Kim, Jeonghwan Kim, Kyungwon Cho, Hanbyul Joo
Abstract:
Recent advances in video generative models enable the synthesis of realistic human‑object interaction videos across a wide range of scenarios and object categories, including complex dexterous manipulations that are difficult to capture with motion capture systems. While the rich interaction knowledge embedded in these synthetic videos holds strong potential for motion planning in dexterous robotic manipulation, their limited physical fidelity and purely 2D nature make them difficult to use directly as imitation targets in physics‑based character control. We present DeVI (Dexterous Video Imitation), a novel framework that leverages text‑conditioned synthetic videos to enable physically plausible dexterous agent control for interacting with unseen target objects. To overcome the imprecision of generative 2D cues, we introduce a hybrid tracking reward that integrates 3D human tracking with robust 2D object tracking. Unlike methods relying on high‑quality 3D kinematic demonstrations, DeVI requires only the generated video, enabling zero‑shot generalization across diverse objects and interaction types. Extensive experiments demonstrate that DeVI outperforms existing approaches that imitate 3D human‑object interaction demonstrations, particularly in modeling dexterous hand‑object interactions. We further validate the effectiveness of DeVI in multi‑object scenes and text‑driven action diversity, showcasing the advantage of using video as an HOI‑aware motion planner.
Authors:Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Ahmed Arafa, Minhaj Nur Alam
Abstract:
Federated learning (FL) enables collaborative model training without sharing raw data; however, the presence of noisy labels across distributed clients can severely degrade the learning performance. In this paper, we propose FedSIR, a multi‑stage framework for robust FL under noisy labels. Different from existing approaches that mainly rely on designing noise‑tolerant loss functions or exploiting loss dynamics during training, our method leverages the spectral structure of client feature representations to identify and mitigate label noise.
Our framework consists of three key components. First, we identify clean and noisy clients by analyzing the spectral consistency of class‑wise feature subspaces with minimal communication overhead. Second, clean clients provide spectral references that enable noisy clients to relabel potentially corrupted samples using both dominant class directions and residual subspaces. Third, we employ a noise‑aware training strategy that integrates logit‑adjusted loss, knowledge distillation, and distance‑aware aggregation to further stabilize federated optimization. Extensive experiments on standard FL benchmarks demonstrate that FedSIR consistently outperforms state‑of‑the‑art methods for FL with noisy labels. The code is available at https://github.com/sinagh72/FedSIR.
Authors:Shelly Golan, Michael Finkelson, Ariel Bereslavsky, Yotam Nitzan, Or Patashnik
Abstract:
Reinforcement Learning (RL) post‑training has become the standard for aligning generative models with human preferences, yet most methods rely on a single scalar reward. When multiple criteria matter, the prevailing practice of ``early scalarization'' collapses rewards into a fixed weighted sum. This commits the model to a single trade‑off point at training time, providing no inference‑time control over inherently conflicting goals ‑‑ such as prompt adherence versus source fidelity in image editing. We introduce ParetoSlider, a multi‑objective RL (MORL) framework that trains a single diffusion model to approximate the entire Pareto front. By training the model with continuously varying preference weights as a conditioning signal, we enable users to navigate optimal trade‑offs at inference time without retraining or maintaining multiple checkpoints. We evaluate ParetoSlider across three state‑of‑the‑art flow‑matching backbones: SD3.5, FluxKontext, and LTX‑2. Our single preference‑conditioned model matches or exceeds the performance of baselines trained separately for fixed reward trade‑offs, while uniquely providing fine‑grained control over competing generative goals.
Authors:Yonatan Haile Medhanie, Yuanhua Ni
Abstract:
Transformer‑based OCR models have shown strong performance on Latin and CJK scripts, but their application to African syllabic writing systems remains limited. We present the first adaptation of TrOCR for printed Tigrinya using the Ge'ez script. Starting from a pre‑trained model, we extend the byte‑level BPE tokenizer to cover 230 Ge'ez characters and introduce Word‑Aware Loss Weighting to resolve systematic word‑boundary failures that arise when applying Latin‑centric BPE conventions to a new script. The unmodified model produces no usable output on Ge'ez text. After adaptation, the TrOCR‑Printed variant achieves 0.22% Character Error Rate and 97.20% exact match accuracy on a held‑out test set of 5,000 synthetic images from the GLOCR dataset. An ablation study confirms that Word‑Aware Loss Weighting is the critical component, reducing CER by two orders of magnitude compared to vocabulary extension alone. The full pipeline trains in under three hours on a single 8 GB consumer GPU. All code, model weights, and evaluation scripts are publicly released.
Authors:Dimitrije Antić, Alvaro Budria, George Paschalidis, Sai Kumar Dwivedi, Dimitrios Tzionas
Abstract:
Reconstructing 3D Human‑Object Interaction from an RGB image is essential for perceptive systems. Yet, this remains challenging as it requires capturing the subtle physical coupling between the body and objects. While current methods rely on sparse, binary contact cues, these fail to model the continuous proximity and dense spatial relationships that characterize natural interactions. We address this limitation via InterFields, a representation that encodes dense, continuous proximity across the entire body and object surfaces. However, inferring these fields from single images is inherently ill‑posed. To tackle this, our intuition is that interaction patterns are characteristically structured by the action and object geometry. We capture this structure in LEXIS, a novel discrete manifold of interaction signatures learned via a VQ‑VAE. We then develop LEXIS‑Flow, a diffusion framework that leverages LEXIS signatures to estimate human and object meshes alongside their InterFields. Notably, these InterFields help in a guided refinement that ensures physically‑plausible, proximity‑aware reconstructions without requiring post‑hoc optimization. Evaluation on Open3DHOI and BEHAVE shows that LEXIS‑Flow significantly outperforms existing SotA baselines in reconstruction, contact, and proximity quality. Our approach not only improves generalization but also yields reconstructions perceived as more realistic, moving us closer to holistic 3D scene understanding. Code & models will be public at https://anticdimi.github.io/lexis.
Authors:Inclusion AI, Tiwei Bie, Haoxing Chen, Tieyuan Chen, Zhenglin Cheng, Long Cui, Kai Gan, Zhicheng Huang, Zhenzhong Lan, Haoquan Li, Jianguo Li, Tao Lin, Qi Qin, Hongjun Wang, Xiaomei Wang, Haoyuan Wu, Yi Xin, Junbo Zhao
Abstract:
We present LLaDA2.0‑Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its architecture combines a fully semantic discrete tokenizer, a MoE‑based dLLM backbone, and a diffusion decoder. By discretizing continuous visual inputs via SigLIP‑VQ, the model enables block‑level masked diffusion for both text and vision inputs within the backbone, while the decoder reconstructs visual tokens into high‑fidelity images. Inference efficiency is enhanced beyond parallel decoding through prefix‑aware optimizations in the backbone and few‑step distillation in the decoder. Supported by carefully curated large‑scale data and a tailored multi‑stage training pipeline, LLaDA2.0‑Uni matches specialized VLMs in multimodal understanding while delivering strong performance in image generation and editing. Its native support for interleaved generation and reasoning establishes a promising and scalable paradigm for next‑generation unified foundation models. Codes and models are available at https://github.com/inclusionAI/LLaDA2.0‑Uni.
Authors:Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr, Federico Tombari, Bernt Schiele
Abstract:
Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large language models (MLLMs). However, the reliance on language‑centric priors and expensive manual annotations prevents MLLMs' intrinsic visual understanding and scalable reward designs. In this work, we introduce SSL‑R1, a generic self‑supervised RL framework that derives verifiable rewards directly from images. To this end, we revisit self‑supervised learning (SSL) in visual domains and reformulate widely‑used SSL tasks into a set of verifiable visual puzzles for RL post‑training, requiring neither human nor external model supervision. Training MLLMs on these tasks substantially improves their performance on multimodal understanding and reasoning benchmarks, highlighting the potential of leveraging vision‑centric self‑supervised tasks for MLLM post‑training. We think this work will provide useful experience in devising effective self‑supervised verifiable rewards to enable RL at scale. Project page: https://github.com/Jiahao000/SSL‑R1.
Authors:Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr, Federico Tombari, Bernt Schiele
Abstract:
Large vision‑language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they still struggle with object hallucinations, i.e., the claim of nonexistent objects in the visual input. To address this challenge, we propose Region‑aware Chain‑of‑Verification (R‑CoV), a visual chain‑of‑verification method to alleviate object hallucinations in LVLMs in a post‑hoc manner. Motivated by how humans comprehend intricate visual information ‑‑ often focusing on specific image regions or details within a given sample ‑‑ we elicit such region‑level processing from LVLMs themselves and use it as a chaining cue to detect and alleviate their own object hallucinations. Specifically, our R‑CoV consists of six steps: initial response generation, entity extraction, coordinate generation, region description, verification execution, and final response generation. As a simple yet effective method, R‑CoV can be seamlessly integrated into various LVLMs in a training‑free manner and without relying on external detection models. Extensive experiments on several widely used hallucination benchmarks across multiple LVLMs demonstrate that R‑CoV can significantly alleviate object hallucinations in LVLMs. Project page: https://github.com/Jiahao000/R‑CoV.
Authors:Qian Chen, Yuehao Chen, Qiang Wang, Lei Zhu, Yanye Lu, Qiushi Ren
Abstract:
Retinal laser speckle contrast imaging (LSCI) is a noninvasive optical modality for monitoring retinal blood flow dynamics. However, conventional temporal LSCI (tLSCI) reconstruction relies on sufficiently long speckle sequences to obtain stable temporal statistics, which makes it vulnerable to acquisition disturbances and limits effective temporal resolution. A physically informed reconstruction framework, termed RetinaDiff (Retinal Diffusion Model), is proposed for retinal tLSCI that is robust to motion and requires only a few frames. In RetinaDiff, registration based on phase correlation is first applied to stabilize the raw speckle sequence before contrast computation, reducing interframe misalignment so that fluctuations at each pixel primarily reflect true flow dynamics. This step provides a physics prior corrected for motion and a high quality multiframe tLSCI reference. Next, guided by the physics prior, a conditional diffusion model performs inverse reconstruction by jointly conditioning on the registered speckle sequence and the corrected prior. Experiments on data acquired with a retinal LSCI system developed in house show improved structural continuity and statistical stability compared with direct reconstruction from few frames and representative baselines. The framework also remains effective in a small number of extremely challenging cases, where both the direct 5‑frame input and the conventional multiframe reconstruction are severely degraded. Overall, this work provides a practical and physically grounded route for reliable retinal tLSCI reconstruction from extremely limited frames. The source code and model weights will be publicly available at https://github.com/QianChen113/RetinaDiff.
Authors:Muzhi Zhu, Shunyao Jiang, Huanyi Zheng, Zekai Luo, Hao Zhong, Anzhou Li, Kaijun Wang, Jintao Rong, Yang Liu, Hao Chen, Tao Lin, Chunhua Shen
Abstract:
Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern generative or unified multimodal models also possess generative spatial intelligence (GSI), the ability to respect and manipulate 3D spatial constraints during image generation, and whether such capability can be measured or improved. We introduce GSI‑Bench, the first benchmark designed to quantify GSI through spatially grounded image editing. It consists of two complementary components: GSI‑Real, a high‑quality real‑world dataset built via a 3D‑prior‑guided generation and filtering pipeline, and GSI‑Syn, a large‑scale synthetic benchmark with controllable spatial operations and fully automated labeling. Together with a unified evaluation protocol, GSI‑Bench enables scalable, model‑agnostic assessment of spatial compliance and editing fidelity. Experiments show that fine‑tuning unified multimodal models on GSI‑Syn yields substantial gains on both synthetic and real tasks and, strikingly, also improves downstream spatial understanding. This provides the first clear evidence that generative training can tangibly strengthen spatial reasoning, establishing a new pathway for advancing spatial intelligence in multimodal models.
Authors:Qizhong Tan, Zhuotao Tian, Guangming Lu, Jun Yu, Wenjie Pei
Abstract:
Existing Video Large Language Models (Video LLMs) struggle with complex video understanding, exhibiting limited reasoning capabilities and potential hallucinations. In particular, these methods tend to perform reasoning solely relying on the pretrained inherent reasoning rationales whilst lacking perception‑aware adaptation to the input video content. To address this, we propose Video‑ToC, a novel video reasoning framework that enhances video understanding through tree‑of‑cue reasoning. Specifically, our approach introduces three key innovations: (1) A tree‑guided visual cue localization mechanism, which endows the model with enhanced fine‑grained perceptual capabilities through structured reasoning patterns; (2) A reasoning‑demand reward mechanism, which dynamically adjusts the reward value for reinforcement learning (RL) based on the estimation of reasoning demands, enabling on‑demand incentives for more effective reasoning strategies; and (3) An automated annotation pipeline that constructs the Video‑ToC‑SFT‑1k and Video‑ToC‑RL‑2k datasets for supervised fine‑tuning (SFT) and RL training, respectively. Extensive evaluations on six video understanding benchmarks and a video hallucination benchmark demonstrate the superiority of Video‑ToC over baselines and recent methods. Code is available at https://github.com/qizhongtan/Video‑ToC.
Authors:Yongji Long, Shijun Liang, Jintao Li, Yun Li
Abstract:
Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing critical long‑range information in complex dynamics. To address this issue, we propose DynamicRad, a unified sparse‑attention paradigm that grounds adaptive selection within a radial locality prior. DynamicRad introduces a dual‑mode strategy: static‑ratio for speed‑optimized execution and dynamic‑threshold for quality‑first filtering. To ensure robustness without online search overhead, we integrate an offline Bayesian Optimization (BO) pipeline coupled with a semantic motion router. This lightweight projection module maps prompt embeddings to optimal sparsity regimes with minimal runtime overhead. Unlike online profiling methods, our offline BO optimizes attention reconstruction error (MSE) on a physics‑based proxy task, ensuring rapid convergence. Experiments on HunyuanVideo and Wan2.1‑14B demonstrate that DynamicRad pushes the efficiency‑‑quality Pareto frontier, achieving 1.7×‑‑2.5× inference speedups with over 80% effective sparsity. In some long‑sequence settings, the dynamic mode even matches or exceeds the dense baseline, while mask‑aware LoRA further improves long‑horizon coherence. Code is available at https://github.com/Adamlong3/DynamicRad.
Authors:Kuanwei Chen, Tingyi Lin
Abstract:
Sign‑language datasets are difficult to preprocess consistently because they vary in annotation schema, clip timing, signer framing, and privacy constraints. Existing work usually reports downstream models, while the preprocessing pipeline that converts raw video into training‑ready pose or video artifacts remains fragmented, backend‑specific, and weakly documented. We present SignDATA, a config‑driven preprocessing toolkit that standardizes heterogeneous sign‑language corpora into comparable outputs for learning. The system supports two end‑to‑end recipes: a pose recipe that performs acquisition, manifesting, person localization, clipping, cropping, landmark extraction, normalization, and WebDataset export, and a video recipe that replaces pose extraction with signer‑cropped video packaging. SignDATA exposes interchangeable MediaPipe and MMPose backends behind a common interface, typed job schemas, experiment‑level overrides, and per‑stage checkpointing with config‑ and manifest‑aware hashes. We validate the toolkit through a research‑oriented evaluation design centered on backend comparison, preprocessing ablations, and privacy‑aware video generation on datasets. Our contribution is a reproducible preprocessing layer for sign‑language research that makes extractor choice, normalization policy, and privacy tradeoffs explicit, configurable, and empirically comparable.Code is available at https://github.com/balaboom123/signdata‑slt.
Authors:Gui Wang, Zehao Zhong, YongSong Zhou, Yudong Li, Ende Wu, Wooi Ping Cheah, Rong Qu, Jianfeng Ren, Linlin Shen
Abstract:
Despite significant progress in Multi‑modal Large Language Models (MLLMs), their clinical reasoning capacity for multi‑modal diagnosis remains largely unexamined. Current benchmarks, mostly single‑modality data, can't evaluate progressive reasoning and cross‑modal integration essential for clinical practice. We introduce the Cross‑Modality Progressive Clinical Reasoning (X‑PCR) benchmark, the first comprehensive evaluation of MLLMs through a complete ophthalmology diagnostic workflow, with two reasoning tasks: 1) a six‑stage progressive reasoning chain spanning image quality assessment to clinical decision‑making, and 2) a cross‑modality reasoning task integrating six imaging modalities. The benchmark comprises 26,415 images and 177,868 expert‑verified VQA pairs curated from 51 public datasets, covering 52 ophthalmic diseases. Evaluation of 21 MLLMs reveals critical gaps in progressive reasoning and cross‑modal integration. Dataset and code: https://github.com/CVI‑SZU/X‑PCR.
Authors:Jiahao Xu, Xiaohan Yuan, Xingchen Wu, Chongyang Xu, Kun Li, Buzhen Huang
Abstract:
Co‑manipulation requires multiple humans to synchronize their motions with a shared object while ensuring reasonable interactions, maintaining natural poses, and preserving stable states. However, most existing motion generation approaches are designed for single‑character scenarios or fail to account for payload‑induced dynamics. In this work, we propose a flow‑matching framework that ensures the generated co‑manipulation motions align with the intended goals while maintaining naturalness and effectiveness. Specifically, we first introduce a generative model that derives explicit manipulation strategies from the object's affordance and spatial configuration, which guide the motion flow toward successful manipulation. To improve motion quality, we then design an adversarial interaction prior that promotes natural individual poses and realistic inter‑person interactions during co‑manipulation. In addition, we also incorporate a stability‑driven simulation into the flow matching process, which refines unstable interaction states through sampling‑based optimization and directly adjusts the vector field regression to promote more effective manipulation. The experimental results demonstrate that our method achieves higher contact accuracy, lower penetration, and better distributional fidelity compared to state‑of‑the‑art human‑object interaction baselines. The code is available at https://github.com/boycehbz/StaCOM.
Authors:Tao Cheng, Shi-Zhe Chen, Hao Zhang, Yixin Qin, Jinwen Luo, Zheng Wei
Abstract:
Chain‑of‑Thought (CoT) reasoning significantly elevates the complex problem‑solving capabilities of multimodal large language models (MLLMs). However, adapting CoT to vision typically discretizes signals to fit LLM inputs, causing early semantic collapse and discarding fine‑grained details. While external tools can mitigate this, they introduce a rigid bottleneck, confining reasoning to predefined operations. Although recent latent reasoning paradigms internalize visual states to overcome these limitations, optimizing the resulting hybrid discrete‑continuous action space remains challenging. In this work, we propose HyLaR (Hybrid Latent Reasoning), a framework that seamlessly interleaves discrete text generation with continuous visual latent representations. Specifically, following an initial cold‑start supervised fine‑tuning (SFT), we introduce DePO (Decoupled Policy Optimization) to enable effective reinforcement learning within this hybrid space. DePO decomposes the policy gradient objective, applying independent trust‑region constraints to the textual and latent components, alongside an exact closed‑form von Mises‑Fisher (vMF) KL regularizer. Extensive experiments demonstrate that HyLaR outperforms standard MLLMs and state‑of‑the‑art latent reasoning approaches across fine‑grained perception and general multimodal understanding benchmarks. Code is available at https://github.com/EthenCheng/HyLaR.
Authors:Gui Wang, YongSong Zhou, Kaijun Deng, Wooi Ping Cheah, Rong Qu, Jianfeng Ren, Linlin Shen
Abstract:
Fine‑grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi‑modal Large Language Models (MLLMs) in this domain remain largely unexplored. To bridge this gap, we introduce SurgCoT, a unified benchmark for evaluating chain‑of‑thought (CoT) reasoning in MLLMs across 7 surgical specialties and 35 diverse procedures. SurgCoT assesses five core reasoning dimensions: Causal Action Ordering, Cue‑Action Alignment, Affordance Mapping, Micro‑Transition Localization, and Anomaly Onset Tracking, through a structured CoT framework with an intensive annotation protocol (Question‑Option‑Knowledge‑Clue‑Answer), where the Knowledge field provides essential background context and Clue provides definitive spatiotemporal evidence. Evaluation of 10 leading MLLMs shows: 1) commercial models outperform open‑source and medical‑specialized variants; 2) significant gaps exist in surgical CoT reasoning; 3) SurgCoT enables effective evaluation and enhances progressive spatiotemporal reasoning. SurgCoT provides a reproducible testbed to narrow the gap between MLLM capabilities and clinical reasoning demands. Code: https://github.com/CVI‑SZU/SurgCoT.
Authors:Md Maklachur Rahman, Soon Ki Jung, Tracy Hammond
Abstract:
Recent segmentation models have demonstrated promising efficiency by aggressively reducing parameter counts and computational complexity. However, these models often struggle to accurately delineate fine lesion boundaries and texture patterns essential for early skin cancer diagnosis and treatment planning. In this paper, we propose MambaLiteUNet, a compact yet robust segmentation framework that integrates Mamba state space modeling into a U‑Net architecture, along with three key modules: Adaptive Multi‑Branch Mamba Feature Fusion (AMF), Local‑Global Feature Mixing (LGFM), and Cross‑Gated Attention (CGA). These modules are designed to enhance local‑global feature interaction, preserve spatial details, and improve the quality of skip connections. MambaLiteUNet achieves an average IoU of 87.12% and average Dice score of 93.09% across ISIC2017, ISIC2018, HAM10000, and PH2 benchmarks, outperforming state‑of‑the‑art models. Compared to U‑Net, our model improves average IoU and Dice by 7.72 and 4.61 points, respectively, while reducing parameters by 93.6% and GFLOPs by 97.6%. Additionally, in domain generalization with six unseen lesion categories, MambaLiteUNet achieves 77.61% IoU and 87.23% Dice, performing best among all evaluated models. Our extensive experiments demonstrate that MambaLiteUNet achieves a strong balance between accuracy and efficiency, making it a competitive and practical solution for dermatological image segmentation. Our code is publicly available at: https://github.com/maklachur/MambaLiteUNet.
Authors:Minghong Wei, Pu Cao, Zhihao Chen, Zhiyuan Zang, Lu Yang, Qing Song
Abstract:
With the rapid advancement of intelligent driving and remote sensing, oriented object detection has gained widespread attention. However, achieving high‑precision performance is fundamentally constrained by the Angle Boundary Discontinuity (ABD) and Cyclic Ambiguity (CA) problems, which typically cause significant angle fluctuations near periodic boundaries. Although recent studies propose continuous angle coders to alleviate these issues, our theoretical and empirical analyses reveal that state‑of‑the‑art methods still suffer from substantial cyclic errors. We attribute this instability to the structural noise amplification within their non‑orthogonal decoding mechanisms. This mathematical vulnerability significantly exacerbates angular deviations, particularly for square‑like objects. To resolve this fundamentally, we propose the Fourier Series Coder (FSC), a lightweight plug‑and‑play component that establishes a continuous, reversible, and mathematically robust angle encoding‑decoding paradigm. By rigorously mapping angles onto a minimal orthogonal Fourier basis and explicitly enforcing a geometric manifold constraint, FSC effectively prevents feature modulus collapse. This structurally stabilized representation ensures highly robust phase unwrapping, intrinsically eliminating the need for heuristic truncations while achieving strict boundary continuity and superior noise immunity. Extensive experiments across three large‑scale datasets demonstrate that FSC achieves highly competitive overall performance, yielding substantial improvements in high‑precision detection. The code will be available at https://github.com/weiminghong/FSC.
Authors:Mobin Habibpour, Niloufar Alipour Talemi, John Spodnik, Camren J. Khoury, Fatemeh Afghah
Abstract:
Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) benchmarks do not evaluate wildfire‑specific multimodal reasoning grounded in thermal measurements. We introduce WildFireVQA, a large‑scale VQA benchmark for aerial wildfire monitoring that integrates RGB imagery with radiometric thermal data. WildFireVQA contains 6,097 RGB‑thermal samples, where each sample includes an RGB image, a color‑mapped thermal visualization, and a radiometric thermal TIFF, and is paired with 34 questions, yielding a total of 207,298 multiple‑choice questions spanning presence and detection, classification, distribution and segmentation, localization and direction, cross‑modal reasoning, and flight planning for operational wildfire intelligence. To improve annotation reliability, we combine multimodal large language model (MLLM)‑based answer generation with sensor‑driven deterministic labeling, manual verification, and intra‑frame and inter‑frame consistency checks. We further establish a comprehensive evaluation protocol for representative MLLMs under RGB, Thermal, and retrieval‑augmented settings using radiometric thermal statistics. Experiments show that across task categories, RGB remains the strongest modality for current models, while retrieved thermal context yields gains for stronger MLLMs, highlighting both the value of temperature‑grounded reasoning and the limitations of existing MLLMs in safety‑critical wildfire scenarios. The dataset and benchmark code are open‑source at https://github.com/mobiiin/WildFire_VQA.
Authors:Byunghyun Kim
Abstract:
We propose Semantic‑Fast‑SAM (SFS), a semantic segmentation framework that combines the Fast Segment Anything model with a semantic labeling pipeline to achieve real‑time performance without sacrificing accuracy. FastSAM is an efficient CNN‑based re‑implementation of the Segment Anything Model (SAM) that runs much faster than the original transformer‑based SAM. Building upon FastSAM's rapid mask generation, we integrate a Semantic‑Segment‑Anything (SSA) labeling strategy to assign meaningful categories to each mask. The resulting SFS model produces high‑quality semantic segmentation maps at a fraction of the computational cost and memory footprint of the original SAM‑based approach. Experiments on Cityscapes and ADE20K benchmarks demonstrate that SFS matches the accuracy of prior SAM‑based methods (mIoU ~ 70.33 on Cityscapes and 48.01 on ADE20K) while achieving approximately 20x faster inference than SSA in the closed‑set setting. We also show that SFS effectively handles open‑vocabulary segmentation by leveraging CLIP‑based semantic heads, outperforming recent open‑vocabulary models on broad class labeling. This work enables practical real‑time semantic segmentation with the "segment‑anything" capability, broadening the applicability of foundation segmentation models in robotics scenarios. The implementation is available at https://github.com/KBH00/Semantic‑Fast‑SAM.
Authors:Xi Chen, Arian Maleki, Shirin Jalali
Abstract:
Multi‑look acquisition is a widely used strategy for reducing speckle noise in coherent imaging systems such as digital holography. By acquiring multiple measurements, speckle can be suppressed through averaging or joint reconstruction, typically under the assumption that speckle realizations across looks are statistically independent. In practice, however, hardware constraints limit measurement diversity, leading to inter‑look correlation that degrades the performance of conventional methods. In this work, we study the reconstruction of speckle‑free reflectivity from complex‑valued multi‑look measurements in the presence of correlated speckle. We model the inter‑look dependence using a first‑order Markov process and derive the corresponding likelihood under a first‑order Markov approximation, resulting in a constrained maximum likelihood estimation problem. To solve this problem, we develop an efficient projected gradient descent framework that combines gradient‑based updates with implicit regularization via deep image priors, and leverages Monte Carlo approximation and matrix‑free operators for scalable computation. Simulation results demonstrate that the proposed approach remains robust under strong inter‑look correlation, achieving performance close to the ideal independent‑look scenario and consistently outperforming methods that ignore such dependencies. These results highlight the importance of explicitly modeling inter‑look correlation and provide a practical framework for multi‑look holographic reconstruction under realistic acquisition conditions. Our code is available at: https://github.com/Computational‑Imaging‑RU/MLE‑Holography‑Markov.
Authors:Weitong Kong, Di Wen, Kunyu Peng, David Schneider, Zeyun Zhong, Alexander Jaus, Zdravko Marinov, Jiale Wei, Ruiping Liu, Junwei Zheng, Yufan Chen, Lei Qi, Rainer Stiefelhagen
Abstract:
Correcting errors in long‑video understanding is disproportionately costly: existing multimodal pipelines produce opaque, end‑to‑end outputs that expose no intermediate state for inspection, forcing annotators to revisit raw video and reconstruct temporal logic from scratch. The core bottleneck is not generation quality alone, but the absence of a supervisory interface through which human effort can be proportional to the scope of each error. We present IMPACT‑CYCLE, a supervisory multi‑agent system that reformulates long‑video understanding as iterative claim‑level maintenance of a shared semantic memory ‑‑ a structured, versioned state encoding typed claims, a claim dependency graph, and a provenance log. Role‑specialized agents operating under explicit authority contracts decompose verification into local object‑relation correctness, cross‑temporal consistency, and global semantic coherence, with corrections confined to structurally dependent claims. When automated evidence is insufficient, the system escalates to human arbitration as the supervisory authority with final override rights; dependency‑closure re‑verification then ensures correction cost remains proportional to error scope. Experiments on VidOR show substantially improved downstream reasoning (VQA: 0.71 to 0.79) and a 4.8x reduction in human arbitration cost, with workload significantly lower than manual annotation. Code will be released at https://github.com/MKong17/IMPACT_CYCLE.
Authors:Tianrong Chen, Jiatao Gu, David Berthelot, Joshua Susskind, Shuangfei Zhai
Abstract:
Normalizing Flows (NFs) are a classical family of likelihood‑based methods that have received revived attention. Recent efforts such as TARFlow have shown that
NFs are capable of achieving promising performance on image modeling tasks, making them viable alternatives to other methods such as diffusion models.
In this work, we further advance the state of Normalizing Flow generative models by introducing iterative TARFlow (iTARFlow). Unlike diffusion models, iTARFlow maintains a fully end‑to‑end, likelihood‑based objective during training. During sampling, it performs autoregressive generation followed by an iterative denoising procedure inspired by diffusion‑style methods. Through extensive experiments, we show that iTARFlow achieves competitive performance across ImageNet resolutions of 64, 128, and 256 pixels, demonstrating its potential as a strong generative model and advancing the frontier of Normalizing Flows. In addition, we analyze the characteristic artifacts produced by iTARFlow, offering insights that may shed light on future improvements. Code is available at https://github.com/apple/ml‑itarflow.
Authors:Xinxuan Lu, Charless Fowlkes, Alexander C. Berg
Abstract:
Current text‑to‑image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global scene understanding in text‑to‑image generation by learning parametric camera tokens. We fine‑tune image generation models for viewpoint‑conditioned text‑to‑image generation on a curated dataset that combines 3D‑rendered images for geometric supervision and photorealistic augmentations for appearance and background diversity. Qualitative and quantitative experiments demonstrate that our method achieves state‑of‑the‑art accuracy while preserving image quality and prompt fidelity. Unlike prior methods that overfit to object‑specific appearance correlations, our viewpoint tokens learn factorized geometric representations that transfer to unseen object categories. Our work shows that text‑vision latent spaces can be endowed with explicit 3D camera structure, offering a pathway toward geometrically‑aware prompts for text‑to‑image generation. Project page: https://randdl.github.io/viewtoken_control/
Authors:Tanuj Sur, Shashank Tripathi, Nikos Athanasiou, Ha Linh Nguyen, Kai Xu, Michael J. Black, Angela Yao
Abstract:
We introduce UniCon3R, a unified feed‑forward framework for online human‑scene 4D reconstruction from monocular video. Current feed‑forward human‑scene reconstruction methods suffer from artifacts, where bodies float above the ground or penetrate parts of the scene. A key reason is the lack of effective interaction modelling between the human and the environment. Our goal is to exploit contact between the human and the scene during inference to actively improve the human mesh reconstruction. To that end, we explicitly model interaction by inferring 4D contact from the human pose and scene geometry and use the contact as a corrective cue for generating the pose. This enables UniCon3R to jointly recover scene geometry and spatially aligned 4D humans within the scene. Experiments on standard human‑centric video benchmarks show that UniCon3R outperforms state‑of‑the‑art baselines on physical plausibility and global human motion estimation while preserving fast, feed‑forward inference speeds. The results validate our central claim: contact serves as a powerful internal prior, thus establishing a new paradigm for physically grounded joint human‑scene reconstruction. Project page is available at https://surtantheta.github.io/UniCon3R .
Authors:Yutian Chen, Shi Guo, Renbiao Jin, Tianshuo Yang, Xin Cai, Yawen Luo, Mingxin Yang, Mulin Yu, Linning Xu, Tianfan Xue
Abstract:
Sparse‑view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non‑generative reconstruction. Existing diffusion‑based approaches mitigates this issues by synthesizing novel views, but they often condition on only one or two capture frames, which restricts geometric consistency and limits scalability to large or diverse scenes. We propose AnyRecon, a scalable framework for reconstruction from arbitrary and unordered sparse inputs that preserves explicit geometric control while supporting flexible conditioning cardinality. To support long‑range conditioning, our method constructs a persistent global scene memory via a prepended capture view cache, and removes temporal compression to maintain frame‑level correspondence under large viewpoint changes. Beyond better generative model, we also find that the interplay between generation and reconstruction is crucial for large‑scale 3D scenes. Thus, we introduce a geometry‑aware conditioning strategy that couples generation and reconstruction through an explicit 3D geometric memory and geometry‑driven capture‑view retrieval. To ensure efficiency, we combine 4‑step diffusion distillation with context‑window sparse attention to reduce quadratic complexity. Extensive experiments demonstrate robust and scalable reconstruction across irregular inputs, large viewpoint gaps, and long trajectories.
Authors:Mario Tuci, Caner Korkmaz, Umut Şimşekli, Tolga Birdal
Abstract:
Training modern neural networks often relies on large learning rates, operating at the edge of stability, where the optimization dynamics exhibit oscillatory and chaotic behavior. Empirically, this regime often yields improved generalization performance, yet the underlying mechanism remains poorly understood. In this work, we represent stochastic optimizers as random dynamical systems, which often converge to a fractal attractor set (rather than a point) with a smaller intrinsic dimension. Building on this connection and inspired by Lyapunov dimension theory, we introduce a novel notion of dimension, coined the `sharpness dimension', and prove a generalization bound based on this dimension. Our results show that generalization in the chaotic regime depends on the complete Hessian spectrum and the structure of its partial determinants, highlighting a complexity that cannot be captured by the trace or spectral norm considered in prior work. Experiments across various MLPs and transformers validate our theory while also providing new insights into the recently observed phenomenon of grokking.
Authors:Jean Mercat, Sedrick Keh, Kushal Arora, Isabella Huang, Paarth Shah, Haruki Nishimura, Shun Iwase, Katherine Liu
Abstract:
We present VLA Foundry, an open‑source framework that unifies LLM, VLM, and VLA training in a single codebase. Most open‑source VLA efforts specialize on the action training stage, often stitching together incompatible pretraining pipelines. VLA Foundry instead provides a shared training stack with end‑to‑end control, from language pretraining to action‑expert fine‑tuning. VLA Foundry supports both from‑scratch training and pretrained backbones from Hugging Face. To demonstrate the utility of our framework, we train and release two types of models: the first trained fully from scratch through our LLM‑‑>VLM‑‑>VLA pipeline and the second built on the pretrained Qwen3‑VL backbone. We evaluate closed‑loop policy performance of both models on LBM Eval, an open‑data, open‑source simulator. We also contribute usability improvements to the simulator and the STEP analysis tools for easier public use. In the nominal evaluation setting, our fully‑open from‑scratch model is on par with our prior closed‑source work and substituting in the Qwen3‑VL backbone leads to a strong multi‑task table top manipulation policy outperforming our baseline by a wide margin. The VLA Foundry codebase is available at https://github.com/TRI‑ML/vla_foundry and all multi‑task model weights are released on https://huggingface.co/collections/TRI‑ML/vla‑foundry. Additional qualitative videos are available on the project website https://tri‑ml.github.io/vla_foundry.
Authors:Zhengwentai Sun, Keru Zheng, Chenghong Li, Hongjie Liao, Xihe Yang, Heyuan Li, Yihao Zhi, Shuliang Ning, Shuguang Cui, Xiaoguang Han
Abstract:
Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi‑view data. Existing methods often address these factors separately, resulting in limited controllability or reduced visual quality. We revisit this problem from an image‑first perspective, where high‑quality human appearance is learned via image generation and used as a prior for video synthesis, decoupling appearance modeling from temporal consistency. We propose a pose‑ and viewpoint‑controllable pipeline that combines a pretrained image backbone with SMPL‑X‑based motion guidance, together with a training‑free temporal refinement stage based on a pretrained video diffusion model. Our method produces high‑quality, temporally consistent videos under diverse poses and viewpoints. We also release a canonical human dataset and an auxiliary model for compositional human image synthesis. Code and data are publicly available at https://github.com/Taited/ReImagine.
Authors:Umut Kocasari, Simon Giebenhain, Richard Shaw, Matthias Nießner
Abstract:
Accurate reconstruction and tracking of dynamic human faces from image sequences is challenging because non‑rigid deformations, expression changes, and viewpoint variations occur simultaneously, creating significant ambiguity in geometry and correspondence estimation. We present a unified method for high‑fidelity 4D facial reconstruction based on canonical facial point prediction, a representation that assigns each pixel a normalized facial coordinate in a shared canonical space. This formulation transforms dense tracking and dynamic reconstruction into a canonical reconstruction problem, enabling temporally consistent geometry and reliable correspondences within a single feed‑forward model. By jointly predicting depth and canonical coordinates, our method enables accurate depth estimation, temporally stable reconstruction, dense 3D geometry, and robust facial point tracking within a single architecture. We implement this formulation using a transformer‑based model that jointly predicts depth and canonical facial coordinates, trained using multi‑view geometry data that non‑rigidly warps into the canonical space. Extensive experiments on image and video benchmarks demonstrate state‑of‑the‑art performance across reconstruction and tracking tasks, achieving approximately 3× lower correspondence error and faster inference than prior dynamic reconstruction methods, while improving depth accuracy by 16%. These results highlight canonical facial point prediction as an effective foundation for unified feed‑forward 4D facial reconstruction.
Authors:Jing Jin, Hao Liu, Yan Bai, Yihang Lou, Zhenke Wang, Tianrun Yuan, Juntong Chen, Yongkang Zhu, Fanhu Zeng, Xuanyu Zhu, Tao Feng, Yige Xu
Abstract:
Multimodal large language models (MLLMs) have shown promising reasoning abilities, yet evaluating their performance in specialized domains remains challenging. STEM reasoning is a particularly valuable testbed because it provides highly verifiable feedback, but existing benchmarks often permit unimodal shortcuts due to modality redundancy and focus mainly on final‑answer accuracy, overlooking the reasoning process itself. To address this challenge, we introduce StepSTEM: a graduate‑level benchmark of 283 problems across mathematics, physics, chemistry, biology, and engineering for fine‑grained evaluation of cross‑modal reasoning in MLLMs. StepSTEM is constructed through a rigorous curation pipeline that enforces strict complementarity between textual and visual inputs. We further propose a general step‑level evaluation framework for both text‑only chain‑of‑thought and interleaved image‑text reasoning, using dynamic programming to align predicted reasoning steps with multiple reference solutions. Experiments across a wide range of models show that current MLLMs still rely heavily on textual reasoning, with even Gemini 3.1 Pro and Claude Opus 4.6 achieving only 38.29% accuracy. These results highlight substantial headroom for genuine cross‑modal STEM reasoning and position StepSTEM as a benchmark for fine‑grained evaluation of multimodal reasoning. Source code is available at https://github.com/lll‑hhh/STEPSTEM.
Authors:Zihao Fan, Xin Lu, Jie Xiao, Dong Li, Jie Huang, Xueyang Fu
Abstract:
In image restoration, single‑step discriminative mappings often lack fine details via expectation learning, whereas generative paradigms suffer from inefficient multi‑step sampling and noise‑residual coupling. To address this dilemma, we propose IR‑Flow, a novel image restoration method based on Rectified Flow that serves as a unified framework bridging the gap between discriminative and generative paradigms. Specifically, we first construct multilevel data distribution flows, which expand the ability of models to learn from and adapt to various levels of degradation. Subsequently, cumulative velocity fields are proposed to learn transport trajectories across varying degradation levels, guiding intermediate states toward the clean target, while a multi‑step consistency constraint is presented to enforce trajectory coherence and boost few‑step restoration performance. We show that directly establishing a linear transport flow between degraded and clean image domains not only enables fast inference but also improves adaptability to out‑of‑distribution degradations. Extensive evaluations on deraining, denoising and raindrop removal tasks demonstrate that IR‑Flow achieves competitive quantitative results with only a few sampling steps, offering an efficient and flexible framework that maintains an excellent distortion‑perception balance. Our code is available at https://github.com/fanzh03/IR‑Flow.
Authors:Liyang Li, Wen Wang, Canyu Zhao, Tianjian Feng, Zhiyue Zhao, Hao Chen, Chunhua Shen
Abstract:
Recent advances in Diffusion Transformers (DiTs) have enabled high‑quality joint audio‑video generation, producing videos with synchronized audio within a single model. However, existing controllable generation frameworks are typically restricted to video‑only control. This restricts comprehensive controllability and often leads to suboptimal cross‑modal alignment. To bridge this gap, we present MMControl, which enables users to perform Multi‑Modal Control in joint audio‑video generation. MMControl introduces a dual‑stream conditional injection mechanism. It incorporates both visual and acoustic control signals, including reference images, reference audio, depth maps, and pose sequences, into a joint generation process. These conditions are injected through bypass branches into a joint audio‑video Diffusion Transformer, enabling the model to simultaneously generate identity‑consistent video and timbre‑consistent audio under structural constraints. Furthermore, we introduce modality‑specific guidance scaling, which allows users to independently and dynamically adjust the influence strength of each visual and acoustic condition at inference time. Extensive experiments demonstrate that MMControl achieves fine‑grained, composable control over character identity, voice timbre, body pose, and scene layout in joint audio‑video generation.
Authors:Yi Zhong, Buqiang Xu, Yijun Wang, Zifei Shan, Shuofei Qiao, Guozhou Zheng, Ningyu Zhang
Abstract:
At present, executable visual workflows have emerged as a mainstream paradigm in real‑world industrial deployments, offering strong reliability and controllability. However, in current practice, such workflows are almost entirely constructed through manual engineering: developers must carefully design workflows, write prompts for each step, and repeatedly revise the logic as requirements evolve ‑‑ making development costly, time‑consuming, and error‑prone. To study whether large language models can automate this multi‑round interaction process, we introduce Chat2Workflow, a benchmark for generating executable visual workflows directly from natural language, and propose a robust agentic baseline to improve performance. The benchmark is built from a large collection of real‑world business workflows, with each instance designed so that the generated workflow can be transformed and directly deployed to practical workflow platforms such as Dify and Coze. Experimental results show that while state‑of‑the‑art language models can often capture high‑level intent, they struggle to generate correct, stable, and executable workflows, especially given complex and evolving requirements. Although our agentic baseline yields up to 6.05% resolve rate gains, the remaining real‑world gap positions Chat2Workflow as a foundation for advancing industrial‑grade automation. Code is available at https://github.com/zjunlp/Chat2Workflow.
Authors:Xiangyang Luo, Xiaozhe Xin, Tao Feng, Xu Guo, Meiguang Jin, Junfeng Ma
Abstract:
Synthesizing human‑‑object interaction (HOI) videos has broad practical value in e‑commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently fail on (i) the structural stability of sensitive regions such as hands and faces and (ii) physically plausible contact (e.g., avoiding hand‑‑object interpenetration). We present CoInteract, an end‑to‑end framework for HOI video synthesis conditioned on a person reference image, a product reference image, text prompts, and speech audio. CoInteract introduces two complementary designs embedded into a Diffusion Transformer (DiT) backbone. First, we propose a Human‑Aware Mixture‑of‑Experts (MoE) that routes tokens to lightweight, region‑specialized experts via spatially supervised routing, improving fine‑grained structural fidelity with minimal parameter overhead. Second, we propose Spatially‑Structured Co‑Generation, a dual‑stream training paradigm that jointly models an RGB appearance stream and an auxiliary HOI structure stream to inject interaction geometry priors. During training, the HOI stream attends to RGB tokens and its supervision regularizes shared backbone weights; at inference, the HOI branch is removed for zero‑overhead RGB generation. Experimental results demonstrate that CoInteract significantly outperforms existing methods in structural stability, logical consistency, and interaction realism.
Authors:Pradyumna YM, Yuxuan Xue, Yue Chen, Nikita Kister, István Sárándi, Gerard Pons-Moll
Abstract:
Reconstructing physically plausible 3D human‑scene interactions (HSI) from a single image currently presents a trade‑off: optimization based methods offer accurate contact but are slow (~20s), while feed‑forward approaches are fast yet lack explicit interaction reasoning, producing floating and interpenetration artifacts.
Our key insight is that geometry‑based human‑‑scene fitting can be amortized into fast feed‑forward inference. We present GRAFT (Geometric Refinement And Fitting Transformer), a learned HSI prior that predicts Interaction Gradients: corrective parameter updates that iteratively refine human meshes by reasoning about their 3D relationship to the surrounding scene.
GRAFT encodes the interaction state into compact body‑anchored tokens, each grounded in the scene geometry via Geometric Probes that capture spatial relationships with nearby surfaces.
A lightweight transformer recurrently updates human meshes and re‑probes the scene, ensuring the final pose aligns with both learned priors and observed geometry.
GRAFT operates either as an end‑to‑end reconstructor using image features, or with geometry alone as a transferable plug‑and‑play HSI prior that improves feed‑forward methods without retraining.
Experiments show GRAFT improves interaction quality by up to 113% over state‑of‑the‑art feed‑forward methods and matches optimization‑based interaction quality at ~50× lower runtime, while generalizing seamlessly to in‑the‑wild multi‑person scenes and being preferred in 64.8% of three‑way user study. Project page: https://pradyumnaym.github.io/graft .
Authors:Ying Zeng, Miaosen Luo, Guangyuan Li, Yang Yang, Ruiyang Fan, Linxiao Shi, Qirui Yang, Jian Zhang, Chengcheng Liu, Siming Zheng, Jinwei Chen, Bo Li, Peng-Tao Jiang
Abstract:
Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and camera parameters. However, this paradigm relies on explicit human instruction of aesthetic intent, which is often ambiguous, incomplete, or inaccessible to non‑expert users. In this work, we propose SmartPhotoCrafter, an automatic photographic image editing method which formulates image editing as a tightly coupled reasoning‑to‑generation process. The proposed model first performs image quality comprehension and identifies deficiencies by the Image Critic module, and then the Photographic Artist module realizes targeted edits to enhance image appeal, eliminating the need for explicit human instructions. A multi‑stage training pipeline is adopted: (i) Foundation pretraining to establish basic aesthetic understanding and editing capabilities, (ii) Adaptation with reasoning‑guided multi‑edit supervision to incorporate rich semantic guidance, and (iii) Coordinated reasoning‑to generation reinforcement learning to jointly optimize reasoning and generation. During training, SmartPhotoCrafter emphasizes photo‑realistic image generation, while supporting both image restoration and retouching tasks with consistent adherence to color‑ and tone‑related semantics. We also construct a stage‑specific dataset, which progressively builds reasoning and controllable generation, effective cross‑module collaboration, and ultimately high‑quality photographic enhancement. Experiments demonstrate that SmartPhotoCrafter outperforms existing generative models on the task of automatic photographic enhancement, achieving photo‑realistic results while exhibiting higher tonal sensitivity to retouching instructions. Project page: https://github.com/vivoCameraResearch/SmartPhotoCrafter.
Authors:Yanshuo Wang, Yuan Xu, Xuesong Li, Jie Hong, Yizhou Wang, Chang Wen Chen, Wentao Zhu
Abstract:
Egocentric assistants often rely on first‑person view data to capture user behavior and context for personalized services. Since different users exhibit distinct habits, preferences, and routines, such personalization is essential for truly effective assistance. However, effectively integrating long‑term user data for personalization remains a key challenge. To address this, we introduce EgoSelf, a system that includes a graph‑based interaction memory constructed from past observations and a dedicated learning task for personalization. The memory captures temporal and semantic relationships among interaction events and entities, from which user‑specific profiles are derived. The personalized learning task is formulated as a prediction problem where the model predicts possible future interactions from individual user's historical behavior recorded in the graph. Extensive experiments demonstrate the effectiveness of EgoSelf as a personalized egocentric assistant. Code is available at https://abie‑e.github.io/EgoSelf/.
Authors:Davide Allegro, Shiyao Li, Stefano Ghidoni, Vincent Lepetit
Abstract:
Current 3D mapping pipelines generally assume static environments, which limits their ability to accurately capture and reconstruct moving objects. To address this limitation, we introduce the novel task of active mapping of moving objects, in which a mapping agent must plan its trajectory while compensating for the object's motion. Our approach, Paparazzo, provides a learning‑free solution that robustly predicts the target's trajectory and identifies the most informative viewpoints from which to observe it, to plan its own path. We also contribute a comprehensive benchmark designed for this new task. Through extensive experiments, we show that Paparazzo significantly improves 3D reconstruction completeness and accuracy compared to several strong baselines, marking an important step toward dynamic scene understanding. Project page: https://davidea97.github.io/paparazzo‑page/
Authors:Hongyu Zhang, Yufan Deng, Zilin Pan, Peng-Tao Jiang, Bo Li, Qibin Hou, Zhiyang Dou, Zhen Dong, Daquan Zhou
Abstract:
Generating high‑quality videos from complex temporal descriptions that contain multiple sequential actions is a key unsolved problem. Existing methods are constrained by an inherent trade‑off: using multiple short prompts fed sequentially into the model improves action fidelity but compromises temporal consistency, while a single complex prompt preserves consistency at the cost of prompt‑following capability. We attribute this problem to two primary causes: 1) temporal misalignment between video content and the prompt, and 2) conflicting attention coupling between motion‑related visual objects and their associated text conditions. To address these challenges, we propose a novel, training‑free attention mechanism, Temporal‑wise Separable Attention (TS‑Attn), which dynamically rearranges attention distribution to ensure temporal awareness and global coherence in multi‑event scenarios. TS‑Attn can be seamlessly integrated into various pre‑trained text‑to‑video models, boosting StoryEval‑Bench scores by 33.5% and 16.4% on Wan2.1‑T2V‑14B and Wan2.2‑T2V‑A14B with only a 2% increase in inference time. It also supports plug‑and‑play usage across models for multi‑event image‑to‑video generation. The source code and project page are available at https://github.com/Hong‑yu‑Zhang/TS‑Attn.
Authors:Xiaoqi Zhuang, Jefersson A. Dos Santos, Jungong Han
Abstract:
Satellite image composition plays a critical role in remote sensing applications such as data augmentation, disaste simulation, and urban planning. We propose HarmoniDiff‑RS, a training‑free diffusion‑based framework for harmonizing composite satellite images under diverse domain conditions. Our method aligns the source and target domains through a Latent Mean Shift operation that transfers radiometric characteristics between them. To balance harmonization and content preservation, we introduce a Timestep‑wise Latent Fusion strategy by leveraging early inverted latents for high harmonization and late latents for semantic consistency to generate a set of composite candidates. A lightweight harmony classifier is trained to further automatically select the most coherent result among them. We also construct RSIC‑H, a benchmark dataset for satellite image harmonization derived from fMoW, providing 500 paired composition samples. Experiments demonstrate that our method effectively performs satellite image composition, showing strong potential for scalable remote‑sensing synthesis and simulation tasks. Code is available at: https://github.com/XiaoqiZhuang/HarmoniDiff‑RS.
Authors:Philipp Weigand, Niels Nawrot, Nikolas Ebert, Carsten Hopf, Oliver Wasenmüller
Abstract:
Peak picking is a fundamental preprocessing step in Mass Spectrometry Imaging (MSI), where each sample is represented by hundreds to thousands of ion images. Existing approaches require careful dataset‑specific hyperparameter tuning, and often fail to generalize across acquisition protocols. We introduce IonMorphNet, a spatial‑structure‑aware representation model for ion images that enables fully data‑driven peak picking without any task‑specific supervision. We curate 53 publicly available MSI datasets and define six structural classes capturing representative spatial patterns in ion images to train standard image backbones for structural pattern classification. Once trained, IonMorphNet can assess ion images and perform peak picking without additional hyperparameter tuning. Using a ConvNeXt V2‑Tiny backbone, our approach improves peak picking performance by +7 % mSCF1 compared to state‑of‑the‑art methods across multiple datasets. Beyond peak picking, we demonstrate that spatially informed channel reduction enables a 3D CNN for patch‑based tumor classification in MSI. This approach matches or exceeds pixel‑wise spectral classifiers by up to +7.3 % Balanced Accuracy on three tumor classification tasks, indicating meaningful ion image selection. The source code and model weights are available at https://github.com/CeMOS‑IS/IonMorphNet.
Authors:Ghadah Alosaimi, Hanadi Alhamdan, Wenke E, Stamos Katsigiannis, Amir Atapour-Abarghouei, Toby P. Breckon
Abstract:
Predicting driver intention from neurophysiological signals offers a promising pathway for enhancing proactive safety in advanced driver assistance systems, yet remains challenging in real‑world driving due to EEG signal non‑stationarity and the complexity of cognitive‑motor preparation. This study proposes and evaluates an EEG‑based driver intention prediction framework using a synchronised multi‑sensor platform integrated into a real electric vehicle. A real‑world on‑road dataset was collected across 32 driving sessions, and 12 deep learning architectures were evaluated under consistent experimental conditions. Among the evaluated architectures, TSCeption achieved the highest average accuracy (0.907) and Macro‑F1 score (0.901). The proposed framework demonstrates strong temporal stability, maintaining robust decoding performance up to 1000 ms before manoeuvre execution with minimal degradation. Furthermore, additional analyses reveal that minimal EEG preprocessing outperforms artefact‑handling pipelines, and prediction performance peaks within a 400‑600 ms interval, corresponding to a critical neural preparatory phase preceding driving manoeuvres. Overall, these findings support the feasibility of early and stable EEG‑based driver intention decoding under real‑world on‑road conditions. Code: https://github.com/galosaimi/Mind2Drive.
Authors:Samyak Sanghvi, Piyush Miglani, Sarvesh Shashikumar, Kaustubh R Borgavi, Veenu Singla, Chetan Arora
Abstract:
Vision Transformers (\textttViT) have become the architecture of choice for many computer vision tasks, yet their performance in computer‑aided diagnostics remains limited. Focusing on breast cancer detection from mammograms, we identify two main causes for this shortfall. First, medical images are high‑resolution with small abnormalities, leading to an excessive number of tokens and making it difficult for the softmax‑based attention to localize and attend to relevant regions. Second, medical image classification is inherently fine‑grained, with low inter‑class and high intra‑class variability, where standard cross‑entropy training is insufficient. To overcome these challenges, we propose a framework with three key components: (1) Region of interest (\textttRoI) based token reduction using an object detection model to guide attention; (2) contrastive learning between selected \textttRoI to enhance fine‑grained discrimination through hard‑negative based training; and (3) a \textttDINOv2 pretrained \textttViT that captures localization‑aware, fine‑grained features instead of global \textttCLIP representations. Experiments on public mammography datasets demonstrate that our method achieves superior performance over existing baselines, establishing its effectiveness and potential clinical utility for large‑scale breast cancer screening. Our code is available for reproducibility here: https://aih‑iitd.github.io/publications/attend‑what‑matters
Authors:Xunpei Sun, Zuoxun Hou, Yi Chang, Gang Chen, Wei-Shi Zheng
Abstract:
Monocular scene flow estimation aims to recover dense 3D motion from image sequences, yet most existing methods are limited to two‑frame inputs, restricting temporal modeling and robustness to occlusions. We propose RAFT‑MSF++, a self‑supervised multi‑frame framework that recurrently fuses temporal features to jointly estimate depth and scene flow. Central to our approach is the Geometry‑Motion Feature (GMF), which compactly encodes coupled motion and geometry cues and is iteratively updated for effective temporal reasoning. To ensure the robustness of this temporal fusion against occlusions, we incorporate relative positional attention to inject spatial priors and an occlusion regularization module to propagate reliable motion from visible regions. These components enable the GMF to effectively propagate information even in ambiguous areas. Extensive experiments show that RAFT‑MSF++ achieves 24.14% SF‑all on the KITTI Scene Flow benchmark, with a 30.99% improvement over the baseline and better robustness in occluded regions. The code is available at https://github.com/sunzunyi/RAFT‑MSF‑PlusPlus.
Authors:Qi Zhang, Jixuan Chen, Kaiyi Zhang, Xinquan Yu, Antoni B. Chan, Hui Huang
Abstract:
Multi‑view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs‑based multi‑view crowd tracking architectures, and most of them are evaluated and compared on relatively small datasets, such as Wildtrack and MultiviewX. Since these two datasets are collected in small scenes and only contain tens of frames in the evaluation stage, it is difficult for the current methods to be applied to real‑world applications where scene size and occlusion are more complicated. In this paper, we propose a Transformer‑based multi‑view crowd tracking model, MVTrackTrans, which adopts interactions between camera views and the ground plane for enhanced multi‑view tracking performance. Besides, for better evaluation, we collect and label two large real‑world multi‑view tracking datasets, MVCrowdTrack and CityTrack, which contain a much larger scene size over a longer time period. Compared with existing methods on the two large and new datasets, the proposed MVTrackTrans model achieves better performance, demonstrating the advantages of the model design in dealing with large scenes. We believe the proposed datasets and model will push the frontiers of the task to more practical scenarios, and the datasets and code are available at: https://github.com/zqyq/MVTrackTrans.
Authors:Bo Li, Jiahao Kang, Yubo Ma, Feng-Lin Liu, Bin Liu, Fang-Lue Zhang, Lin Gao
Abstract:
3D Gaussian representations have emerged as a powerful paradigm for digital head modeling, achieving photorealistic quality with real‑time rendering. However, intuitive and interactive creation or editing of 3D Gaussian head models remains challenging. Although 2D sketches provide an ideal interaction modality for fast, intuitive conceptual design, they are sparse, depth‑ambiguous, and lack high‑frequency appearance cues, making it difficult to infer dense, geometrically consistent 3D Gaussian structures from strokes ‑ especially under real‑time constraints. To address these challenges, we propose SketchFaceGS, the first sketch‑driven framework for real‑time generation and editing of photorealistic 3D Gaussian head models from 2D sketches. Our method uses a feed‑forward, coarse‑to‑fine architecture. A Transformer‑based UV feature‑prediction module first reconstructs a coarse but geometrically consistent UV feature map from the input sketch, and then a 3D UV feature enhancement module refines it with high‑frequency, photorealistic detail to produce a high‑fidelity 3D head. For editing, we introduce a UV Mask Fusion technique combined with a layer‑by‑layer feature‑fusion strategy, enabling precise, real‑time, free‑viewpoint modifications. Extensive experiments show that SketchFaceGS outperforms existing methods in both generation fidelity and editing flexibility, producing high‑quality, editable 3D heads from sketches in a single forward pass.
Authors:Mika Feng, Pierre Gallin-Martel, Koichi Ito, Takafumi Aoki
Abstract:
Face Anti‑Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision‑Language Models (VLMs) for semantic supervision, these multimodal approaches often demand prohibitive computational resources and exhibit high inference latency. Furthermore, their efficacy is inherently limited by the quality of the underlying visual features. This paper revisits the potential of vision‑only foundation models to establish a highly efficient and robust baseline for FAS. We conduct a systematic benchmarking of 15 pre‑trained models, such as supervised CNNs, supervised ViTs, and self‑supervised ViTs, under severe cross‑domain scenarios including the MICO and Limited Source Domains (LSD) protocols. Our comprehensive analysis reveals that self‑supervised vision models, particularly DINOv2 with Registers, significantly suppress attention artifacts and capture critical, fine‑grained spoofing cues. Combined with Face Anti‑Spoofing Data Augmentation (FAS‑Aug), Patch‑wise Data Augmentation (PDA) and Attention‑weighted Patch Loss (APL), our proposed vision‑only baseline achieves state‑of‑the‑art performance in the MICO protocol. This baseline outperforms existing methods under the data‑constrained LSD protocol while maintaining superior computational efficiency. This work provides a definitive vision‑only baseline for FAS, demonstrating that optimized self‑supervised vision transformers can serve as a backbone for both vision‑only and future multimodal FAS systems. The project page is available at: https://gsisaoki.github.io/FAS‑VFMbenchmark‑CVPRW2026/ .
Authors:Johannes Schusterbauer, Ming Gui, Yusong Li, Pingchuan Ma, Felix Krause, Björn Ommer
Abstract:
Diffusion‑ and flow‑based models usually allocate compute uniformly across space, updating all patches with the same timestep and number of function evaluations. While convenient, this ignores the heterogeneity of natural images: some regions are easy to denoise, whereas others benefit from more refinement or additional context. Motivated by this, we explore patch‑level noise scales for image synthesis. We find that naively varying timesteps across image tokens performs poorly, as it exposes the model to overly informative training states that do not occur at inference. We therefore introduce a timestep sampler that explicitly controls the maximum patch‑level information available during training, and show that moving from global to patch‑level timesteps already improves image generation over standard baselines. By further augmenting the model with a lightweight per‑patch difficulty head, we enable adaptive samplers that allocate compute dynamically where it is most needed. Combined with noise levels varying over both space and diffusion time, this yields Patch Forcing (PF), a framework that advances easier regions earlier so they can provide context for harder ones. PF achieves superior results on class‑conditional ImageNet, remains orthogonal to representation alignment and guidance methods, and scales to text‑to‑image synthesis. Our results suggest that patch‑level denoising schedules provide a promising foundation for adaptive image generation.
Authors:Jinglin Xu, Yi Li, Chuxiong Sun, Xiao Xu, Jiangmeng Li, Fanjiang Xu
Abstract:
Multi‑modal test‑time adaptation (TTA) enhances the resilience of benchmark multi‑modal models against distribution shifts by leveraging the unlabeled target data during inference. Despite the documented success, the advancement of multi‑modal TTA methodologies has been impeded by a persistent limitation, i.e., the lack of explicit modeling of category‑conditional distributions, which is crucial for yielding accurate predictions and reliable decision boundaries. Canonical Gaussian discriminant analysis (GDA) provides a vanilla modeling of category‑conditional distributions and achieves moderate advancement in uni‑modal contexts. However, in multi‑modal TTA scenario, the inherent modality distribution asymmetry undermines the effectiveness of modeling the category‑conditional distribution via the canonical GDA. To this end, we introduce a tailored probabilistic Gaussian model for multi‑modal TTA to explicitly model the category‑conditional distributions, and further propose an adaptive contrastive asymmetry rectification technique to counteract the adverse effects arising from modality asymmetry, thereby deriving calibrated predictions and reliable decision boundaries. Extensive experiments across diverse benchmarks demonstrate that our method achieves state‑of‑the‑art performance under a wide range of distribution shifts. The code is available at https://github.com/XuJinglinn/AdaPGC.
Authors:Rongjia Zheng, Shangwei Huang, Lei Zhu, Wei-Shi Zheng, Qing Zhang
Abstract:
We present a generative method for texture filtering, which exhibits surprisingly good performance and generalizability. Our core idea is to empower texture filtering by taking full advantage of the strong learned image prior of pre‑trained generative models. To this end, we propose to fine‑tune a pre‑trained generative model via a two‑stage strategy. Specifically, we first conduct supervised fine‑tuning on a very small set of paired images, and then perform reinforcement fine‑tuning on a large‑scale unlabeled dataset under the guidance of a reward function that quantifies the quality of texture removal and structure preservation. Extensive experiments show that our method clearly outperforms previous methods, and is effective to deal with previously challenging cases. Our code is available at https://github.com/OnlyZZZZ/Generative_Texture_Filtering.
Authors:Jiagao Hu, Daiguo Zhou, Danzhen Fu, Fuhao Li, Zepeng Wang, Fei Wang, Wenhua Liao, Jiayi Xie, Haiyang Sun
Abstract:
Perception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real‑world video data in adverse weather. Existing weather generation approaches struggle to balance visual quality and annotation reusability. We present AutoAWG, a controllable Adverse Weather video Generation framework for Autonomous driving. Our method employs a semantics‑guided adaptive fusion of multiple controls to balance strong weather stylization with high‑fidelity preservation of safety‑critical targets; leverages a vanishing point‑anchored temporal synthesis strategy to construct training sequences from static images, thereby reducing reliance on synthetic data; and adopts masked training to enhance long‑horizon generation stability. On the nuScenes validation set, AutoAWG significantly outperforms prior state‑of‑the‑art methods: without first‑frame conditioning, FID and FVD are relatively reduced by 50.0% and 16.1%; with first‑frame conditioning, they are further reduced by 8.7% and 7.2%, respectively. Extensive qualitative and quantitative results demonstrate advantages in style fidelity, temporal consistency, and semantic‑‑structural integrity, underscoring the practical value of AutoAWG for improving downstream perception in autonomous driving. Our code is available at: https://github.com/higherhu/AutoAWG
Authors:Abdul Mueez, Shruti Vyas
Abstract:
Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervised segmentation. To bridge foundational computer vision with practical metallurgical evaluation, we propose an automated pipeline for dense instance segmentation and grain size estimation that adapts Cellpose‑SAM to microstructures and integrates its topology‑aware gradient tracking with an ASTM E112 Jeffries planimetric module. We systematically benchmark this pipeline against a classical convolutional network (U‑Net), an adaptive‑prompting vision foundation model (MatSAM) and a contemporary vision‑language model (Qwen2.5‑VL‑7B). Our evaluations reveal that while the out‑of‑the‑box vision‑language model struggles with the localized spatial reasoning required for dense microscopic counting and MatSAM suffers from over‑segmentation despite its domain‑specific prompt generation, our adapted pipeline successfully maintains topological separation. Furthermore, experiments across progressively reduced training splits demonstrate exceptional few‑shot scalability; utilizing only two training samples, the proposed system predicts the ASTM grain size number (G) with a mean absolute percentage error (MAPE) as low as 1.50%, while robustness testing across varying target grain counts empirically validates the ASTM 50‑grain sampling minimum. These results highlight the efficacy of application‑level foundation model integration for highly accurate, automated materials characterization. Our project repository is available at https://github.com/mueez‑overflow/ASTM‑Grain‑Size‑Estimator.
Authors:Mohammed Q. Alkhatib
Abstract:
Hyperspectral image (HSI) classification remains challenging due to high spectral dimensionality, redundancy, and limited labeled data. Although convolutional neural networks (CNNs) and Vision Transformers (ViTs) achieve strong performance by exploiting spectral‑spatial information and long‑range dependencies, they often incur high computational cost and large model size, limiting practical use. To address these limitations, a unified hybrid framework, termed ConvVitMamba, is proposed for efficient HSI classification. The architecture integrates three components: a multiscale convolutional feature extractor to capture local spectral, spatial, and joint patterns; a Vision Transformer based tokenization and encoding stage to model global contextual relationships; and a lightweight Mamba inspired gated sequence mixing module for efficient content‑aware refinement without quadratic self‑attention. Principal Component Analysis (PCA) is used as preprocessing to reduce redundancy and improve efficiency. Experiments on four benchmark datasets, including Houston and three UAV borne QUH datasets (Pingan, Qingyun, and Tangdaowan), demonstrate that ConvVitMamba consistently outperforms CNN, Transformer, and Mamba based methods while maintaining a favorable balance between accuracy, model size, and inference efficiency. Ablation studies confirm the complementary contributions of all components. The results indicate that the proposed framework provides an effective and efficient solution for HSI classification in diverse scenarios. The source code is publicly available at https://github.com/mqalkhatib/ConvVitMamba
Authors:Mohammed Q. Alkhatib
Abstract:
This paper presents DDF2Pol, a lightweight dual‑domain convolutional neural network for PolSAR image classification. The proposed architecture integrates two parallel feature extraction streams, one real‑valued and one complex‑valued, designed to capture complementary spatial and polarimetric information from PolSAR data. To further refine the extracted features, a depth‑wise convolution layer is employed for spatial enhancement, followed by a coordinate attention mechanism to focus on the most informative regions. Experimental evaluations conducted on two benchmark datasets, Flevoland and San Francisco, demonstrate that DDF2Pol achieves superior classification performance while maintaining low model complexity. Specifically, it attains an Overall Accuracy (OA) of 98.16% on the Flevoland dataset and 96.12% on the San Francisco dataset, outperforming several state‑of‑the‑art real‑ and complex‑valued models. With only 91,371 parameters, DDF2Pol offers a practical and efficient solution for accurate PolSAR image analysis, even when training data is limited. The source code is publicly available at https://github.com/mqalkhatib/DDF2Pol
Authors:Abrar Majeedi, Zhiyuan Ruan, Ziyi Zhao, Hongcheng Wang, Jianglin Lu, Yin Li
Abstract:
Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain fragile under common degradations, such as fog, blur, or low‑light conditions. Infrared (IR) imaging, a well‑established complement to RGB, offers inherent robustness in these conditions, but its integration into MLLMs remains underexplored. To bridge this gap, we propose DUALVISION, a lightweight fusion module that efficiently incorporates IR‑RGB information into MLLMs via patch‑level localized cross‑attention. To support training and evaluation and to facilitate future research, we also introduce DV‑204K, a dataset of ~25K publicly available aligned IR‑RGB image pairs with 204K modality‑specific QA annotations, and DV‑500, a benchmark of 500 IR‑RGB image pairs with 500 QA pairs designed for evaluating cross‑modal reasoning. Leveraging these datasets, we benchmark both open‑ and closed‑source MLLMs and demonstrate that DUALVISION delivers strong empirical performance under a wide range of visual degradations. Our code and dataset are available at https://abrarmajeedi.github.io/dualvision.
Authors:Yichen Xie, Depu Meng, Chensheng Peng, Yihan Hu, Quentin Herau, Masayoshi Tomizuka, Wei Zhan
Abstract:
Relative position embedding has become a standard mechanism for encoding positional information in Transformers. However, existing formulations are typically limited to a fixed geometric space, namely 1D sequences or regular 2D/3D grids, which restricts their applicability to many computer vision tasks that require geometric reasoning across camera views or between 2D and 3D spaces. To address this limitation, we propose URoPE, a universal extension of Rotary Position Embedding (RoPE) to cross‑view or cross‑dimensional geometric spaces. For each key/value image patch, URoPE samples 3D points along the corresponding camera ray at predefined depth anchors and projects them into the query image plane. Standard 2D RoPE can then be applied using the projected pixel coordinates. URoPE is a parameter‑free and intrinsics‑aware relative position embedding that is invariant to the choice of global coordinate systems, while remaining fully compatible with existing RoPE‑optimized attention kernels. We evaluate URoPE as a plug‑in positional encoding for transformer architectures across a diverse set of tasks, including novel view synthesis, 3D object detection, object tracking, and depth estimation, covering 2D‑2D, 2D‑3D, and temporal scenarios. Experiments show that URoPE consistently improves the performance of transformer‑based models across all tasks, demonstrating its effectiveness and generality for geometric reasoning. Our project website is: https://urope‑pe.github.io/.
Authors:Ruijun Zhang, Hang Su, Kostas Daniilidis, Ziyun Wang
Abstract:
Event cameras have recently shown promising capabilities in instantaneous motion estimation due to their robustness to low light and fast motions. However, computing wide‑baseline correspondence between two arbitrary views remains a significant challenge, since event appearance changes substantially with motion, and learning‑based approaches are constrained by both scalability and limited wide‑baseline supervision. We therefore introduce the first event matching model that achieves cross‑dataset wide‑baseline correspondence in a zero‑shot manner: a single model trained once is deployed on unseen datasets without any target‑domain fine‑tuning or adaptation. To enable this capability, we introduce a motion‑robust and computationally efficient attention backbone that learns multi‑timescale features from event streams, augmented with sparsity‑aware event token selection, making large‑scale training on diverse wide‑baseline supervision computationally feasible. To provide the supervision needed for wide‑baseline generalization, we develop a robust event motion synthesis framework to generate large‑scale event‑matching datasets with augmented viewpoints, modalities, and motions. Extensive experiments across multiple benchmarks show that our framework achieves a 37.7% improvement over the previous best event feature matching methods. Code and data are available at: https://github.com/spikelab‑jhu/Match‑Any‑Events.
Authors:Jay Jung, Ahmad Arrabi, Jax Luo, Scott Raymond, Safwan Wshah
Abstract:
Purpose: Automated C‑arm positioning ensures timely treatment in patients requiring emergent interventions. When a conventional Deep Learning (DL) approach for C‑arm control fails, clinicians must revert to manual operation, resulting in additional delays. Consequently, an agentic C‑arm control framework based on multimodal large language models (MLLMs) is highly desirable, as it can incorporate clinician feedback and use reasoning to make adjustments toward more accurate positioning. Skeletal landmark localization is essential for C‑arm control, and we investigate adapting MLLMs for autonomous landmark localization.
Methods: We used an annotated synthetic X‑ray dataset and a real X‑ray dataset. Each X‑ray in both datasets is paired with several skeletal landmarks. We fine‑tuned two MLLMs and tasked them with retrieving the closest landmarks from each X‑ray. Quantitative evaluations of landmark localization were performed and compared against a leading DL approach. We further conducted qualitative experiments demonstrating: (1) how an MLLM can correct an initially incorrect prediction through reasoning, and (2) how the MLLM can sequentially navigate the C‑arm toward a target location.
Results: On both datasets, fine‑tuned MLLMs demonstrate competitive performance across all localization tasks when compared with the DL approach. In the qualitative experiments, the MLLMs provide evidence of reasoning and spatial awareness.
Conclusion: This study shows that fine‑tuned MLLMs achieve accurate skeletal landmark localization and hold promise for agentic autonomous C‑arm control. Our code is available athttps://github.com/marszzibros/C‑arm‑localization‑LLMs.git
Authors:Cuiling Sun, Linkai Peng, Adam Murphy, Elif Keles, Hiten D. Patel, Ashley Ross, Frank Miller, Baris Turkbey, Andrea Mia Bejar, Halil Ertugrul Aktas, Gorkem Durak, Ulas Bagci
Abstract:
Automated 3D segmentation of prostate lesions from biparametric MRI (bp‑MRI) is essential for reliable algorithmic analysis, but achieving high precision remains challenging. Volumetric methods must combine multiple modalities while ensuring anatomical consistency, but current models struggle to integrate cross‑modal information reliably. While vision‑language models (VLMs) are replacing the currently used architectural designs, they still lack the fine‑grained, lesion‑level semantics required for effective localized guidance. To address these limitations, we propose a new multi‑encoder U‑Net architecture incorporating three key innovations: (1) an alignment loss that enhances foreground text‑image similarity to inject lesion semantics; (2) a heatmap loss that calibrates the similarity map and suppresses spurious background activations; and (3) a final‑stage, confidence‑gated multi‑head cross‑attention refiner that performs localized boundary edits in high‑confidence regions. A phase‑scheduled training regime stabilizes the optimization of these components. Our method consistently outperforms prior approaches, establishing a new state‑of‑the‑art on the PI‑CAI dataset through enhanced multi‑modal fusion and localized text guidance. Our code is available at https://github.com/NUBagciLab/Prostate‑Lesion‑Segmentation.
Authors:Savya Khosla, Sethuraman T, Aryan Chadha, Alex Schwing, Derek Hoiem
Abstract:
Despite recent progress, vision‑language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, which hurts tasks like open‑vocabulary semantic segmentation; and (2) high token counts for fine‑grained visual representations, which limits scalability to long videos. This work addresses both limitations. We propose T‑REN (Text‑aligned Region Encoder Network), an efficient encoder that maps visual data to a compact set of text‑aligned region‑level representations (or region tokens). T‑REN achieves this through a lightweight network added on top of a frozen vision backbone, trained to pool patch‑level representations within each semantic region into region tokens and align them with region‑level text annotations. With only 3.7% additional parameters compared to the vision‑language backbone, this design yields substantially stronger dense cross‑modal understanding while reducing the token count by orders of magnitude. Specifically, T‑REN delivers +5.9 mIoU on ADE20K open‑vocabulary segmentation, +18.4% recall on COCO object‑level text‑image retrieval, +15.6% recall on Ego4D video object localization, and +17.6% mIoU on VSPW video scene parsing, all while reducing token counts by more than 24x for images and 187x for videos compared to the patch‑based vision‑language backbone. The code and model are available at https://github.com/savya08/T‑REN.
Authors:Rui Qian, Chuanhang Deng, Qiang Huang, Jian Xiong, Mingxuan Li, Yingbo Zhou, Wei Zhai, Jintao Chen, Dejing Dou
Abstract:
Reasoning segmentation requires models to ground complex, implicit textual queries into precise pixel‑level masks. Existing approaches rely on a single segmentation token \texttt<SEG>, whose hidden state implicitly encodes both semantic reasoning and spatial localization, limiting the model's ability to explicitly disentangle what to segment from where to segment. We introduce AnchorSeg, which reformulates reasoning segmentation as a structured conditional generation process over image tokens, conditioned on language grounded query banks. Instead of compressing all semantic reasoning and spatial localization into a single embedding, AnchorSeg constructs an ordered sequence of query banks: latent reasoning tokens that capture intermediate semantic states, and a segmentation anchor token that provides explicit spatial grounding. We model spatial conditioning as a factorized distribution over image tokens, where the anchor query determines localization signals while contextual queries provide semantic modulation. To bridge token‑level predictions and pixel‑level supervision, we propose Token‑‑Mask Cycle Consistency (TMCC), a bidirectional training objective that enforces alignment across resolutions. By explicitly decoupling spatial grounding from semantic reasoning through structured language grounded query banks, AnchorSeg achieves state‑of‑the‑art results on ReasonSeg test set (67.7% gIoU and 68.1% cIoU). All code and models are publicly available at https://github.com/rui‑qian/AnchorSeg.
Authors:Wei Yao, Haohan Ma, Hongwen Zhang, Yunlian Sun, Liangjun Xing, Zhile Yang, Yuanjun Guo, Yebin Liu, Jinhui Tang
Abstract:
Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexities in multi‑agent coordination, and limited generalization across objects. In this paper, we present SynAgent, a unified framework that enables scalable and physically plausible cooperative manipulation by leveraging Solo‑to‑Cooperative Agent Synergy to transfer skills from single‑agent human‑object interaction to multi‑agent human‑object‑human scenarios. To maintain semantic integrity during motion transfer, we introduce an interaction‑preserving retargeting method based on an Interact Mesh constructed via Delaunay tetrahedralization, which faithfully maintains spatial relationships among humans and objects. Building upon this refined data, we propose a single‑agent pretraining and adaptation paradigm that bootstraps synergistic collaborative behaviors from abundant single‑human data through decentralized training and multi‑agent PPO. Finally, we develop a trajectory‑conditioned generative policy using a conditional VAE, trained via multi‑teacher distillation from motion imitation priors to achieve stable and controllable object‑level trajectory execution. Extensive experiments demonstrate that SynAgent significantly outperforms existing baselines in both cooperative imitation and trajectory‑conditioned control, while generalizing across diverse object geometries. Codes and data will be available after publication. Project Page: http://yw0208.github.io/synagent
Authors:Jiaqi Wang, Haoge Deng, Ting Pan, Yang Liu, Chengyuan Wang, Fan Zhang, Yonggang Qi, Xinlong Wang
Abstract:
Uniform Discrete Diffusion Model (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with reinforcement learning remains largely unexplored. We observe that naively applying GRPO to UDM leads to training instability and marginal performance gains. To address this, we propose UDM‑GRPO, the first framework to integrate UDM with RL. Our method is guided by two key insights: (i) treating the final clean sample as the action provides more accurate and stable optimization signals; and (ii) reconstructing trajectories via the diffusion forward process better aligns probability paths with the pretraining distribution. Additionally, we introduce two strategies, Reduced‑Step and CFG‑Free, to further improve training efficiency. UDM‑GRPO significantly improves base model performance across multiple T2I tasks. Notably, GenEval accuracy improves from 69% to 96% and PickScore increases from 20.46 to 23.81, achieving state‑of‑the‑art performance in both continuous and discrete settings. On the OCR benchmark, accuracy rises from 8% to 57%, further validating the generalization ability of our method. Code is available at https://github.com/Yovecent/UDM‑GRPO.
Authors:Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li, Guang Li, Lingdong Kong, Yingyan Li, Han Wang, Shaoqing Xu, Yuechen Luo, Fang Li, Chenxu Dang, Junli Wang, Tao Xu, Jing Wu, Jianhua Wu, Xiaoshuai Hao, Wen Zhang, Tianyi Jiang, Lingfeng Zhang, Lei Zhou, Yingbo Tang, Jie Wang, Yinfeng Gao, Xizhou Bu, Haochen Tian, Yihang Qiu, Feiyang Jia, Lin Liu, Yigu Ge, Hanbing Li, Yuannan Shen, Jianwei Cui, Hongwei Xie, Bing Wang, Haiyang Sun, Jingwei Zhao, Jiahui Huang, Pei Liu, Zeyu Zhu, Yuncheng Jiang, Zibin Guo, Chuhong Gong, Hanchao Leng, Kun Ma, Naiyan Wang, Guang Chen, Kuiyuan Yang, Hangjun Ye, Long Chen
Abstract:
Chain‑of‑Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA‑based autonomous driving, yet its autoregressive nature imposes a latency cost that is prohibitive for real‑time deployment. Latent CoT methods attempt to close this gap by compressing reasoning into continuous hidden states, but consistently fall short of their explicit counterparts. We suggest that this is due to purely linguistic latent representations compressing a symbolic abstraction of the world, rather than the causal dynamics that actually govern driving. Thus, we present OneVL (One‑step latent reasoning and planning with Vision‑Language explanations), a unified VLA and World Model framework that routes reasoning through compact latent tokens supervised by dual auxiliary decoders. Alongside a language decoder that reconstructs text CoT, we introduce a visual world model decoder that predicts future‑frame tokens, forcing the latent space to internalize the causal dynamics of road geometry, agent motion, and environmental change. A three‑stage training pipeline progressively aligns these latents with trajectory, language, and visual objectives, ensuring stable joint optimization. In inference, the auxiliary decoders are discarded, and all latent tokens are prefilled in a single parallel pass, matching the speed of answer‑only prediction. Across four benchmarks, OneVL becomes the first latent CoT method to surpass explicit CoT, delivering superior accuracy at answer‑only latency. These results show that with world model supervision, latent CoT produces more generalizable representations than verbose token‑by‑token reasoning. Code has been open‑sourced to the community. Project Page: https://xiaomi‑embodied‑intelligence.github.io/OneVL
Authors:Jiyao Liu, Jianghan Shen, Sida Song, Tianbin Li, Xiaojia Liu, Rongbin Li, Ziyan Huang, Jiashi Lin, Junzhi Ning, Changkai Ji, Siqi Luo, Wenjie Li, Chenglong Ma, Ming Hu, Jing Xiong, Jin Ye, Bin Fu, Ningsheng Xu, Yirong Chen, Lei Jin, Hong Chen, Junjun He
Abstract:
Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large‑scale external knowledge. In medicine, developing clinical guidelines critically depends on such deep evidence integration. However, existing benchmarks fail to evaluate this capability in realistic workflows requiring multi‑step evidence integration and expert‑level judgment. To address this gap, we introduce MedProbeBench, the first benchmark leveraging high‑quality clinical guidelines as expert‑level references. Medical guidelines, with their rigorous standards in neutrality and verifiability, represent the pinnacle of medical expertise and pose substantial challenges for deep research agents. For evaluation, we propose MedProbe‑Eval, a comprehensive evaluation framework featuring: (1) Holistic Rubrics with 1,200+ task‑adaptive rubric criteria for comprehensive quality assessment, and (2) Fine‑grained Evidence Verification for rigorous validation of evidence precision, grounded in 5,130+ atomic claims. Evaluation of 17 LLMs and deep research agents reveals critical gaps in evidence integration and guideline generation, underscoring the substantial distance between current capabilities and expert‑level clinical guideline development. Project: https://github.com/uni‑medical/MedProbeBench
Authors:Zeeshan Nisar, Friedrich Feuerhake, Thomas Lampert
Abstract:
A key challenge in segmentation in digital histopathology is inter‑ and intra‑stain variations as it reduces model performance. Labelling each stain is expensive and time‑consuming so methods using stain transfer via CycleGAN, have been developed for training multi‑stain segmentation models using labels from a single stain. Nevertheless, CycleGAN tends to introduce noise during translation because of the one‑to‑many nature of some stain pairs, which conflicts with its cycle consistency loss. To address this, we propose the Domain Shift Aware CycleGAN, which reduces the presence of such noise. Furthermore, we evaluate several advances from the field of machine learning aimed at resolving similar problems and compare their effectiveness against DSA‑CycleGAN in the context of multi‑stain glomeruli segmentation. Experiments demonstrate that DSA‑CycleGAN not only improves segmentation performance in glomeruli segmentation but also outperforms other methods in reducing noise. This is particularly evident when translating between biologically distinct stains. The code is publicly available at https://github.com/zeeshannisar/DSA‑CycleGAN.
Authors:Jiamin Zheng, Jingwen Yu, Guangcheng Chen, Hong Zhang
Abstract:
Indoor robot navigation is often compromised by glass surfaces, which severely corrupt depth sensor measurements. While foundation models like Depth Anything 3 provide excellent geometric priors, they lack an absolute metric scale. We propose a training‑free framework that leverages depth foundation models as a structural prior, employing a robust local RANSAC‑based alignment to fuse it with raw sensor depth. This naturally avoids contamination from erroneous glass measurements and recovers an accurate metric scale. Furthermore, we introduce \tiGlassRecon, a novel RGB‑D dataset with geometrically derived ground truth for glass regions. Extensive experiments demonstrate that our approach consistently outperforms state‑of‑the‑art baselines, especially under severe sensor depth corruption. The dataset and related code will be released at https://github.com/jarvisyjw/GlassRecon.
Authors:Yongrui Heng, Chaoya Jiang, Han Yang, Shikun Zhang, Wei Ye
Abstract:
Self‑evolution of multimodal large language models (MLLMs) remains a critical challenge: pseudo‑label‑based methods suffer from progressive quality degradation as model predictions drift, while template‑based methods are confined to a static set of transformations that cannot adapt in difficulty or diversity. We contend that robust, continuous self‑improvement requires not only deterministic external feedback independent of the model's internal certainty, but also a mechanism to perpetually diversify the training distribution. To this end, we introduce EVE (Executable Visual transformation‑based self‑Evolution), a novel framework that entirely bypasses pseudo‑labels by harnessing executable visual transformations continuously enriched in both variety and complexity. EVE adopts a Challenger‑Solver dual‑policy architecture. The Challenger maintains and progressively expands a queue of visual transformation code examples, from which it synthesizes novel Python scripts to perform dynamic visual transformations. Executing these scripts yields VQA problems with absolute, execution‑verified ground‑truth answers, eliminating any reliance on model‑generated supervision. A multi‑dimensional reward system integrating semantic diversity and dynamic difficulty calibration steers the Challenger to enrich its code example queue while posing progressively more challenging tasks, preventing mode collapse and fostering reciprocal co‑evolution between the two policies. Extensive experiments demonstrate that EVE consistently surpasses existing self‑evolution methods, establishing a robust and scalable paradigm for verifiable MLLM self‑evolution. The code is available at https://github.com/0001Henry/EVE .
Authors:Claudia Cuttano, Gabriele Trivigno, Carlo Masone, Stefan Roth
Abstract:
Recent advances in semantic correspondence rely on dual‑encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion‑parameter models generalize poorly beyond training keypoints, revealing a gap between benchmark performance and real‑world usability, where queried points rarely match those seen during training. Building upon DINOv2, we introduce MARCO, a unified model for generalizable correspondence driven by a novel training framework that enhances both fine‑grained localization and semantic generalization. By coupling a coarse‑to‑fine objective that refines spatial precision with a self‑distillation framework, which expands sparse supervision beyond annotated regions, our approach transforms a handful of keypoints into dense, semantically coherent correspondences. MARCO sets a new state of the art on SPair‑71k, AP‑10K, and PF‑PASCAL, with gains that amplify at fine‑grained localization thresholds (+8.9 PCK@0.01), strongest generalization to unseen keypoints (+5.1, SPair‑U) and categories (+4.7, MP‑100), while remaining 3x smaller and 10x faster than diffusion‑based approaches. Code is available at https://github.com/visinf/MARCO .
Authors:Svetlana Pavlitska, Malte Stüven, Beyza Keskin, J. Marius Zöllner
Abstract:
Mixture‑of‑Experts (MoE) models provide a structured approach to combining specialized neural networks and offer greater interpretability than conventional ensembles. While MoEs have been successfully applied to image classification and semantic segmentation, their use in object detection remains limited due to challenges in merging dense and structured predictions. In this work, we investigate model‑level mixtures of object detectors and analyze their suitability for improving performance and interpretability in object detection. We propose an MoE architecture that combines YOLO‑based detectors trained on semantically disjoint data subsets, with a learned gating network that dynamically weights expert contributions. We study different strategies for fusing detection outputs and for training the gating mechanism, including balancing losses to prevent expert collapse. Experiments on the BDD100K dataset demonstrate that the proposed MoE consistently outperforms standard ensemble approaches and provides insights into expert specialization across domains, highlighting model‑level MoEs as a viable alternative to traditional ensembling for object detection. Our code is available at https://github.com/KASTEL‑MobilityLab/mixtures‑of‑experts/.
Authors:Andreas Kriegler, Csaba Beleznai, Margrit Gelautz
Abstract:
Symmetric objects are common in daily life and industry, yet their inherent orientation ambiguities that impede the training of deep learning networks for pose estimation are rarely discussed in the literature. To cope with these ambiguities, existing solutions typically require the design of specific loss functions and network architectures or resort to symmetry‑invariant evaluation metrics. In contrast, we focus on the numeric representation of the rotation itself, modifying trigonometric identities with the degrees of symmetry derived from the objects' shapes. We use our representation, SARR, to obtain canonic (symmetry‑resolved) poses for the symmetric objects in two popular 6D pose estimation datasets, T‑LESS and ITODD, where SARR is unique and continuous w.r.t. the visual appearance. This allows us to use a standard CNN for 3D orientation estimation whose performance is evaluated with the symmetry‑sensitive cosine distance \textAR_\textC. Our networks outperform the state of the art using \textAR_\textC and achieve satisfactory performance when using conventional symmetry‑invariant measures. Our method does not require any 3D models but only depth, or, as part of an additional experiment, texture‑less RGB/grayscale images as input. We also show that networks trained on SARR outperform the same networks trained on rotation matrices, Euler angles, quaternions, standard trigonometrics or the recently popular 6d representation ‑‑ even in inference scenarios where no prior knowledge of the objects' symmetry properties is available. Code and a visualization toolkit are available at https://github.com/akriegler/SARR .
Authors:Chenxi Zhao, Chen Zhu, Xiaokun Feng, Aiming Hao, Jiashu Zhu, Jiachen Lei, Jiahong Wu, Xiangxiang Chu, Jufeng Yang
Abstract:
Few‑step generation has been a long‑standing goal, with recent one‑step generation methods exemplified by MeanFlow achieving remarkable results. Existing research on MeanFlow primarily focuses on class‑to‑image generation. However, an intuitive yet unexplored direction is to extend the condition from fixed class labels to flexible text inputs, enabling richer content creation. Compared to the limited class labels, text conditions pose greater challenges to the model's understanding capability, necessitating the effective integration of powerful text encoders into the MeanFlow framework. Surprisingly, although incorporating text conditions appears straightforward, we find that integrating powerful LLM‑based text encoders using conventional training strategies results in unsatisfactory performance. To uncover the underlying cause, we conduct detailed analyses and reveal that, due to the extremely limited number of refinement steps in the MeanFlow generation, such as only one step, the text feature representations are required to possess sufficiently high discriminability. This also explains why discrete and easily distinguishable class features perform well within the MeanFlow framework. Guided by these insights, we leverage a powerful LLM‑based text encoder validated to possess the required semantic properties and adapt the MeanFlow generation process to this framework, resulting in efficient text‑conditioned synthesis for the first time. Furthermore, we validate our approach on the widely used diffusion model, demonstrating significant generation performance improvements. We hope this work provides a general and practical reference for future research on text‑conditioned MeanFlow generation. The code is available at https://github.com/AMAP‑ML/EMF.
Authors:Venkatesh Thirugnana Sambandham, Torsten Schön
Abstract:
Modern text‑to‑image (T2I) models amplify harmful societal biases, challenging their ethical deployment. We introduce an inference‑time method that reliably mitigates social bias while keeping prompt semantics and visual context (background, layout, and style) intact. This ensures context persistency and provides a controllable parameter to adjust mitigation strength, giving practitioners fine‑grained control over fairness‑coherence trade‑offs. Using Embedding Arithmetic, we analyze how bias is structured in the embedding space and correct it without altering model weights, prompts, or datasets. Experiments on FLUX 1.0‑Dev and Stable Diffusion 3.5‑Large show that the conditional embedding space forms a complex, entangled manifold rather than a grid of disentangled concepts. To rigorously assess semantic preservation beyond the circularity and bias limitations of of CLIP scores, we propose the Concept Coherence Score (CCS). Evaluated against this robust metric, our lightweight, tuning‑free method significantly outperforms existing baselines in improving diversity while maintaining high concept coherence, effectively resolving the critical fairness‑coherence trade‑off. By characterizing how models represent social concepts, we establish geometric understanding of latent space as a principled path toward more transparent, controllable, and fair image generation.
Authors:Ammar Bhilwarawala, Mainak Bandyopadhyay
Abstract:
Automated fetal head segmentation in ultrasound images is critical for accurate biometric measurements in prenatal care. While existing deep learning approaches have achieved a reasonable performance, they struggle with issues like low contrast, noise, and complex anatomical boundaries which are inherent to ultrasound imaging. This paper presents Attention‑ResUNet. It is a novel architecture that synergistically combines residual learning with multi‑scale attention mechanisms in order to achieve enhanced fetal head segmentation. Our approach integrates attention gates at four decoder levels to focus selectively on anatomically relevant regions while suppressing the background noise, and complemented by residual connections which facilitates gradient flow and feature reuse. Extensive evaluation on the HC18 Challenge dataset where n = 200 demonstrates that Attention ResUNet achieves a superior performance with a mean Dice score of 99.30 +/‑ 0.14% against similar architectures. It significantly outperforms five baseline architectures including ResUNet (99.26%), Attention U‑Net (98.79%), Swin U‑Net (98.60%), Standard U‑Net (98.58%), and U‑Net++ (97.46%). Through statistical analysis we confirm highly significant improvements (p < 0.001) with effect sizes that range from 0.230 to 13.159 (Cohen's d). Using Saliency map analysis, we reveal that our architecture produces highly concentrated, anatomically consistent activation patterns, which demonstrate an enhanced interpretability which is crucial for clinical deployment. The proposed method establishes a new state of the art performance for automated fetal head segmentation whilst maintaining computational efficiency with 14.7M parameters and a 45 GFLOPs inference cost. Code repository: https://github.com/Ammar‑ss
Authors:Xiao Lingao, Yang He
Abstract:
Large‑scale dataset distillation requires storing auxiliary soft labels that can be 30‑40x larger on ImageNet‑1K and 200x larger on ImageNet‑21K than the condensed images, undermining the goal of dataset compression. We identify two fundamental issues necessitating such extensive labels: (1) insufficient image diversity, where high within‑class similarity in synthetic images requires extensive augmentation, and (2) insufficient supervision diversity, where limited variety in supervisory signals during training leads to performance degradation at high compression rates. To address these challenges, we propose Label Pruning and Quantization for Large‑scale Distillation (LPQLD). We enhance image diversity via class‑wise batching and batch‑normalization supervision during synthesis. For supervision diversity, we introduce Label Pruning with Dynamic Knowledge Reuse to improve label‑per‑augmentation diversity, and Label Quantization with Calibrated Student‑Teacher Alignment to improve augmentation‑per‑image diversity. Our approach reduces soft label storage by 78x on ImageNet‑1K and 500x on ImageNet‑21K while improving accuracy by up to 7.2% and 2.8%, respectively. Extensive experiments validate the superiority of LPQLD across different network architectures and dataset distillation methods. Code is available at https://github.com/he‑y/soft‑label‑pruning‑quantization‑for‑dataset‑distillation.
Authors:Chengan Che, Chao Wang, Jiayuan Huang, Xinyue Chen, Luis C. Garcia-Peraza-Herrera
Abstract:
Recent advancements in self‑supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extending these visual foundations to multi‑modal reasoning tasks is severely bottlenecked by the prohibitive cost of expert textual annotations. To overcome this scalability limitation, we introduce LIME, a large‑scale multi‑modal dataset derived from open‑access surgical videos using human‑free, Large Language Model (LLM)‑generated narratives. While LIME offers immense scalability, unverified generated texts may contain errors, including hallucinations, that could potentially lead to catastrophically degraded pre‑trained medical priors in standard contrastive pipelines. To mitigate this, we propose SurgLIME, a parameter‑efficient Vision‑Language Pre‑training (VLP) framework designed to learn reliable cross‑modal alignments using noisy narratives. SurgLIME preserves foundational medical priors using a LoRA‑adapted dual‑encoder architecture and introduces an automated confidence estimation mechanism that dynamically down‑weights uncertain text during contrastive alignment. Evaluations on the AutoLaparo and Cholec80 benchmarks show that SurgLIME achieves competitive zero‑shot cross‑modal alignment while preserving the robust linear probing performance of the visual foundation model. Dataset, code, and models are publicly available at https://github.com/visurg‑ai/SurgLIME.
Authors:Zehua Zang, Xi Wang, Fuchun Sun, Xiao Xu, Lixiang Lium, Jiahuan Zhou, Jiangmeng Li
Abstract:
Vision‑Language‑Action models (VLAs) achieve remarkable performance in sequential decision‑making but remain fragile to subtle environmental shifts, such as small changes in object pose. We attribute this brittleness to trajectory overfitting, where VLAs over‑attend to the spurious correlation between actions and entities, then reproduce memorized action patterns. We propose Perturbation learning with Delayed Feedback (PDF), a verifier‑free test‑time adaptation framework that improves decision performance without fine‑tuning the base model. PDF mitigates the spurious correlation through uncertainty‑based data augmentation and action voting, while an adaptive scheduler allocates augmentation budgets to balance performance and efficiency. To further improve stability, PDF learns a lightweight perturbation module that retrospectively adjusts action logits guided by delayed feedback, correcting overconfidence issue. Experiments on LIBERO (+7.4% success rate) and Atari (+10.3 human normalized score) demonstrate consistent gains of PDF in task success over vanilla VLA and VLA with test‑time adaptation, establishing a practical path toward reliable test‑time adaptation in multimodal decision‑making agents. The code is available at \hrefhttps://github.com/zhoujiahuan1991/CVPR2026‑PDFhttps://github.com/zhoujiahuan1991/CVPR2026‑PDF.
Authors:Hyeonseo Jang, Hyuk Kwon, Kibok Lee
Abstract:
We investigate recently introduced domain‑class incremental learning scenarios for vision‑language models (VLMs). Recent works address this challenge using parameter‑efficient methods, such as prefix‑tuning or adapters, which facilitate model adaptation to downstream tasks by incorporating task‑specific information into input tokens through additive vectors. However, previous approaches often normalize the weights of these vectors, disregarding the fact that different input tokens require different degrees of adjustment. To overcome this issue, we propose Dynamic Prefix Weighting (DPW), a framework that dynamically assigns weights to prefixes, complemented by adapters. DPW consists of 1) a gating module that adjusts the weights of each prefix based on the importance of the corresponding input token, and 2) a weighting mechanism that derives adapter output weights as a residual of prefix‑tuning weights, ensuring that adapters are utilized only when necessary. Experimental results demonstrate that our method achieves state‑of‑the‑art performance in domain‑class incremental learning scenarios for VLMs. The code is available at: https://github.com/YonseiML/dpw.
Authors:Zixu Li, Yupeng Hu, Zhiwei Chen, Shiqi Zhang, Qinlei Huang, Zhiheng Fu, Yinwei Wei
Abstract:
Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recommendation systems, it encounters a severe challenge in practical scenarios known as the Noise Triplet Correspondence (NTC) problem. This issue primarily arises from the high cost and subjectivity involved in annotating triplet data. To address this problem, we identify two central challenges: the precise estimation of composed semantic discrepancy and the insufficient progressive adaptation to modification discrepancy. To tackle these challenges, we propose a cHrono‑synergiA roBust progressIve learning framework for composed image reTrieval (HABIT), which consists of two core modules. First, the Mutual Knowledge Estimation Module quantifies sample cleanliness by calculating the Transition Rate of mutual information between the composed feature and the target image, thereby effectively identifying clean samples that align with the intended modification semantics. Second, the Dual‑consistency Progressive Learning Module introduces a collaborative mechanism between the historical and current models, simulating human habit formation to retain good habits and calibrate bad habits, ultimately enabling robust learning under the presence of NTC. Extensive experiments conducted on two standard CIR datasets demonstrate that HABIT significantly outperforms most methods under various noise ratios, exhibiting superior robustness and retrieval performance. Codes are available at https://github.com/Lee‑zixu/HABIT
Authors:Julio Silva-Rodríguez, Ender Konukoglu
Abstract:
Super‑resolution (SR) models are attracting growing interest for enhancing minimally invasive surgery and diagnostic videos under hardware constraints. However, valid concerns remain regarding the introduction of hallucinated structures and amplified noise, limiting their reliability in safety‑critical settings. We propose a direct and practical framework to make SR systems more trustworthy by identifying where reconstructions are likely to fail. Our approach integrates a lightweight error‑prediction network that operates on intermediate representations to estimate pixel‑wise reconstruction error. The module is computationally efficient and low‑latency, making it suitable for real‑time deployment. We convert these predictions into operational failure decisions by constructing Conformal Failure Masks (CFM), which localize regions where the SR output should not be trusted. Built on conformal risk control principles, our method provides theoretical guarantees for controlling both the tolerated error limit and the miscoverage in detected failures. We evaluate our approach on image and video SR, demonstrating its effectiveness in detecting unreliable reconstructions in endoscopic and robotic surgery settings. To our knowledge, this is the first study to provide a model‑agnostic, theoretically grounded approach to improving the safety of real‑time endoscopic image SR.
Authors:Koya Sakamoto, Taiki Miyanishi, Daichi Azuma, Shuhei Kurita, Shu Morikuni, Naoya Chiba, Motoaki Kawanabe, Yusuke Iwasawa, Yutaka Matsuo
Abstract:
Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task‑relevant evidence. However, existing visual search and embodied AI benchmarks, including EQA, typically rely on static observations or constrained egocentric motion, and thus do not explicitly evaluate fine‑grained viewpoint‑dependent phenomena that arise under unrestricted 5‑DoF viewpoint control in real‑world 3D environments, such as visibility changes caused by vertical viewpoint shifts, revealing contents inside containers, and disambiguating object attributes that are only observable from specific angles. To address this limitation, we introduce E3VS‑Bench, a benchmark for embodied 3D visual search where agents must control their viewpoints in 5‑DoF to gather viewpoint‑dependent evidence for question answering. E3VS‑Bench consists of 99 high‑fidelity 3D scenes reconstructed using 3D Gaussian Splatting and 2,014 question‑driven episodes. 3D Gaussian Splatting enables photorealistic free‑viewpoint rendering that preserves fine‑grained visual details (e.g., small text and subtle attributes) often degraded in mesh‑based simulators, thereby allowing the construction of questions that cannot be answered from a single view and instead require active inspection across viewpoints in 5‑DoF. We evaluate multiple state‑of‑the‑art VLMs and compare their performance with humans. Despite strong 2D reasoning ability, all models exhibit a substantial gap from humans, highlighting limitations in active perception and coherent viewpoint planning specifically under full 5‑DoF viewpoint changes.
Authors:Qidong Wang, Junjie Hu, Ming Jiang
Abstract:
Recent work has increasingly explored neuron‑level interpretation in vision‑language models (VLMs) to identify neurons critical to final predictions. However, existing neuron analyses generally focus on single tasks, limiting the comparability of neuron importance across tasks. Moreover, ranking strategies tend to score neurons in isolation, overlooking how task‑dependent information pathways shape the write‑in effects of feed‑forward network (FFN) neurons. This oversight can exacerbate neuron polysemanticity in multi‑task settings, introducing noise into the identification and intervention of task‑critical neurons. In this study, we propose HONES (Head‑Oriented Neuron Explanation & Steering), a gradient‑free framework for task‑aware neuron attribution and steering in multi‑task VLMs. HONES ranks FFN neurons by their causal write‑in contributions conditioned on task‑relevant attention heads, and further modulates salient neurons via lightweight scaling. Experiments on four diverse multimodal tasks and two popular VLMs show that HONES outperforms existing methods in identifying task‑critical neurons and improves model performance after steering. Our source code is released at: https://github.com/petergit1/HONES.
Authors:Feixue Shao, Guangze Shi, Xueyu Liu, Yongfei Wu, Mingqiang Wei, Jianan Zhang, Jianbo Lu, Guiying Yan, Weihua Yang
Abstract:
Visual decoding of neurophysiological signals is a critical challenge for brain‑computer interfaces (BCIs) and computational neuroscience. However, current approaches are often constrained by the systematic and stochastic gaps between neural and visual modalities, largely neglecting the intrinsic computational mechanisms of the Human Visual System (HVS). To address this, we propose Brain‑Inspired Capture (BI‑Cap), a neuromimetic perceptual simulation paradigm that aligns these modalities by emulating HVS processing. Specifically, we construct a neuromimetic pipeline comprising four biologically plausible dynamic and static transformations, coupled with Mutual Information (MI)‑guided dynamic blur regulation to simulate adaptive visual processing. Furthermore, to mitigate the inherent non‑stationarity of neural activity, we introduce an evidence‑driven latent space representation. This formulation explicitly models uncertainty, thereby ensuring robust neural embeddings. Extensive evaluations on zero‑shot brain‑to‑image retrieval across two public benchmarks demonstrate that BI‑Cap substantially outperforms state‑of‑the‑art methods, achieving relative gains of 9.2% and 8.0%, respectively. We have released the source code on GitHub through the link https://github.com/flysnow1024/BI‑Cap.
Authors:Yiwei Zhang, Xuesong Chen, Jin Gao, Hanshi Wang, Fudong Ge, Weiming Hu, Shaoshuai Shi, Zhipeng Zhang
Abstract:
Vision‑Language Models(VLMs) excel at autoregressive text generation, yet end‑to‑end autonomous driving requires multi‑task learning with structured outputs and heterogeneous decoding behaviors, such as autoregressive language generation, parallel object detection and trajectory regression. To accommodate these differences, existing systems typically introduce separate or cascaded decoders, resulting in architectural fragmentation and limited backbone reuse. In this work, we present a unified autonomous driving framework built upon a pretrained VLM, where heterogeneous decoding behaviors are reconciled within a single transformer decoder. We demonstrate that pretrained VLM attention exhibits strong transferability beyond pure language modeling. By organizing visual and structured query tokens within a single causal decoder, structured queries can naturally condition on visual context through the original attention mechanism. Textual and structured outputs share a common attention backbone, enabling stable joint optimization across heterogeneous tasks. Trajectory planning is realized within the same causal LLM decoder by introducing structured trajectory queries. This unified formulation enables planning to share the pretrained attention backbone with images and perception tokens. Extensive experiments on end‑to‑end autonomous driving benchmarks demonstrate state‑of‑the‑art performance, including 0.28 L2 and 0.18 collision rate on nuScenes open‑loop evaluation and competitive results (86.8 PDMS) on NAVSIM closed‑loop evaluation. The full model preserves multi‑modal generation capability, while an efficient inference mode achieves approximately 40% lower latency. Code and models are available at https://github.com/Z1zyw/OneDrive
Authors:Yingjie Feng, Yi Wang, Jiaze Wang, Anfeng Liu, Zhuotao Tian
Abstract:
Self‑supervised contrastive learning has emerged as a powerful paradigm for skeleton‑based action recognition by enforcing consistency in the embedding space. However, existing methods rely on binary contrastive objectives that overlook the intrinsic continuity of human motion, resulting in fragmented feature clusters and rigid class boundaries. To address these limitations, we propose TranCLR, a Transitional anchor‑based Contrastive Learning framework that captures the continuous geometry of the action space. Specifically, the proposed Action Transitional Anchor Construction (ATAC) explicitly models the geometric structure of transitional states to enhance the model's perception of motion continuity. Building upon these anchors, a Multi‑Level Geometric Manifold Calibration (MGMC) mechanism is introduced to adaptively calibrate the action manifold across multiple levels of continuity, yielding a smoother and more discriminative representation space. Extensive experiments on the NTU RGB+D, NTU RGB+D 120 and PKU‑MMD datasets demonstrate that TranCLR achieves superior accuracy and calibration performance, effectively learning continuous and uncertainty‑aware skeleton representations. The code is available at https://github.com/Philchieh/TranCLR.
Authors:Zixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang, Guozhi Qiu, Zhiheng Fu, Meng Liu
Abstract:
With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a multi‑modal query consisting of a reference video and a piece of modification text as input. The modification text conveys the user's intended alterations to the reference video. Based on this input, the model aims to retrieve the most relevant target video. In the CVR task, there exists a substantial discrepancy in information density between video and text modalities. Traditional composition methods tend to bias the composed feature toward the reference video, which leads to suboptimal retrieval performance. This limitation is significant due to the presence of three core challenges: (1) modal contribution entanglement, (2) explicit optimization of composed features, and (3) retrieval uncertainty. To address these challenges, we propose the evidence‑dRivRn dual‑sTream diRectionAl anChor calibration networK (ReTrack). ReTrack is the first CVR framework that improves multi‑modal query understanding by calibrating directional bias in composed features. It consists of three key modules: Semantic Contribution Disentanglement, Composition Geometry Calibration, and Reliable Evidence‑driven Alignment. Specifically, ReTrack estimates the semantic contribution of each modality to calibrate the directional bias of the composed feature. It then uses the calibrated directional anchors to compute bidirectional evidence that drives reliable composed‑to‑target similarity estimation. Moreover, ReTrack exhibits strong generalization to the Composed Image Retrieval (CIR) task, achieving SOTA performance across three benchmark datasets in both CVR and CIR scenarios. Codes are available at https://github.com/Lee‑zixu/ReTrack
Authors:Chuanhao Ma, Hanyu Zhou, Shihan Peng, Yan Li, Tao Gu, Luxin Yan
Abstract:
Vision‑language‑action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine‑grained spatiotemporal manipulation. Typically, existing methods mainly embed spatiotemporal knowledge into visual and action representations, and directly perform a cross‑modal mapping for step‑level action prediction. However, such spatiotemporal reasoning remains largely implicit, making it difficult to handle multiple sequential behaviors with explicit spatiotemporal boundaries. In this work, we propose ST‑π, a structured spatiotemporal VLA model for robotic manipulation. Our model is guided by two key designs: 1) Spatiotemporal VLM. We encode 4D observations and task instructions into latent spaces, and feed them into the LLM to generate a sequence of causally ordered chunk‑level action prompts consisting of sub‑tasks, spatial grounding and temporal grounding. 2) Spatiotemporal action expert. Conditioned on chunk‑level action prompts, we design a structured dual‑generator guidance to jointly model spatial dependencies and temporal causality, thus predicting step‑level action parameters. Within this structured framework, the VLM explicitly plans global spatiotemporal behavior, and the action expert further refines local spatiotemporal control. In addition, we propose a real‑world robotic dataset with structured spatiotemporal annotations for fine‑tuning. Extensive experiments have been conducted to demonstrate the effectiveness of our model. Our code link: https://github.com/chuanhaoma/ST‑pi.
Authors:Shivanshu Agnihotri, Snehashis Majhi, Deepak Ranjan Nayak
Abstract:
Automated polyp segmentation is critical for early colorectal cancer detection and its prevention, yet remains challenging due to weak boundaries, large appearance variations, and limited annotated data. Lightweight segmentation models such as U‑Net, U‑Net++, and PraNet offer practical efficiency for clinical deployment but struggle to capture the rich semantic and structural cues required for accurate delineation of complex polyp regions. In contrast, large Vision Foundation Models (VFMs), including SAM, OneFormer, Mask2Former, and DINOv2, exhibit strong generalization but transfer poorly to polyp segmentation due to domain mismatch, insufficient boundary sensitivity, and high computational cost. To bridge this gap, we propose LiteBounD, a \underlineLigh\underlinetw\underlineeight \underlineBoundary‑guided \underlineDistillation framework that transfers complementary semantic and structural priors from multiple VFMs into compact segmentation backbones. LiteBounD introduces (i) a dual‑path distillation mechanism that disentangles semantic and boundary‑aware representations, (ii) a frequency‑aware alignment strategy that supervises low‑frequency global semantics and high‑frequency boundary details separately, and (iii) a boundary‑aware decoder that fuses multi‑scale encoder features with distilled semantically rich boundary information for precise segmentation. Extensive experiments on both seen (Kvasir‑SEG, CVC‑ClinicDB) and unseen (ColonDB, CVC‑300, ETIS) datasets demonstrate that LiteBounD consistently outperforms its lightweight baselines by a significant margin and achieves performance competitive with state‑of‑the‑art methods, while maintaining the efficiency required for real‑time clinical use. Our code is available at https://github.com/lostinrepo/LiteBounD.
Authors:Hongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu, Ehsan Adeli, C. Karen Liu, Jiajun Wu
Abstract:
Reconstructing 3D human motion and human‑object interactions (HOI) from Internet videos is a fundamental step toward building large‑scale datasets of human behavior. Existing methods struggle to recover globally consistent 3D motion under dynamic cameras, especially for motion types underrepresented in current motion‑capture datasets, and face additional difficulty recovering coherent human‑object interactions in 3D. We introduce a two‑stage framework leveraging 2D diffusion that reconstructs 3D human motion and HOI from Internet videos. In the first stage, we synthesize multi‑view 2D motion data for each domain, leveraging 2D keypoints extracted from Internet videos to incorporate human motions that rarely appear in existing MoCap datasets. In the second stage, a camera‑conditioned multi‑view 2D motion diffusion model is trained on the domain‑specific synthetic data to recover 3D human motion and 3D HOI in the world space. We demonstrate the effectiveness of our method on Internet videos featuring challenging motions such as gymnastics, as well as in‑the‑wild HOI videos, and show that it outperforms prior work in producing realistic human motion and human‑object interaction.
Authors:Miaojing Shi, Jun Huang, Zijie Yue, Hanli Wang
Abstract:
Referring video object segmentation (RVOS) aims to segment the target instance in a video, referred by a text expression. Conventional approaches are mostly supervised learning, requiring expensive pixel‑level mask annotations. To tackle it, weakly‑supervised RVOS has recently been proposed to replace mask annotations with bounding boxes or points, which are however still costly and labor‑intensive. In this paper, we design a novel weakly‑supervised RVOS method, namely WSRVOS, to train the model with only text expressions. Given an input video and the referring expression, we first design a contrastive referring expression augmentation scheme that leverages the captioning capabilities of a multimodal large language model to generate both positive and negative expressions. We extract visual and linguistic features from the input video and generated expressions, then perform bi‑directional vision‑language feature selection and interaction to enable fine‑grained multimodal alignment. Next, we propose an instance‑aware expression classification scheme to optimize the model in distinguishing positive from negative expressions. Also, we introduce a positive‑prediction fusion strategy to generate high‑quality pseudo‑masks, which serve as additional supervision to the model. Last, we design a temporal segment ranking constraint such that the overlaps between mask predictions of temporally neighboring frames are required to conform to specific orders. Extensive experiments on four publicly available RVOS datasets, including A2D Sentences, J‑HMDB Sentences, Ref‑YouTube‑VOS, and Ref‑DAVIS17, demonstrate the superiority of our method. Code is available at https://github.com/viscom‑tongji/WSRVOS.
Authors:Haokun Lin, Xinle Jia, Haobo Xu, Bingchen Yao, Xianglong Guo, Yichen Wu, Zhichao Lu, Ying Wei, Qingfu Zhang, Zhenan Sun
Abstract:
The MXFP4 microscaling format, which partitions tensors into blocks of 32 elements sharing an E8M0 scaling factor, has emerged as a promising substrate for efficient LLM inference, backed by native hardware support on NVIDIA Blackwell Tensor Cores. However, activation outliers pose a unique challenge under this format: a single outlier inflates the shared block scale, compressing the effective dynamic range of the remaining elements and causing significant quantization error. Existing rotation‑based remedies, including randomized Hadamard and learnable rotations, are data‑agnostic and therefore unable to specifically target the channels where outliers concentrate. We propose DuQuant++, which adapts the outlier‑aware fine‑grained rotation of DuQuant to the MXFP4 format by aligning the rotation block size with the microscaling group size (B=32). Because each MXFP4 group possesses an independent scaling factor, the cross‑block variance issue that necessitates dual rotations and a zigzag permutation in the original DuQuant becomes irrelevant, enabling DuQuant++ to replace the entire pipeline with a single outlier‑aware rotation, which halves the online rotation cost while simultaneously smoothing the weight distribution. Extensive experiments on the LLaMA‑3 family under MXFP4 W4A4 quantization show that DuQuant++ consistently achieves state‑of‑the‑art performance. Our code is available at https://github.com/Hsu1023/DuQuant‑v2.
Authors:Hongxu Jiang, Fei Li, Boxiao Yu, Ying Zhang, Kaleb Smith, Kuang Gong, Wei Shao
Abstract:
Three‑dimensional (3D) medical image enhancement, including denoising and super‑resolution, is critical for clinical diagnosis in CT, PET, and MRI. Although diffusion models have shown remarkable success in 2D medical imaging, scaling them to high‑resolution 3D volumes remains computationally prohibitive due to lengthy diffusion trajectories over high‑dimensional volumetric data. We observe that in conditional enhancement, strong anatomical priors in the degraded input render dense noise schedules largely redundant. Leveraging this insight, we propose a sparse voxel‑space diffusion framework that trains and samples on a compact set of uniformly subsampled timesteps. The network predicts clean data directly on the data manifold, supervised in velocity space for stable gradient scaling. A lightweight Structure‑aware Trajectory Modulation (STM) module recalibrates time embeddings at each network block based on local anatomical content, enabling structure‑adaptive denoising over the shared sparse schedule. Operating directly in voxel space, our framework preserves fine anatomical detail without lossy compression while achieving up to 10× training acceleration. Experiments on four datasets spanning CT, PET, and MRI demonstrate state‑of‑the‑art performance on both denoising and super‑resolution tasks. Our code is publicly available at: https://github.com/mirthAI/sparse‑3d‑diffusion.
Authors:Anda Cao, Zhuo Gou, Yi Wang, Kaixuan Chen, Yu Wang, Can Wang, Mingli Song, Jie Song
Abstract:
Merging multiple Low‑Rank Adaptation (LoRA) experts into a single backbone is a promising approach for efficient multi‑task deployment. While existing methods strive to alleviate interference via weight interpolation or subspace alignment, they rest upon the implicit assumption that all LoRA matrices contribute constructively to the merged model. In this paper, we uncover a critical bottleneck in current merging paradigms: the existence of negative modules ‑‑ specific LoRA layers that inherently degrade global performance upon merging. We propose Evolutionary Negative Module Pruning (ENMP), a plug‑and‑play LoRA pruning method to locate and exclude these detrimental modules prior to merging. By leveraging an evolutionary search strategy, ENMP effectively navigates the discrete, non‑differentiable landscape of module selection to identify optimal pruning configurations. Extensive evaluations demonstrate that ENMP consistently boosts the performance of existing merging algorithms, achieving a new state‑of‑the‑art across both language and vision domains. Code is available at https://github.com/CaoAnda/ENMP‑LoRAMerging.
Authors:Song Tang, Yunxiang Bai, Wenxin Su, Mao Ye, Jianwei Zhang, Xiatian Zhu
Abstract:
Source‑Free Domain Adaptation (SFDA) seeks to adapt a source model, which is pre‑trained on a supervised source domain, for a target domain, with only access to unlabeled target training data. Relying on pseudo labeling and/or auxiliary supervision, conventional methods are inevitably error‑prone. To mitigate this limitation, in this work we for the first time explore the potentials of off‑the‑shelf vision‑language (ViL) multimodal models (e.g., CLIP) with rich whilst heterogeneous knowledge. We find that directly applying the ViL model to the target domain in a zero‑shot fashion is unsatisfactory, as it is not specialized for this particular task but largely generic. To make it task‑specific, we propose a novel DIFO++ approach. Specifically, DIFO++ alternates between two steps during adaptation: (i) Customizing the ViL model by maximizing the mutual information with the target model in a prompt learning manner, (ii) Distilling the knowledge of this customized ViL model to the target model, centering on gap region reduction. During progressive knowledge adaptation, we first identify and focus on the gap region, where enclosed features are entangled and class‑ambiguous, as it often captures richer task‑specific semantics. Reliable pseudo‑labels are then generated by fusing predictions from the target and ViL models, supported by a memory mechanism. Finally, gap region reduction is guided by category attention and predictive consistency for semantic alignment, complemented by referenced entropy minimization to suppress uncertainty. Extensive experiments show that DIFO++ significantly outperforms the state‑of‑the‑art alternatives. Our code and data are available at https://github.com/tntek/DIFO‑Plus.
Authors:Haotian Qin, Dongliang Chang, Yueying Gao, Yuexuan Tan, Lei Chen, Zhanyu Ma
Abstract:
As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive generators appear almost monthly, making existing watermark, classifier and inversion methods obsolete upon release. The core problem lies not in model recognition, but in the inability to adapt attribution itself. We introduce IncreFA, a framework that redefines attribution as a structured incremental learning problem, allowing the system to learn continuously as new generative models emerge. IncreFA departs from conventional incremental learning by exploiting the hierarchical relationships among generative architectures and coupling them with continual adaptation. It integrates two mutually reinforcing mechanisms: (1) Hierarchical Constraints, which encode architectural hierarchies through learnable orthogonal priors to disentangle family‑level invariants from model‑specific idiosyncrasies; and (2) a Latent Memory Bank, which replays compact latent exemplars and mixes them to generate pseudo‑unseen samples, stabilising representation drift and enhancing open‑set awareness. On the newly constructed Incremental Attribution Benchmark (IABench) covering 28 generative models released between 2022 and 2025, IncreFA achieves state‑of‑the‑art attribution accuracy and 98.93% unseen detection under a temporally ordered open‑set protocol. Code will be available at https://github.com/Ant0ny44/IncreFA.
Authors:Yuzhe Fu, Hancheng Ye, Cong Guo, Junyao Zhang, Qinsi Wang, Yueqian Lin, Changchun Zhou, Hai, Li, Yiran Chen
Abstract:
Point‑based Neural Networks (PNNs) have become a key approach for point cloud processing. However, a core operation in these models, Farthest Point Sampling (FPS), often introduces significant inference latency, especially for large‑scale processing. Despite existing CUDA‑ and hardware‑level optimizations, FPS remains a major bottleneck due to exhaustive computations across multiple network layers in PNNs, which hinders scalability. Through systematic analysis, we identify three substantial redundancies in FPS, including unnecessary full‑cloud computations, redundant late‑stage iterations, and predictable inter‑layer outputs that make later FPS computations avoidable. To address these, we propose FlashFPS, a hardware‑agnostic, plug‑and‑play framework for FPS acceleration, composed of FPS‑Prune and FPS‑Cache. FPS‑Prune introduces candidate pruning and iteration pruning to reduce redundant computations in FPS while preserving sampling quality, and FPS‑Cache eliminates layer‑wise redundancy via cache‑and‑reuse. Integrated into existing CUDA libraries and state‑of‑the‑art PNN accelerators, FlashFPS achieves 5.16× speedup over the standard CUDA baseline on GPU and 2.69× on PNN accelerators, with negligible accuracy loss, enabling efficient and scalable PNN inference. Codes are released at https://github.com/Yuzhe‑Fu/FlashFPS.
Authors:Hyam Omar Ali, Antoine Crosnier, Romain Abraham, Baptiste Combelles, Fabrice Jégou, Bruno Galerne
Abstract:
Sentinel‑5P (S5P) plays a critical role in atmospheric monitoring; however, its spatial resolution limits fine‑scale analysis. Existing super‑resolution (SR) approaches rely on supervised learning with synthetic low‑resolution (LR) data, since true high‑resolution (HR) data do not exist, limiting their applicability to real observations. We propose a self‑supervised hyperspectral SR framework for S5P that enables training without HR ground truth. The method combines Stein's Unbiased Risk Estimator (SURE) with an equivariant imaging constraint, incorporating the S5P degradation operator and noise statistics derived from signal‑to‑noise ratio (SNR) metadata. We also introduce depthwise separable convolution U‑Net architectures designed for efficiency and spectral fidelity. The framework is evaluated in two settings: (i) LR‑HR, where synthetic LR data are used for direct comparison with supervised learning, and (ii) GT‑SHR, where super‑resolved images surpass the native spatial resolution without HR reference. Results across multiple bands show that self‑supervised models achieve performance comparable to supervised methods while maintaining strong consistency. Qualitative analysis shows improved spatial detail over bicubic interpolation, and validation with EMIT data confirms that reconstructed structures are physically meaningful. Code is available at https://github.com/hyamomar/Sentinel‑5P‑Super‑Resolution/tree/main/self_supervised
Authors:Mainak Singha, Tanisha Gupta, Ankit Jha, Muhammad Haris Khan, Sayantani Ghosh, Biplab Banerjee
Abstract:
Pretrained biomedical vision‑language models (VLMs) such as BioMedCLIP perform well on average but often degrade on challenging modalities where inter‑class margins are small and acquisition‑specific variations are pronounced, especially under few‑shot supervision and when modality priors differ from pretraining corpora substantially. We propose BioVLM, a prompt‑learning framework that improves cross‑domain generalization without extensive backbone fine‑tuning. BioVLM learns a diverse prompt bank and introduces dynamic prompt selection: for each input, it selects the most discriminative prompts via a low‑entropy criterion on the predictive distribution, effectively coupling sparse few‑shot evidence with rich LLM semantic priors. To strengthen this coupling, we distill high‑confidence LLM‑derived attributes and enforce robust knowledge transfer through strong/weak augmentation consistency. At test time, BioVLM adapts by choosing modality‑appropriate prompts, enabling transfer to unseen categories and domains, while keeping training lightweight and inference efficient. On 11 MedMNIST+ 2D datasets, BioVLM achieves new state of the art across three distinct generalization settings. Codes are available at https://github.com/mainaksingha01/BioVLM.
Authors:Honglin Chen, Karran Pandey, Rundi Wu, Matheus Gadelha, Yannick Hold-Geoffroy, Ayush Tewari, Niloy J. Mitra, Changxi Zheng, Paul Guerrero
Abstract:
Kinematic rigs provide a structured interface for articulating 3D meshes but lack any associated pose space, i.e., an explicit representation of the plausible manifold of joint configurations for a given mesh. Without such a pose space, stochastic sampling or manual manipulation of raw rig parameters easily results in semantic and/or geometric violations, such as anatomical hyperextension and non‑physical self‑intersections. We propose Video‑informed Pose Spaces (ViPS), a feedforward framework that discovers the latent distribution of valid articulations for auto‑rigged meshes by distilling motion priors from a pretrained video diffusion model. Unlike existing methods that rely on scarce, artist‑authored 4D datasets, or focus on reconstructing instances of individual motions, ViPS transfers generative video model priors into a universal distribution over the given rig parameterization. Differentiable geometric validators applied to the skinned mesh enforce shape‑specific integrity without requiring manual regularizers. Our feedforward model reveals a smooth, compact, and controllable pose space. This, in turn, supports sampling for diverse shape variations, manifold projection for inverse kinematics, and temporally coherent trajectories for animation and keyframing. Further, the distilled 3D pose samples serve as semantic proxies to guide video diffusion, effectively closing the loop between generative 2D priors and structured 3D kinematic control. Our evaluations show that ViPS, trained solely using video priors, matches the performance of state‑of‑the‑art models trained on synthetic artist‑created 4D data in both plausibility and diversity. Additionally, as a universal model, ViPS exhibits robust zero‑shot generalization to out‑of‑distribution species and unseen skeletal topologies.
Authors:Fan Yang, Changsoo Jung, Ryosuke Kawamura, Hon Yung Wong
Abstract:
Multi‑camera systems are widely employed in sports to capture the 3D motion of athletes and equipment, yet calibrating their extrinsic parameters remains costly and labor‑intensive. We introduce an efficient, tool‑free method for multi‑camera extrinsic calibration tailored to sports involving stick‑like implements (e.g., golf clubs, bats, hockey sticks). Our approach jointly exploits two complementary cues from synchronized multi‑camera videos: (i) human body keypoints with unknown metric scale and (ii) a rigid stick‑like implement of known length. We formulate a three‑stage optimization pipeline that refines camera extrinsics, reconstructs human and stick trajectories, and resolves global scale via the stick‑length constraint. Our method achieves accurate extrinsic calibration without dedicated calibration tools. To benchmark this task, we present the first dataset for multi‑camera self‑calibration in stick‑based sports, consisting of synthetic sequences across four sports categories with 3 to 10 cameras. Comprehensive experiments demonstrate that our method delivers SOTA performance, achieving low rotation and translation errors. Our project page: https://fandulu.github.io/sport_stick_multi_cam_calib/.
Authors:Gaozhi Zhou, Hu He, Peng Shen, Jipeng Zhang, Liujue Zhang, Linrui Xu, Zeyuan Wang, Ziyu Li, Xuezhi Cui, Wang Guo, Haifeng Li
Abstract:
Reinforcement learning (RL) post‑training substantially improves remote sensing vision‑language models (RS‑VLMs). However, when handling complex remote sensing imagery (RSI) requiring exhaustive visual scanning, models tend to rely on localized salient cues for rapid inference. We term this RL‑induced bias "perceptual inertia". Driven by reward maximization, models favor quick outcome fitting, leading to two limitations: cognitively, overreliance on specific features impedes complete evidence construction; operationally, models struggle to flexibly shift visual focus across tasks. To address this bias and encourage comprehensive visual evidence mining, we propose RS‑HyRe‑R1, a hybrid reward framework for RSI understanding. It introduces: (1) a spatial reasoning activation reward that enforces structured visual reasoning; (2) a perception correctness reward that provides adaptive quality anchors across RS tasks, ensuring accurate geometric and semantic alignment; and (3) a visual‑semantic path evolution reward that penalizes repetitive reasoning and promotes exploration of complementary cues to build richer evidence chains. Experiments show RS‑HyRe‑R1 effectively mitigates "perceptual inertia", encouraging deeper, more diverse reasoning. With only 3B parameters, it achieves state‑of‑the‑art performance on REC, OVD, and VQA tasks, outperforming models up to 7B parameters. It also demonstrates strong zero‑shot generalization, surpassing the second‑best model by 3.16%, 3.97%, and 2.72% on VQA, OVD, and REC, respectively. Code and datasets are available at https://github.com/geox‑lab/RS‑HyRe‑R1.
Authors:Rongsheng Hu, Runwei Guan, Yicheng Di, Jiayu Bao, Yuan Liu
Abstract:
Manual annotation of high‑quality visual question answering with grounding (VQA‑G) datasets, which pair visual questions with evidential grounding, is crucial for advancing vision‑language models (VLMs), but remains unscalable. Existing automated methods are often hindered by two key issues: (1) inconsistent data fidelity due to model hallucinations; (2) brittle verification mechanisms based on simple heuristics. To address these limitations, we introduce AutoVQA‑G, a self‑improving agentic framework for automated VQA‑G annotation. AutoVQA‑G employs an iterative refinement loop where a Consistency Evaluation module uses Chain‑of‑Thought (CoT) reasoning for fine‑grained visual verification. Based on this feedback, a memory‑augmented Prompt Optimization agent analyzes critiques from failed samples to progressively refine generation prompts. Our experiments show that AutoVQA‑G generates VQA‑G datasets with superior visual grounding accuracy compared to leading multimodal LLMs, offering a promising approach for creating high‑fidelity data to facilitate more robust VLM training and evaluation. Code: https://github.com/rohnson1999/AutoVQA‑G
Authors:Qihao Shen, Jiaxing Xuan, Zhenguang Liu, Sifan Wu, Yutong Xie, Zhaoyan Ming, Yingying Jiao, kui Ren
Abstract:
Advanced deepfake technologies are blurring the lines between real and fake, presenting both revolutionary opportunities and alarming threats. While it unlocks novel applications in fields like entertainment and education, its malicious use has sparked urgent ethical and societal concerns ranging from identity theft to the dissemination of misinformation. To tackle these challenges, feature analysis using frequency features has emergedas a promising direction for deepfake detection. However, oneaspect that has been overlooked so far is that existing methodstend to concentrate on one or a few specific frequency domains,which risks overfitting to particular artifacts and significantlyundermines their robustness when facing diverse forgery patterns. Another underexplored aspect we observe is that different features often attend to the same forged region, resulting in redundant feature representations and limiting the diversity of the extracted clues. This may undermine the ability of a model to capture complementary information across different facets, thereby compromising its generalization capability to diverse manipulations. In this paper, we seek to tackle these challenges from two aspects: (1) we propose a triple‑branch network that jointly captures spatial and frequency features by learning from both original image and image reconstructed by different frequency channels, and (2) we mathematically derive feature decoupling and fusion losses grounded in the mutual information theory, which enhances the model to focus on task‑relevant features across the original image and the image reconstructed by different frequency channels. Extensive experiments on six large‑scale benchmark datasets demonstrate that our method consistently achieves state‑of‑the‑art performance. Our code is released at https://github.com/injooker/Unveiling Deepfake.
Authors:Peng Huang, Yifeng Chen, Zeyu Zhang, Hao Tang
Abstract:
Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D generation (e.g., synthesis, completion, and editing). However, these tasks are often tackled in isolation, resulting in fragmented architectures and representations that hinder knowledge transfer and holistic scene modeling. To address these challenges, we propose UniMesh, a unified framework that jointly learns 3D generation and understanding within a single architecture. First, we introduce a novel Mesh Head that acts as a cross model interface, bridging diffusion based image generation with implicit shape decoders. Second, we develop Chain of Mesh (CoM), a geometric instantiation of iterative reasoning that enables user driven semantic mesh editing through a closed loop latent, prompting, and re generation cycle. Third, we incorporate a self reflection mechanism based on an Actor Evaluator Self reflection triad to diagnose and correct failures in high level tasks like 3D captioning. Experimental results demonstrate that UniMesh not only achieves competitive performance on standard benchmarks but also unlocks novel capabilities in iterative editing and mutual enhancement between generation and understanding. Code: https://github.com/AIGeeksGroup/UniMesh. Website: https://aigeeksgroup.github.io/UniMesh.
Authors:Evren Çetinkaya, Sangmin Lee, Jung Uk Kim, Hong Joo Lee, Nassir Navab
Abstract:
Visual prompting has emerged as a powerful method for adapting pre‑trained models to new domains without updating model parameters. However, existing prompting methods typically optimize a single prompt per domain and apply it uniformly to all inputs, limiting their ability to generalize under intra and inter‑domain variability, which is especially critical in the medical field. To address this, we propose APEX, an Adaptive Prompt EXtraction framework that retrieves input‑specific prompts from a learnable prompt memory. The memory stores diverse, domain‑discriminative prompt representations and is queried via domain features extracted from the Fourier spectrum. To learn robust and discriminative domain features, we introduce a novel Low‑Frequency Feature Contrastive (LFC) learning framework that clusters representations from the same domain while separating those from different domains. Extensive experiments on two medical segmentation tasks demonstrate that APEX significantly improves generalization across both seen and unseen domains. Furthermore, it complements any existing backbones and consistently enhances performance, confirming its effectiveness as a plug‑and‑play prompting solution in medical fields. The code is available at https://github.com/cetinkayaevren/apex/
Authors:Liyang Wang, Zeyu Zhang, Hao Tang
Abstract:
Scene graph representations enable structured visual understanding by modeling objects and their relationships, and have been widely used for multiview and 3D scene reasoning. Existing methods such as MSG learn scene graph embeddings in Euclidean space using contrastive learning and attention based association. However, Euclidean geometry does not explicitly capture hierarchical entailment relationships between places and objects, limiting the structural consistency of learned representations. To address this, we propose Hyperbolic Scene Graph (HSG), which learns scene graph embeddings in hyperbolic space where hierarchical relationships are naturally encoded through geometric distance. Our results show that HSG improves hierarchical structure quality while maintaining strong retrieval performance. The largest gains are observed in graph level metrics: HSG achieves a PP IoU of 33.17 and the highest Graph IoU of 33.51, outperforming the best AoMSG variant (25.37) by 8.14, highlighting the effectiveness of hyperbolic representation learning for scene graph modeling. Code: https://github.com/AIGeeksGroup/HSG.
Authors:Marco Sánchez-Beeckman, Antoni Buades
Abstract:
Being one of the oldest and most basic problems in image processing, image denoising has seen a resurgence spurred by rapid advances in deep learning. Yet, most modern denoising architectures make limited use of the technical knowledge acquired researching the classical denoisers that came before the mainstream use of neural networks, instead relying on depth and large parameter counts. This poses a challenge not only for understanding the properties of such networks, but also for deploying them on real devices which may present resource constraints and diverse noise profiles. Tackling both issues, we propose an architecture dedicated to RAW‑to‑RAW denoising that incorporates the interpretable structure of classical self‑similarity‑based denoisers into a fully learnable neural network. Our design centers on a novel nonlocal block that parallels the established pipeline of neighbor matching, collaborative filtering and aggregation popularized by nonlocal patch‑based methods, operating on learned multiscale feature representations. This built‑in nonlocality efficiently expands the receptive field, sufficing a single block per scale with a moderate number of neighbors to obtain high‑quality results. Training the network on a curated dataset with clean real RAW data and modeled synthetic noise while conditioning it on a noise level map yields a sensor‑agnostic denoiser that generalizes effectively to unseen devices. Both quantitative and visual results on benchmarks and in‑the‑wild photographs position our method as a practical and interpretable solution for real‑world RAW denoising, achieving results competitive with state‑of‑the‑art convolutional and transformer‑based denoisers while using significantly fewer parameters. The code is available at https://github.com/MIA‑UIB/nonlocal‑matchfilter .
Authors:Yihong Yao, Chunlei Li, Canxuan Gang, Wenzhi Hu, Zeyu Zhang, Hao Zhang, Xiaoyan Li
Abstract:
Increasingly advanced data augmentation techniques have greatly aided clinical medical research, increasing data diversity and improving model generalization capabilities. Although most current basic models exhibit strong generalization abilities, image quality varies due to differences in equipment and operators. To address these challenges, we present SegTTA, a framework that improves medical image segmentation without model retraining by combining four augmentations (Gamma correction, Contrast enhancement, Gaussian blur, Gaussian noise) with weighted voting across multiple MedSAM2 checkpoints. Experiments demonstrate consistent improvements across three diverse datasets: healthy uterus segmentation, uterine myoma detection, and multi class hepatic structure segmentation. Ablation studies reveal that large organs benefit from intensity augmentations while small lesions require noise augmentations. The voting threshold controls the coverage precision trade off, enabling task specific optimization for different clinical requirements. Ultimately, on a multiclass hepatic vessel dataset, compared to MedSAM2 baselines, our method achieves an increase of 1.6 in mIoU and 1.9 in aIoU, along with a reduction of approximately 2.0 in HD95. Code will be available at https://github.com/AIGeeksGroup/SegTTA.
Authors:Alexander Saikia, Chiara Di Vece, Zhehua Mao, Sierra Bonilla, Chloe He, Joao Ramalhinho, Tobias Czempiel, Sophia Bano, Danail Stoyanov
Abstract:
Purpose: 3D reconstruction in minimally invasive surgery (MIS) enables enhanced surgical guidance through improved visualisation, tool tracking, and augmented reality. However, traditional RGB‑based keypoint detection and matching pipelines struggle with surgical challenges, such as poor texture and complex illumination. We investigate whether using snapshot hyperspectral imaging (HSI) can provide improved results on keypoint detection and matching surgical scenes. Methods: We developed HyKey, a HYperspectral KEYpoint detection and description model made up of a hybrid 3D‑2D convolutional neural network that jointly extracts spatial‑spectral features from HSI. The model was trained using synthetic homographic augmentation and epipolar geometry constraints on a robotically‑acquired dual‑camera RGB‑HSI laparoscopic dataset of ex‑vivo organs with calibrated camera poses. We benchmarked performance against established RGB‑based methods, including SuperPoint and ALIKE. Results: Our HSI‑based model outperformed RGB baselines on registered RGB frames, achieving 96.62% mean matching accuracy and 67.18% mean average accuracy at 10 degree on pose estimation, demonstrating consistent improvements across multiple evaluation metrics. Conclusion: Integrating spectral information from an HSI cube offers a promising approach for robust monocular 3D reconstruction in MIS, addressing limitations of texture‑poor surgical environments through enhanced spectral‑spatial feature discrimination. Our model and dataset are available at https://github.com/alexsaikia/HyKey‑Hyperspectral‑Keypoint‑Detection
Authors:Meng Zhang, Jinzhong Ning, Xiaolong Wu, Hongfei Lin, Yijia Zhang
Abstract:
Grounded Multimodal Named Entity Recognition (GMNER) aims to jointly identify named entity mentions in text, predict their semantic types, and ground each entity to a corresponding visual region in an associated image. Existing approaches predominantly adopt pipeline‑based architectures that decouple textual entity recognition and visual grounding, leading to error accumulation and suboptimal joint optimization. In this paper, we propose E2E‑GMNER, a fully end‑to‑end generative framework that unifies entity recognition, semantic typing, visual grounding, and implicit knowledge reasoning within a single multimodal large language model. We formulate GMNER as an instruction‑tuned conditional generation task and incorporate chain‑of‑thought reasoning to enable the model to adaptively determine when visual evidence or background knowledge is informative, reducing reliance on noisy cues. To further address the instability of generative bounding box prediction, we introduce Gaussian Risk‑Aware Box Perturbation (GRBP), which replaces hard box supervision with probabilistically perturbed soft targets to improve robustness against annotation noise and discretization errors. Extensive experiments on the Twitter‑GMNER and Twitter‑FMNERG benchmarks demonstrate that E2E‑GMNER achieves highly competitive performance compared with state of the art methods, validating the effectiveness of unified end‑to‑end optimization and noise‑aware grounding supervision. Code is available at:https://github.com/Finch‑coder/E2E‑GMNER
Authors:Enrui Yang, Yuezun Li
Abstract:
Detecting face forgeries using CLIP has recently emerged as a promising and increasingly popular research direction. Owing to its rich visual knowledge acquired through large‑scale pretraining, most existing methods typically rely on the visual encoder of CLIP, while paying limited attention to the text modality. Given the instructive nature of the text modality, we posit that it can be leveraged to instruct Deepfake detection with meticulous design. Accordingly, we shift the focus from the visual modality to the text modality and propose a new Separable Prompt Learning strategy (SePL) that enables CLIP to serve as an effective face forgery detector. The core idea of SePL is to disentangle forgery‑specific and forgery‑irrelevant information in images via two types of prompt learning, with the former enhancing detection. To achieve this disentangle, we describe a cross‑modality alignment strategy and a set of dedicated objectives. Extensive experiments demonstrate that, with this simple adaptation, our method achieves competitive and even superior performance compared to other methods under both cross‑dataset and cross‑method evaluation, highlighting its strong generalizability. The codes have been released at https://github.com/OUC‑YER/SePL‑DeepfakeDetection
Authors:Jiatong Li, Zheng Chen, Kai Liu, Jingkai Wang, Zihan Zhou, Xiaoyang Liu, Libo Zhu, Jue Gong, Radu Timofte, Yulun Zhang, Congyu Wang, Zihao Wang, Ke Wu, Xinzhe Zhu, Fengkai Zhang, Zhongbao Yang, Long Sun, Jiangxin Dong, Jinshan Pan, Jiachen Tu, Yaokun Shi, Guoyi Xu, Yaoxin Jiang, Jiajia Liu, Renyuan Situ, Yixin Yang, Zhaorun Zhou, Junyang Chen, Yuqi Li, Chuanguang Yang, Weilun Feng, Chuanyue Yan, Yuedong Tan, Yingli Tian, Zhenzhong Chen, Tongqi Guo, Ruhan Liu, Sangzi Shi, Huazhang Deng, Jie Yang, Wenzhuo Ma, Yuantong Zhang, Daiqin Yang, Tianrun Chen, Deyi Ji, Yuxiao Jiang, Qi Zhu, Lanyun Zhu, Yuwen Pan, Runze Tian, Mingyu Shi, Zhanfeng Feng, Yuanfei Bao, Jiaming Guo, Renjing Pei, Xin Di, Long Peng, Linfeng Jiang, Xueyang Fu, Yang Cao, Zhengjun Zha, Choulhyouc Lee, Shyang-En Weng, Yi-Cheng Liao, Jorge Tyrakowski, Yu-Syuan Xu, Wei-Chen Chiu, Ching-Chun Huang, Yoonjin Im, Jihye Park, Hyungju Chun, Hyunhee Park, MinKyu Park, Xiaoxuan Yu, Jianxing Zhang, Yuxuan Jiang, Chengxi Zeng, Tianhao Peng, Fan Zhang, David Bull, Watchara Ruangsang, Supavadee Aramvith, JiaHao Deng, Wei Zhou, Hongyu Huang, Shaohui Lin, Zihan Wang, Yilin Chen, Yunchen Li, Junbo Qiao, Wei Li, Jiao Xie, Gaoqi He, Wenxi Li
Abstract:
This paper provides a review of the NTIRE 2026 challenge on mobile real‑world image super‑resolution, highlighting the proposed solutions and the resulting outcomes. The challenge aims to recover high‑resolution (HR) images from low‑resolution (LR) counterparts generated through unknown degradations with a x4 scaling factor while ensuring the models remain executable on mobile devices. The objective is to develop effective and efficient network designs or solutions that achieve state‑of‑the‑art real‑world image super‑resolution performance. The track of the challenge evaluates performance using a weighted combination of image quality assessment (IQA) score and speedup ratios. The competition attracted 108 registrants, with 16 teams achieving a valid score in the final ranking. This collaborative effort advances the performance of mobile real‑world image super‑resolution while offering an in‑depth overview of the latest trends in the field.
Authors:Chunliang Li, Tianze Cao, Sanyuan Zhao
Abstract:
Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high‑resolution images. While existing methods accelerate inference by pruning tokens using frequency maps, their binary hard‑pruning approach is fundamentally limited and fails to improve quality even with better frequency estimation. Observing that VAR models possess significant depth redundancy, we propose a paradigm shift from pruning entire tokens to adaptively allocating per‑token computational depth. To this end, we introduce DepthVAR, a training‑free framework that dynamically allocates computation. It integrates an adaptive depth scheduler, which assigns computational depth via a cyclic rotated schedule for balanced, non‑static refinement, with a dynamic inference process that translates these depths into layer‑major masks, selectively applies transformer blocks, and blends the resulting codes to ensure each token's influence is proportional to its processing depth. Extensive experiments show that DepthVAR achieves 2.3×‑3.1× acceleration with minimal quality loss, offering a competitive compute‑performance trade‑off compared to existing hard‑pruning approaches. Code is available at https://github.com/STOVAGtz/DepthVAR
Authors:Yunkai Dang, Yifan Jiang, Yizhu Jiang, Anqi Chen, Wenbin Li, Yang Gao
Abstract:
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ensuring their reliability in practical deployment necessitates robust confidence estimation. Prior works have predominantly focused on text‑only LLMs, often relying on computationally expensive self‑consistency sampling. In this paper, we extend this to multimodal settings and conduct a comprehensive evaluation of MLLMs' response confidence estimation. Our analysis reveals a significant instinct‑reflection misalignment: the model's implicit token‑level support frequently diverges from its verbal self‑assessment confidence. To address this misalignment, we propose a monotone confidence fusion framework to merge dual‑channel signals and cross‑channel consistency to estimate correctness. Subsequently, an order‑preserving mean alignment step is applied to correct global bias, which improves calibration while preserving the risk‑coverage trade‑off for selective prediction. Experiments on diverse open‑source and closed‑source MLLMs show that our method consistently yields more reliable confidence estimates and improves both calibration and failure prediction. Code will be available at https://github.com/Yunkaidang/Instinct‑vs.‑Reflection.
Authors:Wenwei Xie, Jie Yin, Lu Ma, Xuansong Zhang, Wenjing Zhang
Abstract:
AI‑generated imagery has reached near‑photorealistic fidelity, yet this technology poses significant threats to information security and societal trust. Existing deepfake detection methods often exhibit limited robustness in open‑world scenarios. To address this limitation, this paper investigates intrinsic discrepancies between synthetic and authentic images from a signal‑level perspective. Our analysis reveals that low‑correlation signals serve as distinctive markers for differentiating AI‑generated imagery from real photographs. Building on this insight, we introduce a novel method for quantifying these signals based on fractal theory. By analyzing the fractal characteristics of low‑correlation signals, our method effectively captures the subtle statistical anomalies inherent to the synthesis process. Extensive experimental results demonstrate the method's robustness and superior detection performance. This work emphasizes the need to shift research focus to a new signal‑level direction for deepfake detection. Theoretically, this proposed approach is not limited to face image identification but can be applied to all AI‑generated image detection tasks. This study provides a new research direction for deepfake detection.
Authors:Si Li, Chen-Kai Hu, Zhenhuan Lyu, Yuanqing He
Abstract:
Digital subtraction angiography (DSA) in coronary imaging is fundamentally challenged by physiological motion, forcing reliance on raw angiograms cluttered with anatomical noise. Existing deep learning methods often produced images with two critical clinically unacceptable flaws: persistent boundary artifacts and a loss of native tissue grayscale fidelity that undermined diagnostic confidence. We propose a novel framework termed as CDSA‑Net that for the first time explicitly decouples and jointly optimizes vascular structure preservation and realistic background restoration. CDSA‑Net introduces two core innovations: (i) A hierarchical geometric prior guidance (HGPG) mechanism, embedded in our coronary structure extraction network (CSENet). It synergistically combines integrated geometric prior (IGP) with gated spatial modulation (GSM) and centerline‑aware topology (CAT) loss supervision, ensuring structural continuity. (ii) An adaptive noise module (ANM) within our coronary background restoration network (CBResNet). Unlike standard restoration, ANM uniquely models the stochastic nature of clinical X‑ray noise, bridging the domain gap to enable seamless background intensity estimation and the complete elimination of boundary artifacts. The final subtraction is obtained by removing the restored background from the raw angiogram. Quantitatively, it significantly outperformed state‑of‑the‑art methods in vascular intensity correlation and perceptual quality. A 25.6% improvement in morphology assessment efficiency and a 42.9% gain in hemodynamic evaluation speed set a new benchmark for utility in interventional cardiology, while maintaining diagnostic results consistent with raw angiograms. The project code is available at https://github.com/DrThink‑ai/CDSA‑Net.
Authors:Davie Chen
Abstract:
We present SciDraw‑6K, a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image‑generation models, each paired with prompts in eleven languages (English, Simplified Chinese, Traditional Chinese, Japanese, Korean, German, French, Spanish, Brazilian Portuguese, Italian, and Russian). Images span eight broad scientific categories ‑‑ biomedical, chemistry, materials, electronics, environment, AI systems, physics, and a long "other" tail ‑‑ and are produced primarily by the gemini‑2.5‑flash‑image and gemini‑3‑pro‑image‑preview model families. In contrast to general‑purpose text‑to‑image corpora that dominate the literature, SciDraw‑6K is purpose‑built for the scientific illustration genre: schematic diagrams, mechanism figures, table‑of‑contents graphics, and conceptual posters. We describe the construction pipeline, report dataset statistics, and document its use as the substrate of sci‑draw.com, a public scientific drawing service. The dataset is released to support multilingual text‑to‑image research, domain‑adapted diffusion fine‑tuning, and prompt‑engineering studies for scientific visualization. Dataset: https://huggingface.co/datasets/SciDrawAI/SciDraw‑6K Code: https://github.com/SciDrawAI/scidraw‑6k
Authors:Bo Ma, Weiqi Yan, Jinsong Wu
Abstract:
We propose PPEDCRF, a calibrated selective perturbation framework that protects \emphbackground‑based location privacy in released video frames against gallery‑based retrieval attackers. Even after GPS metadata are stripped, an adversary can geolocate a frame by matching its background visual cues to geo‑tagged reference imagery; PPEDCRF mitigates this threat by estimating location‑sensitive background regions with a dynamic conditional random field (DCRF), rescaling perturbation strength with a normalized control penalty (NCP), and injecting Gaussian noise only inside the inferred regions via a DP‑style calibration rule.
On a controlled paired‑scene retrieval benchmark with eight attacker backbones and three noise seeds, PPEDCRF reduces ResNet18 Top‑1 retrieval accuracy from 0.667 to 0.361\pm0.127 at σ_0=8 while preserving 36.14\,dB PSNR ‑‑ an \approx6\,dB quality advantage over global Gaussian noise. Transfer across the eight‑backbone seed‑averaged benchmark is broadly supportive (23 of 24 backbone‑gallery cells show negative Δ), while appendix‑scale confirmation identifies MixVPR as a remaining adverse‑transfer exception. Matched‑operating‑point analysis shows that PPEDCRF and global Gaussian noise converge in Top‑1 privacy at equal utility, so the practical benefit is spatially concentrated perturbation that preserves higher visual quality at any given noise scale rather than stronger matched‑utility privacy. Code: https://github.com/mabo1215/PPEDCRF
Authors:Zedong Dan, Zijie Wang, Wei Zhang, Xiangru Lin, Weiming Zhang, Xiao Tan, Jingdong Wang, Liang Lin, Guanbin Li
Abstract:
Offline vectorized maps constitute critical infrastructure for high‑precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego‑vehicle trajectories, which fundamentally suffer from viewpoint insufficiency: while memory‑based methods extend observation time by aggregating ego‑trajectory frames, they lack the spatial diversity needed to reveal occluded regions. Incorporating views from surrounding vehicles offers complementary perspectives, yet naive fusion introduces three key challenges: computational cost from large candidate pools, redundancy from near‑collinear viewpoints, and noise from pose errors and occlusion artifacts.
We present OptiMVMap, which reformulates multi‑vehicle mapping as a select‑then‑fuse problem to address these challenges systematically. An Optimal Vehicle Selection (OVS) module strategically identifies a compact subset of helpers that maximally reduce ego‑centric uncertainty in occluded regions, addressing computation and redundancy challenges. Cross‑Vehicle Attention (CVA) and Semantic‑aware Noise Filter (SNF) then perform pose‑tolerant alignment and artifact suppression before BEV‑level fusion, addressing the noise challenge. This targeted pipeline yields more complete and topologically faithful maps with substantially fewer views than indiscriminate aggregation. On nuScenes and Argoverse2, OptiMVMap improves MapTRv2 by +10.5 mAP and +9.3 mAP, respectively, and surpasses memory‑augmented baselines MVMap and HRMapNet by +6.2 mAP and +3.8 mAP on nuScenes. These results demonstrate that uncertainty‑guided selection of helper vehicles is essential for efficient and accurate multi‑vehicle vectorized mapping. The code is released at https://github.com/DanZeDong/OptiMVMap.
Authors:Zihao Zhao, Frederik Hauke, Juliana De Castilhos, Mathis Bode, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn
Abstract:
Developing AI models that are useful in clinical practice, requires efficient collaboration between clinicians and AI developers. This poses a practical challenge: clinicians must repeatedly communicate and refine their requirements with AI developers before those requirements can be translated into executable model development. This iterative process is time‑consuming, and even after repeated discussion, misalignment may still exist because the two sides do not fully share each other's expertise. Coding agents may help close this gap. They can write and refine code on their own, and they carry working knowledge of both medicine and AI to understand commands formulated by both medical experts and developers. We present a prototype that lets clinicians drive AI development directly. A clinician describes the task in plain language, and the system turns the description into a working pipeline, refines it through repeated experiments together with the clinician, and returns a model that meets the stated clinical objective. Across five clinical tasks, the system reliably produces models that matched the clinician's request and reached competitive performance. Most notably, on chest radiographs the system sharply reduced the model's reliance on chest drains, a well‑known shortcut for pneumothorax classification, from 60% to 31% on one dataset and from 50% to 18% on another. Our results suggest that coding agents can shift clinical AI development toward a more clinician‑driven mode, allowing domain experts to shape models directly instead of relaying requirements through specialized AI teams.
Authors:Jidong Kuang, Hongsong Wang, Jie Gui
Abstract:
Human action recognition and motion generation are two active research problems in human‑centric computer vision, both aiming to align motion with textual semantics. However, most existing works study these two problems separately, without uncovering the links between them, namely that motion generation requires semantic comprehension. This work investigates unified action recognition and motion generation by leveraging skeleton coordinates for both motion understanding and generation. We propose Coordinates‑based Autoregressive Motion Diffusion (CoAMD), which synthesizes motion in a coarse‑to‑fine manner. As a core component of CoAMD, we design a Multi‑modal Action Recognizer (MAR) that provides gradient‑based semantic guidance for motion generation. Furthermore, we establish a rigorous benchmark by evaluating baselines on absolute coordinates. Our model can be applied to four important tasks, including skeleton‑based action recognition, text‑to‑motion generation, text‑motion retrieval, and motion editing. Extensive experiments on 13 benchmarks across these tasks demonstrate that our approach achieves state‑of‑the‑art performance, highlighting its effectiveness and versatility for human motion modeling. Code is available at https://github.com/jidongkuang/CoAMD.
Authors:Xingyuan Yu, Yijin Li, Chong Zeng, Yuhang Ming, Hujun Bao, Guofeng Zhang
Abstract:
Capturing both geometry and rigid motion for structured dynamic objects, like multi‑part assemblies or jointed mechanisms, remains a key challenge. Existing dynamic methods, such as deformable meshes or 3DGS, rely on unstructured representations and fail to jointly model suitable geometry and articulated motion. Primitive‑based methods excel at structured static scenes, but their dynamic potential is still unexplored. We propose D‑Prism, the first framework to achieve high‑fidelity structured dynamic modeling by extending differentiable primitives to the dynamic domain. Specifically, we bind 3DGS to primitive surfaces, leveraging their respective strengths in appearance and geometry. We introduce a deformation network to control primitive motion, ensuring it accurately matches the object's movement. Furthermore, we design a novel adaptive control strategy to dynamically adjust primitive counts, better matching objects' true spatial footprint. Experiments confirm that our method excels at structured dynamic modeling, providing both structured geometry and precise motion tracking.
Authors:Kyeong Seon Kim, Baek Seong-Eun, Lee Jung-Mok, Tae-Hyun Oh
Abstract:
Scalable Vector Graphics (SVGs) function both as visual images and as structured code that encode rich geometric and layout information, yet most methods rasterize them and discard this symbolic organization. At the same time, recent sentence embedding methods produce strong text representations but do not naturally extend to visual or structured modalities. We propose a training‑free, instruction‑guided multimodal embedding framework that uses a Multimodal Large Language Model (MLLM) to map text, raster images, and SVG code into an aligned embedding space. We control the direction of embeddings through modality‑specific instructions and structural SVG cues, eliminating the need for learned projection heads or contrastive training. Our method has two key components: (1) Multimodal Explicit One‑word Limitation (mEOL), which instructs the MLLM to summarize any multimodal input into a single token whose hidden state serves as a compact semantic embedding. (2) A semantic SVG rewriting module that assigns meaningful identifiers and simplifies nested SVG elements through visual reasoning over the rendered image, exposing geometric and relational cues hidden in raw code. Using a repurposed VGBench, we build the first text‑to‑SVG retrieval benchmark and show that our training‑free embeddings outperform encoder‑based and training‑based multimodal baselines. These results highlight prompt‑level control as an effective alternative to parameter‑level training for structure‑aware multimodal retrieval. Project page: https://scene‑the‑ella.github.io/meol/
Authors:Zhijia Liang, Jiaming Li, Weikai Chen, Yanhao Zhang, Haonan Lu, Guanbin Li
Abstract:
Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal is like an oasis‑small, critical, and easily lost in a desert of redundancy. Enlarging memory only widens the desert; aggressive compression dries up the oasis. The real difficulty lies in discovering where to look, not how much to remember. We therefore introduce OASIS, a novel framework for streaming video reasoning that tackles this challenge through structured, on‑demand retrieval. It organizes streaming history into hierarchical events and performs reasoning as controlled refinement‑short‑context inference first, followed by semantically grounded retrieval only when uncertainty arises. As the retrieval is driven by high‑level intent rather than embedding similarity, the retrieved memory is substantially more accurate and less noisy. Additionally, the mechanism is plug‑and‑play, training‑free, and readily attaches to different streaming MLLM backbones. Experiments across multiple benchmarks and backbones show that OASIS achieves strong gains in long‑horizon accuracy and compositional reasoning with bounded token cost and low request delay. Code is available at https://github.com/Solus‑sano/OASIS.
Authors:Mehmet Kerem Turkcan
Abstract:
Collisions between cyclists and pedestrians at urban intersections remain a persistent source of injuries, yet few systems attempt real‑time warnings to unequipped road users using commodity hardware. We present a prototype collision warning system that runs on a single edge device with a wide‑angle fisheye camera, producing audible and visual alerts at 30\,fps. The system makes four contributions. First, we develop a calibration pipeline for ultra‑wide fisheye lenses that overcomes corner‑detection failure and optimizer divergence through perspective remapping and direct bundle adjustment. Second, we combine fisheye‑aware object detection with a closed‑form ground‑plane projection via a precomputed lookup table. Third, we introduce a design‑time conformance simulation with 24 scripted hazard scenarios, stochastic size‑aware detection failures, and a latency sweep showing that a first‑order kinematic predictor maintains the mean warning budget above the distracted‑pedestrian reaction time across realistic camera latencies. Fourth, we formalize the decision layer as a separable, auditable testbench with explicit deployment gates, contestability mechanisms, and a residual risk register. Under conformance testing with fisheye localization error, the selected pipeline configuration achieves 93.3% sensitivity and 92.3% specificity, with a mean warning budget of 3.3\,s. The system design was informed by community‑aided design workshops. Code and replication scripts are available at https://github.com/mkturkcan/bikeped.
Authors:Yifei Zhao, Qian Lou, Mengxin Zheng
Abstract:
The public accessibility of large vision‑language models (LVLMs) raises serious concerns about unauthorized model reuse and intellectual property infringement. Existing ownership verification methods often rely on semantically abnormal queries or out‑of‑distribution responses as fingerprints, which can be easily detected and removed by adversaries. We expose this vulnerability through a Semantic Divergence Attack (SDA), which identifies and filters fingerprint queries by measuring semantic divergence between a suspect model and a reference model, showing that existing fingerprints are not semantic‑preserving and are therefore easy to detect and bypass. To address these limitations, we propose SIF (Semantically In‑Distribution Fingerprints), a non‑intrusive ownership verification framework that requires no parameter modification. SIF introduces Semantic‑Aligned Fingerprint Distillation (SAFD), which transfers text watermarking signals into the visual modality to produce semantically coherent yet fingerprinted responses. In addition, Robust‑Fingerprint Optimization (RFO) enhances robustness by simulating worst‑case representation perturbations, making the fingerprints resilient to model modifications such as fine‑tuning and quantization. Extensive experiments on LLaVA‑1.5 and Qwen2.5‑VL demonstrate that SIF achieves strong stealthiness and robustness, providing a practical solution for LVLM copyright protection. Code is available at https://github.com/UCF‑ML‑Research/SIF‑VLM‑Fingerprint
Authors:Jidong Kuang, Hongsong Wang, Jie Gui
Abstract:
With the development of robotics, skeleton‑based action recognition has become increasingly important, as human‑robot interaction requires understanding the actions of humans and humanoid robots. Due to different sources of human skeletons and structures of humanoid robots, skeleton data naturally exhibit heterogeneity. However, previous works overlook the data heterogeneity of skeletons and solely construct models using homogeneous skeletons. Moreover, open‑vocabulary action recognition is also essential for real‑world applications. To this end, this work studies the challenging problem of heterogeneous skeleton‑based action recognition with open vocabularies. We construct a large‑scale Heterogeneous Open‑Vocabulary (HOV) Skeleton dataset by integrating and refining multiple representative large‑scale skeleton‑based action datasets. To address universal skeleton‑based action recognition, we propose a Transformer‑based model that comprises three key components: unified skeleton representation, motion encoder for skeletons, and multi‑grained motion‑text alignment. The motion encoder feeds multi‑modal skeleton embeddings into a two‑stream Transformer‑based encoder to learn spatio‑temporal action representations, which are then mapped to a semantic space to align with text embeddings. Multi‑grained motion‑text alignment incorporates contrastive learning at three levels: global instance alignment, stream‑specific alignment, and fine‑grained alignment. Extensive experiments on popular benchmarks with heterogeneous skeleton data demonstrate both the effectiveness and the generalization ability of the proposed method. Code is available at https://github.com/jidongkuang/Universal‑Skeleton.
Authors:Yiting Wang, Nolwenn Peyratout, Tim Brodermann, Jiahui Wang, Yusi Cao, Michele Cazzola, Elie Tarassov, Takuya Kobayashi, Abderrahim Kasmi, Guillaume Allibert, Cédric Demonceaux, Valentina Donzella, Kurt Debattista, Radu Timofte, Zongwei Wu, Christos Sakaridis
Abstract:
This paper presents the report of the URVIS 2026 challenge on adverse‑to‑extreme panoptic segmentation. As the first challenge of its kind, it attracted 17 registered participants and 47 submissions, with 4 teams reaching the final phase. The challenge is based on the MUSES dataset, a multi‑sensor benchmark for panoptic segmentation in adverse‑to‑extreme weather, including RGB frame camera, LiDAR, radar, and event camera data. Weighted Panoptic Quality (wPQ) is designed and adopted as the official ranking metric for fair evaluation across weather conditions. In this report, we summarise the challenge setting and benchmark results, analyse the performance of the submitted methods, and discuss current progress and remaining challenges for robust multimodal panoptic segmentation. Link: https://urvis‑workshop.github.io/challenge‑Muses.html
Authors:Linyue Zhang, Wenyi Zeng, Zicheng Pan, Yongsheng Gao, Changming Sun, Jun Hu, Lixian Liu, Weichuan Zhang, Tuo Wang
Abstract:
Feature reconstruction techniques are widely applied for few‑shot fine‑grained image classification (FSFGIC). Our research indicates that one of the main challenges facing existing feature‑based FSFGIC methods is how to choose the size of the receptive field to extract feature descriptors (including spatial and frequency feature descriptors) from different category input images, thereby better performing the FSFGIC tasks. To address this, an adaptive receptive field‑based spatial‑frequency feature reconstruction network (ARF‑SFR‑Net) is proposed. The designed ARF‑SFR‑Net has the capability to adaptively determine receptive field sizes for obtaining spatial and frequency features, and effectively fuse them for reconstruction and FSFGIC tasks. The designed ARF‑SFR‑Net can be easily embedded into a given episodic training mechanism for end‑to‑end training from scratch. Extensive experiments on multiple FSFGIC benchmarks demonstrate the effectiveness and superiority of the proposed ARF‑SFR‑Net over state‑of‑the‑art approaches. The code is available at: https://github.com/ICL‑SUST/ARF‑SFR‑Net.git.
Authors:Yingzhi Xia, Setthakorn Tanomkiattikun, Liangli Zhen, Zaiwang Gu
Abstract:
Diffusion models (DMs) have recently shown remarkable performance on inverse problems (IPs). Optimization‑based methods can fast solve IPs using DMs as powerful regularizers, but they are susceptible to local minima and noise overfitting. Although DMs can provide strong priors for Bayesian approaches, enforcing measurement consistency during the denoising process leads to manifold infeasibility issues. We propose Noise‑space Hamiltonian Monte Carlo (N‑HMC), a posterior sampling method that treats reverse diffusion as a deterministic mapping from initial noise to clean images. N‑HMC enables comprehensive exploration of the solution space, avoiding local optima. By moving inference entirely into the initial‑noise space, N‑HMC keeps proposals on the learned data manifold. We provide a comprehensive theoretical analysis of our approach and extend the framework to a noise‑adaptive variant (NA‑NHMC) that effectively handles IPs with unknown noise type and level. Extensive experiments across four linear and three nonlinear inverse problems demonstrate that NA‑NHMC achieves superior reconstruction quality with robust performance across different hyperparameters and initializations, significantly outperforming recent state‑of‑the‑art methods. The code is available at https://github.com/NA‑HMC/NA‑HMC.
Authors:Chen Ma, Yunshu Li, Junhu Fu, Shuyu Liang, Yuanyuan Wang, Yi Guo
Abstract:
Clinical ultrasound analysis demands models that generalize across heterogeneous organs, views, and devices, while supporting interpretable workflow‑level analysis. Existing methods often rely on task‑wise adaptation, and joint learning may be unstable due to cross‑task interference, making it hard to deliver workflow‑level outputs in practice. To address these challenges, we present USTri, a tri‑stage ultrasound intelligence pipeline for unified multi‑organ, multi‑task analysis. Stage I trains a universal generalist USGen on different domains to learn broad, transferable priors that are robust to device and protocol variability. To better handle domain shifts and reach task‑aligned performance while preserving ultrasound shared knowledge, Stage II builds USpec by keeping USGen frozen and finetuning dataset‑specific heads. Stage III introduces USAgent, which mimics clinician workflows by orchestrating USpec specialists for multi‑step inference and deterministic structured reports. On the FMC\_UIA validation set, our model achieves the best overall performance across 4 task types and 27 datasets, outperforming state‑of‑the‑art methods. Moreover, qualitative results show that USAgent produces clinically structured reports with high accuracy and interpretability. Our study suggests a scalable path to ultrasound intelligence that generalizes across heterogeneous ultrasound tasks and supports consistent end‑to‑end clinical workflows. The code is publicly available at: https://github.com/MacDunno/USTri.
Authors:Antonios Kritikos, Nikolaos Spanos, Athanasios Voulodimos
Abstract:
Domain generalization (DG) aims to maintain performance under domain shift, which in computer vision appears primarily as stylistic variations that cause models to overfit to domain‑specific appearance cues rather than class semantics. To overcome this, recent methods use textual representations as stable, domain‑invariant anchors. However, multimodal approaches that rely on cosine similarity‑based contrastive alignment leave a modality gap where image and text embeddings remain geometrically separated despite semantic correspondence. We propose CrossFlowDG, a novel DG framework that addresses this residual gap using noise‑free, cross‑modal flow matching. By learning a continuous transformation in the joint Euclidean latent space, our framework explicitly transports domain‑biased image embeddings toward domain‑invariant text embeddings of the correct class. Using the efficient VMamba image encoder and CLIP's text encoder, CrossFlowDG is tested against four common DG benchmarks, and achieves competitive performance on several benchmarks and state‑of‑the‑art on TerraIncognita. Code is available at: https://github.com/ajkrit/CrossFlowDG
Authors:Zahid Hasan, Masud Ahmed, Nirmalya Roy
Abstract:
Semantic segmentation in hyperbolic space enables compact modeling of hierarchical structure while providing inherent uncertainty quantification. Prior approaches predominantly rely on the Poincaré ball model, which suffers from numerical instability, optimization, and computational challenges. We propose a novel, tractable, architecture‑agnostic semantic segmentation framework (pixel‑wise and mask classification) in the hyperbolic Lorentz model. We employ text embeddings with semantic and visual cues to guide hierarchical pixel‑level representations in Lorentz space. This enables stable and efficient optimization without requiring a Riemannian optimizer, and easily integrates with existing Euclidean architectures. Beyond segmentation, our approach yields free uncertainty estimation, confidence map, boundary delineation, hierarchical and text‑based retrieval, and zero‑shot performance, reaching generalized flatter minima. We introduce a novel uncertainty and confidence indicator in Lorentz cone embeddings. Further, we provide analytical and empirical insights into Lorentz optimization via gradient analysis. Extensive experiments on ADE20K, COCO‑Stuff‑164k, Pascal‑VOC, and Cityscapes, utilizing state‑of‑the‑art per‑pixel classification models (DeepLabV3 and SegFormer) and mask classification models (mask2former and maskformer), validate the effectiveness and generality of our approach. Our results demonstrate the potential of hyperbolic Lorentz embeddings for robust and uncertainty‑aware semantic segmentation. Code is available at https://github.com/mxahan/Lorentz_semantic_segmentation.
Authors:Seungjin Kim, Reza Jafarpourmarzouni, Christopher Neff, Hamed Tabkhi, Vinit Katariya
Abstract:
Vehicle trajectory prediction is central to highway perception, but deployment on roadside edge devices necessitates bounded, deterministic end‑to‑end latency. We present EdgeVTP, an embedded‑first trajectory predictor that combines interaction‑aware graph modeling with a lightweight transformer backbone and a one‑shot curve decoder. By predicting future motion as compact curve parameters (anchored at the last observed position) rather than horizon‑scaled autoregressive waypoints, EdgeVTP reduces decoding overhead while producing smooth trajectories. To keep runtime predictable in crowded scenes, we explicitly bound interaction complexity via a locality graph with a hard neighbor cap. Across three highway benchmarks and two Jetson‑class platforms, EdgeVTP achieves the lowest measured end‑to‑end latency under a protocol that includes graph construction and post‑processing, while attaining state‑of‑the‑art (SotA) prediction accuracy on two of the three datasets and competitive error on other benchmarks. Our code is available at https://github.com/SeungjinStevenKim/EdgeVTP.
Authors:Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi
Abstract:
Imitation learning has enabled robots to acquire complex visuomotor manipulation skills from demonstrations, but deployment failures remain a major obstacle, especially for long‑horizon action‑chunked policies. Once execution drifts off the demonstration manifold, these policies often continue producing locally plausible actions without recovering from the failure. Existing runtime monitors either require failure data, over‑trigger under benign feature drift, or stop at failure detection without providing a recovery mechanism. We present Rewind‑IL, a training‑free online safeguard framework for generative action‑chunked imitation policies. Rewind‑IL combines a zero‑shot failure detector based on Temporal Inter‑chunk Discrepancy Estimate (TIDE), calibrated with split conformal prediction, with a state‑respawning mechanism that returns the robot to a semantically verified safe intermediate state. Offline, a vision‑language model identifies recovery checkpoints in demonstrations, and the frozen policy encoder is used to construct a compact checkpoint feature database. Online, Rewind‑IL monitors self‑consistency in overlapping action chunks, tracks similarity to the checkpoint library, and, upon failure, rewinds execution to the latest verified safe state before restarting inference from a clean policy state. Experiments on real‑world and simulated long‑horizon manipulation tasks, including transfer to flow‑matching action‑chunked policies, demonstrate that policy‑internal consistency coupled with semantically grounded respawning offers a practical route to improved reliability in imitation learning. Supplemental materials are available at https://sjay05.github.io/rewind‑il
Authors:Henry O. Velesaca, David Freire-Obregon, Abel Reyes-Angulo, Steven Araujo, Angel Sappa
Abstract:
Penalty kicks in soccer are decided under extreme time constraints, where goalkeepers benefit from anticipating shot direction from the kickers motion before or around ball contact. In this paper, MambaKick is presented as a learning‑based framework for penalty direction prediction that leverages pretrained human action recognition (HAR) embeddings extracted from contact‑centered short video segments and combines them with a lightweight temporal predictor. Rather than relying on explicit kinematic reconstruction or handcrafted biomechanical features, the approach reuses transferable spatiotemporal representations and utilizes selective state‑spare models (Mamba) for efficient sequence aggregation. Simple contextual metadata (e.g., field side and footedness) are also considered as complementary cues that may reduce ambiguity in real‑world footage. Across a range of HAR backbones, MambaKick consistently improves or matches strong embedding baselines, achieving up to 53.1% accuracy for three classes and 64.5% for two classes under the proposed methodology. Overall, the results indicate that combining pretrained HAR representations with efficient state‑space temporal modeling is a practical direction for low‑latency intention prediction in real‑world sports video. The code will be available at GitHub: https://github.com/hvelesaca/MambaKick/
Authors:Henry O. Velesaca, Andrea Mero, Guillermo A. Castillo, Angel D. Sappa
Abstract:
Pedestrian detection is fundamental to autonomous driving, robotics, and surveillance. Despite progress in deep learning, reliable identification remains challenging due to occlusions, cluttered backgrounds, and degraded visibility. While multispectral detection‑combining visible and thermal sensors‑mitigates poor visibility, the challenge of camouflaged pedestrians remains largely unexplored. Existing Camouflaged Object Detection (COD) benchmarks focus on biological species, leaving a gap in safety‑critical human detection where targets blend into their surroundings. To address this, we introduce Camo‑M3FD (derived from the M3FD dataset), a novel benchmark for cross‑spectral camouflaged pedestrian detection, consisting of registered visible‑thermal image pairs. The dataset is curated using quantitative metrics to ensure high foreground‑background similarity. We provide high‑quality pixel‑level masks and establish a standardized evaluation framework using state‑of‑the‑art COD models. Our results demonstrate that while thermal signals provide indispensable localization cues, multispectral fusion is essential for refining structural details. Camo‑M3FD serves as a foundational resource for developing robust and safety‑critical detection systems. The dataset is available on GitHub: https://cod‑espol.github.io/Camo‑M3FD/
Authors:Xiangkai Wang, Yun Zhao, Dongyi He, Qingling Xia, Gen Li, Nizhuan Wang, Ningxiao Peng, Bin Jiang
Abstract:
Stroke patient cross‑subject electroencephalography (EEG) decoding of motor imagery (MI) brain‑computer interface (BCI) is essential for motor rehabilitation, yet lesion‑related abnormal temporal dynamics and pronounced inter‑patient heterogeneity often undermine generalization. Existing adaptation methods are easily misled by pathological slow‑wave activity and unstable target‑domain pseudo‑labels. To address this challenge, we propose PA‑TCNet, a pathology‑aware temporal calibration framework with physiology‑guided target refinement for stroke motor imagery decoding. PA‑TCNet integrates two coordinated components. The Pathology‑aware Rhythmic State Mamba (PRSM) module decomposes EEG spatiotemporal features into slowly varying rhythmic context and fast transient perturbations, injecting the fused pathological context into selective state propagation to more effectively capture abnormal temporal dynamics. The Physiology‑Guided Target Calibration (PGTC) module constructs source‑domain sensorimotor region‑of‑interest templates, imposing physiological consistency constraints and dynamically refining target‑domain pseudo‑labels, thereby improving adaptation reliability. Leave‑one‑subject‑out experiments on two independent stroke EEG datasets, XW‑Stroke and 2019‑Stroke, yielded mean accuracies of 66.56% and 72.75%, respectively, outperforming state‑of‑the‑art baselines. These results indicate that jointly modeling pathological temporal dynamics and physiology‑constrained pseudo‑supervision can provide more robust cross‑subject initialization for personalized post‑stroke MI‑BCI rehabilitation. The implemented code is available at https://github.com/wxk1224/PA‑TCNet.
Authors:Bo Gao, Chang Liu, Yuyang Miao, Siyuan Ma, Ser-Nam Lim
Abstract:
Recent advancements in Large Generative Models (LGMs) have revolutionized multi‑modal generation. However, generating illustrated storybooks remains an open challenge, where prior works mainly decompose this task into separate stages, and thus, holistic multi‑modal grounding remains limited. Besides, while safety alignment is studied for text‑ or image‑only generation, existing works rarely integrate child‑specific safety constraints into narrative planning and sequence‑level multi‑modal verification. To address these limitations, we propose BookAgent, a safety‑aware multi‑agent collaboration framework designed for high‑quality, safety‑aware visual narratives. Different from prior story visualization models that assume a fixed storyline sequence, BookAgent targets end‑to‑end storybook synthesis from a user draft by jointly planning, scripting, illustrating, and globally repairing inconsistencies. To ensure precise multi‑modal grounding, BookAgent dynamically calibrates page‑level alignment between textual scripts and visual layouts. Furthermore, BookAgent calibrates holistic consistency from the temporal dimension, by verifying‑then‑rectifying global inconsistencies in character identity and storytelling logic. Extensive experiments demonstrate that BookAgent significantly outperforms current methods in narrative coherence, visual consistency, and safety compliance, offering a robust paradigm for reliable agents in complex multi‑modal creation. The implementation will be publicly released at https://github.com/bogao‑code/BookAgent/tree/main.
Authors:Baoyou Chen, Hanchen Xia, Peng Tu, Haojun Shi, Liwei Zhang, Weihao Yuan, Siyu Zhu
Abstract:
Autoregressive vision‑language models (VLMs) deliver strong multimodal capability, but their token‑by‑token decoding imposes a fundamental inference bottleneck. Diffusion VLMs offer a more parallel decoding paradigm, yet directly converting a pretrained autoregressive VLM into a large‑block diffusion VLM (dVLM) often leads to substantial quality degradation. In this work, we present BARD, a simple and effective bridging framework that converts a pretrained autoregressive VLM into a same‑architecture, decoding‑efficient dVLM. Our approach combines progressive supervised block merging, which gradually enlarges the decoding block size, with stage‑wise intra‑dVLM distillation from a fixed small‑block diffusion anchor to recover performance lost at larger blocks. We further incorporate a mixed noise scheduler to improve robustness and token revision during denoising, and memory‑friendly training to enable efficient training on long multimodal sequences. A key empirical finding is that direct autoregressive‑to‑diffusion distillation is poorly aligned and can even hurt performance, whereas distillation within the diffusion regime is consistently effective. Experimental results show that, with \leq 4.4M data, BARD‑VL transfers strong multimodal capability from Qwen3‑VL to a large‑block dVLM. Remarkably, BARD‑VL establishes a new SOTA among comparable‑scale open dVLMs on our evaluation suite at both 4B and 8B scales. At the same time, BARD‑VL achieves up to 3× decoding throughput speedup compared to the source model. Code is available at https://github.com/fudan‑generative‑vision/Bard‑VL.
Authors:Zonghai Yao, Benlu Wang, Yifan Zhang, Junda Wang, Iris Xia, Zhipeng Tang, Shuo Han, Feiyun Ouyang, Zhichao Yang, Arman Cohan, Hong Yu
Abstract:
Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a single view. We introduce MedThinkVQA, an expert‑annotated benchmark for thinking with multiple images, where models must interpret each image, combine cross‑view evidence, and answer diagnostic questions with intermediate supervision and step‑level evaluation. The dataset contains 8,067 cases, including 720 test cases, with an average of 6.62 images per case, substantially denser than prior work, whose expert‑level benchmarks use at most 1.43 images per case. On the test set, the best closed‑source models, Claude‑4.6‑Opus, Gemini‑3‑Pro, and GPT‑5.2‑xhigh, reach only 57.2%, 55.3%, and 54.9% accuracy, while GPT‑5‑mini and GPT‑5‑nano reach 39.7% and 30.8%. Strong open‑source models lag behind, led by Qwen3.5‑397B‑A17B at 52.2% and Qwen3.5‑27B at 50.6%. Further analysis identifies grounded multi‑image reasoning as the main bottleneck: models often fail to extract, align, and compose evidence across views before higher‑level inference can help. Providing expert single‑image cues and cross‑image summaries improves performance, whereas replacing them with self‑generated intermediates reduces accuracy. Step‑level analysis shows that over 70% of errors arise from image reading and cross‑view integration. Scaling results further show that additional inference‑time computation helps only when visual grounding is already reliable; when early evidence extraction is weak, longer reasoning yields limited or unstable gains and can amplify misread cues. These results suggest that the key challenge is not reasoning length alone, but reliable mechanisms for grounding, aligning, and composing distributed evidence across real‑world multimodal clinical inputs.
Authors:Armin Dadras, Robert Sablatnig, Franziska Proksa, Markus Seidl
Abstract:
The reliable computational assessment of photographic composition requires features that are discriminative of spatial layout yet robust to semantic content. This paper proposes a low‑level representation grounded in the assumption that composition can be understood as the flow of visual attention across geometric structure. We introduce VFCNet, which fuses saliency and edge information into a gradient vector flow (GVF) field. The model computes dual‑stream GVF representations, integrates them via attention, and extracts multi‑scale flow features with a DINOv3 backbone. VFCNet achieves state‑of‑the‑art performance on the PICD benchmark (CDA‑1: 0.683, CDA‑2: 0.629), improving by 33.1% and 36.1% over the previous best method. We also show that a simple classifier on self‑supervised DINOv3 features substantially outperforms more sophisticated, composition‑specialized models. Code is available at https://github.com/ADadras/VFCNet
Authors:Devendra Ghori
Abstract:
State‑of‑the‑art deepfake detectors achieve near‑perfect in‑domain accuracy yet degrade under cross‑generator shifts, heavy compression, and adversarial perturbations. The core limitation remains the decoupling of semantic artifact learning from physical invariants: optical‑flow discontinuities, specular‑reflection inconsistencies, and cardiac‑modulated reflectance (rPPG) are treated either as post‑hoc features or ignored.
We introduce PhyLAA‑X, a novel physics‑conditioned extension of Localized Artifact Attention (LAA‑X). PhyLAA‑X injects three end‑to‑end differentiable physics‑derived feature volumes ‑ optical‑flow curl, specular‑reflectance skewness, and spatially‑upsampled rPPG power spectra ‑ directly into the LAA‑X attention computation via cross‑attention gating and a resonance consistency loss. This forces the network to learn manipulation boundaries where semantic inconsistencies and physical violations co‑occur ‑ regions inherently harder for generative models to replicate consistently.
PhyLAA‑X is embedded across an efficient spatiotemporal ensemble (EfficientNet‑B4+BiLSTM, ResNeXt‑101+Transformer, Xception+causal Conv1D) with uncertainty‑aware adaptive weighting. On FaceForensics++ (c23), Aletheia reaches 97.2% accuracy / 0.992 AUC‑ROC; on Celeb‑DF v2, 94.9% / 0.981; on DFDC, 90.8% / 0.966 ‑ outperforming the strongest published baseline (LAA‑Net [1]) by 4.1‑7.3% in cross‑generator settings and maintaining 79.4% accuracy under epsilon = 0.02 PGD‑10 attacks. Single‑backbone ablations confirm PhyLAA‑X alone delivers a 4.2% cross‑dataset AUC gain. The full production system is open‑sourced at https://github.com/devghori1264/Aletheia (v1.2, April 2026) with pretrained weights, the adversarial corpus (referred to as ADC‑2026 in this work), and complete reproducibility artifacts.
Authors:Jiaqi Shi, Yuechan Li, Xulong Zhang, Xiaoyang Qu, Jianzong Wang
Abstract:
High‑resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Existing acceleration strategies, such as token pruning or layer sparsity, suffer from severe "backbone dependency", performing well on Vicuna or Mistral architectures (e.g., LLaVA) but causing significant performance degradation when transferred to architectures like Qwen. To address this, we leverage truncated matrix entropy to uncover a universal three‑stage inference lifecycle, decoupling visual redundancy into universal Intrinsic Visual Redundancy (IVR) and architecture‑dependent Secondary Saturation Redundancy (SSR). Guided by this insight, we propose HalfV, a framework that first mitigates IVR via a unified pruning strategy and then adaptively handles SSR based on its specific manifestation. Experiments demonstrate that HalfV achieves superior efficiency‑performance trade‑offs across diverse backbones. Notably, on Qwen25‑VL, it retains 96.8% performance at a 4.1× FLOPs speedup, significantly outperforming state‑of‑the‑art baselines. Our code is available at https://github.com/civilizwa/HalfV.
Authors:Maksim Zhdanov, Ana Lucic, Max Welling, Jan-Willem van de Meent
Abstract:
We introduce Mosaic, a probabilistic weather forecasting model that addresses three failure modes of spectral degradation in ML‑based weather prediction: spectral damping (statistical), high‑frequency aliasing (architectural), and residual high‑frequency leakage (parametric). Mosaic generates ensemble members through learned functional perturbations and operates on native‑resolution grids via mesh‑aligned block‑sparse attention, a hardware‑aligned mechanism that captures long‑range dependencies at linear cost by sharing keys and values across spatially adjacent queries. At 1.5° resolution with 214M parameters, Mosaic matches or outperforms models trained on 6× finer resolution on key variables and achieves state‑of‑the‑art results among 1.5° models, producing well‑calibrated ensembles whose individual members exhibit near‑perfect spectral alignment across all resolved frequencies. A 24‑member, 10‑day forecast takes under 12s on a single H100~GPU. Code is available at https://github.com/maxxxzdn/mosaic.
Authors:Haoran Feng, Yifan Niu, Zehuan Huang, Yang-Tian Sun, Chunchao Guo, Yuxin Peng, Lu Sheng
Abstract:
We introduce LaviGen, a framework that repurposes 3D generative models for 3D layout generation. Unlike previous methods that infer object layouts from textual descriptions, LaviGen operates directly in the native 3D space, formulating layout generation as an autoregressive process that explicitly models geometric relations and physical constraints among objects, producing coherent and physically plausible 3D scenes. To further enhance this process, we propose an adapted 3D diffusion model that integrates scene, object, and instruction information and employs a dual‑guidance self‑rollout distillation mechanism to improve efficiency and spatial accuracy. Extensive experiments on the LayoutVLM benchmark show LaviGen achieves superior 3D layout generation performance, with 19% higher physical plausibility than the state of the art and 65% faster computation. Our code is publicly available at https://github.com/fenghora/LaviGen.
Authors:Dian Shao, Zhengzheng Xu, Peiyang Wang, Like Liu, Yule Wang, Jieqi Shi, Jing Huo
Abstract:
UAV vision‑language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous multi‑step instructions over long horizons. Existing zero‑shot methods remain limited, as they often rely on large base models, generic prompts, and loosely coordinated modules. In this work, we propose FineCog‑Nav, a top‑down framework inspired by human cognition that organizes navigation into fine‑grained modules for language processing, perception, attention, memory, imagination, reasoning, and decision‑making. Each module is driven by a moderate‑sized foundation model with role‑specific prompts and structured input‑output protocols, enabling effective collaboration and improved interpretability. To support fine‑grained evaluation, we construct AerialVLN‑Fine, a curated benchmark of 300 trajectories derived from AerialVLN, with sentence‑level instruction‑trajectory alignment and refined instructions containing explicit visual endpoints and landmark references. Experiments show that FineCog‑Nav consistently outperforms zero‑shot baselines in instruction adherence, long‑horizon planning, and generalization to unseen environments. These results suggest the effectiveness of fine‑grained cognitive modularization for zero‑shot aerial navigation. Project page: https://smartdianlab.github.io/projects‑FineCogNav.
Authors:Xiangbo Gao, Sicong Jiang, Bangya Liu, Xinghao Chen, Minglai Yang, Siyuan Yang, Mingyang Wu, Jiongze Yu, Qi Zheng, Haozhi Wang, Jiayi Zhang, Jie Yang, Zihan Wang, Qing Yin, Zhengzhong Tu
Abstract:
As AI‑assisted video creation becomes increasingly practical, instruction‑guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large‑scale human‑annotated dataset with complete editing examples and a standardized evaluator for comparing editing systems. Existing resources are limited by small scale, missing edited outputs, or the absence of human quality labels, while current evaluation often relies on expensive manual inspection or generic vision‑language model judges that are not specialized for editing quality. We introduce VEFX‑Dataset, a human‑annotated dataset containing 5,049 video editing examples across 9 major editing categories and 32 subcategories, each labeled along three decoupled dimensions: Instruction Following, Rendering Quality, and Edit Exclusivity. Building on VEFX‑Dataset, we propose VEFX‑Reward, a reward model designed specifically for video editing quality assessment. VEFX‑Reward jointly processes the source video, the editing instruction, and the edited video, and predicts per‑dimension quality scores via ordinal regression. We further release VEFX‑Bench, a benchmark of 300 curated video‑prompt pairs for standardized comparison of editing systems. Experiments show that VEFX‑Reward aligns more strongly with human judgments than generic VLM judges and prior reward models on both standard IQA/VQA metrics and group‑wise preference evaluation. Using VEFX‑Reward as an evaluator, we benchmark representative commercial and open‑source video editing systems, revealing a persistent gap between visual plausibility, instruction following, and edit locality in current models. Our project page is https://xiangbogaobarry.github.io/VEFX‑Bench/.
Authors:Yige Xu, Yongjie Wang, Zizhuo Wu, Kaisong Song, Jun Lin, Zhiqi Shen
Abstract:
Reasoning in vision‑language models (VLMs) has recently attracted significant attention due to its broad applicability across diverse downstream tasks. However, it remains unclear whether the superior performance of VLMs stems from genuine vision‑grounded reasoning or relies predominantly on the reasoning capabilities of their textual backbones. To systematically measure this, we introduce CrossMath, a novel multimodal reasoning benchmark designed for controlled cross‑modal comparisons. Specifically, we construct each problem in text‑only, image‑only, and image+text formats guaranteeing identical task‑relevant information, verified by human annotators. This rigorous alignment effectively isolates modality‑specific reasoning differences while eliminating confounding factors such as information mismatch. Extensive evaluation of state‑of‑the‑art VLMs reveals a consistent phenomenon: a substantial performance gap between textual and visual reasoning. Notably, VLMs excel with text‑only inputs, whereas incorporating visual data (image+text) frequently degrades performance compared to the text‑only baseline. These findings indicate that current VLMs conduct reasoning primarily in the textual space, with limited genuine reliance on visual evidence. To mitigate this limitation, we curate a CrossMath training set for VLM fine‑tuning. Empirical evaluations demonstrate that fine‑tuning on this training set significantly boosts reasoning performance across all individual and joint modalities, while yielding robust gains on two general visual reasoning tasks. Source code is available at https://github.com/xuyige/CrossMath.
Authors:Haojian Huang, Chuanyu Qin, Yinchuan Li, Yingcong Chen
Abstract:
Reinforcement learning has advanced video reasoning in large multi‑modal models, yet dominant pipelines either rely on on‑policy self‑exploration, which plateaus at the model's knowledge boundary, or hybrid replay that mixes policies and demands careful regularization. Dynamic context methods zoom into focused evidence but often require curated pretraining and two‑stage tuning, and their context remains bounded by a small model's capability. In contrast, larger models excel at instruction following and multi‑modal understanding, can supply richer context to smaller models, and rapidly zoom in on target regions via simple tools. Building on this capability, we introduce an observation‑level intervention: a frozen, tool‑integrated teacher identifies the missing spatiotemporal dependency and provides a minimal evidence patch (e.g., timestamps, regions etc.) from the original video while the question remains unchanged. The student answers again with the added context, and training updates with a chosen‑rollout scheme integrated into Group Relative Policy Optimization (GRPO). We further propose a Robust Improvement Reward (RIR) that aligns optimization with two goals: outcome validity through correct answers and dependency alignment through rationales that reflect the cited evidence. Advantages are group‑normalized across the batch, preserving on‑policy exploration while directing it along causally meaningful directions with minimal changes to the training stack. Experiments on various related benchmarks show consistent accuracy gains and strong generalization. Code will be available at https://jethrojames.github.io/FFR/.
Authors:Nishq Poorav Desai, Ali Etemad, Michael Greenspan
Abstract:
Time‑to‑Collision (TTC) forecasting is a critical task in collision prevention, requiring precise temporal prediction and comprehending both local and global patterns encapsulated in a video, both spatially and temporally. To address the multi‑scale nature of video, we introduce a novel spatiotemporal hierarchical transformer‑based architecture called CollideNet, specifically catered for effective TTC forecasting. In the spatial stream, CollideNet aggregates information for each video frame simultaneously at multiple resolutions. In the temporal stream, along with multi‑scale feature encoding, CollideNet also disentangles the non‑stationarity, trend, and seasonality components. Our method achieves state‑of‑the‑art performance in comparison to prior works on three commonly used public datasets, setting a new state‑of‑the‑art by a considerable margin. We conduct cross‑dataset evaluations to analyze the generalization capabilities of our method, and visualize the effects of disentanglement of the trend and seasonality components of the video data. We release our code at https://github.com/DeSinister/CollideNet/.
Authors:Lorenzo Beltrame, Jules Salzinger, Filip Svoboda, Jasmin Lampert, Phillipp Fanta-Jende, Radu Timofte, Marco Körner
Abstract:
We present a three‑stage progressive shadow‑removal pipeline for the CVPR2026 NTIRE WSRD+ challenge. Built on OmniSR, our method treats deshadowing as iterative direct refinement, where later stages correct residual artefacts left by earlier predictions. The model combines RGB appearance with frozen DINOv2 semantic guidance and geometric cues from monocular depth and surface normals, reused across all stages. To stabilise multi‑stage optimisation, we introduce a contraction‑constrained objective that encourages non‑increasing reconstruction error across the cascade. A staged training pipeline transfers from earlier WSRD pretraining to WSRD+ supervision and final WSRD+ 2026 adaptation with cosine‑annealed checkpoint ensembling. On the official WSRD+ 2026 hidden test set, our final ensemble achieved 26.680 PSNR, 0.8740 SSIM, 0.0578 LPIPS, and 26.135 FID, ranked first overall, and won the NTIRE 2026 Image Shadow Removal Challenge. The strong performance of the proposed model is further validated on the ISTD+ and UAV‑SC+ datasets.
Authors:Toby Perrett, Matthew Bouchard, William McCarthy
Abstract:
We introduce neuralCAD‑Edit, the first benchmark for editing 3D CAD models collected from expert CAD engineers. Instead of text conditioning as in prior works, we collect realistic CAD editing requests by capturing videos of professional designers, interacting directly with CAD models in CAD software, while talking, pointing and drawing. We recruited ten consenting designers to contribute to this contained study. We benchmark leading foundation models against human CAD experts carrying out edits, and find a large performance gap in both automatic metrics and human evaluations. Even the best foundation model (GPT 5.2) scores 53% lower (absolute) than CAD experts in human acceptance trials, demonstrating the challenge of neuralCAD‑Edit. We hope neuralCAD‑Edit will provide a solid foundation against which 3D CAD editing approaches and foundation models can be developed. Code/data: https://autodeskailab.github.io/neuralCAD‑Edit
Authors:Henry O. Velesaca, Luigi Miranda, Angel D. Sappa
Abstract:
This paper presents SWNet, a bimodal end‑to‑end cross‑spectral network specifically engineered for the detection of camouflaged weeds in dense agricultural environments. Plant camouflage, characterized by homochromatic blending where invasive species mimic the phenotypic traits of primary crops, poses a significant challenge for traditional computer vision systems. To overcome these limitations, SWNet utilizes a Pyramid Vision Transformer v2 backbone to capture long‑range dependencies and a Bimodal Gated Fusion Module to dynamically integrate Visible and Near‑Infrared information. By leveraging the physiological differences in chlorophyll reflectance captured in the NIR spectrum, the proposed architecture effectively discriminates targets that are otherwise indistinguishable in the visible range. Furthermore, an Edge‑Aware Refinement module is employed to produce sharper object boundaries and reduce structural ambiguity. Experimental results on the Weeds‑Banana dataset indicate that SWNet outperforms ten state‑of‑the‑art methods. The study demonstrates that the integration of cross‑spectral data and boundary‑guided refinement is essential for high segmentation accuracy in complex crop canopies. The code is available on GitHub: https://cod‑espol.github.io/SWNet/
Authors:Federico Nocentini, Kwanggyoon Seo, Qingju Liu, Claudio Ferrari, Stefano Berretti, David Ferman, Hyeongwoo Kim, Pablo Garrido, Akin Caliskan
Abstract:
Speech‑Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, most existing models are trained on single‑language data, limiting their effectiveness in real‑world multilingual scenarios. In this work, we address multilingual SDFA, which is essential for realistic generation since language influences phonetics, rhythm, intonation, and facial expressions. Speaking style is also shaped by individual differences, not only by language. Existing methods typically rely on either language‑specific or speaker‑specific conditioning, but not both, limiting their ability to model their interaction. We introduce Polyglot, a unified diffusion‑based architecture for personalized multilingual SDFA. Our method uses transcript embeddings to encode language information and style embeddings extracted from reference facial sequences to capture individual speaking characteristics. Polyglot does not require predefined language or speaker labels, enabling generalization across languages and speakers through self‑supervised learning. By jointly conditioning on language and style, it captures expressive traits such as rhythm, articulation, and habitual facial movements, producing temporally coherent and realistic animations. Experiments show improved performance in both monolingual and multilingual settings, providing a unified framework for modeling language and personal style in SDFA.
Authors:Laziz Hamdi, Amine Tamasna, Thierry Paquet
Abstract:
Tables condense key transactional and administrative information into compact layouts, but practical extraction requires more than text recognition: systems must also recover structure (rows, columns, merged cells, headers) and interpret roles such as line items, subtotals, and totals under common capture artifacts. Many existing resources for table structure recognition and TableVQA are built from clean digital‑born sources or rendered tables, and therefore only partially reflect noisy administrative conditions.
We introduce DenTab, a dataset of 2,000 cropped table images from dental estimates with high‑quality HTML annotations, enabling evaluation of table recognition (TR) and table visual question answering (TableVQA) on the same inputs. DenTab includes 2,208 questions across eleven categories spanning retrieval, aggregation, and logic/consistency checks. We benchmark 16 systems, including 14 vision‑‑language models (VLMs) and two OCR baselines. Across models, strong structure recovery does not consistently translate into reliable performance on multi‑step arithmetic and consistency questions, and these reasoning failures persist even when using ground‑truth HTML table inputs.
To improve arithmetic reliability without training, we propose the Table Router Pipeline, which routes arithmetic questions to deterministic execution. The pipeline combines (i) a VLM that produces a baseline answer, a structured table representation, and a constrained table program with (ii) a rule‑based executor that performs exact computation over the parsed table. The source code and dataset will be made publicly available at https://github.com/hamdilaziz/DenTab.
Authors:Jieming Yu, Qiuxiao Feng, Zhuohan Wang, Xiaochen Ma
Abstract:
With the rapid advancement of deep generative models, realistic fake images have become increasingly accessible, yet existing localization methods rely on complex designs and still struggle to generalize across manipulation types and imaging conditions. We present a simple but strong baseline based on DINOv3 with LoRA adaptation and a lightweight convolutional decoder. Under the CAT‑Net protocol, our best model improves average pixel‑level F1 by 17.0 points over the previous state of the art on four standard benchmarks using only 9.1\,M trainable parameters on top of a frozen ViT‑L backbone, and even our smallest variant surpasses all prior specialized methods. LoRA consistently outperforms full fine‑tuning across all backbone scales. Under the data‑scarce MVSS‑Net protocol, LoRA reaches an average F1 of 0.774 versus 0.530 for the strongest prior method, while full fine‑tuning becomes highly unstable, suggesting that pre‑trained representations encode forensic information that is better preserved than overwritten. The baseline also exhibits strong robustness to Gaussian noise, JPEG re‑compression, and Gaussian blur. We hope this work can serve as a reliable baseline for the research community and a practical starting point for future image‑forensic applications. Code is available at https://github.com/Irennnne/DINOv3‑IML.
Authors:Laziz Hamdi, Amine Tamasna, Pascal Boisson, Thierry Paquet
Abstract:
We present TableSeq, an image‑only, end‑to‑end framework for joint table structure recognition, content recognition, and cell localization. The model formulates these tasks as a single sequence‑generation problem: one decoder produces an interleaved stream of \textttHTML tags, cell text, and discretized coordinate tokens, thereby aligning logical structure, textual content, and cell geometry within a unified autoregressive sequence. This design avoids external OCR, auxiliary decoders, and complex multi‑stage post‑processing. TableSeq combines a lightweight high‑resolution FCN‑H16 encoder with a minimal structure‑prior head and a single‑layer transformer encoder, yielding a compact architecture that remains effective on challenging layouts. Across standard benchmarks, TableSeq achieves competitive or state‑of‑the‑art results while preserving architectural simplicity. It reaches 95.23 TEDS / 96.83 S‑TEDS on PubTabNet, 97.45 TEDS / 98.69 S‑TEDS on FinTabNet, and 99.79 / 99.54 / 99.66 precision / recall / F1 on SciTSR under the CAR protocol, while remaining competitive on PubTables‑1M under GriTS. Beyond TSR/TCR, the same sequence interface generalizes to index‑based table querying without task‑specific heads, achieving the best IRDR score and competitive ICDR/ICR performance. We also study multi‑token prediction for faster blockwise decoding and show that it reduces inference latency with only limited accuracy degradation. Overall, TableSeq provides a practical and reproducible single‑stream baseline for unified table recognition, and the source code will be made publicly available at https://github.com/hamdilaziz/TableSeq.
Authors:Meng Yu, Lei Sun, Jianhao Zeng, Xiangxiang Chu, Kun Zhan
Abstract:
Diffusion Probabilistic Models have demonstrated remarkable performance across a wide range of generative tasks. However, we have observed that these models often suffer from a Signal‑to‑Noise Ratio‑timestep (SNR‑t) bias. This bias refers to the misalignment between the SNR of the denoising sample and its corresponding timestep during the inference phase. Specifically, during training, the SNR of a sample is strictly coupled with its timestep. However, this correspondence is disrupted during inference, leading to error accumulation and impairing the generation quality. We provide comprehensive empirical evidence and theoretical analysis to substantiate this phenomenon and propose a simple yet effective differential correction method to mitigate the SNR‑t bias. Recognizing that diffusion models typically reconstruct low‑frequency components before focusing on high‑frequency details during the reverse denoising process, we decompose samples into various frequency components and apply differential correction to each component individually. Extensive experiments show that our approach significantly improves the generation quality of various diffusion models (IDDPM, ADM, DDIM, A‑DPM, EA‑DPM, EDM, PFGM++, and FLUX) on datasets of various resolutions with negligible computational overhead. The code is at https://github.com/AMAP‑ML/DCW.
Authors:Chenye Wang, Qingyuan Cai, Saihui Hou, Aoqi Li, Yongzhen Huang
Abstract:
Gait recognition has emerged as a powerful biometric technique for identifying individuals at a distance without requiring user cooperation. Most existing methods focus primarily on RGB‑derived modalities, which fall short in real‑world scenarios requiring multi‑modal collaboration and cross‑modal retrieval. To overcome these challenges, we present MMGait, a comprehensive multi‑modal gait benchmark integrating data from five heterogeneous sensors, including an RGB camera, a depth camera, an infrared camera, a LiDAR scanner, and a 4D Radar system. MMGait contains twelve modalities and 334,060 sequences from 725 subjects, enabling systematic exploration across geometric, photometric, and motion domains. Based on MMGait, we conduct extensive evaluations on single‑modal, cross‑modal, and multi‑modal paradigms to analyze modality robustness and complementarity. Furthermore, we introduce a new task, Omni Multi‑Modal Gait Recognition, which aims to unify the above three gait recognition paradigms within a single model. We also propose a simple yet powerful baseline, OmniGait, which learns a shared embedding space across diverse modalities and achieves promising recognition performance. The MMGait benchmark, codebase, and pretrained checkpoints are publicly available at https://github.com/BNU‑IVC/MMGait.
Authors:Jinhao Shen, Haoqian Du, Xulu Zhang, Xiao-Yong Wei, Qing Li
Abstract:
Text‑guided image editing, a pivotal task in modern multimedia content creation, has seen remarkable progress with training‑free methods that eliminate the need for additional optimization. Despite recent progress, existing methods are typically constrained by a competitive paradigm in which the editing and reconstruction branches are independently driven by their respective objectives to maximize alignment with target and source prompts. The adversarial strategy causes semantic conflicts and unpredictable outcomes due to the lack of coordination between branches. To overcome these issues, we propose Coopetitive Training‑Free Image Editing (CoEdit), a novel zero‑shot framework that transforms attention control from competition to coopetitive negotiation, achieving editing harmony across spatial and temporal dimensions. Spatially, CoEdit introduces Dual‑Entropy Attention Manipulation, which quantifies directional entropic interactions between branches to reformulate attention control as a harmony‑maximization problem, eventually improving the localization of editable and preservable regions. Temporally, we present Entropic Latent Refinement mechanism to dynamically adjust latent representations over time, minimizing accumulated editing errors and ensuring consistent semantic transitions throughout the denoising trajectory. Additionally, we propose the Fidelity‑Constrained Editing Score, a composite metric that jointly evaluates semantic editing and background fidelity. Extensive experiments on standard benchmarks demonstrate that CoEdit achieves superior performance in both editing quality and structural preservation, enhancing multimedia information utilization by enabling more effective interaction between visual and textual modalities. The code will be available at https://github.com/JinhaoShen/CoEdit.
Authors:Jiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen, Zhaoyang Li, Boyuan Cao, Hongming Shan
Abstract:
Video‑to‑Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the hierarchical nature of speech, which spans coarse speaker‑aware semantics to fine‑grained prosodic details. This oversight hinders direct alignment between visual and speech features at specific hierarchical levels during property matching. In this paper, leveraging the hierarchical structure of Residual Vector Quantization (RVQ)‑based codec, we propose HiCoDiT, a novel Hierarchical Codec Diffusion Transformer that exploits the inherent hierarchy of discrete speech tokens to achieve strong audio‑visual alignment. Specifically, since lower‑level tokens encode coarse speaker‑aware semantics and higher‑level tokens capture fine‑grained prosody, HiCoDiT employs low‑level and high‑level blocks to generate tokens at different levels. The low‑level blocks condition on lip‑synchronized motion and facial identity to capture speaker‑aware content, while the high‑level blocks use facial expression to modulate prosodic dynamics. Finally, to enable more effective coarse‑to‑fine conditioning, we propose a dual‑scale adaptive instance layer normalization that jointly captures global vocal style through channel‑wise normalization and local prosody dynamics through temporal‑wise normalization. Extensive experiments demonstrate that HiCoDiT outperforms baselines in fidelity and expressiveness, highlighting the potential of discrete modelling for VTS. The code and speech demo are both available at https://github.com/Jiaxin‑Ye/HiCoDiT.
Authors:Wei Lu, Zi-Yang Bo, Fei-Fei Sang, Yi Liu, Xue Yang, Si-Bao Chen
Abstract:
Shadows are prevalent in high‑resolution aerospace imagery (ASI). They often cause spectral distortion and information loss, which degrade downstream interpretation tasks. While deep learning methods have advanced natural‑image shadow removal, their direct application to ASI faces two primary challenges. First, strictly paired training data are severely lacking. Second, homogeneous shadow assumptions fail to handle the broad penumbra transition zones inherent in aerospace scenes. To address these issues, we propose AeroDeshadow, a unified two‑stage framework integrating physics‑guided shadow synthesis and penumbra‑aware restoration. In the first stage, a Physics‑aware Degradation Shadow Synthesis Network (PDSS‑Net) explicitly models illumination decay and spatial attenuation. This process constructs AeroDS‑Syn, a large‑scale paired dataset featuring soft boundary transitions. Constrained by this physical formulation, a Penumbra‑aware Cascaded DeShadowing Network (PCDS‑Net) then decouples the input into umbra and penumbra components. By restoring these regions progressively, PCDS‑Net alleviates boundary artifacts and over‑correction. Trained solely on the synthetic AeroDS‑Syn, the network generalizes to real‑world ASI without requiring paired real annotations. Experimental results indicate that AeroDeshadow achieves state‑of‑the‑art quantitative accuracy and visual fidelity across synthetic and real‑world datasets. The datasets and code will be made publicly available at: https://github.com/AeroVILab‑AHU/AeroDeshadow.
Authors:Lifan Jiang, Tianrun Wu, Yuhang Pei, Chenyang Wang, Boxi Wu, Deng Cai
Abstract:
The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross‑paradigm comparisons difficult, while video editing lacks reliable evaluation benchmarks. Furthermore, common automatic metrics often misalign with human preference, yet directly deploying large multimodal models (MLLMs) as evaluators incurs prohibitive computational and financial costs. We present UniEditBench, a unified benchmark for image and video editing that supports reconstruction‑based and instruction‑driven methods under a shared protocol. UniEditBench includes a structured taxonomy of nine image operations (Add, Remove, Replace, Change, Stroke‑based, Extract, Adjust, Count, Reorder) and eight video operations, with coverage of challenging compositional tasks such as counting and spatial reordering. To enable scalable evaluation, we distill a high‑capacity MLLM judge (Qwen3‑VL‑235B‑A22B Instruct) into lightweight 4B/8B evaluators that provide multi‑dimensional scoring over structural fidelity, text alignment, background consistency, naturalness, and temporal‑spatial consistency (for videos). Experiments show that the distilled evaluators maintain strong agreement with human judgments and substantially reduce deployment cost relative to the teacher model. UniEditBench provides a practical and reproducible protocol for benchmarking modern visual editing methods. Our benchmark and the associated reward models are publicly available at https://github.com/wesar1/UniEditBench.
Authors:Taewoong Kang, Hyojin Jang, Sohyun Jeong, Seunggi Moon, Gihwi Kim, Hoon Jin Jung, Jaegul choo
Abstract:
Recent digital media advancements have created increasing demands for sophisticated portrait manipulation techniques, particularly head swapping, where one's head is seamlessly integrated with another's body. However, current approaches predominantly rely on face‑centered cropped data with limited view angles, significantly restricting their real‑world applicability. They struggle with diverse head expressions, varying hairstyles, and natural blending beyond facial regions. To address these limitations, we propose Adaptive Head Synthesis (AHS), which effectively handles full upper‑body images with varied head poses and expressions. AHS incorporates a novel head reenacted synthetic data augmentation strategy to overcome self‑supervised training constraints, enhancing generalization across diverse facial expressions and orientations without requiring paired training data. Comprehensive experiments demonstrate that AHS achieves superior performance in challenging real‑world scenarios, producing visually coherent results that preserve identity and expression fidelity across various head orientations and hairstyles. Notably, AHS shows exceptional robustness in maintaining facial identity while drastic expression changes and faithfully preserving accessories while significant head pose variations.
Authors:Irem Ulku, Erdem Akagündüz, Ömer Özgür Tanrıöver
Abstract:
Multimodal remote sensing data provide complementary information for semantic segmentation, but in real‑world deployments, some modalities may be unavailable due to sensor failures, acquisition issues, or challenging atmospheric conditions. Existing multimodal segmentation models typically address missing modalities by learning a shared representation across inputs. However, this approach can introduce a trade‑off by compromising modality‑specific complementary information and reducing performance when all modalities are available. In this paper, we tackle this limitation with CBC‑SLP, a multimodal semantic segmentation model designed to preserve both modality‑invariant and modality‑specific information. Inspired by the theoretical results on modality alignment, which state that perfectly aligned multimodal representations can lead to sub‑optimal performance in downstream prediction tasks, we propose a novel structured latent projection approach as an architectural inductive bias. Rather than enforcing this strategy through a loss term, we incorporate it directly into the architecture. In particular, to use the complementary information effectively while maintaining robustness under random modality dropout, we structure the latent representations into shared and modality‑specific components and adaptively transfer them to the decoder according to the random modality availability mask. Extensive experiments on three multimodal remote sensing image sets demonstrate that CBC‑SLP consistently outperforms state‑of‑the‑art multimodal models across full and missing modality scenarios. Besides, we empirically demonstrate that the proposed strategy can recover the complementary information that may not be preserved in a shared representation. The code is available at https://github.com/iremulku/Multispectral‑Semantic‑Segmentation‑via‑Structured‑Latent‑Projection‑CBC‑SLP‑.
Authors:Liwen Yu, Chi Liu, Xiaotong Han, Congcong Zhu, Minghao Wang, Sheng Shen
Abstract:
Automated Aesthetic Quality Assessment (AQA) treats images primarily as static pixel vectors, aligning predictions with human‑rating scores largely through semantic perception. However, this paradigm diverges from human aesthetic cognition, which arises from dynamic visual exploration shaped by scanning paths, processing fluency, and the interplay between bottom‑up salience and top‑down intention. We introduce AestheticNet, a novel cognitive‑inspired AQA paradigm that integrates human‑like visual cognition and semantic perception with a two‑pathway architecture. The visual attention pathway, implemented as a gaze‑aligned visual encoder (GAVE) pre‑trained offline on eye‑tracking data using resource‑efficient contrast gaze alignment, models attention from human vision system. This pathway augments the semantic pathway, which uses a fixed semantic encoder such as CLIP, through cross‑attention fusion. Visual attention provides a cognitive prior reflecting foreground/background structure, color cascade, brightness, and lighting, all of which are determinants of aesthetic perception beyond semantics. Experiments validated by hypothesis testing show a consistent improvement over the semantic‑alone baselines, and demonstrate the gaze module as a model‑agnostic corrector compatible with diverse AQA backbones, supporting the necessity and modularity of human‑like visual cognition for AQA. Our code is available at https://github.com/keepgallop/AestheticNet.
Authors:Jun Li, Lizhi Xiong, Ziqiang Li, Weiwei Jiang, Zhangjie Fu, Yong Li, Guo-Sen Xie
Abstract:
Text‑to‑image generative models have achieved impressive fidelity and diversity, but can inadvertently produce unsafe or undesirable content due to implicit biases embedded in large‑scale training datasets. Existing concept erasure methods, whether text‑only or image‑assisted, face trade‑offs: textual approaches often fail to fully suppress concepts, while naive image‑guided methods risk over‑erasing unrelated content. We propose TICoE, a text‑image Collaborative Erasing framework that achieves precise and faithful concept removal through a continuous convex concept manifold and hierarchical visual representation learning. TICoE precisely removes target concepts while preserving unrelated semantic and visual content. To objectively assess the quality of erasure, we further introduce a fidelity‑oriented evaluation strategy that measures post‑erasure usability. Experiments on multiple benchmarks show that TICoE surpasses prior methods in concept removal precision and content fidelity, enabling safer, more controllable text‑to‑image generation. Our code is available at https://github.com/OpenAscent‑L/TICoE.git
Authors:Chengxin Liu, Wonseok Choi, Chenshuang Zhang, Tae-Hyun Oh
Abstract:
Vision‑Language Models (VLMs) have demonstrated strong capability in a wide range of tasks such as visual recognition, document parsing, and visual grounding. Nevertheless, recent work shows that while VLMs often manage to capture the correct image region corresponding to the question, they do not necessarily produce the correct answers. In this work, we demonstrate that this misalignment could be attributed to suboptimal information flow within VLMs, where text tokens distribute too much attention to irrelevant visual tokens, leading to incorrect answers. Based on the observation, we show that modulating the information flow during inference can improve the perception capability of VLMs. The idea is that text tokens should only be associated with important visual tokens during decoding, eliminating the interference of irrelevant regions. To achieve this, we propose a token dynamics‑based method to determine the importance of visual tokens, where visual tokens that exhibit distinct activation patterns during different decoding stages are viewed as important. We apply our approach to representative open‑source VLMs and evaluate on various datasets, including visual question answering, visual grounding and counting, optical character recognition, and object hallucination. The results show that our approach significantly improves the performance of baselines. Project page: https://cxliu0.github.io/AIF/.
Authors:Junjie Wen, Junlin He, Fei Ma, Jinqiang Cui
Abstract:
Accurate open‑vocabulary 3D scene understanding requires semantic representations that are both language‑aligned and spatially precise at the pixel level, while remaining scalable when lifted to 3D space. However, existing representations struggle to jointly satisfy these requirements, and densely propagating pixel‑wise semantics to 3D often results in substantial redundancy, leading to inefficient storage and querying in large‑scale scenes. To address these challenges, we present \emphPLAF, a Pixel‑wise Language‑Aligned Feature extraction framework that enables dense and accurate semantic alignment in 2D without sacrificing open‑vocabulary expressiveness. Building upon this representation, we further design an efficient semantic storage and querying scheme that significantly reduces redundancy across both 2D and 3D domains. Experimental results show that \emphPLAF provides a strong semantic foundation for accurate and efficient open‑vocabulary 3D scene understanding. The codes are publicly available at https://github.com/RockWenJJ/PLAF.
Authors:Jinlun Ye, Jiang Liao, Runhe Lai, Xinhua Lu, Jiaxin Zhuang, Zhiyong Gan, Ruixuan Wang
Abstract:
Vision‑language models (VLMs) such as CLIP exhibit strong Out‑of‑distribution (OOD) detection capabilities by aligning visual and textual representations. Recent CLIP‑based test‑time adaptation methods further improve detection performance by incorporating external OOD labels. However, such labels are finite and fixed, while the real OOD semantic space is inherently open‑ended. Consequently, fixed labels fail to represent the diverse and evolving OOD semantics encountered in test streams. To address this limitation, we introduce Test‑time Textual Learning (TTL), a framework that dynamically learns OOD textual semantics from unlabeled test streams, without relying on external OOD labels. TTL updates learnable prompts using pseudo‑labeled test samples to capture emerging OOD knowledge. To suppress noise introduced by pseudo‑labels, we introduce an OOD knowledge purification strategy that selects reliable OOD samples for adaptation while suppressing noise. In addition, TTL maintains an OOD Textual Knowledge Bank that stores high‑quality textual features, providing stable score calibration across batches. Extensive experiments on two standard benchmarks with nine OOD datasets demonstrate that TTL consistently achieves state‑of‑the‑art performance, highlighting the value of textual adaptation for robust test‑time OOD detection. Our code is available at https://github.com/figec/TTL.
Authors:Junguang Yao, Wenye Liu, Stjepan Picek, Yue Zheng
Abstract:
Visual speaker recognition based on lip motion offers a silent, hands‑free, and behavior‑driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely heavily on appearance‑dependent representations, lip motion encodes subject‑specific behavioral dynamics driven by consistent articulation patterns and muscle coordination, offering inherent stability across environmental changes. However, capturing these robust, fine‑grained dynamics is challenging for conventional frame‑based cameras due to motion blur and low dynamic range. To exploit the intrinsic stability of lip motion and address these sensing limitations, we propose NeuroLip, an event‑based framework that captures fine‑grained lip dynamics under a strict yet practical cross‑scene protocol: training is performed under a single controlled condition, while recognition must generalize to unseen viewing and lighting conditions. NeuroLip features a 1) Temporal‑aware Voxel Encoding module with adaptive event weighting, 2) Structure‑aware Spatial Enhancer that amplifies discriminative behavioral patterns by suppressing noise while preserving vertically structured motion information, and 3) Polarity Consistency Regularization mechanism to retain motion‑direction cues encoded in event polarities. To facilitate systematic evaluation, we introduce DVSpeaker, a comprehensive event‑based lip‑motion dataset comprising 50 subjects recorded under four distinct viewpoint and illumination scenarios. Extensive experiments demonstrate that NeuroLip achieves near‑perfect matched‑scene accuracy and robust cross‑scene generalization, attaining over 71% accuracy on unseen viewpoints and nearly 76% under low‑light conditions, outperforming representative existing methods by at least 8.54%. The dataset and code are publicly available at https://github.com/JiuZeongit/NeuroLip.
Authors:Geunyoung Jung, Soohong Kim, Inseok Kong, Jiyoung Jung
Abstract:
The advent of deep neural networks has led to remarkable progress in 3D point cloud recognition, but they remain vulnerable to adversarial attacks. Although various defense methods have been studied, they suffer from a trade‑off between robustness and transferability. We propose Adversarial Point Counterattack (APC) to achieve both simultaneously. APC is a lightweight input‑level purification module that generates instance‑specific counter‑perturbations for each point, effectively neutralizing attacks. Leveraging clean‑adversarial pairs, APC enforces geometric consistency in data space and semantic consistency in feature space. To improve generalizability across diverse attacks, we adopt a hybrid training strategy using adversarial point clouds from multiple attack types. Since APC operates purely on input point clouds, it directly transfers to unseen models and defends against attacks targeting them without retraining. At inference, a single APC forward pass provides purified point clouds with negligible time and parameter overhead. Extensive experiments on two 3D recognition benchmarks demonstrate that the APC achieves state‑of‑the‑art defense performance. Furthermore, cross‑model evaluations validate its superior transferability. The code is available at https://github.com/gyjung975/APC.
Authors:Ruxin Ding, Jianfeng Ren, Heng Yu, Jiawei Li, Xudong Jiang
Abstract:
Spatiotemporal Local Binary Pattern (STLBP) is a widely used dynamic texture descriptor, but it suffers from extremely high dimensionality. To tackle this, STLBP features are often extracted on three orthogonal planes, which sacrifice inter‑plane correlation. In this work, we propose a Locality‑Preserving Pixel‑Difference Hashing (LP^2DH) framework that jointly encodes pixel differences in the full spatiotemporal neighbourhood. LP^2DH transforms Pixel‑Difference Vectors (PDVs) into compact binary codes with maximal discriminative power. Furthermore, we incorporate a locality‑preserving embedding to maintain the PDVs' local structure before and after hashing. Then, a curvilinear search strategy is utilized to jointly optimize the hashing matrix and binary codes via gradient descent on the Stiefel manifold. After hashing, dictionary learning is applied to encode the binary vectors into codewords, and the resulting histogram is utilized as the final feature representation. The proposed LP^2DH achieves state‑of‑the‑art performance on three major dynamic texture recognition benchmarks: 99.80% against DT‑GoogleNet's 98.93% on UCLA, 98.52% against HoGF^3D's 97.63% on DynTex++, and 96.19% compared to STS's 95.00% on YUPENN. The source code is available at: https://github.com/drx770/LP2DH.
Authors:Geunyoung Jung, Soohong Kim, Kyungwoo Song, Jiyoung Jung
Abstract:
With the rise of pre‑trained models in the 3D point cloud domain for a wide range of real‑world applications, adapting them to downstream tasks has become increasingly important. However, conventional full fine‑tuning methods are computationally expensive and storage‑intensive. Although prompt tuning has emerged as an efficient alternative, it often suffers from overfitting, thereby compromising generalization capability. To address this issue, we propose Prototypical Point‑level Prompt Tuning (P^3T), a parameter‑efficient prompt tuning method designed for pre‑trained 3D vision‑language models (VLMs). P^3T consists of two components: 1) Point Prompter, which generates instance‑aware point‑level prompts for the input point cloud, and 2) Text Prompter, which employs learnable prompts into the input text instead of hand‑crafted ones. Since both prompters operate directly on input data, P^3T enables task‑specific adaptation of 3D VLMs without sacrificing generalizability. Furthermore, to enhance embedding space alignment, which is key to fine‑tuning 3D VLMs, we introduce a prototypical loss that reduces intra‑category variance. Extensive experiments demonstrate that our method matches or outperforms full fine‑tuning in classification and few‑shot learning, and further exhibits robust generalization under data shift in the cross‑dataset setting. The code is available at \textcolorviolethttps://github.com/gyjung975/P3T.
Authors:Chen Zhao, Yunzhe Xu, Zhizhou Chen, Enxuan Gu, Kai Zhang, Xiaoming Liu, Jian Yang, Ying Tai
Abstract:
Ultra‑high‑definition (UHD) image restoration poses unique challenges due to the high spatial resolution, diverse content, and fine‑grained structures present in UHD images. To address these issues, we introduce a progressive spectral decomposition for the restoration process, decomposing it into three stages: zero‑frequency enhancement, low‑frequency restoration, and high‑frequency refinement. Based on this formulation, we propose a novel framework, ERR, which integrates three cooperative sub‑networks: the zero‑frequency enhancer (ZFE), the low‑frequency restorer (LFR), and the high‑frequency refiner (HFR). The ZFE incorporates global priors to learn holistic mappings, the LFR reconstructs the main content by focusing on coarse‑scale information, and the HFR adopts our proposed frequency‑windowed Kolmogorov‑Arnold Network (FW‑KAN) to recover fine textures and intricate details for high‑fidelity restoration. To further advance research in UHD image restoration, we also construct a large‑scale, high‑quality benchmark dataset, LSUHDIR, comprising 82,126 UHD images with diverse scenes and rich content. Our proposed methods demonstrate superior performance across a range of UHD image restoration tasks, and extensive ablation studies confirm the contribution and necessity of each module. Project page: https://github.com/NJU‑PCALab/ERR.
Authors:Bingyu Li, Tao Huo, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li
Abstract:
Open‑vocabulary remote sensing image segmentation (OVRSIS) remains underexplored due to fragmented datasets, limited training diversity, and the lack of evaluation benchmarks that reflect realistic geospatial application demands. Our previous OVRSISBenchV1 established an initial cross‑dataset evaluation protocol, but its limited scope is insufficient for assessing realistic open‑world generalization. To address this issue, we propose OVRSISBenchV2, a large‑scale and application‑oriented benchmark for OVRSIS. We first construct OVRSIS95K, a balanced dataset of about 95K image‑‑mask pairs covering 35 common semantic categories across diverse remote sensing scenes. Built upon OVRSIS95K and 10 downstream datasets, OVRSISBenchV2 contains 170K images and 128 categories, substantially expanding scene diversity, semantic coverage, and evaluation difficulty. Beyond standard open‑vocabulary segmentation, it further includes downstream protocols for building extraction, road extraction, and flood detection, thereby better reflecting realistic geospatial application demands and complex deployment scenarios. We also propose Pi‑Seg, a baseline for OVRSIS. Pi‑Seg improves transferability through a positive‑incentive noise mechanism, where learnable and semantically guided perturbations broaden the visual‑text feature space during training. Extensive experiments on OVRSISBenchV1, OVRSISBenchV2, and downstream tasks show that Pi‑Seg delivers strong and consistent results, particularly on the more challenging OVRSISBenchV2 benchmark. Our results highlight both the importance of realistic benchmark design and the effectiveness of perturbation‑based transfer for OVRSIS. The code and datasets are available at \hrefhttps://github.com/LiBingyu01/RSKT‑Seg/tree/Pi‑SegLiBingyu01/RSKT‑Seg/tree/Pi‑Seg.
Authors:Dong-Uk Seo, Jinwoo Jeon, Eungchang Mason Lee, Hyun Myung
Abstract:
Gaussian splatting has recently gained traction as a compelling map representation for SLAM systems, enabling dense and photo‑realistic scene modeling. However, its application to monocular SLAM remains challenging due to the lack of reliable geometric cues from monocular input. Without geometric supervision, mapping or tracking could fall in local‑minima, resulting in structural degeneracies and inaccuracies. To address this challenge, we propose GaussianFlow SLAM, a monocular 3DGS‑SLAM that leverages optical flow as a geometry‑aware cue to guide the optimization of both the scene structure and camera poses. By encouraging the projected motion of Gaussians, termed GaussianFlow, to align with the optical flow, our method introduces consistent structural cues to regularize both map reconstruction and pose estimation. Furthermore, we introduce normalized error‑based densification and pruning modules to refine inactive and unstable Gaussians, thereby contributing to improved map quality and pose accuracy. Experiments conducted on public datasets demonstrate that our method achieves superior rendering quality and tracking accuracy compared with state‑of‑the‑art algorithms. The source code is available at: https://github.com/url‑kaist/gaussianflow‑slam.
Authors:Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan Yuille, Liang-Chieh Chen
Abstract:
Flow matching models have emerged as a powerful framework for realistic image generation by learning to reverse a corruption process that progressively adds Gaussian noise. However, because noise is injected in the latent domain, its impact on different frequency components is non‑uniform. As a result, during inference, flow matching models tend to generate low‑frequency components (global structure) in the early stages, while high‑frequency components (fine details) emerge only later in the reverse process. Building on this insight, we propose Frequency‑Aware Flow Matching (FreqFlow), a novel approach that explicitly incorporates frequency‑aware conditioning into the flow matching framework via time‑dependent adaptive weighting. We introduce a two‑branch architecture: (1) a frequency branch that separately processes low‑ and high‑frequency components to capture global structure and refine textures and edges, and (2) a spatial branch that synthesizes images in the latent domain, guided by the frequency branch's output. By explicitly integrating frequency information into the generation process, FreqFlow ensures that both large‑scale coherence and fine‑grained details are effectively modeled low‑frequency conditioning reinforces global structure, while high‑frequency conditioning enhances texture fidelity and detail sharpness. On the class‑conditional ImageNet‑256 generation benchmark, our method achieves state‑of‑the‑art performance with an FID of 1.38, surpassing the prior diffusion model DiT and flow matching model SiT by 0.79 and 0.58 FID, respectively. Code is available at https://github.com/OliverRensu/FreqFlow.
Authors:Mohammad Mahdi Abootorabi, Parvin Mousavi, Purang Abolmaesumi, Evan Shelhamer
Abstract:
Deep networks that rely on prototypes‑interpretable representations that can be related to the model input‑have gained significant attention for balancing high accuracy with inherent interpretability, which makes them suitable for critical domains such as healthcare. However, these models are limited by their reliance on training data, which hampers their robustness to distribution shifts. While test‑time adaptation (TTA) improves the robustness of deep networks by updating parameters and statistics, the prototypes of interpretable models have not been explored for this purpose. We introduce ProtoTTA, a general framework for prototypical models that leverages intermediate prototype signals rather than relying solely on model outputs. ProtoTTA minimizes the entropy of the prototype‑similarity distribution to encourage more confident and prototype‑specific activations on shifted data. To maintain stability, we employ geometric filtering to restrict updates to samples with reliable prototype activations, regularized by prototype‑importance weights and model‑confidence scores. Experiments across four prototypical backbones on four diverse benchmarks spanning fine‑grained vision, histopathology, and NLP demonstrate that ProtoTTA improves robustness over standard output entropy minimization while restoring correct semantic focus in prototype activations. We also introduce novel interpretability metrics and a vision‑language model (VLM) evaluation framework to explain TTA dynamics, confirming ProtoTTA restores human‑aligned semantic focus and correlates reliably with VLM‑rated reasoning quality. Code is available at: https://github.com/DeepRCL/ProtoTTA.
Authors:Sanjeev Panta, Rhett M Morvant, Xu Yuan, Li Chen, Nian-Feng Tzeng
Abstract:
Accurate and timely rainfall nowcasting is crucial for disaster mitigation and water resource management. Despite recent advances in deep learning, precipitation prediction remains challenging due to limitations in effectively leveraging diverse multimedia data sources. We introduce M3R, a Meteorology‑informed MultiModal attention‑based architecture for direct Rainfall prediction that synergistically combines visual NEXRAD radar imagery with numerical Personal Weather Station (PWS) measurements, using a comprehensive pipeline for temporal alignment of heterogeneous meteorological data. With specialized multimodal attention mechanisms, M3R novelly leverages weather station time series as queries to selectively attend to spatial radar features, enabling focused extraction of precipitation signatures. Experimental results for three spatial areas of 100 km 100 km centered at NEXRAD radar stations demonstrate that M3R outperforms existing approaches, achieving substantial improvements in accuracy, efficiency, and precipitation detection capabilities. Our work establishes new benchmarks for multimedia‑based precipitation nowcasting and provides practical tools for operational weather prediction systems. The source code is available at https://github.com/Sanjeev97/M3Rain
Authors:Keon Kim, Krish Chelikavada
Abstract:
Multi‑step zoom‑in pipelines are widely used for GUI grounding, yet the intermediate predictions they produce are typically discarded after coordinate remapping. We observe that these intermediate outputs contain a useful confidence signal for free: zoom consistency, the distance between a model's step‑2 prediction and the crop center. Unlike log‑probabilities or token‑level uncertainty, zoom consistency is a geometric quantity in a shared coordinate space, making it directly comparable across architecturally different VLMs without calibration. We prove this quantity is a linear estimator of step‑1 spatial error under idealized conditions (perfect step‑2, target within crop) and show it correlates with prediction correctness across two VLMs (AUC = 0.60; Spearman rho = ‑0.14, p < 10^‑6 for KV‑Ground‑8B; rho = ‑0.11, p = 0.0003 for Qwen3.5‑27B). The correlation is small but consistent across models, application categories, and operating systems. As a proof‑of‑concept, we use zoom consistency to route between a specialist and generalist model, capturing 16.5% of the oracle headroom between them (+0.8%, McNemar p = 0.19). Code is available at https://github.com/omxyz/zoom‑consistency‑routing.
Authors:Ninghui Xu, Fabio Tosi, Lihui Wang, Jiawei Han, Luca Bartolomei, Zhiting Yao, Matteo Poggi, Stefano Mattoccia
Abstract:
Conventional frame‑based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternative visual representation with higher dynamic range free from such limitations. The complementary characteristics of the two modalities make event‑frame asymmetric stereo promising for reliable 3D perception under fast motion and challenging illumination. However, the modality gap often leads to marginalization of domain‑specific cues essential for cross‑modal stereo matching. In this paper, we introduce Bi‑CMPStereo, a novel bidirectional cross‑modal prompting framework that fully exploits semantic and structural features from both domains for robust matching. Our approach learns finely aligned stereo representations within a target canonical space and integrates complementary representations by projecting each modality into both event and frame domains. Extensive experiments demonstrate that our approach significantly outperforms state‑of‑the‑art methods in accuracy and generalization.
Authors:Zhanhao Liang, Tao Yang, Jie Wu, Chengjian Feng, Liang Zheng
Abstract:
This paper focuses on the alignment of flow matching models with human preferences. A promising way is fine‑tuning by directly backpropagating reward gradients through the differentiable generation process of flow matching. However, backpropagating through long trajectories results in prohibitive memory costs and gradient explosion. Therefore, direct‑gradient methods struggle to update early generation steps, which are crucial for determining the global structure of the final image. To address this issue, we introduce LeapAlign, a fine‑tuning method that reduces computational cost and enables direct gradient propagation from reward to early generation steps. Specifically, we shorten the long trajectory into only two steps by designing two consecutive leaps, each skipping multiple ODE sampling steps and predicting future latents in a single step. By randomizing the start and end timesteps of the leaps, LeapAlign leads to efficient and stable model updates at any generation step. To better use such shortened trajectories, we assign higher training weights to those that are more consistent with the long generation path. To further enhance gradient stability, we reduce the weights of gradient terms with large magnitude, instead of completely removing them as done in previous works. When fine‑tuning the Flux model, LeapAlign consistently outperforms state‑of‑the‑art GRPO‑based and direct‑gradient methods across various metrics, achieving superior image quality and image‑text alignment.
Authors:Sumit Chaturvedi, Yannick Hold-Geoffroy, Mengwei Ren, Jingyuan Liu, He Zhang, Yiqun Mei, Julie Dorsey, Zhixin Shu
Abstract:
This paper presents a method for image relighting that enables precise and continuous control over multiple illumination attributes in a photograph. We formulate relighting as a conditional image generation task and introduce attribute tokens to encode distinct lighting factors such as intensity, color, ambient illumination, diffuse level, and 3D light positions. The model is trained on a large‑scale synthetic dataset with ground‑truth lighting annotations, supplemented by a small set of real captures to enhance realism and generalization. We validate our approach across a variety of relighting tasks, including controlling in‑scene lighting fixtures and editing environment illumination using virtual light sources, on synthetic and real images. Our method achieves state‑of‑the‑art quantitative and qualitative performance compared to prior work. Remarkably, without explicit inverse rendering supervision, the model exhibits an inherent understanding of how light interacts with scene geometry, occlusion, and materials, yielding convincing lighting effects even in traditionally challenging scenarios such as placing lights within objects or relighting transparent materials plausibly. Project page: vrroom.github.io/tokenlight/
Authors:Hao Gao, Shaoyu Chen, Yifan Zhu, Yuehao Song, Wenyu Liu, Qian Zhang, Xinggang Wang
Abstract:
High‑level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed‑loop interactions. Although diffusion‑based planners are effective at modeling complex trajectory distributions, they often suffer from stochastic instabilities and the lack of corrective negative feedback when trained purely with imitation learning. To address these issues, we propose RAD‑2, a unified generator‑discriminator framework for closed‑loop planning. Specifically, a diffusion‑based generator is used to produce diverse trajectory candidates, while an RL‑optimized discriminator reranks these candidates according to their long‑term driving quality. This decoupled design avoids directly applying sparse scalar rewards to the full high‑dimensional trajectory space, thereby improving optimization stability. To further enhance reinforcement learning, we introduce Temporally Consistent Group Relative Policy Optimization, which exploits temporal coherence to alleviate the credit assignment problem. In addition, we propose On‑policy Generator Optimization, which converts closed‑loop feedback into structured longitudinal optimization signals and progressively shifts the generator toward high‑reward trajectory manifolds. To support efficient large‑scale training, we introduce BEV‑Warp, a high‑throughput simulation environment that performs closed‑loop evaluation directly in Bird's‑Eye View feature space via spatial warping. RAD‑2 reduces the collision rate by 56% compared with strong diffusion‑based planners. Real‑world deployment further demonstrates improved perceived safety and driving smoothness in complex urban traffic.
Authors:Yiyang Jiang, Li Zhang, Xiao-Yong Wei, Li Qing
Abstract:
Many SLT systems quietly assume that brief chunks of signing map directly to spoken‑language words. That assumption breaks down because signers often create meaning on the fly using context, space, and movement. We revisit SLT and argue that it is mainly a cross‑modal reasoning task, not just a straightforward video‑to‑text conversion. We thus introduce a reasoning‑driven SLT framework that uses an ordered sequence of latent thoughts as an explicit middle layer between the video and the generated text. These latent thoughts gradually extract and organize meaning over time. On top of this, we use a plan‑then‑ground decoding method: the model first decides what it wants to say, and then looks back at the video to find the evidence. This separation improves coherence and faithfulness. We also built and released a new large‑scale gloss‑free SLT dataset with stronger context dependencies and more realistic meanings. Experiments across several benchmarks show consistent gains over existing gloss‑free methods. Our code and data are available at https://github.com/fletcherjiang/SignThought.
Authors:Leyi Wu, Pengjun Fang, Kai Sun, Yazhou Xing, Yinwei Wu, Songsong Wang, Ziqi Huang, Dan Zhou, Yingqing He, Ying-Cong Chen, Qifeng Chen
Abstract:
Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks‑largely designed for realistic videos‑struggle to evaluate animation‑style generation with its stylized appearance, exaggerated motion, and character‑centric consistency. Moreover, they also rely on fixed prompt sets and rigid pipelines, offering limited flexibility for open‑domain content and custom evaluation needs. To address this gap, we introduce AnimationBench, the first systematic benchmark for evaluating animation image‑to‑video generation. AnimationBench operationalizes the Twelve Basic Principles of Animation and IP Preservation into measurable evaluation dimensions, together with Broader Quality Dimensions including semantic consistency, motion rationality, and camera motion consistency. The benchmark supports both a standardized close‑set evaluation for reproducible comparison and a flexible open‑set evaluation for diagnostic analysis, and leverages visual‑language models for scalable assessment. Extensive experiments show that AnimationBench aligns well with human judgment and exposes animation‑specific quality differences overlooked by realism‑oriented benchmarks, leading to more informative and discriminative evaluation of state‑of‑the‑art I2V models.
Authors:Roni Itkin, Noam Issachar, Yehonatan Keypur, Xingyu Chen, Anpei Chen, Sagie Benaim
Abstract:
The efficient spatial allocation of primitives serves as the foundation of 3D Gaussian Splatting, as it directly dictates the synergy between representation compactness, reconstruction speed, and rendering fidelity. Previous solutions, whether based on iterative optimization or feed‑forward inference, suffer from significant trade‑offs between these goals, mainly due to the reliance on local, heuristic‑driven allocation strategies that lack global scene awareness. Specifically, current feed‑forward methods are largely pixel‑aligned or voxel‑aligned. By unprojecting pixels into dense, view‑aligned primitives, they bake redundancy into the 3D asset. As more input views are added, the representation size increases and global consistency becomes fragile. To this end, we introduce GlobalSplat, a framework built on the principle of align first, decode later. Our approach learns a compact, global, latent scene representation that encodes multi‑view input and resolves cross‑view correspondences before decoding any explicit 3D geometry. Crucially, this formulation enables compact, globally consistent reconstructions without relying on pretrained pixel‑prediction backbones or reusing latent features from dense baselines. Utilizing a coarse‑to‑fine training curriculum that gradually increases decoded capacity, GlobalSplat natively prevents representation bloat. On RealEstate10K and ACID, our model achieves competitive novel‑view synthesis performance while utilizing as few as 16K Gaussians, significantly less than required by dense pipelines, obtaining a light 4MB footprint. Further, GlobalSplat enables significantly faster inference than the baselines, operating under 78 milliseconds in a single forward pass. Project page is available at https://r‑itk.github.io/globalsplat/
Authors:Tianhao Fu, Austin Wang, Charles Chen, Roby Aldave-Garza, Yucheng Chen
Abstract:
Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and clinical decision support. Many strong uncertainty methods require repeated inference, while efficient single‑forward‑pass alternatives often provide weaker failure ranking or rely on restrictive feature‑space assumptions. We present SegWithU, a post‑hoc framework that augments a frozen pretrained segmentation backbone with a lightweight uncertainty head. SegWithU taps intermediate backbone features and models uncertainty as perturbation energy in a compact probe space using rank‑1 posterior probes. It produces two voxel‑wise uncertainty maps: a calibration‑oriented map for probability tempering and a ranking‑oriented map for error detection and selective prediction. Across ACDC, BraTS2024, and LiTS, SegWithU is the strongest and most consistent single‑forward‑pass baseline, achieving AUROC/AURC of 0.9838/2.4885, 0.9946/0.2660, and 0.9925/0.8193, respectively, while preserving segmentation quality. These results suggest that perturbation‑based uncertainty modeling is an effective and practical route to reliability‑aware medical segmentation.
Source code is available at https://github.com/ProjectNeura/SegWithU.
Authors:Onno Niemann, Gonzalo Martínez Muñoz, Alberto Suárez Gonzalez
Abstract:
Recent work has shown that diffusion models trained with the denoising score matching (DSM) objective often violate the Fokker‑‑Planck (FP) equation that governs the evolution of the true data density. Directly penalizing these deviations in the objective function reduces their magnitude but introduces a significant computational overhead. It is also observed that enforcing strict adherence to the FP equation does not necessarily lead to improvements in the quality of the generated samples, as often the best results are obtained with weaker FP regularization. In this paper, we investigate whether simpler penalty terms can provide similar benefits. We empirically analyze several lightweight regularizers, study their effect on FP residuals and generation quality, and show that the benefits of FP regularization are available at substantially lower computational cost. Our code is available at https://github.com/OnnoNiemann/fp_diffusion_analysis.
Authors:Youngjin Oh, Junyoung Park, Junhyeong Kwon, Nam Ik Cho
Abstract:
Adverse lighting conditions, such as cast shadows and irregular illumination, pose significant challenges to computer vision systems by degrading visibility and color fidelity. Consequently, effective shadow removal and ALN are critical for restoring underlying image content, improving perceptual quality, and facilitating robust performance in downstream tasks. However, while achieving state‑of‑the‑art results on specific benchmarks is a primary goal in image restoration challenges, real‑world applications often demand robust models capable of handling diverse domains. To address this, we present a comprehensive study on lighting‑related image restoration by exploring two contrasting strategies. We leverage a robust framework for ALN, DINOLight, as a specialized baseline to exploit the characteristics of each individual dataset, and extend it to OmniLight, a generalized alternative incorporating our proposed Wavelet Domain Mixture‑of‑Experts (WD‑MoE) that is trained across all provided datasets. Through a comparative analysis of these two methods, we discuss the impact of data distribution on the performance of specialized and unified architectures in lighting‑related image restoration. Notably, both approaches secured top‑tier rankings across all three lighting‑related tracks in the NTIRE 2026 Challenge, demonstrating their outstanding perceptual quality and generalization capabilities. Our codes are available at https://github.com/OBAKSA/Lighting‑Restoration.
Authors:Arman Hatami, Romina Aalishah, Ilya E. Monosov
Abstract:
Machine unlearning aims to remove targeted knowledge from a trained model without the cost of retraining from scratch. In class unlearning, however, reducing accuracy on forget classes does not necessarily imply true forgetting: forgotten information can remain encoded in internal representations, and apparent forgetting may arise from classifier‑head suppression rather than representational removal. We show that existing class‑unlearning methods often exhibit weak or negative selectivity, preserve forget‑class structure in deep representations, or rely heavily on final‑layer bias shifts. We then introduce DAMP (Depth‑Aware Modulation by Projection), a one‑shot, closed‑form weight‑surgery method that removes forget‑specific directions from a pretrained network without gradient‑based optimization. At each stage, DAMP computes class prototypes in the input space of the next learnable operator, extracts forget directions as residuals relative to retain‑class prototypes, and applies a projection‑based update to reduce downstream sensitivity to those directions. To preserve utility, DAMP uses a parameter‑free depth‑aware scaling rule derived from probe separability, applying smaller edits in early layers and larger edits in deeper layers. The method naturally extends to multi‑class forgetting through low‑rank subspace removal. Across MNIST, CIFAR‑10, CIFAR‑100, and Tiny ImageNet, and across convolutional and transformer architectures, DAMP more closely resembles the retraining gold standard than some of the prior methods, improving selective forgetting while better preserving retain‑class performance and reducing residual forget‑class structure in deep layers.
Authors:Kanzhi Cheng, Zehao Li, Zheng Ma, Nuo Chen, Jialin Cao, Qiushi Sun, Zichen Ding, Fangzhi Xu, Hang Yan, Jiajun Chen, Anh Tuan Luu, Jianbing Zhang, Lewei Lu, Dahua Lin
Abstract:
Mobile agents powered by vision‑language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achieving a marked performance leap, e.g., nearly 70% success on AndroidWorld. However, these systems keep their training data closed and remain opaque about their task and trajectory synthesis recipes. We present OpenMobile, an open‑source framework that synthesizes high‑quality task instructions and agent trajectories, with two key components: (1) The first is a scalable task synthesis pipeline that constructs a global environment memory from exploration, then leverages it to generate diverse and grounded instructions. and (2) a policy‑switching strategy for trajectory rollout. By alternating between learner and expert models, it captures essential error‑recovery data often missing in standard imitation learning. Agents trained on our data achieve competitive results across three dynamic mobile agent benchmarks: notably, our fine‑tuned Qwen2.5‑VL and Qwen3‑VL reach 51.7% and 64.7% on AndroidWorld, far surpassing existing open‑data approaches. Furthermore, we conduct transparent analyses on the overlap between our synthetic instructions and benchmark test sets, and verify that performance gains stem from broad functionality coverage rather than benchmark overfitting. We release data and code at https://njucckevin.github.io/openmobile/ to bridge the data gap and facilitate broader mobile agent research.
Authors:Feifei Sang, Wei Lu, Hongruixuan Chen, Sibao Chen, Bin Luo
Abstract:
Building extraction from optical Remote Sensing (RS) imagery suffers from performance degradation under real‑world hazy and low‑light conditions. However, existing optical methods and benchmarks focus primarily on ideal clear‑weather conditions. While SAR offers all‑weather sensing, its side‑looking geometry causes geometric distortions. To address these challenges, we introduce HaLoBuilding, the first optical benchmark specifically designed for building extraction under hazy and low‑light conditions. By leveraging a same‑scene multitemporal pairing strategy, we ensure pixel‑level label alignment and high fidelity even under extreme degradation. Building upon this benchmark, we propose HaLoBuild‑Net, a novel end‑to‑end framework for building extraction in adverse RS scenarios. At its core, we develop a Spatial‑Frequency Focus Module (SFFM) to effectively mitigate meteorological interference on building features by coupling large receptive field attention with frequency‑aware channel reweighting guided by stable low‑frequency anchors. Additionally, a Global Multi‑scale Guidance Module (GMGM) provides global semantic constraints to anchor building topologies, while a Mutual‑Guided Fusion Module (MGFM) implements bidirectional semantic‑spatial calibration to suppress shallow noise and sharpen weather‑induced blurred boundaries. Extensive experiments demonstrate that HaLoBuild‑Net significantly outperforms state‑of‑the‑art methods and conventional cascaded restoration‑segmentation paradigms on the HaLoBuilding dataset, while maintaining robust generalization on WHU, INRIA, and LoveDA datasets. The source code and datasets are publicly available at: https://github.com/AeroVILab‑AHU/HaLoBuilding.
Authors:Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang, Jinjie Hu, Qiang Ji, Yihua Cao, Yihao Meng, Zhaoyue Cui, Mengmei Liu, Meng Meng, Jian Luan
Abstract:
Recent advances in video‑to‑audio (V2A) generation enable high‑quality audio synthesis from visual content, yet achieving robust and fine‑grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual‑text conflict and imprecise stylistic control due to entangled temporal and timbre information in reference audio. Moreover, the lack of standardized benchmarks limits systematic evaluation.
We propose ControlFoley, a unified multimodal V2A framework that enables precise control over video, text, and reference audio. We introduce a joint visual encoding paradigm that integrates CLIP with a spatio‑temporal audio‑visual encoder to improve alignment and textual controllability. We further propose temporal‑timbre decoupling to suppress redundant temporal cues while preserving discriminative timbre features. In addition, we design a modality‑robust training scheme with unified multimodal representation alignment (REPA) and random modality dropout. We also present VGGSound‑TVC, a benchmark for evaluating textual controllability under varying degrees of visual‑text conflict.
Extensive experiments demonstrate state‑of‑the‑art performance across multiple V2A tasks, including text‑guided, text‑controlled, and audio‑controlled generation. ControlFoley achieves superior controllability under cross‑modal conflict while maintaining strong synchronization and audio quality, and shows competitive or better performance compared to an industrial V2A system.
Code, models, datasets, and demos are available at: https://yjx‑research.github.io/ControlFoley/.
Authors:Yangchen Zeng, Zhenyu Yu, Dongming Jiang, Wenbo Zhang, Yifan Hong, Zhanhua Hu, Jiao Luo, Kangning Cui
Abstract:
Transformer‑based detectors have advanced small‑object detection, but they often remain inefficient and vulnerable to background‑induced query noise, which motivates deep decoders to refine low‑quality queries. We present HELP (Heatmap‑guided Embedding Learning Paradigm), a noise‑aware positional‑semantic fusion framework that studies where to embed positional information by selectively preserving positional encodings in foreground‑salient regions while suppressing background clutter. Within HELP, we introduce Heatmap‑guided Positional Embedding (HPE) as the core embedding mechanism and visualize it with a heatbar for interpretable diagnosis and fine‑tuning. HPE is integrated into both the encoder and decoder: it guides noise‑suppressed feature encoding by injecting heatmap‑aware positional encoding, and it enables high‑quality query retrieval by filtering background‑dominant embeddings via a gradient‑based mask filter before decoding. To address feature sparsity in complex small targets, we integrate Linear‑Snake Convolution to enrich retrieval‑relevant representations. The gradient‑based heatmap supervision is used during training only, incurring no additional gradient computation at inference. As a result, our design reduces decoder layers from eight to three and achieves a 59.4% parameter reduction (66.3M vs. 163M) while maintaining consistent accuracy gains under a reduced compute budget across benchmarks. Code Repository: https://github.com/yidimopozhibai/Noise‑Suppressed‑Query‑Retrieval
Authors:Fabrizio Guillaro, Vincenzo De Rosa, Davide Cozzolino, Luisa Verdoliva
Abstract:
Significant progress has been made in detecting synthetic images, however most existing approaches operate on a single image instance and overlook a key characteristic of real‑world dissemination: as viral images circulate on the web, multiple near‑duplicate versions appear and lose quality due to repeated operations like recompression, resizing and cropping. As a consequence, the same image may yield inconsistent forensic predictions based on which version has been analyzed. In this work, to address this issue we propose QuAD (Quality‑Aware calibration with near‑Duplicates) a novel framework that makes decisions based on all available near‑duplicates of the same image. Given a query, we retrieve its online near‑duplicates and feed them to a detector: the resulting scores are then aggregated based on the estimated quality of the corresponding instance. By doing so, we take advantage of all pieces of information while accounting for the reduced reliability of images impaired by multiple processing steps. To support large‑scale evaluation, we introduce two datasets: AncesTree, an in‑lab dataset of 136k images organized in stochastic degradation trees that simulate online reposting dynamics, and ReWIND, a real‑world dataset of nearly 10k near‑duplicate images collected from viral web content. Experiments on several state‑of‑the‑art detectors show that our quality‑aware fusion improves their performance consistently, with an average gain of around 8% in terms of balanced accuracy compared to plain average. Our results highlight the importance of jointly processing all the images available online to achieve reliable detection of AI‑generated content in real‑world applications. Code and data are publicly available at https://grip‑unina.github.io/QuAD/
Authors:Jianchao Huang, Fengming Zhang, Haibo Zhu, Tao Yan
Abstract:
Small object detection remains a significant challenge due to feature degradation from downsampling, mutual occlusion in dense clusters, and complex background interference. To address these issues, this paper proposes FSDETR, a frequency‑spatial feature enhancement framework built upon the RT‑DETR baseline. By establishing a collaborative modeling mechanism, the method effectively leverages complementary structural information. Specifically, a Spatial Hierarchical Attention Block (SHAB) captures both local details and global dependencies to strengthen semantic representation. Furthermore, to mitigate occlusion in dense scenes, the Deformable Attention‑based Intra‑scale Feature Interaction (DA‑AIFI) focuses on informative regions via dynamic sampling. Finally, the Frequency‑Spatial Feature Pyramid Network (FSFPN) integrates frequency filtering with spatial edge extraction via the Cross‑domain Frequency‑Spatial Block (CFSB) to preserve fine‑grained details. Experimental results show that with only 14.7M parameters, FSDETR achieves 13.9% APS on VisDrone 2019 and 48.95% AP50 tiny on TinyPerson, showing strong performance on small‑object benchmarks. The code and models are available at https://github.com/YT3DVision/FSDETR.
Authors:Meng-Xun Li, Wen-Hui Deng, Zhi-Xing Wu, Chun-Xiao Jin, Jia-Min Wu, Yue Han, James Kit Hon Tsoi, Gui-Song Xia, Cui Huang
Abstract:
Vision‑Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains largely underexplored due to the lack of fine‑grained, annotated datasets and comprehensive benchmarks. To address this, we present MetaDent, a comprehensive resource that includes (1) a novel and large‑scale dentistry image dataset collected from clinical, public, and web sources; (2) a semi‑structured annotation framework designed to capture the hierarchical and clinically nuanced nature of dental photography; and (3) comprehensive benchmark suites for evaluating state‑of‑the‑art VLMs on clinical image understanding. Our labeling approach combines a high‑level image summary with point‑by‑point, free‑text descriptions of abnormalities. This method enables rich, scalable, and task‑agnostic representations. We curated 60,669 dental images from diverse sources and annotated a representative subset of 2,588 images using this meta‑labeling scheme. Leveraging Large Language Models (LLMs), we derive standardized benchmarks: approximately 15K Visual Question Answering (VQA) pairs and an 18‑class multi‑label classification dataset, which we validated with human review and error analysis to justify that the LLM‑driven transition reliably preserves fidelity and semantic accuracy. We then evaluate state‑of‑the‑art VLMs across VQA, classification, and image captioning tasks. Quantitative results reveal that even the most advanced models struggle with a fine‑grained understanding of intraoral scenes, achieving moderate accuracy and producing inconsistent or incomplete descriptions in image captioning. We publicly release our dataset, annotations, and tools to foster reproducible research and accelerate the development of vision‑language systems for dental applications.
Authors:Haileab Yagersew
Abstract:
Retail theft costs the global economy over \100 billion annually, yet existing AI‑based detection systems require expensive custom model training on proprietary datasets and charge \200‑500/month per store. We present Paza, a zero‑shot retail theft detection framework that achieves practical concealment detection without training any model. Our approach orchestrates multiple existing models in a layered pipeline ‑ cheap object detection and pose estimation running continuously, with an expensive vision‑language model (VLM) invoked only when behavioral pre‑filters trigger. A multi‑signal suspicion pre‑filter (requiring dwell time plus at least one behavioral signal) reduces VLM invocations by 240x compared to per‑frame analysis, bounding calls to <=10/minute and enabling a single GPU to serve 10‑20 stores. The architecture is model‑agnostic: the VLM component accepts any OpenAI‑compatible endpoint, enabling operators to swap between models such as Gemma 4, Qwen3.5‑Omni, GPT‑4o, or future releases without code changes ‑ ensuring the system improves as the VLM landscape evolves. We evaluate the VLM component on the DCSASS synthesized shoplifting dataset (169 clips, controlled environment), achieving 89.5% precision and 92.8% specificity at 59.3% recall zero‑shot ‑ where the recall gap is attributable to sparse frame sampling in offline evaluation rather than VLM reasoning failures, as precision and specificity are the operationally critical metrics determining false alarm rates. We present a detailed cost model showing viability at \50‑100/month per store (3‑10x cheaper than commercial alternatives), and introduce a privacy‑preserving design that obfuscates faces in the detection pipeline. The source code is available at https://github.com/xHaileab/Paza‑AI.
Authors:Andrey Moskalenko, Alexey Bryncev, Ivan Kosmynin, Kira Shilovskaya, Mikhail Erofeev, Dmitry Vatolin, Radu Timofte, Kun Wang, Yupeng Hu, Zhiran Li, Hao Liu, Qianlong Xiang, Liqiang Nie, Konstantinos Chaldaiopoulos, Niki Efthymiou, Athanasia Zlatintsi, Panagiotis Filntisis, Katerina Pastra, Petros Maragos, Li Yang, Gen Zhan, Yiting Liao, Yabin Zhang, Yuxin Liu, Xu Wu, Yunheng Zheng, Linze Li, Kun He, Cong Wu, Xuefeng Zhu, Tianyang Xu, Xiaojun Wu, Wenzhuo Zhao, Keren Fu, Gongyang Li, Shixiang Shi, Jianlin Chen, Haibin Ling, Yaoxin Jiang, Guoyi Xu, Jiajia Liu, Yaokun Shi, Jiachen Tu
Abstract:
This paper presents an overview of the NTIRE 2026 Challenge on Video Saliency Prediction. The goal of the challenge participants was to develop automatic saliency map prediction methods for the provided video sequences. The novel dataset of 2,000 diverse videos with an open license was prepared for this challenge. The fixations and corresponding saliency maps were collected using crowdsourced mouse tracking and contain viewing data from over 5,000 assessors. Evaluation was performed on a subset of 800 test videos using generally accepted quality metrics. The challenge attracted over 20 teams making submissions, and 7 teams passed the final phase with code review. All data used in this challenge is made publicly available ‑ https://github.com/msu‑video‑group/NTIRE26_Saliency_Prediction.
Authors:Yuan Sun, Xuan Wang, WeiLi Zhang, Wenxuan Zhang, Yu Guo, Fei Wang
Abstract:
We propose a compositional method for constructing a complete 3D head avatar from a single image. Prior one‑shot holistic approaches frequently fail to produce realistic hair dynamics during animation, largely due to inadequate decoupling of hair from the facial region, resulting in entangled geometry and unnatural deformations. Our method explicitly decouples hair from the face, modeling these components using distinct deformation paradigms while integrating them into a unified rendering pipeline. Furthermore, by leveraging image‑to‑3D lifting techniques, we preserve fine‑grained textures from the input image to the greatest extent possible, effectively mitigating the common issue of high‑frequency information loss in generalized models. Specifically, given a frontal portrait image, we first perform hair removal to obtain a bald image. Both the original image and the bald image are then lifted to dense, detail‑rich 3D Gaussian Splatting (3DGS) representations. For the bald 3DGS, we rig it to a FLAME mesh via non‑rigid registration with a prior model, enabling natural deformation that follows the mesh triangles during animation. For the hair component, we employ semantic label supervision combined with a boundary‑aware reassignment strategy to extract a clean and isolated set of hair Gaussians. To control hair deformation, we introduce a cage structure that supports Position‑Based Dynamics (PBD) simulation, allowing realistic and physically plausible transformations of the hair Gaussian primitives under head motion, gravity, and inertial effects. Striking qualitative results, including dynamic animations under diverse head motions, gravity effects, and expressions, showcase substantially more realistic hair behavior alongside faithfully preserved facial details, outperforming state‑of‑the‑art one‑shot methods in perceptual realism.
Authors:Jordan Shipard, Arnold Wiliem, Kien Nguyen Thanh, Wei Xiang, Clinton Fookes
Abstract:
Generalized Category Discovery (GCD) challenges methods to identify known and novel classes using partially labeled data, mirroring human category learning. Unlike prior GCD methods, which operate within a single modality and require dataset‑specific fine‑tuning, we propose a modality‑agnostic GCD approach inspired by the human brain's abstract category formation. Our OmniGCD leverages modality‑specific encoders (e.g., vision, audio, text, remote sensing) to process inputs, followed by dimension reduction to construct a GCD latent space, which is transformed at test‑time into a representation better suited for clustering using a novel synthetically trained Transformer‑based model. To evaluate OmniGCD, we introduce a zero‑shot GCD setting where no dataset‑specific fine‑tuning is allowed, enabling modality‑agnostic category discovery. Trained once on synthetic data, OmniGCD performs zero‑shot GCD across 16 datasets spanning four modalities, improving classification accuracy for known and novel classes over baselines (average percentage point improvement of +6.2, +17.9, +1.5 and +12.7 for vision, text, audio and remote sensing). This highlights the importance of strong encoders while decoupling representation learning from category discovery. Improving modality‑agnostic methods will propagate across modalities, enabling encoder development independent of GCD. Our work serves as a benchmark for future modality‑agnostic GCD works, paving the way for scalable, human‑inspired category discovery. All code is available \hrefhttps://github.com/Jordan‑HS/OmniGCDhere
Authors:Yanguang Sun, Hengmin Zhang, Jianjun Qian, Jian Yang, Lei Luo
Abstract:
Early identification and removal of polyps can reduce the risk of developing colorectal cancer. However, the diverse morphologies, complex backgrounds and often concealed nature of polyps make polyp segmentation in colonoscopy images highly challenging. Despite the promising performance of existing deep learning‑based polyp segmentation methods, their perceptual capabilities remain biased toward local regions, mainly because of the strong spatial correlations between neighboring pixels in the spatial domain. This limitation makes it difficult to capture the complete polyp structures, ultimately leading to sub‑optimal segmentation results. In this paper, we propose a novel adaptive spectrum guidance network, called ASGNet, which addresses the limitations of spatial perception by integrating spectral features with global attributes. Specifically, we first design a spectrum‑guided non‑local perception module that jointly aggregates local and global information, therefore enhancing the discriminability of polyp structures, and refining their boundaries. Moreover, we introduce a multi‑source semantic extractor that integrates rich high‑level semantic information to assist in the preliminary localization of polyps. Furthermore, we construct a dense cross‑layer interaction decoder that effectively integrates diverse information from different layers and strengthens it to generate high‑quality representations for accurate polyp segmentation. Extensive quantitative and qualitative results demonstrate the superiority of our ASGNet approach over 21 state‑of‑the‑art methods across five widely‑used polyp segmentation benchmarks. The code will be publicly available at: https://github.com/CSYSI/ASGNet.
Authors:Jiyoung Lim, Heejae Yang, Jee-Hyong Lee
Abstract:
Composed Image Retrieval (CIR) aims to retrieve target images by integrating a reference image with a corresponding modification text. CIR requires jointly considering the explicit semantics specified in the query and the implicit semantics embedded within its bi‑modal composition. Recent training‑free Zero‑Shot CIR (ZS‑CIR) methods leverage Multimodal Large Language Models (MLLMs) to generate detailed target descriptions, converting the implicit information into explicit textual expressions. However, these methods rely heavily on the textual modality and fail to capture the fuzzy retrieval nature that requires considering diverse combinations of candidates. This leads to reduced diversity and accuracy in retrieval results. To address this limitation, we propose a novel training‑free method, Geodesic Mixup‑based Implicit semantic eXpansion and Explicit semantic Re‑ranking for ZS‑CIR (G‑MIXER). G‑MIXER constructs composed query features that reflect the implicit semantics of reference image‑text pairs through geodesic mixup over a range of mixup ratios, and builds a diverse candidate set. The generated candidates are then re‑ranked using explicit semantics derived from MLLMs, improving both retrieval diversity and accuracy. Our proposed G‑MIXER achieves state‑of‑the‑art performance across multiple ZS‑CIR benchmarks, effectively handling both implicit and explicit semantics without additional training. Our code will be available at https://github.com/maya0395/gmixer.
Authors:Yi He, Tao Wang, Yi Jin, Congyan Lang, Yidong Li, Haibin Ling
Abstract:
Recent advances in 3D Gaussian Splatting (3DGS) have enabled highly efficient and photorealistic novel view synthesis. However, segmenting objects accurately in 3DGS remains challenging due to the discrete nature of Gaussian representations, which often leads to aliasing and artifacts at object boundaries. In this paper, we introduce NG‑GS, a novel framework for high‑quality object segmentation in 3DGS that explicitly addresses boundary discretization. Our approach begins by automatically identifying ambiguous Gaussians at object boundaries using mask variance analysis. We then apply radial basis function (RBF) interpolation to construct a spatially continuous feature field, enhanced by multi‑resolution hash encoding for efficient multi‑scale representation. A joint optimization strategy aligns 3DGS with a lightweight NeRF module through alignment and spatial continuity losses, ensuring smooth and consistent segmentation boundaries. Extensive experiments on NVOS, LERF‑OVS, and ScanNet benchmarks demonstrate that our method achieves state‑of‑the‑art performance, with significant gains in boundary mIoU. Code is available at https://github.com/BJTU‑KD3D/NG‑GS.
Authors:Amir El-Ghoussani, Marc Hölle, Gustavo Carneiro, Vasileios Belagiannis
Abstract:
We address the problem of prompt‑guided image editing in visual autoregressive models. Given a source image and a target text prompt, we aim to modify the source image according to the target prompt, while preserving all regions which are unrelated to the requested edit. To this end, we present Masked Logit Nudging, which uses the source image token maps to introduce a guidance step that aligns the model's predictions under the target prompt with these source token maps. Specifically, we convert the fixed source encodings into logits using the VAR encoding, nudging the model's predicted logits towards the targets along a semantic trajectory defined by the source‑target prompts. Edits are applied only within spatial masks obtained through a dedicated masking scheme that leverages cross‑attention differences between the source and edited prompts. Then, we introduce a refinement to correct quantization errors and improve reconstruction quality. Our approach achieves the best image editing performance on the PIE benchmark at 512px and 1024px resolutions. Beyond editing, our method delivers faithful reconstructions and outperforms previous methods on COCO at 512px and OpenImages at 1024px. Overall, our method outperforms VAR‑related approaches and achieves comparable or even better performance than diffusion models, while being much faster. Code is available at 'https://github.com/AmirMaEl/MLN'.
Authors:Ruiqi Wang, Qi Yu, Jie Ma, Hanlin Wu
Abstract:
High‑resolution (HR) land‑cover mapping is often constrained by the high cost of dense HR annotations. We revisit this problem from the perspective of map super‑resolution, which enhances coarse low‑resolution (LR) land‑cover products into HR maps at the resolution of the input imagery. Existing weakly supervised methods can leverage LR labels, but they typically use them to retrain dense predictors with substantial computational cost. We propose MapSR, a prompt‑driven framework that decouples supervision from model training. MapSR uses LR labels once to extract class prompts from frozen vision foundation model features through a lightweight linear probe, after which HR mapping proceeds via training‑free metric inference and graph‑based prediction refinement. Specifically, class prompts are estimated by aggregating high‑confidence HR features identified by the linear probe, and HR predictions are obtained by cosine‑similarity matching followed by graph‑based propagation for spatial refinement. Experiments on the Chesapeake Bay dataset show that MapSR achieves 59.64% mIoU without any HR labels, remaining competitive with the strongest weakly supervised baseline and surpassing a fully supervised baseline. Notably, MapSR reduces trainable parameters by four orders of magnitude and shortens training time from hours to minutes, enabling scalable HR mapping under limited annotation and compute budgets. The code is available at https://github.com/rikirikirikiriki/MapSR.
Authors:Yixu Huang, Tinghui Zhu, Muhao Chen
Abstract:
Visual reasoning models (VRMs) have recently shown strong cross‑modal reasoning capabilities by integrating visual perception with language reasoning. However, they often suffer from overthinking, producing unnecessarily long reasoning chains for any tasks. We attribute this issue to Reasoning Path Redundancy in visual reasoning: many visual questions do not require the full reasoning process. To address this, we propose AVR, an adaptive visual reasoning framework that decomposes visual reasoning into three cognitive functions: visual perception, logical reasoning, and answer application. It further enables models to dynamically choose among three response formats: Full Format, Perception‑Only Format, and Direct Answer. AVR is trained with FS‑GRPO, an adaptation of Group Relative Policy Optimization that encourages the model to select the most efficient reasoning format while preserving correctness. Experiments on multiple vision‑language benchmarks show that AVR reduces token usage by 50‑‑90% while maintaining overall accuracy, especially in perception‑intensive tasks. These results demonstrate that adaptive visual reasoning can effectively mitigate overthinking in VRMs. Code and data are available at: https://github.com/RunRiotComeOn/AVR.
Authors:Mingqian Ji, Shanshan Zhang, Jian Yang
Abstract:
Vision Transformer (ViT)‑based sparse multi‑view 3D object detectors have achieved remarkable accuracy but still suffer from high inference latency due to heavy token processing. To accelerate these models, token compression has been widely explored. However, our revisit of existing strategies, such as token pruning, merging, and patch size enlargement, reveals that they often discard informative background cues, disrupt contextual consistency, and lose fine‑grained semantics, negatively affecting 3D detection. To overcome these limitations, we propose SEPatch3D, a novel framework that dynamically adjusts patch sizes while preserving critical semantic information within coarse patches. Specifically, we design Spatiotemporal‑aware Patch Size Selection (SPSS) that assigns small patches to scenes containing nearby objects to preserve fine details and large patches to background‑dominated scenes to reduce computation cost. To further mitigate potential detail loss, Informative Patch Selection (IPS) selects the informative patches for feature refinement, and Cross‑Granularity Feature Enhancement (CGFE) injects fine‑grained details into selected coarse patches, enriching semantic features. Experiments on the nuScenes and Argoverse 2 validation sets show that SEPatch3D achieves up to 57% faster inference than the StreamPETR baseline and 20% higher efficiency than the state‑of‑the‑art ToC3D‑faster, while preserving comparable detection accuracy. Code is available at https://github.com/Mingqj/SEPatch3D.
Authors:Zheng Chen, Bowen Chai, Rongjun Gao, Mingtao Nie, Xi Li, Bingnan Duan, Jianping Fang, Xiaohong Liu, Linghe Kong, Yulun Zhang
Abstract:
Video face restoration aims to enhance degraded face videos into high‑quality results with realistic facial details, stable identity, and temporal coherence. Recent diffusion‑based methods have brought strong generative priors to restoration and enabled more realistic detail synthesis. However, existing approaches for face videos still rely heavily on generic diffusion priors and multi‑step sampling, which limit both facial adaptation and inference efficiency. These limitations motivate the use of one‑step diffusion for video face restoration, yet achieving faithful facial recovery alongside temporally stable outputs remains challenging. In this paper, we propose, DVFace, a one‑step diffusion framework for real‑world video face restoration. Specifically, we introduce a spatio‑temporal dual‑codebook design to extract complementary spatial and temporal facial priors from degraded videos. We further propose an asymmetric spatio‑temporal fusion module to inject these priors into the diffusion backbone according to their distinct roles. Evaluation on various benchmarks shows that DVFace delivers superior restoration quality, temporal consistency, and identity preservation compared to recent methods. Code: https://github.com/zhengchen1999/DVFace.
Authors:Zheng Chen, Kai Liu, Jingkai Wang, Xianglong Yan, Jianze Li, Ziqing Zhang, Jue Gong, Jiatong Li, Lei Sun, Xiaoyang Liu, Radu Timofte, Yulun Zhang, Jihye Park, Yoonjin Im, Hyungju Chun, Hyunhee Park, MinKyu Park, Zheng Xie, Xiangyu Kong, Weijun Yuan, Zhan Li, Qiurong Song, Luen Zhu, Fengkai Zhang, Xinzhe Zhu, Junyang Chen, Congyu Wang, Yixin Yang, Zhaorun Zhou, Jiangxin Dong, Jinshan Pan, Shengwei Wang, Jiajie Ou, Baiang Li, Sizhuo Ma, Qiang Gao, Jusheng Zhang, Jian Wang, Keze Wang, Yijiao Liu, Yingsi Chen, Hui Li, Yu Wang, Congchao Zhu, Saeed Ahmad, Ik Hyun Lee, Jun Young Park, Ji Hwan Yoon, Kainan Yan, Zian Wang, Weibo Wang, Shihao Zou, Chao Dong, Wei Zhou, Linfeng Li, Jaeseong Lee, Jaeho Chae, Jinwoo Kim, Seonjoo Kim, Yucong Hong, Zhenming Yan, Junye Chen, Ruize Han, Song Wang, Yuxuan Jiang, Chengxi Zeng, Tianhao Peng, Fan Zhang, David Bull, Tongyao Mu, Qiong Cao, Yifan Wang, Youwei Pan, Leilei Cao, Xiaoping Peng, Wei Deng, Yifei Chen, Wenbo Xiong, Xian Hu, Yuxin Zhang, Xiaoyun Cheng, Yang Ji, Zonghao Chen, Zhihao Xue, Junqin Hu, Nihal Kumar, Snehal Singh Tomar, Klaus Mueller, Surya Vashisth, Prateek Shaily, Jayant Kumar, Hardik Sharma, Ashish Negi, Sachin Chaudhary, Akshay Dudhane, Praful Hambarde, Amit Shukla, Shijun Shi, Jiangning Zhang, Yong Liu, Kai Hu, Jing Xu, Xianfang Zeng, Amitesh M, Hariharan S, Chia-Ming Lee, Yu-Fan Lin, Chih-Chung Hsu, Nishalini K, Sreenath K A, Bilel Benjdira, Anas M. Ali, Wadii Boulila, Shuling Zheng, Zhiheng Fu, Feng Zhang, Zhanglu Chen, Boyang Yao, Nikhil Pathak, Aagam Jain, Milan Kumar, Kishor Upla, Vivek Chavda, Sarang N S, Raghavendra Ramachandra, Zhipeng Zhang, Qi Wang, Shiyu Wang, Jiachen Tu, Guoyi Xu, Yaoxin Jiang, Jiajia Liu, Yaokun Shi, Yuqi Li, Chuanguang Yang, Weilun Feng, Zhuzhi Hong, Hao Wu, Junming Liu, Yingli Tian, Amish Bhushan Kulkarni, Tejas R R Shet, Saakshi M Vernekar, Nikhil Akalwadi, Kaushik Mallibhat, Ramesh Ashok Tabib, Uma Mudenagudi, Yuwen Pan, Tianrun Chen, Deyi Ji, Qi Zhu, Lanyun Zhu, Heyan Zhangyi
Abstract:
This paper presents the NTIRE 2026 image super‑resolution (×4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high‑resolution (HR) images from low‑resolution (LR) inputs generated through bicubic downsampling with a ×4 scaling factor. The objective is to develop effective super‑resolution solutions and analyze recent advances in the field. To reflect the evolving objectives of image super‑resolution, the challenge includes two tracks: (1) a restoration track, which emphasizes pixel‑wise fidelity and ranks submissions based on PSNR; and (2) a perceptual track, which focuses on visual realism and evaluates results using a perceptual score. A total of 194 participants registered for the challenge, with 31 teams submitting valid entries. This report summarizes the challenge design, datasets, evaluation protocol, main results, and methods of participating teams. The challenge provides a unified benchmark and offers insights into current progress and future directions in image super‑resolution.
Authors:Biwei Dai, Po-Wen Chang, Wahid Bhimji, Paolo Calafiura, Ragansu Chakkappai, Yuan-Tang Chou, Sascha Diefenbacher, Jordan Dudley, Ibrahim Elsharkawy, Steven Farrell, Isabelle Guyon, Chris Harris, Elham E Khoda, Benjamin Nachman, David Rousseau, Uroš Seljak, Ihsan Ullah, Yulei Zhang
Abstract:
Weak gravitational lensing, the correlated distortion of background galaxy shapes by foreground structures, is a powerful probe of the matter distribution in our universe and allows accurate constraints on the cosmological model. In recent years, high‑order statistics and machine learning (ML) techniques have been applied to weak lensing data to extract the nonlinear information beyond traditional two‑point analysis. However, these methods typically rely on cosmological simulations, which poses several challenges: simulations are computationally expensive, limiting most realistic setups to a low training data regime; inaccurate modeling of systematics in the simulations create distribution shifts that can bias cosmological parameter constraints; and varying simulation setups across studies make method comparison difficult. To address these difficulties, we present the first weak lensing benchmark dataset with several realistic systematics and launch the FAIR Universe Weak Lensing Machine Learning Uncertainty Challenge. The challenge focuses on measuring the fundamental properties of the universe from weak lensing data with limited training set and potential distribution shifts, while providing a standardized benchmark for rigorous comparison across methods. Organized in two phases, the challenge will bring together the physics and ML communities to advance the methodologies for handling systematic uncertainties, data efficiency, and distribution shifts in weak lensing analysis with ML, ultimately facilitating the deployment of ML approaches into upcoming weak lensing survey analysis.
Authors:Team HY-World, Chenjie Cao, Xuhui Zuo, Zhenwei Wang, Yisu Zhang, Junta Wu, Zhenyang Liu, Yuning Gong, Yang Liu, Bo Yuan, Chao Zhang, Coopers Li, Dongyuan Guo, Fan Yang, Haiyu Zhang, Hang Cao, Jianchen Zhu, Jiaxin Lin, Jie Xiao, Jihong Zhang, Junlin Yu, Lei Wang, Lifu Wang, Lilin Wang, Linus, Minghui Chen, Peng He, Penghao Zhao, Qi Chen, Rui Chen, Rui Shao, Sicong Liu, Wangchen Qin, Xiaochuan Niu, Xiang Yuan, Yi Sun, Yifei Tang, Yifu Sun, Yihang Lian, Yonghao Tan, Yuhong Liu, Yuyang Yin, Zhiyuan Min, Tengfei Wang, Chunchao Guo
Abstract:
We introduce HY‑World 2.0, a multi‑modal world model framework that advances our prior project HY‑World 1.0. HY‑World 2.0 accommodates diverse input modalities, including text prompts, single‑view images, multi‑view images, and videos, and produces 3D world representations. With text or single‑view image inputs, the model performs world generation, synthesizing high‑fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four‑stage method: a) Panorama Generation with HY‑Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe‑based view generation model with consistent memory. We also upgrade WorldMirror, a feed‑forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi‑view images or videos. Also, we introduce WorldLens, a high‑performance 3DGS rendering platform featuring a flexible engine‑agnostic architecture, automatic IBL lighting, efficient collision detection, and training‑rendering co‑design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY‑World 2.0 achieves state‑of‑the‑art performance on several benchmarks among open‑source approaches, delivering results comparable to the closed‑source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models.
Authors:Lin-Zhuo Chen, Jian Gao, Yihang Chen, Ka Leong Cheng, Yipengjing Sun, Liangxiao Hu, Nan Xue, Xing Zhu, Yujun Shen, Yao Yao, Yinghao Xu
Abstract:
Streaming 3D reconstruction aims to recover 3D information, such as camera poses and point clouds, from a video stream, which necessitates geometric accuracy, temporal consistency, and computational efficiency. Motivated by the principles of Simultaneous Localization and Mapping (SLAM), we introduce LingBot‑Map, a feed‑forward 3D foundation model for reconstructing scenes from streaming data, built upon a geometric context transformer (GCT) architecture. A defining aspect of LingBot‑Map lies in its carefully designed attention mechanism, which integrates an anchor context, a pose‑reference window, and a trajectory memory to address coordinate grounding, dense geometric cues, and long‑range drift correction, respectively. This design keeps the streaming state compact while retaining rich geometric context, enabling stable efficient inference at around 20 FPS on 518 x 378 resolution inputs over long sequences exceeding 10,000 frames. Extensive evaluations across a variety of benchmarks demonstrate that our approach achieves superior performance compared to both existing streaming and iterative optimization‑based approaches.
Authors:Tianshuo Yang, Guanyu Chen, Yutian Chen, Zhixuan Liang, Yitian Liu, Zanxin Chen, Chunpu Xu, Haotian Liang, Jiangmiao Pang, Yao Mu, Ping Luo
Abstract:
While end‑to‑end Vision‑Language‑Action (VLA) models offer a promising paradigm for robotic manipulation, fine‑tuning them on narrow control data often compromises the profound reasoning capabilities inherited from their base Vision‑Language Models (VLMs). To resolve this fundamental trade‑off, we propose HiVLA, a visual‑grounded‑centric hierarchical framework that explicitly decouples high‑level semantic planning from low‑level motor control. In high‑level part, a VLM planner first performs task decomposition and visual grounding to generate structured plans, comprising a subtask instruction and a precise target bounding box. Then, to translate this plan into physical actions, we introduce a flow‑matching Diffusion Transformer (DiT) action expert in low‑level part equipped with a novel cascaded cross‑attention mechanism. This design sequentially fuses global context, high‑resolution object‑centric crops and skill semantics, enabling the DiT to focus purely on robust execution. Our decoupled architecture preserves the VLM's zero‑shot reasoning while allowing independent improvement of both components. Extensive experiments in simulation and the real world demonstrate that HiVLA significantly outperforms state‑of‑the‑art end‑to‑end baselines, particularly excelling in long‑horizon skill composition and the fine‑grained manipulation of small objects in cluttered scenes.
Authors:Fei Tang, Bofan Chen, Zhengxi Lu, Tongbo Chen, Songqin Nong, Tao Jiang, Wenhao Xu, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
Abstract:
GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layouts. Test‑time zoom‑in methods improve localization by cropping and re‑running inference at higher resolution, but apply cropping uniformly across all instances with fixed crop sizes, ignoring whether the model is actually uncertain on each case. We propose UI‑Zoomer, a training‑free adaptive zoom‑in framework that treats both the trigger and scale of zoom‑in as a prediction uncertainty quantification problem. A confidence‑aware gate fuses spatial consensus among stochastic candidates with token‑level generation confidence to selectively trigger zoom‑in only when localization is uncertain. When triggered, an uncertainty‑driven crop sizing module decomposes prediction variance into inter‑sample positional spread and intra‑sample box extent, deriving a per‑instance crop radius via the law of total variance. Extensive experiments on ScreenSpot‑Pro, UI‑Vision, and ScreenSpot‑v2 demonstrate consistent improvements over strong baselines across multiple model architectures, achieving gains of up to +13.4%, +10.3%, and +4.2% respectively, with no additional training required.
Authors:Francesco Tonini, Alessandro Conti, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci
Abstract:
Human‑Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time, limiting their applicability to static environments. With the advent of Multimodal Large Language Models (MLLMs), it has become feasible to explore more flexible paradigms for interaction recognition. In this work, we revisit HOI detection through the lens of MLLMs and apply them to in‑the‑wild HOI detection. We define the Unconstrained HOI (U‑HOI) task, a novel HOI domain that removes the requirement for a predefined list of interactions at both training and inference. We evaluate a range of MLLMs on this setting and introduce a pipeline that includes test‑time inference and language‑to‑graph conversion to extract structured interactions from free‑form text. Our findings highlight the limitations of current HOI detectors and the value of MLLMs for U‑HOI. Code will be available at https://github.com/francescotonini/anyhoi
Authors:Jiun Tian Hoe, Weipeng Hu, Xudong Jiang, Yap-Peng Tan, Chee Seng Chan
Abstract:
Human‑Object Interaction (HOI) modelling captures how humans act upon and relate to objects, typically expressed as <person, action, object> triplets. Existing approaches split into two disjoint families: HOI generation synthesises scenes from structured triplets and layout, but fails to integrate mixed conditions like HOI and object‑only entities; and HOI editing modifies interactions via text, yet struggles to decouple pose from physical contact and scale to multiple interactions. We introduce OneHOI, a unified diffusion transformer framework that consolidates HOI generation and editing into a single conditional denoising process driven by shared structured interaction representations. At its core, the Relational Diffusion Transformer (R‑DiT) models verb‑mediated relations through role‑ and instance‑aware HOI tokens, layout‑based spatial Action Grounding, a Structured HOI Attention to enforce interaction topology, and HOI RoPE to disentangle multi‑HOI scenes. Trained jointly with modality dropout on our HOI‑Edit‑44K, along with HOI and object‑centric datasets, OneHOI supports layout‑guided, layout‑free, arbitrary‑mask, and mixed‑condition control, achieving state‑of‑the‑art results across both HOI generation and editing. Code is available at https://jiuntian.github.io/OneHOI/.
Authors:Yuhang Dai, Xingyi Yang
Abstract:
Feed‑forward 3D reconstruction models are efficient but rigid: once trained, they perform inference in a zero‑shot manner and cannot adapt to the test scene. As a result, visually plausible reconstructions often contain errors, particularly under occlusions, specularities, and ambiguous cues. To address this, we introduce Free Geometry, a framework that enables feed‑forward 3D reconstruction models to self‑evolve at test time without any 3D ground truth. Our key insight is that, when the model receives more views, it produces more reliable and view‑consistent reconstructions. Leveraging this property, given a testing sequence, we mask a subset of frames to construct a self‑supervised task. Free Geometry enforces cross‑view feature consistency between representations from full and partial observations, while maintaining the pairwise relations implied by the held‑out frames. This self‑supervision allows for fast recalibration via lightweight LoRA updates, taking less than 2 minutes per dataset on a single GPU. Our approach consistently improves state‑of‑the‑art foundation models, including Depth Anything 3 and VGGT, across 4 benchmark datasets, yielding an average improvement of 3.73% in camera pose accuracy and 2.88% in point map prediction. Code is available at https://github.com/hiteacherIamhumble/Free‑Geometry .
Authors:Xiaomin Li, Tala Wang, Zichen Zhong, Ying Zhang, Zirui Zheng, Takashi Isobe, Dezhuang Li, Huchuan Lu, You He, Xu Jia
Abstract:
Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual clues for accurate reasoning. Yet, current benchmarks predominantly aim at evaluating MLLMs' pre‑existing knowledge or perceptual understanding, often neglecting the critical capability of reasoning. To bridge this gap, we introduce DailyClue, a benchmark designed for visual clue‑driven reasoning in daily scenarios. Our construction is guided by two core principles: (1) strict grounding in authentic daily activities, and (2) challenging query design that necessitates more than surface‑level perception. Instead of simple recognition, our questions compel MLLMs to actively explore suitable visual clues and leverage them for subsequent reasoning. To this end, we curate a comprehensive dataset spanning four major daily domains and 16 distinct subtasks. Comprehensive evaluation across MLLMs and agentic models underscores the formidable challenge posed by our benchmark. Our analysis reveals several critical insights, emphasizing that the accurate identification of visual clues is essential for robust reasoning.
Authors:Weijie Wang, Qihang Cao, Sensen Gao, Donny Y. Chen, Haofei Xu, Wenjing Bian, Songyou Peng, Tat-Jen Cham, Chuanxia Zheng, Andreas Geiger, Jianfei Cai, Jia-Wang Bian, Bohan Zhuang
Abstract:
Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and interacting with the physical world. While traditional methods achieve high fidelity, they are limited by slow per‑scene optimization or category‑specific training, which hinders their practical deployment and scalability. Hence, generalizable feed‑forward 3D reconstruction has witnessed rapid development in recent years. By learning a model that maps images directly to 3D representations in a single forward pass, these methods enable efficient reconstruction and robust cross‑scene generalization. Our survey is motivated by a critical observation: despite the diverse geometric output representations, ranging from implicit fields to explicit primitives, existing feed‑forward approaches share similar high‑level architectural patterns, such as image feature extraction backbones, multi‑view information fusion mechanisms, and geometry‑aware design principles. Consequently, we abstract away from these representation differences and instead focus on model design, proposing a novel taxonomy centered on model design strategies that are agnostic to the output format. Our proposed taxonomy organizes the research directions into five key problems that drive recent research development: feature enhancement, geometry awareness, model efficiency, augmentation strategies and temporal‑aware models. To support this taxonomy with empirical grounding and standardized evaluation, we further comprehensively review related benchmarks and datasets, and extensively discuss and categorize real‑world applications based on feed‑forward 3D models. Finally, we outline future directions to address open challenges such as scalability, evaluation standards, and world modeling.
Authors:Enzhuo Zhang, Sijie Zhao, Dilxat Muhtar, Zhenshi Li, Xueliang Zhang, Pengfeng Xiao
Abstract:
Generative diffusion priors have recently achieved state‑of‑the‑art performance in natural image super‑resolution, demonstrating a powerful capability to synthesize photorealistic details. However, their direct application to remote sensing image super‑resolution (RSISR) reveals significant shortcomings. Unlike natural images, remote sensing images exhibit a unique texture distribution where ground objects are globally stochastic yet locally clustered, leading to highly imbalanced textures. This imbalance severely hinders the model's spatial perception. To address this, we propose TexADiff, a novel framework that begins by estimating a Relative Texture Density Map (RTDM) to represent the texture distribution. TexADiff then leverages this RTDM in three synergistic ways: as an explicit spatial conditioning to guide the diffusion process, as a loss modulation term to prioritize texture‑rich regions, and as a dynamic adapter for the sampling schedule. These modifications are designed to endow the model with explicit texture‑aware capabilities. Experiments demonstrate that TexADiff achieves superior or competitive quantitative metrics. Furthermore, qualitative results show that our model generates faithful high‑frequency details while effectively suppressing texture hallucinations. This improved reconstruction quality also results in significant gains in downstream task performance. The source code of our method can be found at https://github.com/ZezFuture/TexAdiff.
Authors:Jianlin Xiang, Linhui Dai, Xue Yang, Chaolei Yang, Yanshan Li
Abstract:
Interpretability is essential for deploying object detection systems in critical applications, especially under low‑quality imaging conditions that degrade visual information and increase prediction uncertainty. Existing methods either enhance image quality or design complex architectures, but often lack interpretability and fail to improve semantic discrimination. In contrast, prototype learning enables interpretable modeling by associating features with class‑centered semantics, which can provide more stable and interpretable representations under degradation. Motivated by this, we propose HiProto, a new paradigm for interpretable object detection based on hierarchical prototype learning. By constructing structured prototype representations across multiple feature levels, HiProto effectively models class‑specific semantics, thereby enhancing both semantic discrimination and interpretability. Building upon prototype modeling, we first propose a Region‑to‑Prototype Contrastive Loss (RPC‑Loss) to enhance the semantic focus of prototypes on target regions. Then, we propose a Prototype Regularization Loss (PR‑Loss) to improve the distinctiveness among class prototypes. Finally, we propose a Scale‑aware Pseudo Label Generation Strategy (SPLGS) to suppress mismatched supervision for RPC‑Loss, thereby preserving the robustness of low‑level prototype representations. Experiments on ExDark, RTTS, and VOC2012‑FOG demonstrate that HiProto achieves competitive results while offering clear interpretability through prototype responses, without relying on image enhancement or complex architectures. Our code will be available at https://github.com/xjlDestiny/HiProto.git.
Authors:Felicia Bader, Philipp Seeböck, Anastasia Bartashova, Ulrike Attenberger, Georg Langs
Abstract:
In diagnostic reports, experts encode complex imaging data into clinically actionable information. They describe subtle pathological findings that are meaningful in their anatomical context. Reports follow relatively consistent structures, expressing diagnostic information with few words that are often associated with tiny but consequential image observations. Standard vision language models struggle to identify the associations between these informative text components and small locations in the images. Here, we propose "MApLe", a multi‑task, multi‑instance vision language alignment approach that overcomes these limitations. It disentangles the concepts of anatomical region and diagnostic finding, and links local image information to sentences in a patch‑wise approach. Our method consists of a text embedding trained to capture anatomical and diagnostic concepts in sentences, a patch‑wise image encoder conditioned on anatomical structures, and a multi‑instance alignment of these representations. We demonstrate that MApLe can successfully align different image regions and multiple diagnostic findings in free‑text reports. We show that our model improves the alignment performance compared to state‑of‑the‑art baseline models when evaluated on several downstream tasks. The code is available at https://github.com/cirmuw/MApLe.
Authors:Songlin Du, Xiaoyong Lu, Yaping Yan, Guobao Xiao, Xiaobo Lu, Takeshi Ikenaga
Abstract:
Local feature matching plays a critical role in understanding the correspondence between cross‑view images. However, traditional methods are constrained by the inherent local nature of feature descriptors, limiting their ability to capture non‑local scene information that is essential for accurate cross‑view correspondence. In this paper, we introduce SceneGlue, a scene‑aware feature matching framework designed to overcome these limitations. SceneGlue leverages a hybridizable matching paradigm that integrates implicit parallel attention and explicit cross‑view visibility estimation. The parallel attention mechanism simultaneously exchanges information among local descriptors within and across images, enhancing the scene's global context. To further enrich the scene awareness, we propose the Visibility Transformer, which explicitly categorizes features into visible and invisible regions, providing an understanding of cross‑view scene visibility. By combining explicit and implicit scene‑level awareness, SceneGlue effectively compensates for the local descriptor constraints. Notably, SceneGlue is trained using only local feature matches, without requiring scene‑level groundtruth annotations. This scene‑aware approach not only improves accuracy and robustness but also enhances interpretability compared to traditional methods. Extensive experiments on applications such as homography estimation, pose estimation, image matching, and visual localization validate SceneGlue's superior performance. The source code is available at https://github.com/songlin‑du/SceneGlue.
Authors:Shuyun Wang, Hu Zhang, Xin Shen, Dadong Wang, Xin Yu
Abstract:
Bitstream‑corrupted video recovery aims to restore realistic content degraded during video storage or transmission. Existing methods typically assume that predefined masks of corrupted regions are available, but manually annotating these masks is labor‑intensive and impractical in real‑world scenarios. To address this limitation, we introduce a new blind video recovery setting that removes the reliance on predefined masks. This setting presents two major challenges: accurately identifying corrupted regions and recovering content from extensive and irregular degradations. We propose a Metadata‑Guided Diffusion Model (M‑GDM) to tackle these challenges. Specifically, intrinsic video metadata are leveraged as corruption indicators through a dual‑stream metadata encoder that separately embeds motion vectors and frame types before fusing them into a unified representation. This representation interacts with corrupted latent features via cross‑attention at each diffusion step. To preserve intact regions, we design a prior‑driven mask predictor that generates pseudo masks using both metadata and diffusion priors, enabling the separation and recombination of intact and recovered regions through hard masking. To mitigate boundary artifacts caused by imperfect masks, a post‑refinement module enhances consistency between intact and recovered regions. Extensive experiments demonstrate the effectiveness of our method and its superiority in blind video recovery. Code is available at: https://github.com/Shuyun‑Wang/M‑GDM.
Authors:Zhiyuan Xu, Jiuming Liu, Yuxin Chen, Masayoshi Tomizuka, Chenfeng Xu, Chensheng Peng
Abstract:
We present SparseGen, a novel framework for efficient image‑to‑3D generation, which exhibits low input‑view bias while being significantly faster. Unlike traditional approaches that rely on dense volumetric grids, triplanes, or pixel‑aligned primitives, we model scenes with a compact sparse set of learned 3D anchor queries and a learned expansion operator that decodes each transformed query into a small local set of 3D Gaussian primitives. Trained under a rectified‑flow reconstruction objective without 3D supervision, our model learns to allocate representation capacity where geometry and appearance matter, achieving significant reductions in memory and inference time while preserving multi‑view fidelity. We introduce quantitative measures of input‑view bias and utilization to show that sparse queries reduce overfitting to conditioning views while being representationally efficient. Our results argue that sparse set‑latent expansion is a principled, practical alternative for efficient 3D generative modeling.
Authors:Hye Jin Rhee, Joseph Damilola Akinyemi
Abstract:
Accurate and resource‑efficient automated diagnosis is a cornerstone of modern agricultural expert systems. While Convolutional Neural Networks (CNNs) have established benchmarks in plant pathology, their ability to capture long‑range spatial dependencies is often limited by standard pooling layers, and their high memory footprint hinders deployment on portable devices. This paper proposes a lightweight hybrid CNN‑LSTM system for bean leaf disease classification. By integrating an LSTM layer to model the spatial‑sequential relationships within feature maps, our hybrid architecture achieves a 94.38% accuracy while maintaining an exceptionally small footprint of 1.86 MB; a 70% reduction in size compared to traditional CNN‑based systems. Furthermore, we provide a systematic evaluation of image augmentation strategies, demonstrating that tailored transformations are superior to generic combinations for maintaining the integrity of diagnostic patterns. Results on the ibean dataset confirm that the proposed system achieves state‑of‑the‑art F1 scores of 99.22% with EfficientNet‑B7+LSTM, providing a robust and scalable framework for real‑time agricultural decision support in resource‑constrained environments. The code and augmented datasets used in this study are publicly available on this \hrefhttps://github.com/HJin‑R/bean_diseaseGithub repo.
Authors:Arya Shah, Vaibhav Tripathi, Mayank Singh, Chaklam Silpasuwanchai
Abstract:
Vision‑language models are increasingly deployed in high‑stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood, particularly in relation to how these models represent visual information internally. Whether models whose visual representations more closely mirror human neural processing are also more resistant to adversarial pressure is an open question with implications for both neuroscience and AI safety. We investigate this question by evaluating 12 open‑weight vision‑language models spanning 6 architecture families and a 40× parameter range (256M‑‑10B) along two axes: brain alignment, measured by predicting fMRI responses from the Natural Scenes Dataset across 8 human subjects and 6 visual cortex regions of interest, and sycophancy, measured through 76,800 two‑turn gaslighting prompts spanning 5 categories and 10 difficulty levels. Region‑of‑interest analysis reveals that alignment specifically in early visual cortex (V1‑‑V3) is a reliable negative predictor of sycophancy (r = ‑0.441, BCa 95% CI [‑0.740, ‑0.031]), with all 12 leave‑one‑out correlations negative and the strongest effect for existence denial attacks (r = ‑0.597, p = 0.040). This anatomically specific relationship is absent in higher‑order category‑selective regions, suggesting that faithful low‑level visual encoding provides a measurable anchor against adversarial linguistic override in vision‑language models. We release our code on \hrefhttps://github.com/aryashah2k/Gaslight‑Gatekeep‑Sycophantic‑ManipulationGitHub and dataset on \hrefhttps://huggingface.co/datasets/aryashah00/Gaslight‑Gatekeep‑V1‑V3Hugging Face
Authors:Chen Wang, Yixin Zhu, Yongbin Zhu, Fengyuan Shi, Qi Li, Jun Wang, Zuozhu Liu, Keli Hu
Abstract:
Accurate lesion segmentation in ultrasound images is essential for preventive screening and clinical diagnosis, yet remains challenging due to low contrast, blurry boundaries, and significant scale variations. Although existing deep learning‑based methods have achieved remarkable performance, these methods still struggle with scale variations and indistinct tumor boundaries. To address these challenges, we propose a progressive boundary enhanced U‑Net (PBE‑UNet). Specially, we first introduce a scale‑aware aggregation module (SAAM) that dynamically adjusts its receptive field to capture robust multi‑scale contextual information. Then, we propose a boundary‑guided feature enhancement (BGFE) module to enhance the feature representations. We find that there are large gaps between the narrow boundary and the wide segmentation error areas. Unlike existing methods that treat boundaries as static masks, the BGFE module progressively expands the narrow boundary prediction into broader spatial attention maps. Thus, broader spatial attention maps could effectively cover the wider segmentation error regions and enhance the model's focus on these challenging areas. We conduct expensive experiments on four benchmark ultrasound datasets, BUSI, Dataset B, TN3K, and BP. The experimental results how that our proposed PBE‑UNet outperforms state‑of‑the‑art ultrasound image segmentation methods. The code is at https://github.com/cruelMouth/PBE‑UNet.
Authors:Jaejoon Yoo, SuBeen Lee, Yerim Jeon, Miso Lee, Jae-Pil Heo
Abstract:
3D Single Object Tracking (3D‑SOT) aims to localize a target object across a sequence of LiDAR point clouds, given its 3D bounding box in the first frame. Recent methods have adopted a memory‑based approach to utilize previously observed features of the target object, but remain limited to only a few recent frames. This work reveals that their temporal capacity is fundamentally constrained to short‑term context due to severe temporal feature inconsistency and excessive memory overhead. To this end, we propose a robust long‑term 3D‑SOT framework, ChronoTrack, which preserves the temporal feature consistency while efficiently aggregating the diverse target features via long‑term memory. Based on a compact set of learnable memory tokens, ChronoTrack leverages long‑term information through two complementary objectives: a temporal consistency loss and a memory cycle consistency loss. The former enforces feature alignment across frames, alleviating temporal drift and improving the reliability of proposed long‑term memory. In parallel, the latter encourages each token to encode diverse and discriminative target representations observed throughout the sequence via memory‑point‑memory cyclic walks. As a result, ChronoTrack achieves new state‑of‑the‑art performance on multiple 3D‑SOT benchmarks, demonstrating its effectiveness in long‑term target modeling with compact memory while running at real‑time speed of 42 FPS on a single RTX 4090 GPU. The code is available at https://github.com/ujaejoon/ChronoTrack
Authors:Svetlana Pavlitska, Haixi Fan, Konstantin Ditschuneit, J. Marius Zöllner
Abstract:
Sparse mixture‑of‑experts (MoE) layers have been shown to substantially increase model capacity without a proportional increase in computational cost and are widely used in transformer architectures, where they typically replace feed‑forward network blocks. In contrast, integrating sparse MoE layers into convolutional neural networks (CNNs) remains inconsistent, with most prior work focusing on fine‑grained MoEs operating at the filter or channel levels. In this work, we investigate a coarser, patch‑wise formulation of sparse MoE layers for semantic segmentation, where local regions are routed to a small subset of convolutional experts. Through experiments on the Cityscapes and BDD100K datasets using encoder‑decoder and backbone‑based CNNs, we conduct a design analysis to assess how architectural choices affect routing dynamics and expert specialization. Our results demonstrate consistent, architecture‑dependent improvements (up to +3.9 mIoU) with little computational overhead, while revealing strong design sensitivity. Our work provides empirical insights into the design and internal dynamics of sparse MoE layers in CNN‑based dense prediction. Our code is available at https://github.com/KASTEL‑MobilityLab/moe‑layers/.
Authors:Zhijie Bao, Fangke Chen, Licheng Bao, Chenhui Zhang, Wei Chen, Jiajie Peng, Zhongyu Wei
Abstract:
The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation frameworks that are aligned with the real‑world medical imaging practice. Existing practices that report single or coarse‑grained metrics are lack the granularity required for specialized clinical support and fail to assess the reliability of reasoning mechanisms. To address this, we propose a paradigm shift toward multidimensional, fine‑grained and in‑depth evaluation. Based on a two‑stage systematic construction pipeline designed for this paradigm, we instantiate it with MedRCube. We benchmark 33 MLLMs, Lingshu‑32B achieve top‑tier performance. Crucially, MedRCube exposes a series of pronounced insights inaccessible under prior evaluation settings. Furthermore, we introduce a credibility evaluation subset to quantify reasoning credibility, uncover a highly significant positive association between shortcut behavior and diagnostic task performance, raising concerns for clinically trustworthy deployment. The resources of this work can be found at https://github.com/F1mc/MedRCube.
Authors:Jie Liang, Jiahao Wu, Chao Wang, Jiayu Yang, Xiaoyun Zheng, Kaiqiang Xiong, Zhanke Wang, Jinbo Yan, Feng Gao, Ronggang Wang
Abstract:
Dynamic 3D scene reconstruction is essential for immersive media such as VR, MR, and XR, yet remains challenging for long multi‑view sequences with large‑scale motion. Existing dynamic Gaussian approaches are either Frame‑Stream, offering scalability but poor temporal stability, or Clip, achieving local consistency at the cost of high memory and limited sequence length. We propose ClipGStream, a hybrid reconstruction framework that performs stream optimization at the clip level rather than the frame level. The sequence is divided into short clips, where dynamic motion is modeled using clip‑independent spatio‑temporal fields and residual anchor compensation to capture local variations efficiently, while inter‑clip inherited anchors and decoders maintain structural consistency across clips. This Clip‑Stream design enables scalable, flicker‑free reconstruction of long dynamic videos with high temporal coherence and reduced memory overhead. Extensive experiments demonstrate that ClipGStream achieves state‑of‑the‑art reconstruction quality and efficiency. The project page is available at: https://liangjie1999.github.io/ClipGStreamWeb/
Authors:Muhammad Ahmed Ullah Khan, Muhammad Haris Bin Amir, Didier Stricker, Muhammad Zeshan Afzal
Abstract:
Continual learning enables models to acquire new knowledge over time while retaining previously learned capabilities. However, its application to text‑to‑3D generation remains unexplored. We present ReConText3D, the first framework for continual text‑to‑3D generation. We first demonstrate that existing text‑to‑3D models suffer from catastrophic forgetting under incremental training. ReConText3D enables generative models to incrementally learn new 3D categories from textual descriptions while preserving the ability to synthesize previously seen assets. Our method constructs a compact and diverse replay memory through text‑embedding k‑Center selection, allowing representative rehearsal of prior knowledge without modifying the underlying architecture. To systematically evaluate continual text‑to‑3D learning, we introduce Toys4K‑CL, a benchmark derived from the Toys4K dataset that provides balanced and semantically diverse class‑incremental splits. Extensive experiments on the Toys4K‑CL benchmark show that ReConText3D consistently outperforms all baselines across different generative backbones, maintaining high‑quality generation for both old and new classes. To the best of our knowledge, this work establishes the first continual learning framework and benchmark for text‑to‑3D generation, opening a new direction for incremental 3D generative modeling. Project page is available at: https://mauk95.github.io/ReConText3D/.
Authors:Haoran Lou, Ziyan Liu, Chunxiao Fan, Yuexin Wu, Yue Ming, Hao Wu, Kai Zuo, Yibo Chen, Xu Tang
Abstract:
Multimodal Large Language Models (MLLMs) possess intrinsic reasoning and world‑knowledge capabilities, yet adapting them for dense retrieval remains challenging. Existing approaches rely on invasive parameter updates, such as full fine‑tuning and LoRA, which may disrupt the pre‑trained semantic space and impair the structured knowledge essential for reasoning. To address this, we propose SLQ, a parameter‑efficient tuning framework that adapts MLLMs for retrieval while keeping the backbone entirely frozen. SLQ introduces a small set of Shared Latent Queries that are appended to both text and image tokens, leveraging the model's native causal attention to aggregate multimodal context into a unified embedding space. Furthermore, to better evaluate retrieval beyond superficial pattern matching, we construct KARR‑Bench, a benchmark designed for knowledge‑aware reasoning retrieval. Extensive experiments show that SLQ outperforms full fine‑tuning and LoRA on COCO and Flickr30K, while achieving competitive performance on MMEB and yielding substantial gains on KARR‑Bench, validating that preserving the pre‑trained representations via non‑invasive adaptation is an effective strategy for MLLM‑based retrieval. The code is available under: https://github.com/CnFaker/SLQ.
Authors:Yulu Gao, Bohao Zhang, Zongheng Tang, Jitong Liao, Wenjun Wu, Si Liu
Abstract:
Instance‑level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for applications in embodied AI and remote collaboration. This task is exceptionally difficult due to severe changes in scale, perspective, and occlusion, which destabilize direct pixel‑level matching. While recent geometry‑aware models like VGGT provide a strong foundation for feature alignment, we find they often fail at dense prediction tasks due to significant pixel‑level projection drift, even when their internal object‑level attention remains consistent. To bridge this gap, we introduce VGGT‑Segmentor (VGGT‑S), a framework that unifies robust geometric modeling with pixel‑accurate semantic segmentation. VGGT‑S leverages VGGT's powerful cross‑view feature representation and introduces a novel Union Segmentation Head. This head operates in three stages: mask prompt fusion, point‑guided prediction, and iterative mask refinement, effectively translating high‑level feature alignment into a precise segmentation mask. Furthermore, we propose a single‑image self‑supervised training strategy that eliminates the need for paired annotations and enables strong generalization. On the Ego‑Exo4D benchmark, VGGT‑S sets a new state‑of‑the‑art, achieving 67.7% and 68.0% average IoU for Ego to Exo and Exo to Ego tasks, respectively, significantly outperforming prior methods. Notably, our correspondence‑free pretrained model surpasses most fully‑supervised baselines, demonstrating the effectiveness and scalability of our approach. Code is publicly available at: https://github.com/buaa‑colalab/VGGT‑S.
Authors:Bingxue Xu, Emil Hedemalm, Ajinkya Khoche, Patric Jensfelt
Abstract:
The challenge of 3D multi‑object tracking is achieving robustness in real‑world applications, for example under adverse conditions and maintaining consistency as distance increases. To overcome these challenges, sensor fusion approaches that combine LiDAR, cameras, and radar have emerged. However, existing multimodal methods usually treat radar as another learned feature inside the network. When the overall model degrades in difficult environments, the robustness advantages that radar could provide are also reduced. In this paper we propose RadarMOT, a radar‑informed 3D multi‑object tracking framework that explicitly uses radar point clouds as additional observations to refine state estimation and recover objects missed by the detector at long ranges. Evaluations on the MAN‑TruckScenes dataset show that RadarMOT consistently improves the Average Multi‑Object Tracking Accuracy (AMOTA) by 12.7% at long range and up to 10.3% in adverse weather. The code will be available at https://github.com/bingxue‑xu/radarmot
Authors:Yunkai Dang, Minxin Dai, Yuekun Yang, Zhangnan Li, Wenbin Li, Feng Miao, Yang Gao
Abstract:
Ultra‑high‑resolution (UHR) remote sensing imagery couples kilometer‑scale context with query‑critical evidence that may occupy only a few pixels. Such vast spatial scale leads to a quadratic explosion of visual tokens and hinders the extraction of information from small objects. Previous works utilize direct downsampling, dense tiling, or global top‑k pruning, which either compromise query‑critical image details or incur unpredictable compute. In this paper, we propose UHR‑BAT, a query‑guided and region‑faithful token compression framework to efficiently select visual tokens under a strict context budget. Specifically, we leverage text‑guided, multi‑scale importance estimation for visual tokens, effectively tackling the challenge of achieving precise yet low‑cost feature extraction. Furthermore, by introducing region‑wise preserve and merge strategies, we mitigate visual token redundancy, further driving down the computational budget. Experimental results show that UHR‑BAT achieves state‑of‑the‑art performance across various benchmarks. Code will be available at https://github.com/Yunkaidang/UHR.
Authors:Sanghyeok Chu, Pyunghwan Ahn, Gwangmo Song, SeungHwan Kim, Honglak Lee, Bohyung Han
Abstract:
Sparse Upcycling provides an efficient way to initialize a Mixture‑of‑Experts (MoE) model from pretrained dense weights instead of training from scratch. However, since all experts start from identical weights and the router is randomly initialized, the model suffers from expert symmetry and limited early specialization. We propose Cluster‑aware Upcycling, a strategy that incorporates semantic structure into MoE initialization. Our method first partitions the dense model's input activations into semantic clusters. Each expert is then initialized using the subspace representations of its corresponding cluster via truncated SVD, while setting the router's initial weights to the cluster centroids. This cluster‑aware initialization breaks expert symmetry and encourages early specialization aligned with the data distribution. Furthermore, we introduce an expert‑ensemble self‑distillation loss that stabilizes training by providing reliable routing guidance using an ensemble teacher. When evaluated on CLIP ViT‑B/32 and ViT‑B/16, Cluster‑aware Upcycling consistently outperforms existing methods across both zero‑shot and few‑shot benchmarks. The proposed method also produces more diverse and disentangled expert representations, reduces inter‑expert similarity, and leads to more confident routing behavior. Project page: https://sanghyeokchu.github.io/cluster‑aware‑upcycling/
Authors:Simin Huo, Ning Li
Abstract:
Token compression is crucial for mitigating the quadratic complexity of self‑attention mechanisms in Vision Transformers (ViTs), which often involve numerous input tokens. Existing methods, such as ToMe, rely on GPU‑inefficient operations (e.g., sorting, scattered writes), introducing overheads that limit their effectiveness. We introduce MaMe, a training‑free, differentiable token merging method based entirely on matrix operations, which is GPU‑friendly to accelerate ViTs. Additionally, we present MaRe, its inverse operation, for token restoration, forming a MaMe+MaRe pipeline for image synthesis. When applied to pre‑trained models, MaMe doubles ViT‑B throughput with a 2% accuracy drop. Notably, fine‑tuning the last layer with MaMe boosts ViT‑B accuracy by 1.0% at 1.1x speed. In SigLIP2‑B@512 zero‑shot classification, MaMe provides 1.3x acceleration with negligible performance degradation. In video tasks, MaMe accelerates VideoMAE‑L by 48.5% on Kinetics‑400 with only a 0.84% accuracy loss. Furthermore, MaMe achieves simultaneous improvements in both performance and speed on some tasks. In image synthesis, the MaMe+MaRe pipeline enhances quality while reducing Stable Diffusion v2.1 generation latency by 31%. Collectively, these results demonstrate MaMe's and MaRe's effectiveness in accelerating vision models. The code is available at https://github.com/cominder/mamehttps://github.com/cominder/mame.
Authors:Yifan Li, Pei Cheng, Bin Fu, Shuai Yang, Jiaying Liu
Abstract:
Video chroma‑lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge. Existing methods typically rely on expensive supervised training with synthetic paired data. This paper proposes VibeFlow, a novel self‑supervised framework that unleashes the intrinsic physical understanding of pre‑trained video generation models. Instead of learning color and light transitions from scratch, we introduce a disentangled data perturbation pipeline that enforces the model to adaptively recombine structure from source videos and color‑illumination cues from reference images, enabling robust disentanglement in a self‑supervised manner. Furthermore, to rectify discretization errors inherent in flow‑based models, we introduce Residual Velocity Fields alongside a Structural Distortion Consistency Regularization, ensuring rigorous structural preservation and temporal coherence. Our framework eliminates the need for costly training resources and generalizes in a zero‑shot manner to diverse applications, including video relighting, recoloring, low‑light enhancement, day‑night translation, and object‑specific color editing. Extensive experiments demonstrate that VibeFlow achieves impressive visual quality with significantly reduced computational overhead. Our project is publicly available at https://lyf1212.github.io/VibeFlow‑webpage.
Authors:Yu Wang, Sharon Li
Abstract:
In‑context learning (ICL) enables models to adapt to new tasks via inference‑time demonstrations. Despite its success in large language models, the extension of ICL to multimodal settings remains poorly understood in terms of its internal mechanisms and how it differs from text‑only ICL. In this work, we conduct a systematic analysis of ICL in multimodal large language models. Using identical task formulations across modalities, we show that multimodal ICL performs comparably to text‑only ICL in zero‑shot settings but degrades significantly under few‑shot demonstrations. To understand this gap, we decompose multimodal ICL into task mapping construction and task mapping transfer, and analyze how models establish cross‑modal task mappings, and transfer them to query samples across layers. Our analysis reveals that current models lack reasoning‑level alignment between visual and textual representations, and fail to reliably transfer learned task mappings to queries. Guided by these findings, we further propose a simple inference‑stage enhancement method that reinforces task mapping transfer. Our results provide new insights into the mechanisms and limitations of multimodal ICL and suggest directions for more effective multimodal adaptation. Our code is available \hrefhttps://github.com/deeplearning‑wisc/Multimocal‑ICL‑Analysis‑Framework‑MGIhere.
Authors:Iris Zheng, Guojun Tang, Alexander Doronin, Paul Teal, Fang-Lue Zhang
Abstract:
We present SSD‑GS, a physically‑based relighting framework built upon 3D Gaussian Splatting (3DGS) that achieves high‑quality reconstruction and photorealistic relighting under novel lighting conditions. In physically‑based relighting, accurately modeling light‑material interactions is essential for faithful appearance reproduction. However, existing 3DGS‑based relighting methods adopt coarse shading decompositions, either modeling only diffuse and specular reflections or relying on neural networks to approximate shadows and scattering. This leads to limited fidelity and poor physical interpretability, particularly for anisotropic metals and translucent materials. To address these limitations, SSD‑GS decomposes reflectance into four components: diffuse, specular, shadow, and subsurface scattering. We introduce a learnable dipole‑based scattering module for subsurface transport, an occlusion‑aware shadow formulation that integrates visibility estimates with a refinement network, and an enhanced specular component with an anisotropic Fresnel‑based model. Through progressive integration of all components during training, SSD‑GS effectively disentangles lighting and material properties, even for unseen illumination conditions, as demonstrated on the challenging OLAT dataset. Experiments demonstrate superior quantitative and perceptual relighting quality compared to prior methods and pave the way for downstream tasks, including controllable light source editing and interactive scene relighting. The source code is available at: https://github.com/irisfreesiri/SSD‑GS.
Authors:Akshit Achara, Yovin Yathathugoda, Nick Byrne, Michela Antonelli, Esther Puyol Anton, Alexander Hammers, Andrew P. King
Abstract:
The robustness of machine learning models can be compromised by spurious correlations between non‑causal features in the input data and target labels. A common way to test for such correlations is to train on data where the label is strongly tied to some non‑causal cue, then evaluate on examples where that tie no longer holds. This idea is well established for classification tasks, but for semantic segmentation the specific failure modes are not well understood. We show that a model may achieve reasonable overlap while assigning the wrong semantic label, swapping one plausible foreground class for another, even when object boundaries are largely correct. We focus on this semantic label‑flip behaviour and quantify it with a simple diagnostic (Flip) that counts how often ground truth foreground pixels are assigned the wrong foreground identity while remaining predicted as foreground. In a setting where category and scene are correlated during training, increasing the correlation consistently widens the gap between common and rare test conditions and increases these within‑object label swaps on counterfactual groups. Overall, our results motivate assessing segmentation robustness under distribution shift beyond overlap by decomposing foreground errors into correct pixels, flipped‑identity pixels, and missed‑to‑background pixels. We also propose an entropy‑based, ground truth label‑free `flip‑risk' score, which is computed from foreground identity uncertainty, and show that it can flag flip‑prone cases at inference time. Code is available at https://github.com/acharaakshit/label‑flips.
Authors:Shivam Chand Kaushik
Abstract:
Semiconductor failure analysis (FA) requires engineers to examine inspection images, correlate equipment telemetry, consult historical defect records, and write structured reports, a process that can consume several hours of expert time per case. We present SemiFA, an agentic multi‑modal framework that autonomously generates structured FA reports from semiconductor inspection images in under one minute. SemiFA decomposes FA into a four‑agent LangGraph pipeline: a DefectDescriber that classifies and narrates defect morphology using DINOv2 and LLaVA‑1.6, a RootCauseAnalyzer that fuses SECS/GEM equipment telemetry with historically similar defects retrieved from a Qdrant vector database, a SeverityClassifier that assigns severity and estimates yield impact, and a RecipeAdvisor that proposes corrective process adjustments. A fifth node assembles a PDF report. We introduce SemiFA‑930, a dataset of 930 annotated semiconductor defect images paired with structured FA narratives across nine defect classes, drawn from procedural synthesis, WM‑811K, and MixedWM38. Our DINOv2‑based classifier achieves 92.1% accuracy on 140 validation images (macro F1 = 0.917), and the full pipeline produces complete FA reports in 48 seconds on an NVIDIA A100‑SXM4‑40 GB GPU. A GPT‑4o judge ablation across four modality conditions demonstrates that multi‑modal fusion improves root cause reasoning by +0.86 composite points (1‑5 scale) over an image‑only baseline, with equipment telemetry as the more load‑bearing modality. To our knowledge, SemiFA is the first system to integrate SECS/GEM equipment telemetry into a vision‑language model pipeline for autonomous FA report generation.
Authors:Kathakoli Sengupta, Kai Ao, Paola Cascante-Bonilla
Abstract:
Large Language Models (LLMs) and Vision‑Language Models (VLMs) increasingly generate indoor scenes through intermediate structures such as layouts and scene graphs, yet evaluation still relies on LLM or VLM judges that score rendered views, making judgments sensitive to viewpoint, prompt phrasing, and hallucination. When the evaluator is unstable, it becomes difficult to determine whether a model has produced a spatially plausible scene or whether the output score reflects the choice of viewpoint, rendering, or prompt. We introduce SceneCritic, a symbolic evaluator for floor‑plan‑level layouts. SceneCritic's constraints are grounded in SceneOnto, a structured spatial ontology we construct by aggregating indoor scene priors from 3D‑FRONT, ScanNet, and Visual Genome. SceneOnto traverses this ontology to jointly verify semantic, orientation, and geometric coherence across object relationships, providing object‑level and relationship‑level assessments that identify specific violations and successful placements. Furthermore, we pair SceneCritic with an iterative refinement test bed that probes how models build and revise spatial structure under different critic modalities: a rule‑based critic using collision constraints as feedback, an LLM critic operating on the layout as text, and a VLM critic operating on rendered observations. Through extensive experiments, we show that (a) SceneCritic aligns substantially better with human judgments than VLM‑based evaluators, (b) text‑only LLMs can outperform VLMs on semantic layout quality, and (c) image‑based VLM refinement is the most effective critic modality for semantic and orientation correction.
Authors:Jian Han, Jinlai Liu, Jiahuan Wang, Bingyue Peng, Zehuan Yuan
Abstract:
While diffusion models dominate the field of visual generation, they are computationally inefficient, applying a uniform computational effort regardless of different complexity. In contrast, autoregressive (AR) models are inherently complexity‑aware, as evidenced by their variable likelihoods, but are often hindered by lossy discrete tokenization and error accumulation. In this work, we introduce Generative Refinement Networks (GRN), a next‑generation visual synthesis paradigm to address these issues. At its core, GRN addresses the discrete tokenization bottleneck through a theoretically near‑lossless Hierarchical Binary Quantization (HBQ), achieving a reconstruction quality comparable to continuous counterparts. Built upon HBQ's latent space, GRN fundamentally upgrades AR generation with a global refinement mechanism that progressively perfects and corrects artworks ‑‑ like a human artist painting. Besides, GRN integrates an entropy‑guided sampling strategy, enabling complexity‑aware, adaptive‑step generation without compromising visual quality. On the ImageNet benchmark, GRN establishes new records in image reconstruction (0.56 rFID) and class‑conditional image generation (1.81 gFID). We also scale GRN to more challenging text‑to‑image and text‑to‑video generation, delivering superior performance on an equivalent scale. We release all models and code to foster further research on GRN.
Authors:Himangi Mittal, Gaurav Mittal, Nelson Daniel Troncoso, Yu Hu
Abstract:
Computer Use Agents (CUAs) fundamentally rely on graphical user interface (GUI) grounding to translate language instructions into executable screen actions, but editing‑level grounding in dense coding interfaces (such as VS Code and Cursor), where sub‑pixel accuracy is required to interact with dense IDE elements, remains underexplored. Existing approaches typically rely on single‑shot coordinate prediction, which lacks a mechanism for error correction and often fails in high‑density interfaces. In this technical report, we conduct an empirical study of pixel‑precise cursor localization in coding environments. Instead of a single‑step execution, our agent engages in an iterative refinement process, utilizing visual feedback from previous attempts to reach the target element. This closed‑loop grounding mechanism allows the agent to self‑correct displacement errors and adapt to dynamic UI changes. We evaluate our approach across Claude, Qwen, and GPT on a suite of complex coding benchmarks, demonstrating that multi‑turn refinement significantly outperforms state‑of‑the‑art single‑shot models in both click precision and overall task success rate. Our results suggest that iterative visual reasoning is a critical component for the next generation of reliable software engineering agents. Code: https://github.com/microsoft/precision‑cua‑bench/tree/main.
Authors:Amir Hossein Kargaran, Nafiseh Nikeghbal, Jana Diesner, François Yvon, Hinrich Schütze
Abstract:
Optical character recognition (OCR) has advanced rapidly with the rise of vision‑language models, yet evaluation has remained concentrated on a small cluster of high‑ and mid‑resource scripts. We introduce GlotOCR Bench, a comprehensive benchmark evaluating OCR generalization across 100+ Unicode scripts. Our benchmark comprises clean and degraded image variants rendered from real multilingual texts. Images are rendered using fonts from the Google Fonts repository, shaped with HarfBuzz and rasterized with FreeType, supporting both LTR and RTL scripts. Samples of rendered images were manually reviewed to verify correct rendering across all scripts. We evaluate a broad suite of open‑weight and proprietary vision‑language models and find that most perform well on fewer than ten scripts, and even the strongest frontier models fail to generalize beyond thirty scripts. Performance broadly tracks script‑level pretraining coverage, suggesting that current OCR systems rely on language model pretraining as much as on visual recognition. Models confronted with unfamiliar scripts either produce random noise or hallucinate characters from similar scripts they already know. We release the benchmark and pipeline for reproducibility. Pipeline Code: https://github.com/cisnlp/glotocr‑bench, Benchmark: https://hf.co/datasets/cis‑lmu/glotocr‑bench.
Authors:Tong Zhang, Jiangning Zhang, Zhucun Xue, Juntao Jiang, Yicheng Xu, Chengming Xu, Teng Hu, Xingyu Xie, Xiaobin Hu, Yabiao Wang, Yong Liu, Shuicheng Yan
Abstract:
Balancing convergence speed, generalization capability, and computational efficiency remains a core challenge in deep learning optimization. First‑order gradient descent methods, epitomized by stochastic gradient descent (SGD) and Adam, serve as the cornerstone of modern training pipelines. However, large‑scale model training, stringent differential privacy requirements, and distributed learning paradigms expose critical limitations in these conventional approaches regarding privacy protection and memory efficiency. To mitigate these bottlenecks, researchers explore second‑order optimization techniques to surpass first‑order performance ceilings, while zeroth‑order methods reemerge to alleviate memory constraints inherent to large‑scale training. Despite this proliferation of methodologies, the field lacks a cohesive framework that unifies underlying principles and delineates application scenarios for these disparate approaches. In this work, we retrospectively analyze the evolutionary trajectory of deep learning optimization algorithms and present a comprehensive empirical evaluation of mainstream optimizers across diverse model architectures and training scenarios. We distill key emerging trends and fundamental design trade‑offs, pinpointing promising directions for future research. By synthesizing theoretical insights with extensive empirical evidence, we provide actionable guidance for designing next‑generation highly efficient, robust, and trustworthy optimization methods. The code is available at https://github.com/APRIL‑AIGC/Awesome‑Optimizer.
Authors:Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris
Abstract:
Multimodal large language models (MLLMs) perform well on many vision‑language tasks but often struggle with vision‑centric problems that require fine‑grained visual reasoning. Recent evidence suggests that this limitation arises not from weak visual representations, but from under‑utilization of visual information during instruction tuning, where many tasks can be partially solved using language priors alone. We propose a simple and lightweight approach that augments visual instruction tuning with a small number of visually grounded self‑supervised tasks expressed as natural language instructions. By reformulating classical self‑supervised pretext tasks, such as rotation prediction, color matching, and cross‑view correspondence, as image‑instruction‑response triplets, we introduce supervision that cannot be solved without relying on visual evidence. Our approach requires no human annotations, no architectural modifications, and no additional training stages. Across multiple models, training regimes, and benchmarks, injecting only a small fraction (3‑10%) of such visually grounded instructions consistently improves performance on vision‑centric evaluations. Our findings highlight instruction tuning with visually grounded SSL tasks as a powerful lever for improving visual reasoning in MLLMs through simple adjustments to the training data distribution. Code available at: https://github.com/sirkosophia/V‑GIFT
Authors:Yiyang Huang, Yitian Zhang, Yizhou Wang, Mingyuan Zhang, Liang Shi, Huimin Zeng, Yun Fu
Abstract:
Despite significant progress in video‑language modeling, hallucinations remain a persistent challenge in Video Large Language Models (Vid‑LLMs), referring to outputs that appear plausible yet contradict the content of the input video. This survey presents a comprehensive analysis of hallucinations in Vid‑LLMs and introduces a systematic taxonomy that categorizes them into two core types: dynamic distortion and content fabrication, each comprising two subtypes with representative cases. Building on this taxonomy, we review recent advances in the evaluation and mitigation of hallucinations, covering key benchmarks, metrics, and intervention strategies. We further analyze the root causes of dynamic distortion and content fabrication, which often result from limited capacity for temporal representation and insufficient visual grounding. These insights inform several promising directions for future work, including the development of motion‑aware visual encoders and the integration of counterfactual learning techniques. This survey consolidates scattered progress to foster a systematic understanding of hallucinations in Vid‑LLMs, laying the groundwork for building robust and reliable video‑language systems. An up‑to‑date curated list of related works is maintained at https://github.com/hukcc/Awesome‑Video‑Hallucination .
Authors:Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, Christian Theobalt
Abstract:
We present Grasp in Gaussians (GraG), a fast and robust method for reconstructing dynamic 3D hand‑object interactions from a single monocular video. Unlike recent approaches that optimize heavy neural representations, our method focuses on tracking the hand and the object efficiently, once initialized from pretrained large models. Our key insight is that accurate and temporally stable hand‑object motion can be recovered using a compact Sum‑of‑Gaussians (SoG) representation, revived from classical tracking literature and integrated with generative Gaussian‑based initializations. We initialize object pose and geometry using a video‑adapted SAM3D pipeline, then convert the resulting dense Gaussian representation into a lightweight SoG via subsampling. This compact representation enables efficient and fast tracking while preserving geometric fidelity. For the hand, we adopt a complementary strategy: starting from off‑the‑shelf monocular hand pose initialization, we refine hand motion using simple yet effective 2D joint and depth alignment losses, avoiding per‑frame refinement of a detailed 3D hand appearance model while maintaining stable articulation. Extensive experiments on public benchmarks demonstrate that GraG reconstructs temporally coherent hand‑object interactions on long sequences 6.4x faster than prior work while improving object reconstruction by 13.4% and reducing hand's per‑joint position error by over 65%.
Authors:Yifan Du, Zikang Liu, Jinbiao Peng, Jie Wu, Junyi Li, Jinyang Li, Wayne Xin Zhao, Ji-Rong Wen
Abstract:
Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, managing the heterogeneous information and high token costs associated with multimodal inputs over long horizons remains a critical challenge, as existing methods often suffer from context explosion or the loss of crucial visual signals. To address this, we propose a novel Long‑horizon MultiModal deep search framework, named LMM‑Searcher, centered on a file‑based visual representation mechanism. By offloading visual assets to an external file system and mapping them to lightweight textual identifiers (UIDs), our approach mitigates context overhead while preserving multimodal information for future access. We equip the agent with a tailored fetch‑image tool, enabling a progressive, on‑demand visual loading strategy for active perception. Furthermore, we introduce a data synthesis pipeline designed to generate queries requiring complex cross‑modal multi‑hop reasoning. Using this pipeline, we distill 12K high‑quality trajectories to fine‑tune Qwen3‑VL‑Thinking‑30A3B into a specialized multimodal deep search agent. Extensive experiments across four benchmarks demonstrate that our method successfully scales to 100‑turn search horizons, achieving state‑of‑the‑art performance among open‑source models on challenging long‑horizon benchmarks like MM‑BrowseComp and MMSearch‑Plus, while also exhibiting strong generalizability across different base models. Our code will be released in https://github.com/RUCAIBox/LMM‑Searcher.
Authors:Feiyu Tan, Heran Yang, Qihong Duan, Kai Ye, Qi Xie, Deyu Meng
Abstract:
Image‑to‑image translation (I2I) is a fundamental task in computer vision, focused on mapping an input image from a source domain to a corresponding image in a target domain while preserving domain‑invariant features and adapting domain‑specific attributes. Despite the remarkable success of deep learning‑based I2I approaches, the lack of paired data and unsupervised learning framework still hinder their effectiveness. In this work, we address the challenge by incorporating transformation symmetry priors into image‑to‑image translation networks. Specifically, we introduce rotation group equivariant convolutions to achieve rotation equivariant I2I framework, a novel contribution, to the best of our knowledge, along this research direction. This design ensures the preservation of rotation symmetry, one of the most intrinsic and domain‑invariant properties of natural and scientific images, throughout the network. Furthermore, we conduct a systematic study on image symmetry priors on real dataset and propose a novel transformation learnable equivariant convolutions (TL‑Conv) that adaptively learns transformation groups, enhancing symmetry preservation across diverse datasets. We also provide a theoretical analysis of the equivariance error of TL‑Conv, proving that it maintains exact equivariance in continuous domains and provide a bound for the error in discrete cases. Through extensive experiments across a range of I2I tasks, we validate the effectiveness and superior performance of our approach, highlighting the potential of equivariant networks in enhancing generation quality and its broad applicability. Our code is available at https://github.com/tanfy929/Equivariant‑I2I
Authors:Yunkai Dang, Yizhu Jiang, Yifan Jiang, Qi Fan, Yinghuan Shi, Wenbin Li, Yang Gao
Abstract:
Multimodal Large Language Models (MLLMs) suffer from substantial computational overhead due to the high redundancy in visual token sequences. Existing approaches typically address this issue using single‑layer Vision Transformer (ViT) features and static pruning strategies. However, such fixed configurations are often brittle under diverse instructions. To overcome these limitations, we propose CLASP, a plug‑and‑play token reduction framework based on class‑adaptive layer fusion and dual‑stage pruning. Specifically, CLASP first constructs category‑specific visual representations through multi‑layer vision feature fusion. It then performs dual‑stage pruning, allocating the token budget between attention‑salient pivot tokens for relevance and redundancy‑aware completion tokens for coverage. Through class‑adaptive pruning, CLASP enables prompt‑conditioned feature fusion and budget allocation, allowing aggressive yet robust visual token reduction. Extensive experiments demonstrate that CLASP consistently outperforms existing methods across a wide range of benchmarks, pruning ratios, and MLLM architectures. Code will be available at https://github.com/Yunkaidang/CLASP.
Authors:T. Camaret Ndir, Marco Reisert, Robin T. Schirrmeister
Abstract:
In‑context learning (ICL) enables medical image segmentation models to adapt to new anatomical structures from limited examples, reducing the clinical annotation burden. However, standard ICL methods typically rely on dense, global cross‑attention, which scales poorly with image resolution. While recent approaches have introduced localized attention mechanisms, they often lack explicit supervision on the selection process, leading to redundant computation in non‑informative regions. We propose PatchICL, a hierarchical framework that combines selective image patching with multi‑level supervision. Our approach learns to actively identify and attend only to the most informative anatomical regions. Compared to UniverSeg, a strong global‑attention baseline, PatchICL achieves competitive in‑domain CT segmentation accuracy while reducing compute by 44% at 512×512 resolution. On 35 out‑of‑domain datasets spanning diverse imaging modalities, PatchICL outperforms the baseline on 6 of 13 modality categories, with particular strength on modalities dominated by localized pathology such as OCT and dermoscopy. Training and evaluation code are available at https://github.com/tidiane‑camaret/ic_segmentation
Authors:Zeheng Wang, Zitong Yu, Yijie Zhu, Bo Zhao, Haochen Liang, Taorui Wang, Wei Xia, Jiayu Zhang, Zhishu Liu, Hui Ma, Fei Ma, Qi Tian
Abstract:
LLM‑based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this paper, given that single‑round retrieval‑augmented generation is highly susceptible to modal ambiguity and therefore struggles to capture complex affective dependencies across modalities, we introduce AffectAgent, an affect‑oriented multi‑agent retrieval‑augmented generation framework that leverages collaborative decision‑making among agents for fine‑grained affective understanding. Specifically, AffectAgent comprises three jointly optimized specialized agents, namely a query planner, an evidence filter, and an emotion generator, which collaboratively perform analytical reasoning to retrieve cross‑modal samples, assess evidence, and generate predictions. These agents are optimized end‑to‑end using Multi‑Agent Proximal Policy Optimization (MAPPO) with a shared affective reward to ensure consistent emotion understanding. Furthermore, we introduce Modality‑Balancing Mixture of Experts (MB‑MoE) and Retrieval‑Augmented Adaptive Fusion (RAAF), where MB‑MoE dynamically regulates the contributions of different modalities to mitigate representation mismatch caused by cross‑modal heterogeneity, while RAAF enhances semantic completion under missing‑modality conditions by incorporating retrieved audiovisual embeddings. Extensive experiments on MER‑UniBench demonstrate that AffectAgent achieves superior performance across complex scenarios. Our code will be released at: https://github.com/Wz1h1NG/AffectAgent.
Authors:Junfeng Xia, Wenhao Ye, Xuanye Pan, Xinke Shen, Mo Wang, Quanying Liu
Abstract:
Current fMRI foundation models primarily rely on a limited range of brain states and mismatched pretraining tasks, restricting their ability to learn generalized representations across diverse brain states. We present Brain‑DiT, a universal multi‑state fMRI foundation model pretrained on 349,898 sessions from 24 datasets spanning resting, task, naturalistic, disease, and sleep states. Unlike prior fMRI foundation models that rely on masked reconstruction in the raw‑signal space or a latent space, Brain‑DiT adopts metadata‑conditioned diffusion pretraining with a Diffusion Transformer (DiT), enabling the model to learn multi‑scale representations that capture both fine‑grained functional structure and global semantics. Across extensive evaluations and ablations on 7 downstream tasks, we find consistent evidence that diffusion‑based generative pretraining is a stronger proxy than reconstruction or alignment, with metadata‑conditioned pretraining further improving downstream performance by disentangling intrinsic neural dynamics from population‑level variability. We also observe that downstream tasks exhibit distinct preferences for representational scale: ADNI classification benefits more from global semantic representations, whereas age/sex prediction comparatively relies more on fine‑grained local structure. Code and parameters of Brain‑DiT are available at \hrefhttps://github.com/REDMAO4869/Brain‑DiTLink.
Authors:Ziyuan Xia, Jingyi Xu, Chong Cui, Yuanhong Yu, Jiazhao Zhang, Qingsong Yan, Tao Ni, Junbo Chen, Xiaowei Zhou, Hujun Bao, Ruizhen Hu, Sida Peng
Abstract:
Training embodied AI agents depends critically on the visual fidelity of simulation environments and the ability to model dynamic humans. Current simulators rely on mesh‑based rasterization with limited visual realism, and their support for dynamic human avatars, where available, is constrained to mesh representations, hindering agent generalization to human‑populated real‑world scenarios. We present Habitat‑GS, a navigation‑centric embodied AI simulator extended from Habitat‑Sim that integrates 3D Gaussian Splatting scene rendering and drivable gaussian avatars while maintaining full compatibility with the Habitat ecosystem. Our system implements a 3DGS renderer for real‑time photorealistic rendering and supports scalable 3DGS asset import from diverse sources. For dynamic human modeling, we introduce a gaussian avatar module that enables each avatar to simultaneously serve as a photorealistic visual entity and an effective navigation obstacle, allowing agents to learn human‑aware behaviors in realistic settings. Experiments on point‑goal navigation demonstrate that agents trained on 3DGS scenes achieve stronger cross‑domain generalization, with mixed‑domain training being the most effective strategy. Evaluations on avatar‑aware navigation further confirm that gaussian avatars enable effective human‑aware navigation. Finally, performance benchmarks validate the system's scalability across varying scene complexity and avatar counts.
Authors:Yuhao Liu, Dingju Wang, Ziyang Zheng
Abstract:
This paper presents our approach to the NTIRE 2026 3D Restoration and Reconstruction Challenge (Track 1), which focuses on reconstructing high‑quality 3D representations from degraded multi‑view inputs. The challenge involves recovering geometrically consistent and photorealistic 3D scenes in extreme low‑light environments. To address this task, we propose Extreme Low‑light Optimized Gaussian Splatting (ELoG‑GS), a robust low‑light 3D reconstruction pipeline that integrates learning‑based point cloud initialization and luminance‑guided color enhancement for stable and photorealistic Gaussian Splatting. Our method incorporates both geometry‑aware initialization and photometric adaptation strategies to improve reconstruction fidelity under challenging conditions. Extensive experiments on the NTIRE Track 1 benchmark demonstrate that our approach significantly improves reconstruction quality over the baselines, achieving superior visual fidelity and geometric consistency. The proposed method provides a practical solution for robust 3D reconstruction in real‑world degraded scenarios. In the final testing phase, our method achieved a PSNR of 18.6626 and an SSIM of 0.6855 on the official platform leaderboard. Code is available at https://github.com/lyh120/FSGS_EAPGS.
Authors:Kangmin Seo, MinKyu Lee, Tae-Young Kim, ByeongCheol Lee, JoonSeoung An, Jae-Pil Heo
Abstract:
Recent advances in 3D Gaussian Splatting (3DGS) have enabled impressive real‑time photorealistic rendering. However, conventional training pipelines inherently assume full multi‑view consistency among input images, which makes them sensitive to distractors that violate this assumption and cause visual artifacts. In this work, we revisit an underexplored aspect of 3DGS: its inherent ability to suppress inconsistent signals. Building on this insight, we propose PDF‑GS (Progressive Distractor Filtering for Robust 3D Gaussian Splatting), a framework that amplifies this self‑filtering property through a progressive multi‑phase optimization. The progressive filtering phases gradually remove distractors by exploiting discrepancy cues, while the following reconstruction phase restores fine‑grained, view‑consistent details from the purified Gaussian representation. Through this iterative refinement, PDF‑GS achieves robust, high‑fidelity, and distractor‑free reconstructions, consistently outperforming baselines across diverse datasets and challenging real‑world conditions. Moreover, our approach is lightweight and easily adaptable to existing 3DGS frameworks, requiring no architectural changes or additional inference overhead, leading to a new state‑of‑the‑art performance. The code is publicly available at https://github.com/kangrnin/PDF‑GS.
Authors:Yinxi He, Kang Liao, Chunyu Lin, Tianyi Wei, Yao Zhao
Abstract:
This paper introduces StructDiff, a generative framework based on a single‑scale diffusion model for single‑image generation. Single‑image generation aims to synthesize diverse samples with similar visual content to the source image by capturing its internal statistics, without relying on external data. However, existing methods often struggle to preserve the structural layout, especially for images with large rigid objects or strict spatial constraints. Moreover, most approaches lack spatial controllability, making it difficult to guide the structure or placement of generated content. To address these challenges, StructDiff introduces an adaptive receptive field module to maintain both global and local distributions. Building on this foundation, StructDiff incorporates 3D positional encoding (PE) as a spatial prior, allowing flexible control over positions, scale, and local details of generated objects. To our knowledge, this spatial control capability represents the first exploration of PE‑based manipulation in single‑image generation. Furthermore, we propose a novel evaluation criterion for single‑image generation based on large language models (LLMs). This criterion specifically addresses the limitations of existing objective metrics and the high labor costs associated with user studies. StructDiff also demonstrates broad applicability across downstream tasks, such as text‑guided image generation, image editing, outpainting, and paint‑to‑image synthesis. Extensive experiments demonstrate that StructDiff outperforms existing methods in structural consistency, visual quality, and spatial controllability. The project page is available at https://butter‑crab.github.io/StructDiff/.
Authors:Francesco Chiumento, Julia Dietlmeier, Ronan P. Killeen, Kathleen M. Curran, Noel E. O'Connor, Mingming Liu
Abstract:
Detecting amyloid‑β (Aβ) positivity is crucial for early diagnosis of Alzheimer's disease but typically requires PET imaging, which is costly, invasive, and not widely accessible, limiting its use for population‑level screening. We address this gap by proposing a PET‑guided knowledge distillation framework that enables Aβ prediction from MRI alone, without requiring non‑imaging clinical covariates or PET at inference. Our approach employs a BiomedCLIP‑based teacher model that learns PET‑MRI alignment via cross‑modal attention and triplet contrastive learning with PET‑informed (Centiloid‑aware) online negative sampling. An MRI‑only student then mimics the teacher via feature‑level and logit‑level distillation. Evaluated across four MRI contrasts (T1w, T2w, FLAIR, T2) and two independent datasets, our approach demonstrates effective knowledge transfer (best AUC: 0.74 on OASIS‑3, 0.68 on ADNI) while maintaining interpretability and eliminating the need for clinical variables. Saliency analysis confirms that predictions focus on anatomically relevant cortical regions, supporting the clinical viability of PET‑free Aβ screening. Code is available at https://github.com/FrancescoChiumento/pet‑guided‑mri‑amyloid‑detection.
Authors:Zhaoyang Jia, Naifu Xue, Zihan Zheng, Jiahao Li, Bin Li, Xiaoyi Zhang, Zongyu Guo, Yuan Zhang, Houqiang Li, Yan Lu
Abstract:
Recent advanced diffusion methods typically derive strong generative priors by scaling diffusion transformers. However, scaling fails to generalize when adapted for real‑time compression scenarios that demand lightweight models. In this paper, we explore the design of real‑time and lightweight diffusion codecs by addressing two pivotal questions. First, does diffusion pre‑training benefit lightweight diffusion codecs? Through systematic analysis, we find that generation‑oriented pre‑training is less effective at small model scales whereas compression‑oriented pre‑training yields consistently better performance. Second, are transformers essential? We find that while global attention is crucial for standard generation, lightweight convolutions suffice for compression‑oriented diffusion when paired with distillation. Guided by these findings, we establish a one‑step lightweight convolution diffusion codec that achieves real‑time 60~FPS encoding and 42~FPS decoding at 1080p. Further enhanced by distillation and adversarial learning, the proposed codec reduces bitrate by 85% at a comparable FID to MS‑ILLM, bridging the gap between generative compression and practical real‑time deployment. Codes are released at https://github.com/microsoft/GenCodec/tree/main/CoD_Lite
Authors:Guanyi Qin, Jie Liang, Bingbing Zhang, Lishen Qu, Ya-nan Guan, Hui Zeng, Lei Zhang, Radu Timofte, Jianhui Sun, Xinli Yue, Tao Shao, Huan Hou, Wenjie Liao, Shuhao Han, Jieyu Yuan, Chunle Guo, Chongyi Li, Zewen Chen, Yunze Liu, Jian Guo, Juan Wang, Yun Zeng, Bing Li, Weiming Hu, Hesong Li, Dehua Liu, Xinjie Zhang, Qiang Li, Li Yan, Wei Dong, Qingsen Yan, Xingcan Li, Shenglong Zhou, Manjiang Yin, Yinxiang Zhang, Hongbo Wang, Jikai Xu, Zhaohui Fan, Dandan Zhu, Wei Sun, Weixia Zhang, Kun Zhu, Nana Zhang, Kaiwei Zhang, Qianqian Zhang, Zhihan Zhang, William Gordon, Linwei Wu, Jiachen Tu, Guoyi Xu, Yaoxin Jiang, Cici Liu, Yaokun Shi
Abstract:
In this paper, we present an overview of the NTIRE 2026 challenge on the 3rd Restore Any Image Model in the Wild, specifically focusing on Track 1: Professional Image Quality Assessment. Conventional Image Quality Assessment (IQA) typically relies on scalar scores. By compressing complex visual characteristics into a single number, these methods fundamentally struggle to distinguish subtle differences among uniformly high‑quality images. Furthermore, they fail to articulate why one image is superior, lacking the reasoning capabilities required to provide guidance for vision tasks. To bridge this gap, recent advancements in Multimodal Large Language Models (MLLMs) offer a promising paradigm. Inspired by this potential, our challenge establishes a novel benchmark exploring the ability of MLLMs to mimic human expert cognition in evaluating high‑quality image pairs. Participants were tasked with overcoming critical bottlenecks in professional scenarios, centering on two primary objectives: (1) Comparative Quality Selection: reliably identifying the visually superior image within a high‑quality pair; and (2) Interpretative Reasoning: generating grounded, expert‑level explanations that detail the rationale behind the selection. In total, the challenge attracted nearly 200 registrations and over 2,500 submissions. The top‑performing methods significantly advanced the state of the art in professional IQA. The challenge dataset is available at https://github.com/narthchin/RAIM‑PIQA, and the official homepage is accessible at https://www.codabench.org/competitions/12789/.
Authors:Junbin Su, Ziteng Xue, Shihui Zhang, Kun Chen, Weiming Hu, Zhipeng Zhang
Abstract:
Parameter‑efficient fine‑tuning (PEFT) in multimodal tracking reveals a concerning trend where recent performance gains are often achieved at the cost of inflated parameter budgets, which fundamentally erodes PEFT's efficiency promise. In this work, we introduce SEATrack, a Simple, Efficient, and Adaptive two‑stream multimodal tracker that tackles this performance‑efficiency dilemma from two complementary perspectives. We first prioritize cross‑modal alignment of matching responses, an underexplored yet pivotal factor that we argue is essential for breaking the trade‑off. Specifically, we observe that modality‑specific biases in existing two‑stream methods generate conflicting matching attention maps, thereby hindering effective joint representation learning. To mitigate this, we propose AMG‑LoRA, which seamlessly integrates Low‑Rank Adaptation (LoRA) for domain adaptation with Adaptive Mutual Guidance (AMG) to dynamically refine and align attention maps across modalities. We then depart from conventional local fusion approaches by introducing a Hierarchical Mixture of Experts (HMoE) that enables efficient global relation modeling, effectively balancing expressiveness and computational efficiency in cross‑modal fusion. Equipped with these innovations, SEATrack advances notable progress over state‑of‑the‑art methods in balancing performance with efficiency across RGB‑T, RGB‑D, and RGB‑E tracking tasks. \hrefhttps://github.com/AutoLab‑SAI‑SJTU/SEATrack\textcolorcyanCode is available.
Authors:Nihal Jaiswal, Siddhartha Arjaria, Gyanendra Chaubey, Ankush Kumar, Aditya Singh, Anchal Chaurasiya
Abstract:
Text‑to‑image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in training data. We introduce T2I‑BiasBench, a unified evaluation framework of thirteen complementary metrics that jointly captures demographic bias, element omission, and cultural collapse in diffusion models ‑ the first framework to address all three dimensions simultaneously.
We evaluate three open‑source models ‑ Stable Diffusion v1.5, BK‑SDM Base, and Koala Lightning ‑ against Gemini 2.5 Flash (RLHF‑aligned) as a reference baseline. The benchmark comprises 1,574 generated images across five structured prompt categories. T2I‑BiasBench integrates six established metrics with seven additional measures: four newly proposed (Composite Bias Score, Grounded Missing Rate, Implicit Element Missing Rate, Cultural Accuracy Ratio) and three adapted (Hallucination Score, Vendi Score, CLIP Proxy Score).
Three key findings emerge: (1) Stable Diffusion v1.5 and BK‑SDM exhibit bias amplification (>1.0) in beauty‑related prompts; (2) contextual constraints such as surgical PPE substantially attenuate professional‑role gender bias (Doctor CBS = 0.06 for SD v1.5); and (3) all models, including RLHF‑aligned Gemini, collapse to a narrow set of cultural representations (CAS: 0.54‑1.00), confirming that alignment techniques do not resolve cultural coverage gaps.
T2I‑BiasBench is publicly released to support standardized, fine‑grained bias evaluation of generative models. The project page is available at: https://gyanendrachaubey.github.io/T2I‑BiasBench/
Authors:Paschalis Giakoumoglou, Symeon Papadopoulos
Abstract:
Modern diffusion‑based inpainting models pose significant challenges for image forgery localization (IFL), as their full regeneration pipelines reconstruct the entire image via a latent decoder, disrupting the camera‑level noise patterns that existing forensic methods rely on. We propose DiffusionPrint, a patch‑level contrastive learning framework that learns a forensic signal robust to the spectral distortions introduced by latent decoding. It exploits the fact that inpainted regions generated by the same model share a consistent generative fingerprint, using this as a self‑supervisory signal. DiffusionPrint trains a convolutional backbone via a MoCo‑style objective with cross‑category hard negative mining and a generator‑aware classification head, producing a forensic feature map that serves as a highly discriminative secondary modality in fusion‑based IFL frameworks. Integrated into TruFor, MMFusion, and a lightweight fusion baseline, DiffusionPrint consistently improves localization across multiple generative models, with gains of up to +28% on mask types unseen during fine‑tuning and confirmed generalization to unseen generative architectures. Code is available at https://github.com/mever‑team/diffusionprint
Authors:Jiawei Fan, Shigeng Wang, Chao Li, Xiaolong Liu, Anbang Yao
Abstract:
In this paper, we present Chain‑of‑Models Pre‑Training (CoM‑PT), a novel performance‑lossless training acceleration method for vision foundation models (VFMs). This approach fundamentally differs from existing acceleration methods in its core motivation: rather than optimizing each model individually, CoM‑PT is designed to accelerate the training pipeline at the model family level, scaling efficiently as the model family expands. Specifically, CoM‑PT establishes a pre‑training sequence for the model family, arranged in ascending order of model size, called model chain. In this chain, only the smallest model undergoes standard individual pre‑training, while the other models are efficiently trained through sequential inverse knowledge transfer from their smaller predecessors by jointly reusing the knowledge in the parameter space and the feature space. As a result, CoM‑PT enables all models to achieve performance that is mostly superior to standard individual training while significantly reducing training cost, and this is extensively validated across 45 datasets spanning zero‑shot and fine‑tuning tasks. Notably, its efficient scaling property yields a remarkable phenomenon: training more models even results in higher efficiency. For instance, when pre‑training on CC3M: i) given ViT‑L as the largest model, progressively prepending smaller models to the model chain reduces computational complexity by up to 72%; ii) within a fixed model size range, as the VFM family scales across 3, 4, and 7 models, the acceleration ratio of CoM‑PT exhibits a striking leap: from 4.13X to 5.68X and 7.09X. Since CoM‑PT is naturally agnostic to specific pre‑training paradigms, we open‑source the code to spur further extensions in more computationally intensive scenarios, such as large language model pre‑training.
Authors:Jiwan Kim, Kibum Kim, Wonjoong Kim, Byung-Kwan Lee, Chanyoung Park
Abstract:
Recently, visual token pruning has been studied to handle the vast number of visual tokens in Multimodal Large Language Models. However, we observe that while existing pruning methods perform reliably on simple visual understanding, they struggle to effectively generalize to complex visual reasoning tasks, a critical gap underexplored in previous studies. Through a systematic analysis, we identify Relevant Visual Information Shift (RVIS) during decoding as the primary failure driver. To address this, we propose Decoding‑stage Shift‑aware Token Pruning (DSTP), a training‑free add‑on framework that enables existing pruning methods to align visual tokens with shifting reasoning requirements during the decoding stage. Extensive experiments demonstrate that DSTP significantly mitigates performance degradation of pruning methods in complex reasoning tasks, while consistently yielding performance gains even across visual understanding benchmarks. Furthermore, DSTP demonstrates effectiveness across diverse state‑of‑the‑art architectures, highlighting its generalizability and efficiency with minimal computational overhead.
Authors:Dongjian Yu, Weiqing Min, Qian Jiang, Xing Lin, Xin Jin, Shuqiang Jiang
Abstract:
Accurate estimation of food nutrition plays a vital role in promoting healthy dietary habits and personalized diet management. Most existing food datasets primarily focus on Western cuisines and lack sufficient coverage of Chinese dishes, which restricts accurate nutritional estimation for Chinese meals. Moreover, many state‑of‑the‑art nutrition prediction methods rely on depth sensors, restricting their applicability in daily scenarios. To address these limitations, we introduce OmniFood8K, a comprehensive multimodal dataset comprising 8,036 food samples, each with detailed nutritional annotations and multi‑view images. In addition, to enhance models' capability in nutritional prediction, we construct NutritionSynth‑115K, a large‑scale synthetic dataset that introduces compositional variations while preserving precise nutritional labels. Moreover, we propose an end‑to‑end framework for nutritional prediction from a single RGB image. First, we predict a depth map from a single RGB image and design the Scale‑Shift Residual Adapter (SSRA) to refine it for global scale consistency and local structural preservation. Second, we propose the Frequency‑Aligned Fusion Module (FAFM) to hierarchically align and fuse RGB and depth features in the frequency domain. Finally, we design a Mask‑based Prediction Head (MPH) to emphasize key ingredient regions via dynamic channel selection for more accurate prediction. Extensive experiments on multiple datasets demonstrate the superiority of our method over existing approaches. Project homepage: https://yudongjian.github.io/OmniFood8K‑food/
Authors:Deyuan Liu, Peng Sun, Yansen Han, Zhenglin Cheng, Chuyan Chen, Tao Lin
Abstract:
The push for efficient text to image synthesis has moved the field toward one step sampling, yet existing methods still face a three way tradeoff among fidelity, inference speed, and training efficiency. Approaches that rely on external discriminators can sharpen one step performance, but they often introduce training instability, high GPU memory overhead, and slow convergence, which complicates scaling and parameter efficient tuning. In contrast, regression based distillation and consistency objectives are easier to optimize, but they typically lose fine details when constrained to a single step. We present APEX, built on a key theoretical insight: adversarial correction signals can be extracted endogenously from a flow model through condition shifting. Using a transformation creates a shifted condition branch whose velocity field serves as an independent estimator of the model's current generation distribution, yielding a gradient that is provably GAN aligned, replacing the sample dependent discriminator terms that cause gradient vanishing. This discriminator free design is architecture preserving, making APEX a plug and play framework compatible with both full parameter and LoRA based tuning. Empirically, our 0.6B model surpasses FLUX‑Schnell 12B (20× more parameters) in one step quality. With LoRA tuning on Qwen‑Image 20B, APEX reaches a GenEval score of 0.89 at NFE=1 in 6 hours, surpassing the original 50‑step teacher (0.87) and providing a 15.33× inference speedup. Code is available https://github.com/LINs‑lab/APEX.
Authors:Clara Xue, Zizheng Yan, Zhenning Shi, Yuhang Yu, Jingyu Zhuang, Qi Zhang, Jinwei Chen, Qingnan Fan
Abstract:
Live Photo captures both a high‑quality key photo and a short video clip to preserve the precious dynamics around the captured moment. While users may choose alternative frames as the key photo to capture better expressions or timing, these frames often exhibit noticeable quality degradation, as the photo capture ISP pipeline delivers significantly higher image quality than the video pipeline. This quality gap highlights the need for dedicated restoration techniques to enhance the reselected key photo. To this end, we propose LiveMoments, a reference‑guided image restoration framework tailored for the reselected key photo in Live Photos. Our method employs a two‑branch neural network: a reference branch that extracts structural and textural information from the original high‑quality key photo, and a main branch that restores the reselected frame using the guidance provided by the reference branch. Furthermore, we introduce a unified Motion Alignment module that incorporates motion guidance for spatial alignment at both the latent and image levels. Experiments on real and synthetic Live Photos demonstrate that LiveMoments significantly improves perceptual quality and fidelity over existing solutions, especially in scenes with fast motion or complex structures. Our code is available at https://github.com/OpenVeraTeam/LiveMoments.
Authors:Yexiong Lin, Jia Shi, Shanshan Ye, Wanyu Wang, Yu Yao, Tongliang Liu
Abstract:
Flow matching has emerged as a powerful generative framework, with recent few‑step methods achieving remarkable inference acceleration. However, we identify a critical yet overlooked limitation: these models suffer from severe diversity degradation, concentrating samples on dominant modes while neglecting rare but valid variations of the target distribution. We trace this degradation to averaging distortion: when trained with MSE objectives, class‑conditional flows learn a frequency‑weighted mean over intra‑class sub‑modes, causing the model to over‑represent high‑density modes while systematically neglecting low‑density ones. To address this, we propose SubFlow, Sub‑mode Conditioned Flow Matching, which eliminates averaging distortion by decomposing each class into fine‑grained sub‑modes via semantic clustering and conditioning the flow on sub‑mode indices. Each conditioned sub‑distribution is approximately unimodal, so the learned flow accurately targets individual modes with no averaging distortion, restoring full mode coverage in a single inference step. Crucially, SubFlow is entirely plug‑and‑play: it integrates seamlessly into existing one‑step models such as MeanFlow and Shortcut Models without any architectural modifications. Extensive experiments on ImageNet‑256 demonstrate that SubFlow yields substantial gains in generation diversity (Recall) while maintaining competitive image quality (FID), confirming its broad applicability across different one‑step generation frameworks. Project page: https://yexionglin.github.io/subflow.
Authors:Hang Xu, Chen Long, Bing Wang, Hao Chen, Zhen Dong
Abstract:
Underwater Image Enhancement (UIE) is essential for robust visual perception in marine applications. However, existing methods predominantly rely on uniform mapping tailored to average dataset distributions, leading to over‑processing mildly degraded images or insufficient recovery for severe ones. To address this challenge, we propose a novel adaptive enhancement framework, SDAR‑Net. Unlike existing uniform paradigms, it first decouples specific degradation styles from the input and subsequently modulates the enhancement process adaptively. Specifically, since underwater degradation primarily shifts the appearance while keeping the scene structure, SDAR‑Net formulates image features into dynamic degradation style embeddings and static scene structural representations through a carefully designed training framework. Subsequently, we introduce an adaptive routing mechanism. By evaluating style features and adaptively predicting soft weights at different enhancement states, it guides the weighted fusion of the corresponding image representations, accurately satisfying the adaptive restoration demands of each image. Extensive experiments show that SDAR‑Net achieves a new state‑of‑the‑art (SOTA) performance with a PSNR of 25.72 dB on real‑world benchmark, and demonstrates its utility in downstream vision tasks. Our code is available at https://github.com/WHU‑USI3DV/SDAR‑Net.
Authors:Sandra Gómez-Gálvez, Tobias Olenyi, Gillian Dobbie, Katerina Taškova
Abstract:
Deep neural networks, despite their high accuracy, often exhibit poor confidence calibration, limiting their reliability in high‑stakes applications. Current ad‑hoc confidence calibration methods attempt to fix this during training but face a fundamental trade‑off: two‑phase training methods achieve strong classification performance at the cost of training instability and poorer confidence calibration, while single‑loss methods are stable but underperform in classification. This paper addresses and mitigates this stability‑performance trade‑off. We propose Socrates Loss, a novel, unified loss function that explicitly leverages uncertainty by incorporating an auxiliary unknown class, whose predictions directly influence the loss function and a dynamic uncertainty penalty. This unified objective allows the model to be optimized for both classification and confidence calibration simultaneously, without the instability of complex, scheduled losses. We provide theoretical guarantees that our method regularizes the model to prevent miscalibration and overfitting. Across four benchmark datasets and multiple architectures, our comprehensive experiments demonstrate that Socrates Loss consistently improves training stability while achieving more favorable accuracy‑calibration trade‑off, often converging faster than existing methods.
Authors:Qingyuan Cai, Saihui Hou, Xuecai Hu, Yongzhen Huang
Abstract:
Gait recognition, as a reliable biometric technology, has seen rapid development in recent years while it faces significant challenges caused by diverse clothing styles in the real world. This paper introduces BarbieGait, a synthetic gait dataset where real‑world subjects are uniquely mapped into a virtual engine to simulate extensive clothing changes while preserving their gait identity information. As a pioneering work, BarbieGait provides a controllable gait data generation method, enabling the production of large datasets to validate cross‑clothing issues that are difficult to verify with real‑world data. However, the diversity of clothing increases intra‑class variance and makes one of the biggest challenges to learning cloth‑invariant features under varying clothing conditions. Therefore, we propose GaitCLIF (Gait‑oriented CLoth‑Invariant Feature) as a robust baseline model for cross‑clothing gait recognition. Through extensive experiments, we validate that our method significantly improves cross‑clothing performance on BarbieGait and the existing popular gait benchmarks. We believe that BarbieGait, with its extensive cross‑clothing gait data, will further advance the capabilities of gait recognition in cross‑clothing scenarios and promote progress in related research.
Authors:Parth Parag Kulkarni, Rohit Gupta, Prakash Chandra Chhipa, Mubarak Shah
Abstract:
The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and exploration. Existing classification‑based approaches operate at a coarse city‑level granularity and fail to capture fine‑grained details, while image retrieval methods are impractical on a global scale due to the need for extensive image galleries which are infeasible to compile. Comparatively, constructing a gallery of GPS coordinates is straightforward and inexpensive. We propose VidTAG, a dual‑encoder framework that performs frame‑to‑GPS retrieval using both self‑supervised and language‑aligned features. To address temporal inconsistencies in video predictions, we introduce the TempGeo module, which aligns frame embeddings, and the GeoRefiner module, an encoder‑decoder architecture that refines GPS features using the aligned frame embeddings. Evaluations on Mapillary (MSLS) and GAMa datasets demonstrate our model's ability to generate temporally consistent trajectories and outperform baselines, achieving a 20% improvement at the 1 km threshold over GeoCLIP. We also beat current State‑of‑the‑Art by 25% on global coarse grained video geolocalization (CityGuessr68k). Our approach enables fine‑grained video geolocalization and lays a strong foundation for future research. More details on the project webpage: https://parthpk.github.io/vidtag_webpage/
Authors:Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi
Abstract:
Latent diffusion models for medical image super‑resolution universally inherit variational autoencoders designed for natural photographs. We show that this default choice, not the diffusion architecture, is the dominant constraint on reconstruction quality. In a controlled experiment holding all other pipeline components fixed, replacing the generic Stable Diffusion VAE with MedVAE, a domain‑specific autoencoder pretrained on more than 1.6 million medical images, yields +2.91 to +3.29 dB PSNR improvement across knee MRI, brain MRI, and chest X‑ray (n = 1,820; Cohen's d = 1.37 to 1.86, all p < 10^‑20, Wilcoxon signed‑rank). Wavelet decomposition localises the advantage to the finest spatial frequency bands encoding anatomically relevant fine structure. Ablations across inference schedules, prediction targets, and generative architectures confirm the gap is stable within plus or minus 0.15 dB, while hallucination rates remain comparable between methods (Cohen's h < 0.02 across all datasets), establishing that reconstruction fidelity and generative hallucination are governed by independent pipeline components. These results provide a practical screening criterion: autoencoder reconstruction quality, measurable without diffusion training, predicts downstream SR performance (R^2 = 0.67), suggesting that domain‑specific VAE selection should precede diffusion architecture search. Code and trained model weights are publicly available at https://github.com/sebasmos/latent‑sr.
Authors:Md Tanvirul Alam
Abstract:
Large vision‑language models (VLMs) often rely on familiar semantic priors, but existing evaluations do not cleanly separate perception failures from rule‑mapping failures. We study this behavior as semantic fixation: preserving a default interpretation even when the prompt specifies an alternative, equally valid mapping. To isolate this effect, we introduce VLM‑Fix, a controlled benchmark over four abstract strategy games that evaluates identical terminal board states under paired standard and inverse rule formulations. Across 14 open and closed VLMs, accuracy consistently favors standard rules, revealing a robust semantic‑fixation gap. Prompt interventions support this mechanism: neutral alias prompts substantially narrow the inverse‑rule gap, while semantically loaded aliases reopen it. Post‑training is strongly rule‑aligned: training on one rule improves same‑rule transfer but hurts opposite‑rule transfer, while joint‑rule training improves broader transfer. To test external validity beyond synthetic games, we evaluate analogous defamiliarization interventions on VLMBias and observe the same qualitative pattern. Finally, late‑layer activation steering partially recovers degraded performance, indicating that semantic‑fixation errors are at least partly editable in late representations. Project page, code, and dataset available at https://maveryn.github.io/vlm‑fix/.
Authors:Arun Sharma
Abstract:
We introduce compute‑grounded reasoning (CGR), a design paradigm for spatial‑aware research agents in which every answerable sub‑problem is resolved by deterministic computation before a language model is asked to generate. Spatial Atlas instantiates CGR as a single Agent‑to‑Agent (A2A) server that handles two challenging benchmarks: FieldWorkArena, a multimodal spatial question‑answering benchmark spanning factory, warehouse, and retail environments, and MLE‑Bench, a suite of 75 Kaggle machine learning competitions requiring end‑to‑end ML engineering. A structured spatial scene graph engine extracts entities and relations from vision descriptions, computes distances and safety violations deterministically, then feeds computed facts to large language models, thereby avoiding hallucinated spatial reasoning. Entropy‑guided action selection maximizes information gain per step and routes queries across a three‑tier frontier model stack (OpenAI + Anthropic). A self‑healing ML pipeline with strategy‑aware code generation, a score‑driven iterative refinement loop, and a prompt‑based leak audit registry round out the system. We evaluate across both benchmarks and show that CGR yields competitive accuracy while maintaining interpretability through structured intermediate representations and deterministic spatial computations.
Authors:Xingyu Qiu, Yuqian Fu, Jiawei Geng, Bin Ren, Jiancheng Pan, Zongwei Wu, Hao Tang, Yanwei Fu, Radu Timofte, Nicu Sebe, Mohamed Elhoseiny, Lingyi Hong, Mingxi Cheng, Xingqi He, Runze Li, Xingdong Sheng, Wenqiang Zhang, Jiacong Liu, Shu Luo, Yikai Qin, Yaze Zhao, Yongwei Jiang, Yixiong Zou, Zhe Zhang, Yang Yang, Kaiyu Li, Bowen Fu, Zixuan Jiang, Ke Li, Hui Qiao, Xiangyong Cao, Xuanlong Yu, Youyang Sha, Longfei Liu, Di Yang, Xi Shen, Kyeongryeol Go, Taewoong Jang, Saiprasad Meesiyawar, Ravi Kirasur, Rakshita Kulkarni, Bhoomi Deshpande, Harsh Patil, Uma Mudenagudi, Shuming Hu, Chao Chen, Tao Wang, Wei Zhou, Qi Xu, Zhenzhao Xing, Dandan Zhao, Hanzhe Xia, Dongdong Lu, Zhe Zhang, Jingru Wang, Guangwei Huang, Jiachen Tu, Yaokun Shi, Guoyi Xu, Yaoxin Jiang, Jiajia Liu, Liwei Zhou, Bei Dou, Tao Wu, Zekang Fan, Junjie Liu, Adhémar de Senneville, Flavien Armangeon, Mengbers, Yazhe Lyu, Zhimeng Xin, Zijian Zhuang, Hongchun Zhu, Li Wang
Abstract:
Cross‑domain few‑shot object detection (CD‑FSOD) remains a challenging problem for existing object detectors and few‑shot learning approaches, particularly when generalizing across distinct domains. As part of NTIRE 2026, we hosted the second CD‑FSOD Challenge to systematically evaluate and promote progress in detecting objects in unseen target domains under limited annotation conditions. The challenge received strong community interest, with 128 registered participants and a total of 696 submissions. Among them, 31 teams actively participated, and 19 teams submitted valid final results. Participants explored a wide range of strategies, introducing innovative methods that push the performance frontier under both open‑source and closed‑source tracks. This report presents a detailed overview of the NTIRE 2026 CD‑FSOD Challenge, including a summary of the submitted approaches and an analysis of the final results across all participating teams. Challenge Codes: https://github.com/ohMargin/NTIRE2026_CDFSOD.
Authors:Chengkun Yue, Chuanzhi Xu, Jiangpeng He
Abstract:
Nutrition estimation of meals from visual data is an important problem for dietary monitoring and computational health, but existing approaches largely rely on single images of the finally completed dish. This setting is fundamentally limited because many nutritionally relevant ingredients and transformations, such as oils, sauces, and mixed components, become visually ambiguous after cooking, making accurate calorie and macronutrient estimation difficult. In this paper, we investigate whether the cooking process information from egocentric cooking videos can contribute to dish‑level nutrition estimation. First, we further manually annotated the HD‑EPIC dataset and established the first benchmark for video‑based nutrition estimation. Most importantly, we propose V‑Nutri, a staged framework that combines Nutrition5K‑pretrained visual backbones with a lightweight fusion module that aggregates features from the final dish frame and cooking process keyframes extracted from the egocentric videos. V‑Nutri also includes a cooking keyframes selection module, a VideoMamba‑based event‑detection model that targets ingredient‑addition moments. Experiments on the HD‑EPIC dataset show that process cues can provide complementary nutritional evidence, improving nutrition estimation under controlled conditions. Our results further indicate that the benefit of process keyframes depends strongly on backbone representation capacity and event detection quality. Our code and annotated dataset is available at https://github.com/K624‑YCK/V‑Nutri.
Authors:David Nordström, Johan Edstedt, Fredrik Kahl, Georg Bökman
Abstract:
Finding matching keypoints between images is a core problem in 3D computer vision. However, modern matchers struggle with large in‑plane rotations. A straightforward mitigation is to learn rotation invariance via data augmentation. However, it remains unclear at which stage rotation invariance should be incorporated. In this paper, we study this in the context of a modern sparse matching pipeline. We perform extensive experiments by training on a large collection of 3D vision datasets and evaluating on popular image matching benchmarks. Surprisingly, we find that incorporating rotation invariance already in the descriptor yields similar performance to handling it in the matcher. However, rotation invariance is achieved earlier in the matcher when it is learned in the descriptor, allowing for a faster rotation‑invariant matcher. Further, we find that enforcing rotation invariance does not hurt upright performance when trained at scale. Finally, we study the emergence of rotation invariance through scale and find that increasing the training data size substantially improves generalization to rotated images. We release two matchers robust to in‑plane rotations that achieve state‑of‑the‑art performance on e.g. multi‑modal (WxBS), extreme (HardMatch), and satellite image matching (SatAst). Code is available at https://github.com/davnords/loma.
Authors:Mihir Prabhudesai, Aryan Satpathy, Yangmin Li, Zheyang Qin, Nikash Bhardwaj, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak
Abstract:
We have witnessed remarkable advances in LLM reasoning capabilities with the advent of DeepSeek‑R1. However, much of this progress has been fueled by the abundance of internet question‑answer (QA) pairs, a major bottleneck going forward, since such data is limited in scale and concentrated mainly in domains like mathematics. In contrast, other sciences such as physics lack large‑scale QA datasets to effectively train reasoning‑capable models. In this work, we show that physics simulators can serve as a powerful alternative source of supervision for training LLMs for physical reasoning. We generate random scenes in physics engines, create synthetic question‑answer pairs from simulated interactions, and train LLMs using reinforcement learning on this synthetic data. Our models exhibit zero‑shot sim‑to‑real transfer to real‑world physics benchmarks: for example, training solely on synthetic simulated data improves performance on IPhO (International Physics Olympiad) problems by 5‑10 percentage points across model sizes. These results demonstrate that physics simulators can act as scalable data generators, enabling LLMs to acquire deep physical reasoning skills beyond the limitations of internet‑scale QA data. Code available at: https://sim2reason.github.io/.
Authors:Donghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li, Jingyu Lin, Xiaohu Huang, Yichen Liu, Xin Gao, Cunjian Chen, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng
Abstract:
In this work, we study Human‑Object Interaction Video Generation (HOIVG), which aims to synthesize high‑quality human‑object interaction videos conditioned on text, reference images, audio, and pose. This task holds significant practical value for automating content creation in real‑world applications, such as e‑commerce demonstrations, short video production, and interactive entertainment. However, existing approaches fail to accommodate all these requisite conditions. We present OmniShow, an end‑to‑end framework tailored for this practical yet challenging task, capable of harmonizing multimodal conditions and delivering industry‑grade performance. To overcome the trade‑off between controllability and quality, we introduce Unified Channel‑wise Conditioning for efficient image and pose injection, and Gated Local‑Context Attention to ensure precise audio‑visual synchronization. To effectively address data scarcity, we develop a Decoupled‑Then‑Joint Training strategy that leverages a multi‑stage training process with model merging to efficiently harness heterogeneous sub‑task datasets. Furthermore, to fill the evaluation gap in this field, we establish HOIVG‑Bench, a dedicated and comprehensive benchmark for HOIVG. Extensive experiments demonstrate that OmniShow achieves overall state‑of‑the‑art performance across various multimodal conditioning settings, setting a solid standard for the emerging HOIVG task.
Authors:Jinhui Ye, Ning Gao, Senqiao Yang, Jinliang Zheng, Zixuan Wang, Yuxin Chen, Pengguang Chen, Yilun Chen, Shu Liu, Jiaya Jia
Abstract:
Vision‑Language‑Action (VLA) models have recently emerged as a promising paradigm for building general‑purpose robotic agents. However, the VLA landscape remains highly fragmented and complex: as existing approaches vary substantially in architectures, training data, embodiment configurations, and benchmark‑specific engineering. In this work, we introduce StarVLA‑α, a simple yet strong baseline designed to study VLA design choices under controlled conditions. StarVLA‑α deliberately minimizes architectural and pipeline complexity to reduce experimental confounders and enable systematic analysis. Specifically, we re‑evaluate several key design axes, including action modeling strategies, robot‑specific pretraining, and interface engineering. Across unified multi‑benchmark training on LIBERO, SimplerEnv, RoboTwin, and RoboCasa, the same simple baseline remains highly competitive, indicating that a strong VLM backbone combined with minimal design is already sufficient to achieve strong performance without relying on additional architectural complexity or engineering tricks. Notably, our single generalist model outperforms π_0.5 by 20% on the public real‑world RoboChallenge benchmark. We expect StarVLA‑α to serve as a solid starting point for future research in the VLA regime. Code will be released at https://github.com/starVLA/starVLA.
Authors:Nick Stracke, Kolja Bauer, Stefan Andreas Baumann, Miguel Angel Bautista, Josh Susskind, Björn Ommer
Abstract:
Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multiple possible futures through full video synthesis remains prohibitively inefficient. We model scene dynamics orders of magnitude more efficiently by directly operating on a long‑term motion embedding that is learned from large‑scale trajectories obtained from tracker models. This enables efficient generation of long, realistic motions that fulfill goals specified via text prompts or spatial pokes. To achieve this, we first learn a highly compressed motion embedding with a temporal compression factor of 64x. In this space, we train a conditional flow‑matching model to generate motion latents conditioned on task descriptions. The resulting motion distributions outperform those of both state‑of‑the‑art video models and specialized task‑specific approaches.
Authors:Junwoo Park, Jangho Lee, Sunho Lim
Abstract:
Pretrained detectors perform well on benchmarks but often suffer performance degradation in real‑world deployments due to distribution gaps between training data and target environments. COCO‑like benchmarks emphasize category diversity rather than instance density, causing detectors trained under per‑class sparsity to struggle in dense, single‑ or few‑class scenes such as surveillance and traffic monitoring. In fixed‑camera environments, the quasi‑static background provides a stable, label‑free prior that can be exploited at inference to suppress spurious detections. To address the issue, we propose Background Embedding Memory (BEM), a lightweight, training‑free, weight‑frozen module that can be attached to pretrained detectors during inference. BEM estimates clean background embeddings, maintains a prototype memory, and re‑scores detection logits with an inverse‑similarity, rank‑weighted penalty, effectively reducing false positives while maintaining recall. Empirically, background‑frame cosine similarity correlates negatively with object count and positively with Precision‑Confidence AUC (P‑AUC), motivating its use as a training‑free control signal. Across YOLO and RT‑DETR families on LLVIP and simulated surveillance streams, BEM consistently reduces false positives while preserving real‑time performance. Our code is available at https://github.com/Leo‑Park1214/Background‑Embedding‑Memory.git
Authors:Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
Abstract:
Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as autonomous driving. We present Re2Pix, a hierarchical video prediction framework that decomposes forecasting into two stages: semantic representation prediction and representation‑guided visual synthesis. Instead of directly predicting future RGB frames, our approach first forecasts future scene structure in the feature space of a frozen vision foundation model, and then conditions a latent diffusion model on these predicted representations to render photorealistic frames. This decomposition enables the model to focus first on scene dynamics and then on appearance generation. A key challenge arises from the train‑test mismatch between ground‑truth representations available during training and predicted ones used at inference. To address this, we introduce two conditioning strategies, nested dropout and mixed supervision, that improve robustness to imperfect autoregressive predictions. Experiments on challenging driving benchmarks demonstrate that the proposed semantics‑first design significantly improves temporal semantic consistency, perceptual quality, and training efficiency compared to strong diffusion baselines. We provide the implementation code at https://github.com/Sta8is/Re2Pix
Authors:Dujun Nie, Fengjiao Chen, Qi Lv, Jun Kuang, Xiaoyu Li, Xuezhi Cao, Xunliang Cai
Abstract:
While the shortage of explicit action data limits Vision‑Language‑Action (VLA) models, human action videos offer a scalable yet unlabeled data source. A critical challenge in utilizing large‑scale human video datasets lies in transforming visual signals into ontology‑independent representations, known as latent actions. However, the capacity of latent action representation to derive robust control from visual observations has yet to be rigorously evaluated. We introduce the Latent Action Representation Yielding (LARY) Benchmark, a unified framework for evaluating latent action representations on both high‑level semantic actions (what to do) and low‑level robotic control (how to do). The comprehensively curated dataset encompasses over one million videos (1,000 hours) spanning 151 action categories, alongside 620K image pairs and 595K motion trajectories across diverse embodiments and environments. Our experiments reveal two crucial insights: (i) General visual foundation models, trained without any action supervision, consistently outperform specialized embodied latent action models. (ii) Latent‑based visual space is fundamentally better aligned to physical action space than pixel‑based space. These results suggest that general visual representations inherently encode action‑relevant knowledge for physical control, and that semantic‑level abstraction serves as a fundamentally more effective pathway from vision to action than pixel‑level reconstruction.
Authors:Guillaume Astruc, Eduard Trulls, Jan Hosang, Loic Landrieu, Paul-Edouard Sarlin
Abstract:
The growing availability of co‑located geospatial data spanning aerial imagery, street‑level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all‑to‑all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude‑longitude encoder that improves spatial representation by capturing multi‑scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single‑modality contrastive models and coordinate‑only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip.
Authors:Wenhao Li, Xueying Jiang, Gongjie Zhang, Xiaoqin Zhang, Ling Shao, Shijian Lu
Abstract:
4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the spatiotemporal domain where the underlying geometric characteristics of 4D point cloud videos are hard to capture, leading to degraded representation learning and understanding of 4D point cloud videos. We address the above challenge from a complementary spectral perspective. By transforming 4D point cloud videos into graph spectral signals, we can decompose them into multiple frequency bands each of which captures distinct geometric structures of point cloud videos. Our spectral analysis reveals that the decomposed low‑frequency signals capture more coarse shapes while high‑frequency signals encode more fine‑grained geometry details. Building on these observations, we design Spatio‑Temporal‑Spectral Mixer (STS‑Mixer), a unified framework that mixes spatial, temporal, and spectral representations of point cloud videos. STS‑Mixer integrates multi‑band delineated spectral signals with spatiotemporal information to capture rich geometries and temporal dynamics, while enabling fine‑grained and holistic understanding of 4D point cloud videos. Extensive experiments show that STS‑Mixer achieves superior performance consistently across multiple widely adopted benchmarks on both 3D action recognition and 4D semantic segmentation tasks. Code and models are available at https://github.com/Vegetebird/STS‑Mixer.
Authors:Songlong Xing, Weijie Wang, Zhengyu Zhao, Jindong Gu, Philip Torr, Nicu Sebe
Abstract:
Despite their impressive zero‑shot abilities, vision‑language models such as CLIP have been shown to be susceptible to adversarial attacks. To enhance its adversarial robustness, recent studies finetune the pretrained vision encoder of CLIP with adversarial examples on a proxy dataset such as ImageNet by aligning adversarial images with correct class labels. However, these methods overlook the important roles of training data distributions and learning objectives, resulting in reduced zero‑shot capabilities and limited transferability of robustness across domains and datasets. In this work, we propose a simple yet effective paradigm AdvFLYP, which follows the training recipe of CLIP's pretraining process when performing adversarial finetuning to the model. Specifically, AdvFLYP finetunes CLIP with adversarial images created based on image‑text pairs collected from the web, and match them with their corresponding texts via a contrastive loss. To alleviate distortion of adversarial image embeddings of noisy web images, we further propose to regularise AdvFLYP by penalising deviation of adversarial image features. We show that logit‑ and feature‑level regularisation terms benefit robustness and clean accuracy, respectively. Extensive experiments on 14 downstream datasets spanning various domains show the superiority of our paradigm over mainstream practices. Our code and model weights are released at https://github.com/Sxing2/AdvFLYP.
Authors:Sohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun Oh
Abstract:
Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolithic metric that cannot incorporate multiple conditions simultaneously. To address this, we propose CLAY, an adaptive similarity computation method that reframes the embedding space of pretrained Vision‑Language Models (VLMs) as a text‑conditional similarity space without additional training. This design separates the textual conditioning process and visual feature extraction, allowing highly efficient and multi‑conditioned retrieval with fixed visual embeddings. We also construct a synthetic evaluation dataset CLAY‑EVAL, for comprehensive assessment under diverse conditioned retrieval settings. Experiments on standard datasets and our proposed dataset show that CLAY achieves high retrieval accuracy and notable computational efficiency compared to previous works.
Authors:Yang Ji, Zonghao Chen, Zhihao Xue, Junqin Hu
Abstract:
Real‑world image super‑resolution is particularly challenging for diffusion models because real degradations are complex, heterogeneous, and rarely modeled explicitly. We propose a degradation‑aware and structure‑preserving diffusion framework for real‑world SR. Specifically, we introduce Degradation‑aware Token Injection, which encodes lightweight degradation statistics from low‑resolution inputs and fuses them with semantic conditioning features, enabling explicit degradation‑aware restoration. We further propose Spatially Asymmetric Noise Injection, which modulates diffusion noise with local edge strength to better preserve structural regions during training. Both modules are lightweight add‑ons to the adopted diffusion SR framework, requiring only minor modifications to the conditioning pipeline. Experiments on DIV2K and RealSR show that our method delivers competitive no‑reference perceptual quality and visually more realistic restoration results than recent baselines, while maintaining a favorable perception‑‑distortion trade‑off. Ablations confirm the effectiveness of each module and their complementary gains when combined. The code and model are publicly available at https://github.com/jiyang0315/DASP‑SR.git.
Authors:Qilin Zhang, Jinyu Zhu, Olaf Wysocki, Benjamin Busam, Boris Jutzi
Abstract:
Recent semantic 3D Gaussian Splatting (3DGS) methods primarily rely on 2D foundation models, often yielding ambiguous boundaries and limited support for structured urban semantics. While city models such as CityGML encode hierarchically organized semantics together with building geometry, these labels cannot be directly mapped to Gaussian primitives. We present GS4City, a hierarchical semantic Gaussian Splatting method that incorporates city‑model priors for urban scene understanding. GS4City derives reliable image‑aligned masks from Level of Detail (LoD) 3 CityGML models via two‑pass raycasting, explicitly using parent‑child relations to validate and recover fine‑grained facade elements. It then fuses these geometry‑grounded masks with foundation‑model predictions to establish scene‑consistent instance correspondences, and learns a compact identity encoding for each Gaussian under joint 2D identity supervision and 3D spatial regularization. Experiments on the TUM2TWIN and Gold Coast datasets show that GS4City effectively incorporates structured building semantics into Gaussian scene representations, outperforming existing 2D‑driven semantic 3DGS baselines, including LangSplat and Gaga, by up to 15.8 IoU points in coarse building segmentation and 14.2 mIoU points in fine‑grained semantic segmentation. By bridging structured city models and photorealistic Gaussian scene representations, GS4City enables semantically queryable and structure‑aware urban reconstruction. Code is available at https://github.com/Jinyzzz/GS4City.
Authors:Jijun Xiang, Tao Wang, Jiayi Wang, Pengxiang Wang, Cheng Chen, Nian Wang
Abstract:
While Hyperspectral Anomaly Detection (HAD) excels at identifying sparse targets in complex scenes, existing models remain trapped in a scalar "reconstruction‑as‑endpoint" paradigm. This reliance on ambiguous scalar residuals consistently triggers sub‑pixel anomaly vanishing during spatial downsampling, alongside severe confirmation bias when unpurified anomalies corrupt training weights. In this paper, we propose Reconstruction‑to‑Vector Diffusion (R2VD), which fundamentally redefines reconstruction as a manifold purification origin to establish a novel residual‑guided generative dynamics paradigm. Our framework introduces a four‑stage pipeline: (1) a Physical Prior Extraction (PPE) stage that mitigates early confirmation bias via dual‑stream statistical guidance; (2) a Guided Manifold Purification (GMP) stage utilizing an OmniContext Autoencoder (OCA) to extract purified residual maps while preserving fragile sub‑pixel topologies; (3) a Residual Score Modeling (RSM) stage where a Diffusion Transformer (DiT), guarded by a Physical Spectral Firewall (PSF), effectively isolates cross‑spectral leakage; and (4) a Vector Dynamics Inference (VDI) stage that robustly decouples targets from backgrounds by evaluating high‑dimensional vector interference patterns instead of conventional scalar errors. Comprehensive evaluations on eight datasets confirm that R2VD establishes a new state‑of‑the‑art, delivering exceptional target detectability and background suppression. The code is available at https://github.com/Bondojijun/R2VD.
Authors:Yiran Qin, Jiahua Ma, Li Kang, Wenzhan Li, Yihang Jiao, Xin Wen, Xiufeng Song, Heng Zhou, Jiwen Yu, Zhenfei Yin, Xihui Liu, Philip Torr, Yilun Du, Ruimao Zhang
Abstract:
Recent advancements in foundational models, such as large language models and world models, have greatly enhanced the capabilities of robotics, enabling robots to autonomously perform complex tasks. However, acquiring large‑scale, high‑quality training data for robotics remains a challenge, as it often requires substantial manual effort and is limited in its coverage of diverse real‑world environments. To address this, we propose a novel hybrid approach called Compositional Simulation, which combines classical simulation and neural simulation to generate accurate action‑video pairs while maintaining real‑world consistency. Our approach utilizes a closed‑loop real‑sim‑real data augmentation pipeline, leveraging a small amount of real‑world data to generate diverse, large‑scale training datasets that cover a broader spectrum of real‑world scenarios. We train a neural simulator to transform classical simulation videos into real‑world representations, improving the accuracy of policy models trained in real‑world environments. Through extensive experiments, we demonstrate that our method significantly reduces the sim2real domain gap, resulting in higher success rates in real‑world policy model training. Our approach offers a scalable solution for generating robust training data and bridging the gap between simulated and real‑world robotics.
Authors:Koki Ryu, Hitomi Yanaka
Abstract:
Personalized image aesthetics assessment (PIAA) is an important research problem with practical real‑world applications. While methods based on vision‑language models (VLMs) are promising candidates for PIAA, it remains unclear whether they internally encode rich, multi‑level aesthetic attributes required for effective personalization. In this paper, we first analyze the internal representations of VLMs to examine the presence and distribution of such aesthetic attributes, and then leverage them for lightweight, individual‑level personalization without model fine‑tuning. Our analysis reveals that VLMs encode diverse aesthetic attributes that propagate into the language decoder layers. Building on these representations, we demonstrate that simple linear models can perform PIAA effectively. We further analyze how aesthetic information is transferred across layers in different VLM architectures and across image domains. Our findings provide insights into how VLMs can be utilized for modeling subjective, individual aesthetic preferences. Our code is available at https://github.com/ynklab/vlm‑latent‑piaa.
Authors:Jianshi Wu, Minghang Zhu, Dunqiang Liu, Wen Li, Sheng Ao, Siqi Shen, Chenglu Wen, Cheng Wang
Abstract:
LiDAR relocalization has attracted increasing attention as it can deliver accurate 6‑DoF pose estimation in complex 3D environments. Recent learning‑based regression methods offer efficient solutions by directly predicting global poses without the need for explicit map storage. However, these methods often struggle in challenging scenes due to their equal treatment of all predicted points, which is vulnerable to noise and outliers. In this paper, we propose LEADER, a robust LiDAR‑based relocalization framework enhanced by a simple, yet effective geometric encoder. Specifically, a Robust Projection‑based Geometric Encoder architecture which captures multi‑scale geometric features is first presented to enhance descriptiveness in geometric representation. A Truncated Relative Reliability loss is then formulated to model point‑wise ambiguity and mitigate the influence of unreliable predictions. Extensive experiments on the Oxford RobotCar and NCLT datasets demonstrate that LEADER outperforms state‑of‑the‑art methods, achieving 24.1% and 73.9% relative reductions in position error over existing techniques, respectively. The source code is released on https://github.com/JiansW/LEADER.
Authors:Dongxu Wei, Qi Xu, Zhiqi Li, Hangning Zhou, Cong Qiu, Hailong Qin, Mu Yang, Zhaopeng Cui, Peidong Liu
Abstract:
3D scene generation has long been dominated by 2D multi‑view or video diffusion models. This is due not only to the lack of scene‑level 3D latent representation, but also to the fact that most scene‑level 3D visual data exists in the form of multi‑view images or videos, which are naturally compatible with 2D diffusion architectures. Typically, these 2D‑based approaches degrade 3D spatial extrapolation to 2D temporal extension, which introduces two fundamental issues: (i) representing 3D scenes via 2D views leads to significant representation redundancy, and (ii) latent space rooted in 2D inherently limits the spatial consistency of the generated 3D scenes. In this paper, we propose, for the first time, to perform 3D scene generation directly within an implicit 3D latent space to address these limitations. First, we repurpose frozen 2D representation encoders to construct our 3D Representation Autoencoder (3DRAE), which grounds view‑coupled 2D semantic representations into a view‑decoupled 3D latent representation. This enables representing 3D scenes observed from arbitrary numbers of views‑‑at any resolution and aspect ratio‑‑with fixed complexity and rich semantics. Then we introduce 3D Diffusion Transformer (3DDiT), which performs diffusion modeling in this 3D latent space, achieving remarkably efficient and spatially consistent 3D scene generation while supporting diverse conditioning configurations. Moreover, since our approach directly generates a 3D scene representation, it can be decoded to images and optional point maps along arbitrary camera trajectories without requiring per‑trajectory diffusion sampling pass, which is common in 2D‑based approaches.
Authors:Jiaqi Wu, Zhen Wang, Enhao Huang, Kangqing Shen, Yulin Wang, Yang Yue, Yifan Pu, Gao Huang
Abstract:
Text‑guided multispectral object detection uses text semantics to guide semantic‑aware cross‑modal interaction between RGB and IR for more robust perception. However, notable limitations remain: (1) existing methods often use text only as an auxiliary semantic enhancement signal, without exploiting its guiding role to bridge the inherent granularity asymmetry between RGB and IR; and (2) conventional data‑driven attention‑based fusion tends to emphasize stable consensus while overlooking potentially valuable cross‑modal discrepancies. To address these issues, we propose a semantic bridge fusion framework with bi‑support modeling for multispectral object detection. Specifically, text is used as a shared semantic bridge to align RGB and IR responses under a unified category condition, while the recalibrated thermal semantic prior is projected onto the RGB branch for semantic‑level mapping fusion. We further formulate RGB‑IR interaction evidence into the regular consensus support and the complementary discrepancy support that contains potentially discriminative cues, and introduce them into fusion via dynamic recalibration as a structured inductive bias. In addition, we design a bidirectional semantic alignment module for closed‑loop vision‑text guidance enhancement. Extensive experiments demonstrate the effectiveness of the proposed fusion framework and its superior detection performance on multispectral benchmarks. Code is available at https://github.com/zhenwang5372/Bridging‑RGB‑IR‑Gap.
Authors:You Su, Yonghong Song, Jingqi Chen, Zehan Wen
Abstract:
Change detection is a fundamental task in remote sensing, aiming to quantify the impacts of human activities and ecological dynamics on land‑cover changes. Existing change detection methods are limited to predefined classes in training datasets, which constrains their scalability in real‑world scenarios. In recent years, numerous advanced open‑vocabulary semantic segmentation models have emerged for remote sensing imagery. However, there is still a lack of an effective framework for directly applying these models to open‑vocabulary change detection (OVCD), a novel task that integrates vision and language to detect changes across arbitrary categories. To address these challenges, we first construct a category‑agnostic change detection dataset, termed CA‑CDD. Further, we design a category‑agnostic change head to detect the transitions of arbitrary categories and index them to specific classes. Based on them, we propose Seg2Change, an adapter designed to adapt open‑vocabulary semantic segmentation models to change detection task. Without bells and whistles, this simple yet effective framework achieves state‑of‑the‑art OVCD performance (+9.52 IoU on WHU‑CD and +5.50 mIoU on SECOND). Our code is released at https://github.com/yogurts‑sy/Seg2Change.
Authors:Ya-nan Guan, Shaonan Zhang, Hang Guo, Yawen Wang, Xinying Fan, Tianqu Zhuang, Jie Liang, Hui Zeng, Guanyi Qin, Lishen Qu, Tao Dai, Shu-Tao Xia, Lei Zhang, Radu Timofte, Bin Chen, Yuanbo Zhou, Hongwei Wang, Qinquan Gao, Tong Tong, Yanxin Qian, Lizhao You, Jingru Cong, Lei Xiong, Shuyuan Zhu, Zhi-Qiang Zhong, Kan Lv, Yang Yang, Kailing Tang, Minjian Zhang, Zhipei Lei, Zhe Xu, Liwen Zhang, Dingyong Gou, Yanlin Wu, Cong Li, Xiaohui Cui, Jiajia Liu, Guoyi Xu, Yaoxin Jiang, Yaokun Shi, Jiachen Tu, Liqing Wang, Shihang Li, Bo Zhang, Biao Wang, Haiming Xu, Xiang Long, Xurui Liao, Yanqiao Zhai, Haozhe Li, Shijun Shi, Jiangning Zhang, Yong Liu, Kai Hu, Jing Xu, Xianfang Zeng, Yuyang Liu, Minchen Wei
Abstract:
In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI Flash Portrait. Despite significant advancements in deep learning for image restoration, existing models still encounter substantial challenges in real‑world low‑light portrait scenarios. Specifically, they struggle to achieve an optimal balance among noise suppression, detail preservation, and faithful illumination and color reproduction. To bridge this gap, this challenge aims to establish a novel benchmark for real‑world low‑light portrait restoration. We comprehensively evaluate the proposed algorithms utilizing a hybrid evaluation system that integrates objective quantitative metrics with rigorous subjective assessment protocols. For this competition, we provide a dataset containing 800 groups of real‑captured low‑light portrait data. Each group consists of a 1K‑resolution low‑light input image, a 1K ground truth (GT), and a 1K person mask. This challenge has garnered widespread attention from both academia and industry, attracting over 100 participating teams and receiving more than 3,000 valid submissions. This report details the motivation behind the challenge, the dataset construction process, the evaluation metrics, and the various phases of the competition. The released dataset and baseline code for this track are publicly available from the same \hrefhttps://github.com/zsn1434/AI_Flash‑BaseLine/tree/mainGitHub repository, and the official challenge webpage is hosted on \hrefhttps://www.codabench.org/competitions/12885/CodaBench.
Authors:Julien Walther, Rémi Giraud, Michaël Clément
Abstract:
Superpixels offer a compact image representation by grouping pixels into coherent regions. Recent methods have reached a plateau in terms of segmentation accuracy by generating noisy superpixel shapes. Moreover, most existing approaches produce a single fixed‑scale partition that limits their use in vision pipelines that would benefit multi‑scale representations. In this work, we introduce H‑SPAM (Hierarchical Superpixel Anything Model), a unified framework for generating accurate, regular, and perfectly nested hierarchical superpixels. Starting from a fine partition, guided by deep features and external object priors, H‑SPAM constructs the hierarchy through a two‑phase region merging process that first preserves object consistency and then allows controlled inter‑object grouping. The hierarchy can also be modulated using visual attention maps or user input to preserve important regions longer in the hierarchy. Experiments on standard benchmarks show that H‑SPAM strongly outperforms existing hierarchical methods in both accuracy and regularity, while performing on par with most recent state‑of‑the‑art non‑hierarchical methods. Code and pretrained models are available: https://github.com/waldo‑j/hspam.
Authors:Stefan Schulz, Fernando Edelstein, Hannah Dröge, Matthias B. Hullin, Markus Plack
Abstract:
Real‑time free‑viewpoint rendering requires balancing multi‑camera redundancy with the latency constraints of interactive applications. We address this challenge by combining lightweight geometry with learning and propose 3DTV, a feedforward network for real‑time sparse‑view interpolation. A Delaunay‑based triplet selection ensures angular coverage for each target view. Building on this, we introduce a pose‑aware depth module that estimates a coarse‑to‑fine depth pyramid, enabling efficient feature reprojection and occlusion‑aware blending. Unlike methods that require scene‑specific optimization, 3DTV runs feedforward without retraining, making it practical for AR/VR, telepresence, and interactive applications. Our experiments on challenging multi‑view video datasets demonstrate that 3DTV consistently achieves a strong balance of quality and efficiency, outperforming recent real‑time novel‑view baselines. Crucially, 3DTV avoids explicit proxies, enabling robust rendering across diverse scenes. This makes it a practical solution for low‑latency multi‑view streaming and interactive rendering.
Project Page: https://stefanmschulz.github.io/3DTV_webpage/
Authors:Tuo Liu, Shuijin Lin, Shaozhen Yan, Haifeng Wang, Jie Lu, Jianhua Ma, Chunfeng Lian
Abstract:
The biological definition of Alzheimer's disease (AD) relies on multi‑modal neuroimaging, yet the clinical utility of positron emission tomography (PET) is limited by cost and radiation exposure, hindering early screening at preclinical or prodromal stages. While generative models offer a promising alternative by synthesizing PET from magnetic resonance imaging (MRI), achieving subject‑specific precision remains a primary challenge. Here, we introduce DIReCT++, a Domain‑Informed ReCTified flow model for synthesizing multi‑tracer PET from MRI combined with fundamental clinical information. Our approach integrates a 3D rectified flow architecture to capture complex cross‑modal and cross‑tracer relationships with a domain‑adapted vision‑language model (BiomedCLIP) that provides text‑guided, personalized generation using clinical scores and imaging knowledge. Extensive evaluations on multi‑center datasets demonstrate that DIReCT++ not only produces synthetic PET images (^18F‑AV‑45 and ^18F‑FDG) of superior fidelity and generalizability but also accurately recapitulates disease‑specific patterns. Crucially, combining these synthesized PET images with MRI enables precise personalized stratification of mild cognitive impairment (MCI), advancing a scalable, data‑efficient tool for the early diagnosis and prognostic prediction of AD. The source code will be released on https://github.com/ladderlab‑xjtu/DIReCT‑PLUS.
Authors:Camile Lendering, Erkut Akdag, Egor Bondarev
Abstract:
Accurate defect segmentation is critical for industrial inspection, yet dense pixel‑level annotations are rarely available. A common workaround is to convert inexpensive bounding boxes into pseudo‑masks using foundation segmentation models such as the Segment Anything Model (SAM). However, these pseudo‑labels are systematically noisy on industrial surfaces, often hallucinating background structure while missing sparse defects.
To address this limitation, a noise‑robust box‑to‑pixel distillation framework, Boxes2Pixels, is proposed that treats SAM as a noisy teacher rather than a source of ground‑truth supervision. Bounding boxes are converted into pseudo‑masks offline by SAM, and a compact student is trained with (i) a hierarchical decoder over frozen DINOv2 features for semantic stability, (ii) an auxiliary binary localization head to decouple sparse foreground discovery from class prediction, and (iii) a one‑sided online self‑correction mechanism that relaxes background supervision when the student is confident, targeting teacher false negatives.
On a manually annotated wind turbine inspection benchmark, the proposed Boxes2Pixels improves anomaly mIoU by +6.97 and binary IoU by +9.71 over the strongest baseline trained under identical weak supervision. Moreover, online self‑correction increases the binary recall by +18.56, while the model employs 80% fewer trainable parameters. Code is available at https://github.com/CLendering/Boxes2Pixels.
Authors:Tianyang Dai, Ming Chang, Yan Chen, Yang Hu
Abstract:
Unsupervised remote photoplethysmography (rPPG) promises to leverage unlabeled video data, but its potential is hindered by a critical challenge: training on low‑quality "in‑the‑wild" videos severely degrades model performance. An essential step missing here is to assess the suitability of the videos for rPPG model learning before using them for the task. Existing video quality assessment (VQA) methods are mainly designed for human perception and not directly applicable to the above purpose. In this work, we propose rPPG‑VQA, a novel framework for assessing video suitability for rPPG. We integrate signal‑level and scene‑level analyses and design a dual‑branch assessment architecture. The signal‑level branch evaluates the physiological signal quality of the videos via robust signal‑to‑noise ratio (SNR) estimation with a multi‑method consensus mechanism, and the scene‑level branch uses a multimodal large language model (MLLM) to identify interferences like motion and unstable lighting. Furthermore, we propose a two‑stage adaptive sampling (TAS) strategy that utilizes the quality score to curate optimal training datasets. Experiments show that by training on large‑scale, "in‑the‑wild" videos filtered by our framework, we can develop unsupervised rPPG models that achieve a substantial improvement in accuracy on standard benchmarks. Our code is available at https://github.com/Tianyang‑Dai/rPPG‑VQA.
Authors:Runyu Zhu, SiXun Dong, Zhiqiang Zhang, Qingxia Ye, Zhihua Xu
Abstract:
Low‑light conditions severely hinder 3D restoration and reconstruction by degrading image visibility, introducing color distortions, and contaminating geometric priors for downstream optimization. We present NAKA‑GS, a bionics‑inspired framework for low‑light 3D Gaussian Splatting that jointly improves photometric restoration and geometric initialization. Our method starts with a Naka‑guided chroma‑correction network, which combines physics‑prior low‑light enhancement, dual‑branch input modeling, frequency‑decoupled correction, and mask‑guided optimization to suppress bright‑region chromatic artifacts and edge‑structure errors. The enhanced images are then fed into a feed‑forward multi‑view reconstruction model to produce dense scene priors. To further improve Gaussian initialization, we introduce a lightweight Point Preprocessing Module (PPM) that performs coordinate alignment, voxel pooling, and distance‑adaptive progressive pruning to remove noisy and redundant points while preserving representative structures. Without introducing heavy inference overhead, NAKA‑GS improves restoration quality, training stability, and optimization efficiency for low‑light 3D reconstruction. The proposed method was presented in the NTIRE 3D Restoration and Reconstruction (3DRR) Challenge, and outperformed the baseline methods by a large margin. The code is available at https://github.com/RunyuZhu/Naka‑GS
Authors:Jinsung Lee, Jaemin Oh, Namhun Kim, Dongwon Kim, Byung-Jun Yoon, Suha Kwak
Abstract:
Image tokenizers play a central role in modern generative models, where the structure of the latent space critically determines the downstream generation performance. A key but underexplored property of effective latent representations is spectral organization, the ability to encode information across frequency components. In this work, we introduce structured state‑space regularization, a principled approach to inducing spectral structure in latent spaces. We derive a regularization objective by revisiting state‑space models (SSMs) as systems mimicking a basis function's behavior. This perspective reveals that hidden states of SSMs are induced to capture the frequency components, resulting in a novel regularizer that enforces the latent space to capture spectral structure of images. Experiments demonstrate that our regularizer improves the generative performance of image tokenizers while incurring only minimal loss in their reconstruction fidelity.
Authors:Yakun Yu, Ashley Wiens, Adrián Barahona-Ríos, Benedict Wilkins, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer
Abstract:
Visual glitches in video games degrade player experience and perceived quality, yet manual quality assurance cannot scale to the growing test surface of modern game development. Prior automation efforts, particularly those using vision‑language models (VLMs), largely operate on single frames or rely on limited video‑level baselines that struggle under realistic scene variation, making robust video‑level glitch detection challenging. We present RESP, a practical multi‑frame framework for gameplay glitch detection with VLMs. Our key idea is reference‑guided prompting: for each test frame, we select a reference frame from earlier in the same video, establishing a visual baseline and reframing detection as within‑video comparison rather than isolated classification. RESP sequentially prompts the VLM with reference/test pairs and aggregates noisy frame predictions into a stable video‑level decision without fine‑tuning the VLM. To enable controlled analysis of reference effects, we introduce RefGlitch, a synthetic dataset of manually labeled reference/test frame pairs with balanced coverage across five glitch types. Experiments across five VLMs and three datasets (one synthetic, two real‑world) show that reference guidance consistently strengthens frame‑level detection and that the improved frame‑level evidence reliably transfers to stronger video‑level triage under realistic QA conditions. Code and data are available at: \hrefhttps://github.com/PipiZong/RESP_code.gitthis https URL.
Authors:Weikun Peng, Denys Iliash, Manolis Savva
Abstract:
We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling from readily available real‑world videos valuable. Our task focuses on obtaining simulation‑ready interactive 3D objects from egocentric video input. While prior work largely focuses on articulations, we capture general cross‑part functional mappings (e.g., rotation of stove knob controls stove burner temperature) through function templates, a structured computational representation. Function templates enable precise evaluation and direct compilation into executable code across simulation platforms. To enable comprehensive benchmarking, we introduce a dataset of 271 egocentric videos featuring challenging real‑world interactions with paired 3D geometry, segmentation over 2D and 3D, articulation and function template annotations. To tackle the task, we propose a 4‑stage pipeline consisting of: 2D part segmentation, reconstruction, articulation estimation, and function template inference. Comprehensive benchmarking shows that the task is challenging for off‑the‑shelf methods, highlighting avenues for future work.
Authors:Zhiyuan Zhang, Zijian Zhou, Linjun Li, Long Chen, Hao Tang, Yichen Gong
Abstract:
3D texture generation is receiving increasing attention, as it enables the creation of realistic and aesthetic texture materials for untextured 3D meshes. However, existing 3D texture generation methods are limited to producing only a few types of non‑emissive PBR materials (e.g., albedo, metallic maps and roughness maps), making them difficult to replicate highly popular styles, such as cyberpunk, failing to achieve effects like realistic LED emissions. To address this limitation, we propose a novel task, emission texture generation, which enables the synthesized 3D objects to faithfully reproduce the emission materials from input reference images. Our key contributions include: first, We construct the Objaverse‑Emission dataset, the first dataset that contains 40k 3D assets with high‑quality emission materials. Second, we propose EmissionGen, a novel baseline for the emission texture generation task. Third, we define detailed evaluation metrics for the emission texture generation task. Our results demonstrate significant potential for future industrial applications. Dataset will be available at https://github.com/yx345kw/EmissionGen.
Authors:Joanna Kaleta, Piotr Wójcik, Kacper Marzol, Tomasz Trzciński, Kacper Kania, Marek Kowalski
Abstract:
In 3D reconstruction, the problem of inverse rendering, namely recovering the illumination of the scene and the material properties, is fundamental. Existing Gaussian Splatting‑based methods primarily target static scenes and often assume simplified or moderate lighting to avoid entangling shadows with surface appearance. This limits their ability to accurately separate lighting effects from material properties, particularly in real‑world conditions. We address this limitation by leveraging dynamic elements ‑ regions of the scene that undergo motion ‑ as a supervisory signal for inverse rendering. Motion reveals the same surfaces under varying lighting conditions, providing stronger cues for disentangling material and illumination. This thesis is supported by our experimental results which show we improve LPIPS by 23% for albedo estimation and by 15% for scene relighting relative to next‑best baseline. To this end, we introduce LumiMotion, the first Gaussian‑based approach that leverages dynamics for inverse rendering and operates in arbitrary dynamic scenes. Our method learns a dynamic 2D Gaussian Splatting representation that employs a set of novel constraints which encourage the dynamic regions of the scene to deform, while keeping static regions stable. As we demonstrate, this separation is crucial for correct optimization of the albedo. Finally, we release a new synthetic benchmark comprising five scenes under four lighting conditions, each in both static and dynamic variants, for the first time enabling systematic evaluation of inverse rendering methods in dynamic environments and challenging lighting. Link to project page: https://joaxkal.github.io/LumiMotion/
Authors:Peng Yuan, Yuyang Yin, Yuxuan Cai, Zheng Wei
Abstract:
Existing browser agent benchmarks face a fundamental trilemma: real‑website benchmarks lack reproducibility due to content drift, controlled environments sacrifice realism by omitting real‑web noise, and both require costly manual curation that limits scalability. We present WebForge, the first fully automated framework that resolves this trilemma through a four‑agent pipeline ‑‑ Plan, Generate, Refine, and Validate ‑‑ that produces interactive, self‑contained web environments end‑to‑end without human annotation. A seven‑dimensional difficulty control framework structures task design along navigation depth, visual complexity, reasoning difficulty, and more, enabling systematic capability profiling beyond single aggregate scores. Using WebForge, we construct WebForge‑Bench, a benchmark of 934 tasks spanning 7 domains and 3 difficulty levels. Multi‑model experiments show that difficulty stratification effectively differentiates model capabilities, while cross‑domain analysis exposes capability biases invisible to aggregate metrics. Together, these results confirm that multi‑dimensional evaluation reveals distinct capability profiles that a single aggregate score cannot capture. Code and benchmark are publicly available at https://github.com/yuandaxia2001/WebForge.
Authors:Jinhui Hou, Zhiyu Zhu, Junhui Hou
Abstract:
Diffusion bridge models have shown great promise in image restoration by explicitly connecting clean and degraded image distributions. However, they often rely on complex and high‑cost trajectories, which limit both sampling efficiency and final restoration quality. To address this, we propose an Energy‑oriented diffusion Bridge (E‑Bridge) framework to approximate a set of low‑cost manifold geodesic trajectories to boost the performance of the proposed method. We achieve this by designing a novel bridge process that evolves over a shorter time horizon and makes the reverse process start from an entropy‑regularized point that mixes the degraded image and Gaussian noise, which theoretically reduces the required trajectory energy. To solve this process efficiently, we draw inspiration from consistency models to learn a single‑step mapping function, optimized via a continuous‑time consistency objective tailored for our trajectory, so as to analytically map any state on the trajectory to the target image. Notably, the trajectory length in our framework becomes a tunable task‑adaptive knob, allowing the model to adaptively balance information preservation against generative power for tasks of varying degradation, such as denoising versus super‑resolution. Extensive experiments demonstrate that our E‑Bridge achieves state‑of‑the‑art performance across various image restoration tasks while enabling high‑quality recovery with a single or fewer sampling steps. Our project page is https://jinnh.github.io/E‑Bridge/.
Authors:Qiang Gao, Yi Wang, Yong Zhang, Yong Li, Yongbing Deng, Lan Du, Cunjian Chen
Abstract:
Medical image segmentation remains challenging due to limited fine‑grained annotations, complex anatomical structures, and image degradation from noise, low contrast, or illumination variation. We propose TAMISeg, a text‑guided segmentation framework that incorporates clinical language prompts and semantic distillation as auxiliary semantic cues to enhance visual understanding and reduce reliance on pixel‑level fine‑grained annotations. TAMISeg integrates three core components: a consistency‑aware encoder pretrained with strong perturbations for robust feature extraction, a semantic encoder distillation module with supervision from a frozen DINOv3 teacher to enhance semantic discriminability, and a scale‑adaptive decoder that segments anatomical structures across different spatial scales. Experiments on the Kvasir‑SEG, MosMedData+, and QaTa‑COV19 datasets demonstrate that TAMISeg consistently outperforms existing uni‑modal and multi‑modal methods in both qualitative and quantitative evaluations. Code will be made publicly available at https://github.com/qczggaoqiang/TAMISeg.
Authors:Ye Wang, Kai Huang, Sumin Shen, Chenyang Ma
Abstract:
Referring Camouflaged Object Detection (Ref‑COD) focuses on segmenting specific camouflaged targets in a query image using category‑aligned references. Despite recent advances, existing methods struggle with reference‑target semantic alignment, explicit uncertainty modeling, and robust boundary preservation. To address these issues, we propose EviRCOD, an integrated framework consisting of three core components: (1) a Reference‑Guided Deformable Encoder (RGDE) that employs hierarchical reference‑driven modulation and multi‑scale deformable aggregation to inject semantic priors and align cross‑scale representations; (2) an Uncertainty‑Aware Evidential Decoder (UAED) that incorporates Dirichlet evidence estimation into hierarchical decoding to model uncertainty and propagate confidence across scales; and (3) a Boundary‑Aware Refinement Module (BARM) that selectively enhances ambiguous boundaries by exploiting low‑level edge cues and prediction confidence. Experiments on the Ref‑COD benchmark demonstrate that EviRCOD achieves state‑of‑the‑art detection performance while providing well‑calibrated uncertainty estimates. Code is available at: https://github.com/blueecoffee/EviRCOD.
Authors:Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Jiankang Deng, Cordelia Schmid, Stefanos Zafeiriou
Abstract:
Generating realistic 3D hand‑object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high‑fidelity physical plausibility. Existing methods remain limited in their ability to learn expressive motion representations for generation and perform temporal reasoning. In this paper, we present HO‑Flow, a framework for synthesizing realistic hand‑object motion sequences from texts and canoncial 3D objects. HO‑Flow first employs an interaction‑aware variational autoencoder to encode sequences of hand and object motions into a unified latent manifold by incorporating hand and object kinematics, enabling the representation to capture rich interaction dynamics. It then leverages a masked flow matching model that combines auto‑regressive temporal reasoning with continuous latent generation, improving temporal coherence. To further enhance generalization, HO‑Flow predicts object motions relative to the initial frame, enabling effective pre‑training on large‑scale synthetic data. Experiments on the GRAB, OakInk, and DexYCB benchmarks demonstrate that HO‑Flow achieves state‑of‑the‑art performance in both physical plausibility and motion diversity for interaction motion synthesis.
Authors:Mingyu Dong, Chong Xia, Mingyuan Jia, Weichen Lyu, Long Xu, Zheng Zhu, Yueqi Duan
Abstract:
Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal for the advancement of Spatial Intelligence and Embodied AI. However, existing methods struggle to achieve practical deployment due to the insufficient integration of cross‑modal information, leaving them dependent on manual object prompting, reliant on auxiliary visual inputs, and restricted to overly simplistic scenes by training biases. To address these limitations, we propose ReplicateAnyScene, a framework capable of fully automated and zero‑shot transformation of casually captured videos into compositional 3D scenes. Specifically, our pipeline incorporates a five‑stage cascade to extract and structurally align generic priors from vision foundation models across textual, visual, and spatial dimensions, grounding them into structured 3D representations and ensuring semantic coherence and physical plausibility of the constructed scenes. To facilitate a more comprehensive evaluation of this task, we further introduce the C3DR benchmark to assess reconstruction quality from diverse aspects. Extensive experiments demonstrate the superiority of our method over existing baselines in generating high‑quality compositional 3D scenes.
Authors:Said Ohamouddou, Hanaa El Afia, Abdellatif El Afia, Raddouane Chiheb
Abstract:
Three‑dimensional (3D) point cloud analysis has become central to applications ranging from autonomous driving and robotics to forestry and ecological monitoring.
Although numerous deep learning methods have been proposed for point cloud understanding, including supervised backbones, self‑supervised pre‑training (SSL), and parameter‑efficient fine‑tuning (PEFT), their implementations are scattered across incompatible codebases with differing data pipelines, evaluation protocols, and configuration formats, making fair comparisons difficult.
We introduce \lib, a unified, extensible PyTorch library that integrates over 55 model configurations covering 29 supervised architectures, seven SSL pre‑training methods, and five PEFT strategies, all within a single registry‑based framework supporting classification, semantic segmentation, part segmentation, and few‑shot learning.
\lib provides standardised training runners, cross‑validation with stratified K‑fold splitting, automated LaTeX/CSV table generation, built‑in Friedman/Nemenyi statistical testing with critical‑difference diagrams for rigorous multi‑model comparison, and a comprehensive test suite with 2\,200+ automated tests validating every configuration end‑to‑end.
The code is available at https://github.com/said‑ohamouddou/LIDARLearn under the MIT licence.
Authors:Yuqi Chen, Xiaohan Zhang, Ahmad Arrabi, Waqas Sultani, Chen Chen, Safwan Wshah
Abstract:
Natural‑language Guided Cross‑view Geo‑localization (NGCG) aims to retrieve geo‑tagged satellite imagery using textual descriptions of ground scenes. While recent NGCG methods commonly rely on CLIP‑style dual‑encoder architectures, they often suffer from weak cross‑modal generalization and require complex architectural designs. In contrast, Multimodal Large Language Models (MLLMs) offer powerful semantic reasoning capabilities but are not directly optimized for retrieval tasks. In this work, we present a simple yet effective framework to adapt MLLMs for NGCG via parameter‑efficient finetuning. Our approach optimizes latent representations within the MLLM while preserving its pretrained multimodal knowledge, enabling strong cross‑modal alignment without redesigning model architectures. Through systematic analysis of diverse variables, from model backbone to feature aggregation, we provide practical and generalizable insights for leveraging MLLMs in NGCG. Our method achieves SOTA on GeoText‑1652 with a 12.2% improvement in Text‑to‑Image Recall@1 and secures top performance in 5 out of 12 subtasks on CVG‑Text, all while surpassing baselines with far fewer trainable parameters. These results position MLLMs as a robust foundation for semantic cross‑view retrieval and pave the way for MLLM‑based NGCG to be adopted as a scalable, powerful alternative to traditional dual‑encoder designs. Project page and code are available at https://yuqichen888.github.io/NGCG‑MLLMs‑web/.
Authors:Wei Zhang, Xinyu Chang, Xiao Li, Yiming Zhu, Xiaolin Hu
Abstract:
Adversarial examples present significant challenges to the security of Deep Neural Network (DNN) applications. Specifically, there are patch‑based and texture‑based attacks that are usually used to craft physical‑world adversarial examples, posing real threats to security‑critical applications such as person detection in surveillance and autonomous systems, because those attacks are physically realizable. Existing defense mechanisms face challenges in the adaptive attack setting, i.e., the attacks are specifically designed against them. In this paper, we propose Adversarial Spectrum Defense (ASD), a defense mechanism that leverages spectral decomposition via Discrete Wavelet Transform (DWT) to analyze adversarial patterns across multiple frequency scales. The multi‑resolution and localization capability of DWT enables ASD to capture both high‑frequency (fine‑grained) and low‑frequency (spatially pervasive) perturbations. By integrating this spectral analysis with the off‑the‑shelf Adversarial Training (AT) model, ASD provides a comprehensive defense strategy against both patch‑based and texture‑based adversarial attacks. Extensive experiments demonstrate that ASD+AT achieved state‑of‑the‑art (SOTA) performance against various attacks, outperforming the APs of previous defense methods by 21.73%, in the face of strong adaptive adversaries specifically designed against ASD. Code available at https://github.com/weiz0823/adv‑spectral‑defense .
Authors:Zeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang, Ruibin Yuan, Hubery Yin, Qifeng Chen, Chen Li, Jing Lyu, Wei Xue, Yike Guo
Abstract:
Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized models, leaving the development of a truly unified framework that can seamlessly integrate all three tasks underexplored. While some pioneering works have explored unifying audio understanding and generation, they often remain confined to specific domains. To address this, we introduce Audio‑Omni, the first end‑to‑end framework to unify generation and editing across general sound, music, and speech domains, with integrated multi‑modal understanding capabilities. Our architecture synergizes a frozen Multimodal Large Language Model for high‑level reasoning with a trainable Diffusion Transformer for high‑fidelity synthesis. To overcome the critical data scarcity in audio editing, we construct AudioEdit, a new large‑scale dataset comprising over one million meticulously curated editing pairs. Extensive experiments demonstrate that Audio‑Omni achieves state‑of‑the‑art performance across a suite of benchmarks, outperforming prior unified approaches while achieving performance on par with or superior to specialized expert models. Beyond its core capabilities, Audio‑Omni exhibits remarkable inherited capabilities, including knowledge‑augmented reasoning generation, in‑context generation, and zero‑shot cross‑lingual control for audio generation, highlighting a promising direction toward universal generative audio intelligence. The code, model, and dataset will be publicly released on https://zeyuet.github.io/Audio‑Omni.
Authors:Yifan Gao, Haoyue Li, Feng Yuan, Xin Gao, Weiran Huang, Xiaosong Wang
Abstract:
We present Camyla, a system for fully autonomous research within the scientific domain of medical image segmentation. Camyla transforms raw datasets into literature‑grounded research proposals, executable experiments, and complete manuscripts without human intervention. Autonomous experimentation over long horizons poses three interrelated challenges: search effort drifts toward unpromising directions, knowledge from earlier trials degrades as context accumulates, and recovery from failures collapses into repetitive incremental fixes. To address these challenges, the system combines three coupled mechanisms: Quality‑Weighted Branch Exploration for allocating effort across competing proposals, Layered Reflective Memory for retaining and compressing cross‑trial knowledge at multiple granularities, and Divergent Diagnostic Feedback for diversifying recovery after underperforming trials. The system is evaluated on CamylaBench, a contamination‑free benchmark of 31 datasets constructed exclusively from 2025 publications, under a strict zero‑intervention protocol across two independent runs within a total of 28 days on an 8‑GPU cluster. Across the two runs, Camyla generates more than 2,700 novel model implementations and 40 complete manuscripts, and surpasses the strongest per‑dataset baseline selected from 14 established architectures, including nnU‑Net, on 22 and 18 of 31 datasets under identical training budgets, respectively (union: 24/31). Senior human reviewers score the generated manuscripts at the T1/T2 boundary of contemporary medical imaging journals. Relative to automated baselines, Camyla outperforms AutoML and NAS systems on aggregate segmentation performance and exceeds six open‑ended research agents on both task completion and baseline‑surpassing frequency. These results suggest that domain‑scale autonomous research is achievable in medical image segmentation.
Authors:Bo Ma, Jinsong Wu, Weiqi Yan
Abstract:
Mamba selective state space models (SSMs) provide linear‑time sequence modeling but remain sensitive to selective‑scan chunk scheduling. We present COREY, a \emphconcept‑and‑feasibility runtime scheduler that maps fixed‑bin activation entropy to chunk size. We evaluate COREY in three tiers: a prototype cost model, real‑checkpoint kernel timing, and routed end‑to‑end ablations on modern GPUs.
At the kernel level, a calibrated rule, \(H_\mathrmref=\log K\), recovers the locally optimal chunk and matches a one‑time static oracle, yielding \(4.41×\) lower latency than an unoptimized baseline on a consumer GPU and \(3.90×\)‑‑\(4.04×\) lower latency on a data‑center accelerator. Routing this choice into a patched live scan kernel closes the engineering loop without improving end‑to‑end speed: in unified routed ablations, the best static chunk outperforms all entropy‑guided and proxy schedulers.
Sampled‑histogram COREY adds \(+4.6%\) overhead; a guarded fallback to Static‑512 reduces this to \(+1.3%\); and a lightweight sequence‑length‑keyed table further reduces it to \(+0.7%\). However, both remain slower than the static oracle because they retain scheduling cost. On an 80‑prompt LongBench subset, passive and routed inference are exactly output‑equivalent, with \(100%\) greedy‑token agreement and zero metric deltas.
A mixed‑regime study shows that a single sequence‑length rule matches the per‑regime chunk oracle for balanced serving. COREY is therefore validated as a quality‑preserving scheduling prototype, but current entropy statistics are not a robust throughput win over static chunk tuning on measured SSM checkpoint workloads. SourceCode: https://github.com/mabo1215/COREY_Transformer/.
Authors:Shiyin Jiang, Wei Long, Minghao Han, Zhenghao Chen, Ce Zhu, Shuhang Gu
Abstract:
The rapid growth of visual data under stringent storage and bandwidth constraints makes extremely low‑bitrate image compression increasingly important. While Vector Quantization (VQ) offers strong structural fidelity, existing methods lack a principled mechanism for joint rate‑distortion (RD) optimization due to the disconnect between representation learning and entropy modeling. We propose RDVQ, a unified framework that enables end‑to‑end RD optimization for VQ‑based compression via a differentiable relaxation of the codebook distribution, allowing the entropy loss to directly shape the latent prior. We further develop an autoregressive entropy model that supports accurate entropy modeling and test‑time rate control. Extensive experiments demonstrate that RDVQ achieves strong performance at extremely low bitrates with a lightweight architecture, attaining competitive or superior perceptual quality with significantly fewer parameters. Compared with RDEIC, RDVQ reduces bitrate by up to 75.71% on DISTS and 37.63% on LPIPS on DIV2K‑val. Beyond empirical gains, RDVQ introduces an entropy‑constrained formulation of VQ, highlighting the potential for a more unified view of image tokenization and compression. The code will be available at https://github.com/CVL‑UESTC/RDVQ.
Authors:Jia Li, Yu Zhang, Yin Chen, Zhenzhen Hu, Yong Li, Richang Hong, Shiguang Shan, Meng Wang
Abstract:
Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine‑grained muscular activations and coarse‑grained holistic affective states, respectively. Despite their inherent semantic correlation, existing studies predominantly focus on knowledge transfer from AUs to FEs, while bidirectional learning remains insufficiently explored. In practice, this challenge is further compounded by heterogeneous data conditions, where AU and FE datasets differ in annotation paradigms (frame‑level vs.\ clip‑level), label granularity, and data availability and diversity, hindering effective joint learning. To address these issues, we propose a Structured Semantic Mapping (SSM) framework for bidirectional AU‑‑FE learning under different data domains and heterogeneous supervision. SSM consists of three key components: (1) a shared visual backbone that learns unified facial representations from dynamic AU and FE videos; (2) semantic mediation via a Textual Semantic Prototype (TSP) module, which constructs structured semantic prototypes from fixed textual descriptions augmented with learnable context prompts, serving as supervision signals and cross‑task alignment anchors in a shared semantic space; and (3) a Dynamic Prior Mapping (DPM) module that incorporates prior knowledge derived from the Facial Action Coding System and learns a data‑driven association matrix in a high‑level feature space, enabling explicit and bidirectional knowledge transfer. Extensive experiments on popular AU detection and FE recognition benchmarks show that SSM achieves state‑of‑the‑art performance on both tasks simultaneously, and demonstrate that holistic expression semantics can in turn enhance fine‑grained AU learning even across heterogeneous datasets.
Authors:Hung-Ting Su, Ting-Jun Wang, Jia-Fong Yeh, Min Sun, Winston H. Hsu
Abstract:
Conventional Vision‑and‑Language Navigation (VLN) benchmarks assume instructions are feasible and the referenced target exists, leaving agents ill‑equipped to handle false‑premise goals. We introduce VLN‑NF, a benchmark with false‑premise instructions where the target is absent from the specified room and agents must navigate, gather evidence through in‑room exploration, and explicitly output NOT‑FOUND. VLN‑NF is constructed via a scalable pipeline that rewrites VLN instructions using an LLM and verifies target absence with a VLM, producing plausible yet factually incorrect goals. We further propose REV‑SPL to jointly evaluate room reaching, exploration coverage, and decision correctness. To address this challenge, we present ROAM, a two‑stage hybrid that combines supervised room‑level navigation with LLM/VLM‑driven in‑room exploration guided by a free‑space clearance prior. ROAM achieves the best REV‑SPL among compared methods, while baselines often under‑explore and terminate prematurely under unreliable instructions. VLN‑NF project page can be found at https://vln‑nf.github.io/.
Authors:Jingkai Wang, Jue Gong, Zheng Chen, Kai Liu, Jiatong Li, Yulun Zhang, Radu Timofte, Jiachen Tu, Yaokun Shi, Guoyi Xu, Yaoxin Jiang, Jiajia Liu, Yingsi Chen, Yijiao Liu, Hui Li, Yu Wang, Congchao Zhu, Alexandru-Gabriel Lefterache, Anamaria Radoi, Chuanyue Yan, Tao Lu, Yanduo Zhang, Kanghui Zhao, Jiaming Wang, Yuqi Li, WenBo Xiong, Yifei Chen, Xian Hu, Wei Deng, Daiguo Zhou, Sujith Roy, Claudia Jesuraj, Vikas B, Spoorthi LC, Nikhil Akalwadi, Ramesh Ashok Tabib, Uma Mudenagudi, Yuxuan Jiang, Chengxi Zeng, Tianhao Peng, Fan Zhang, David Bull Wei Zhou, Linfeng Li, Hongyu Huang, Hoyoung Lee, SangYun Oh, ChangYoung Jeong, Axi Niu, Jinyang Zhang, Zhenguo Wu, Senyan Qing, Jinqiu Sun, Yanning Zhang
Abstract:
This paper provides a review of the NTIRE 2026 challenge on real‑world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on generating natural and realistic outputs while maintaining identity consistency. Its goal is to advance state‑of‑the‑art solutions for perceptual quality and realism, without imposing constraints on computational resources or training data. Performance is evaluated using a weighted image quality assessment (IQA) score and employs the AdaFace model as an identity checker. The competition attracted 96 registrants, with 10 teams submitting valid models; ultimately, 9 teams achieved valid scores in the final ranking. This collaborative effort advances the performance of real‑world face restoration while offering an in‑depth overview of the latest trends in the field.
Authors:Zijia Lu, Jingru Yi, Jue Wang, Yuxiao Chen, Junwen Chen, Xinyu Li, Davide Modolo
Abstract:
Referring multi‑object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMOT approaches decompose object grounding and tracking into separated modules and exhibit limited performance due to the scarcity of training videos, ambiguous annotations, and restricted domains. In this work, we introduce STORM, an end‑to‑end MLLM that jointly performs grounding and tracking within a unified framework, eliminating external detectors and enabling coherent reasoning over appearance, motion, and language. To improve data efficiency, we propose a task‑composition learning (TCL) strategy that decomposes RMOT into image grounding and object tracking, allowing STORM to leverage data‑rich sub‑tasks and learn structured spatial‑‑temporal reasoning. We further construct STORM‑Bench, a new RMOT dataset with accurate trajectories and diverse, unambiguous referring expressions generated through a bottom‑up annotation pipeline. Extensive experiments show that STORM achieves state‑of‑the‑art performance on image grounding, single‑object tracking, and RMOT benchmarks, demonstrating strong generalization and robust spatial‑‑temporal grounding in complex real‑world scenarios. STORM‑Bench is released at https://github.com/amazon‑science/storm‑referring‑multi‑object‑grounding.
Authors:Lincoln Spencer, Song Wang, Chen Chen
Abstract:
Surgical phase segmentation is central to computer‑assisted surgery, yet robust models remain difficult to develop when labeled surgical videos are scarce. We study data‑efficient phase segmentation for manual small‑incision cataract surgery (SICS) through a controlled comparison of visual representations. To isolate representation quality, we pair each visual encoder with the same temporal model (MS‑TCN++) under identical training and evaluation settings on SICS‑155 (19 phases). We compare supervised encoders (ResNet‑50, I3D) against large self‑supervised foundation models (DINOv3, V‑JEPA2), and use a cached‑feature pipeline that decouples expensive visual encoding from lightweight temporal learning. Foundation‑model features improve segmentation performance in this setup, with DINOv3 ViT‑7B achieving the best overall results (83.4% accuracy, 87.0 edit score). We further examine cataract‑domain transfer using unlabeled videos and lightweight adaptation, and analyze when it helps or hurts. Overall, the study indicates strong transferability of modern vision foundation models to surgical workflow understanding and provides practical guidance for low‑label medical video settings. The project website is available at: https://sl2005.github.io/DataEfficient‑sics‑phase‑seg/
Authors:Chenhan Jiang, Yu Chen, Qingwen Zhang, Jifei Song, Songcen Xu, Dit-Yan Yeung, Jiankang Deng
Abstract:
The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large‑scale training data featuring diverse and precise camera trajectories. While real‑world captures are photorealistic, they are typically sparse and discrete. Conversely, synthetic data scales but suffers from a domain gap and often lacks realistic semantics. We introduce FreeScale, a novel framework that leverages the power of scene reconstruction to transform limited real‑world image sequences into a scalable source of high‑quality training data. Our key insight is that an imperfect reconstructed scene serves as a rich geometric proxy, but naively sampling from it amplifies artifacts. To this end, we propose a certainty‑aware free‑view sampling strategy identifying novel viewpoints that are both semantically meaningful and minimally affected by reconstruction errors. We demonstrate FreeScale's effectiveness by scaling up the training of feedforward NVS models, achieving a notable gain of 2.7 dB in PSNR on challenging out‑of‑distribution benchmarks. Furthermore, we show that the generated data can actively enhance per‑scene 3D Gaussian Splatting optimization, leading to consistent improvements across multiple datasets. Our work provides a practical and powerful data generation engine to overcome a fundamental bottleneck in 3D vision. Project page: https://mvp‑ai‑lab.github.io/FreeScale.
Authors:Haopeng Chen, Yihao Ai, Kabeen Kim, Robby T. Tan, Yixin Chen, Bo Wang
Abstract:
Low‑visibility scenarios, such as low‑light conditions, pose significant challenges to human pose estimation due to the scarcity of annotated low‑light datasets and the loss of visual information under poor illumination. Recent domain adaptation techniques attempt to utilize well‑lit labels by augmenting well‑lit images to mimic low‑light conditions. But handcrafted augmentations oversimplify noise patterns, while learning‑based methods often fail to preserve high‑frequency low‑light characteristics, producing unrealistic images that lead pose models to generalize poorly to real low‑light scenes. Moreover, recent pose estimators rely on image cues through image‑to‑keypoint cross‑attention, but these cues become unreliable under low‑light conditions. To address these issues, we propose Unsupervised Domain Adaptation for Pose Estimation (UDAPose), a novel framework that synthesizes low‑light images and dynamically fuses visual cues with pose priors for improved pose estimation. Specifically, our synthesis method incorporates a Direct‑Current‑based High‑Pass Filter (DHF) and a Low‑light Characteristics Injection Module (LCIM) to inject high‑frequency details from input low‑light images, overcoming rigidity or the detail loss in existing approaches. Furthermore, we introduce a Dynamic Control of Attention (DCA) module that adaptively balances image cues with learned pose priors in the Transformer architecture. Experiments show that UDAPose outperforms state‑of‑the‑art methods, with notable AP gains of 10.1 (56.4%) on the ExLPose‑test hard set (LL‑H) and 7.4 (31.4%) in cross‑dataset validation on EHPT‑XC. Code: https://github.com/Vision‑and‑Multimodal‑Intelligence‑Lab/UDAPose
Authors:Xinlei Guan, David Arosemena, Tejaswi Dhandu, Kuan Huang, Meng Xu, Miles Q. Li, Bingyu Shen, Ruiyang Qin, Umamaheswara Rao Tida, Boyang Li
Abstract:
The rapid growth of generative AI has introduced new challenges in content moderation and digital forensics. In particular, benign AI‑generated images can be paired with harmful or misleading text, creating difficult‑to‑detect misuse. This contextual misuse undermines the traditional moderation framework and complicates attribution, as synthetic images typically lack persistent metadata or device signatures. We introduce a steganography enabled attribution framework that embeds cryptographically signed identifiers into images at creation time and uses multimodal harmful content detection as a trigger for attribution verification. Our system evaluates five watermarking methods across spatial, frequency, and wavelet domains. It also integrates a CLIP‑based fusion model for multimodal harmful‑content detection. Experiments demonstrate that spread‑spectrum watermarking, especially in the wavelet domain, provides strong robustness under blur distortions, and our multimodal fusion detector achieves an AUC‑ROC of 0.99, enabling reliable cross‑modal attribution verification. These components form an end‑to‑end forensic pipeline that enables reliable tracing of harmful deployments of AI‑generated imagery, supporting accountability in modern synthetic media environments. Our code is available at GitHub: https://github.com/bli1/steganography
Authors:Song Jin, Juntian Zhang, Xun Zhang, Zeying Tian, Fei Jiang, Guojun Yin, Wei Lin, Yong Liu, Rui Yan
Abstract:
Recent advancements in Vision‑Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain remains constrained by benchmarks that rely on coarse‑grained categories, single‑view imagery, and inaccurate metadata. To bridge this gap, we introduce DiningBench, a hierarchical, multi‑view benchmark designed to evaluate VLMs across three levels of cognitive complexity: Fine‑Grained Classification, Nutrition Estimation, and Visual Question Answering. Unlike previous datasets, DiningBench comprises 3,021 distinct dishes with an average of 5.27 images per entry, incorporating fine‑grained "hard" negatives from identical menus and rigorous, verification‑based nutritional data. We conduct an extensive evaluation of 29 state‑of‑the‑art open‑source and proprietary models. Our experiments reveal that while current VLMs excel at general reasoning, they struggle significantly with fine‑grained visual discrimination and precise nutritional reasoning. Furthermore, we systematically investigate the impact of multi‑view inputs and Chain‑of‑Thought reasoning, identifying five primary failure modes. DiningBench serves as a challenging testbed to drive the next generation of food‑centric VLM research. All codes are released in https://github.com/meituan/DiningBench.
Authors:Di Wen, Zeyun Zhong, David Schneider, Manuel Zaremski, Linus Kunzmann, Yitian Shi, Ruiping Liu, Yufan Chen, Junwei Zheng, Jiahang Li, Jonas Hemmerich, Qiyi Tong, Patric Grauberger, Arash Ajoudani, Danda Pani Paudel, Sven Matthiesen, Barbara Deml, Jürgen Beyerer, Luc Van Gool, Rainer Stiefelhagen, Kunyu Peng
Abstract:
We introduce IMPACT, a synchronized five‑view RGB‑D dataset for deployment‑oriented industrial procedural understanding, built around real assembly and disassembly of a commercial angle grinder with professional‑grade tools. To our knowledge, IMPACT is the first real industrial assembly benchmark that jointly provides synchronized ego‑exo RGB‑D capture, decoupled bimanual annotation, compliance‑aware state tracking, and explicit anomaly‑‑recovery supervision within a single real industrial workflow. It comprises 112 trials from 13 participants totaling 39.5 hours, with multi‑route execution governed by a partial‑order prerequisite graph, a six‑category anomaly taxonomy, and operator cognitive load measured via NASA‑TLX. The annotation hierarchy links hand‑specific atomic actions to coarse procedural steps, component assembly states, and per‑hand compliance phases, with synchronized null spans across views to decouple perceptual limitations from algorithmic failure. Systematic baselines reveal fundamental limitations that remain invisible to single‑task benchmarks, particularly under realistic deployment conditions that involve incomplete observations, flexible execution paths, and corrective behavior. The full dataset, annotations, and evaluation code are available at https://github.com/Kratos‑Wen/IMPACT.
Authors:Yizheng Xie, Lennart Bastian, Congyue Deng, Thomas W. Mitchel, Maolin Gao, Daniel Cremers
Abstract:
Deep functional maps, leveraging learned feature extractors and spectral correspondence solvers, are fundamental to non‑rigid 3D shape matching. Based on an analysis of open‑source implementations, we find that standard functional map implementations solve k independent linear systems serially, which is a computational bottleneck at higher spectral resolution. We thus propose a vectorized reformulation that solves all systems in a single kernel call, achieving up to a 33x speedup while preserving the exact solution. Furthermore, we identify and document a previously unnoticed implementation divergence in the spatial gradient features of the mainstay DiffusionNet: two variants that parameterize distinct families of tangent‑plane transformations, and present experiments analyzing their respective behaviors across diverse benchmarks. We additionally revisit overlap prediction evaluation for partial‑to‑partial matching and show that balanced accuracy provides a useful complementary metric under varying overlap ratios. To share these advancements with the wider community, we present an open‑source codebase, DeepShapeMatchingKit, that incorporates these improvements and standardizes training, evaluation, and data pipelines for common deep shape matching methods. The codebase is available at: https://github.com/xieyizheng/DeepShapeMatchingKit
Authors:Alexandru Brateanu, Tingting Mu, Codruta Ancuti, Cosmin Ancuti
Abstract:
Low‑light image enhancement (LLIE) aims to restore natural visibility, color fidelity, and structural detail under severe illumination degradation. State‑of‑the‑art (SOTA) LLIE techniques often rely on large models and multi‑stage training, limiting practicality for edge deployment. Moreover, their dependence on a single color space introduces instability and visible exposure or color artifacts. To address these, we propose Multinex, an ultra‑lightweight structured framework that integrates multiple fine‑grained representations within a principled Retinex residual formulation. It decomposes an image into illumination and color prior stacks derived from distinct analytic representations, and learns to fuse these representations into luminance and reflectance adjustments required to correct exposure. By prioritizing enhancement over reconstruction and exploiting lightweight neural operations, Multinex significantly reduces computational cost, exemplified by its lightweight (45K parameters) and nano (0.7K parameters) versions. Extensive benchmarks show that all lightweight variants significantly outperform their corresponding lightweight SOTA models, and reach comparable performance to heavy models. Paper page available at https://albrateanu.github.io/multinex.
Authors:Jie Cai, Kangning Yang, Zhiyuan Li, Florin-Alexandru Vasluianu, Radu Timofte, Jinlong Li, Jinglin Shen, Zibo Meng, Junyan Cao, Lu Zhao, Pengwei Liu, Yuyi Zhang, Fengjun Guo, Jiagao Hu, Zepeng Wang, Fei Wang, Daiguo Zhou, Yi'ang Chen, Honghui Zhu, Mengru Yang, Yan Luo, Kui Jiang, Jin Guo, Jonghyuk Park, Jae-Young Sim, Wei Zhou, Hongyu Huang, Linfeng Li, Lindong Kong, Saiprasad Meesiyawar, Misbha Falak Khanpagadi, Nikhil Akalwadi, Ramesh Ashok Tabib, Uma Mudenagudi, Bilel Benjdira, Anas M. Ali, Wadii Boulila, Kosuke Shigematsu, Hiroto Shirono, Asuka Shin, Guoyi Xu, Yaoxin Jiang, Jiajia Liu, Yaokun Shi, Jiachen Tu, Shreeniketh Joshi, Jin-Hui Jiang, Yu-Fan Lin, Yu-Jou Hsiao, Chia-Ming Lee, Fu-En Yang, Yu-Chiang Frank Wang, Chih-Chung Hsu
Abstract:
In this paper, we review the NTIRE 2026 challenge on single‑image reflection removal (SIRR) in the wild. SIRR is a fundamental task in image restoration. Despite progress in academic research, most methods are tested on synthetic images or limited real‑world images, creating a gap in real‑world applications. In this challenge, we provide participants with the OpenRR‑5k dataset. This dataset requires participants to process real‑world images covering a range of reflection scenarios and intensities, aiming to generate clean images without reflections. The challenge attracted more than 100 registrations, with eleven of them participating in the final testing phase. The top‑ranked methods advanced the state‑of‑the‑art reflection removal performance and earned unanimous recognition from five experts in the field. The proposed OpenRR‑5k dataset is available at https://huggingface.co/datasets/qiuzhangTiTi/OpenRR‑5k, and the homepage of this challenge is at https://github.com/caijie0620/OpenRR‑5k.
Authors:Jingru Li, Wei Ren, Tianqing Zhu
Abstract:
Large Vision‑Language Models (LVLMs) rely on attention‑based retrieval of safety instructions to maintain alignment during generation. Existing attacks typically optimize image perturbations to maximize harmful output likelihood, but suffer from slow convergence due to gradient conflict between adversarial objectives and the model's safety‑retrieval mechanism. We propose Attention‑Guided Visual Jailbreaking, which circumvents rather than overpowers safety alignment by directly manipulating attention patterns. Our method introduces two simple auxiliary objectives: (1) suppressing attention to alignment‑relevant prefix tokens and (2) anchoring generation on adversarial image features. This simple yet effective push‑pull formulation reduces gradient conflict by 45% and achieves 94.4% attack success rate on Qwen‑VL (vs. 68.8% baseline) with 40% fewer iterations. At tighter perturbation budgets (ε=8/255), we maintain 59.0% ASR compared to 45.7% for standard methods. Mechanistic analysis reveals a failure mode we term safety blindness: successful attacks suppress system‑prompt attention by 80%, causing models to generate harmful content not by overriding safety rules, but by failing to retrieve them.
Authors:Peng Yuan, Bingyin Mei, Hui Zhang
Abstract:
Composed Image Retrieval (CIR) retrieves target images using a reference image paired with modification text. Despite rapid advances, all existing methods and datasets operate at the image level ‑‑ a single reference image plus modification text in, a single target image out ‑‑ while real e‑commerce users reason about products shown from multiple viewpoints. We term this mismatch View Incompleteness and formally define a new Multi‑View CIR task that generalizes standard CIR from image‑level to product‑level retrieval. To support this task, we construct FashionMV, the first large‑scale multi‑view fashion dataset for product‑level CIR, comprising 127K products, 472K multi‑view images, and over 220K CIR triplets, built through a fully automated pipeline leveraging large multimodal models. We further propose ProCIR (Product‑level Composed Image Retrieval), a modeling framework built upon a multimodal large language model that employs three complementary mechanisms ‑‑ two‑stage dialogue, caption‑based alignment, and chain‑of‑thought guidance ‑‑ together with an optional supervised fine‑tuning (SFT) stage that injects structured product knowledge prior to contrastive training. Systematic ablation across 16 configurations on three fashion benchmarks reveals that: (1) alignment is the single most critical mechanism; (2) the two‑stage dialogue architecture is a prerequisite for effective alignment; and (3) SFT and chain‑of‑thought serve as partially redundant knowledge injection paths. Our best 0.8B‑parameter model outperforms all baselines, including general‑purpose embedding models 10x its size. The dataset, model, and code are publicly available at https://github.com/yuandaxia2001/FashionMV.
Authors:Devdoot Chatterjee, Zakaria Laskar, C. V. Jawahar
Abstract:
We present a generalizable feed‑forward Gaussian splatting framework for human 3D reconstruction and real‑time animation that operates directly on multi‑view RGB images and their associated SMPL‑X poses. Unlike prior methods that rely on depth supervision, fixed input views, UV map, or repeated feed‑forward inference for each target view or pose, our approach predicts, in a canonical pose, a set of 3D Gaussian primitives associated with each SMPL‑X vertex. One Gaussian is regularized to remain close to the SMPL‑X surface, providing a strong geometric prior and stable correspondence to the parametric body model, while an additional small set of unconstrained Gaussians per vertex allows the representation to capture geometric structures that deviate from the parametric surface, such as clothing and hair. In contrast to recent approaches such as HumanRAM, which require repeated network inference to synthesize novel poses, our method produces an animatable human representation from a single forward pass; by explicitly associating Gaussian primitives with SMPL‑X vertices, the reconstructed model can be efficiently animated via linear blend skinning without further network evaluation. We evaluate our method on the THuman 2.1, AvatarReX and THuman 4.0 datasets, where it achieves reconstruction quality comparable to state‑of‑the‑art methods while uniquely supporting real‑time animation and interactive applications. Code and pre‑trained models are available at https://github.com/Devdoot57/HumanGS .
Authors:Meng'en Qin, Yu Song, Quanling Zhao, Xiaodong Yang, Yingtao Che, Xiaohui Yang
Abstract:
Learning multi‑scale representations is the common strategy to tackle object scale variation in dense prediction tasks. Although existing feature pyramid networks have greatly advanced visual recognition, inherent design defects inhibit them from capturing discriminative features and recognizing small objects. In this work, we propose Asymptotic Content‑Aware Pyramid Attention Network (A3‑FPN), to augment multi‑scale feature representation via the asymptotically disentangled framework and content‑aware attention modules. Specifically, A3‑FPN employs a horizontally‑spread column network that enables asymptotically global feature interaction and disentangles each level from all hierarchical representations. In feature fusion, it collects supplementary content from the adjacent level to generate position‑wise offsets and weights for context‑aware resampling, and learns deep context reweights to improve intra‑category similarity. In feature reassembly, it further strengthens intra‑scale discriminative feature learning and reassembles redundant features based on information content and spatial variation of feature maps. Extensive experiments on MS COCO, VisDrone2019‑DET and Cityscapes demonstrate that A3‑FPN can be easily integrated into state‑of‑the‑art CNN and Transformer‑based architectures, yielding remarkable performance gains. Notably, when paired with OneFormer and Swin‑L backbone, A3‑FPN achieves 49.6 mask AP on MS COCO and 85.6 mIoU on Cityscapes. Codes are available at https://github.com/mason‑ching/A3‑FPN.
Authors:Vasiliki Vasileiou, Panagiotis P. Filntisis, Petros Maragos, Kostas Daniilidis
Abstract:
Monocular head pose estimation is traditionally formulated as direct regression from a single image to an absolute pose. This paradigm forces the network to implicitly internalize a dataset‑specific canonical reference frame. In this work, we argue that predicting the relative rigid transformation between two observed head configurations is a fundamentally easier and more robust formulation. We introduce VGGT‑HPE, a relative head pose estimator built upon a general‑purpose geometry foundation model. Finetuned exclusively on synthetic facial renderings, our method sidesteps the need for an implicit anchor by reducing the problem to estimating a geometric displacement from an explicitly provided anchor with a known pose. As a practical benefit, the relative formulation also allows the anchor to be chosen at test time ‑ for instance, a near‑neutral frame or a temporally adjacent one ‑ so that the prediction difficulty can be controlled by the application. Despite zero real‑world training data, VGGT‑HPE achieves state‑of‑the‑art results on the BIWI benchmark, outperforming established absolute regression methods trained on mixed and real datasets. Through controlled easy‑ and hard‑pair benchmarks, we also systematically validate our core hypothesis: relative prediction is intrinsically more accurate than absolute regression, with the advantage scaling alongside the difficulty of the target pose. Project page and code: https://vasilikivas.github.io/VGGT‑HPE
Authors:Ruibin Li, Tao Yang, Fangzhou Ai, Tianhe Wu, Shilei Wen, Bingyue Peng, Lei Zhang
Abstract:
Streaming video generation (SVG) distills a pretrained bidirectional video diffusion model into an autoregressive model equipped with sliding window attention (SWA). However, SWA inevitably loses distant history during long video generation, and its computational overhead remains a critical challenge to real‑time deployment. In this work, we propose Hybrid Forcing, which jointly optimizes temporal information retention and computational efficiency through a hybrid attention design. First, we introduce lightweight linear temporal attention to preserve long‑range dependencies beyond the sliding window. In particular, we maintain a compact key‑value state to incrementally absorb evicted tokens, retaining temporal context with negligible memory and computational overhead. Second, we incorporate block‑sparse attention into the local sliding window to reduce redundant computation within short‑range modeling, reallocating computational capacity toward more critical dependencies. Finally, we introduce a decoupled distillation strategy tailored to the hybrid attention design. A few‑step initial distillation is performed under dense attention, then the distillation of our proposed linear temporal and block‑sparse attention is activated for streaming modeling, ensuring stable optimization. Extensive experiments on both short‑ and long‑form video generation benchmarks demonstrate that Hybrid Forcing consistently achieves state‑of‑the‑art performance. Notably, our model achieves real‑time, unbounded 832x480 video generation at 29.5 FPS on a single NVIDIA H100 GPU without quantization or model compression. The source code and trained models are available at https://github.com/leeruibin/hybrid‑forcing.
Authors:Yu Jiang, Hanwen Jiang, Ahmed Abdelkader, Wen-Sheng Chu, Brandon Y. Feng, Zhangyang Wang, Qixing Huang
Abstract:
With the emergence of 3D foundation models, there is growing interest in fine‑tuning them for downstream tasks, where LoRA is the dominant fine‑tuning paradigm. As 3D datasets exhibit distinct variations in texture, geometry, camera motion, and lighting, there are interesting fundamental questions: 1) Are there LoRA subspaces associated with each type of variation? 2) Are these subspaces disentangled (i.e., orthogonal to each other)? 3) How do we compute them effectively? This paper provides answers to all these questions. We introduce a robust approach that generates synthetic datasets with controlled variations, fine‑tunes a LoRA adapter on each dataset, and extracts a LoRA sub‑space associated with each type of variation. We show that these subspaces are approximately disentangled. Integrating them leads to a reduced LoRA subspace that enables efficient LoRA fine‑tuning with improved prediction accuracy for downstream tasks. In particular, we show that such a reduced LoRA subspace, despite being derived entirely from synthetic data, generalizes to real datasets. An ablation study validates the effectiveness of the choices in our approach.
Authors:Kunal Purkayastha, Ayan Banerjee, Josep Llados, Umapada Pal
Abstract:
In Document Understanding, the challenge of reconstructing damaged, occluded, or incomplete text remains a critical yet unexplored problem. Subsequent document understanding tasks can benefit from a document reconstruction process. In response, this paper presents a novel unified pipeline combining state‑of‑the‑art Optical Character Recognition (OCR), advanced image analysis, masked language modeling, and diffusion‑based models to restore and reconstruct text while preserving visual integrity. We create a synthetic dataset of 30,078 degraded document images that simulates diverse document degradation scenarios, setting a benchmark for restoration tasks. Our pipeline detects and recognizes text, identifies degradation with an occlusion detector, and uses an inpainting model for semantically coherent reconstruction. A diffusion‑based module seamlessly reintegrates text, matching font, size, and alignment. To evaluate restoration quality, we propose a Unified Context Similarity Metric (UCSM), incorporating edit, semantic, and length similarities with a contextual predictability measure that penalizes deviations when the correct text is contextually obvious. Our work advances document restoration, benefiting archival research and digital preservation while setting a new standard for text reconstruction. The OPRB dataset and code are available at \hrefhttps://huggingface.co/datasets/kpurkayastha/OPRBHugging Face and \hrefhttps://github.com/kunalpurkayastha/DocReviveGithub respectively.
Authors:Xunpei Sun, Wenwei Lin, Yi Chang, Gang Chen
Abstract:
Unsupervised optical flow methods typically lack reliable uncertainty estimation, limiting their robustness and interpretability. We propose U^2Flow, the first recurrent unsupervised framework that jointly estimates optical flow and per‑pixel uncertainty. The core innovation is a decoupled learning strategy that derives uncertainty supervision from augmentation consistency via a Laplace‑based maximum likelihood objective, enabling stable training without ground truth. The predicted uncertainty is further integrated into the network to guide adaptive flow refinement and dynamically modulate the regional smoothness loss. Furthermore, we introduce an uncertainty‑guided bidirectional flow fusion mechanism that enhances robustness in challenging regions. Extensive experiments on KITTI and Sintel demonstrate that U^2Flow achieves state‑of‑the‑art performance among unsupervised methods while producing highly reliable uncertainty maps, validating the effectiveness of our joint estimation paradigm. The code is available at https://github.com/sunzunyi/U2FLOW.
Authors:Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo
Abstract:
Vision‑‑language models (VLMs) have achieved impressive performance on complex multimodal reasoning tasks, yet they still fail on simple grounding skills such as object counting. Existing evaluations mostly assess only final outputs, offering limited insight into where these failures arise inside the model. In this work, we present an empirical study of VLM counting behavior through both behavioral and mechanistic analysis. We introduce COUNTINGTRICKS, a controlled evaluation suite of simple shape‑based counting cases designed to expose vulnerabilities under different patchification layouts and adversarial prompting conditions. Using attention analysis and component‑wise probing, we show that count‑relevant visual evidence is strongest in the modality projection stage but degrades substantially in later language layers, where models become more susceptible to text priors. Motivated by this finding, we further evaluate Modality Attention Share (MAS), a lightweight intervention that encourages a minimum budget of visual attention during answer generation. Our results suggest that counting failures in VLMs stem not only from visual perception limits, but also from the underuse of visual evidence during language‑stage reasoning. Code and dataset will be released at https://github.com/leduy99/‑CVPRW26‑Modality‑Attention‑Share.
Authors:Xu Liu, Guikun Chen, Wenguan Wang
Abstract:
Large language models (LLMs) suffer from hallucination and context forgetting. Prior studies suggest that attention drift is a primary cause of these problems, where LLMs' focus shifts towards newly generated tokens and away from the initial input context. To counteract this, we make use of a related, intrinsic characteristic of LLMs: attention sink ‑‑ the tendency to consistently allocate high attention to the very first token (i.e., <BOS>) of a sequence. Concretely, we propose an advanced context anchoring method, SinkTrack, which treats <BOS> as an information anchor and injects key contextual features (such as those derived from the input image or instruction) into its representation. As such, LLM remains anchored to the initial input context throughout the entire generation process. SinkTrack is training‑free, plug‑and‑play, and introduces negligible inference overhead. Experiments demonstrate that SinkTrack mitigates hallucination and context forgetting across both textual (e.g., +21.6% on SQuAD2.0 with Llama3.1‑8B‑Instruct) and multi‑modal (e.g., +22.8% on M3CoT with Qwen2.5‑VL‑7B‑Instruct) tasks. Its consistent gains across different architectures and scales underscore the robustness and generalizability. We also analyze its underlying working mechanism from the perspective of information delivery. Our source code is available at https://github.com/67L1/SinkTrack.
Authors:Shaobo Liu, Haobo Xiong, Kai Liu, Yuna Lin
Abstract:
Parameter‑efficient fine‑tuning of pre‑trained codecs is a promising direction in image compression for human and machine vision. While most existing works have primarily focused on tuning the feature structure within the encoder‑decoder backbones, the adaptation of the statistical semantics within the entropy model has received limited attention despite its function of predicting the probability distribution of latent features. Our analysis reveals that naive adapter insertion into the entropy model can lead to suboptimal outcomes, underscoring that the effectiveness of adapter‑based tuning depends critically on the coordination between adapter type and placement across the compression pipeline. Therefore, we introduce Structure‑Semantics Co‑Tuning (S2‑CoT), a novel framework that achieves this coordination via two specialized, synergistic adapters: the Structural Fidelity Adapter (SFA) and the Semantic Context Adapter (SCA). SFA is integrated into the encoder‑decoder to preserve high‑fidelity representations by dynamically fusing spatial and frequency information; meanwhile, the SCA adapts the entropy model to align with SFA‑tuned features by refining the channel context for more efficient statistical coding. Through joint optimization, S2‑CoT turns potential performance degradation into synergistic gains, achieving state‑of‑the‑art results across four diverse base codecs with only a small fraction of trainable parameters, closely matching full fine‑tuning performance. Code is available at https://github.com/Brock‑bit4/S2‑CoT.
Authors:Kening Wang, Di Wen, Yufan Chen, Ruiping Liu, Junwei Zheng, Jiale Wei, Kailun Yang, Rainer Stiefelhagen, Kunyu Peng
Abstract:
Automatic sleep staging is a multimodal learning problem involving heterogeneous physiological signals such as EEG and EOG, which often suffer from domain shifts across institutions, devices, and populations. In practice, these data are also affected by noisy annotations, yet label‑noise‑robust multi‑source domain generalization remains underexplored. We present the first benchmark for Noisy Labels in Multi‑Source Domain‑Generalized Sleep Staging (NL‑DGSS) and show that existing noisy‑label learning methods degrade substantially when domain shifts and label noise coexist. To address this challenge, we propose FF‑TRUST, a domain‑invariant multimodal sleep staging framework with Joint Time‑Frequency Early Learning Regularization (JTF‑ELR). By jointly exploiting temporal and spectral consistency together with confidence‑diversity regularization, FF‑TRUST improves robustness under noisy supervision. Experiments on five public datasets demonstrate consistent state‑of‑the‑art performance under diverse symmetric and asymmetric noise settings. The benchmark and code will be made publicly available at https://github.com/KNWang970918/FF‑TRUST.git.
Authors:Yuchen Zou, Huikai Shao, Lihuang Fang, Zhipeng Xiong, Dexing Zhong
Abstract:
Recently, synthetic palmprints have been increasingly used as substitutes for real data to train recognition models. To be effective, such synthetic data must reflect the diversity of real palmprints, including both style variation and geometric variation. However, existing palmprint generation methods mainly focus on style translation, while geometric variation is either ignored or approximated by simple handcrafted augmentations. In this work, we propose FlowPalm, an optical‑flow‑driven palmprint generation framework capable of simulating the complex non‑rigid deformations observed in real palms. Specifically, FlowPalm estimates optical flows between real palmprint pairs to capture the statistical patterns of geometric deformations. Building on these priors, we design a progressive sampling process that gradually introduces the geometric deformations during diffusion while maintaining identity consistency. Extensive experiments on six benchmark datasets demonstrate that FlowPalm significantly outperforms state‑of‑the‑art palmprint generation approaches in downstream recognition tasks. Project page: https://yuchenzou.github.io/FlowPalm/
Authors:Yiyu Liu, Shuo Ye, Chao Hao, Zitong Yu
Abstract:
Video Camouflaged Object Detection (VCOD) is currently constrained by the scarcity of challenging benchmarks and the limited robustness of models against erratic motion dynamics. Existing methods often struggle with Motion‑Induced Appearance Instability and Temporal Feature Misalignment caused by complex motion scenarios. To address the data bottleneck, we present YUV20K, a pixel‑level annoated complexity‑driven VCOD benchmark. Comprising 24,295 annotated frames across 91 scenes and 47 kinds of species, it specifically targets challenging scenarios like large‑displacement motion, camera motion and other 4 types scenarios. On the methodological front, we propose a novel framework featuring two key modules: Motion Feature Stabilization (MFS) and Trajectory‑Aware Alignment (TAA). The MFS module utilizes frame‑agnostic Semantic Basis Primitives to stablize features, while the TAA module leverages trajectory‑guided deformable sampling to ensure precise temporal alignment. Extensive experiments demonstrate that our method significantly outperforms state‑of‑the‑art competitors on existing datasets and establishes a new baseline on the challenging YUV20K. Notably, our framework exhibits superior cross‑domain generalization and robustness when confronting complex spatiotemporal scenarios. Our code and dataset will be available at https://github.com/K1NSA/YUV20K
Authors:Yimin Zhu, Lincoln Linlin Xu
Abstract:
Although hyperspectral image (HSI) classification is critical for supporting various environmental applications, it is a challenging task due to the spectral‑mixture effect, the spatial‑spectral heterogeneity and the difficulty to preserve class boundaries and details. This letter presents a novel unmixing‑guided spatial‑spectral Mamba with clustering tokens for improved HSI classification, with the following contributions. First, to disentangle the spectral mixture effect in HSI for improved pattern discovery, we design a novel spectral unmixing network that not only automatically learns endmembers and abundance maps from HSI but also accounts for endmember variabilities. Second, to generate Mamba token sequences, based on the clusters defined by abundance maps, we design an efficient Top‑K token selection strategy to adaptively sequence the tokens for improved representational capability. Third, to improve spatial‑spectral feature learning and detail preservation, based on the Top‑K token sequences, we design a novel unmixing‑guided spatial‑spectral Mamba module that greatly improves traditional Mamba models in terms of token learning and sequencing. Fourth, to learn simultaneously the endmember‑abundance patterns and classification labels, a multi‑task scheme is designed for model supervision, leading to a new unmixing‑classification framework that outputs not only accurate classification maps but also a comprehensive spectral‑library and abundance maps. Comparative experiments on four HSI datasets demonstrate that our model can greatly outperform the other state‑of‑the‑art approaches. Code is available at https://github.com/GSIL‑UCalgary/Unmixing_guided_Mamba.git
Authors:Guillermo Auza Banegas, Diego Calvimontes Vera, Sergio Castro Sandoval, Natalia Condori Peredo, Edwin Salcedo
Abstract:
Robust license plate recognition in unconstrained environments remains a significant challenge, particularly in underrepresented regions with limited data availability and unique visual characteristics, such as Bolivia. Recognition accuracy in real‑world conditions is often degraded by factors such as illumination changes and viewpoint distortion. To address these challenges, we introduce BLPR, a novel deep learning‑based License Plate Detection and Recognition (LPDR) framework specifically designed for Bolivian license plates. The proposed system follows a two‑stage pipeline where a YOLO‑based detector is pretrained on synthetic data generated in Blender to simulate extreme perspectives and lighting conditions, and subsequently fine‑tuned on street‑level data collected in La Paz, Bolivia. Detected plates are geometrically rectified and passed to a character recognition model. To improve robustness under ambiguous scenarios, a lightweight vision‑language model (Gemma3 4B) is selectively triggered as a confidence‑based fallback mechanism. The proposed framework further leverages synthetic‑to‑real domain adaptation to improve robustness under diverse real‑world conditions. We also introduce the first publicly available Bolivian LPDR dataset, enabling evaluation under diverse viewpoint and illumination conditions. The system achieves a character‑level recognition accuracy of 89.6% on real‑world data, demonstrating its effectiveness for deployment in challenging urban environments. Our project is publicly available at https://github.com/EdwinTSalcedo/BLPR.
Authors:Bochu Ding, Brinnae Bent, Augustus Wendell
Abstract:
Text‑to‑image (T2I) models, and their encoded biases, increasingly shape the visual media the public encounters. While researchers have produced a rich body of work on bias measurement, auditing, and mitigation in T2I systems, those methods largely target technical stakeholders, leaving a gap in public legibility. We introduce GLEaN (Generative Likeness Evaluation at N‑Scale), a portrait‑based explainability pipeline designed to make T2I model biases visually understandable to a broad audience. GLEaN comprises three stages: automated large‑scale image generation from identity prompts, facial landmark‑based filtering and spatial alignment, and median‑pixel composition that distills a model's central tendency into a single representative portrait. The resulting composites require no statistical background to interpret; a viewer can see, at a glance, who a model 'imagines' when prompted with 'a doctor' versus a 'felon.' We demonstrate GLEaN on Stable Diffusion XL across 40 social and occupational identity prompts, producing composites that reproduce documented biases and surface new associations between skin tone and predicted emotion. We find in a between‑subjects user study (N = 291) that GLEaN portraits communicate biases as effectively as conventional data tables, but require significantly less viewing time. Because the method relies solely on generated outputs, it can also be replicated on any black‑box and closed‑weight systems without access to model internals. GLEaN offers a scalable, model‑agnostic approach to bias explainability, purpose‑built for public comprehension, and is publicly available at https://github.com/cultureiolab/GLEaN.
Authors:Chaoyi Zhou, Run Wang, Feng Luo, Mert D. Pesé, Zhiwen Fan, Yiqi Zhong, Siyu Huang
Abstract:
Recent advances in vision foundation models have revolutionized geometry reconstruction and semantic understanding. Yet, most of the existing approaches treat these capabilities in isolation, leading to redundant pipelines and compounded errors. This paper introduces FF3R, a fully annotation‑free feed‑forward framework that unifies geometric and semantic reasoning from unconstrained multi‑view image sequences. Unlike previous methods, FF3R does not require camera poses, depth maps, or semantic labels, relying solely on rendering supervision for RGB and feature maps, establishing a scalable paradigm for unified 3D reasoning. In addition, we address two critical challenges in feedforward feature reconstruction pipelines, namely global semantic inconsistency and local structural inconsistency, through two key innovations: (i) a Token‑wise Fusion Module that enriches geometry tokens with semantic context via cross‑attention, and (ii) a Semantic‑Geometry Mutual Boosting mechanism combining geometry‑guided feature warping for global consistency with semantic‑aware voxelization for local coherence. Extensive experiments on ScanNet and DL3DV‑10K demonstrate FF3R's superior performance in novel‑view synthesis, open‑vocabulary semantic segmentation, and depth estimation, with strong generalization to in‑the‑wild scenarios, paving the way for embodied intelligence systems that demand both spatial and semantic understanding.
Authors:Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew, Jan van Gemert, Xucong Zhang
Abstract:
Multi‑person social interactions are inherently built on coherence and relationships among all individuals within the group, making multi‑person localization and body pose estimation essential to understanding these social dynamics. One promising approach is 2D‑to‑3D pose lifting which provides a 3D human pose consisting of rich spatial details by building on the significant advances in 2D pose estimation. However, the existing 2D‑to‑3D pose lifting methods often neglect inter‑person relationships or cannot handle varying group sizes, limiting their effectiveness in multi‑person settings. We propose MuPPet, a novel multi‑person 2D‑to‑3D pose lifting framework that explicitly models inter‑person correlations. To leverage these inter‑person dependencies, our approach introduces Person Encoding to structure individual representations, Permutation Augmentation to enhance training diversity, and Dynamic Multi‑Person Attention to adaptively model correlations between individuals. Extensive experiments on group interaction datasets demonstrate MuPPet significantly outperforms state‑of‑the‑art single‑ and multi‑person 2D‑to‑3D pose lifting methods, and improves robustness in occlusion scenarios. Our findings highlight the importance of modeling inter‑person correlations, paving the way for accurate and socially‑aware 3D pose estimation. Our code is available at: https://github.com/Thomas‑Markhorst/MuPPet
Authors:Justin Li, Daniel Ding, Asmita Yuki Pritha, Aryana Hou, Xin Wang, Shu Hu
Abstract:
Automated diagnosis from chest CT has improved considerably with deep learning, but models trained on skewed datasets tend to perform unevenly across patient demographics. However, the situation is worse than simple demographic bias. In clinical data, class imbalance and group underrepresentation often coincide, creating compound failure modes that neither standard rebalancing nor fairness corrections can fix alone. We introduce a two‑level objective that targets both axes of this problem. Logit‑adjusted cross‑entropy loss operates at the sample level, shifting decision margins by class frequency with provable consistency guarantees. Conditional Value at Risk aggregation operates at the group level, directing optimization pressure toward whichever demographic group currently has the higher loss. We evaluate on the Fair Disease Diagnosis benchmark using a 3D ResNet‑18 pretrained on Kinetics‑400, classifying CT volumes into Adenocarcinoma, Squamous Cell Carcinoma, COVID‑19, and Normal groups with patient sex annotations. The training set illustrates the compound problem concretely: squamous cell carcinoma has 84 samples total, 5 of them female. The combined loss reaches a gender‑averaged macro F1 of 0.8403 with a fairness gap of 0.0239, a 13.3% improvement in score and 78% reduction in demographic disparity over the baseline. Ablations show that each component alone falls short. The code is publicly available at https://github.com/Purdue‑M2/Fair‑Disease‑Diagnosis.
Authors:Taminul Islam, Abdellah Lakhssassi, Toqi Tahamid Sarker, Mohamed Embaby, Khaled R Ahmed, Amer AbuGhazaleh
Abstract:
Quantifying exhaled CO2 from free‑roaming cattle is both a direct indicator of rumen metabolic state and a prerequisite for farm‑scale carbon accounting, yet no existing system can deliver continuous, spatially resolved measurements without physical confinement or contact. We present TRACE (Thermal Recognition Attentive‑Framework for CO2 Emissions from Livestock), the first unified framework to jointly address per‑frame CO2 plume segmentation and clip‑level emission flux classification from mid‑wave infrared (MWIR) thermal video. TRACE contributes three domain‑specific advances: a Thermal Gas‑Aware Attention (TGAA) encoder that incorporates per‑pixel gas intensity as a spatial supervisory signal to direct self‑attention toward high‑emission regions at each encoder stage; an Attention‑based Temporal Fusion (ATF) module that captures breath‑cycle dynamics through structured cross‑frame attention for sequence‑level flux classification; and a four‑stage progressive training curriculum that couples both objectives while preventing gradient interference. Benchmarked against fifteen state‑of‑the‑art models on the CO2 Farm Thermal Gas Dataset, TRACE achieves an mIoU of 0.998 and the best result on every segmentation and classification metric simultaneously, outperforming domain‑specific gas segmenters with several times more parameters and surpassing all baselines in flux classification. Ablation studies confirm that each component is individually essential: gas‑conditioned attention alone determines precise plume boundary localization, and temporal reasoning is indispensable for flux‑level discrimination. TRACE establishes a practical path toward non‑invasive, continuous, per‑animal CO2 monitoring from overhead thermal cameras at commercial scale. Codes are available at https://github.com/taminulislam/trace.
Authors:Shuang Li, Jian Gao, Chulhong Kim, Seongwook Choi, Qian Chen, Yibing Wang, Shuang Wu, Yu Zhang, Tingting Huang, Yucheng Zhou, Boxin Yao, Yao Yao, Changhui Li
Abstract:
Three‑dimensional (3D) handheld photoacoustic tomography typically relies on bulky and expensive external positioning sensors to correct motion artifacts, which severely limits its clinical flexibility and accessibility. To address this challenge, we present PA‑SFM, a tracker‑free framework that leverages exclusively single‑modality photoacoustic data for both sensor pose recovery and high‑fidelity 3D reconstruction via differentiable acoustic radiation modeling. Unlike traditional structure‑from‑motion (SFM) methods based on visual features, PA‑SFM integrates the acoustic wave equation into a differentiable programming pipeline. By leveraging a high‑performance, GPU‑accelerated acoustic radiation kernel, the framework simultaneously optimizes the 3D photoacoustic source distribution and the sensor array pose via gradient descent. To ensure robust convergence in freehand scenarios, we introduce a coarse‑to‑fine optimization strategy that incorporates geometric consistency checks and rigid‑body constraints to eliminate motion outliers. We validated the proposed method through both numerical simulations and in‑vivo rat experiments. The results demonstrate that PA‑SFM achieves sub‑millimeter positioning accuracy and restores high‑resolution 3D vascular structures comparable to ground‑truth benchmarks, offering a low‑cost, software‑defined solution for clinical freehand photoacoustic imaging. The source code is publicly available at \hrefhttps://github.com/JaegerCQ/PA‑SFMhttps://github.com/JaegerCQ/PA‑SFM.
Authors:Tianfu Wang, Leilei Ding, Ziyang Tao, Yi Zhan, Zhiyuan Ma, Wei Wu, Yuxuan Lei, Yuan Feng, Junyang Wang, Yin Wu, Yizhao Xu, Hongyuan Zhu, Qi Liu, Nicholas Jing Yuan, Yanyong Zhang, Hui Xiong
Abstract:
High‑fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel‑based models often lack precise control, while code‑based synthesis limits intuitive flexibility. To bridge this gap, we introduce EvoDiagram, an agentic framework that generates object‑level editable diagrams via an intermediate canvas schema. EvoDiagram employs a coordinated multi‑agent system to decouple semantic intent from rendering logic, resolving conflicts across heterogeneous design layers. Additionally, we propose a design knowledge evolution mechanism that distills execution traces into a hierarchical memory of domain guidelines, enabling agents to retrieve context‑aware expertise adaptively. We further release CanvasBench, a benchmark consisting of both data and metrics for canvas‑based diagramming. Extensive experiments demonstrate that EvoDiagram exhibits excellent performance and balance against baselines in generating editable, structurally consistent, and aesthetically coherent diagrams. Our code is available at https://github.com/AuraX‑AI/EvoDiagram.
Authors:Shukang Yin, Sirui Zhao, Hanchao Wang, Baozhi Jia, Xianquan Wang, Chaoyou Fu, Enhong Chen
Abstract:
Token pruning has emerged as a mainstream approach for developing efficient Video Large Language Models (Video LLMs). This work revisits and advances the two predominant token‑pruning paradigms: attention‑based selection and similarity‑based clustering. Our study reveals two critical limitations in existing methods: (1) conventional top‑k selection strategies fail to fully account for the attention distribution, which is often spatially multi‑modal and long‑tailed in magnitude; and (2) direct similarity‑based clustering frequently generates fragmented clusters, resulting in distorted representations after pooling. To address these bottlenecks, we propose Tango, a novel framework designed to optimize the utilization of visual signals. Tango integrates a diversity‑driven strategy to enhance attention‑based token selection, and introduces Spatio‑temporal Rotary Position Embedding (ST‑RoPE) to preserve geometric structure via locality priors. Comprehensive experiments across various Video LLMs and video understanding benchmarks demonstrate the effectiveness and generalizability of our approach. Notably, when retaining only 10% of the video tokens, Tango preserves 98.9% of the original performance on LLaVA‑OV while delivering a 1.88× inference speedup.
Authors:Zibin Geng, Xuefeng Jiang, Jia Li, Zheng Li, Tian Wen, Lvhua Wu, Sheng Sun, Yuwei Wang, Min Liu
Abstract:
Prompt learning is a parameter‑efficient approach for vision‑language models, yet its robustness under label noise is less investigated. Visual content contains richer and more reliable semantic information, which remains more robust under label noise. However, the prompt itself is highly susceptible to label noise. Motivated by this intuition, we propose VisPrompt, a lightweight and robust vision‑guided prompt learning framework for noisy‑label settings. Specifically, we exploit a cross‑modal attention mechanism to reversely inject visual semantics into prompt representations. This enables the prompt tokens to selectively aggregate visual information relevant to the current sample, thereby improving robustness by anchoring prompt learning to stable instance‑level visual evidence and reducing the influence of noisy supervision. To address the instability caused by using the same way of injecting visual information for all samples, despite differences in the quality of their visual cues, we further introduce a lightweight conditional modulation mechanism to adaptively control the strength of visual information injection, which strikes a more robust balance between text‑side semantic priors and image‑side instance evidence. The proposed framework effectively suppresses the noise‑induced disturbances, reduce instability in prompt updates, and alleviate memorization of mislabeled samples. VisPrompt significantly improves robustness while keeping the pretrained VLM backbone frozen and introducing only a small amount of additional trainable parameters. Extensive experiments under synthetic and real‑world label noise demonstrate that VisPrompt generally outperforms existing baselines on seven benchmark datasets and achieves stronger robustness. Our code is publicly available at https://github.com/gezbww/Vis_Prompt.
Authors:Guanyu Zhou, Yida Yin, Wenhao Chai, Shengbang Tong, Xingyu Fu, Zhuang Liu
Abstract:
Vision‑language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contributing factor is that natural image datasets provide limited supervision for low‑level visual skills. This motivates a practical question: can targeted synthetic supervision, generated from only a task keyword such as Depth Order, address these weaknesses? To investigate this question, we introduce VisionFoundry, a task‑aware synthetic data generation pipeline that takes only the task name as input and uses large language models (LLMs) to generate questions, answers, and text‑to‑image (T2I) prompts, then synthesizes images with T2I models and verifies consistency with a proprietary VLM, requiring no reference images or human annotation. Using VisionFoundry, we construct VisionFoundry‑10K, a synthetic visual question answering (VQA) dataset containing 10k image‑question‑answer triples spanning 10 tasks. Models trained on VisionFoundry‑10K achieve substantial improvements on visual perception benchmarks: +7% on MMVP and +10% on CV‑Bench‑3D, while preserving broader capabilities and showing favorable scaling behavior as data size increases. Our results suggest that limited task‑targeted supervision is an important contributor to this bottleneck and that synthetic supervision is a promising path toward more systematic training for VLMs.
Authors:Wenyi Xiao, Xinchi Xu, Leilei Gan
Abstract:
Large Vision Language Models (LVLMs) achieve strong multimodal reasoning but frequently exhibit hallucinations and incorrect responses with high certainty, which hinders their usage in high‑stakes domains. Existing verbalized confidence calibration methods, largely developed for text‑only LLMs, typically optimize a single holistic confidence score using binary answer‑level correctness. This design is mismatched to LVLMs: an incorrect prediction may arise from perceptual failures or from reasoning errors given correct perception, and a single confidence conflates these sources while visual uncertainty is often dominated by language priors. To address these issues, we propose VL‑Calibration, a reinforcement learning framework that explicitly decouples confidence into visual and reasoning confidence. To supervise visual confidence without ground‑truth perception labels, we introduce an intrinsic visual certainty estimation that combines (i) visual grounding measured by KL‑divergence under image perturbations and (ii) internal certainty measured by token entropy. We further propose token‑level advantage reweighting to focus optimization on tokens based on visual certainty, suppressing ungrounded hallucinations while preserving valid perception. Experiments on thirteen benchmarks show that VL‑Calibration effectively improves calibration while boosting visual reasoning accuracy, and it generalizes to out‑of‑distribution benchmarks across model scales and architectures.
Authors:Stefan Andreas Baumann, Jannik Wiese, Tommaso Martorella, Mahdi M. Kalayeh, Björn Ommer
Abstract:
Accurately anticipating how complex, diverse scenes will evolve requires models that represent uncertainty, simulate along extended interaction chains, and efficiently explore many plausible futures. Yet most existing approaches rely on dense video or latent‑space prediction, expending substantial capacity on dense appearance rather than on the underlying sparse trajectories of points in the scene. This makes large‑scale exploration of future hypotheses costly and limits performance when long‑horizon, multi‑modal motion is essential. We address this by formulating the prediction of open‑set future scene dynamics as step‑wise inference over sparse point trajectories. Our autoregressive diffusion model advances these trajectories through short, locally predictable transitions, explicitly modeling the growth of uncertainty over time. This dynamics‑centric representation enables fast rollout of thousands of diverse futures from a single image, optionally guided by initial constraints on motion, while maintaining physical plausibility and long‑range coherence. We further introduce OWM, a benchmark for open‑set motion prediction based on diverse in‑the‑wild videos, to evaluate accuracy and variability of predicted trajectory distributions under real‑world uncertainty. Our method matches or surpasses dense simulators in predictive accuracy while achieving orders‑of‑magnitude higher sampling speed, making open‑set future prediction both scalable and practical. Project page: http://compvis.github.io/myriad.
Authors:Shunkai Zhou, Zike Yan, Fei Xue, Dong Wu, Yuchen Deng, Hongbin Zha
Abstract:
We present Online3R, a new sequential reconstruction framework that is capable of adapting to new scenes through online learning, effectively resolving inconsistency issues. Specifically, we introduce a set of learnable lightweight visual prompts into a pretrained, frozen geometry foundation model to capture the knowledge of new environments while preserving the fundamental capability of the foundation model for geometry prediction. To solve the problems of missing groundtruth and the requirement of high efficiency when updating these visual prompts at test time, we introduce a local‑global self‑supervised learning strategy by enforcing the local and global consistency constraints on predictions. The local consistency constraints are conducted on intermediate and previously local fused results, enabling the model to be trained with high‑quality pseudo groundtruth signals; the global consistency constraints are operated on sparse keyframes spanning long distances rather than per frame, allowing the model to learn from a consistent prediction over a long trajectory in an efficient way. Our experiments demonstrate that Online3R outperforms previous state‑of‑the‑art methods on various benchmarks. Project page: https://shunkaizhou.github.io/online3r‑1.0/
Authors:Zhengxian Yang, Shengqi Wang, Shi Pan, Hongshuai Li, Haoxiang Wang, Lin Li, Guanjun Li, Zhengqi Wen, Borong Lin, Jianhua Tao, Tao Yu
Abstract:
Fully immersive experiences that tightly integrate 6‑DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved through computer‑generated content, constructing them directly from real‑world captured videos remains largely unexplored. We introduce Immersive Volumetric Videos, a new volumetric media format designed to provide large 6‑DoF interaction spaces, audiovisual feedback, and high‑resolution, high‑frame‑rate dynamic content. To support IVV construction, we present ImViD, a multi‑view, multi‑modal dataset built upon a space‑oriented capture philosophy. Our custom capture rig enables synchronized multi‑view video‑audio acquisition during motion, facilitating efficient capture of complex indoor and outdoor scenes with rich foreground‑‑background interactions and challenging dynamics. The dataset provides 5K‑resolution videos at 60 FPS with durations of 1‑5 minutes, offering richer spatial, temporal, and multimodal coverage than existing benchmarks. Leveraging this dataset, we develop a dynamic light field reconstruction framework built upon a Gaussian‑based spatio‑temporal representation, incorporating flow‑guided sparse initialization, joint camera temporal calibration, and multi‑term spatio‑temporal supervision for robust and accurate modeling of complex motion. We further propose, to our knowledge, the first method for sound field reconstruction from such multi‑view audiovisual data. Together, these components form a unified pipeline for immersive volumetric video production. Extensive benchmarks and immersive VR experiments demonstrate that our pipeline generates high‑quality, temporally stable audiovisual volumetric content with large 6‑DoF interaction spaces. This work provides both a foundational definition and a practical construction methodology for immersive volumetric videos.
Authors:Siyuan Zhou, Hejun Wang, Hu Cheng, Jinxi Li, Dongsheng Wang, Junwei Jiang, Yixiao Jin, Jiayue Huang, Shiwei Mao, Shangjia Liu, Yafei Yang, Hongkang Song, Shenxing Wei, Zihui Zhang, Peng Huang, Shijie Liu, Zhengli Hao, Hao Li, Yitian Li, Wenqi Zhou, Zhihan Zhao, Zongqi He, Hongtao Wen, Shouwang Huang, Peng Yun, Bowen Cheng, Pok Kazaf Fu, Wai Kit Lai, Jiahao Chen, Kaiyuan Wang, Zhixuan Sun, Ziqi Li, Haochen Hu, Di Zhang, Chun Ho Yuen, Bing Wang, Zhihua Wang, Chuhang Zou, Bo Yang
Abstract:
We present PhysInOne, a large‑scale synthetic dataset addressing the critical scarcity of physically‑grounded training data for AI systems. Unlike existing datasets limited to merely hundreds or thousands of examples, PhysInOne provides 2 million videos across 153,810 dynamic 3D scenes, covering 71 basic physical phenomena in mechanics, optics, fluid dynamics, and magnetism. Distinct from previous works, our scenes feature multiobject interactions against complex backgrounds, with comprehensive ground‑truth annotations including 3D geometry, semantics, dynamic motion, physical properties, and text descriptions. We demonstrate PhysInOne's efficacy across four emerging applications: physics‑aware video generation, long‑/short‑term future frame prediction, physical property estimation, and motion transfer. Experiments show that fine‑tuning foundation models on PhysInOne significantly enhances physical plausibility, while also exposing critical gaps in modeling complex physical dynamics and estimating intrinsic properties. As the largest dataset of its kind, orders of magnitude beyond prior works, PhysInOne establishes a new benchmark for advancing physics‑grounded world models in generation, simulation, and embodied AI.
Authors:Qingwen Zhang, Xiaomeng Zhu, Chenhan Jiang, Patric Jensfelt
Abstract:
Reliable 3D dynamic perception requires models that can anticipate motion beyond predefined categories, yet progress is hindered by the scarcity of dense, high‑quality motion annotations. While self‑supervision on unlabeled real data offers a path forward, empirical evidence suggests that scaling unlabeled data fails to close the performance gap due to noisy proxy signals. In this paper, we propose a shift in paradigm: learning robust real‑world motion priors entirely from scalable simulation. We introduce SynFlow, a data generation pipeline that generates large‑scale synthetic dataset specifically designed for LiDAR scene flow. Unlike prior works that prioritize sensor‑specific realism, SynFlow employs a motion‑oriented strategy to synthesize diverse kinematic patterns across 4,000 sequences (~940k frames), termed SynFlow‑4k. This represents a 34x scale‑up in annotated volume over existing real‑world benchmarks. Our experiments demonstrate that SynFlow‑4k provides a highly domain‑invariant motion prior. In a zero‑shot regime, models trained exclusively on our synthetic data generalize across multiple real‑world benchmarks, rivaling in‑domain supervised baselines on nuScenes and outperforming state‑of‑the‑art methods on TruckScenes by 31.8%. Furthermore, SynFlow‑4k serves as a label‑efficient foundation: fine‑tuning with only 5% of real‑world labels surpasses models trained from scratch on the full available budget. We open‑source the pipeline and dataset to facilitate research in generalizable 3D motion estimation. More detail can be found at https://kin‑zhang.github.io/SynFlow.
Authors:Shipeng Zhu, Ang Chen, Na Nie, Pengfei Fang, Min-Ling Zhang, Hui Xue
Abstract:
Ancient inscriptions, as repositories of cultural memory, have suffered from centuries of environmental and human‑induced degradation. Restoring their intertwined visual and textual integrity poses one of the most demanding challenges in digital heritage preservation. However, existing AI‑based approaches often rely on rigid pipelines, struggling to generalize across such complex and heterogeneous real‑world degradations. Inspired by the skill‑coordinated workflow of human epigraphers, we propose EpiAgent, an agent‑centric system that formulates inscription restoration as a hierarchical planning problem. Following an Observe‑Conceive‑Execute‑Reevaluate paradigm, an LLM‑based central planner orchestrates collaboration among multimodal analysis, historical experience, specialized restoration tools, and iterative self‑refinement. This agent‑centric coordination enables a flexible and adaptive restoration process beyond conventional single‑pass methods. Across real‑world degraded inscriptions, EpiAgent achieves superior restoration quality and stronger generalization compared to existing methods. Our work marks an important step toward expert‑level agent‑driven restoration of cultural heritage. The code is available at https://github.com/blackprotoss/EpiAgent.
Authors:Zengbin Wang, Feng Xiong, Liang Lin, Xuecai Hu, Yong Wang, Yanlin Wang, Man Zhang, Xiangxiang Chu
Abstract:
Reinforcement learning with verifiable rewards (RLVR) has significantly advanced the reasoning ability of vision‑language models (VLMs). However, the inherent text‑dominated nature of VLMs often leads to insufficient visual faithfulness, characterized by sparse attention activation to visual tokens. More importantly, our empirical analysis reveals that temporal visual forgetting along reasoning steps exacerbates this deficiency. To bridge this gap, we propose Visually‑Guided Policy Optimization (VGPO), a novel framework to reinforce visual focus during policy optimization. Specifically, VGPO initially introduces a Visual Attention Compensation mechanism that leverages visual similarity to localize and amplify visual cues, while progressively elevating visual expectations in later steps to counteract visual forgetting. Building on this mechanism, we implement a dual‑grained advantage re‑weighting strategy: the intra‑trajectory level highlights tokens exhibiting relatively high visual activation, while the inter‑trajectory level prioritizes trajectories demonstrating superior visual accumulation. Extensive experiments demonstrate that VGPO achieves better visual activation and superior performance in mathematical multimodal reasoning and visual‑dependent tasks. The code has been released at https://github.com/wzb‑bupt/VGPO.
Authors:Narges Rashvand, Shanle Yao, Armin Danesh Pazho, Babak Rahimi Ardabili, Hamed Tabkhi
Abstract:
Pose‑based Video Anomaly Detection (VAD) has gained significant attention for its privacy‑preserving nature and robustness to environmental variations. However, traditional frame‑level evaluations treat video as a collection of isolated frames, fundamentally misaligned with how anomalies manifest and are acted upon in the real world. In operational surveillance systems, what matters is not the flagging of individual frames, but the reliable detection, localization, and reporting of a coherent anomalous event, a contiguous temporal episode with an identifiable onset and duration. Frame‑level metrics are blind to this distinction, and as a result, they systematically overestimate model performance for any deployment that requires actionable, event‑level alerts. In this work, we propose a shift toward an event‑centric perspective in VAD. We first audit widely used VAD benchmarks, including SHT[19], CHAD[6], NWPUC[4], and HuVAD[25], to characterize their event structure. We then introduce two strategies for temporal event localization: a score‑refinement pipeline with hierarchical Gaussian smoothing and adaptive binarization, and an end‑to‑end Dual‑Branch Model that directly generates event‑level detections. Finally, we establish the first event‑based evaluation standard for VAD by adapting Temporal Action Localization metrics, including tIoU‑based event matching and multi‑threshold F1 evaluation. Our results quantify a substantial performance gap: while all SoTA models achieve frame‑level AUC‑ROC exceeding 52% on the NWPUC[4], their event‑level localization precision falls below 10% even at a minimal tIoU=0.2, with an average event‑level F1 of only 0.11 across all thresholds. The code base for this work is available at https://github.com/TeCSAR‑UNCC/EventCentric‑VAD.
Authors:Yuze Su, Hongsong Wang, Jie Gui, Liang Wang
Abstract:
Reconstructing photorealistic and topology‑aware human avatars from monocular videos remains a significant challenge in the fields of computer vision and graphics. While existing 3D human avatar modeling approaches can effectively capture body motion, they often fail to accurately model fine details such as hand movements and facial expressions. To address this, we propose Structure‑aware Fine‑grained Gaussian Splatting (SFGS), a novel method for reconstructing expressive and coherent full‑body 3D human avatars from a monocular video sequence. The SFGS use both spatial‑only triplane and time‑aware hexplane to capture dynamic features across consecutive frames. A structure‑aware gaussian module is designed to capture pose‑dependent details in a spatially coherent manner and improve pose and texture expression. To better model hand deformations, we also propose a residual refinement module based on fine‑grained hand reconstruction. Our method requires only a single‑stage training and outperforms state‑of‑the‑art baselines in both quantitative and qualitative evaluations, generating high‑fidelity avatars with natural motion and fine details. The code is on Github: https://github.com/Su245811YZ/SFGS
Authors:Agniva Sengupta, Dilara Kuş, Jianning Li, Stefan Zachow
Abstract:
We solve the problem of determining the pose of known shapes in \mathbbR^3 from their unoccluded silhouettes. The pose is determined up to global optimality using a simple yet under‑explored property of the area‑of‑silhouette: its continuity w.r.t trajectories in the rotation space. The proposed method utilises pre‑computed silhouette‑signatures, modelled as a response surface of the area‑of‑silhouettes. Querying this silhouette‑signature response surface for pose estimation leads to a strong branching of the rotation search space, making resolution‑guided candidate search feasible. Additionally, we utilise the aspect ratio of 2D ellipses fitted to projected silhouettes as an auxiliary global shape signature to accelerate the pose search. This combined strategy forms the first method to efficiently estimate globally optimal pose from just the silhouettes, without being guided by correspondences, for any shape, irrespective of its convexity and genus. We validate our method on synthetic and real examples, demonstrating significantly improved accuracy against comparable approaches.
Code, data, and supplementary in: https://agnivsen.github.io/pose‑from‑silhouette/
Authors:Nazir Nayal, Christopher Wewer, Jan Eric Lenssen
Abstract:
Diffusion models and their variations, such as rectified flows, generate diverse and high‑quality images, but they are still hindered by slow iterative sampling caused by the highly curved generative paths they learn. An important cause of high curvature, as shown by previous work, is independence between the source distribution (standard Gaussian) and the data distribution. In this work, we tackle this limitation by two complementary contributions. First, we attempt to break away from the standard Gaussian assumption by introducing κ\texttt‑FC, a general formulation that conditions the source distribution on an arbitrary signal κ that aligns it better with the data distribution. Then, we present MixFlow, a simple but effective training strategy that reduces the generative path curvatures and considerably improves sampling efficiency. MixFlow trains a flow model on linear mixtures of a fixed unconditional distribution and a κ\texttt‑FC‑based distribution. This simple mixture improves the alignment between the source and data, provides better generation quality with less required sampling steps, and accelerates the training convergence considerably. On average, our training procedure improves the generation quality by 12% in FID compared to standard rectified flow and 7% compared to previous baselines under a fixed sampling budget. Code available at: \hrefhttps://github.com/NazirNayal8/MixFlowhttps://github.com/NazirNayal8/MixFlow
Authors:Le-Van Thai, Tien Dat Nguyen, Hoai Nhan Pham, Lan Anh Dinh Thi, Duy-Dong Nguyen, Ngoc Lam Quang Bui
Abstract:
Semi‑supervised semantic segmentation in computational pathology remains challenging due to scarce pixel‑level annotations and unreliable pseudo‑label supervision. We propose UniSemAlign, a dual‑modal semantic alignment framework that enhances visual segmentation by injecting explicit class‑level structure into pixel‑wise learning. Built upon a pathology‑pretrained Transformer encoder, UniSemAlign introduces complementary prototype‑level and text‑level alignment branches in a shared embedding space, providing structured guidance that reduces class ambiguity and stabilizes pseudo‑label refinement. The aligned representations are fused with visual predictions to generate more reliable supervision for unlabeled histopathology images. The framework is trained end‑to‑end with supervised segmentation, cross‑view consistency, and cross‑modal alignment objectives. Extensive experiments on the GlaS and CRAG datasets demonstrate that UniSemAlign substantially outperforms recent semi‑supervised baselines under limited supervision, achieving Dice improvements of up to 2.6% on GlaS and 8.6% on CRAG with only 10% labeled data, and strong improvements at 20% supervision. Code is available at: https://github.com/thailevann/UniSemAlign
Authors:Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu, Wen-Kai Kuo, Jun-Wei Hsieh
Abstract:
Lightweight face recognition is increasingly important for deployment on edge and mobile devices, where strict constraints on latency, memory, and energy consumption must be met alongside reliable accuracy. Although recent hybrid CNN‑Transformer architectures have advanced global context modeling, striking an effective balance between recognition performance and computational efficiency remains an open challenge. In this work, we present FaceLiVTv2, an improved version of our FaceLiVT hybrid architecture designed for efficient global‑‑local feature interaction in mobile face recognition. At its core is Lite MHLA, a lightweight global token interaction module that replaces the original multi‑layer attention design with multi‑head linear token projections and affine rescale transformations, reducing redundancy while preserving representational diversity across heads. We further integrate Lite MHLA into a unified RepMix block that coordinates local and global feature interactions and adopts global depthwise convolution for adaptive spatial aggregation in the embedding stage. Under our experimental setup, results on LFW, CA‑LFW, CP‑LFW, CFP‑FP, AgeDB‑30, and IJB show that FaceLiVTv2 consistently improves the accuracy‑efficiency trade‑off over existing lightweight methods. Notably, FaceLiVTv2 reduces mobile inference latency by 22% relative to FaceLiVTv1, achieves speedups of up to 30.8% over GhostFaceNets on mobile devices, and delivers 20‑41% latency improvements over EdgeFace and KANFace across platforms while maintaining higher recognition accuracy. These results demonstrate that FaceLiVTv2 offers a practical and deployable solution for real‑time face recognition. Code is available at https://github.com/novendrastywn/FaceLiVT.
Authors:Mengxin Fu, Yuezun Li
Abstract:
Diffusion models are known for generating high‑quality images, causing serious security concerns. To combat this, most efforts rely on deep neural networks (e.g., CNNs and Transformers), while largely overlooking the potential of traditional machine learning models. In this paper, we freshly investigate such alternatives and proposes a novel Dynamic Assembly Forest model (DAF) to detect diffusion‑generated images. Built upon the deep forest paradigm, DAF addresses the inherent limitations in feature learning and scalable training, making it an effective diffusion‑generated image detector. Compared to existing DNN‑based methods, DAF has significantly fewer parameters, much lower computational cost, and can be deployed without GPUs, while achieving competitive performance under standard evaluation protocols. These results highlight the strong potential of the proposed method as a practical substitute for heavyweight DNN models in resource‑constrained scenarios. Our code and models are available at https://github.com/OUC‑VAS/DAF.
Authors:Yutong Zhang, Jiaxin Chen, Honglin Chen, Kaiqi Zheng, Shengcai Liao, Hanwen Zhong, Weixin Li, Yunhong Wang
Abstract:
Memory‑efficient transfer learning (METL) approaches have recently achieved promising performance in adapting pre‑trained models to downstream tasks. They avoid applying gradient backpropagation in large backbones, thus significantly reducing the number of trainable parameters and high memory consumption during fine‑tuning. However, since they typically employ a lightweight and learnable side network, these methods inevitably introduce additional memory and time overhead during inference, which contradicts the ultimate goal of efficient transfer learning. To address the above issue, we propose a novel approach dubbed Masked Dual Path Distillation (MDPD) to accelerate inference while retaining parameter and memory efficiency in fine‑tuning with fading side networks. Specifically, MDPD develops a framework that enhances the performance by mutually distilling the frozen backbones and learnable side networks in fine‑tuning, and discard the side network during inference without sacrificing accuracy. Moreover, we design a novel feature‑based knowledge distillation method for the encoder structure with multiple layers. Extensive experiments on distinct backbones across vision/language‑only and vision‑and‑language tasks demonstrate that our method not only accelerates inference by at least 25.2% while keeping parameter and memory consumption comparable, but also remarkably promotes the accuracy compared to SOTA approaches. The source code is available at https://github.com/Zhang‑VKk/MDPD.
Authors:Lishen Qu, Yao Liu, Jie Liang, Hui Zeng, Wen Dai, Guanyi Qin, Ya-nan Guan, Shihao Zhou, Jufeng Yang, Lei Zhang, Radu Timofte, Xiyuan Yuan, Wanjie Sun, Shihang Li, Bo Zhang, Bin Chen, Jiannan Lin, Yuxu Chen, Qinquan Gao, Tong Tong, Song Gao, Jiacong Tang, Tao Hu, Xiaowen Ma, Qingsen Yan, Sunhan Xu, Juan Wang, Xinyu Sun, Lei Qi, He Xu, Jiachen Tu, Guoyi Xu, Yaoxin Jiang, Jiajia Liu, Yaokun Shi
Abstract:
This paper presents NTIRE 2026, the 3rd Restore Any Image Model (RAIM) challenge on multi‑exposure image fusion in dynamic scenes. We introduce a benchmark that targets a practical yet difficult HDR imaging setting, where exposure bracketing must be fused under scene motion, illumination variation, and handheld camera jitter. The challenge data contains 100 training sequences with 7 exposure levels and 100 test sequences with 5 exposure levels, reflecting real‑world scenarios that frequently cause misalignment and ghosting artefacts. We evaluate submissions with a leaderboard score derived from PSNR, SSIM, and LPIPS, while also considering perceptual quality, efficiency, and reproducibility during the final review. This track attracted 114 participating teams and received 987 submissions. The winning methods significantly improved the ability to remove artifacts from multi‑exposure fusion and recover fine details. The dataset and the code of each team can be found at the repository: https://github.com/qulishen/RAIM‑HDR.
Authors:Junxi Wang, Te Sun, Jiayi Zhu, Junxian Li, Haowen Xu, Zichen Wen, Xuming Hu, Zhiyu Li, Linfeng Zhang
Abstract:
Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memory overhead, leading to high costs in both storage and computation. To address this issue, we propose StreamMeCo, an efficient Stream Agent Memory Compression framework. Specifically, based on the connectivity of the memory graph, StreamMeCo introduces edge‑free minmax sampling for the isolated nodes and an edge‑aware weight pruning for connected nodes, evicting the redundant memory nodes while maintaining the accuracy. In addition, we introduce a time‑decay memory retrieval mechanism to further eliminate the performance degradation caused by memory compression. Extensive experiments on three challenging benchmark datasets (M3‑Bench‑robot, M3‑Bench‑web and Video‑MME‑Long) demonstrate that under 70% memory graph compression, StreamMeCo achieves a 1.87 speedup in memory retrieval while delivering an average accuracy improvement of 1.0%. Our code is available at https://github.com/Celina‑love‑sweet/StreamMeCo.
Authors:Zhiyu Zhou, Peilin Liu, Ruoxuan Zhang, Luyang Zhang, Cheng Zhang, Hongxia Xie, Wen-Huang Cheng
Abstract:
Small object‑centric spatial understanding in indoor videos remains a significant challenge for multimodal large language models (MLLMs), despite its practical value for object search and assistive applications. Although existing benchmarks have advanced video spatial intelligence, embodied reasoning, and diagnostic perception, no existing benchmark directly evaluates whether a model can localize a target object in video and express its position with sufficient precision for downstream use. In this work, we introduce PinpointQA, the first dataset and benchmark for small object‑centric spatial understanding in indoor videos. Built from ScanNet++ and ScanNet200, PinpointQA comprises 1,024 scenes and 10,094 QA pairs organized into four progressively challenging tasks: Target Presence Verification (TPV), Nearest Reference Identification (NRI), Fine‑Grained Spatial Description (FSD), and Structured Spatial Prediction (SSP). The dataset is built from intermediate spatial representations, with QA pairs generated automatically and further refined through quality control. Experiments on representative MLLMs reveal a consistent capability gap along the progressive chain, with SSP remaining particularly difficult. Supervised fine‑tuning on PinpointQA yields substantial gains, especially on the harder tasks, demonstrating that PinpointQA serves as both a diagnostic benchmark and an effective training dataset. The dataset and project page are available at https://rainchowz.github.io/PinpointQA.
Authors:Langzhe Gu, Hung-Jui Huang, Mohamad Qadri, Michael Kaess, Wenzhen Yuan
Abstract:
Accurate object geometry estimation is essential for many downstream tasks, including robotic manipulation and physical interaction. Although vision is the dominant modality for shape perception, it becomes unreliable under occlusions or challenging lighting conditions. In such scenarios, tactile sensing provides direct geometric information through physical contact. However, reconstructing global 3D geometry from sparse local touches alone is fundamentally underconstrained. We present TouchAnything, a framework that leverages a pretrained large‑scale 2D vision diffusion model as a semantic and geometric prior for 3D reconstruction from sparse tactile measurements. Unlike prior work that trains category‑specific reconstruction networks or learns diffusion models directly from tactile data, we transfer the geometric knowledge encoded in pretrained visual diffusion models to the tactile domain. Given sparse contact constraints and a coarse class‑level description of the object, we formulate reconstruction as an optimization problem that enforces tactile consistency while guiding solutions toward shapes consistent with the diffusion prior. Our method reconstructs accurate geometries from only a few touches, outperforms existing baselines, and enables open‑world 3D reconstruction of previously unseen object instances. Our project page is https://grange007.github.io/touchanything .
Authors:Zengyi Yang, Yu Liu, Juan Cheng, Zhiqin Zhu, Yafei Zhang, Huafeng Li
Abstract:
Infrared‑visible image fusion aims to integrate complementary information for robust visual understanding, but existing fusion methods struggle with simultaneously adapting to multiple downstream tasks. To address this issue, we propose a Closed‑Loop Dynamic Network (CLDyN) that can adaptively respond to the semantic requirements of diverse downstream tasks for task‑customized image fusion. Specifically, CLDyN introduces a closed‑loop optimization mechanism that establishes a semantic transmission chain to achieve explicit feedback from downstream tasks to the fusion network through a Requirement‑driven Semantic Compensation (RSC) module. The RSC module leverages a Basis Vector Bank (BVB) and an Architecture‑Adaptive Semantic Injection (A2SI) block to customize the network architecture according to task requirements, thereby enabling task‑specific semantic compensation and allowing the fusion network to actively adapt to diverse tasks without retraining. To promote semantic compensation, a reward‑penalty strategy is introduced to reward or penalize the RSC module based on task performance variations. Experiments on the M3FD, FMB, and VT5000 datasets demonstrate that CLDyN not only maintains high fusion quality but also exhibits strong multi‑task adaptability. The code is available at https://github.com/YR0211/CLDyN.
Authors:Ao Li, Yonggen Ling, Yiyang Lin, Yuji Wang, Yong Deng, Yansong Tang
Abstract:
Accurate 3D human keypoints localization is a critical technology enabling robots to achieve natural and safe physical interaction with users. Conventional 3D human keypoints estimation methods primarily focus on the whole‑body reconstruction quality relative to the root joint. However, in practical human‑robot interaction (HRI) scenarios, robots are more concerned with the precise metric‑scale spatial localization of task‑relevant body parts under the egocentric camera 3D coordinate. We propose TAIHRI, the first Vision‑Language Model (VLM) tailored for close‑range HRI perception, capable of understanding users' motion commands and directing the robot's attention to the most task‑relevant keypoints. By quantizing 3D keypoints into a finite interaction space, TAIHRI precisely localize the 3D spatial coordinates of critical body parts by 2D keypoint reasoning via next token prediction, and seamlessly adapt to downstream tasks such as natural language control or global space human mesh recovery. Experiments on egocentric interaction benchmarks demonstrate that TAIHRI achieves superior estimation accuracy for task‑critical body parts. We believe TAIHRI opens new research avenues in the field of embodied human‑robot interaction. Code is available at: https://github.com/Tencent/TAIHRI.
Authors:Yuanting Fan, Jun Liu, Bin-Bin Gao, Xiaochen Chen, Yuhuan Lin, Zhewei Dai, Jiawei Zhan, Chengjie Wang
Abstract:
Existing defect/anomaly generation methods often rely on few‑shot learning, which overfits to specific defect categories due to the lack of large‑scale paired defect editing data. This issue is aggravated by substantial variations in defect scale and morphology, resulting in limited generalization, degraded realism, and category consistency. We address these challenges by introducing UDG, a large‑scale dataset of 300K normal‑abnormal‑mask‑caption quadruplets spanning diverse domains, and by presenting UniDG, a universal defect generation foundation model that supports both reference‑based defect generation and text instruction‑based defect editing without per‑category fine‑tuning. UniDG performs Defect‑Context Editing via adaptive defect cropping and structured diptych input format, and fuses reference and target conditions through MM‑DiT multimodal attention. A two‑stage training strategy, Diversity‑SFT followed by Consistency‑RFT, further improves diversity while enhancing realism and reference consistency. Extensive experiments on MVTec‑AD and VisA show that UniDG outperforms prior few‑shot anomaly generation and image insertion/editing baselines in synthesis quality and downstream single‑ and multi‑class anomaly detection/localization. Code will be available at https://github.com/RetoFan233/UniDG.
Authors:Xinyu Zhang, Zurong Mai, Qingmei Li, Zjin Liao, Yibin Wen, Yuhang Chen, Xiaoya Fan, Chan Tsz Ho, Bi Tianyuan, Haoyuan Liang, Ruifeng Su, Zihao Qian, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
Abstract:
While multimodal large language models (MLLMs) have made significant strides in natural image understanding, their ability to perceive and reason over hyperspectral image (HSI) remains underexplored, which is a vital modality in remote sensing. The high dimensionality and intricate spectral‑spatial properties of HSI pose unique challenges for models primarily trained on RGB data.To address this gap, we introduce Hyperspectral Multimodal Benchmark (HM‑Bench), the first benchmark designed specifically to evaluate MLLMs in HSI understanding. We curate a large‑scale dataset of 19,337 question‑answer pairs across 13 task categories, ranging from basic perception to spectral reasoning. Given that existing MLLMs are not equipped to process raw hyperspectral cubes natively, we propose a dual‑modality evaluation framework that transforms HSI data into two complementary representations: PCA‑based composite images and structured textual reports. This approach facilitates a systematic comparison of different representation for model performance. Extensive evaluations on 18 representative MLLMs reveal significant difficulties in handling complex spatial‑spectral reasoning tasks. Furthermore, our results demonstrate that visual inputs generally outperform textual inputs, highlighting the importance of grounding in spectral‑spatial evidence for effective HSI understanding. Dataset and appendix can be accessed at https://github.com/HuoRiLi‑Yu/HM‑Bench.
Authors:Jinqi Luo, Jinyu Yang, Tal Neiman, Lei Fan, Bing Yin, Son Tran, Mubarak Shah, René Vidal
Abstract:
Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prompt engineering, response classification, or finetuning to improve MLLM safety. Nevertheless, such approaches are often ineffective against evolving malicious patterns, may require rerunning the query, or demand heavy computational resources. Steering the activations of a frozen model at inference time has recently emerged as a flexible and effective solution. However, existing steering methods for MLLMs typically handle only a narrow set of safety‑related concepts or struggle to adjust specific concepts without affecting others. To address these challenges, we introduce Dictionary‑Aligned Concept Control (DACO), a framework that utilizes a curated concept dictionary and a Sparse Autoencoder (SAE) to provide granular control over MLLM activations. First, we curate a dictionary of 15,000 multimodal concepts by retrieving over 400,000 caption‑image stimuli and summarizing their activations into concept directions. We name the dataset DACO‑400K. Second, we show that the curated dictionary can be used to intervene activations via sparse coding. Third, we propose a new steering approach that uses our dictionary to initialize the training of an SAE and automatically annotate the semantics of the SAE atoms for safeguarding MLLMs. Experiments on multiple MLLMs (e.g., QwenVL, LLaVA, InternVL) across safety benchmarks (e.g., MM‑SafetyBench, JailBreakV) show that DACO significantly improves MLLM safety while maintaining general‑purpose capabilities.
Authors:Zewei Zhou, Jiajun Zou, Jiajia Zhang, Ao Yang, Ruichao He, Haozheng Zhou, Ao Liu, Jiawei Liu, Leilei Jin, Shan Shen, Daying Sun
Abstract:
Graph neural networks (GNNs) are increasingly applied to physical design tasks such as congestion prediction and wirelength estimation, yet progress is hindered by inconsistent circuit representations and the absence of controlled evaluation protocols. We present R2G (RTL‑to‑GDSII), a multi‑view circuit‑graph benchmark suite that standardizes five stage‑aware views with information parity (every view encodes the same attribute set, differing only in where features attach) over 30 open‑source IP cores (up to 10^6 nodes/edges). R2G provides an end‑to‑end DEF‑to‑graph pipeline spanning synthesis, placement, and routing stages, together with loaders, unified splits, domain metrics, and reproducible baselines. By decoupling representation choice from model choice, R2G isolates a confound that prior EDA and graph‑ML benchmarks leave uncontrolled. In systematic studies with GINE, GAT, and ResGatedGCN, we find: (i) view choice dominates model choice, with Test R^2 varying by more than 0.3 across representations for a fixed GNN; (ii) node‑centric views generalize best across both placement and routing; and (iii) decoder‑head depth (3‑‑4 layers) is the primary accuracy driver, turning divergent training into near‑perfect predictions (R^2>0.99). Code and datasets are available at https://github.com/ShenShan123/R2G.
Authors:Hyunwoo Kim, Itai Lang, Hadar Averbuch-Elor, Silvia Sellán, Rana Hanocka
Abstract:
We propose MeshOn, a method that finds physically and semantically realistic compositions of two input meshes. Given an accessory, a base mesh with a user‑defined target region, and optional text strings for both meshes, MeshOn uses a multi‑step optimization framework to realistically fit the meshes onto each other while preventing intersections. We initialize the shapes' rigid configuration via a structured alignment scheme using Vision‑to‑Language Models, which we then optimize using a combination of attractive geometric losses, and a physics‑inspired barrier loss that prevents surface intersections. We then obtain a final deformation of the object, assisted by a diffusion prior. Our method successfully fits accessories of various materials over a breadth of target regions, and is designed to fit directly into existing digital artist workflows. We demonstrate the robustness and accuracy of our pipeline by comparing it with generative approaches and traditional registration algorithms.
Authors:Yi-Hua Huang, Zi-Xin Zou, Yuting He, Chirui Chang, Cheng-Feng Pu, Ziyi Yang, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
Abstract:
Animatable 3D assets, defined as geometry equipped with an articulated skeleton and skinning weights, are fundamental to interactive graphics, embodied agents, and animation production. While recent 3D generative models can synthesize visually plausible shapes from images, the results are typically static. Obtaining usable rigs via post‑hoc auto‑rigging is brittle and often produces skeletons that are topologically inconsistent with the generated geometry. We present AniGen, a unified framework that directly generates animate‑ready 3D assets conditioned on a single image. Our key insight is to represent shape, skeleton, and skinning as mutually consistent S^3 Fields (Shape, Skeleton, Skin) defined over a shared spatial domain. To enable the robust learning of these fields, we introduce two technical innovations: (i) a confidence‑decaying skeleton field that explicitly handles the geometric ambiguity of bone prediction at Voronoi boundaries, and (ii) a dual skin feature field that decouples skinning weights from specific joint counts, allowing a fixed‑architecture network to predict rigs of arbitrary complexity. Built upon a two‑stage flow‑matching pipeline, AniGen first synthesizes a sparse structural scaffold and then generates dense geometry and articulation in a structured latent space. Extensive experiments demonstrate that AniGen substantially outperforms state‑of‑the‑art sequential baselines in rig validity and animation quality, generalizing effectively to in‑the‑wild images across diverse categories including animals, humanoids, and machinery. Homepage: https://yihua7.github.io/AniGen‑web/
Authors:Lucas Wojcik, Eduardo A. F. Machoski, Eduil Nascimento, Rayson Laroca, David Menotti
Abstract:
Modern Automatic License Plate Recognition (ALPR) systems achieve outstanding performance in controlled, well‑defined scenarios. However, large‑scale real‑world usage remains challenging due to low‑quality imaging devices, compression artifacts, and suboptimal camera installation. Identifying illegible license plates (LPs) has recently become feasible through a dedicated benchmark; however, its impact has been limited by its small size and annotation errors. In this work, we expand the original benchmark to over three times the size with two extra capture days, revise its annotations and introduce novel labels. LP‑level annotations include bounding boxes, text, and legibility level, while vehicle‑level annotations comprise make, model, type, and color. Image‑level annotations feature camera identity, capture conditions (e.g., rain and faulty cameras), acquisition time, and day ID. We present a novel training procedure featuring an Exponential Moving Average‑based loss function and a refined learning rate scheduler, addressing common mistakes in testing. These improvements enable a baseline model to achieve an 89.5% F1‑score on the test set, considerably surpassing the previous state of the art. We further introduce a novel protocol to explicitly addresses camera contamination between training and evaluation splits, where results show a small impact. Dataset and code are publicly available at https://github.com/lmlwojcik/LPLCv2‑Dataset.
Authors:Ruixiang Jiang, Changwen Chen
Abstract:
Interpretation is essential to deciphering the language of art: audiences communicate with artists by recovering meaning from visual artifacts. However, current Generative Art (GenArt) evaluators remain fixated on surface‑level image quality or literal prompt adherence, failing to assess the deeper symbolic or abstract meaning intended by the creator. We address this gap by formalizing a Peircean computational semiotic theory that models Human‑GenArt Interaction (HGI) as cascaded semiosis. This framework reveals that artistic meaning is conveyed through three modes ‑ iconic, symbolic, and indexical ‑ yet existing evaluators operate heavily within the iconic mode, remaining structurally blind to the latter two. To overcome this structural blindness, we propose SemJudge. This evaluator explicitly assesses symbolic and indexical meaning in HGI via a Hierarchical Semiosis Graph (HSG) that reconstructs the meaning‑making process from prompt to generated artifact. Extensive quantitative experiments show that SemJudge aligns more closely with human judgments than prior evaluators on an interpretation‑intensive fine‑art benchmark. User studies further demonstrate that SemJudge produces deeper, more insightful artistic interpretations, thereby paving the way for GenArt to move beyond the generation of "pretty" images toward a medium capable of expressing complex human experience. Project page: https://github.com/songrise/SemJudge.
Authors:Weikai Huang, Jieyu Zhang, Sijun Li, Taoyang Jia, Jiafei Duan, Yunqian Cheng, Jaemin Cho, Matthew Wallingford, Rustin Soraki, Chris Dongjoo Kim, Shuo Liu, Donovan Clay, Taira Anderson, Winson Han, Ali Farhadi, Bharath Hariharan, Zhongzheng Ren, Ranjay Krishna
Abstract:
Understanding objects in 3D from a single image is a cornerstone of spatial intelligence. A key step toward this goal is monocular 3D object detection‑‑recovering the extent, location, and orientation of objects from an input RGB image. To be practical in the open world, such a detector must generalize beyond closed‑set categories, support diverse prompt modalities, and leverage geometric cues when available. Progress is hampered by two bottlenecks: existing methods are designed for a single prompt type and lack a mechanism to incorporate additional geometric cues, and current 3D datasets cover only narrow categories in controlled environments, limiting open‑world transfer. In this work we address both gaps. First, we introduce WildDet3D, a unified geometry‑aware architecture that natively accepts text, point, and box prompts and can incorporate auxiliary depth signals at inference time. Second, we present WildDet3D‑Data, the largest open 3D detection dataset to date, constructed by generating candidate 3D boxes from existing 2D annotations and retaining only human‑verified ones, yielding over 1M images across 13.5K categories in diverse real‑world scenes. WildDet3D establishes a new state‑of‑the‑art across multiple benchmarks and settings. In the open‑world setting, it achieves 22.6/24.8 AP3D on our newly introduced WildDet3D‑Bench with text and box prompts. On Omni3D, it reaches 34.2/36.4 AP3D with text and box prompts, respectively. In zero‑shot evaluation, it achieves 40.3/48.9 ODS on Argoverse 2 and ScanNet. Notably, incorporating depth cues at inference time yields substantial additional gains (+20.7 AP on average across settings).
Authors:Xingming Liao, Ning Chen, Muying Shu, Yunpeng Yin, Peijian Zeng, Zhuowei Wang, Nankai Lin, Lianglun Cheng
Abstract:
Fine‑grained visual understanding and high‑level reasoning in real‑world open‑water environments remain under‑explored due to the lack of dedicated benchmarks. We introduce MARINER, a comprehensive benchmark built under the novel Entity‑Environment‑Event (3E) paradigm. MARINER contains 16,629 multi‑source maritime images with 63 fine‑grained vessel categories, diverse adverse environments, and 5 typical dynamic maritime incidents, covering fine‑grained classification, object detection, and visual question answering tasks. We conduct extensive evaluations on mainstream Multimodal Large language models (MLLMs) and establish baselines, revealing that even advanced models struggle with fine‑grained discrimination and causal reasoning in complex marine scenes. As a dedicated maritime benchmark, MARINER fills the gap of realistic and cognitive‑level evaluation for maritime multimodal understanding, and promotes future research on robust vision‑language models for open‑water applications. Appendix and supplementary materials are available at https://lxixim.github.io/MARINER.
Authors:Kun Wang, Yupeng Hu, Zhiran Li, Hao Liu, Qianlong Xiang, Liqiang Nie
Abstract:
In this report, we present our champion solution for the NTIRE 2026 Challenge on Video Saliency Prediction held in conjunction with CVPR 2026. To exploit complementary inductive biases for video saliency, we propose Video Saliency with Adaptive Gated Experts (ViSAGE), a multi‑expert ensemble framework. Each specialized decoder performs adaptive gating and modulation to refine spatio‑temporal features. The complementary predictions from different experts are then fused at inference. ViSAGE thereby aggregates diverse inductive biases to capture complex spatio‑temporal saliency cues in videos. On the Private Test set, ViSAGE ranked first on two out of four evaluation metrics, and outperformed most competing solutions on the other two metrics, demonstrating its effectiveness and generalization ability. Our code has been released at https://github.com/iLearn‑Lab/CVPRW26‑ViSAGE.
Authors:Jiahao Zhang, Shaofei Huang, Yaxiong Wang, Zhedong Zheng
Abstract:
Text‑based person search faces inherent limitations due to data scarcity, driven by stringent privacy constraints and the high cost of manual annotation. To mitigate this, existing methods usually rely on a Pretrain‑then‑Finetune paradigm, where models are first pretrained on synthetic person‑caption data to establish cross‑modal alignment, followed by fine‑tuning on labeled real‑world datasets. However, this paradigm lacks practicality in real‑world deployment scenarios, where large‑scale annotated target‑domain data is typically inaccessible. In this work, we propose a new Pretrain‑then‑Adapt paradigm that eliminates reliance on extensive target‑domain supervision through an offline test‑time adaptation manner, enabling dynamic model adaptation using only unlabeled test data with minimal post‑train time cost. To mitigate overconfidence with false positives of previous entropy‑based test‑time adaptation, we propose an Uncertainty‑Aware Test‑Time Adaptation (UATTA) framework, which introduces a bidirectional retrieval disagreement mechanism to estimate uncertainty, i.e., low uncertainty is assigned when an image‑text pair ranks highly in both image‑to‑text and text‑to‑image retrieval, indicating high alignment; otherwise, high uncertainty is detected. This indicator drives offline test‑time model recalibration without labels, effectively mitigating domain shift. We validate UATTA on four benchmarks, i.e., CUHK‑PEDES, ICFG‑PEDES, RSTPReid, and PAB, showing consistent improvements across both CLIP‑based (one‑stage) and XVLM‑based (two‑stage) frameworks. Ablation studies confirm that UATTA outperforms existing offline test‑time adaptation strategies, establishing a new benchmark for label‑efficient, deployable person search systems. Our code is available at https://github.com/nkuzjh/UATTA.
Authors:Gianluca Guglielmo, Marc Masana
Abstract:
State‑of‑the‑art post‑hoc out‑of‑distribution detection methods rely on intermediate layer activation editing. However, they exhibit inconsistent performance across datasets and models. We show that this instability is driven by differences in the activation distributions, and identify a failure mode of scaling‑based methods that arises when penultimate layer activations are not rectified. Motivated by this analysis, we propose \ours, a hyperparameter‑free post‑hoc method that replaces sorted activation magnitudes with a fixed in‑distribution reference profile. Our simple plug‑and‑play method shows strong and consistent performance across datasets and architectures without assumptions on the penultimate layer activation function, and without requiring any hyperparameter tuning, while preserving in‑distribution classification accuracy by construction. We further analyze what drives the improvement, showing that both inhibiting and exciting activation shifts independently contribute to better out‑of‑distribution discrimination.
Authors:Xiaoben Li, Jingyi Wu, Zeyu Cai, Siyuan Yu, Boqian Li, Yuliang Xiu
Abstract:
Human body fitting, which aligns parametric body models such as SMPL to raw 3D point clouds of clothed humans, serves as a crucial first step for downstream tasks like animation and texturing. An effective fitting method should be both locally expressive‑capturing fine details such as hands and facial features‑and globally robust to handle real‑world challenges, including clothing dynamics, pose variations, and noisy or partial inputs. Existing approaches typically excel in only one aspect, lacking an all‑in‑one solution. We upgrade ETCH to ETCH‑X, which leverages a tightness‑aware fitting paradigm to filter out clothing dynamics ("undress"), extends expressiveness with SMPL‑X, and replaces explicit sparse markers (which are highly sensitive to partial data) with implicit dense correspondences ("dense fit") for more robust and fine‑grained body fitting. Our disentangled "undress" and "dense fit" modular stages enable separate and scalable training on composable data sources, including diverse simulated garments (CLOTH3D), large‑scale full‑body motions (AMASS), and fine‑grained hand gestures (InterHand2.6M), improving outfit generalization and pose robustness of both bodies and hands. Our approach achieves robust and expressive fitting across diverse clothing, poses, and levels of input completeness, delivering a substantial performance improvement over ETCH on both: 1) seen data, such as 4D‑Dress (MPJPE‑All, 33.0% ) and CAPE (V2V‑Hands, 35.8% ), and 2) unseen data, such as BEDLAM2.0 (MPJPE‑All, 80.8% ; V2V‑All, 80.5% ). Code and models will be released at https://xiaobenli00.github.io/ETCH‑X/.
Authors:Zhengyang Sun, Yu Chen, Xin Zhou, Xiaofan Li, Xiwu Chen, Dingkang Liang, Xiang Bai
Abstract:
Text‑to‑video diffusion models have enabled open‑ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA , a training‑free identify‑then‑guide framework for improved numerical alignment. NUMINA identifies prompt‑layout inconsistencies by selecting discriminative self‑ and cross‑attention heads to derive a countable latent layout. It then refines this layout conservatively and modulates cross‑attention to guide regeneration. On the introduced CountBench, NUMINA improves counting accuracy by up to 7.4% on Wan2.1‑1.3B, and by 4.9% and 5.5% on 5B and 14B models, respectively. Furthermore, CLIP alignment is improved while maintaining temporal consistency. These results demonstrate that structural guidance complements seed search and prompt enhancement, offering a practical path toward count‑accurate text‑to‑video diffusion. The code is available at https://github.com/H‑EmbodVis/NUMINA.
Authors:Shilin Yan, Jintao Tong, Hongwei Xue, Xiaojun Tang, Yangyang Wang, Kunyu Shi, Guannan Zhang, Ruixuan Li, Yixiong Zou
Abstract:
The advent of agentic multimodal models has empowered systems to actively interact with external environments. However, current agents suffer from a profound meta‑cognitive deficit: they struggle to arbitrate between leveraging internal knowledge and querying external utilities. Consequently, they frequently fall prey to blind tool invocation, resorting to reflexive tool execution even when queries are resolvable from the raw visual context. This pathological behavior precipitates severe latency bottlenecks and injects extraneous noise that derails sound reasoning. Existing reinforcement learning protocols attempt to mitigate this via a scalarized reward that penalizes tool usage. Yet, this coupled formulation creates an irreconcilable optimization dilemma: an aggressive penalty suppresses essential tool use, whereas a mild penalty is entirely subsumed by the variance of the accuracy reward during advantage normalization, rendering it impotent against tool overuse. To transcend this bottleneck, we propose HDPO, a framework that reframes tool efficiency from a competing scalar objective to a strictly conditional one. By eschewing reward scalarization, HDPO maintains two orthogonal optimization channels: an accuracy channel that maximizes task correctness, and an efficiency channel that enforces execution economy exclusively within accurate trajectories via conditional advantage estimation. This decoupled architecture naturally induces a cognitive curriculum‑compelling the agent to first master task resolution before refining its self‑reliance. Extensive evaluations demonstrate that our resulting model, Metis, reduces tool invocations by orders of magnitude while simultaneously elevating reasoning accuracy.
Authors:Yunsong Zhou, Hangxu Liu, Xuekun Jiang, Xing Shen, Yuanzhen Zhou, Hui Wang, Baole Fang, Yang Tian, Mulin Yu, Qiaojun Yu, Li Ma, Hengjie Li, Hanqing Wang, Jia Zeng, Jiangmiao Pang
Abstract:
Robotic manipulation with deformable objects represents a data‑intensive regime in embodied learning, where shape, contact, and topology co‑evolve in ways that far exceed the variability of rigids. Although simulation promises relief from the cost of real‑world data acquisition, prevailing sim‑to‑real pipelines remain rooted in rigid‑body abstractions, producing mismatched geometry, fragile soft dynamics, and motion primitives poorly suited for cloth interaction. We posit that simulation fails not for being synthetic, but for being ungrounded. To address this, we introduce SIM1, a physics‑aligned real‑to‑sim‑to‑real data engine that grounds simulation in the physical world. Given limited demonstrations, the system digitizes scenes into metric‑consistent twins, calibrates deformable dynamics through elastic modeling, and expands behaviors via diffusion‑based trajectory generation with quality filtering. This pipeline transforms sparse observations into scaled synthetic supervision with near‑demonstration fidelity. Experiments show that policies trained on purely synthetic data achieve parity with real‑data baselines at a 1:15 equivalence ratio, while delivering 90% zero‑shot success and 50% generalization gains in real‑world deployment. These results validate physics‑aligned simulation as scalable supervision for deformable manipulation and a practical pathway for data‑efficient policy learning.
Authors:Tao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai, Wei Yin, Weiqiang Ren, Qian Zhang, Wei Hua, Sida Peng, Xiaoyang Guo, Xiaowei Zhou
Abstract:
This paper addresses the task of large‑scale 3D scene reconstruction from long video sequences. Recent feed‑forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D priors or geometric constraints. However, these methods often struggle to maintain reconstruction accuracy and consistency over long sequences due to limited memory capacity and the inability to effectively capture global contextual cues. In contrast, humans can naturally exploit the global understanding of the scene to inform local perception. Motivated by this, we propose a novel neural global context representation that efficiently compresses and retains long‑range scene information, enabling the model to leverage extensive contextual cues for enhanced reconstruction accuracy and consistency. The context representation is realized through a set of lightweight neural sub‑networks that are rapidly adapted during test time via self‑supervised objectives, which substantially increases memory capacity without incurring significant computational overhead. The experiments on multiple large‑scale benchmarks, including the KITTI Odometry~\citeGeiger2012CVPR and Oxford Spires~\citetao2025spires datasets, demonstrate the effectiveness of our approach in handling ultra‑large scenes, achieving leading pose accuracy and state‑of‑the‑art 3D reconstruction accuracy while maintaining efficiency. Code is available at https://zju3dv.github.io/scal3r.
Authors:Wenbo Hu, Xin Chen, Yan Gao-Tian, Yihe Deng, Nanyun Peng, Kai-Wei Chang
Abstract:
Group Relative Policy Optimization (GRPO) has emerged as the de facto Reinforcement Learning (RL) objective driving recent advancements in Multimodal Large Language Models. However, extending this success to open‑source multimodal generalist models remains heavily constrained by two primary challenges: the extreme variance in reward topologies across diverse visual tasks, and the inherent difficulty of balancing fine‑grained perception with multi‑step reasoning capabilities. To address these issues, we introduce Gaussian GRPO (G^2RPO), a novel RL training objective that replaces standard linear scaling with non‑linear distributional matching. By mathematically forcing the advantage distribution of any given task to strictly converge to a standard normal distribution, \mathcalN(0,1), G^2RPO theoretically ensures inter‑task gradient equity, mitigates vulnerabilities to heavy‑tail outliers, and offers symmetric update for positive and negative rewards. Leveraging the enhanced training stability provided by G^2RPO, we introduce two task‑level shaping mechanisms to seamlessly balance perception and reasoning. First, response length shaping dynamically elicits extended reasoning chains for complex queries while enforce direct outputs to bolster visual grounding. Second, entropy shaping tightly bounds the model's exploration zone, effectively preventing both entropy collapse and entropy explosion. Integrating these methodologies, we present OpenVLThinkerV2, a highly robust, general‑purpose multimodal model. Extensive evaluations across 18 diverse benchmarks demonstrate its superior performance over strong open‑source and leading proprietary frontier models.
Authors:Boyang Zhang, Sebastián G. Acosta, Preston Carlson, Sacha Bron, Pierre-Loïc Doulcet, Daniel B. Ospina, Simon Suo
Abstract:
AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and meaning needed for autonomous decisions, including correct table structure, precise chart data, semantically meaningful formatting, and visual grounding. Existing benchmarks do not fully capture this setting for enterprise automation, relying on narrow document distributions and text‑similarity metrics that miss agent‑critical failures. We introduce ParseBench, a benchmark of ~2,000 human‑verified pages from enterprise documents spanning insurance, finance, and government, organized around five capability dimensions: tables, charts, content faithfulness, semantic formatting, and visual grounding. Across 14 methods spanning vision‑language models, specialized document parsers, and LlamaParse, the benchmark reveals a fragmented capability landscape: no method is consistently strong across all five dimensions. LlamaParse Agentic achieves the highest overall score at 84.9%, and the benchmark highlights the remaining capability gaps across current systems. Dataset and evaluation code are available on https://huggingface.co/datasets/llamaindex/ParseBench and https://github.com/run‑llama/ParseBench.
Authors:Onkar Susladkar, Dong-Hwan Jang, Tushar Prakash, Adheesh Juvekar, Vedant Shah, Ayush Barik, Nabeel Bashir, Muntasir Wahed, Ritish Shrirao, Ismini Lourentzou
Abstract:
We introduce RewardFlow, an inversion‑free framework that steers pretrained diffusion and flow‑matching models at inference time through multi‑reward Langevin dynamics. RewardFlow unifies complementary differentiable rewards for semantic alignment, perceptual fidelity, localized grounding, object consistency, and human preference, and further introduces a differentiable VQA‑based reward that provides fine‑grained semantic supervision through language‑vision reasoning. To coordinate these heterogeneous objectives, we design a prompt‑aware adaptive policy that extracts semantic primitives from the instruction, infers edit intent, and dynamically modulates reward weights and step sizes throughout sampling. Across several image editing and compositional generation benchmarks, RewardFlow delivers state‑of‑the‑art edit fidelity and compositional alignment.
Authors:Simon Gerstenecker, Andreas Geiger, Katrin Renz
Abstract:
Generalization under distribution shift remains a central bottleneck for closed‑loop autonomous driving. Although simulators like CARLA enable safe and scalable testing, existing benchmarks rarely measure true generalization: they typically reuse training scenarios at test time. Success can therefore reflect memorization rather than robust driving behavior. We introduce Fail2Drive, the first paired‑route benchmark for closed‑loop generalization in CARLA, with 200 routes and 17 new scenario classes spanning appearance, layout, behavioral, and robustness shifts. Each shifted route is matched with an in‑distribution counterpart, isolating the effect of the shift and turning qualitative failures into quantitative diagnostics. Evaluating multiple state‑of‑the‑art models reveals consistent degradation, with an average success‑rate drop of 22.8%. Our analysis uncovers unexpected failure modes, such as ignoring objects clearly visible in the LiDAR and failing to learn the fundamental concepts of free and occupied space. To accelerate follow‑up work, Fail2Drive includes an open‑source toolbox for creating new scenarios and validating solvability via a privileged expert policy. Together, these components establish a reproducible foundation for benchmarking and improving closed‑loop driving generalization. We open‑source all code, data, and tools at https://github.com/autonomousvision/fail2drive .
Authors:Nan Huang, Pengcheng Yu, Weijia Zeng, James M. Rehg, Angjoo Kanazawa, Haiwen Feng, Qianqian Wang
Abstract:
Large‑scale multi‑view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with ground‑truth 3D/4D annotations. Such annotations are expensive and particularly scarce for dynamic scenes, limiting scalability. We propose SelfEvo, a self‑improving framework that continually improves pretrained multi‑view reconstruction models using unlabeled videos. SelfEvo introduces a self‑distillation scheme using spatiotemporal context asymmetry, enabling self‑improvement for learning‑based 4D perception without external annotations. We systematically study design choices that make self‑improvement effective, including loss signals, forms of asymmetry, and other training strategies. Across eight benchmarks spanning diverse datasets and domains, SelfEvo consistently improves pretrained baselines and generalizes across base models (e.g. VGGT and π^3), with significant gains on dynamic scenes. Overall, SelfEvo achieves up to 36.5% relative improvement in video depth estimation and 20.1% in camera estimation, without using any labeled data. Project Page: https://self‑evo.github.io/.
Authors:Johanna Karras, Yuanhao Wang, Yingwei Li, Ira Kemelmacher-Shlizerman
Abstract:
Given a person and a garment image, virtual try‑on (VTO) aims to synthesize a realistic image of the person wearing the garment, while preserving their original pose and identity. Although recent VTO methods excel at visualizing garment appearance, they largely overlook a crucial aspect of the try‑on experience: the accuracy of garment fit ‑‑ for example, depicting how an extra‑large shirt looks on an extra‑small person. A key obstacle is the absence of datasets that provide precise garment and body size information, particularly for "ill‑fit" cases, where garments are significantly too large or too small. Consequently, current VTO methods default to generating well‑fitted results regardless of the garment or person size.
In this paper, we take the first steps towards solving this open problem. We introduce FIT (Fit‑Inclusive Try‑on), a large‑scale VTO dataset comprising over 1.13M try‑on image triplets accompanied by precise body and garment measurements. We overcome the challenges of data collection via a scalable synthetic strategy: (1) We programmatically generate 3D garments using GarmentCode and drape them via physics simulation to capture realistic garment fit. (2) We employ a novel re‑texturing framework to transform synthetic renderings into photorealistic images while strictly preserving geometry. (3) We introduce person identity preservation into our re‑texturing model to generate paired person images (same person, different garments) for supervised training. Finally, we leverage our FIT dataset to train a baseline fit‑aware virtual try‑on model. Our data and results set the new state‑of‑the‑art for fit‑aware virtual try‑on, as well as offer a robust benchmark for future research. We will make all data and code publicly available on our project page: https://johannakarras.github.io/FIT.
Authors:Hang Ye, Xiaoxuan Ma, Fan Lu, Wayne Wu, Kwan-Yee Lin, Yizhou Wang
Abstract:
Digital human generation has been studied for decades and supports a wide range of real‑world applications. However, most existing systems are passively animated, relying on privileged state or scripted control, which limits scalability to novel environments. We instead ask: how can digital humans actively behave using only visual observations and specified goals in novel scenes? Achieving this would enable populating any 3D environments with digital humans at scale that exhibit spontaneous, natural, goal‑directed behaviors. To this end, we introduce Visually‑grounded Humanoid Agents, a coupled two‑layer (world‑agent) paradigm that replicates humans at multiple levels: they look, perceive, reason, and behave like real people in real‑world 3D scenes. The World Layer reconstructs semantically rich 3D Gaussian scenes from real‑world videos via an occlusion‑aware pipeline and accommodates animatable Gaussian‑based human avatars. The Agent Layer transforms these avatars into autonomous humanoid agents, equipping them with first‑person RGB‑D perception and enabling them to perform accurate, embodied planning with spatial awareness and iterative reasoning, which is then executed at the low level as full‑body actions to drive their behaviors in the scene. We further introduce a benchmark to evaluate humanoid‑scene interaction in diverse reconstructed environments. Experiments show our agents achieve robust autonomous behavior, yielding higher task success rates and fewer collisions than ablations and state‑of‑the‑art planning methods. This work enables active digital human population and advances human‑centric embodied AI. Data, code, and models will be open‑sourced.
Authors:Jingjing Wang, Zhengdong Hong, Chong Bao, Yuke Zhu, Junhan Sun, Guofeng Zhang
Abstract:
Human‑like generalization in open‑world remains a fundamental challenge for robotic manipulation. Existing learning‑based methods, including reinforcement learning, imitation learning, and vision‑language‑action‑models (VLAs), often struggle with novel tasks and unseen environments. Another promising direction is to explore generalizable representations that capture fine‑grained spatial and geometric relations for open‑world manipulation. While large‑language‑model (LLMs) and vision‑language‑model (VLMs) provide strong semantic reasoning based on language or annotated 2D representations, their limited 3D awareness restricts their applicability to fine‑grained manipulation. To address this, we propose LAMP, which lifts image‑editing as 3D priors to extract inter‑object 3D transformations as continuous, geometry‑aware representations. Our key insight is that image‑editing inherently encodes rich 2D spatial cues, and lifting these implicit cues into 3D transformations provides fine‑grained and accurate guidance for open‑world manipulation. Extensive experiments demonstrate that \codename delivers precise 3D transformations and achieves strong zero‑shot generalization in open‑world manipulation. Project page: https://zju3dv.github.io/LAMP/.
Authors:Rui Gan, Junyi Ma, Pei Li, Xingyou Yang, Kai Chen, Sikai Chen, Bin Ran
Abstract:
Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision‑language models (VLMs) show strong general reasoning capabilities, their performance in safety‑critical traffic scenarios remains insufficiently evaluated due to the ego‑vehicle focus of existing benchmarks. To bridge this gap, we present CrashSight, a large‑scale vision‑language benchmark for roadway crash understanding using real‑world roadside camera data. The dataset comprises 250 crash videos, annotated with 13K multiple‑choice question‑answer pairs organized under a two‑tier taxonomy. Tier 1 evaluates the visual grounding of scene context and involved parties, while Tier 2 probes higher‑level reasoning, including crash mechanics, causal attribution, temporal progression, and post‑crash outcomes. We benchmark 8 state‑of‑the‑art VLMs and show that, despite strong scene description capabilities, current models struggle with temporal and causal reasoning in safety‑critical scenarios. We provide a detailed analysis of failure scenarios and discuss directions for improving VLM crash understanding. The benchmark provides a standardized evaluation framework for infrastructure‑assisted perception in cooperative autonomous driving. The CrashSight benchmark, including the full dataset and code, is accessible at https://mcgrche.github.io/crashsight.
Authors:Fan Yang, Wenrui Chen, Guorun Yan, Ruize Liao, Wanjun Jia, Dongsheng Luo, Jiacheng Lin, Kailun Yang, Zhiyong Li, Yaonan Wang
Abstract:
In unstructured environments, functional dexterous grasping calls for the tight integration of semantic understanding, precise 3D functional localization, and physically interpretable execution. Modular hierarchical methods are more controllable and interpretable than end‑to‑end VLA approaches, but existing ones still rely on predefined affordance labels and lack the tight semantic‑‑pose coupling needed for functional dexterous manipulation. To address this, we propose BLaDA (Bridging Language to Dexterous Actions in 3DGS fields), an interpretable zero‑shot framework that grounds open‑vocabulary instructions as perceptual and control constraints for functional dexterous manipulation. BLaDA establishes an interpretable reasoning chain by first parsing natural language into a structured sextuple of manipulation constraints via a Knowledge‑guided Language Parsing (KLP) module. To achieve pose‑consistent spatial reasoning, we introduce the Triangular Functional Point Localization (TriLocation) module, which utilizes 3D Gaussian Splatting as a continuous scene representation and identifies functional regions under triangular geometric constraints. Finally, the 3D Keypoint Grasp Matrix Transformation Execution (KGT3D+) module decodes these semantic‑geometric constraints into physically plausible wrist poses and finger‑level commands. Extensive experiments on complex benchmarks demonstrate that BLaDA significantly outperforms existing methods in both affordance grounding precision and the success rate of functional manipulation across diverse categories and tasks. Code will be publicly available at https://github.com/PopeyePxx/BLaDA.
Authors:Wenli Zhang, Xianglong Shi, Sirui Zhao, Xinqi Chen, Guo Cheng, Yifan Xu, Tong Xu, Yong Liao
Abstract:
Diffusion‑based audio‑driven talking‑head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinformation. Existing protection methods are largely limited to a single modality, and neither image‑only nor audio‑only attacks can effectively suppress speech‑driven facial dynamics. To address this gap, we propose SyncBreaker, a stage‑aware multimodal protection framework that jointly perturbs portrait and audio inputs under modality‑specific perceptual constraints. Our key contributions are twofold. First, for the image stream, we introduce nullifying supervision with Multi‑Interval Sampling (MIS) across diffusion stages to steer the generation toward the static reference portrait by aggregating guidance from multiple denoising intervals. Second, for the audio stream, we propose Cross‑Attention Fooling (CAF), which suppresses interval‑specific audio‑conditioned cross‑attention responses. Both streams are optimized independently and combined at inference time to enable flexible deployment. We evaluate SyncBreaker in a white‑box proactive protection setting. Extensive experiments demonstrate that SyncBreaker more effectively degrades lip synchronization and facial dynamics than strong single‑modality baselines, while preserving input perceptual quality and remaining robust under purification. Code: https://github.com/kitty384/SyncBreaker.
Authors:Chensheng Dai, Shengjun Zhang, Min Chen, Yueqi Duan
Abstract:
3D Gaussian Splatting (3DGS) has demonstrated impressive performance in 3D scene reconstruction. Beyond novel view synthesis, it shows great potential for multi‑view surface reconstruction. Existing methods employ optimization‑based reconstruction pipelines that achieve precise and complete surface extractions. However, these approaches typically require dense input views and high time consumption for per‑scene optimization. To address these limitations, we propose SurfelSplat, a feed‑forward framework that generates efficient and generalizable pixel‑aligned Gaussian surfel representations from sparse‑view images. We observe that conventional feed‑forward structures struggle to recover accurate geometric attributes of Gaussian surfels because the spatial frequency of pixel‑aligned primitives exceeds Nyquist sampling rates. Therefore, we propose a cross‑view feature aggregation module based on the Nyquist sampling theorem. Specifically, we first adapt the geometric forms of Gaussian surfels with spatial sampling rate‑guided low‑pass filters. We then project the filtered surfels across all input views to obtain cross‑view feature correlations. By processing these correlations through a specially designed feature fusion network, we can finally regress Gaussian surfels with precise geometry. Extensive experiments on DTU reconstruction benchmarks demonstrate that our model achieves comparable results with state‑of‑the‑art methods, and predict Gaussian surfels within 1 second, offering a 100x speedup without costly per‑scene training.
Authors:Junyao Gao, Sibo Liu, Jiaxing Li, Yanan Sun, Yuanpeng Tu, Fei Shen, Weidong Zhang, Cairong Zhao, Jun Zhang
Abstract:
In this paper, we introduce MegaStyle, a novel and scalable data curation pipeline that constructs an intra‑style consistent, inter‑style diverse and high‑quality style dataset. We achieve this by leveraging the consistent text‑to‑image style mapping capability of current large generative models, which can generate images in the same style from a given style description. Building on this foundation, we curate a diverse and balanced prompt gallery with 170K style prompts and 400K content prompts, and generate a large‑scale style dataset MegaStyle‑1.4M via content‑style prompt combinations. With MegaStyle‑1.4M, we propose style‑supervised contrastive learning to fine‑tune a style encoder MegaStyle‑Encoder for extracting expressive, style‑specific representations, and we also train a FLUX‑based style transfer model MegaStyle‑FLUX. Extensive experiments demonstrate the importance of maintaining intra‑style consistency, inter‑style diversity and high‑quality for style dataset, as well as the effectiveness of the proposed MegaStyle‑1.4M. Moreover, when trained on MegaStyle‑1.4M, MegaStyle‑Encoder and MegaStyle‑FLUX provide reliable style similarity measurement and generalizable style transfer, making a significant contribution to the style transfer community. More results are available at our project website https://jeoyal.github.io/MegaStyle/.
Authors:Arnav Devalapally, Poornima Jain, Kartik Srinivas, Vineeth N. Balasubramanian
Abstract:
The increasing adaptation of vision models across domains, such as satellite imagery and medical scans, has raised an emerging privacy risk: models may inadvertently retain and leak sensitive source‑domain specific information in the target domain. This creates a compelling use case for machine unlearning to protect the privacy of sensitive source‑domain data. Among adaptation techniques, source‑free domain adaptation (SFDA) calls for an urgent need for machine unlearning (MU), where the source data itself is protected, yet the source model exposed during adaptation encodes its influence. Our experiments reveal that existing SFDA methods exhibit strong zero‑shot performance on source‑exclusive classes in the target domain, indicating they inadvertently leak knowledge of these classes into the target domain, even when they are not represented in the target data. We identify and address this risk by proposing an MU setting called SCADA‑UL: Unlearning Source‑exclusive ClAsses in Domain Adaptation. Existing MU methods do not address this setting as they are not designed to handle data distribution shifts. We propose a new unlearning method, where an adversarially generated forget class sample is unlearned by the model during the domain adaptation process using a novel rescaled labeling strategy and adversarial optimization. We also extend our study to two variants: a continual version of this problem setting and to one where the specific source classes to be forgotten may be unknown. Alongside theoretical interpretations, our comprehensive empirical results show that our method consistently outperforms baselines in the proposed setting while achieving retraining‑level unlearning performance on benchmark datasets. Our code is available at https://github.com/D‑Arnav/SCADA
Authors:You Hu, Chenzhuo Zhao, Changfa Mo, Haotian Liu, Xiaobai Li
Abstract:
Modern multimodal generators can now produce scientific figures at near‑publishable quality, creating a new challenge for visual forensics and research integrity. Unlike conventional AI‑generated natural images, scientific figures are structured, text‑dense, and tightly aligned with scholarly semantics, making them a distinct and difficult detection target. However, existing AI‑generated image detection benchmarks and methods are almost entirely developed for open‑domain imagery, leaving this setting largely unexplored. We present the first benchmark for AI‑generated scientific figure detection. To construct it, we develop an agent‑based data pipeline that retrieves licensed source papers, performs multimodal understanding of paper text and figures, builds structured prompts, synthesizes candidate figures, and filters them through a review‑driven refinement loop. The resulting benchmark covers multiple figure categories, multiple generation sources and aligned real‑‑synthetic pairs. We benchmark representative detectors under zero‑shot, cross‑generator, and degraded‑image settings. Results show that current methods fail dramatically in zero‑shot transfer, exhibit strong generator‑specific overfitting, and remain fragile under common post‑processing corruptions. These findings reveal a substantial gap between existing AIGI detection capabilities and the emerging distribution of high‑quality scientific figures. We hope this benchmark can serve as a foundation for future research on robust and generalizable scientific‑figure forensics. The dataset is available at https://github.com/Joyce‑yoyo/SciFigDetect.
Authors:Yiduo Jia, Muzhi Zhu, Hao Zhong, Mingyu Liu, Yuling Xi, Hao Chen, Bin Qin, Yongjie Yang, Zhenbo Luo, Chunhua Shen
Abstract:
To extend the reinforcement learning post‑training paradigm to omni‑modal models for concurrently bolstering video‑audio understanding and collaborative reasoning, we propose OmniJigsaw, a generic self‑supervised framework built upon a temporal reordering proxy task. Centered on the chronological reconstruction of shuffled audio‑visual clips, this paradigm strategically orchestrates visual and auditory signals to compel cross‑modal integration through three distinct strategies: Joint Modality Integration, Sample‑level Modality Selection, and Clip‑level Modality Masking. Recognizing that the efficacy of such proxy tasks is fundamentally tied to puzzle quality, we design a two‑stage coarse‑to‑fine data filtering pipeline, which facilitates the efficient adaptation of OmniJigsaw to massive unannotated omni‑modal data. Our analysis reveals a ``bi‑modal shortcut phenomenon'' in joint modality integration and demonstrates that fine‑grained clip‑level modality masking mitigates this issue while outperforming sample‑level modality selection. Extensive evaluations on 15 benchmarks show substantial gains in video, audio, and collaborative reasoning, validating OmniJigsaw as a scalable paradigm for self‑supervised omni‑modal learning.
Authors:Yunxiang Peng, Mengmeng Ma, Ziyu Yao, Xi Peng
Abstract:
Reliable generalization metrics are fundamental to the evaluation of machine learning models. Especially in high‑stakes applications where labeled target data are scarce, evaluation of models' generalization performance under distribution shift is a pressing need. We focus on two practical scenarios: (1) Before deployment, how to select the best model for unlabeled target data? (2) After deployment, how to monitor model performance under distribution shift? The central need in both cases is a reliable and label‑free proxy metric. Yet existing proxy metrics, such as model confidence or accuracy‑on‑the‑line, are often unreliable as they only assess model output while ignoring the internal mechanisms that produce them. We address this limitation by introducing a new perspective: using the inner workings of a model, i.e., circuits, as a predictive metric of generalization performance. Leveraging circuit discovery, we extract the causal interactions between internal representations as a circuit, from which we derive two metrics tailored to the two practical scenarios. (1) Before deployment, we introduce Dependency Depth Bias, which measures different models' generalization capability on target data. (2) After deployment, we propose Circuit Shift Score, which predicts a model's generalization under different distribution shifts. Across various tasks, both metrics demonstrate significantly improved correlation with generalization performance, outperforming existing proxies by an average of 13.4% and 34.1%, respectively. Our code is available at https://github.com/deep‑real/GenCircuit.
Authors:Linge Wang, Yingying Chen, Bingke Zhu, Lu Zhou, Jinqiao Wang
Abstract:
Recent advances in audio‑visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, jointly optimizing these objectives in a single forward pass forces the contrastive branch to rely on randomly visible patches designed for reconstruction rather than cross‑modal alignment, introducing semantic noise and optimization interference. We propose TG‑DP, a Teacher‑Guided Dual‑Path framework that decouples reconstruction and alignment into separate optimization paths. By disentangling the masking regimes of the two branches, TG‑DP enables the contrastive pathway to use a visibility pattern better suited to cross‑modal alignment. A teacher model further provides auxiliary guidance for organizing visible tokens in this branch, helping reduce interference and stabilize cross‑modal representation learning. TG‑DP achieves state‑of‑the‑art performance in zero‑shot retrieval. On AudioSet, it improves R@1 from 35.2% to 37.4% for video‑to‑audio retrieval and from 27.9% to 37.1% for audio‑to‑video retrieval. The learned representations also remain semantically robust, achieving state‑of‑the‑art linear‑probe performance on AS20K and VGGSound. Taken together, our results suggest that decoupling multimodal objectives and introducing teacher‑guided structure into the contrastive pathway provide an effective framework for improving large‑scale audio‑visual pretraining. Code is available at https://github.com/wanglg20/TG‑DP.
Authors:Sharva Gogawale, Gal Grudka, Daria Vasyutinsky-Shapira, Omer Ventura, Berat Kurar-Barakat, Nachum Dershowitz
Abstract:
A join is a set of manuscript fragments identified as originally emanating from the same manuscript. We study manuscript join retrieval: Given a query image of a fragment, retrieve other fragments originating from the same physical manuscript. We propose Bag of Bags (BoB), an image‑level representation that replaces the global‑level visual codebook of classical Bag of Words (BoW) with a fragment‑specific vocabulary of local visual words. Our pipeline trains a sparse convolutional autoencoder on binarized fragment patches, encodes connected components from each page, clusters the resulting embeddings with per‑image k‑means, and compares images using set‑to‑set distances between their local vocabularies. Evaluated on fragments from the Cairo Genizah, the best BoB variant (viz. Chamfer) achieves Hit@1 of 0.78 and MRR of 0.84, compared to 0.74 and 0.80, respectively, for the strongest BoW baseline (BoW‑RawPatches‑χ^2), a 6.1% relative improvement in top‑1 accuracy. We furthermore study a mass‑weighted BoB‑OT variant that incorporates cluster population into prototype matching and present a formal approximation guarantee bounding its deviation from full component‑level optimal transport. A two‑stage pipeline using a BoW shortlist followed by BoB‑OT reranking provides a practical compromise between retrieval strength and computational cost, supporting applicability to larger manuscript collections. The code and dataset are available at https://github.com/TAU‑CH/midrash_bob.
Authors:Luozheng Qin, Jia Gong, Qian Qiao, Tianjiao Li, Li Xu, Haoyu Pan, Chao Qu, Zhiyu Tan, Hao Li
Abstract:
Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This imbalance motivates us to invert the conventional paradigm: rather than extending understanding‑centric MLLMs to support generation, we propose Uni‑ViGU, a framework that unifies video generation and understanding by extending a video generator as the foundation. We introduce a unified flow method that performs continuous flow matching for video and discrete flow matching for text within a single process, enabling coherent multimodal generation. We further propose a modality‑driven MoE‑based framework that augments Transformer blocks with lightweight layers for text generation while preserving generative priors. To repurpose generation knowledge for understanding, we design a bidirectional training mechanism with two stages: Knowledge Recall reconstructs input prompts to leverage learned text‑video correspondences, while Capability Refinement fine‑tunes on detailed captions to establish discriminative shared representations. Experiments demonstrate that Uni‑ViGU achieves competitive performance on both video generation and understanding, validating generation‑centric architectures as a scalable path toward unified multimodal intelligence. Project Page and Code: https://fr0zencrane.github.io/uni‑vigu‑page/.
Authors:Junjie Fei, Jun Chen, Zechun Liu, Yunyang Xiong, Chong Zhou, Wei Wen, Junlin Han, Mingchen Zhuge, Saksham Suri, Qi Qian, Shuming Liu, Lemeng Wu, Raghuraman Krishnamoorthi, Vikas Chandra, Mohamed Elhoseiny, Chenchen Zhu
Abstract:
Adapting Multimodal Large Language Models (MLLMs) for hour‑long videos is bottlenecked by context limits. Dense visual streams saturate token budgets and exacerbate the lost‑in‑the‑middle phenomenon. Existing heuristics, like sparse sampling or uniform pooling, blindly sacrifice fidelity by discarding decisive moments and wasting bandwidth on irrelevant backgrounds. We propose Tempo, an efficient query‑aware framework compressing long videos for downstream understanding. Tempo leverages a Small Vision‑Language Model (SVLM) as a local temporal compressor, casting token reduction as an early cross‑modal distillation process to generate compact, intent‑aligned representations in a single forward pass. To enforce strict budgets without breaking causality, we introduce Adaptive Token Allocation (ATA). Exploiting the SVLM's zero‑shot relevance prior and semantic front‑loading, ATA acts as a training‑free O(1) dynamic router. It allocates dense bandwidth to query‑critical segments while compressing redundancies into minimal temporal anchors to maintain the global storyline. Extensive experiments show our 6B architecture achieves state‑of‑the‑art performance with aggressive dynamic compression (0.5‑16 tokens/frame). On the extreme‑long LVBench (4101s), Tempo scores 52.3 under a strict 8K visual budget, outperforming GPT‑4o and Gemini 1.5 Pro. Scaling to 2048 frames reaches 53.7. Crucially, Tempo compresses hour‑long videos substantially below theoretical limits, proving true long‑form video understanding relies on intent‑driven efficiency rather than greedily padded context windows.
Authors:Kang Ding, Hongsong Wang, Jie Gui, Liang Wang
Abstract:
Text‑to‑motion generation has attracted increasing attention in the research community recently, with potential applications in animation, virtual reality, robotics, and human‑computer interaction. Diffusion and autoregressive models are two popular and parallel research directions for text‑to‑motion generation. However, diffusion models often suffer from error amplification during noise prediction, while autoregressive models exhibit mode collapse due to motion discretization. To address these limitations, we propose a flexible, high‑fidelity, and semantically faithful text‑to‑motion framework, named Coordinate‑based Dual‑constrained Autoregressive Motion Generation (CDAMD). With motion coordinates as input, CDAMD follows the autoregressive paradigm and leverages diffusion‑inspired multi‑layer perceptrons to enhance the fidelity of predicted motions. Furthermore, a Dual‑Constrained Causal Mask is introduced to guide autoregressive generation, where motion tokens act as priors and are concatenated with textual encodings. Since there is limited work on coordinate‑based motion synthesis, we establish new benchmarks for both text‑to‑motion generation and motion editing. Experimental results demonstrate that our approach achieves state‑of‑the‑art performance in terms of both fidelity and semantic consistency on these benchmarks.
Authors:Christof Leitgeb, Thomas Puchleitner, Max Peter Ronecker, Daniel Watzenig
Abstract:
Reliable and weather‑robust perception systems are essential for safe autonomous driving and typically employ multi‑modal sensor configurations to achieve comprehensive environmental awareness. While recent automotive FMCW Radar‑based approaches achieved remarkable performance on detection tasks in adverse weather conditions, they exhibited limitations in resolving fine‑grained spatial details particularly critical for detecting smaller and vulnerable road users (VRUs). Furthermore, existing research has not adequately addressed VRU detection in adverse weather datasets such as K‑Radar. We present DinoRADE, a Radar‑centered detection pipeline that processes dense Radar tensors and aggregates vision features around transformed reference points in the camera perspective via deformable cross‑attention. Vision features are provided by a DINOv3 Vision Foundation Model. We present a comprehensive performance evaluation on the K‑Radar dataset in all weather conditions and are among the first to report detection performance individually for five object classes. Additionally, we compare our method with existing single‑class detection approaches and outperform recent Radar‑camera approaches by 12.1%. The code is available under https://github.com/chr‑is‑tof/RADE‑Net.
Authors:Francesca Fati, Alberto Rota, Adriana V. Gregory, Anna Catozzo, Maria C. Giuliano, Mrinal Dhar, Luigi De Vitis, Annie T. Packard, Francesco Multinu, Elena De Momi, Carrie L. Langstraat, Timothy L. Kline
Abstract:
Adnexal mass evaluation via ultrasound is a challenging clinical task, often hindered by subjective interpretation and significant inter‑observer variability. While automated segmentation is a foundational step for quantitative risk assessment, traditional fully supervised convolutional architectures frequently require large amounts of pixel‑level annotations and struggle with domain shifts common in medical imaging. In this work, we propose a label‑efficient segmentation framework that leverages the robust semantic priors of a pretrained DINOv3 foundational vision transformer backbone. By integrating this backbone with a Dense Prediction Transformer (DPT)‑style decoder, our model hierarchically reassembles multi‑scale features to combine global semantic representations with fine‑grained spatial details. Evaluated on a clinical dataset of 7,777 annotated frames from 112 patients, our method achieves state‑of‑the‑art performance compared to established fully supervised baselines, including U‑Net, U‑Net++, DeepLabV3, and MAnet. Specifically, we obtain a Dice score of 0.945 and improved boundary adherence, reducing the 95th‑percentile Hausdorff Distance by 11.4% relative to the strongest convolutional baseline. Furthermore, we conduct an extensive efficiency analysis demonstrating that our DINOv3‑based approach retains significantly higher performance under data starvation regimes, maintaining strong results even when trained on only 25% of the data. These results suggest that leveraging large‑scale self‑supervised foundations provides a promising and data‑efficient solution for medical image segmentation in data‑constrained clinical environments. Project Repository: https://github.com/FrancescaFati/MESA
Authors:Jun Li, Yingying Shi, Zhixuan Ruan, Nan Guo, Jianhua Xu
Abstract:
In a real‑world traffic scenario, varying‑scale objects are usually distributed in a cluttered background, which poses great challenges to accurate detection. Although current Mamba‑based methods can efficiently model long‑range dependencies, they still struggle to capture small objects with abundant local details, which hinders joint modeling of local structures and global semantics. Moreover, state‑space models exhibit limited hierarchical feature representation and weak cross‑scale interaction due to flat sequential modeling and insufficient spatial inductive biases, leading to sub‑optimal performance in complex scenes. To address these issues, we propose a Mamba with Deformable Dilated Convolutions Network (MDDCNet) for accurate traffic object detection in this study. In MDDCNet, a well‑designed hybrid backbone with successive Multi‑Scale Deformable Dilated Convolution (MSDDC) blocks and Mamba blocks enables hierarchical feature representation from local details to global semantics. Meanwhile, a Channel‑Enhanced Feed‑Forward Network (CE‑FFN) is further devised to overcome the limited channel interaction capability of conventional feed‑forward networks, whilst a Mamba‑based Attention‑Aggregating Feature Pyramid Network (A^2FPN) is constructed to achieve enhanced multi‑scale feature fusion and interaction. Extensive experimental results on public benchmark and real‑world datasets demonstrate the superiority of our method over various advanced detectors. The code is available at https://github.com/Bettermea/MDDCNet.
Authors:Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta
Abstract:
Talking‑head generation has advanced rapidly with diffusion‑based generative models, but training usually depends on centralized face‑video and speech datasets, raising major privacy concerns. The problem is more acute for personalized talking‑head generation, where identity‑specific data are highly sensitive and often cannot be pooled across users or devices. PrivFedTalk is presented as a privacy‑aware federated framework for personalized talking‑head generation that combines conditional latent diffusion with parameter‑efficient identity adaptation. A shared diffusion backbone is trained across clients, while each client learns lightweight LoRA identity adapters from local private audio‑visual data, avoiding raw data sharing and reducing communication cost. To address heterogeneous client distributions, Identity‑Stable Federated Aggregation (ISFA) weights client updates using privacy‑safe scalar reliability signals computed from on‑device identity consistency and temporal stability estimates. Temporal‑Denoising Consistency (TDC) regularization is introduced to reduce inter‑frame drift, flicker, and identity drift during federated denoising. To limit update‑side privacy risk, secure aggregation and client‑level differential privacy are applied to adapter updates. The implementation supports both low‑memory GPU execution and multi‑GPU client‑parallel training on heterogeneous shared hardware. Comparative experiments on the present setup across multiple training and aggregation conditions with PrivFedTalk, FedAvg, and FedProx show stable federated optimization and successful end‑to‑end training and evaluation under constrained resources. The results support the feasibility of privacy‑aware personalized talking‑head training in federated environments, while suggesting that stronger component‑wise, privacy‑utility, and qualitative claims need further standardized evaluation.
Authors:Minh Sao Khue Luu, Evgeniy N. Pavlovskiy, Bair N. Tuchinov
Abstract:
We propose a unified objective function, termed CATMIL, that augments the base segmentation loss with two auxiliary supervision terms operating at different levels. The first term, Component‑Adaptive Tversky, reweights voxel contributions based on connected components to balance the influence of lesions of different sizes. The second term, based on Multiple Instance Learning, introduces lesion‑level supervision by encouraging the detection of each lesion instance. These terms are combined with the standard nnU‑Net loss to jointly optimize voxel‑level segmentation accuracy and lesion‑level detection. We evaluate the proposed objective on the MSLesSeg dataset using a consistent nnU‑Net framework and 5‑fold cross‑validation. The results show that CATMIL achieves the most balanced performance across segmentation accuracy, lesion detection, and error control. It improves Dice score (0.7834) and reduces boundary error compared to standard losses. More importantly, it substantially increases small lesion recall and reduces false negatives, while maintaining the lowest false positive volume among compared methods. These findings demonstrate that integrating component‑level and lesion‑level supervision within a unified objective provides an effective and practical approach for improving small lesion segmentation in highly imbalanced settings. All code and pretrained models are available at https://github.com/luumsk/SmallLesionMRI.
Authors:Felix Embacher, Jonas Uhrig, Marius Cordts, Markus Enzweiler
Abstract:
Retrieving rare and safety‑critical driving scenarios from large‑scale datasets is essential for building robust autonomous driving (AD) systems. As dataset sizes continue to grow, the key challenge shifts from collecting more data to efficiently identifying the most relevant samples. We introduce SearchAD, a large‑scale rare image retrieval dataset for AD containing over 423k frames drawn from 11 established datasets. SearchAD provides high‑quality manual annotations of more than 513k bounding boxes covering 90 rare categories. It specifically targets the needle‑in‑a‑haystack problem of locating extremely rare classes, with some appearing fewer than 50 times across the entire dataset. Unlike existing benchmarks, which focused on instance‑level retrieval, SearchAD emphasizes semantic image retrieval with a well‑defined data split, enabling text‑to‑image and image‑to‑image retrieval, few‑shot learning, and fine‑tuning of multi‑modal retrieval models. Comprehensive evaluations show that text‑based methods outperform image‑based ones due to stronger inherent semantic grounding. While models directly aligning spatial visual features with language achieve the best zero‑shot results, and our fine‑tuning baseline significantly improves performance, absolute retrieval capabilities remain unsatisfactory. With a held‑out test set on a public benchmark server, SearchAD establishes the first large‑scale dataset for retrieval‑driven data curation and long‑tail perception research in AD: https://iis‑esslingen.github.io/searchad/
Authors:Yun Zhu, Jianjun Qian, Jian Yang, Jin Xie, Na Zhao
Abstract:
Incremental 3D object perception is a critical step toward embodied intelligence in dynamic indoor environments. However, existing incremental 3D detection methods rely on extensive annotations of novel classes for satisfactory performance. To address this limitation, we propose FI3Det, a Few‑shot Incremental 3D Detection framework that enables efficient 3D perception with only a few novel samples by leveraging vision‑language models (VLMs) to learn knowledge of unseen categories. FI3Det introduces a VLM‑guided unknown object learning module in the base stage to enhance perception of unseen categories. Specifically, it employs VLMs to mine unknown objects and extract comprehensive representations, including 2D semantic features and class‑agnostic 3D bounding boxes. To mitigate noise in these representations, a weighting mechanism is further designed to re‑weight the contributions of point‑ and box‑level features based on their spatial locations and feature consistency within each box. Moreover, FI3Det proposes a gated multimodal prototype imprinting module, where category prototypes are constructed from aligned 2D semantic and 3D geometric features to compute classification scores, which are then fused via a multimodal gating mechanism for novel object detection. As the first framework for few‑shot incremental 3D object detection, we establish both batch and sequential evaluation settings on two datasets, ScanNet V2 and SUN RGB‑D, where FI3Det achieves strong and consistent improvements over baseline methods. Code is available at https://github.com/zyrant/FI3Det.
Authors:Zile Guo, Zhan Chen, Enze Zhu, Kan Wei, Yongkang Zou, Xiaoxuan Liu, Lei Wang
Abstract:
Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex 3D dynamics is essential for autonomous navigation and robust decision‑making in unconstrained environments. However, under the highly dynamic camera trajectories typical of UAV views, existing world models often struggle to maintain spatiotemporal physical consistency. A key reason lies in the distribution bias of current training data: most existing datasets exhibit restricted 2.5D motion patterns, such as ground‑constrained autonomous driving scenes or relatively smooth human‑centric egocentric videos, and therefore lack realistic high‑dynamic 6‑DoF UAV motion priors. To address this gap, we present MotionScape, a large‑scale real‑world UAV‑view video dataset with highly dynamic motion for world modeling. MotionScape contains over 30 hours of 4K UAV‑view videos, totaling more than 4.5M frames. This novel dataset features semantically and geometrically aligned training samples, where diverse real‑world UAV videos are tightly coupled with accurate 6‑DoF camera trajectories and fine‑grained natural language descriptions. To build the dataset, we develop an automated multi‑stage processing pipeline that integrates CLIP‑based relevance filtering, temporal segmentation, robust visual SLAM for trajectory recovery, and large‑language‑model‑driven semantic annotation. Extensive experiments show that incorporating such semantically and geometrically aligned annotations effectively improves the ability of existing world models to simulate complex 3D dynamics and handle large viewpoint shifts, thereby benefiting decision‑making and planning for UAV agents in complex environments. The dataset is publicly available at https://github.com/Thelegendzz/MotionScape
Authors:Tao Han, Zhibin Wen, Zhenghao Chen, Fenghua Lin, Junyu Gao, Song Guo, Lei Bai
Abstract:
While AI‑based numerical weather prediction (NWP) enables rapid forecasting, generating high‑resolution outputs remains computationally demanding due to limited multi‑scale adaptability and inefficient data representations. We propose the 3D Gaussian splatting‑based scale‑aware vision transformer (GSSA‑ViT), a novel framework for arbitrary‑resolution forecasting and flexible downscaling of high‑dimensional atmospheric fields. Specifically, latitude‑longitude grid points are treated as centers of 3D Gaussians. A generative 3D Gaussian prediction scheme is introduced to estimate key parameters, including covariance, attributes, and opacity, for unseen samples, improving generalization and mitigating overfitting. In addition, a scale‑aware attention module is designed to capture cross‑scale dependencies, enabling the model to effectively integrate information across varying downscaling ratios and support continuous resolution adaptation. To our knowledge, this is the first NWP approach that combines generative 3D Gaussian modeling with scale‑aware attention for unified multi‑scale prediction. Experiments on ERA5 show that the proposed method accurately forecasts 87 atmospheric variables at arbitrary resolutions, while evaluations on ERA5 and CMIP6 demonstrate its superior performance in downscaling tasks. The proposed framework provides an efficient and scalable solution for high‑resolution, multi‑scale atmospheric prediction and downscaling. Code is available at: https://github.com/binbin2xs/weather‑GS.
Authors:Wenkui Yang, Chao Jin, Haisu Zhu, Weilin Luo, Derek Yuen, Kun Shao, Huaibo Huang, Junxian Duan, Jie Cao, Ran He
Abstract:
Existing red‑teaming studies on GUI agents have important limitations. Adversarial perturbations typically require white‑box access, which is unavailable for commercial systems, while prompt injection is increasingly mitigated by stronger safety alignment. To study robustness under a more practical threat model, we propose Semantic‑level UI Element Injection, a red‑teaming setting that overlays safety‑aligned and harmless UI elements onto screenshots to misdirect the agent's visual grounding. Our method uses a modular Editor‑Overlapper‑Victim pipeline and an iterative search procedure that samples multiple candidate edits, keeps the best cumulative overlay, and adapts future prompt strategies based on previous failures. Across five victim models, our optimized attacks improve attack success rate by up to 4.4x over random injection on the strongest victims. Moreover, elements optimized on one source model transfer effectively to other target models, indicating model‑agnostic vulnerabilities. After the first successful attack, the victim still clicks the attacker‑controlled element in more than 15% of later independent trials, versus below 1% for random injection, showing that the injected element acts as a persistent attractor rather than simple visual clutter.
Authors:Hazza Mahmood, Yongqiang Yu, Rao Anwer
Abstract:
Accurate and interpretable plant disease diagnosis remains a major challenge for vision‑language models (VLMs) in real‑world agriculture. We introduce AgriChain, a dataset of approximately 11,000 expert‑curated leaf images spanning diverse crops and pathologies, each paired with (i) a disease label, (ii) a calibrated confidence score (High/Medium/Low), and (iii) an expert‑verified chain‑of‑thought (CoT) rationale. Draft explanations were first generated by GPT‑4o and then verified by a professional agricultural engineer using standardized descriptors (e.g., lesion color, margin, and distribution). We fine‑tune Qwen2.5‑VL‑3B on AgriChain, resulting in a specialized model termed AgriChain‑VL3B, to jointly predict diseases and generate visually grounded reasoning. On a 1,000‑image test set, our CoT‑supervised model achieves 73.1% top‑1 accuracy (macro F1 = 0.466; weighted F1 = 0.655), outperforming strong baselines including Gemini 1.5 Flash, Gemini 2.5 Pro, and GPT‑4o Mini. The generated explanations align closely with expert reasoning, consistently referencing key visual cues. These findings demonstrate that expert‑verified reasoning supervision significantly enhances both accuracy and interpretability, bridging the gap between generic multimodal models and human expertise, and advancing trustworthy, globally deployable AI for sustainable agriculture. The dataset and code are publicly available at: https://github.com/hazzanabeel12‑netizen/agrichain
Authors:Qihui Zhu, Tao Zhang, Yuchen Wang, Zijian Wen, Mengjie Zhang, Shuangwu Chen, Xiaobin Tan, Jian Yang, Yang Liu, Zhenhua Dong, Xianzhi Yu, Yinfei Pan
Abstract:
In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real‑time or resource‑constrained applications. Visual token pruning is a promising strategy for reducing the cost of MLLM inference by removing redundant visual tokens. Existing research usually assumes that all attention heads contribute equally to the visual interpretation. However, our study reveals that different heads may capture distinct visual semantics and inherently play distinct roles in visual processing. In light of this observation, we propose HAWK, a head importance‑aware visual token pruning method that perceives the varying importance of attention heads in visual tasks to maximize the retention of crucial tokens. By leveraging head importance weights and text‑guided attention to assess visual token significance, HAWK effectively retains task‑relevant visual tokens while removing redundant ones. The proposed HAWK is entirely training‑free and can be seamlessly applied to various MLLMs. Extensive experiments on multiple mainstream vision‑language benchmarks demonstrate that HAWK achieves state‑of‑the‑art accuracy. When applied to Qwen2.5‑VL, HAWK retains 96.0% of the original accuracy after pruning 80.2% of the visual tokens. Additionally, it reduces end‑to‑end latency to 74.4% of the original and further decreases GPU memory usage across the tested models. The code is available at https://github.com/peppery77/HAWK.git.
Authors:Changwoon Choi, Hyunsoo Lee, Clément Jambon, Yael Vinker, Young Min Kim
Abstract:
Recent generative models can create visually plausible 3D representations of objects. However, the generation process often allows for implicit control signals, such as contextual descriptions, and rarely supports bold geometric distortions beyond existing data distributions. We propose a geometric stylization framework that deforms a 3D mesh, allowing it to express the style of an image. While style is inherently ambiguous, we utilize pre‑trained diffusion models to extract an abstract representation of the provided image. Our coarse‑to‑fine stylization pipeline can drastically deform the input 3D model to express a diverse range of geometric variations while retaining the valid topology of the original mesh and part‑level semantics. We also propose an approximate VAE encoder that provides efficient and reliable gradients from mesh renderings. Extensive experiments demonstrate that our method can create stylized 3D meshes that reflect unique geometric features of the pictured assets, such as expressive poses and silhouettes, thereby supporting the creation of distinctive artistic 3D creations. Project page: https://changwoonchoi.github.io/GeoStyle
Authors:Chanhyuk Choi, Taesoo Kim, Donggyu Lee, Siyeol Jung, Taehwan Kim
Abstract:
Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role. However, existing approaches often limit expressive flexibility and struggle to generate extended emotions. Label‑based methods represent emotions with discrete categories, which fail to capture a wide range of emotions. Audio‑based methods can leverage emotionally rich speech signals ‑ and even benefit from expressive text‑to‑speech (TTS) synthesis ‑ but they fail to express the target emotions because emotions and linguistic contents are entangled in emotional speeches. Images‑based methods, on the other hand, rely on target reference images to guide emotion transfer, yet they require high‑quality frontal views and face challenges in acquiring reference data for extended emotions (e.g., sarcasm). To address these limitations, we propose Cross‑Modal Emotion Transfer (C‑MET), a novel approach that generates facial expressions based on speeches by modeling emotion semantic vectors between speech and visual feature spaces. C‑MET leverages a large‑scale pretrained audio encoder and a disentangled facial expression encoder to learn emotion semantic vectors that represent the difference between two different emotional embeddings across modalities. Extensive experiments on the MEAD and CREMA‑D datasets demonstrate that our method improves emotion accuracy by 14% over state‑of‑the‑art methods, while generating expressive talking face videos ‑ even for unseen extended emotions. Code, checkpoint, and demo are available at https://chanhyeok‑choi.github.io/C‑MET/
Authors:Alvin Kimbowa, Arjun Parmar, Ibrahim Mujtaba, Will Wei, Maziar Badii, Matthew Harkey, David Liu, Ilker Hacihaliloglu
Abstract:
Objective: To develop a robust and compact deep learning model for automated knee cartilage segmentation on point‑of‑care ultrasound (POCUS) devices.
Methods: We propose MonoUNet, a novel, highly compact segmentation model consisting of (i) an aggressively reduced U‑Net backbone, (ii) a trainable monogenic block that extracts multi‑scale local phase features from the input, and (iii) a gating mechanism that injects these features into the encoder stages to reduce sensitivity to variations in ultrasound image appearance. MonoUNet segmentation performance was evaluated on a multi‑site, multi‑device knee cartilage ultrasound dataset using Dice score and mean average surface distance (MASD). Agreement between MonoUNet and manual cartilage outcomes (thickness and echo intensity) was assessed using Bland‑Altman analysis with 95% limits of agreement, and reliability was assessed using intraclass correlation coefficient (ICC_2,k).
Results: Overall, MonoUNet outperformed existing lightweight segmentation models, with average Dice scores ranging from 92.62% to 94.82% and MASD values between 0.133 mm and 0.254 mm. MonoUNet reduces the number of parameters by 10x‑‑700x and computational cost by 14x‑‑2000x relative to existing lightweight models. MonoUNet cartilage outcomes showed excellent reliability and agreement with the manual outcomes: intraclass correlation coefficients (ICC_2,k)=0.96 and bias=2.00% (0.047 mm) for average thickness, and ICC_2,k=0.99 and bias=0.80% (0.328 a.u.) for echo intensity.
Conclusion: Incorporating trainable local phase features improves the robustness of highly compact neural networks for knee cartilage segmentation across varying acquisition settings and could support scalable ultrasound‑based assessment and monitoring of knee osteoarthritis using POCUS devices. The code is publicly available at https://github.com/alvinkimbowa/monounet.
Authors:Peiran Xu, Jiaqi Zheng, Yadong Mu
Abstract:
This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accomplish a given task. Although recent Vision‑Language Models (VLMs) have achieved impressive results in multimodal understanding and reasoning, their performance remains limited when applied to embodied planning that involves multi‑turn interaction, long‑horizon reasoning, and extended context analysis. To bridge this gap, we propose RoboAgent, a capability‑driven planning pipeline in which the model actively invokes different sub‑capabilities. Each capability maintains its own context, and produces intermediate reasoning results or interacts with the environment according to the query given by a scheduler. This framework decomposes complex planning into a sequence of basic vision‑language problems that VLMs can better address, enabling a more transparent and controllable reasoning process. The scheduler and all capabilities are implemented with a single VLM, without relying on external tools. To train this VLM, we adopt a multi‑stage paradigm that consists of: (1) behavior cloning with expert plans, (2) DAgger training using trajectories collected by the model, and (3) reinforcement learning guided by an expert policy. Across these stages, we exploit the internal information of the environment simulator to construct high‑quality supervision for each capability, and we further introduce augmented and synthetic data to enhance the model's performance in more diverse scenarios. Extensive experiments on widely used embodied task planning benchmarks validate the effectiveness of the proposed approach. Our codes will be available at https://github.com/woyut/RoboAgent_CVPR26.
Authors:Zihao Liu, Xiaoyu Wu, Wenna Li, Jianqin Wu, Linlin Yang
Abstract:
Open‑world video anomaly detection (OWVAD) aims to detect and explain abnormal events under different anomaly definitions, which is important for applications such as intelligent surveillance and live‑streaming content moderation. Recent MLLM‑based methods have shown promising open‑world generalization, but still suffer from three major limitations: inefficiency for practical deployment, lack of streaming processing adaptation, and limited support for dynamic anomaly definitions in both modeling and evaluation. To address these issues, this paper proposes ESOM, an efficient streaming OWVAD model that operates in a training‑free manner. ESOM includes a Definition Normalization module to structure user prompts for reducing hallucination, an Inter‑frame‑matched Intra‑frame Token Merging module to compress redundant visual tokens, a Hybrid Streaming Memory module for efficient causal inference, and a Probabilistic Scoring module that converts interval‑level textual outputs into frame‑level anomaly scores. In addition, this paper introduces OpenDef‑Bench, a new benchmark with clean surveillance videos and diverse natural anomaly definitions for evaluating performance under varying conditions. Extensive experiments show that ESOM achieves real‑time efficiency on a single GPU and state‑of‑the‑art performance in anomaly temporal localization, classification, and description generation. The code and benchmark will be released at https://github.com/Kamino666/ESOM_OpenDef‑Bench.
Authors:Junxiong Liang, Mengwei Bao, Tianxiang Wang, Xinggang Wang, An-An Liu, Ryan Wen Liu
Abstract:
Ship detection for navigation is a fundamental perception task in intelligent waterway transportation systems. However, existing public ship detection datasets remain limited in terms of scale, the proportion of small‑object instances, and scene diversity, which hinders the systematic evaluation and generalization study of detection algorithms in complex maritime environments. To this end, we construct WUTDet, a large‑scale ship detection dataset. WUTDet contains 100,576 images and 381,378 annotated ship instances, covering diverse operational scenarios such as ports, anchorages, navigation, and berthing, as well as various imaging conditions including fog, glare, low‑lightness, and rain, thereby exhibiting substantial diversity and challenge. Based on WUTDet, we systematically evaluate 20 baseline models from three mainstream detection architectures, namely CNN, Transformer, and Mamba. Experimental results show that the Transformer architecture achieves superior overall detection accuracy (AP) and small‑object detection performance (APs), demonstrating stronger adaptability to complex maritime scenes; the CNN architecture maintains an advantage in inference efficiency, making it more suitable for real‑time applications; and the Mamba architecture achieves a favorable balance between detection accuracy and computational efficiency. Furthermore, we construct a unified cross‑dataset test set, Ship‑GEN, to evaluate model generalization. Results on Ship‑GEN show that models trained on WUTDet exhibit stronger generalization under different data distributions. These findings demonstrate that WUTDet provides effective data support for the research, evaluation, and generalization analysis of ship detection algorithms in complex maritime scenarios. The dataset is publicly available at: https://github.com/MAPGroup/WUTDet.
Authors:Hang Zhang, Qijian Tian, Jingyu Gong, Daoguo Dong, Xuhong Wang, Yuan Xie, Xin Tan
Abstract:
Articulated objects are essential for embodied AI and world models, yet inferring their kinematics from a single closed‑state image remains challenging because crucial motion cues are often occluded. Existing methods either require multi‑state observations or rely on explicit part priors, retrieval, or other auxiliary inputs that partially expose the structure to be inferred. In this work, we present DailyArt, which formulates articulated joint estimation from a single static image as a synthesis‑mediated reasoning problem. Instead of directly regressing joints from a heavily occluded observation, DailyArt first synthesizes a maximally articulated opened state under the same camera view to expose articulation cues, and then estimates the full set of joint parameters from the discrepancy between the observed and synthesized states. Using a set‑prediction formulation, DailyArt recovers all joints simultaneously without requiring object‑specific templates, multi‑view inputs, or explicit part annotations at test time. Taking estimated joints as conditions, the framework further supports part‑level novel state synthesis as a downstream capability. Extensive experiments show that DailyArt achieves strong performance in articulated joint estimation and supports part‑level novel state synthesis conditioned on joints. Project page is available at https://rangooo123.github.io/DaliyArt.github.io/.
Authors:Huibin Bai, Shuai Li, Hanxiao Zhai, Yanbo Gao, Chong Lv, Yibo Wang, Haipeng Ping, Wei Hua, Xingyu Gao
Abstract:
Monocular Depth Estimation (MDE) is a fundamental computer vision task with important applications in 3D vision. The current mainstream MDE methods employ an encoder‑decoder architecture with multi‑level/scale feature processing. However, the limitations of the current architecture and the effects of different‑level features on the prediction accuracy are not evaluated. In this paper, we first investigate the above problem and show that there is still substantial potential in the current framework if encoder features can be improved. Therefore, we propose to formulate the depth estimation problem from the feature restoration perspective, by treating pretrained encoder features as degraded features of an assumed ground truth feature that yields the ground truth depth map. Then an Invertible Transform‑enhanced Indirect Diffusion (InvT‑IndDiffusion) module is developed for feature restoration. Due to the absence of direct supervision on feature, only indirect supervision from the final sparse depth map is used. During the iterative procedure of diffusion, this results in feature deviations among steps. The proposed InvT‑IndDiffusion solves this problem by using an invertible transform‑based decoder under the bi‑Lipschitz condition. Finally, a plug‑and‑play Auxiliary Viewpoint‑based Low‑level Feature Enhancement module (AV‑LFE) is developed to enhance local details with auxiliary viewpoint when available. Experiments demonstrate that the proposed method achieves better performance than the state‑of‑the‑art methods on various datasets. Specifically on the KITTI benchmark, compared with the baseline, the performance is improved by 4.09% and 37.77% under different training settings in terms of RMSE. Code is available at https://github.com/whitehb1/IID‑RDepth.
Authors:Pavan Kumar Anasosalu Vasu, Cem Koc, Fartash Faghri, Chun-Liang Li, Bo Feng, Zhengfeng Lai, Meng Cao, Oncel Tuzel, Hadi Pouransari
Abstract:
Streaming vision‑language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real‑time visual assistants. Existing VLM frameworks predominantly assess models in offline settings. In contrast, the performance of a streaming VLM depends on additional metrics beyond pure video understanding, including proactiveness, which reflects the timeliness of the model's responses, and consistency, which captures the robustness of its responses over time. To address this limitation, we propose VSAS‑Bench, a new framework and benchmark for Visual Streaming Assistants. In contrast to prior benchmarks that primarily employ single‑turn question answering on video inputs, VSAS‑Bench features temporally dense annotations with over 18,000 annotations across diverse input domains and task types. We introduce standardized synchronous and asynchronous evaluation protocols, along with metrics that isolate and measure distinct capabilities of streaming VLMs. Using this framework, we conduct large‑scale evaluations of recent video and streaming VLMs, analyzing the accuracy‑latency trade‑off under key design factors such as memory buffer length, memory access policy, and input resolution, yielding several practical insights. Finally, we show empirically that conventional VLMs can be adapted to streaming settings without additional training, and demonstrate that these adapted models outperform recent streaming VLMs. For example, Qwen3‑VL‑4B surpasses Dispider, the best streaming VLM on our benchmark, by 3% under the asynchronous protocol. The benchmark and code will be available at https://github.com/apple/ml‑vsas‑bench.
Authors:Tencent Robotics X, HY Vision Team, :, Xumin Yu, Zuyan Liu, Ziyi Wang, He Zhang, Yongming Rao, Fangfu Liu, Yani Zhang, Ruowen Zhao, Oran Wang, Yves Liang, Haitao Lin, Minghui Wang, Yubo Dong, Kevin Cheng, Bolin Ni, Rui Huang, Han Hu, Zhengyou Zhang, Linus, Shunyu Yao
Abstract:
We introduce HY‑Embodied‑0.5, a family of foundation models specifically designed for real‑world embodied agents. To bridge the gap between general Vision‑Language Models (VLMs) and the demands of embodied agents, our models are developed to enhance the core capabilities required by embodied intelligence: spatial and temporal visual perception, alongside advanced embodied reasoning for prediction, interaction, and planning. The HY‑Embodied‑0.5 suite comprises two primary variants: an efficient model with 2B activated parameters designed for edge deployment, and a powerful model with 32B activated parameters targeted for complex reasoning. To support the fine‑grained visual perception essential for embodied tasks, we adopt a Mixture‑of‑Transformers (MoT) architecture to enable modality‑specific computing. By incorporating latent tokens, this design effectively enhances the perceptual representation of the models. To improve reasoning capabilities, we introduce an iterative, self‑evolving post‑training paradigm. Furthermore, we employ on‑policy distillation to transfer the advanced capabilities of the large model to the smaller variant, thereby maximizing the performance potential of the compact model. Extensive evaluations across 22 benchmarks, spanning visual perception, spatial reasoning, and embodied understanding, demonstrate the effectiveness of our approach. Our MoT‑2B model outperforms similarly sized state‑of‑the‑art models on 16 benchmarks, while the 32B variant achieves performance comparable to frontier models such as Gemini 3.0 Pro. In downstream robot control experiments, we leverage our robust VLM foundation to train an effective Vision‑Language‑Action (VLA) model, achieving compelling results in real‑world physical evaluations. Code and models are open‑sourced at https://github.com/Tencent‑Hunyuan/HY‑Embodied.
Authors:Xiangru Jian, Hao Xu, Wei Pang, Xinjian Zhao, Chengyu Tao, Qixin Zhang, Xikun Zhang, Chao Zhang, Guanzhi Deng, Alex Xue, Juan Du, Tianshu Yu, Garth Tarr, Linqi Song, Qiuzhuang Sun, Dacheng Tao
Abstract:
The manufacturing sector is increasingly adopting Multimodal Large Language Models (MLLMs) to transition from simple perception to autonomous execution, yet current evaluations fail to reflect the rigorous demands of real‑world manufacturing environments. Progress is hindered by data scarcity and a lack of fine‑grained domain semantics in existing datasets. To bridge this gap, we introduce FORGE. Wefirst construct a high‑quality multimodal dataset that combines real‑world 2D images and 3D point clouds, annotated with fine‑grained domain semantics (e.g., exact model numbers). We then evaluate 18 state‑of‑the‑art MLLMs across three manufacturing tasks, namely workpiece verification, structural surface inspection, and assembly verification, revealing significant performance gaps. Counter to conventional understanding, the bottleneck analysis shows that visual grounding is not the primary limiting factor. Instead, insufficient domain‑specific knowledge is the key bottleneck, setting a clear direction for future research. Beyond evaluation, we show that our structured annotations can serve as an actionable training resource: supervised fine‑tuning of a compact 3B‑parameter model on our data yields up to 90.8% relative improvement in accuracy on held‑out manufacturing scenarios, providing preliminary evidence for a practical pathway toward domain‑adapted manufacturing MLLMs. The code and datasets are available at https://ai4manufacturing.github.io/forge‑web.
Authors:Wenze Wang, Mehdi Hosseinzadeh, Feras Dayoub
Abstract:
Robotic manipulation systems that follow language instructions often execute grasp primitives in a largely single‑shot manner: a model proposes an action, the robot executes it, and failures such as empty grasps, slips, stalls, timeouts, or semantically wrong grasps are not surfaced to the decision layer in a structured way. Inspired by agentic loops in digital tool‑using agents, we reformulate language‑guided grasping as a bounded embodied agent operating over grounded execution states, where physical actions expose an explicit tool‑state stream. We introduce a physical agentic loop that wraps an unmodified learned manipulation primitive (grasp‑and‑lift) with (i) an event‑based interface and (ii) an execution monitoring layer, Watchdog, which converts noisy gripper telemetry into discrete outcome labels using contact‑aware fusion and temporal stabilization. These outcome events, optionally combined with post‑grasp semantic verification, are consumed by a deterministic bounded policy that finalizes, retries, or escalates to the user for clarification, guaranteeing finite termination. We validate the resulting loop on a mobile manipulator with an eye‑in‑hand D405 camera, keeping the underlying grasp model unchanged and evaluating representative scenarios involving visual ambiguity, distractors, and induced execution failures. Results show that explicit execution‑state monitoring and bounded recovery enable more robust and interpretable behavior than open‑loop execution, while adding minimal architectural overhead. For the source code and demo refer to our project page: https://wenzewwz123.github.io/Agentic‑Loop/
Authors:Diego Gomez, Antoine Guédon, Nissim Maruani, Bingchen Gong, Maks Ovsjanikov
Abstract:
3D Gaussian Splatting (3DGS) has revolutionized fast novel view synthesis, yet its opacity‑based formulation makes surface extraction fundamentally difficult. Unlike implicit methods built on Signed Distance Fields or occupancy, 3DGS lacks a global geometric field, forcing existing approaches to resort to heuristics such as TSDF fusion of blended depth maps.
Inspired by the Objects as Volumes framework, we derive a principled occupancy field for Gaussian Splatting and show how it can be used to extract highly accurate watertight meshes of complex scenes. Our key contribution is to introduce a learnable oriented normal at each Gaussian element and to define an adapted attenuation formulation, which leads to closed‑form expressions for both the normal and occupancy fields at arbitrary locations in space. We further introduce a novel consistency loss and a dedicated densification strategy to enforce Gaussians to wrap the entire surface by closing geometric holes, ensuring a complete shell of oriented primitives. We modify the differentiable rasterizer to output depth as an isosurface of our continuous model, and introduce Primal Adaptive Meshing for Region‑of‑Interest meshing at arbitrary resolution.
We additionally expose fundamental biases in standard surface evaluation protocols and propose two more rigorous alternatives. Overall, our method Gaussian Wrapping sets a new state‑of‑the‑art on DTU and Tanks and Temples, producing complete, watertight meshes at a fraction of the size of concurrent work‑recovering thin structures such as the notoriously elusive bicycle spokes.
Authors:Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Kai Wang, Zheng Wang, Peng Hu, Xi Peng, Hongyuan Zhu
Abstract:
Existing dynamic data pruning methods often fail under noisy‑label settings, as they typically rely on per‑sample loss as the ranking criterion. This could mistakenly lead to preserving noisy samples due to their high loss values, resulting in significant performance drop. To address this, we propose AlignPrune, a noise‑robust module designed to enhance the reliability of dynamic pruning under label noise. Specifically, AlignPrune introduces the Dynamic Alignment Score (DAS), which is a loss‑trajectory‑based criterion that enables more accurate identification of noisy samples, thereby improving pruning effectiveness. As a simple yet effective plug‑and‑play module, AlignPrune can be seamlessly integrated into state‑of‑the‑art dynamic pruning frameworks, consistently outperforming them without modifying either the model architecture or the training pipeline. Extensive experiments on five widely‑used benchmarks across various noise types and pruning ratios demonstrate the effectiveness of AlignPrune, boosting accuracy by up to 6.3% over state‑of‑the‑art baselines. Our results offer a generalizable solution for pruning under noisy data, encouraging further exploration of learning in real‑world scenarios. Code is available at: https://github.com/leonqin430/AlignPrune.
Authors:Changkun Liu, Jiezhi Yang, Zeman Li, Yuan Deng, Jiancong Guo, Luca Ballan
Abstract:
Streaming 3D perception is well suited to robotics and augmented reality, where long visual streams must be processed efficiently and consistently. Recent recurrent models offer a promising solution by maintaining fixed‑size states and enabling linear‑time inference, but they often suffer from drift accumulation and temporal forgetting over long sequences due to the limited capacity of compressed latent memories. We propose Mem3R, a streaming 3D reconstruction model with a hybrid memory design that decouples camera tracking from geometric mapping to improve temporal consistency over long sequences. For camera tracking, Mem3R employs an implicit fast‑weight memory implemented as a lightweight Multi‑Layer Perceptron updated via Test‑Time Training. For geometric mapping, Mem3R maintains an explicit token‑based fixed‑size state. Compared with CUT3R, this design not only significantly improves long‑sequence performance but also reduces the model size from 793M to 644M parameters. Mem3R supports existing improved plug‑and‑play state update strategies developed for CUT3R. Specifically, integrating it with TTT3R decreases Absolute Trajectory Error by up to 39% over the base implementation on 500 to 1000 frame sequences. The resulting improvements also extend to other downstream tasks, including video depth estimation and 3D reconstruction, while preserving constant GPU memory usage and comparable inference throughput. Project page: https://lck666666.github.io/Mem3R/
Authors:Ruihang Xu, Dewei Zhou, Xiaolong Shen, Fan Ma, Yi Yang
Abstract:
Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However, existing visual generative models often fail at precise spatial manipulation, resulting in incorrect scaling and positioning of objects. This limitation primarily stems from the lack of explicit mechanisms to incorporate 3D geometry and perspective projection. To achieve accurate manipulation, we develop PhyEdit, an image editing framework that leverages explicit geometric simulation as contextual 3D‑aware visual guidance. By combining this plug‑and‑play 3D prior with joint 2D‑‑3D supervision, our method effectively improves physical accuracy and manipulation consistency. To support this method and evaluate performance, we present a real‑world dataset, RealManip‑10K, for 3D‑aware object manipulation featuring paired images and depth annotations. We also propose ManipEval, a benchmark with multi‑dimensional metrics to evaluate 3D spatial control and geometric consistency. Extensive experiments show that our approach outperforms existing methods, including strong closed‑source models, in both 3D geometric accuracy and manipulation consistency.
Authors:Mohamed Darwish Mounis, Mohamed Mahmoud, Shaimaa Sedek, Mahmoud Abdalla, Mahmoud SalahEldin Kasem, Abdelrahman Abdallah, Hyun-Soo Kang
Abstract:
Multimodal retrieval systems struggle to resolve image‑text queries against text‑only corpora: the best vision‑language encoder achieves only 27.6 nDCG@10 on MM‑BRIGHT, underperforming strong text‑only retrievers. We argue the bottleneck is not the retriever but the query ‑‑ raw multimodal queries entangle visual descriptions, conversational noise, and retrieval intent in ways that systematically degrade embedding similarity. We present BRIDGE, a two‑component system that resolves this mismatch without multimodal encoders. FORGE (Focused Retrieval Query Generator) is a query alignment model trained via reinforcement learning, which distills noisy multimodal queries into compact, retrieval‑optimized search strings. LENS (Language‑Enhanced Neural Search) is a reasoning‑enhanced dense retriever fine‑tuned on reasoning‑intensive retrieval data to handle the intent‑rich queries FORGE produces. Evaluated on MM‑BRIGHT (2,803 queries, 29 domains), BRIDGE achieves 29.7 nDCG@10, surpassing all multimodal encoder baselines including Nomic‑Vision (27.6). When FORGE is applied as a plug‑and‑play aligner on top of Nomic‑Vision, the combined system reaches 33.3 nDCG@10 ‑‑ exceeding the best text‑only retriever (32.2) ‑‑ demonstrating that query alignment is the key bottleneck in multimodal‑to‑text retrieval. https://github.com/mm‑bright/multimodal‑reasoning‑retrieval
Authors:Kartikay Tehlan, Lukas Förner, Sina Wendrich, Nico Schmutzenhofer, Michael Frühwald, Matthias Wagner, Nassir Navab, Thomas Wendler
Abstract:
We propose a geometric framework for longitudinal multi‑parametric MRI analysis based on patient‑specific energy modelling in sequence space. Rather than operating on images with spatial networks, each voxel is represented by its multi‑sequence intensity vector (T1, T1c, T2, FLAIR, ADC), and a compact implicit neural representation is trained via denoising score matching to learn an energy function E_θ(\mathbfu) over \mathbbR^d from a single baseline scan. The learned energy landscape provides a differential‑geometric description of tissue regimes without segmentation labels. Local minima define tissue basins, gradient magnitude reflects proximity to regime boundaries, and Laplacian curvature characterises local constraint structure. Importantly, this baseline energy manifold is treated as a fixed geometric reference: it encodes the set of contrast combinations observed at diagnosis and is not retrained at follow‑up. Longitudinal assessment is therefore formulated as evaluation of subsequent scans relative to this baseline geometry. Rather than comparing anatomical segmentations, we analyse how the distribution of MRI sequence vectors evolves under the baseline energy function. In a paediatric case with later recurrence, follow‑up scans show progressive deviation in energy and directional displacement in sequence space toward the baseline tumour‑associated regime before clear radiological reappearance. In a case with stable disease, voxel distributions remain confined to established low‑energy basins without systematic drift. The presented cases serve as proof‑of‑concept that patient‑specific energy manifolds can function as geometric reference systems for longitudinal mpMRI analysis without explicit segmentation or supervised classification, providing a foundation for further investigation of manifold‑based tissue‑at‑risk tracking in neuro‑oncology.
Authors:Sonja Adomeit, Kartikay Tehlan, Lukas Förner, Katharina Weisser, Helen Scholtiseek, David Kaufmann, Julie Steinestel, Constantin Lapa, Thomas Kröncke, Thomas Wendler
Abstract:
Multimodal imaging analysis often relies on joint latent representations, yet these approaches rarely define what information is shared versus modality‑specific. Clarifying this distinction is clinically relevant, as it delineates the irreducible contribution of each modality and informs rational acquisition strategies. We propose a subspace decomposition framework that reframes multimodal fusion as a problem of orthogonal subspace separation rather than translation. We decompose Prostate‑Specific Membrane Antigen (PSMA) PET uptake into an MRI‑explainable physiological envelope and an orthogonal residual reflecting signal components not expressible within the MRI feature manifold. Using multiparametric MRI, we train an intensity‑based, non‑spatial implicit neural representation (INR) to map MRI feature vectors to PET uptake. We introduce a projection‑based regularization using singular value decomposition to penalize residual components lying within the span of the MRI feature manifold. This enforces mathematical orthogonality between tissue‑level physiological properties (structure, diffusion, perfusion) and intracellular PSMA expression. Tested on 13 prostate cancer patients, the model demonstrates that residual components spanned by MRI features are absorbed into the learned envelope, while the orthogonal residual is largest in tumour regions. This indicates that PSMA PET contains signal components not recoverable from MRI‑derived physiological descriptors. The resulting decomposition provides a structured characterization of modality complementarity grounded in representation geometry rather than image translation.
Authors:Wei Zhang, Vincent Ress, David Skuddis, Uwe Soergel, Norbert Haala
Abstract:
RTK‑SLAM systems integrate simultaneous localization and mapping (SLAM) with real‑time kinematic (RTK) GNSS positioning, promising both relative consistency and globally referenced coordinates for efficient georeferenced surveying. A critical and underappreciated issue is that the standard evaluation metric, Absolute Trajectory Error (ATE), first fits an optimal rigid‑body transformation between the estimated trajectory and reference before computing errors. This so‑called SE(3) alignment absorbs global drift and systematic errors, making trajectories appear more accurate than they are in practice, and is unsuitable for evaluating the global accuracy of RTK‑SLAM. We present a geodetically referenced dataset and evaluation methodology that expose this gap. A key design principle is that the RTK receiver is used solely as a system input, while ground truth is established independently via a geodetic total station. This separation is absent from all existing datasets, where GNSS typically serves as (part of) the ground truth. The dataset is collected with a handheld RTK‑SLAM device, comprising two scenes. We evaluate LiDAR‑inertial, visual‑inertial, and LiDAR‑visual‑inertial RTK‑SLAM systems alongside standalone RTK, reporting direct global accuracy and SE(3)‑aligned relative accuracy to make the gap explicit. Results show that SE(3) alignment can underestimate absolute positioning error by up to 76%. RTK‑SLAM achieves centimeter‑level absolute accuracy in open‑sky conditions and maintains decimeter‑level global accuracy indoors, where standalone RTK degrades to tens of meters. The dataset, calibration files, and evaluation scripts are publicly available at https://rtk‑slam‑dataset.github.io/.
Authors:Changmiao Wang, Songqi Zhang, Yongquan Zhang, Yifei Wang, Liya Liu, Nannan Li, Xingzhi Li, Jiexin Pan, Yi Jiang, Xiang Wan, Hai Wang, Ahmed Elazab
Abstract:
Kidney stone disease ranks among the most prevalent conditions in urology, and understanding the composition of these stones is essential for creating personalized treatment plans and preventing recurrence. Current methods for analyzing kidney stones depend on postoperative specimens, which prevents rapid classification before surgery. To overcome this limitation, we introduce a new approach called the Urinary Stone Segmentation and Classification Network (USCNet). This innovative method allows for precise preoperative classification of kidney stones by integrating Computed Tomography (CT) images with clinical data from Electronic Health Records (EHR). USCNet employs a Transformer‑based multimodal fusion framework with CT‑EHR attention and segmentation‑guided attention modules for accurate classification. Moreover, a dynamic loss function is introduced to effectively balance the dual objectives of segmentation and classification. Experiments on an in‑house kidney stone dataset show that USCNet demonstrates outstanding performance across all evaluation metrics, with its classification efficacy significantly surpassing existing mainstream methods. This study presents a promising solution for the precise preoperative classification of kidney stones, offering substantial clinical benefits. The source code has been made publicly available: https://github.com/ZhangSongqi0506/KidneyStone.
Authors:Reiji Saito, Satoshi Kamiya, Kazuhiro Hotta
Abstract:
In conventional anomaly detection, training data consist of only normal samples. However, in real‑world scenarios, the definition of a normal sample is often ambiguous. For example, there are cases where a sample has small scratches or stains but is still acceptable for practical usage. On the other hand, higher precision is required when manufacturing equipment is upgraded. In such cases, normal samples may include small scratches, tiny dust particles, or a foreign object that we would prefer to classify as an anomaly. Such cases frequently occur in industrial settings, yet they have not been discussed until now. Thus, we propose novel scenarios and an evaluation metric to accommodate specification changes in real‑world applications. Furthermore, to address the ambiguity of normal samples, we propose the RePaste, which enhances learning by re‑pasting regions with high anomaly scores from the previous step into the input for the next step. On our scenarios using the MVTec AD benchmark, RePaste achieved the state‑of‑the‑art performance with respect to the proposed evaluation metric, while maintaining high AUROC and PRO scores. Code: https://github.com/ReijiSoftmaxSaito/Scenario
Authors:Mojgan Madadikhaljan, Jonathan Prexl, Isabelle Wittmann, Conrad M Albrecht, Michael Schmitt
Abstract:
In this work, we present LIANet (Location Is All You Need Network), a coordinate‑based neural representation that models multi‑temporal spaceborne Earth observation (EO) data for a given region of interest as a continuous spatiotemporal neural field. Given only spatial and temporal coordinates, LIANet reconstructs the corresponding satellite imagery. Once pretrained, this neural representation can be adapted to various EO downstream tasks, such as semantic segmentation or pixel‑wise regression, importantly, without requiring access to the original satellite data. LIANet intends to serve as a user‑friendly alternative to Geospatial Foundation Models (GFMs) by eliminating the overhead of data access and preprocessing for end‑users and enabling fine‑tuning solely based on labels. We demonstrate the pretraining of LIANet across target areas of varying sizes and show that fine‑tuning it for downstream tasks achieves competitive performance compared to training from scratch or using established GFMs. The source code and datasets are publicly available at https://github.com/mojganmadadi/LIANet/tree/v1.0.1.
Authors:Mehdi Hosseinzadeh, King Hang Wong, Feras Dayoub
Abstract:
We present KITE, a training‑free, keyframe‑anchored, layout‑grounded front‑end that converts long robot‑execution videos into compact, interpretable tokenized evidence for vision‑language models (VLMs). KITE distills each trajectory into a small set of motion‑salient keyframes with open‑vocabulary detections and pairs each keyframe with a schematic bird's‑eye‑view (BEV) representation that encodes relative object layout, axes, timestamps, and detection confidence. These visual cues are serialized with robot‑profile and scene‑context tokens into a unified prompt, allowing the same front‑end to support failure detection, identification, localization, explanation, and correction with an off‑the‑shelf VLM. On the RoboFAC benchmark, KITE with Qwen2.5‑VL substantially improves over vanilla Qwen2.5‑VL in the training‑free setting, with especially large gains on simulation failure detection, identification, and localization, while remaining competitive with a RoboFAC‑tuned baseline. A small QLoRA fine‑tune further improves explanation and correction quality. We also report qualitative results on real dual‑arm robots, demonstrating the practical applicability of KITE as a structured and interpretable front‑end for robot failure analysis. Code and models are released on our project page: https://m80hz.github.io/kite/
Authors:Qingze He, Fagui Liu, Dengke Zhang, Qingmao Wei, Quan Tang
Abstract:
Weakly supervised semantic segmentation aims to achieve pixel‑level predictions using image‑level labels. Existing methods typically entangle semantic recognition and object localization, which often leads models to focus exclusively on sparse discriminative regions. Although foundation models show immense potential, many approaches still follow the tightly coupled optimization paradigm, struggling to effectively alleviate pseudo‑label noise and often relying on time‑consuming multi‑stage retraining or unstable end‑to‑end joint optimization. To address the above challenges, we present ModuSeg, a training‑free weakly supervised semantic segmentation framework centered on explicitly decoupling object discovery and semantic assignment. Specifically, we integrate a general mask proposer to extract geometric proposals with reliable boundaries, while leveraging semantic foundation models to construct an offline feature bank, transforming segmentation into a non‑parametric feature retrieval process. Furthermore, we propose semantic boundary purification and soft‑masked feature aggregation strategies to effectively mitigate boundary ambiguity and quantization errors, thereby extracting high‑quality category prototypes. Extensive experiments demonstrate that the proposed decoupled architecture better preserves fine boundaries without parameter fine‑tuning and achieves highly competitive performance on standard benchmark datasets. Code is available at https://github.com/Autumnair007/ModuSeg.
Authors:Jaeyoung Chung, Hyunjin Son, Kyoung Mu Lee
Abstract:
We present the first generative approach to photomosaic creation. Traditional photomosaic methods rely on a large number of tile images and color‑based matching, which limits both diversity and structural consistency. Our generative photomosaic framework synthesizes tile images using diffusion‑based generation conditioned on reference images. A low‑frequency conditioned diffusion mechanism aligns global structure while preserving prompt‑driven details. This generative formulation enables photomosaic composition that is both semantically expressive and structurally coherent, effectively overcoming the fundamental limitations of matching‑based approaches. By leveraging few‑shot personalized diffusion, our model is able to produce user‑specific or stylistically consistent tiles without requiring an extensive collection of images.
Authors:Renyang Liu, Jiale Li, Jie Zhang, Cong Wu, Xiaojun Jia, Shuxin Li, Wei Zhou, Kwok-Yan Lam, See-kiong Ng
Abstract:
Palmprint recognition is deployed in security‑critical applications, including access control and palm‑based payment, due to its contactless acquisition and highly discriminative ridge‑and‑crease textures. However, the robustness of deep palmprint recognition systems against physically realizable attacks remains insufficiently understood. Existing studies are largely confined to the digital setting and do not adequately account for the texture‑dominant nature of palmprint recognition or the distortions introduced during physical acquisition. To address this gap, we propose CAAP, a capture‑aware adversarial patch framework for palmprint recognition. CAAP learns a universal patch that can be reused across inputs while remaining effective under realistic acquisition variation. To match the structural characteristics of palmprints, the framework adopts a cross‑shaped patch topology, which enlarges spatial coverage under a fixed pixel budget and more effectively disrupts long‑range texture continuity. CAAP further integrates three modules: ASIT for input‑conditioned patch rendering, RaS for stochastic capture‑aware simulation, and MS‑DIFE for feature‑level identity‑disruptive guidance. We evaluate CAAP on the Tongji, IITD, and AISEC datasets against generic CNN backbones and palmprint‑specific recognition models. Experiments show that CAAP achieves strong untargeted and targeted attack performance with favorable cross‑model and cross‑dataset transferability. The results further show that, although adversarial training can partially reduce the attack success rate, substantial residual vulnerability remains. These findings indicate that deep palmprint recognition systems remain vulnerable to physically realizable, capture‑aware adversarial patch attacks, underscoring the need for more effective defenses in practice. Code available at https://github.com/ryliu68/CAAP.
Authors:Xiaoxiao Ma, Jiachen Lei, Tianfei Ren, Jie Huang, Siming Fu, Aiming Hao, Jiahong Wu, Xiangxiang Chu, Feng Zhao
Abstract:
Reinforcement learning (RL) has been successfully applied to autoregressive (AR) and diffusion models. However, extending RL to hybrid AR‑diffusion frameworks remains challenging due to interleaved inference and noisy log‑probability estimation. In this work, we study masked autoregressive models (MAR) and show that the diffusion head plays a critical role in training dynamics, often introducing noisy gradients that lead to instability and early performance saturation. To address this issue, we propose a stabilized RL framework for MAR. We introduce multi‑trajectory expectation (MTE), which estimates the optimization direction by averaging over multiple diffusion trajectories, thereby reducing diffusion‑induced gradient noise. To avoid over‑smoothing, we further estimate token‑wise uncertainty from multiple trajectories and apply multi‑trajectory optimization only to the top‑k% uncertain tokens. In addition, we introduce a consistency‑aware token selection strategy that filters out AR tokens that are less aligned with the final generated content. Extensive experiments across multiple benchmarks demonstrate that our method consistently improves visual quality, training stability, and spatial structure understanding over baseline GRPO and pre‑RL models. Code is available at: https://github.com/AMAP‑ML/mar‑grpo.
Authors:Zhiheng Li, Zongyang Ma, Yuntong Pan, Ziqi Zhang, Xiaolei Lv, Bo Li, Jun Gao, Jianing Zhang, Chunfeng Yuan, Bing Li, Weiming Hu
Abstract:
Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critical threat: Adversarial Smuggling Attacks. Unlike adversarial perturbations (for misclassification) and adversarial jailbreaks (for harmful output generation), adversarial smuggling exploits the Human‑AI capability gap. It encodes harmful content into human‑readable visual formats that remain AI‑unreadable, thereby evading automated detection and enabling the dissemination of harmful content. We classify smuggling attacks into two pathways: (1) Perceptual Blindness, disrupting text recognition; and (2) Reasoning Blockade, inhibiting semantic understanding despite successful text recognition. To evaluate this threat, we constructed SmuggleBench, the first comprehensive benchmark comprising 1,700 adversarial smuggling attack instances. Evaluations on SmuggleBench reveal that both proprietary (e.g., GPT‑5) and open‑source (e.g., Qwen3‑VL) state‑of‑the‑art models are vulnerable to this threat, producing Attack Success Rates (ASR) exceeding 90%. By analyzing the vulnerability through the lenses of perception and reasoning, we identify three root causes: the limited capabilities of vision encoders, the robustness gap in OCR, and the scarcity of domain‑specific adversarial examples. We conduct a preliminary exploration of mitigation strategies, investigating the potential of test‑time scaling (via CoT) and adversarial training (via SFT) to mitigate this threat. Our code is publicly available at https://github.com/zhihengli‑casia/smugglebench.
Authors:Jiyun Won, Heemin Yang, Woohyeok Kim, Jungseul Ok, Sunghyun Cho
Abstract:
Recent work has explored optimizing image signal processing (ISP) pipelines for various tasks by composing predefined modules and adapting them to task‑specific objectives. However, jointly optimizing module sequences and parameters remains challenging. Existing approaches rely on neural architecture search (NAS) or step‑wise reinforcement learning (RL), but NAS suffers from a training‑inference mismatch, while step‑wise RL leads to unstable training and high computational overhead due to stage‑wise decision‑making. We propose POS‑ISP, a sequence‑level RL framework that formulates modular ISP optimization as a global sequence prediction problem. Our method predicts the entire module sequence and its parameters in a single forward pass and optimizes the pipeline using a terminal task reward, eliminating the need for intermediate supervision and redundant executions. Experiments across multiple downstream tasks show that POS‑ISP improves task performance while reducing computational cost, highlighting sequence‑level optimization as a stable and efficient paradigm for task‑aware ISP. The project page is available at https://w1jyun.github.io/POS‑ISP
Authors:Yuheng Shi, Xiaohuan Pei, Linfeng Wen, Minjing Dong, Chang Xu
Abstract:
MLLMs require high‑resolution visual inputs for fine‑grained tasks like document understanding and dense scene perception. However, current global resolution scaling paradigms indiscriminately flood the quadratic self‑attention mechanism with visually redundant tokens, severely bottlenecking inference throughput while ignoring spatial sparsity and query intent. To overcome this, we propose Q‑Zoom, a query‑aware adaptive high‑resolution perception framework that operates in an efficient coarse‑to‑fine manner. First, a lightweight Dynamic Gating Network safely bypasses high‑resolution processing when coarse global features suffice. Second, for queries demanding fine‑grained perception, a Self‑Distilled Region Proposal Network (SD‑RPN) precisely localizes the task‑relevant Region‑of‑Interest (RoI) directly from intermediate feature spaces. To optimize these modules efficiently, the gating network uses a consistency‑aware generation strategy to derive deterministic routing labels, while the SD‑RPN employs a fully self‑supervised distillation paradigm. A continuous spatio‑temporal alignment scheme and targeted fine‑tuning then seamlessly fuse the dense local RoI with the coarse global layout. Extensive experiments demonstrate that Q‑Zoom establishes a dominant Pareto frontier. Using Qwen2.5‑VL‑7B as a primary testbed, Q‑Zoom accelerates inference by 2.52 times on Document & OCR benchmarks and 4.39 times in High‑Resolution scenarios while matching the baseline's peak accuracy. Furthermore, when configured for maximum perceptual fidelity, Q‑Zoom surpasses the baseline's peak performance by 1.1% and 8.1% on these respective benchmarks. These robust improvements transfer seamlessly to Qwen3‑VL, LLaVA, and emerging RL‑based thinking‑with‑image models. Project page is available at https://yuhengsss.github.io/Q‑Zoom/.
Authors:Dewei Zhou, You Li, Zongxin Yang, Yi Yang
Abstract:
We introduce region‑specific image refinement as a dedicated problem setting: given an input image and a user‑specified region (e.g., a scribble mask or a bounding box), the goal is to restore fine‑grained details while keeping all non‑edited pixels strictly unchanged. Despite rapid progress in image generation, modern models still frequently suffer from local detail collapse (e.g., distorted text, logos, and thin structures). Existing instruction‑driven editing models emphasize coarse‑grained semantic edits and often either overlook subtle local defects or inadvertently change the background, especially when the region of interest occupies only a small portion of a fixed‑resolution input. We present RefineAnything, a multimodal diffusion‑based refinement model that supports both reference‑based and reference‑free refinement. Building on a counter‑intuitive observation that crop‑and‑resize can substantially improve local reconstruction under a fixed VAE input resolution, we propose Focus‑and‑Refine, a region‑focused refinement‑and‑paste‑back strategy that improves refinement effectiveness and efficiency by reallocating the resolution budget to the target region, while a blended‑mask paste‑back guarantees strict background preservation. We further introduce a boundary‑aware Boundary Consistency Loss to reduce seam artifacts and improve paste‑back naturalness. To support this new setting, we construct Refine‑30K (20K reference‑based and 10K reference‑free samples) and introduce RefineEval, a benchmark that evaluates both edited‑region fidelity and background consistency. On RefineEval, RefineAnything achieves strong improvements over competitive baselines and near‑perfect background preservation, establishing a practical solution for high‑precision local refinement. Project Page: https://limuloo.github.io/RefineAnything/.
Authors:Jiajun Yang, Keyan Chen, Zhengxia Zou, Zhenwei Shi
Abstract:
Cloud detection in remote sensing imagery is a fundamental, critical, and highly challenging problem. Existing deep learning‑based cloud detection methods generally formulate it as a single‑stage pixel‑wise binary segmentation task with one forward pass. However, such single‑stage approaches exhibit ambiguity and uncertainty in thin‑cloud regions and struggle to accurately handle fragmented clouds and boundary details. In this paper, we propose a novel deep learning framework termed CloudMamba. To address the ambiguity in thin‑cloud regions, we introduce an uncertainty‑guided two‑stage cloud detection strategy. An embedded uncertainty estimation module is proposed to automatically quantify the confidence of thin‑cloud segmentation, and a second‑stage refinement segmentation is introduced to improve the accuracy in low‑confidence hard regions. To better handle fragmented clouds and fine‑grained boundary details, we design a dual‑scale Mamba network based on a CNN‑Mamba hybrid architecture. Compared with Transformer‑based models with quadratic computational complexity, the proposed method maintains linear computational complexity while effectively capturing both large‑scale structural characteristics and small‑scale boundary details of clouds, enabling accurate delineation of overall cloud morphology and precise boundary segmentation. Extensive experiments conducted on the GF1_WHU and Levir_CS public datasets demonstrate that the proposed method outperforms existing approaches across multiple segmentation accuracy metrics, while offering high efficiency and process transparency. Our code is available at https://github.com/jayoungo/CloudMamba.
Authors:Subin Park, Jung Uk Kim
Abstract:
Sound source localization task aims to identify the locations of sound‑emitting objects by leveraging correlations between audio and visual modalities. Most existing SSL methods rely on contrastive learning‑based feature matching, but lack explicit reasoning and verification, limiting their effectiveness in complex acoustic scenes. Inspired by human meta‑cognitive processes, we propose a training‑free SSL framework that exploits the intrinsic reasoning capabilities of Multimodal Large Language Models (MLLMs). Our Generation‑Analysis‑Refinement (GAR) pipeline consists of three stages: Generation produces initial bounding boxes and audio classifications; Analysis quantifies Audio‑Visual Consistency via open‑set role tagging and anchor voting; and Refinement applies adaptive gating to prevent unnecessary adjustments. Extensive experiments on single‑source and multi‑source benchmarks demonstrate competitive performance. The source code is available at https://github.com/VisualAIKHU/GAR‑SSL.
Authors:Huy Q. Le, Loc X. Nguyen, Yu Qiao, Seong Tae Kim, Eui-Nam Huh, Choong Seon Hong
Abstract:
Federated Learning (FL) enables decentralized model training across multiple clients without exposing private data, making it ideal for privacy‑sensitive applications. However, in real‑world FL scenarios, clients often hold data from distinct domains, leading to severe domain shift and degraded global model performance. To address this, prototype learning has been emerged as a promising solution, which leverages class‑wise feature representations. Yet, existing methods face two key limitations: (1) Existing prototype‑based FL methods typically construct a single global prototype per class by aggregating local prototypes from all clients without preserving domain information. (2) Current feature‑prototype alignment is domain‑agnostic, forcing clients to align with global prototypes regardless of domain origin. To address these challenges, we propose Federated Domain‑Aware Prototypes (FedDAP) to construct domain‑specific global prototypes by aggregating local client prototypes within the same domain using a similarity‑weighted fusion mechanism. These global domain‑specific prototypes are then used to guide local training by aligning local features with prototypes from the same domain, while encouraging separation from prototypes of different domains. This dual alignment enhances domain‑specific learning at the local level and enables the global model to generalize across diverse domains. Finally, we conduct extensive experiments on three different datasets: DomainNet, Office‑10, and PACS to demonstrate the effectiveness of our proposed framework to address the domain shift challenges. The code is available at https://github.com/quanghuy6997/FedDAP.
Authors:Bohao Xing, Deng Li, Rong Gao, Xin Liu, Heikki Kälviäinen
Abstract:
Recently, Transformer has made significant progress in various vision tasks. To balance computation and efficiency in video tasks, recent works heavily rely on factorized or window‑based self‑attention. However, these approaches split spatiotemporal correlations between regions of interest in videos, limiting the models' ability to capture motion and long‑range dependencies. In this paper, we argue that, similar to the human visual system, the importance of temporal and spatial information varies across different time scales, and attention is allocated sparsely over time through glance and gaze behavior. Is equal consideration of time and space crucial for success in video tasks? Motivated by this understanding, we propose a dual‑path network called the Overall Glance and Refined Gaze (OG‑ReG) Transformer. The Glance path extracts coarse‑grained overall spatiotemporal information, while the Gaze path supplements the Glance path by providing local details. Our model achieves state‑of‑the‑art results on the Kinetics‑400, Something‑Something v2, and Diving‑48, demonstrating its competitive performance. The code will be available at https://github.com/linuxsino/OG‑ReG.
Authors:Guillermo Gil de Avalle, Laura Maruster, Eric Sloot, Christos Emmanouilidis
Abstract:
Maintenance procedures in manufacturing facilities are often documented as flowcharts in static PDFs or scanned images. They encode procedural knowledge essential for asset lifecycle management, yet inaccessible to modern operator support systems. Vision‑language models, the dominant paradigm for image understanding, struggle to reconstruct connection topology from such diagrams. We present FlowExtract, a pipeline for extracting directed graphs from ISO 5807‑standardized flowcharts. The system separates element detection from connectivity reconstruction, using YOLOv8 and EasyOCR for standard domain‑aligned node detection and text extraction, combined with a novel edge detection method that analyzes arrowhead orientations and traces connecting lines backward to source nodes. Evaluated on industrial troubleshooting guides, FlowExtract achieves very high node detection and substantially outperforms vision‑language model baselines on edge extraction, offering organizations a practical path toward queryable procedural knowledge representations. The implementation is available athttps://github.com/guille‑gil/FlowExtract.
Authors:Junchao Yi, Rui Zhao, Jiahao Tang, Weixian Lei, Linjie Li, Qisheng Su, Zhengyuan Yang, Lijuan Wang, Xiaofeng Zhu, Alex Jinpeng Wang
Abstract:
Multimodal generation has long been dominated by text‑driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descriptions, spatial layouts, and editing instructions, can be unified into a single visual representation. We present FlowInOne, a framework that reformulates multimodal generation as a purely visual flow, converting all inputs into visual prompts and enabling a clean image‑in, image‑out pipeline governed by a single flow matching model. This vision‑centric formulation naturally eliminates cross‑modal alignment bottlenecks, noise scheduling, and task‑specific architectural branches, unifying text‑to‑image generation, layout‑guided editing, and visual instruction following under one coherent paradigm. To support this, we introduce VisPrompt‑5M, a large‑scale dataset of 5 million visual prompt pairs spanning diverse tasks including physics‑aware force dynamics and trajectory prediction, alongside VP‑Bench, a rigorously curated benchmark assessing instruction faithfulness, spatial precision, visual realism, and content consistency. Extensive experiments demonstrate that FlowInOne achieves state‑of‑the‑art performance across all unified generation tasks, surpassing both open‑source models and competitive commercial systems, establishing a new foundation for fully vision‑centric generative modeling where perception and creation coexist within a single continuous visual space. Our code and models are released on https://csu‑jpg.github.io/FlowInOne.github.io/
Authors:Roberto Brusnicki, Mattia Piccinini, Johannes Betz
Abstract:
Vision‑Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poorly characterized, particularly regarding how input configurations affect their capabilities. We introduce VENUSS (VLM Evaluation oN Understanding Sequential Scenes), a framework for systematic sensitivity analysis of VLM performance on sequential driving scenes, establishing baselines for future research. Building upon existing datasets, VENUSS extracts temporal sequences from driving videos, and generates structured evaluations across custom categories. By comparing 25+ existing VLMs across 2,600+ scenarios, we reveal how even top models achieve only 57% accuracy, not matching human performance under similar constraints (65%) and exposing significant capability gaps. Our analysis shows that VLMs excel with static object detection but struggle with understanding vehicle dynamics and temporal relations. VENUSS offers the first systematic sensitivity analysis of VLMs focused on how input image configurations ‑ resolution, frame count, temporal intervals, spatial layouts, and presentation modes ‑ affect performance on sequential driving scenes. Supplementary material available at https://TUM‑AVS.github.io/VENUSS/.
Authors:Pedro Quesado, Erkut Akdag, Yasaman Kashefbahrami, Willem Menu, Egor Bondarev
Abstract:
Live‑streaming Novel View Synthesis (NVS) from unposed multi‑view video remains an open challenge in a wide range of applications. Existing methods for dynamic scene representation typically require ground‑truth camera parameters and involve lengthy optimizations (\approx 2.67s), which makes them unsuitable for live streaming scenarios. To address this issue, we propose a novel viewpoint video live‑streaming method (LiveStre4m), a feed‑forward model for real‑time NVS from unposed sparse multi‑view inputs. LiveStre4m introduces a multi‑view vision transformer for keyframe 3D scene reconstruction coupled with a diffusion‑transformer interpolation module that ensures temporal consistency and stable streaming. In addition, a Camera Pose Predictor module is proposed to efficiently estimate both poses and intrinsics directly from RGB images, removing the reliance on known camera calibration information. Our approach enables temporally consistent novel‑view video streaming in real‑time using as few as two synchronized unposed input streams. LiveStre4m attains an average reconstruction time of 0.07s per‑frame at 1024 × 768 resolution, outperforming the optimization‑based dynamic scene representation methods by orders of magnitude in runtime. These results demonstrate that LiveStre4m makes real‑time NVS streaming feasible in practical settings, marking a substantial step toward deployable live novel‑view synthesis systems. Code available at: https://github.com/pedro‑quesado/LiveStre4m
Authors:Ethan Nguyen, Javier Carmona, Arisa Matsuzaki, Naoki Kaneko, Katsushi Arisaka
Abstract:
Introduction: Mechanical thrombectomy can cause vessel deformation and procedure‑related injury. Benchtop models are widely used for device testing, but time‑resolved, full‑field 3D vessel‑motion measurements remain limited.
Methods: We developed a nine‑camera, low‑cost multi‑view workflow for benchtop thrombectomy in silicone middle cerebral artery phantoms (2160p, 20 fps). Multi‑view videos were calibrated, segmented, and reconstructed with 4D Gaussian Splatting. Reconstructed point clouds were converted to fixed‑connectivity edge graphs for region‑of‑interest (ROI) displacement tracking and a relative surface‑based stress proxy. Stress‑proxy values were derived from edge stretch using a Neo‑Hookean mapping and reported as comparative surface metrics. A synthetic Blender pipeline with known deformation provided geometric and temporal validation.
Results: In synthetic bulk translation, the stress proxy remained near zero for most edges (median \approx 0 MPa; 90th percentile 0.028 MPa), with sparse outliers. In synthetic pulling (1‑5 mm), reconstruction showed close geometric and temporal agreement with ground truth, with symmetric Chamfer distance of 1.714‑1.815 mm and precision of 0.964‑0.972 at τ= 1 mm. In preliminary benchtop comparative trials (one trial per condition), cervical aspiration catheter placement showed higher max‑median ROI displacement and stress‑proxy values than internal carotid artery terminus placement.
Conclusion: The proposed protocol provides standardized, time‑resolved surface kinematics and comparative relative displacement and stress proxy measurements for thrombectomy benchtop studies. The framework supports condition‑to‑condition comparisons and methods validation, while remaining distinct from absolute wall‑stress estimation. Implementation code and example data are available at https://ethanuser.github.io/vessel4D
Authors:Daewon Yoon, Injun Baek, Sangyu Han, Yearim Kim, Nojun Kwak
Abstract:
Video depth estimation is essential for providing 3D scene structure in applications ranging from autonomous driving to mixed reality. Current end‑to‑end video depth models have established state‑of‑the‑art performance. Although current end‑to‑end (E2E) models have achieved state‑of‑the‑art performance, they function as tightly coupled systems that suffer from a significant adaptation lag whenever superior single‑image depth estimators are released. To mitigate this issue, post‑processing methods such as NVDS offer a modular plug‑and‑play alternative to incorporate any evolving image depth model without retraining. However, existing post‑processing methods still struggle to match the efficiency and practicality of E2E systems due to limited speed, accuracy, and RGB reliance. In this work, we revitalize the role of post‑processing by proposing VDPP (Video Depth Post‑Processing), a framework that improves the speed and accuracy of post‑processing methods for video depth estimation. By shifting the paradigm from computationally expensive scene reconstruction to targeted geometric refinement, VDPP operates purely on geometric refinements in low‑resolution space. This design achieves exceptional speed (>43.5 FPS on NVIDIA Jetson Orin Nano) while matching the temporal coherence of E2E systems, with dense residual learning driving geometric representations rather than full reconstructions. Furthermore, our VDPP's RGB‑free architecture ensures true scalability, enabling immediate integration with any evolving image depth model. Our results demonstrate that VDPP provides a superior balance of speed, accuracy, and memory efficiency, making it the most practical solution for real‑time edge deployment. Our project page is at https://github.com/injun‑baek/VDPP
Authors:Weikai Qu, Sijun Liang, Cheng Pan, Zikuan Yang, Guanchi Zhou, Xianjun Fu, Bo Liu, Changmiao Wang, Ahmed Elazab
Abstract:
Photographs taken in adverse weather conditions often suffer from blurriness, occlusion, and low brightness due to interference from rain, snow, and fog. These issues can significantly hinder the performance of subsequent computer vision tasks, making the removal of weather effects a crucial step in image enhancement. Existing methods primarily target specific weather conditions, with only a few capable of handling multiple weather scenarios. However, mainstream approaches often overlook performance considerations, resulting in large parameter sizes, long inference times, and high memory costs. In this study, we introduce the WeatherRemover model, designed to enhance the restoration of images affected by various weather conditions while balancing performance. Our model adopts a UNet‑like structure with a gating mechanism and a multi‑scale pyramid vision Transformer. It employs channel‑wise attention derived from convolutional neural networks to optimize feature extraction, while linear spatial reduction helps curtail the computational demands of attention. The gating mechanisms, strategically placed within the feed‑forward and downsampling phases, refine the processing of information by selectively addressing redundancy and mitigating its influence on learning. This approach facilitates the adaptive selection of essential data, ensuring superior restoration and maximizing efficiency. Additionally, our lightweight model achieves an optimal balance between restoration quality, parameter efficiency, computational overhead, and memory usage, distinguishing it from other multi‑weather models, thereby meeting practical application demands effectively. The source code is available at https://github.com/RICKand‑MORTY/WeatherRemover.
Authors:Weikai Qu, Sijun Liang, Xianfeng Li, Cheng Pan, An Yan, Ahmed Elazab, Shanzhou Niu, Dong Zeng, Xiang Wan, Changmiao Wang
Abstract:
In computed tomography imaging, metal implants frequently generate severe artifacts that compromise image quality and hinder diagnostic accuracy. There are three main challenges in the existing methods: the deterioration of organ and tissue structures, dependence on sinogram data, and an imbalance between resource use and restoration efficiency. Addressing these issues, we introduce MARMamba, which effectively eliminates artifacts caused by metals of different sizes while maintaining the integrity of the original anatomical structures of the image. Furthermore, this model only focuses on CT images affected by metal artifacts, thus negating the requirement for additional input data. The model is a streamlined UNet architecture, which incorporates multi‑scale Mamba (MS‑Mamba) as its core module. Within MS‑Mamba, a flip mamba block captures comprehensive contextual information by analyzing images from multiple orientations. Subsequently, the average maximum feed‑forward network integrates critical features with average features to suppress the artifacts. This combination allows MARMamba to reduce artifacts efficiently. The experimental results demonstrate that our model excels in reducing metal artifacts, offering distinct advantages over other models. It also strikes an optimal balance between computational demands, memory usage, and the number of parameters, highlighting its practical utility in the real world. The code of the presented model is available at: https://github.com/RICKand‑MORTY/MARMamba.
Authors:Yiyang Li, Yanbo Gao, Shuai Li, Zhenyu Du, Jinglin Zhang, Hui Yuan, Mao Ye, Xingyu Gao
Abstract:
Implicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN‑based multi‑scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN‑INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found at https://github.com/yiyang‑sdu/CWRNN‑INVR.githttps://github.com/yiyang‑sdu/CWRNN‑INVR.git.
Authors:Tomas Guija-Valiente, Iago Suárez
Abstract:
AI‑driven content generation has made remarkable progress in recent years. However, neural networks and human designers operate in fundamentally different ways, making collaboration between them challenging. We address this gap for Scalable Vector Graphics (SVG) by equipping neural networks with tools commonly used by designers, such as axis alignment and explicit continuity control at command junctions. We introduce DesigNet, a hierarchical Transformer‑VAE that operates directly on SVG sequences with a continuous command parameterization. Our main contributions are two differentiable modules: a continuity self‑refinement module that predicts C^0, G^1, and C^1 continuity for each curve point and enforces it by modifying Bézier control points, and an alignment self‑refinement module with snapping capabilities for horizontal or vertical lines. DesigNet produces editable outlines and achieves competitive results against state‑of‑the‑art methods, with notably higher accuracy in continuity and alignment. These properties ensure the outputs are easier to refine and integrate into professional design workflows. Source Code: https://github.com/TomasGuija/DesigNet.
Authors:Berna Kabadayi, Vanessa Sklyarova, Wojciech Zielonka, Justus Thies, Gerard Pons-Moll
Abstract:
Realistic digital avatars require expressive and dynamic hair motion; however, most existing head avatar methods assume rigid hair movement. These methods often fail to disentangle hair from the head, representing it as a simple outer shell and failing to capture its natural volumetric behavior. In this paper, we address these limitations by introducing PhysHead, a hybrid representation for animatable head avatars with realistic hair dynamics learned from multi‑view video. At the core is a 3D Gaussian‑based layered representation of the head. Our approach combines a 3D parametric mesh for the head with strand‑based hair, which can be directly simulated using physics engines. For the appearance model, we employ Gaussian primitives attached to both the head mesh and hair segments. This representation enables the creation of photorealistic head avatars with dynamic hair behavior, such as wind‑blown motion, overcoming the constraints of rigid hair in existing methods. However, these animation capabilities also require new training schemes. In particular, we propose the use of VLM‑based models to generate appearance of regions that are occluded in the dynamic training sequences. In quantitative and qualitative studies, we demonstrate the capabilities of the proposed model and compare it with existing baselines. We show that our method can synthesize physically plausible hair motion besides expression and camera control.
Authors:Teng Hu, Jiangning Zhang, Hongrui Huang, Ran Yi, Zihan Su, Jieyu Weng, Zhucun Xue, Lizhuang Ma, Ming-Hsuan Yang, Dacheng Tao
Abstract:
The rapid advancement of Artificial Intelligence Generated Content (AIGC) has revolutionized video generation, enabling systems ranging from proprietary pioneers like OpenAI's Sora, Google's Veo3, and Bytedance's Seedance to powerful open‑source contenders like Wan and HunyuanVideo to synthesize temporally coherent and semantically rich videos. These advancements pave the way for building "world models" that simulate real‑world dynamics, with applications spanning entertainment, education, and virtual reality. However, existing reviews on video generation often focus on narrow technical fields, e.g., Generative Adversarial Networks (GAN) and diffusion models, or specific tasks (e. g., video editing), lacking a comprehensive perspective on the field's evolution, especially regarding Auto‑Regressive (AR) models and integration of multimodal information. To address these gaps, this survey firstly provides a systematic review of the development of video generation technology, tracing its evolution from early GANs to dominant diffusion models, and further to emerging AR‑based and multimodal techniques. We conduct an in‑depth analysis of the foundational principles, key advancements, and comparative strengths/limitations. Then, we explore emerging trends in multimodal video generation, emphasizing the integration of diverse data types to enhance contextual awareness. Finally, by bridging historical developments and contemporary innovations, this survey offers insights to guide future research in video generation and its applications, including virtual/augmented reality, personalized education, autonomous driving simulations, digital entertainment, and advanced world models, in this rapidly evolving field. For more details, please refer to the project at https://github.com/sjtuplayer/Awesome‑Video‑Foundations.
Authors:Ashmal Vayani, Parth Parag Kulkarni, Joseph Fioresi, Song Wang, Mubarak Shah
Abstract:
Medical diagnosis using Large Multimodal Models (LMMs) has gained increasing attention due to capability of these models in providing precise diagnoses. These models generally combine medical questions with visual inputs to generate diagnoses or treatments. However, they are often overly general and unsuitable under the wide range of medical conditions in real‑world healthcare. In clinical practice, diagnosis is performed by multiple specialists, each contributing domain‑specific expertise. To emulate this process, a potential solution is to deploy a dynamic multi‑agent LMM framework, where each agent functions as a medical specialist. Current approaches in this emerging area, typically relying on static or predefined selection of various specialists, cannot be adapted to the changing practical scenario. In this paper, we propose MedRoute, a flexible and dynamic multi‑agent framework that comprises of a collaborative system of specialist LMM agents. Furthermore, we add a General Practitioner with an RL‑trained router for dynamic specialist selection, and a Moderator that produces the final decision. In this way, our framework closely mirrors real clinical workflows. Extensive evaluations on text and image‑based medical datasets demonstrate improved diagnostic accuracy, outperforming the state‑of‑the‑art baselines. Our work lays a strong foundation for future research. Code and models are available at https://github.com/UCF‑CRCV/MedRoute/.
Authors:David Picard, Nicolas Dufour, Lucas Degeorge, Arijit Ghosh, Davide Allegro, Tom Ravaud, Yohann Perron, Corentin Sautier, Zeynep Sonat Baltaci, Fei Meng, Syrine Kalleli, Marta López-Rauhut, Thibaut Loiseau, Ségolène Albouy, Raphael Baena, Elliot Vincent, Loic Landrieu
Abstract:
This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop‑in replacement for self‑attention. PoM aggregates input tokens into a compact representation through a learned polynomial function, from which each token retrieves contextual information. We prove that PoM satisfies the contextual mapping property, ensuring that transformers equipped with PoM remain universal sequence‑to‑sequence approximators. We replace standard self‑attention with PoM across five diverse domains: text generation, handwritten text recognition, image generation, 3D modeling, and Earth observation. PoM matches the performance of attention‑based models while drastically reducing computational cost when working with long sequences. The code is available at https://github.com/davidpicard/pom.
Authors:Junbin Zhang, Meng Cao, Feng Tan, Yikai Lin, Yuexian Zou
Abstract:
Achieving fine‑grained and structurally sound controllability is a cornerstone of advanced visual generation. Existing part‑based frameworks treat user‑provided parts as an unordered set and therefore ignore their intrinsic spatial and semantic relationships, which often results in compositions that lack structural integrity. To bridge this gap, we propose Graph‑PiT, a framework that explicitly models the structural dependencies of visual components using a graph prior. Specifically, we represent visual parts as nodes and their spatial‑semantic relationships as edges. At the heart of our method is a Hierarchical Graph Neural Network (HGNN) module that performs bidirectional message passing between coarse‑grained part‑level super‑nodes and fine‑grained IP+ token sub‑nodes, refining part embeddings before they enter the generative pipeline. We also introduce a graph Laplacian smoothness loss and an edge‑reconstruction loss so that adjacent parts acquire compatible, relation‑aware embeddings. Quantitative experiments on controlled synthetic domains (character, product, indoor layout, and jigsaw), together with qualitative transfer to real web images, show that Graph‑PiT improves structural coherence over vanilla PiT while remaining compatible with the original IP‑Prior pipeline. Ablation experiments confirm that explicit relational reasoning is crucial for enforcing user‑specified adjacency constraints. Our approach not only enhances the plausibility of generated concepts but also offers a scalable and interpretable mechanism for complex, multi‑part image synthesis. The code is available at https://github.com/wolf‑bailang/Graph‑PiT.
Authors:Tao Hu, Varun Jampani
Abstract:
Despite tremendous recent progress in human video generation, generative video diffusion models still struggle to capture the dynamics and physics of human motions faithfully. In this paper, we propose a new framework for human video generation, HumANDiff, which enhances the human motion control with three key designs: 1) Articulated motion‑consistent noise sampling that correlates the spatiotemporal distribution of latent noise and replaces the unstructured random Gaussian noise with 3D articulated noise sampled on the dense surface manifold of a statistical human body template. It inherits body topology priors for spatially and temporally consistent noise sampling. 2) Joint appearance‑motion learning that enhances the standard training objective of video diffusion models by jointly predicting pixel appearances and corresponding physical motions from the articulated noises. It enables high‑fidelity human video synthesis, e.g., capturing motion‑dependent clothing wrinkles. 3) Geometric motion consistency learning that enforces physical motion consistency across frames via a novel geometric motion consistency loss defined in the articulated noise space. HumANDiff enables scalable controllable human video generation by fine‑tuning video diffusion models with articulated noise sampling. Consequently, our method is agnostic to diffusion model design, and requires no modifications to the model architecture. During inference, HumANDiff enables image‑to‑video generation within a single framework, achieving intrinsic motion control without requiring additional motion modules. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance in rendering motion‑consistent, high‑fidelity humans with diverse clothing styles. Project page: https://taohuumd.github.io/projects/HumANDiff/
Authors:Ioannis Nasios
Abstract:
Landslides represent a major geohazard with severe impacts on human life, infrastructure, and ecosystems, underscoring the need for accurate and timely detection approaches to support disaster risk reduction. This study proposes a modular, multi‑model framework that fuses Sentinel‑2 optical imagery with Sentinel‑1 Synthetic Aperture Radar (SAR) data, for robust landslide detection. The methodology leverages multi‑encoder vision transformers, where each data modality is processed through separate lightweight pretrained encoders, achieving strong performance in landslide detection. In addition, the integration of multiple models, particularly the combination of neural networks and gradient boosting models (LightGBM and XGBoost), demonstrates the power of ensemble learning to further enhance accuracy and robustness. Derived spectral indices, such as NDVI, are integrated alongside original bands to enhance sensitivity to vegetation and surface changes. The proposed methodology achieves a state‑of‑the‑art F1 score of 0.919 on landslide detection, addressing a patch‑based classification task rather than pixel‑level segmentation and operating without pre‑event Sentinel‑2 data, highlighting its effectiveness in a non‑classical change detection setting. It also demonstrated top performance in a machine learning competition, achieving a strong balance between precision and recall and highlighting the advantages of explicitly leveraging the complementary strengths of optical and radar data. The conducted experiments and research also emphasize scalability and operational applicability, enabling flexible configurations with optical‑only, SAR‑only, or combined inputs, and offering a transferable framework for broader natural hazard monitoring and environmental change applications. Full training and inference code can be found in https://github.com/IoannisNasios/sentinel‑landslide‑cls.
Authors:Ahmet Rasim Emirdagi, Süleyman Aslan, Mısra Yavuz, Görkay Aydemir, Yunus Bilge Kurt, Nasrin Rahimi, Burak Can Biner, M. Akın Yılmaz
Abstract:
Metal artifacts from high‑attenuation implants severely degrade CT image quality, obscuring critical anatomical structures and posing a challenge for standard deep learning methods that require extensive paired training data. We propose a paradigm shift: reframing artifact reduction as an in‑context reasoning task by adapting a general‑purpose vision‑language diffusion foundation model via parameter‑efficient Low‑Rank Adaptation (LoRA). By leveraging rich visual priors, our approach achieves effective artifact suppression with only 16 to 128 paired training examples reducing data requirements by two orders of magnitude. Crucially, we demonstrate that domain adaptation is essential for hallucination mitigation; without it, foundation models interpret streak artifacts as erroneous natural objects (e.g., waffles or petri dishes). To ground the restoration, we propose a multi‑reference conditioning strategy where clean anatomical exemplars from unrelated subjects are provided alongside the corrupted input, enabling the model to exploit category‑specific context to infer uncorrupted anatomy. Extensive evaluation on the AAPM CT‑MAR benchmark demonstrates that our method achieves state‑of‑the‑art performance on perceptual and radiological‑feature metrics . This work establishes that foundation models, when appropriately adapted, offer a scalable alternative for interpretable, data‑efficient medical image reconstruction. Code is available at https://github.com/ahmetemirdagi/CT‑EditMAR.
Authors:Yangyi Xiao, Siting Zhu, Baoquan Yang, Tianchen Deng, Yongbo Chen, Hesheng Wang
Abstract:
Multi‑traversal scene reconstruction is important for high‑fidelity autonomous driving simulation and digital twin construction. This task involves integrating multiple sequences captured from the same geographical area at different times. In this context, a primary challenge is the significant appearance inconsistency across traversals caused by varying illumination and environmental conditions, despite the shared underlying geometry. This paper presents ADM‑GS (Appearance Decomposition Gaussian Splatting for Multi‑Traversal Reconstruction), a framework that applies an explicit appearance decomposition to the static background to alleviate appearance entanglement across traversals. For the static background, we decompose the appearance into traversal‑invariant material, representing intrinsic material properties, and traversal‑dependent illumination, capturing lighting variations. Specifically, we propose a neural light field that utilizes a frequency‑separated hybrid encoding strategy. By incorporating surface normals and explicit reflection vectors, this design separately captures low‑frequency diffuse illumination and high‑frequency specular reflections. Quantitative evaluations on the Argoverse 2 and Waymo Open datasets demonstrate the effectiveness of ADM‑GS. In multi‑traversal experiments, our method achieves a +0.98 dB PSNR improvement over existing latent‑based baselines while producing more consistent appearance across traversals. Code will be available at https://github.com/IRMVLab/ADM‑GS.
Authors:Yingjian Zhu, Xinming Wang, Kun Ding, Ying Wang, Bin Fan, Shiming Xiang
Abstract:
Multi‑modal Retrieval‑Augmented Generation (RAG) has emerged as a highly effective paradigm for Knowledge‑Based Visual Question Answering (KB‑VQA). Despite recent advancements, prevailing methods still primarily depend on images as the retrieval key, and often overlook or misplace the role of Vision‑Language Models (VLMs), thereby failing to leverage their potential fully. In this paper, we introduce WikiSeeker, a novel multi‑modal RAG framework that bridges these gaps by proposing a multi‑modal retriever and redefining the role of VLMs. Rather than serving merely as answer generators, we assign VLMs two specialized agents: a Refiner and an Inspector. The Refiner utilizes the capability of VLMs to rewrite the textual query according to the input image, significantly improving the performance of the multimodal retriever. The Inspector facilitates a decoupled generation strategy by selectively routing reliable retrieved context to another LLM for answer generation, while relying on the VLM's internal knowledge when retrieval is unreliable. Extensive experiments on EVQA, InfoSeek, and M2KR demonstrate that WikiSeeker achieves state‑of‑the‑art performance, with substantial improvements in both retrieval accuracy and answer quality. Our code will be released on https://github.com/zhuyjan/WikiSeeker.
Authors:Bo Ma, Jinsong Wu, Weiqi Yan
Abstract:
In LLM/VLM agents, prompt privacy risk propagates beyond a single model call because raw user content can flow into retrieval queries, memory writes, tool calls, and logs. Existing de‑identification pipelines address document boundaries but not this cross‑stage propagation. We propose BodhiPromptShield, a policy‑aware framework that detects sensitive spans, routes them via typed placeholders, semantic abstraction, or secure symbolic mapping, and delays restoration to authorized boundaries. Relative to enterprise redaction, this adds explicit propagation‑aware mediation and restoration timing as a security variable. Under controlled evaluation on the Controlled Prompt‑Privacy Benchmark (CPPB), stage‑wise propagation suppresses from 10.7% to 7.1% across retrieval, memory, and tool stages; PER reaches 9.3% with 0.94 AC and 0.92 TSR, outperforming generic de‑identification. These are controlled systems results on CPPB rather than formal privacy guarantees or public‑benchmark transfer claims.
The project repository is available at https://github.com/mabo1215/BodhiPromptShield.git.
Authors:Amadou S. Sangare, Adrien Maglo, Mohamed Chaouch, Bertrand Luvison
Abstract:
Text‑to‑Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisely control image layouts, something that natural language alone cannot reliably express. Controllable generation methods augment the initial T2I model with additional conditions that more easily describe the scene. Prior works straightforwardly train the augmented network with the same loss as the initial network. Although natural at first glance, this can lead to very long training times in some cases before convergence. In this work, we revisit the training objective of controllable diffusion models through a detailed analysis of their denoising dynamics. We show that direct supervision on the clean target image, dubbed x_0‑supervision, or an equivalent re‑weighting of the diffusion loss, yields faster convergence. Experiments on multiple control settings demonstrate that our formulation accelerates convergence by up to 2× according to our novel metric (mean Area Under the Convergence Curve ‑ mAUCC), while also improving both visual quality and conditioning accuracy. Our code is available at https://github.com/CEA‑LIST/x0‑supervision
Authors:Dongliang Zhu, Zhiyi Niu, Bo Zhao, Jiajian Huang, Shuo Ye, Xun Lin, Hui Ma, Taorui Wang, Jiayu Zhang, Chunmei Zhu, Junzhe Cao, Yingjie Ma, Rencheng Song, Albert Clapés, Sergio Escalera, Dan Guo, Zitong Yu
Abstract:
Subtle visual signals, although difficult to perceive with the naked eye, contain important information that can reveal hidden patterns in visual data. These signals play a key role in many applications, including biometric security, multimedia forensics, medical diagnosis, industrial inspection, and affective computing. With the rapid development of computer vision and representation learning techniques, detecting and interpreting such subtle signals has become an emerging research direction. However, existing studies often focus on specific tasks or modalities, and models still face challenges in robustness, representation ability, and generalization when handling subtle and weak signals in real‑world environments. To promote research in this area, we organize the Subtle visual Challenge, which aims to learn robust representations for subtle visual signals. The challenge includes two tasks: cross‑domain multimodal deception detection and remote photoplethysmography (rPPG) estimation. We hope that this challenge will encourage the development of more robust and generalizable models for subtle visual understanding, and further advance research in computer vision and multimodal learning. A total of 22 teams submitted their final results to this workshop competition, and the corresponding baseline models have been released on the \hrefhttps://sites.google.com/view/svc‑cvpr26MMDD2026 platform\footnotehttps://sites.google.com/view/svc‑cvpr26
Authors:Mengtian Li, Kunyan Dai, Yi Ding, Ruobing Ni, Ying Zhang, Wenwu Wang, Zhifeng Xie
Abstract:
Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio‑temporally aligned audio remains labor‑intensive. We propose FoleyDesigner, a novel framework inspired by professional Foley workflows, integrating film clip analysis, spatio‑temporally controllable Foley generation, and professional audio mixing capabilities. FoleyDesigner employs a multi‑agent architecture for precise spatio‑temporal analysis. It achieves spatio‑temporal alignment through latent diffusion models trained on spatio‑temporal cues extracted from video frames, combined with large language model (LLM)‑driven hybrid mechanisms that emulate post‑production practices in film industry. To address the lack of high‑quality stereo audio datasets in film, we introduce FilmStereo, the first professional stereo audio dataset containing spatial metadata, precise timestamps, and semantic annotations for eight common Foley categories. For applications, the framework supports interactive user control while maintaining seamless integration with professional pipelines, including 5.1‑channel Dolby Atmos systems compliant with ITU‑R BS.775 standards, thereby offering extensive creative flexibility. Extensive experiments demonstrate that our method achieves superior spatio‑temporal alignment compared to existing baselines, with seamless compatibility with professional film production standards. The project page is available at https://gekiii996.github.io/FoleyDesigner/ .
Authors:Weiqi Zhang, Junsheng Zhou, Haotian Geng, Kanle Shi, Shenkun Xu, Yi Fang, Yu-Shen Liu
Abstract:
3D Gaussian Splatting has demonstrated superior performance in rendering efficiency and quality, yet the generation of 3D Gaussians still remains a challenge without proper geometric priors. Existing methods have explored predicting point maps as geometric references for inferring Gaussian primitives, while the unreliable estimated geometries may lead to poor generations. In this work, we introduce GaussianGrow, a novel approach that generates 3D Gaussians by learning to grow them from easily accessible 3D point clouds, naturally enforcing geometric accuracy in Gaussian generation. Specifically, we design a text‑guided Gaussian growing scheme that leverages a multi‑view diffusion model to synthesize consistent appearances from input point clouds for supervision. To mitigate artifacts caused by fusing neighboring views, we constrain novel views generated at non‑preset camera poses identified in overlapping regions across different views. For completing the hard‑to‑observe regions, we propose to iteratively detect the camera pose by observing the largest un‑grown regions in point clouds and inpainting them by inpainting the rendered view with a pretrained 2D diffusion model. The process continues until complete Gaussians are generated. We extensively evaluate GaussianGrow on text‑guided Gaussian generation from synthetic and even real‑scanned point clouds. Project Page: https://weiqi‑zhang.github.io/GaussianGrow
Authors:Xuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li, Xichao Teng
Abstract:
We present Consistent‑Recurrent Feature Flow Transformer (CRFT), a unified coarse‑to‑fine framework based on feature flow learning for robust cross‑modal image registration. CRFT learns a modality‑independent feature flow representation within a transformer‑based architecture that jointly performs feature alignment and flow estimation. The coarse stage establishes global correspondences through multi‑scale feature correlation, while the fine stage refines local details via hierarchical feature fusion and adaptive spatial reasoning. To enhance geometric adaptability, an iterative discrepancy‑guided attention mechanism with a Spatial Geometric Transform (SGT) recurrently refines the flow field, progressively capturing subtle spatial inconsistencies and enforcing feature‑level consistency. This design enables accurate alignment under large affine and scale variations while maintaining structural coherence across modalities. Extensive experiments on diverse cross‑modal datasets demonstrate that CRFT consistently outperforms state‑of‑the‑art registration methods in both accuracy and robustness. Beyond registration, CRFT provides a generalizable paradigm for multimodal spatial correspondence, offering broad applicability to remote sensing, autonomous navigation, and medical imaging. Code and datasets are publicly available at https://github.com/NEU‑Liuxuecong/CRFT.
Authors:Yongchuan Cui, Peng Liu
Abstract:
Remote sensing imagery suffers from clouds, haze, noise, resolution limits, and sensor heterogeneity. Existing restoration and fusion approaches train separate models per degradation type. In this work, we present Language‑conditioned Large‑scale Remote Sensing restoration model (LLaRS), the first unified foundation model for multi‑modal and multi‑task remote sensing low‑level vision. LLaRS employs Sinkhorn‑Knopp optimal transport to align heterogeneous bands into semantically matched slots, routes features through three complementary mixture‑of‑experts layers (convolutional experts for spatial patterns, channel‑mixing experts for spectral fidelity, and attention experts with low‑rank adapters for global context), and stabilizes joint training via step‑level dynamic weight adjustment. To train LLaRS, we construct LLaRS1M, a million‑scale multi‑task dataset spanning eleven restoration and enhancement tasks, integrating real paired observations and controlled synthetic degradations with diverse natural language prompts. Experiments show LLaRS consistently outperforms seven competitive models, and parameter‑efficient finetuning experiments demonstrate strong transfer capability and adaptation efficiency on unseen data. Repo: https://github.com/yc‑cui/LLaRS
Authors:Xinran Wang, Yuxuan Zhang, Xiao Zhang, Haolong Yan, Muxi Diao, Songyu Xu, Zhonghao Yan, Hongbing Li, Kongming Liang, Zhanyu Ma
Abstract:
Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response‑level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high‑quality images across five distinct domains. With an average caption length of over 200 words and dense, token‑level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx‑hhnkh.github.io/DetailVerifyBench/.
Authors:Rongfei Chen, Tingting Zhang, Xiaoyu Shen, Wei Zhang
Abstract:
The missing modality problem poses a fundamental challenge in multimodal sentiment analysis, significantly degrading model accuracy and generalization in real world scenarios. Existing approaches primarily improve robustness through prompt learning and pre trained models. However, two limitations remain. First, the necessity of generating missing modalities lacks rigorous evaluation. Second, the structural dependencies among multimodal prompts and their global coherence are insufficiently explored. To address these issues, a Prompt based Missing Modality Adaptation framework is proposed. A Missing Modality Evaluator is introduced at the input stage to dynamically assess the importance of missing modalities using pretrained models and pseudo labels, thereby avoiding low quality data imputation. Building on this, a Modality invariant Prompt Disentanglement module decomposes shared prompts into modality specific private prompts to capture intrinsic local correlations and improve representation quality. In addition, a Dynamic Prompt Weighting module computes mutual information based weights from cross attention outputs to adaptively suppress interference from missing modalities. To enhance global consistency, a Multi level Prompt Dynamic Connection module integrates shared prompts with self attention outputs through residual connections, leveraging global prompt priors to strengthen key guidance features. Extensive experiments on three public benchmarks, including CMU MOSI, CMU MOSEI, and CH SIMS, demonstrate that the proposed framework achieves state of the art performance and stable results under diverse missing modality settings. The implementation is available at https://github.com/rongfei‑chen/ProMMA
Authors:Xuanguang Liu, Lei Ding, Yujie Li, Chenguang Dai, Zhenchao Zhang, Mengmeng Li, Ziyi Yang, Yifan Sun, Yongqi Sun, Hanyun Wang
Abstract:
Multimodal change detection (MMCD) identifies changed areas in multimodal remote sensing (RS) data, demonstrating significant application value in land use monitoring, disaster assessment, and urban sustainable development. However, literature MMCD approaches exhibit limitations in cross‑modal interaction and exploiting modality‑specific characteristics. This leads to insufficient modeling of fine‑grained change information, thus hindering the precise detection of semantic changes in multimodal data. To address the above problems, we propose STSF‑Net, a framework designed for MMCD between optical and SAR images. STSF‑Net jointly models modality‑specific and spatio‑temporal common features to enhance change representations. Specifically, modality‑specific features are exploited to capture genuine semantic change signals, while spatio‑temporal common features are embedded to suppress pseudo‑changes caused by differences in imaging mechanisms. Furthermore, we introduce an optical and SAR feature fusion strategy that adaptively adjusts feature importance based on semantic priors obtained from pre‑trained foundational models, enabling semantic‑guided adaptive fusion of multi‑modal information. In addition, we introduce the Delta‑SN6 dataset, the first openly‑accessible multiclass MMCD benchmark consisting of very‑high‑resolution (VHR) fully polarimetric SAR and optical images. Experimental results on Delta‑SN6, BRIGHT, and Wuhan‑Het datasets demonstrate that our method outperforms the state‑of‑the‑art (SOTA) by 3.21%, 1.08%, and 1.32% in mIoU, respectively. The associated code and Delta‑SN6 dataset will be released at: https://github.com/liuxuanguang/STSF‑Net.
Authors:Li Kang, Yutao Fan, Rui Li, Heng Zhou, Yiran Qin, Zhemeng Zhang, Songtao Huang, Xiufeng Song, Zaibin Zhang, Bruno N. Y. Chen, Zhenfei Yin, Dongzhan Zhou, Wangmeng Zuo, Lei Bai
Abstract:
Multi‑agent embodied systems hold promise for complex collaborative manipulation, yet face critical challenges in spatial coordination, temporal reasoning, and shared workspace awareness. Inspired by human collaboration where cognitive planning occurs separately from physical execution, we introduce the concept of compositional environment ‑‑ a synergistic integration of real‑world and simulation components that enables multiple robotic agents to perceive intentions and operate within a unified decision‑making space. Building on this concept, we present CoEnv, a framework that leverages simulation for safe strategy exploration while ensuring reliable real‑world deployment. CoEnv operates through three stages: real‑to‑sim scene reconstruction that digitizes physical workspaces, VLM‑driven action synthesis supporting both real‑time planning with high‑level interfaces and iterative planning with code‑based trajectory generation, and validated sim‑to‑real transfer with collision detection for safe deployment. Extensive experiments on challenging multi‑arm manipulation benchmarks demonstrate CoEnv's effectiveness in achieving high task success rates and execution efficiency, establishing a new paradigm for multi‑agent embodied AI.
Authors:Pu Wang, Zhixuan Mao, Jialu Li, Zhuoran Zheng, Dianjie Lu, Youshan Zhang
Abstract:
Automatic diagnosis of canine pneumothorax is challenged by data scarcity and the need for trustworthy models. To address this, we first introduce a public, pixel‑level annotated dataset to facilitate research. We then propose a novel diagnostic paradigm that reframes the task as a synergistic process of signal localization and spectral detection. For localization, our method employs a Vision‑Language Model (VLM) to guide an iterative Flow Matching process, which progressively refines segmentation masks to achieve superior boundary accuracy. For detection, the segmented mask is used to isolate features from the suspected lesion. We then apply Random Matrix Theory (RMT), a departure from traditional classifiers, to analyze these features. This approach models healthy tissue as predictable random noise and identifies pneumothorax by detecting statistically significant outlier eigenvalues that represent a non‑random pathological signal. The high‑fidelity localization from Flow Matching is crucial for purifying the signal, thus maximizing the sensitivity of our RMT detector. This synergy of generative segmentation and first‑principles statistical analysis yields a highly accurate and interpretable diagnostic system (source code is available at: https://github.com/Pu‑Wang‑alt/Canine‑pneumothorax).
Authors:Gwanghyun Kim, Junghun James Kim, Suh Yoon Jeon, Jason Park, Se Young Chun
Abstract:
Reconstructing textured 3D human models from a single image is fundamental for AR/VR and digital human applications. However, existing methods mostly focus on single individuals and thus fail in multi‑human scenes, where naive composition of individual reconstructions often leads to artifacts such as unrealistic overlaps, missing geometry in occluded regions, and distorted interactions. These limitations highlight the need for approaches that incorporate group‑level context and interaction priors. We introduce a holistic method that explicitly models both group‑ and instance‑level information. To mitigate perspective‑induced geometric distortions, we first transform the input into a canonical orthographic space. Our primary component, Human Group‑Instance Multi‑View Diffusion (HUG‑MVD), then generates complete multi‑view normals and images by jointly modeling individuals and group context to resolve occlusions and proximity. Subsequently, the Human Group‑Instance Geometric Reconstruction (HUG‑GR) module optimizes the geometry by leveraging explicit, physics‑based interaction priors to enforce physical plausibility and accurately model inter‑human contact. Finally, the multi‑view images are fused into a high‑fidelity texture. Together, these components form our complete framework, HUG3D. Extensive experiments show that HUG3D significantly outperforms both single‑human and existing multi‑human methods, producing physically plausible, high‑fidelity 3D reconstructions of interacting people from a single image. Project page: https://jongheean11.github.io/HUG3D_project
Authors:Yi-Jen Tsai, Yen-Yu Lin, Chien-Yao Wang
Abstract:
Few‑Shot Semantic Segmentation (FSS) focuses on segmenting novel object categories from only a handful of annotated examples. Most existing approaches rely on extensive episodic training to learn transferable representations, which is both computationally demanding and sensitive to distribution shifts. In this work, we revisit FSS from the perspective of modern vision foundation models and explore the potential of Segment Anything Model 3 (SAM3) as a training‑free solution. By repurposing its Promptable Concept Segmentation (PCS) capability, we adopt a simple spatial concatenation strategy that places support and query images into a shared canvas, allowing a fully frozen SAM3 to perform segmentation without any fine‑tuning or architectural changes. Experiments on PASCAL‑5^i and COCO‑20^i show that this minimal design already achieves state‑of‑the‑art performance, outperforming many heavily engineered methods. Beyond empirical gains, we uncover that negative prompts can be counterproductive in few‑shot settings, where they often weaken target representations and lead to prediction collapse despite their intended role in suppressing distractors. These findings suggest that strong cross‑image reasoning can emerge from simple spatial formulations, while also highlighting limitations in how current foundation models handle conflicting prompt signals. Code at: https://github.com/WongKinYiu/FSS‑SAM3
Authors:Honghao Fu, Miao Xu, Yiwei Wang, Dailing Zhang, Jun Liu, Yujun Cai
Abstract:
Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval‑augmented generation (RAG) is a promising remedy by organizing query‑relevant visual evidence into a compact context, most existing methods (i) flatten videos into independent segments, breaking their inherent spatio‑temporal structure, and (ii) depend on explicit semantic matching, which can miss cues that are implicitly relevant to the query's intent. To overcome these limitations, we propose VideoStir, a structured and intent‑aware long‑video RAG framework. It firstly structures a video as a spatio‑temporal graph at clip level, and then performs multi‑hop retrieval to aggregate evidence across distant yet contextually related events. Furthermore, it introduces an MLLM‑backed intent‑relevance scorer that retrieves frames based on their alignment with the query's reasoning intent. To support this capability, we curate IR‑600K, a large‑scale dataset tailored for learning frame‑query intent alignment. Experiments show that VideoStir is competitive with state‑of‑the‑art baselines without relying on auxiliary information, highlighting the promise of shifting long‑video RAG from flattened semantic matching to structured, intent‑aware reasoning. Codes and checkpoints are available at https://github.com/RomGai/VideoStir.
Authors:Xiang Zhang, Tengfei Wang, Fang Xu, Xin Wang, Zongqian Zhan
Abstract:
Visual localization in large‑scale UAV scenarios is a critical capability for autonomous systems, yet it remains challenging due to geometric complexity and environmental variations. While 3D Gaussian Splatting (3DGS) has emerged as a promising scene representation, existing 3DGS‑based visual localization methods struggle with robust pose initialization and sensitivity to rendering artifacts in large‑scale settings. To address these limitations, we propose LSGS‑Loc, a novel visual localization pipeline tailored for large‑scale 3DGS scenes. Specifically, we introduce a scale‑aware pose initialization strategy that combines scene‑agnostic relative pose estimation with explicit 3DGS scale constraints, enabling geometrically grounded localization without scene‑specific training. Furthermore, in the pose refinement, to mitigate the impact of reconstruction artifacts such as blur and floaters, we develop a Laplacian‑based reliability masking mechanism that guides photometric refinement toward high‑quality regions. Extensive experiments on large‑scale UAV benchmarks demonstrate that our method achieves state‑of‑the‑art accuracy and robustness for unordered image queries, significantly outperforming existing 3DGS‑based approaches. Code is available at: https://github.com/xzhang‑z/LSGS‑Loc
Authors:Yuxin Yang, Yinan Zhou, Yuxin Chen, Ziqi Zhang, Zongyang Ma, Chunfeng Yuan, Bing Li, Jun Gao, Weiming Hu
Abstract:
Composed Image Retrieval (CIR) has demonstrated significant potential by enabling flexible multimodal queries that combine a reference image and modification text. However, CIR inherently prioritizes semantic matching, struggling to reliably retrieve a user‑specified instance across contexts. In practice, emphasizing concrete instance fidelity over broad semantics is often more consequential. In this work, we propose Object‑Anchored Composed Image Retrieval (OACIR), a novel fine‑grained retrieval task that mandates strict instance‑level consistency. To advance research on this task, we construct OACIRR (OACIR on Real‑world images), the first large‑scale, multi‑domain benchmark comprising over 160K quadruples and four challenging candidate galleries enriched with hard‑negative instance distractors. Each quadruple augments the compositional query with a bounding box that visually anchors the object in the reference image, providing a precise and flexible way to ensure instance preservation. To address the OACIR task, we propose AdaFocal, a framework featuring a Context‑Aware Attention Modulator that adaptively intensifies attention within the specified instance region, dynamically balancing focus between the anchored instance and the broader compositional context. Extensive experiments demonstrate that AdaFocal substantially outperforms existing compositional retrieval models, particularly in maintaining instance‑level fidelity, thereby establishing a robust baseline for this challenging task while opening new directions for more flexible, instance‑aware retrieval systems.
Authors:Jae Joong Lee
Abstract:
Every existing method for compressing 3D Gaussian Splatting, NeRF, or transformer‑based 3D reconstructors requires learning a data‑dependent codebook through per‑scene fine‑tuning. We show this is unnecessary. The parameter vectors that dominate storage in these models, 45‑dimensional spherical harmonics in 3DGS and 1024‑dimensional key‑value vectors in DUSt3R, fall in a dimension range where a single random rotation transforms any input into coordinates with a known Beta distribution. This makes precomputed, data‑independent Lloyd‑Max quantization near‑optimal, within a factor of 2.7 of the information‑theoretic lower bound. We develop 3D, deriving (1) a dimension‑dependent criterion that predicts which parameters can be quantized and at what bit‑width before running any experiment, (2) norm‑separation bounds connecting quantization MSE to rendering PSNR per scene, (3) an entry‑grouping strategy extending rotation‑based quantization to 2‑dimensional hash grid features, and (4) a composable pruning‑quantization pipeline with a closed‑form compression ratio. On NeRF Synthetic, 3DTurboQuant compresses 3DGS by 3.5x with 0.02dB PSNR loss and DUSt3R KV caches by 7.9x with 39.7dB pointmap fidelity. No training, no codebook learning, no calibration data. Compression takes seconds. The code will be released (https://github.com/JaeLee18/3DTurboQuant)
Authors:Rixiang Ni, Boyang Li, Jun Chen, Yonghao Li, Feiyu Ren, Yuji Wang, Haoyang Yuan, Wujiao He, Wei An
Abstract:
Infrared small target detection (IRSTD) aims to separate small targets from clutter backgrounds. Extensive research is dedicated to the pixel‑level supervision‑guided "encoder‑decoder" segmentation paradigm. Although having achieved promising performance, they neglect the fact that small targets only occupy a few pixels and are usually accompanied with blurred boundary caused by clutter backgrounds. Based on this observation, we argue that the first principle of IRSTD should be target localization instead of separating all target region accompanied with indistinguishable background noise. In this paper, we reformulate IRSTD as a centroid regression task and propose a novel Single‑Point Supervision guided Infrared Probabilistic Response Encoding method (namely, SPIRE), which is indeed challenging due to the mismatch between reduced supervision network and equivalent output. Specifically, we first design a Point‑Response Prior Supervision (PRPS), which transforms single‑point annotations into probabilistic response map consistent with infrared point‑target response characteristics, with a High‑Resolution Probabilistic Encoder (HRPE) that enables encoder‑only, end‑to‑end regression without decoder reconstruction. By preserving high‑resolution features and increasing effective supervision density, SPIRE alleviates optimization instability under sparse target distributions. Finally, extensive experiments on various IRSTD benchmarks, including SIRST‑UAVB and SIRST4 demonstrate that SPIRE achieves competitive target‑level detection performance with consistently low false alarm rate (Fa) and significantly reduced computational cost. Code is publicly available at: https://github.com/NIRIXIANG/SPIRE‑IRSTD.
Authors:Yang Yi, Xieyuanli Chen, Jinpu Zhang, Hui Shen, Dewen Hu
Abstract:
Robust local feature detection and description are foundational tasks in computer vision. Existing methods primarily rely on single appearance cues for modeling, leading to unstable keypoints and insufficient descriptor discriminability. In this paper, we propose a multi‑cue guided local feature learning framework that leverages semantic and geometric cues to synergistically enhance detection robustness and descriptor discriminability. Specifically, we construct a joint semantic‑normal prediction head and a depth stability prediction head atop a lightweight backbone. The former leverages a shared 3D vector field to deeply couple semantic and normal cues, thereby resolving optimization interference from heterogeneous inconsistencies. The latter quantifies the reliability of local regions from a geometric consistency perspective, providing deterministic guidance for robust keypoint selection. Based on these predictions, we introduce the Semantic‑Depth Aware Keypoint (SDAK) mechanism for feature detection. By coupling semantic reliability with depth stability, SDAK reweights keypoint responses to suppress spurious features in unreliable regions. For descriptor construction, we design a Unified Triple‑Cue Fusion (UTCF) module, which employs a semantic‑scheduled gating mechanism to adaptively inject multi‑attribute features, improving descriptor discriminability. Extensive experiments on four benchmarks validate the effectiveness of the proposed framework. The source code and pre‑trained model will be available at: https://github.com/yiyscut/GESS.git.
Authors:Yijie Deng, Shuaihang Yuan, Yi Fang
Abstract:
Image Goal Navigation (ImageNav) is evaluated by a coarse success criterion, the agent must stop within 1m of the target, which is sufficient for finding objects but falls short for downstream tasks such as grasping that require precise positioning. We introduce AnyImageNav, a training‑free system that pushes ImageNav toward this more demanding setting. Our key insight is that the goal image can be treated as a geometric query: any photo of an object, a hallway, or a room corner can be registered to the agent's observations via dense pixel‑level correspondences, enabling recovery of the exact 6‑DoF camera pose. Our method realizes this through a semantic‑to‑geometric cascade: a semantic relevance signal guides exploration and acts as a proximity gate, invoking a 3D multi‑view foundation model only when the current view is highly relevant to the goal image; the model then self‑certifies its registration in a loop for an accurate recovered pose. Our method sets state‑of‑the‑art navigation success rates on Gibson (93.1%) and HM3D (82.6%), and achieves pose recovery that prior methods do not provide: a position error of 0.27m and heading error of 3.41 degrees on Gibson, and 0.21m / 1.23 degrees on HM3D, a 5‑10x improvement over adapted baselines.Our project page: https://yijie21.github.io/ain/
Authors:Xueming Fu, Lixia Han
Abstract:
Real‑world smoke simultaneously attenuates scene radiance, adds airlight, and destabilizes multi‑view appearance consistency, making robust 3D reconstruction particularly difficult. We present SmokeGS‑R, a practical pipeline developed for the NTIRE 2026 3D Restoration and Reconstruction Track 2 challenge. The key idea is to decouple geometry recovery from appearance correction: we generate physics‑guided pseudo‑clean supervision with a refined dark channel prior and guided filtering, train a sharp clean‑only 3D Gaussian Splatting source model, and then harmonize its renderings with a donor ensemble using geometric‑mean reference aggregation, LAB‑space Reinhard transfer, and light Gaussian smoothing. On the official challenge testing leaderboard, the final submission achieved \mboxPSNR =15.217 and \mboxSSIM =0.666. After the public release of RealX3D, we re‑evaluated the same frozen result on the seven released challenge scenes without retraining and obtained \mboxPSNR =15.209, \mboxSSIM =0.644, and \mboxLPIPS =0.551, outperforming the strongest official baseline average on the same scenes by +3.68 dB PSNR. These results suggest that a geometry‑first reconstruction strategy combined with stable post‑render appearance harmonization is an effective recipe for real‑world multi‑view smoke restoration. The code is available at https://github.com/windrise/3drr_Track2_SmokeGS‑R.
Authors:Gabriel E. Lima, Valfride Nascimento, Eduardo Santos, Eduil Nascimento, Rayson Laroca, David Menotti
Abstract:
Extracting vehicle information from surveillance images is essential for intelligent transportation systems, enabling applications such as traffic monitoring and criminal investigations. While Automatic License Plate Recognition (ALPR) is widely used, Fine‑Grained Vehicle Classification (FGVC) offers a complementary approach by identifying vehicles based on attributes such as color, make, model, and type. Although there have been advances in this field, existing studies often assume well‑controlled conditions, explore limited attributes, and overlook FGVC integration with ALPR. To address these gaps, we introduce UFPR‑VeSV, a dataset comprising 24,945 images of 16,297 unique vehicles with annotations for 13 colors, 26 makes, 136 models, and 14 types. Collected from the Military Police of Paraná (Brazil) surveillance system, the dataset captures diverse real‑world conditions, including partial occlusions, nighttime infrared imaging, and varying lighting. All FGVC annotations were validated using license plate information, with text and corner annotations also being provided. A qualitative and quantitative comparison with established datasets confirmed the challenging nature of our dataset. A benchmark using five deep learning models further validated this, revealing specific challenges such as handling multicolored vehicles, infrared images, and distinguishing between vehicle models that share a common platform. Additionally, we apply two optical character recognition models to license plate recognition and explore the joint use of FGVC and ALPR. The results highlight the potential of integrating these complementary tasks for real‑world applications. The UFPR‑VeSV dataset is publicly available at: https://github.com/Lima001/UFPR‑VeSV‑Dataset.
Authors:Chan-Wei Hu, Zhengzhong Tu
Abstract:
Multi‑modal retrieval‑augmented generation (MM‑RAG) relies heavily on re‑rankers to surface the most relevant evidence for image‑question queries. However, standard re‑rankers typically process the full query image as a global embedding, making them susceptible to visual distractors (e.g., background clutter) that skew similarity scores. We propose Region‑R1, a query‑side region cropping framework that formulates region selection as a decision‑making problem during re‑ranking, allowing the system to learn to retain the full image or focus only on a question‑relevant region before scoring the retrieved candidates. Region‑R1 learns a policy with a novel region‑aware group relative policy optimization (r‑GRPO) to dynamically crop a discriminative region. Across two challenging benchmarks, E‑VQA and InfoSeek, Region‑R1 delivers consistent gains, achieving state‑of‑the‑art performances by increasing conditional Recall@1 by up to 20%. These results show the great promise of query‑side adaptation as a simple but effective way to strengthen MM‑RAG re‑ranking.
Authors:Timothy Chen, Adam Dai, Maximilian Adang, Grace Gao, Mac Schwager
Abstract:
What makes a good viewpoint? The quality of the data used to learn 3D reconstructions is crucial for enabling efficient and accurate scene modeling. We study the active view selection problem and develop a principled analysis that yields a simple and interpretable criterion for selecting informative camera poses. Our key insight is that informative views can be obtained by minimizing a tractable approximation of the Fisher Information Gain, which reduces to favoring viewpoints that cover geometry that has been insufficiently observed by past cameras. This leads to a lightweight coverage‑based view selection metric that avoids expensive transmittance estimation and is robust to noise and training dynamics. We call this metric COVER (Camera Optimization for View Exploration and Reconstruction). We integrate our method into the Nerfstudio framework and evaluate it on real datasets within fixed and embodied data acquisition scenarios. Across multiple datasets and radiance‑field baselines, our method consistently improves reconstruction quality compared to state‑of‑the‑art active view selection methods. Additional visualizations and our Nerfstudio package can be found at https://chengine.github.io/nbv_gym/.
Authors:Daniel DeTone, Tianwei Shen, Fan Zhang, Lingni Ma, Julian Straub, Richard Newcombe, Jakob Engel
Abstract:
Detecting and localizing objects in space is a fundamental computer vision problem. While much progress has been made to solve 2D object detection, 3D object localization is much less explored and far from solved, especially for open‑world categories. To address this research challenge, we propose Boxer, an algorithm to estimate static 3D bounding boxes (3DBBs) from 2D open‑vocabulary object detections, posed images and optional depth either represented as a sparse point cloud or dense depth. At its core is BoxerNet, a transformer‑based network which lifts 2D bounding box (2DBB) proposals into 3D, followed by multi‑view fusion and geometric filtering to produce globally consistent de‑duplicated 3DBBs in metric world space. Boxer leverages the power of existing 2DBB detection algorithms (e.g. DETIC, OWLv2, SAM3) to localize objects in 2D. This allows the main BoxerNet model to focus on lifting to 3D rather than detecting, ultimately reducing the demand for costly annotated 3DBB training data. Extending the CuTR formulation, we incorporate an aleatoric uncertainty for robust regression, a median depth patch encoding to support sparse depth inputs, and large‑scale training with over 1.2 million unique 3DBBs. BoxerNet outperforms state‑of‑the‑art baselines in open‑world 3DBB lifting, including CuTR in egocentric settings without dense depth (0.532 vs. 0.010 mAP) and on CA‑1M with dense depth available (0.412 vs. 0.250 mAP).
Authors:Ali Aliev, Kamil Garifullin, Nikolay Yudin, Vera Soboleva, Alexander Molozhavenko, Ivan Oseledets, Aibek Alanov, Maxim Rakhuba
Abstract:
In a rapidly growing field of model training there is a constant practical interest in parameter‑efficient fine‑tuning and various techniques that use a small amount of training data to adapt the model to a narrow task. However, there is an open question: how to combine several adapters tuned for different tasks into one which is able to yield adequate results on both tasks? Specifically, merging subject and style adapters for generative models remains unresolved. In this paper we seek to show that in the case of orthogonal fine‑tuning (OFT), we can use structured orthogonal parametrization and its geometric properties to get the formulas for training‑free adapter merging. In particular, we derive the structure of the manifold formed by the recently proposed Group‑and‑Shuffle (\mathcalGS) orthogonal matrices, and obtain efficient formulas for the geodesics approximation between two points. Additionally, we propose a \textspectra restoration transform that restores spectral properties of the merged adapter for higher‑quality fusion. We conduct experiments in subject‑driven generation tasks showing that our technique to merge two \mathcalGS orthogonal matrices is capable of uniting concept and style features of different adapters. To the best of our knowledge, this is the first training‑free method for merging multiplicative orthogonal adapters. Code is available via the \hrefhttps://github.com/ControlGenAI/OrthoFuselink.
Authors:Ziqian Liu, Stephan Alaniz
Abstract:
Instruction‑guided image editing has seen remarkable progress with models like FLUX.2 and Qwen‑Image‑Edit, yet they still struggle with complex scenarios with multiple similar instances each requiring individual edits. We observe that state‑of‑the‑art models suffer from severe over‑editing and spatial misalignment when faced with multiple identical instances and composite instructions. To this end, we introduce a comprehensive benchmark specifically designed to evaluate fine‑grained consistency in multi‑instance and multi‑instruction settings. To address the failures of existing methods observed in our benchmark, we propose Multi‑Instance Regional Alignment via Guided Editing (MIRAGE), a training‑free framework that enables precise, localized editing. By leveraging a vision‑language model to parse complex instructions into regional subsets, MIRAGE employs a multi‑branch parallel denoising strategy. This approach injects latent representations of target regions into the global representation space while maintaining background integrity through a reference trajectory. Extensive evaluations on MIRA‑Bench and RefEdit‑Bench demonstrate that our framework significantly outperforms existing methods in achieving precise instance‑level modifications while preserving background consistency. Our benchmark and code are available at https://github.com/ZiqianLiu666/MIRAGE.
Authors:Yasaman Kashefbahrami, Erkut Akdag, Panagiotis Meletis, Evgeniya Balmashnova, Dip Goswami, Egor Bondarau
Abstract:
Accurate Point Cloud Registration (PCR) is an important task in 3D data processing, involving the estimation of a rigid transformation between two point clouds. While deep‑learning methods have addressed key limitations of traditional non‑learning approaches, such as sensitivity to noise, outliers, occlusion, and initialization, they are developed and evaluated on clean, dense, synthetic datasets (limiting their generalizability to real‑world industrial scenarios). This paper introduces R3PM‑Net, a lightweight, global‑aware, object‑level point matching network designed to bridge this gap by prioritizing both generalizability and real‑time efficiency. To support this transition, two datasets, Sioux‑Cranfield and Sioux‑Scans, are proposed. They provide an evaluation ground for registering imperfect photogrammetric and event‑camera scans to digital CAD models, and have been made publicly available. Extensive experiments demonstrate that R3PM‑Net achieves competitive accuracy with unmatched speed. On ModelNet40, it reaches a perfect fitness score of 1 and inlier RMSE of 0.029 cm in only 0.007s, approximately 7 times faster than the state‑of‑the‑art method RegTR. This performance carries over to the Sioux‑Cranfield dataset, maintaining a fitness of 1 and inlier RMSE of 0.030 cm with similarly low latency. Furthermore, on the highly challenging Sioux‑Scans dataset, R3PM‑Net successfully resolves edge cases in under 50 ms. These results confirm that R3PM‑Net offers a robust, high‑speed solution for critical industrial applications, where precision and real‑time performance are indispensable. The code and datasets are available at https://github.com/YasiiKB/R3PM‑Net.
Authors:Julia Chae, Nicholas Kolkin, Jui-Hsien Wang, Richard Zhang, Sara Beery, Cusuh Ham
Abstract:
Humans have remarkable selective sensitivity to identities ‑‑ easily distinguishing between highly similar identities, even across significantly different contexts such as diverse viewpoints or lighting. Vision models have struggled to match this capability, and progress toward identity‑focused tasks such as personalized image generation is slowed by a lack of identity‑focused evaluation metrics. To help facilitate progress, we propose ID‑Sim, a feed‑forward metric designed to faithfully reflect human selective sensitivity. To build ID‑Sim, we curate a high‑quality training set of images spanning diverse real‑world domains, augmented with generative synthetic data that provides controlled, fine‑grained identity and contextual variations. We evaluate our metric on a new unified evaluation benchmark for assessing consistency with human annotations across identity‑focused recognition, retrieval, and generative tasks.
Authors:StarVLA Community
Abstract:
Building generalist embodied agents requires integrating perception, language understanding, and action, which are core capabilities addressed by Vision‑Language‑Action (VLA) approaches based on multimodal foundation models, including recent advances in vision‑language models and world models. Despite rapid progress, VLA methods remain fragmented across incompatible architectures, codebases, and evaluation protocols, hindering principled comparison and reproducibility. We present StarVLA, an open‑source codebase for VLA research. StarVLA addresses these challenges in three aspects. First, it provides a modular backbone‑‑action‑head architecture that supports both VLM backbones (e.g., Qwen‑VL) and world‑model backbones (e.g., Cosmos) alongside representative action‑decoding paradigms, all under a shared abstraction in which backbone and action head can each be swapped independently. Second, it provides reusable training strategies, including cross‑embodiment learning and multimodal co‑training, that apply consistently across supported paradigms. Third, it integrates major benchmarks, including LIBERO, SimplerEnv, RoboTwin~2.0, RoboCasa‑GR1, and BEHAVIOR‑1K, through a unified evaluation interface that supports both simulation and real‑robot deployment. StarVLA also ships simple, fully reproducible single‑benchmark training recipes that, despite minimal data engineering, already match or surpass prior methods on multiple benchmarks with both VLM and world‑model backbones. To our best knowledge, StarVLA is one of the most comprehensive open‑source VLA frameworks available, and we expect it to lower the barrier for reproducing existing methods and prototyping new ones. StarVLA is being actively maintained and expanded; we will update this report as the project evolves. The code and documentation are available at https://github.com/starVLA/starVLA.
Authors:Hyunsoo Cha, Wonjung Woo, Byungjun Kim, Hanbyul Joo
Abstract:
We present Vanast, a unified framework that generates garment‑transferred human animation videos directly from a single human image, garment images, and a pose guidance video. Conventional two‑stage pipelines treat image‑based virtual try‑on and pose‑driven animation as separate processes, which often results in identity drift, garment distortion, and front‑back inconsistency. Our model addresses these issues by performing the entire process in a single unified step to achieve coherent synthesis. To enable this setting, we construct large‑scale triplet supervision. Our data generation pipeline includes generating identity‑preserving human images in alternative outfits that differ from garment catalog images, capturing full upper and lower garment triplets to overcome the single‑garment‑posed video pair limitation, and assembling diverse in‑the‑wild triplets without requiring garment catalog images. We further introduce a Dual Module architecture for video diffusion transformers to stabilize training, preserve pretrained generative quality, and improve garment accuracy, pose adherence, and identity preservation while supporting zero‑shot garment interpolation. Together, these contributions allow Vanast to produce high‑fidelity, identity‑consistent animation across a wide range of garment types.
Authors:Siyuan Liu, Chaoqun Zheng, Xin Zhou, Tianrui Feng, Dingkang Liang, Xiang Bai
Abstract:
Scene‑level point cloud understanding remains challenging due to diverse geometries, imbalanced category distributions, and highly varied spatial layouts. Existing methods improve object‑level performance but rely on static network parameters during inference, limiting their adaptability to dynamic scene data. We propose PointTPA, a Test‑time Parameter Adaptation framework that generates input‑aware network parameters for scene‑level point clouds. PointTPA adopts a Serialization‑based Neighborhood Grouping (SNG) to form locally coherent patches and a Dynamic Parameter Projector (DPP) to produce patch‑wise adaptive weights, enabling the backbone to adjust its behavior according to scene‑specific variations while maintaining a low parameter overhead. Integrated into the PTv3 structure, PointTPA demonstrates strong parameter efficiency by introducing two lightweight modules of less than 2% of the backbone's parameters. Despite this minimal parameter overhead, PointTPA achieves 78.4% mIoU on ScanNet validation, surpassing existing parameter‑efficient fine‑tuning (PEFT) methods across multiple benchmarks, highlighting the efficacy of our test‑time dynamic network parameter adaptation mechanism in enhancing 3D scene understanding. The code is available at https://github.com/H‑EmbodVis/PointTPA.
Authors:David Nordström, Johan Edstedt, Georg Bökman, Jonathan Astermark, Anders Heyden, Viktor Larsson, Mårten Wadenbäck, Michael Felsberg, Fredrik Kahl
Abstract:
Local feature matching has long been a fundamental component of 3D vision systems such as Structure‑from‑Motion (SfM), yet progress has lagged behind the rapid advances of modern data‑driven approaches. The newer approaches, such as feed‑forward reconstruction models, have benefited extensively from scaling dataset sizes, whereas local feature matching models are still only trained on a few mid‑sized datasets. In this paper, we revisit local feature matching from a data‑driven perspective. In our approach, which we call LoMa, we combine large and diverse data mixtures, modern training recipes, scaled model capacity, and scaled compute, resulting in remarkable gains in performance. Since current standard benchmarks mainly rely on collecting sparse views from successful 3D reconstructions, the evaluation of progress in feature matching has been limited to relatively easy image pairs. To address the resulting saturation of benchmarks, we collect 1000 highly challenging image pairs from internet data into a new dataset called HardMatch. Ground truth correspondences for HardMatch are obtained via manual annotation by the authors. In our extensive benchmarking suite, we find that LoMa makes outstanding progress across the board, outperforming the state‑of‑the‑art method ALIKED+LightGlue by +18.6 mAA on HardMatch, +29.5 mAA on WxBS, +21.4 (1m, 10^\circ) on InLoc, +24.2 AUC on RUBIK, and +12.4 mAA on IMC 2022. We release our code and models publicly at https://github.com/davnords/LoMa.
Authors:Zeyu Ma, Alexander Raistrick, Jia Deng
Abstract:
In this paper, we explore the design space of procedural rules for multi‑view stereo (MVS). We demonstrate that we can generate effective training data using SimpleProc: a new, fully procedural generator driven by a very small set of rules using Non‑Uniform Rational Basis Splines (NURBS), as well as basic displacement and texture patterns. At a modest scale of 8,000 images, our approach achieves superior results compared to manually curated images (at the same scale) sourced from games and real‑world objects. When scaled to 352,000 images, our method yields performance comparable to‑‑and in several benchmarks, exceeding‑‑models trained on over 692,000 manually curated images. The source code and the data are available at https://github.com/princeton‑vl/SimpleProc.
Authors:Sudarshan Rajagopalan, Vishal M. Patel
Abstract:
Pre‑trained diffusion models have enabled significant advancements in All‑in‑One Restoration (AiOR), offering improved perceptual quality and generalization. However, diffusion‑based restoration methods primarily rely on fine‑tuning or Control‑Net style modules to leverage the pre‑trained diffusion model's priors for AiOR. In this work, we show that these pre‑trained diffusion models inherently possess restoration behavior, which can be unlocked by directly learning prompt embeddings at the output of the text encoder. Interestingly, this behavior is largely inaccessible through text prompts and text‑token embedding optimization. Furthermore, we observe that naive prompt learning is unstable because the forward noising process using degraded images is misaligned with the reverse sampling trajectory. To resolve this, we train prompts within a diffusion bridge formulation that aligns training and inference dynamics, enforcing a coherent denoising path from noisy degraded states to clean images. Building on these insights, we introduce our lightweight learned prompts on the pre‑trained WAN video model and FLUX image models, converting them into high‑performing restoration models. Extensive experiments demonstrate that our approach achieves competitive performance and generalization across diverse degradations, while avoiding fine‑tuning and restoration‑specific control modules.
Authors:Weian Mao, Xi Lin, Wei Huang, Yuxin Xie, Tianfu Fu, Bohan Zhuang, Song Han, Yukang Chen
Abstract:
Extended reasoning in large language models (LLMs) creates severe KV cache memory bottlenecks. Leading KV cache compression methods estimate KV importance using attention scores from recent post‑RoPE queries. However, queries rotate with position during RoPE, making representative queries very few, leading to poor top‑key selection and unstable reasoning. To avoid this issue, we turn to the pre‑RoPE space, where we observe that Q and K vectors are highly concentrated around fixed non‑zero centers and remain stable across positions ‑‑ Q/K concentration. We show that this concentration causes queries to preferentially attend to keys at specific distances (e.g., nearest keys), with the centers determining which distances are preferred via a trigonometric series. Based on this, we propose TriAttention to estimate key importance by leveraging these centers. Via the trigonometric series, we use the distance preference characterized by these centers to score keys according to their positions, and also leverage Q/K norms as an additional signal for importance estimation. On AIME25 with 32K‑token generation, TriAttention matches Full Attention reasoning accuracy while achieving 2.5x higher throughput or 10.7x KV memory reduction, whereas leading baselines achieve only about half the accuracy at the same efficiency. TriAttention enables OpenClaw deployment on a single consumer GPU, where long context would otherwise cause out‑of‑memory with Full Attention.
Authors:Yicheng Xiao, Wenhu Zhang, Lin Song, Yukang Chen, Wenbo Li, Nan Jiang, Tianhe Ren, Haokun Lin, Wei Huang, Haoyang Huang, Xiu Li, Nan Duan, Xiaojuan Qi
Abstract:
Image spatial editing performs geometry‑driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine‑grained spatial manipulations, motivating a dedicated assessment suite. Our contributions are listed: (i) We introduce SpatialEdit‑Bench, a complete benchmark that evaluates spatial editing by jointly measuring perceptual plausibility and geometric fidelity via viewpoint reconstruction and framing analysis. (ii) To address the data bottleneck for scalable training, we construct SpatialEdit‑500k, a synthetic dataset generated with a controllable Blender pipeline that renders objects across diverse backgrounds and systematic camera trajectories, providing precise ground‑truth transformations for both object‑ and camera‑centric operations. (iii) Building on this data, we develop SpatialEdit‑16B, a baseline model for fine‑grained spatial editing. Our method achieves competitive performance on general editing while substantially outperforming prior methods on spatial manipulation tasks. All resources will be made public at https://github.com/EasonXiao‑888/SpatialEdit.
Authors:Shuai Liu, Shulin Tian, Kairui Hu, Yuhao Dong, Zhe Yang, Bo Li, Jingkang Yang, Chen Change Loy, Ziwei Liu
Abstract:
Coworking AI agents operating within local file systems are rapidly emerging as a paradigm in human‑AI interaction; however, effective personalization remains limited by severe data constraints, as strict privacy barriers and the difficulty of jointly collecting multimodal real‑world traces prevent scalable training and evaluation, and existing methods remain interaction‑centric while overlooking dense behavioral traces in file‑system operations; to address this gap, we propose FileGram, a comprehensive framework that grounds agent memory and personalization in file‑system behavioral traces, comprising three core components: (1) FileGramEngine, a scalable persona‑driven data engine that simulates realistic workflows and generates fine‑grained multimodal action sequences at scale; (2) FileGramBench, a diagnostic benchmark grounded in file‑system behavioral traces for evaluating memory systems on profile reconstruction, trace disentanglement, persona drift detection, and multimodal grounding; and (3) FileGramOS, a bottom‑up memory architecture that builds user profiles directly from atomic actions and content deltas rather than dialogue summaries, encoding these traces into procedural, semantic, and episodic channels with query‑time abstraction; extensive experiments show that FileGramBench remains challenging for state‑of‑the‑art memory systems and that FileGramEngine and FileGramOS are effective, and by open‑sourcing the framework, we hope to support future research on personalized memory‑centric file‑system agents.
Authors:Mauricio Soroco, Francesco Pittaluga, Zaid Tasneem, Abhishek Aich, Bingbing Zhuang, Wuyang Chen, Manmohan Chandraker, Ziyu Jiang
Abstract:
Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real‑world testing provides. Yet existing instruction guided image editors, trained on object‑centric or artistic data, struggle with dense, safety‑critical driving layouts. We propose HorizonWeaver, which tackles three fundamental challenges in driving scene editing: (1) multi‑level granularity, requiring coherent object‑ and scene‑level edits in dense environments; (2) rich high‑level semantics, preserving diverse objects while following detailed instructions; and (3) ubiquitous domain shifts, handling changes in climate, layout, and traffic across unseen environments. The core of HorizonWeaver is a set of complementary contributions across data, model, and training: (1) Data: Large‑scale dataset generation, where we build a paired real/synthetic dataset from Boreas, nuScenes, and Argoverse2 to improve generalization; (2) Model: Language‑Guided Masks for fine‑grained editing, where semantics‑enriched masks and prompts enable precise, language‑guided edits; and (3) Training: Content preservation and instruction alignment, where joint losses enforce scene consistency and instruction fidelity. Together, HorizonWeaver provides a scalable framework for photorealistic, instruction‑driven editing of complex driving scenes, collecting 255K images across 13 editing categories and outperforming prior methods in L1, CLIP, and DINO metrics, achieving +46.4% user preference and improving BEV segmentation IoU by +33%. Project page: https://msoroco.github.io/horizonweaver/
Authors:Ke Li, Maoliang Li, Jialiang Chen, Jiayu Chen, Zihao Zheng, Shaoqi Wang, Xiang Chen
Abstract:
Video mashup creation represents a complex video editing paradigm that recomposes existing footage to craft engaging audio‑visual experiences, demanding intricate orchestration across semantic, visual, and auditory dimensions and multiple levels. However, existing automated editing frameworks often overlook the cross‑level multimodal orchestration to achieve professional‑grade fluidity, resulting in disjointed sequences with abrupt visual transitions and musical misalignment. To address this, we formulate video mashup creation as a Multimodal Coherency Satisfaction Problem (MMCSP) and propose the DIRECT framework. Simulating a professional production pipeline, our hierarchical multi‑agent framework decomposes the challenge into three cascade levels: the Screenwriter for source‑aware global structural anchoring, the Director for instantiating adaptive editing intent and guidance, and the Editor for intent‑guided shot sequence editing with fine‑grained optimization. We further introduce Mashup‑Bench, a comprehensive benchmark with tailored metrics for visual continuity and auditory alignment. Extensive experiments demonstrate that DIRECT significantly outperforms state‑of‑the‑art baselines in both objective metrics and human subjective evaluation. Project page and code: https://github.com/AK‑DREAM/DIRECT
Authors:Kaede Shiohara, Toshihiko Yamasaki
Abstract:
Automatic residential floorplan generation has long been a central challenge bridging architecture and computer graphics, aiming to make spatial design more efficient and accessible. While early methods based on constraint satisfaction or combinatorial optimization ensure feasibility, they lack diversity and flexibility. Recent generative models achieve promising results but struggle to generalize across heterogeneous conditional tasks, such as generation from site boundaries, room adjacency graphs, or partial layouts, due to their suboptimal representations. To address this gap, we introduce Floorplan Markup Language (FML), a general representation that encodes floorplan information within a single structured grammar, which casts the entire floorplan generation problem into a next token prediction task. Leveraging FML, we develop a transformer‑based generative model, FMLM, capable of producing high‑fidelity and functional floorplans under diverse conditions. Comprehensive experiments on the RPLAN dataset demonstrate that FMLM, despite being a single model, surpasses the previous task‑specific state‑of‑the‑art methods.
Authors:Yude Zou, Junji Gong, Xing Gao, Zixuan Li, Tianxing Chen, Guanjie Zheng
Abstract:
Human‑object‑scene interactions (HOSI) generation has broad applications in embodied AI, simulation, and animation. Unlike human‑object interaction (HOI) and human‑scene interaction (HSI), HOSI generation requires reasoning over dynamic object‑scene changes, yet suffers from limited annotated data. To address these issues, we propose a coarse‑to‑fine instruction‑conditioned interaction generation framework that is explicitly aligned with the iterative denoising process of a consistency model. In particular, we adopt a dynamic perception strategy that leverages trajectories from the preceding refinement to update scene context and condition subsequent refinement at each denoising step of consistency model, yielding consistent interactions. To further reduce physical artifacts, we introduce a bump‑aware guidance that mitigates collisions and penetrations during sampling without requiring fine‑grained scene geometry, enabling real‑time generation. To overcome data scarcity, we design a hybrid training startegy that synthesizes pseudo‑HOSI samples by injecting voxelized scene occupancy into HOI datasets and jointly trains with high‑fidelity HSI data, allowing interaction learning while preserving realistic scene awareness. Extensive experiments demonstrate that our method achieves state‑of‑the‑art performance in both HOSI and HOI generation, and strong generalization to unseen scenes. Project page: https://yudezou.github.io/InfBaGel‑page/
Authors:Haoxuan Han, Weijie Wang, Zeyu Zhang, Yefei He, Bohan Zhuang
Abstract:
Recent advancements in Vision‑Language Models (VLMs) have significantly pushed the boundaries of Visual Question Answering (VQA).However,high‑resolution details can sometimes become noise that leads to hallucinations or reasoning errors. In this paper,we propose Degradation‑Driven Prompting (DDP), a novel framework that improves VQA performance by strategically reducing image fidelity to force models to focus on essential structural information. We evaluate DDP across two distinct tasks. Physical attributes targets images prone to human misjudgment, where DDP employs a combination of 80p downsampling, structural visual aids (white background masks and orthometric lines), and In‑Context Learning (ICL) to calibrate the model's focus. Perceptual phenomena addresses various machine‑susceptible visual anomalies and illusions, including Visual Anomaly (VA), Color (CI), Motion(MI),Gestalt (GI), Geometric (GSI), and Visual Illusions (VI).For this task, DDP integrates a task‑classification stage with specialized tools such as blur masks and contrast enhancement alongside downsampling. Our experimental results demonstrate that less is more: by intentionally degrading visual inputs and providing targeted structural prompts, DDP enables VLMs to bypass distracting textures and achieve superior reasoning accuracy on challenging visual benchmarks.
Authors:Jiajun Zhai, Hao Shi, Shangwei Guo, Kailun Yang, Kaiwei Wang
Abstract:
Robotic Vision‑Language‑Action (VLA) models generalize well for open‑ended manipulation, but their perception is fragile under sensing‑stage degradations such as extreme low light, motion blur, and black clipping. We present E‑VLA, an event‑augmented VLA framework that improves manipulation robustness when conventional frame‑based vision becomes unreliable. Instead of reconstructing images from events, E‑VLA directly leverages motion and structural cues in event streams to preserve semantic perception and perception‑action consistency under adverse conditions. We build an open‑source teleoperation platform with a DAVIS346 event camera and collect a real‑world synchronized RGB‑event‑action manipulation dataset across diverse tasks and illumination settings. We also propose lightweight, pretrained‑compatible event integration strategies and study event windowing and fusion for stable deployment. Experiments show that even a simple parameter‑free fusion, i.e., overlaying accumulated event maps onto RGB images, could substantially improve robustness in dark and blur‑heavy scenes: on Pick‑Place at 20 lux, success increases from 0% (image‑only) to 60% with overlay fusion and to 90% with our event adapter; under severe motion blur (1000 ms exposure), Pick‑Place improves from 0% to 20‑25%, and Sorting from 5% to 32.5%. Overall, E‑VLA provides systematic evidence that event‑driven perception can be effectively integrated into VLA models, pointing toward robust embodied intelligence beyond conventional frame‑based imaging. Code and dataset will be available at https://github.com/JJayzee/E‑VLA.
Authors:Hongyu Liu, Xuan Wang, Zijian Wu, Yating Wang, Ziyu Wan, Yue Ma, Runtao Liu, Boyao Zhou, Yujun Shen, Qifeng Chen
Abstract:
We introduce AvatarPointillist, a novel framework for generating dynamic 4D Gaussian avatars from a single portrait image. At the core of our method is a decoder‑only Transformer that autoregressively generates a point cloud for 3D Gaussian Splatting. This sequential approach allows for precise, adaptive construction, dynamically adjusting point density and the total number of points based on the subject's complexity. During point generation, the AR model also jointly predicts per‑point binding information, enabling realistic animation. After generation, a dedicated Gaussian decoder converts the points into complete, renderable Gaussian attributes. We demonstrate that conditioning the decoder on the latent features from the AR generator enables effective interaction between stages and markedly improves fidelity. Extensive experiments validate that AvatarPointillist produces high‑quality, photorealistic, and controllable avatars. We believe this autoregressive formulation represents a new paradigm for avatar generation, and we will release our code inspire future research.
Authors:DataFlow Team, Bohan Zeng, Daili Hua, Kaixin Zhu, Yifan Dai, Bozhou Li, Yuran Wang, Chengzhuo Tong, Yifan Yang, Mingkun Chang, Jianbin Zhao, Zhou Liu, Hao Liang, Xiaochen Ma, Ruichuan An, Junbo Niu, Zimo Meng, Tianyi Bai, Meiyi Qiang, Huanyao Zhang, Zhiyou Xiao, Tianyu Guo, Qinhan Yu, Runhao Zhao, Zhengpin Li, Xinyi Huang, Yisheng Pan, Yiwen Tang, Juanxi Tian, Yang Shi, Yue Ding, Xinlong Chen, Hongcheng Gao, Minglei Shi, Jialong Wu, Zekun Wang, Yuanxing Zhang, Xintao Wang, Pengfei Wan, Yiren Song, Mike Zheng Shou, Wentao Zhang
Abstract:
World models have garnered significant attention as a promising research direction in artificial intelligence, yet a clear and unified definition remains lacking. In this paper, we introduce OpenWorldLib, a comprehensive and standardized inference framework for Advanced World Models. Drawing on the evolution of world models, we propose a clear definition: a world model is a model or framework centered on perception, equipped with interaction and long‑term memory capabilities, for understanding and predicting the complex world. We further systematically categorize the essential capabilities of world models. Based on this definition, OpenWorldLib integrates models across different tasks within a unified framework, enabling efficient reuse and collaborative inference. Finally, we present additional reflections and analyses on potential future directions for world model research. Code link: https://github.com/OpenDCAI/OpenWorldLib
Authors:Jaeyoon Jung, Yejun Yoon, Kunwoo Park
Abstract:
Automated fact‑checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text‑only to multimodal fact‑checking, a prevailing assumption is that incorporating visual evidence universally improves performance. In this work, we challenge this assumption and show that the indiscriminate use of multimodal evidence can reduce accuracy. To address this challenge, we propose AMuFC, a multimodal fact‑checking framework that employs two collaborative vision‑language models with distinct roles for the adaptive use of visual evidence: an Analyzer determines whether visual evidence is necessary for claim verification, and a Verifier predicts claim veracity conditioned on both the retrieved evidence and the Analyzer's assessment. Experimental results on three datasets show that incorporating the Analyzer's assessment of visual evidence necessity into the Verifier's prediction yields substantial improvements in verification performance. We will release all code and datasets at https://github.com/ssu‑humane/AMuFC.
Authors:Qing Zhou, Bingxuan Zhao, Tao Yang, Hongyuan Zhang, Junyu Gao, Qi Wang
Abstract:
Dynamic data pruning accelerates deep learning by selectively omitting less informative samples during training. While per‑sample loss is a common importance metric, obtaining it can be challenging or infeasible for complex models or loss functions, often requiring significant implementation effort. This work proposes the Batch Loss Score (BLS), a computationally efficient alternative using an Exponential Moving Average (EMA) of readily available batch losses to assign scores to individual samples. We frame the batch loss, from the perspective of a single sample, as a noisy measurement of its scaled individual loss, with noise originating from stochastic batch composition. It is formally shown that the EMA mechanism functions as a first‑order low‑pass filter, attenuating high‑frequency batch composition noise. This yields a score approximating the smoothed and persistent contribution of the individual sample to the loss, providing a theoretical grounding for BLS as a proxy for sample importance. BLS demonstrates remarkable code integration simplicity (three‑line injection) and readily adapts existing per‑sample loss‑based methods (one‑line proxy). Its effectiveness is demonstrated by enhancing two such methods to losslessly prune 20%‑50% of samples across 14 datasets, 11 tasks and 18 models, highlighting its utility and broad applicability, especially for complex scenarios where per‑sample loss is difficult to access. Code is available at https://github.com/mrazhou/BLS.
Authors:Yihan Sun, Yuqi Cheng, Junjie Zu, Yuxiang Tan, Guoyang Xie, Yucheng Wang, Yunkang Cao, Weiming Shen
Abstract:
Industrial 3D anomaly detection performance is fundamentally constrained by the scarcity and long‑tailed distribution of abnormal samples. To address this challenge, we propose Synthesis4AD, an end‑to‑end paradigm that leverages large‑scale, high‑fidelity synthetic anomalies to learn more discriminative representations for 3D anomaly detection. At the core of Synthesis4AD is 3D‑DefectStudio, a software platform built upon the controllable synthesis engine MPAS, which injects geometrically realistic defects guided by higher‑dimensional support primitives while simultaneously generating accurate point‑wise anomaly masks. Furthermore, Synthesis4AD incorporates a multimodal large language model (MLLM) to interpret product design information and automatically translate it into executable anomaly synthesis instructions, enabling scalable and knowledge‑driven anomalous data generation. To improve the robustness and generalization of the downstream detector on unstructured point clouds, Synthesis4AD further introduces a training pipeline based on spatial‑distribution normalization and geometry‑faithful data augmentations, which alleviates the sensitivity of Point Transformer architectures to absolute coordinates and improves feature learning under realistic data variations. Extensive experiments demonstrate state‑of‑the‑art performance on Real3D‑AD, MulSen‑AD, and a real‑world industrial parts dataset. The proposed synthesis method MPAS and the interactive system 3D‑DefectStudio will be publicly released at https://github.com/hustCYQ/Synthesis4AD.
Authors:Yeonwoo Cha, Jaehoon Yoo, Semin Kim, Yunseo Park, Jinhyeon Kwon, Seunghoon Hong
Abstract:
Flow‑based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample‑wise velocities connecting each sample from a simple prior to the target data. When sample‑wise velocities conflict at the same intermediate state, however, this averaged velocity can misguide samples toward low‑density regions, degrading generation quality. To address this issue, we propose the Flow Divergence Sampler (FDS), a training‑free framework that refines intermediate states before each solver step. Our key finding reveals that the severity of this misguidance is quantified by the divergence of the marginal velocity field that is readily computable during inference with a well‑optimized model. FDS exploits this signal to steer states toward less ambiguous regions. As a plug‑and‑play framework compatible with standard solvers and off‑the‑shelf flow backbones, FDS consistently improves fidelity across various generation tasks including text‑to‑image synthesis, and inverse problems.
Authors:Inseong Choi, Siwoo Lee, Seung-Hun Nam, Soohwan Song
Abstract:
Diffusion models are promising for sparse‑view novel view synthesis (NVS), as they can generate pseudo‑ground‑truth views to aid 3D reconstruction pipelines like 3D Gaussian Splatting (3DGS). However, these synthesized images often contain photometric and geometric inconsistencies, and their direct use for supervision can impair reconstruction. To address this, we propose Partial‑Reference Image Quality Assessment (PR‑IQA), a framework that evaluates diffusion‑generated views using reference images from different poses, eliminating the need for ground truth. PR‑IQA first computes a geometrically consistent partial quality map in overlapping regions. It then performs quality completion to inpaint this partial map into a dense, full‑image map. This completion is achieved via a cross‑attention mechanism that incorporates reference‑view context, ensuring cross‑view consistency and enabling thorough quality assessment. When integrated into a diffusion‑augmented 3DGS pipeline, PR‑IQA restricts supervision to high‑confidence regions identified by its quality maps. Experiments demonstrate that PR‑IQA outperforms existing IQA methods, achieving full‑reference‑level accuracy without ground‑truth supervision. Thus, our quality‑aware 3DGS approach more effectively filters inconsistencies, producing superior 3D reconstructions and NVS results. The project page is available at https://kakaomacao.github.io/pr‑iqa‑project‑page/.
Authors:Jiwon Kim, Ikbeom Jang
Abstract:
Medical imaging archives are growing rapidly in both size and resolution, making efficient compression increasingly important for storage and data transfer. Most existing codecs compress full images/volumes(including non‑diagnostic background) or apply differential ROI coding that still preserves background bits. We propose MedROI, a codec‑agnostic, plug‑and‑play ROI‑centric framework that discards background voxels prior to compression. MedROI extracts a tight tissue bounding box via lightweight intensity‑based thresholding and stores a fixed 54byte meta data record to enable spatial restoration during decompression. The cropped ROI is then compressed using any existing 2D or 3D codec without architectural modifications or retraining. We evaluate MedROI on 200 T1‑weighted brain MRI volumes from ADNI using 6 codec configurations spanning conventional codecs (JPEG2000 2D/3D, HEIF) and neural compressors (LIC_TCM, TCM+AuxT, BCM‑Net, SirenMRI). MedROI yields statistically significant improvements in compression ratio and encoding/decoding time for most configurations (two‑sided t‑test with multiple‑comparison correction), while maintaining comparable reconstruction quality when measured within the ROI; HEIF is the primary exception in compression‑ratio gains. For example, on JPEG20002D (lv3), MedROI improves CR from 20.35 to 27.37 while reducing average compression time from 1.701s to 1.380s. Code is available at https://github.com/labhai/MedROI.
Authors:Jianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang, Simon Jenni, Yun Fu
Abstract:
Recent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training objectives, or data modalities. However, these representations are essentially internal abstractions of samples that characterize samples independently, leading to limited expressiveness. In this paper, we propose The Indra Representation Hypothesis, inspired by the philosophical metaphor of Indra's Net. We argue that representations from unimodal foundation models are converging to implicitly reflect a shared relational structure underlying reality, akin to the relational ontology of Indra's Net. We formalize this hypothesis using the V‑enriched Yoneda embedding from category theory, defining the Indra representation as a relational profile of each sample with respect to others. This formulation is shown to be unique, complete, and structure‑preserving under a given cost function. We instantiate the Indra representation using angular distance and evaluate it in cross‑model and cross‑modal scenarios involving vision, language, and audio. Extensive experiments demonstrate that Indra representations consistently enhance robustness and alignment across architectures and modalities, providing a theoretically grounded and practical framework for training‑free alignment of unimodal foundation models. Our code is available at https://github.com/Jianglin954/Indra.
Authors:Junyoung Park, Youngjin Oh, Nam Ik Cho
Abstract:
Blind‑spot networks (BSNs) enable self‑supervised image denoising by preventing access to the target pixel, allowing clean signal estimation without ground‑truth supervision. However, this approach assumes pixel‑wise noise independence, which is violated in real‑world sRGB images due to spatially correlated noise from the camera's image signal processing (ISP) pipeline. While several methods employ downsampling to decorrelate noise, they alter noise statistics and limit the network's ability to utilize full contextual information. In this paper, we propose the Triangular‑Masked Blind‑Spot Network (TM‑BSN), a novel blind‑spot architecture that accurately models the spatial correlation of real sRGB noise. This correlation originates from demosaicing, where each pixel is reconstructed from neighboring samples with spatially decaying weights, resulting in a diamond‑shaped pattern. To align the receptive field with this geometry, we introduce a triangular‑masked convolution that restricts the kernel to its upper‑triangular region, creating a diamond‑shaped blind spot at the original resolution. This design excludes correlated pixels while fully leveraging uncorrelated context, eliminating the need for downsampling or post‑processing. Furthermore, we use knowledge distillation to transfer complementary knowledge from multiple blind‑spot predictions into a lightweight U‑Net, improving both accuracy and efficiency. Extensive experiments on real‑world benchmarks demonstrate that our method achieves state‑of‑the‑art performance, significantly outperforming existing self‑supervised approaches. Our code is available at https://github.com/parkjun210/TM‑BSN.
Authors:Ryuki Tezuka, Chihiro Nakatani, Norimichi Ukita
Abstract:
This paper proposes Group Activity Feature (GAF) learning without group activity annotations. Unlike prior work, which uses low‑level static local features to learn GAFs, we propose leveraging dynamics‑aware and group‑aware pretext tasks, along with local and global features provided by DINO, for group‑dynamics‑aware GAF learning. To adapt DINO and GAF learning to local dynamics and global group features, our pretext tasks use person flow estimation and group‑relevant object location estimation, respectively. Person flow estimation is used to represent the local motion of each person, which is an important cue for understanding group activities. In contrast, group‑relevant object location estimation encourages GAFs to learn scene context (e.g., spatial relations of people and objects) as global features. Comprehensive experiments on public datasets demonstrate the state‑of‑the‑art performance of our method in group activity retrieval and recognition. Our ablation studies verify the effectiveness of each component in our method. Code: https://github.com/tezuka0001/Group‑DINOmics.
Authors:Ze-Xin Yin, Liu Liu, Xinjie Wang, Wei Sui, Zhizhong Su, Jian Yang, Jin Xie
Abstract:
Compositional 3D scene generation from a single view requires the simultaneous recovery of scene layout and 3D assets. Existing approaches mainly fall into two categories: feed‑forward generation methods and per‑instance generation methods. The former directly predict 3D assets with explicit 6DoF poses through efficient network inference, but they generalize poorly to complex scenes. The latter improve generalization through a divide‑and‑conquer strategy, but suffer from time‑consuming pose optimization. To bridge this gap, we introduce 3D‑Fixer, a novel in‑place completion paradigm. Specifically, 3D‑Fixer extends 3D object generative priors to generate complete 3D assets conditioned on the partially visible point cloud at the original locations, which are cropped from the fragmented geometry obtained from the geometry estimation methods. Unlike prior works that require explicit pose alignment, 3D‑Fixer uses fragmented geometry as a spatial anchor to preserve layout fidelity. At its core, we propose a coarse‑to‑fine generation scheme to resolve boundary ambiguity under occlusion, supported by a dual‑branch conditioning network and an Occlusion‑Robust Feature Alignment (ORFA) strategy for stable training. Furthermore, to address the data scarcity bottleneck, we present ARSG‑110K, the largest scene‑level dataset to date, comprising over 110K diverse scenes and 3M annotated images with high‑fidelity 3D ground truth. Extensive experiments show that 3D‑Fixer achieves state‑of‑the‑art geometric accuracy, which significantly outperforms baselines such as MIDI and Gen3DSR, while maintaining the efficiency of the diffusion process. Code and data will be publicly available at https://zx‑yin.github.io/3dfixer.
Authors:Pei Yang, Hai Ci, Beibei Lin, Yiren Song, Mike Zheng Shou
Abstract:
Nighttime video deraining is uniquely challenging because raindrops interact with artificial lighting. Unlike daytime white rain, nighttime rain takes on various colors and appears locally illuminated. Existing small‑scale synthetic datasets rely on 2D rain overlays and fail to capture these physical properties, causing models to generalize poorly to real‑world night rain. Meanwhile, capturing real paired nighttime videos remains impractical because rain effects cannot be isolated from other degradations like sensor noise. To bridge this gap, we introduce UENR‑600K, a large‑scale, physically grounded dataset containing 600,000 1080p frame pairs. We utilize Unreal Engine to simulate rain as 3D particles within virtual environments. This approach guarantees photorealism and physically real raindrops, capturing correct details like color refractions, scene occlusions, rain curtains. Leveraging this high‑quality data, we establish a new state‑of‑the‑art baseline by adapting the Wan 2.2 video generation model. Our baseline treat deraining as a video‑to‑video generation task, exploiting strong generative priors to almost entirely bridge the sim‑to‑real gap. Extensive benchmarking demonstrates that models trained on our dataset generalize significantly better to real‑world videos. Project page: https://showlab.github.io/UENR‑600K/.
Authors:Weiguo Pian, Saksham Singh Kushwaha, Zhimin Chen, Shijian Deng, Kai Wang, Yunhui Guo, Yapeng Tian
Abstract:
In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on‑screen and off‑screen sounds across diverse domains (e.g., ambient events, musical instruments, and human speech). Prior video‑conditioned audio generation models typically focus on producing on‑screen environmental sounds that correspond to visible sounding events, neglecting off‑screen auditory events. While recent holistic joint text‑video‑to‑audio generation models aim to produce auditory scenes with both on‑ and off‑screen sound but they are limited to non‑speech sounds, lacking the ability to generate or integrate human speech. To overcome these limitations, we introduce OmniSonic, a flow‑matching‑based diffusion framework jointly conditioned on video and text. It features a TriAttn‑DiT architecture that performs three cross‑attention operations to process on‑screen environmental sound, off‑screen environmental sound, and speech conditions simultaneously, with a Mixture‑of‑Experts (MoE) gating mechanism that adaptively balances their contributions during generation. Furthermore, we construct UniHAGen‑Bench, a new benchmark with over one thousand samples covering three representative on/off‑screen speech‑environment scenarios. Extensive experiments show that OmniSonic consistently outperforms state‑of‑the‑art approaches on both objective metrics and human evaluations, establishing a strong baseline for universal and holistic audio generation. Project page: https://weiguopian.github.io/OmniSonic_webpage/
Authors:Xu Yan, Jun Yin, Shiliang Sun, Minghua Wan
Abstract:
Although multi‑view multi‑label learning has been extensively studied, research on the dual‑missing scenario, where both views and labels are incomplete, remains largely unexplored. Existing methods mainly rely on contrastive learning or information bottleneck theory to learn consistent representations under missing‑view conditions, but loss‑based alignment without explicit structural constraints limits the ability to capture stable and discriminative shared semantics. To address this issue, we introduce a more structured mechanism for consistent representation learning: we learn discrete consistent representations through a multi‑view shared codebook and cross‑view reconstruction, which naturally align different views within the limited shared codebook embeddings and reduce feature redundancy. At the decision level, we design a weight estimation method that evaluates the ability of each view to preserve label correlation structures, assigning weights accordingly to enhance the quality of the fused prediction. In addition, we introduce a fused‑teacher self‑distillation framework, where the fused prediction guides the training of view‑specific classifiers and feeds the global knowledge back into the single‑view branches, thereby enhancing the generalization ability of the model under missing‑label conditions. The effectiveness of our proposed method is thoroughly demonstrated through extensive comparative experiments with advanced methods on five benchmark datasets. Code is available at https://github.com/xuy11/SCSD.
Authors:Ao Li, Jiawei Sun, Le Dong, Zhenyu Wang, Weisheng Dong
Abstract:
Real‑world exposure correction is fundamentally challenged by spatially non‑uniform degradations, where diverse exposure errors frequently coexist within a single image. However, existing exposure correction methods are still largely developed under a predominantly uniform assumption. Architecturally, they typically rely on globally aggregated modulation signals that capture only the overall exposure trend. From the optimization perspective, conventional reconstruction losses are usually derived under a shared global scale, thus overlooking the spatially varying correction demands across regions. To address these limitations, we propose a new exposure correction paradigm explicitly designed for spatial non‑uniformity. Specifically, we introduce a Spatial Signal Encoder to predict spatially adaptive modulation weights, which are used to guide multiple look‑up tables for image transformation, together with an HSL‑based compensation module for improved color fidelity. Beyond the architectural design, we propose an uncertainty‑inspired non‑uniform loss that dynamically allocates the optimization focus based on local restoration uncertainties, better matching the heterogeneous nature of real‑world exposure errors. Extensive experiments demonstrate that our method achieves superior qualitative and quantitative performance compared with state‑of‑the‑art methods. Code is available at https://github.com/FALALAS/rethinkingEC.
Authors:Dat Nguyen, Enjie Ghorbel, Anis Kacem, Marcella Astrid, Djamila Aouada
Abstract:
In this paper, we propose Localized Artifact Attention X (LAA‑X), a novel deepfake detection framework that is both robust to high‑quality forgeries and capable of generalizing to unseen manipulations. Existing approaches typically rely on binary classifiers coupled with implicit attention mechanisms, which often fail to generalize beyond known manipulations. In contrast, LAA‑X introduces an explicit attention strategy based on a multi‑task learning framework combined with blending‑based data synthesis. Auxiliary tasks are designed to guide the model toward localized, artifact‑prone (i.e., vulnerable) regions. The proposed framework is compatible with both CNN and transformer backbones, resulting in two different versions, namely, LAA‑Net and LAA‑Former, respectively. Despite being trained only on real and pseudo‑fake samples, LAA‑X competes with state‑of‑the‑art methods across multiple benchmarks. Code and pre‑trained weights for LAA‑Net\footnotehttps://github.com/10Ring/LAA‑Net and LAA‑Former\footnotehttps://github.com/10Ring/LAA‑Former are publicly available.
Authors:Taiping Qu, Hongkai Zhang, Lantian Zhang, Can Zhao, Nan Zhang, Hui Wang, Zhen Zhou, Mingye Zou, Kairui Bo, Pengfei Zhao, Xingxing Jin, Zixian Su, Kun Jiang, Huan Liu, Yu Du, Maozhou Wang, Ruifang Yan, Zhongyuan Wang, Tiejun Huang, Lei Xu, Henggui Zhang
Abstract:
Cardiac magnetic resonance (CMR) is a cornerstone for diagnosing cardiovascular disease. However, it remains underutilized due to complex, time‑consuming interpretation across multi‑sequences, phases, quantitative measures that heavily reliant on specialized expertise. Here, we present BAAI Cardiac Agent, a multimodal intelligent system designed for end‑to‑end CMR interpretation. The agent integrates specialized cardiac expert models to perform automated segmentation of cardiac structures, functional quantification, tissue characterization and disease diagnosis, and generates structured clinical reports within a unified workflow. Evaluated on CMR datasets from two hospitals (2413 patients) spanning 7‑types of major cardiovascular diseases, the agent achieved an area under the receiver‑operating‑characteristic curve exceeding 0.93 internally and 0.81 externally. In the task of estimating left ventricular function indices, the results generated by this system for core parameters such as ejection fraction, stroke volume, and left ventricular mass are highly consistent with clinical reports, with Pearson correlation coefficients all exceeding 0.90. The agent outperformed state‑of‑the‑art models in segmentation and diagnostic tasks, and generated clinical reports showing high concordance with expert radiologists (six readers across three experience levels). By dynamically orchestrating expert models for coordinated multimodal analysis, this agent framework enables accurate, efficient CMR interpretation and highlights its potentials for complex clinical imaging workflows. Code is available at https://github.com/plantain‑herb/Cardiac‑Agent.
Authors:Junsheng Zhou, Zhifan Yang, Liang Han, Wenyuan Zhang, Kanle Shi, Shenkun Xu, Yu-Shen Liu
Abstract:
This paper tackles the challenge of recovering 4D dynamic scenes from videos captured by as few as four portable cameras. Learning to model scene dynamics for temporally consistent novel‑view rendering is a foundational task in computer graphics, where previous works often require dense multi‑view captures using camera arrays of dozens or even hundreds of views. We propose 4C4D, a novel framework that enables high‑fidelity 4D Gaussian Splatting from video captures of extremely sparse cameras. Our key insight lies that the geometric learning under sparse settings is substantially more difficult than modeling appearance. Driven by this observation, we introduce a Neural Decaying Function on Gaussian opacities for enhancing the geometric modeling capability of 4D Gaussians. This design mitigates the inherent imbalance between geometry and appearance modeling in 4DGS by encouraging the 4DGS gradients to focus more on geometric learning. Extensive experiments across sparse‑view datasets with varying camera overlaps show that 4C4D achieves superior performance over prior art. Project page at: https://junshengzhou.github.io/4C4D.
Authors:Nahyuk Lee, Zhiang Chen, Marc Pollefeys, Sunghwan Hong
Abstract:
Flow‑matching methods for 3D shape assembly learn point‑wise velocity fields that transport parts toward assembled configurations, yet they receive no explicit guidance about which cross‑part interactions should drive the motion. We introduce TORA, a topology‑first representation alignment framework that distills relational structure from a frozen pretrained 3D encoder into the flow‑matching backbone during training. We first realize this via simple instantiation, token‑wise cosine matching, which injects the learned geometric descriptors from the teacher representation. We then extend to employ a Centered Kernel Alignment (CKA) loss to match the similarity structure between student and teacher representations for enhanced topological alignment. Through systematic probing of diverse 3D encoders, we show that geometry‑ and contact‑centric teacher properties, not semantic classification ability, govern alignment effectiveness, and that alignment is most beneficial at later transformer layers where spatial structure naturally emerges. TORA introduces zero inference overhead while yielding two consistent benefits: faster convergence (up to 6.9×) and improved accuracy in‑distribution, along with greater robustness under domain shift. Experiments on five benchmarks spanning geometric, semantic, and inter‑object assembly demonstrate state‑of‑the‑art performance, with particularly pronounced gains in zero‑shot transfer to unseen real‑world and synthetic datasets. Project page: https://nahyuklee.github.io/tora.
Authors:Hang Wang, Chao Shen, Lei Zhang, Zhi-Qi Cheng
Abstract:
AI‑generated videos (AIGVs) have achieved unprecedented photorealism, posing severe threats to digital forensics. Existing AIGV detectors focus mainly on localized artifacts or short‑term temporal inconsistencies, thus often fail to capture the underlying generative logic governing global temporal evolution, limiting AIGV detection performance. In this paper, we identify a distinctive fingerprint in AIGVs, termed anomalous temporal self‑similarity (ATSS). Unlike real videos that exhibit stochastic natural dynamics, AIGVs follow deterministic anchor‑driven trajectories (e.g., text or image prompts), inducing unnaturally repetitive correlations across visual and semantic domains. To exploit this, we propose the ATSS method, a multimodal detection framework that exploits this insight via a triple‑similarity representation and a cross‑attentive fusion mechanism. Specifically, ATSS reconstructs semantic trajectories by leveraging frame‑wise descriptions to construct visual, textual, and cross‑modal similarity matrices, which jointly quantify the inherent temporal anomalies. These matrices are encoded by dedicated Transformer encoders and integrated via a bidirectional cross‑attentive fusion module to effectively model intra‑ and inter‑modal dynamics. Extensive experiments on four large‑scale benchmarks, including GenVideo, EvalCrafter, VideoPhy, and VidProM, demonstrate that ATSS significantly outperforms state‑of‑the‑art methods in terms of AP, AUC, and ACC metrics, exhibiting superior generalization across diverse video generation models. Code and models of ATSS will be released at https://github.com/hwang‑cs‑ime/ATSS.
Authors:Haoyu Li, Tingyan Wen, Lin Qi, Zhe Wu, Yihuang Chen, Xing Zhou, Lifei Zhu, Xueqian Wang, Kai Zhang
Abstract:
Diffusion models produce high‑quality text‑to‑image results, but their iterative denoising is computationally expensive.Distribution Matching Distillation (DMD) emerges as a promising path to few‑step distillation, but suffers from diversity collapse and fidelity degradation when reduced to two steps or fewer. We present 1.x‑Distill, the first fractional‑step distillation framework that breaks the integer‑step constraint of prior few‑step methods and establishes 1.x‑step generation as a practical regime for distilled diffusion models.Specifically, we first analyze the overlooked role of teacher CFG in DMD and introduce a simple yet effective modification to suppress mode collapse. Then, to improve performance under extreme steps, we introduce Stagewise Focused Distillation, a two‑stage strategy that learns coarse structure through diversity‑preserving distribution matching and refines details with inference‑consistent adversarial distillation. Furthermore, we design a lightweight compensation module for Distill‑‑Cache co‑Training, which naturally incorporates block‑level caching into our distillation pipeline.Experiments on SD3‑Medium and SD3.5‑Large show that 1.x‑Distill surpasses prior few‑step methods, achieving better quality and diversity at 1.67 and 1.74 effective NFEs, respectively, with up to 33x speedup over original 28x2 NFE sampling.
Authors:Indar Kumar, Girish Karhana, Sai Krishna Jasti, Ankit Hemant Lade
Abstract:
Frozen pretrained image representations are widely used for transfer learning: a backbone is kept fixed, feature vectors are extracted, and a lightweight classifier is trained on top. This pipeline usually feeds the full feature vector to the classifier, even when the target task has far fewer classes than the pretraining task. We revisit a classical alternative: supervised dimensionality reduction with Linear Discriminant Analysis (LDA) before linear probing.
We evaluate ten dimensionality‑reduction strategies on frozen features from six backbones ‑‑ ResNet‑18, ResNet‑50, MobileNetV3‑Small, EfficientNet‑B0, ViT‑B/16, and DINOv2‑ViT‑S/14 ‑‑ across CIFAR‑100, Tiny ImageNet, and CUB‑200‑2011. Under a fixed logistic‑regression protocol, LDA improves accuracy over full features in 11 of 12 coarse‑grained configurations, with gains up to 4.5 percentage points while reducing feature dimensionality by 48‑87%. The same projection consistently hurts on fine‑grained CUB‑200, where full features win across all six backbones. This establishes a practical boundary condition: LDA is useful when class‑level structure is coarse enough to be captured by mean‑separating directions, but it can discard subtle cues needed for fine‑grained recognition.
We also compare LDA with PCA, PCA+LDA, regularized LDA, Local Fisher Discriminant Analysis, Neighbourhood Components Analysis, and three lightweight LDA extensions. The results show that plain LDA offers the best accuracy‑cost tradeoff for most coarse‑grained settings, while more complex supervised reduction methods rarely justify their additional cost. Overall, the study provides concrete guidance for when post‑hoc supervised projection should, and should not, be inserted into frozen‑feature image classification pipelines.
Authors:Lei Zhou, Haoyu Wu, Akshat Dave, Dimitris Samaras
Abstract:
We introduce a test‑time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or modifying pre‑trained image‑only networks. Rather than feeding priors into the architecture, we cast them as constraints on the predictions and optimize the network at inference time. The optimization loss consists of a self‑supervised objective and prior penalty terms. The self‑supervised objective captures the compatibility among multi‑view predictions and is implemented using photometric or geometric loss between renderings from other views and each view itself. Any available priors are converted into penalty terms on the corresponding output modalities. Across a series of 3D vision benchmarks, including point map estimation and camera pose estimation, our method consistently improves performance over base MVTs by a large margin. On the ETH3D, 7‑Scenes, and NRGBD datasets, our method reduces the point‑map distance error by more than half compared with the base image‑only models. Our method also outperforms retrained prior‑aware feed‑forward methods, demonstrating the effectiveness of our test‑time constrained optimization (TCO) framework for incorporating priors into 3D vision tasks.
Authors:Hessen Bougueffa Eutamene, Abdellah Zakaria Sellam, Abdelmalik Taleb-Ahmed, Abdenour Hadid
Abstract:
Detecting AI‑generated images remains a significant challenge because detectors trained on specific generators often fail to generalize to unseen models; however, while pixel‑level artifacts vary across models, frequency‑domain signatures exhibit greater consistency, providing a promising foundation for cross‑generator detection. To address this, we propose SPARK‑IL, a retrieval‑augmented framework that combines dual‑path spectral analysis with incremental learning by utilizing a partially frozen ViT‑L/14 encoder for semantic representations alongside a parallel path for raw RGB pixel embeddings. Both paths undergo multi‑band Fourier decomposition into four frequency bands, which are individually processed by Kolmogorov‑Arnold Networks (KAN) with mixture‑of‑experts for band‑specific transformations before the resulting spectral embeddings are fused via cross‑attention with residual connections. During inference, this fused embedding retrieves the k nearest labeled signatures from a Milvus database using cosine similarity to facilitate predictions via majority voting, while an incremental learning strategy expands the database and employs elastic weight consolidation to preserve previously learned transformations. Evaluated on the UniversalFakeDetect benchmark across 19 generative models ‑‑ including GANs, face‑swapping, and diffusion methods ‑‑ SPARK‑IL achieves a 94.6% mean accuracy, with the code to be publicly released at https://github.com/HessenUPHF/SPARK‑IL.
Authors:Felix Stillger, Lukas Hahn, Frederik Hasecke, Tobias Meisen
Abstract:
Camera extrinsic calibration is a fundamental task in computer vision. However, precise relative pose estimation in constrained, highly distorted environments, such as in‑cabin automotive monitoring (ICAM), remains challenging. We present InCaRPose, a Transformer‑based architecture designed for robust relative pose prediction between image pairs, which can be used for camera extrinsic calibration. By leveraging frozen backbone features such as DINOv3 and a Transformer‑based decoder, our model effectively captures the geometric relationship between a reference and a target view. Unlike traditional methods, our approach achieves absolute metric‑scale translation within the physically plausible adjustment range of in‑cabin camera mounts in a single inference step, which is critical for ICAM, where accurate real‑world distances are required for safety‑relevant perception. We specifically address the challenges of highly distorted fisheye cameras in automotive interiors by training exclusively on synthetic data. Our model is capable of generalization to real‑world cabin environments without relying on the exact same camera intrinsics and additionally achieves competitive performance on the public 7‑Scenes dataset. Despite having limited training data, InCaRPose maintains high precision in both rotation and translation, even with a ViT‑Small backbone. This enables real‑time performance for time‑critical inference, such as driver monitoring in supervised autonomous driving. We release our real‑world In‑Cabin‑Pose test dataset consisting of highly distorted vehicle‑interior images and our code at https://github.com/felixstillger/InCaRPose.
Authors:Mohammad Heydari, Wei Dong, Shahram Shirani, Jun Chen, Han Zhou
Abstract:
Nighttime image dehazing remains a challenging low‑level vision problem due to the joint presence of haze, glow, non‑uniform illumination, color distortion, and sensor noise, which often invalidate assumptions commonly used in daytime dehazing. To address these challenges, we propose HistoFusionNet, a transformer‑enhanced architecture tailored for nighttime image dehazing by combining histogram‑guided representation learning with frequency‑adaptive feature refinement. Built upon a multi‑scale encoder‑decoder backbone, our method introduces histogram transformer blocks that model long‑range dependencies by grouping features according to their dynamic‑range characteristics, enabling more effective aggregation of similarly degraded regions under complex nighttime lighting. To further improve restoration fidelity, we incorporate a frequency‑aware refinement branch that adaptively exploits complementary low‑ and high‑frequency cues, helping recover scene structures, suppress artifacts, and enhance local details. This design yields a unified framework that is particularly well suited to the heterogeneous degradations encountered in real nighttime hazy scenes. Extensive experiments and highly competitive performance of our method on the NTIRE 2026 Nighttime Image Dehazing Challenge benchmark demonstrate the effectiveness of the proposed method. Our team ranked 1st among 22 participating teams, highlighting the robustness and competitive performance of HistoFusionNet. The code is available at: https://github.com/heydarimo/Night‑Time‑Dehazing
Authors:Guiyu Zhang, Yabo Chen, Xunzhi Xiang, Junchao Huang, Zhongyu Wang, Li Jiang
Abstract:
Controlling both camera motion and object dynamics is essential for coherent and expressive video generation, yet current methods typically handle only one motion type or rely on ambiguous 2D cues that entangle camera‑induced parallax with true object movement. We present SymphoMotion, a unified motion‑control framework that jointly governs camera trajectories and object dynamics within a single model. SymphoMotion features a Camera Trajectory Control mechanism that integrates explicit camera paths with geometry‑aware cues to ensure stable, structurally consistent viewpoint transitions, and an Object Dynamics Control mechanism that combines 2D visual guidance with 3D trajectory embeddings to enable depth‑aware, spatially coherent object manipulation. To support large‑scale training and evaluation, we further construct RealCOD‑25K, a comprehensive real‑world dataset containing paired camera poses and object‑level 3D trajectories across diverse indoor and outdoor scenes, addressing a key data gap in unified motion control. Extensive experiments and user studies show that SymphoMotion significantly outperforms existing methods in visual fidelity, camera controllability, and object‑motion accuracy, establishing a new benchmark for unified motion control in video generation. Codes and data are publicly available at https://grenoble‑zhang.github.io/SymphoMotion/.
Authors:Tianci Luo, Haohao Pan, Jinpeng Wang, Niu Lian, Xinrui Chen, Bin Chen, Shu-Tao Xia, Chun Yuan
Abstract:
Visual in‑context learning (VICL) enables visual foundation models to handle multiple tasks by steering them with demonstrative prompts. The choice of such prompts largely influences VICL performance, standing out as a key challenge. Prior work has made substantial progress on prompt retrieval and reranking strategies, but mainly focuses on prompt images while overlooking labels. We reveal these approaches sometimes get visually similar but label‑inconsistent prompts, which potentially degrade VICL performance. On the other hand, higher label consistency between query and prompts preferably indicates stronger VICL results. Motivated by these findings, we develop a framework named LaPR (Label‑aware Prompt Retrieval), which highlights the role of labels in prompt selection. Our framework first designs an image‑label joint representation for prompts to incorporate label cues explicitly. Besides, to handle unavailable query labels at test time, we introduce a mixture‑of‑expert mechanism to the dual encoders with query‑adaptive routing. Each expert is expected to capture a specific label mode, while the router infers query‑adaptive mixture weights and helps to learn label‑aware representation. We carefully design alternative optimization for experts and router, with a VICL performance‑guided contrastive loss and a label‑guided contrastive loss, respectively. Extensive experiments show promising and consistent improvement of LaPR on in‑context segmentation, detection, and colorization tasks. Moreover, LaPR generalizes well across feature extractors and cross‑fold scenarios, suggesting the importance of label utilization in prompt retrieval for VICL. Code is available at https://github.com/luotc‑why/CVPR26‑LaPR.
Authors:Jun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang, Yaowei Wang, Shu-Tao Xia, Bin Chen
Abstract:
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos based on text queries that describe only partial events. Existing methods suffer from incomplete global contextual perception, struggling with query ambiguity and local noise induced by spurious responses. To address these issues, we propose DreamPRVR, which adopts a coarse‑to‑fine representation learning paradigm. The model first generates global contextual semantic registers as coarse‑grained highlights spanning the entire video and then concentrates on fine‑grained similarity optimization for precise cross‑modal matching. Concretely, these registers are generated by initializing from the video‑centric distribution produced by a probabilistic variational sampler and then iteratively refined via a text‑supervised truncated diffusion model. During this process, textual semantic structure learning constructs a well‑formed textual latent space, enhancing the reliability of global perception. The registers are then adaptively fused with video tokens through register‑augmented Gaussian attention blocks, enabling context‑aware feature learning. Extensive experiments show that DreamPRVR outperforms state‑of‑the‑art methods. Code is released at https://github.com/lijun2005/CVPR26‑DreamPRVR.
Authors:Yunyao Yu, Zhengxian Wu, Zhuohong Chen, Hangrui Xu, Zirui Liao, Xiangwen Deng, Zhifang Liu, Senyuan Shi, Haoqian Wang
Abstract:
In the unsupervised self‑evolution of Multimodal Large Language Models, the quality of feedback signals during post‑training is pivotal for stable and effective learning. However, existing self‑evolution methods predominantly rely on majority voting to select the most frequent output as the pseudo‑golden answer, which may stem from the model's intrinsic biases rather than guaranteeing the objective correctness of the reasoning paths. To counteract the degradation, we propose Continuous Softened Retracing reSampling (CSRS) in MLLM self‑evolution. Specifically, we introduce a Retracing Re‑inference Mechanism (RRM) that the model re‑inferences from anchor points to expand the exploration of long‑tail reasoning paths. Simultaneously, we propose Softened Frequency Reward (SFR), which replaces binary rewards with continuous signals, calibrating reward based on the answers' frequency across sampled reasoning sets. Furthermore, incorporated with Visual Semantic Perturbation (VSP), CSRS ensures the model prioritizes mathematical logic over visual superficiality. Experimental results demonstrate that CSRS significantly enhances the reasoning performance of Qwen2.5‑VL‑7B on benchmarks such as MathVision. We achieve state‑of‑the‑art (SOTA) results in unsupervised self‑evolution on geometric tasks. Our code is avaible at https://github.com/yyy195/CSRS.
Authors:Haofeng Liu, Ziyue Wang, Alex Y. W. Kong, Guanyi Qin, Yunqiu Xu, Chang Han Low, Mingqi Gao, Lap Yan Lennon Chan, Yueming Jin
Abstract:
Surgical video segmentation is fundamental to computer‑assisted surgery. In practice, surgeons need to dynamically specify targets throughout extended procedures, using heterogeneous cues such as visual selections, textual expressions, or audio instructions. However, existing Promptable Video Object Segmentation (PVOS) methods are typically restricted to a single prompt modality and rely on coupled frameworks that cause optimization interference between target initialization and tracking. Moreover, these methods produce hallucinated predictions when the target is absent and suffer from accumulated mask drift without failure recovery. To address these challenges, we present UniSurgSAM, a unified PVOS model enabling reliable surgical video segmentation through visual, textual, or audio prompts. Specifically, UniSurgSAM employs a decoupled two‑stage framework that independently optimizes initialization and tracking to resolve the optimization interference. Within this framework, we introduce three key designs for reliability: presence‑aware decoding that models target absence to suppress hallucinations; boundary‑aware long‑term tracking that prevents mask drift over extended sequences; and adaptive state transition that closes the loop between stages for failure recovery. Furthermore, we establish a multi‑modal and multi‑granular benchmark from four public surgical datasets with precise instance‑level masklets. Extensive experiments demonstrate that UniSurgSAM achieves state‑of‑the‑art performance in real time across all prompt modalities and granularities, providing a practical foundation for computer‑assisted surgery. Code and datasets will be available at https://jinlab‑imvr.github.io/UniSurgSAM.
Authors:Peter Yongho Kim, Juhyeon Park, Jungwoo Park, Jubin Choi, Jungwoo Seo, Jiook Cha, Taesup Moon
Abstract:
Modeling long‑range spatiotemporal dynamics in functional Magnetic Resonance Imaging (fMRI) remains a key challenge due to the high dimensionality of the four‑dimensional signals. Prior voxel‑based models, although demonstrating excellent performance and interpretation capabilities, are constrained by prohibitive memory demands and thus can only capture limited temporal windows. To address this, we propose TABLeT (Two‑dimensionally Autoencoded Brain Latent Transformer), a novel approach that tokenizes fMRI volumes using a pre‑trained 2D natural image autoencoder. Each 3D fMRI volume is compressed into a compact set of continuous tokens, enabling long‑sequence modeling with a simple Transformer encoder with limited VRAM. Across large‑scale benchmarks including the UK‑Biobank (UKB), Human Connectome Project (HCP), and ADHD‑200 datasets, TABLeT outperforms existing models in multiple tasks, while demonstrating substantial gains in computational and memory efficiency over the state‑of‑the‑art voxel‑based method given the same input. Furthermore, we develop a self‑supervised masked token modeling approach to pre‑train TABLeT, which improves the model's performance for various downstream tasks. Our findings suggest a promising approach for scalable and interpretable spatiotemporal modeling of brain activity. Our code is available at https://github.com/beotborry/TABLeT.
Authors:Viet Dung Nguyen, Yuhang Song, Anh Nguyen, Jamison Heard, Reynold Bailey, Alexander Ororbia
Abstract:
Robot reinforcement learning from demonstrations (RLfD) assumes that expert data is abundant; this is usually unrealistic in the real world given data scarcity as well as high collection cost. Furthermore, imitation learning algorithms assume that the data is independently and identically distributed, which ultimately results in poorer performance as gradual errors emerge and compound within test‑time trajectories. We address these issues by introducing the "master your own expertise" (MYOE) framework, a self‑imitation framework that enables robotic agents to learn complex behaviors from limited demonstration data samples. Inspired by human perception and action, we propose and design what we call the queryable mixture‑of‑preferences state space model (QMoP‑SSM), which estimates the desired goal at every time step. These desired goals are used in computing the "preference regret", which is used to optimize the robot control policy. Our experiments demonstrate the robustness, adaptability, and out‑of‑sample performance of our agent compared to other state‑of‑the‑art RLfD schemes. The GitHub repository that supports this work can be found at: https://github.com/rxng8/neurorobot‑preference‑regret‑learning.
Authors:Zilin Huang, Zhengyang Wan, Zihao Sheng, Boyue Wang, Junwei You, Yue Leng, Sikai Chen
Abstract:
Deploying reinforcement learning policies trained in simulation to real autonomous vehicles remains a fundamental challenge, particularly for VLM‑guided RL frameworks whose policies are typically learned with simulator‑native observations and simulator‑coupled action semantics that are unavailable on physical platforms. This paper presents Sim2Real‑AD, a modular framework for zero‑shot sim‑to‑real transfer of CARLA‑trained VLM‑guided RL policies to full‑scale vehicles without any real‑world RL training data. The framework decomposes the transfer problem into four components: a Geometric Observation Bridge (GOB) that converts monocular front‑view images into simulator‑compatible bird's‑eye‑view (BEV) observations, a Physics‑Aware Action Mapping (PAM) that translates policy outputs into platform‑agnostic physical commands, a Two‑Phase Progressive Training (TPT) strategy that stabilizes adaptation by separating action‑space and observation‑space transfer, and a Real‑time Deployment Pipeline (RDP) that integrates perception, policy inference, control conversion, and safety monitoring for closed‑loop execution. Simulation experiments show that the framework preserves the relative performance ordering of representative RL algorithms across different reward paradigms and validate the contribution of each module. Zero‑shot deployment on a full‑scale Ford E‑Transit achieves success rates of 90%, 80%, and 75% in car‑following, obstacle avoidance, and stop‑sign interaction scenarios, respectively. To the best of our knowledge, this study is among the first to demonstrate zero‑shot closed‑loop deployment of a CARLA‑trained VLM‑guided RL policy on a full‑scale real vehicle without any real‑world RL training data. The demo video and code are available at: https://zilin‑huang.github.io/Sim2Real‑AD‑website/.
Authors:Meng'en Qin, Zhe Li, Xiaohui Yang
Abstract:
Genotype‑by‑Environment (GxE) interactions influence the performance of genotypes across diverse environments, reducing the predictability of phenotypes in target environments. In‑depth analysis of GxE interactions facilitates the identification of how genetic advantages or defects are expressed or suppressed under specific environmental conditions, thereby enabling genetic selection and enhancing breeding practices. This paper introduces two key models for GxE interaction research. Specifically, it includes significance analysis based on the mixed effect model to determine whether genes or GxE interactions significantly affect phenotypic traits; stability analysis, which further investigates the interactive relationships between genes and environments, as well as the relative superiority or inferiority of genotypes across environments. Additionally, this paper presents RGxEStat, a lightweight interactive tool, which is developed by the authors and integrates the construction, solution, and visualization of the aforementioned models. Designed to eliminate the need for breeders and agronomists to learn complex SAS or R programming, RGxEStat provides a user‑friendly interface for streamlined breeding data analysis, significantly accelerating research cycles. Codes and datasets are available at https://github.com/mason‑ching/RGxEStat.
Authors:Asmita Yuki Pritha, Jason Xu, Daniel Ding, Justin Li, Aryana Hou, Xin Wang, Shu Hu
Abstract:
Deep learning models for COVID‑19 detection from chest CT scans generally perform well when the training and test data originate from the same institution, but they often struggle when scans are drawn from multiple centres with differing scanners, imaging protocols, and patient populations. One key reason is that existing methods treat COVID‑19 classification as the sole training objective, without accounting for the data source of each scan. As a result, the learned representations tend to be biased toward centres that contribute more training data. To address this, we propose a multi‑task learning approach in which the model is trained to predict both the COVID‑19 diagnosis and the originating data centre. The two tasks share an EfficientNet‑B7 backbone, which encourages the feature extractor to learn representations that hold across all four participating centres. Since the training data is not evenly distributed across sources, we apply a logit‑adjusted cross‑entropy loss [1] to the source classification head to prevent underrepresented centres from being overlooked. Our pre‑processing follows the SSFL framework with KDS [2], selecting eight representative slices per scan. Our method achieves an F1 score of 0.9098 and an AUC‑ROC of 0.9647 on a validation set of 308 scans. The code is publicly available at https://github.com/Purdue‑M2/‑multisource‑covid‑ct.
Authors:Zhenghao Chen, Huiqun Wang, Di Huang
Abstract:
Multimodal large language models (MLLMs) are increasingly being applied to spatial cognition tasks, where they are expected to understand and interact with complex environments. Most existing works improve spatial reasoning by introducing 3D priors or geometric supervision, which enhances performance but incurs substantial data preparation and alignment costs. In contrast, purely 2D approaches often struggle with multi‑frame spatial reasoning due to their limited ability to capture cross‑frame spatial relationships. To address these limitations, we propose EgoMind, a Chain‑of‑Thought framework that enables geometry‑free spatial reasoning through Role‑Play Caption, which jointly constructs a coherent linguistic scene graph across frames, and Progressive Spatial Analysis, which progressively reasons toward task‑specific questions. With only 5K auto‑generated SFT samples and 20K RL samples, EgoMind achieves competitive results on VSI‑Bench, SPAR‑Bench, SITE‑Bench, and SPBench, demonstrating its effectiveness in strengthening the spatial reasoning capabilities of MLLMs and highlighting the potential of linguistic reasoning for spatial cognition. Code and data are released at https://github.com/Hyggge/EgoMind.
Authors:Bingliang Li, Zhenhong Sun, Jiaming Bian, Yuehao Wu, Yifu Wang, Hongdong Li, Yatao Bian, Huadong Mo, Daoyi Dong
Abstract:
Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: inter‑shot consistency and explicit editability. While 2D diffusion‑based generators produce vivid imagery, they often suffer from identity drift along with limited geometric control; conversely, traditional 3D animation workflows are consistent and editable but require expert‑heavy, labor‑intensive authoring. We present StoryBlender, a grounded 3D storyboard generation framework governed by a Story‑centric Reflection Scheme. At its core, we propose the StoryBlender system, which is built on a three‑stage pipeline: (1) Semantic‑Spatial Grounding, to construct a continuity memory graph to decouple global assets from shot‑specific variables for long‑horizon consistency; (2) Canonical Asset Materialization, to instantiate entities in a unified coordinate space to maintain visual identity; and (3) Spatial‑Temporal Dynamics, to achieve layout design and cinematic evolution through visual metrics. By orchestrating multiple agents in a hierarchical manner within a verification loop, StoryBlender iteratively self‑corrects spatial hallucinations via engine‑verified feedback. The resulting native 3D scenes support direct, precise editing of cameras and visual assets while preserving unwavering multi‑shot continuity. Experiments demonstrate that StoryBlender significantly improves consistency and editability over both diffusion‑based and 3D‑grounded baselines. Code, data, and demonstration video will be available on https://engineeringai‑lab.github.io/StoryBlender/
Authors:Nanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang, Hong Li, Yong-Lu Li
Abstract:
While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high‑level physics reasoning. In this work, we investigate the first step of physical reasoning, i.e., intuitive physics understanding, revealing substantial limitations in understanding the dynamics of continuum objects. To isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state‑of‑the‑art MLLMs perform poorly on these foundational tasks. To address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi‑task fine‑tuning framework. SDF substantially improves performance, achieving up to 20.7% gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost‑efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene‑Dynamic‑Field.
Authors:Chushan Zhang, Ruihan Lu, Jinguang Tong, Yikai Wang, Hongdong Li
Abstract:
Leveraging 3D information within Multimodal Large Language Models (MLLMs) has recently shown significant advantages for indoor scene understanding. However, existing methods, including those using explicit ground‑truth 3D positional encoding and those grafting external 3D foundation models for implicit geometry, struggle with the trade‑off in 2D‑3D representation fusion, leading to suboptimal deployment. To this end, we propose 3D‑Implicit Depth Emergence, a method that reframes 3D perception as an emergent property derived from geometric self‑supervision rather than explicit encoding. Our core insight is the Implicit Geometric Emergence Principle: by strategically leveraging privileged geometric supervision through mechanisms like a fine‑grained geometry validator and global representation constraints, we construct an information bottleneck. This bottleneck forces the model to maximize the mutual information between visual features and 3D structures, allowing 3D awareness to emerge naturally within a unified visual representation. Unlike existing approaches, our method enables 3D perception to emerge implicitly, disentangling features in dense regions and, crucially, eliminating depth and pose dependencies during inference with zero latency overhead. This paradigm shift from external grafting to implicit emergence represents a fundamental rethinking of 3D knowledge integration in visual‑language models. Extensive experiments demonstrate that our method surpasses SOTA on multiple 3D scene understanding benchmarks. Our approach achieves a 55% reduction in inference latency while maintaining strong performance across diverse downstream tasks, underscoring the effectiveness of meticulously designed auxiliary objectives for dependency‑free 3D understanding. Source code can be found at github.com/ChushanZhang/3D‑IDE.
Authors:Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, Xiangtao Kong, Jixin Zhao, Shihao Wang, Lei Zhang
Abstract:
Most of the recent generative image super‑resolution (SR) methods rely on adapting large text‑to‑image (T2I) diffusion models pretrained on web‑scale text‑image data. While effective, this paradigm starts from a generic T2I generator, despite that SR is fundamentally a low‑resolution (LR) input‑conditioned image restoration task. In this work, we investigate whether an SR model trained purely on visual data can rival T2I‑based ones. To this end, we propose VOSR, a Vision‑Only generative framework for SR. We first extract semantically rich and spatially grounded features from the LR input using a pretrained vision encoder as visual semantic guidance. We then revisit classifier‑free guidance for training generative models and show that the standard unconditional branch is ill‑suited to restoration models trained from scratch. We therefore replace it with a restoration‑oriented guidance strategy that preserves weak LR anchors. Built upon these designs, we first train a multi‑step VOSR model from scratch and then distill it into a one‑step model for efficient inference. VOSR requires less than one‑tenth of the training cost of representative T2I‑based SR methods, yet in both multi‑step and one‑step settings, it achieves competitive or even better perceptual quality and efficiency, while producing more faithful structures with fewer hallucinations on both synthetic and real‑world benchmarks. Our results, for the first time, show that high‑quality generative SR can be achieved without multimodal pretraining. The code and models can be found at https://github.com/cswry/VOSR.
Authors:Fengbei Liu, Sunwoo Kwak, Hao Phung, Nusrat Binta Nizam, Ilan Richter, Nir Uriel, Hadar Averbuch-Elor, Daborah Estrin, Mert R. Sabuncu
Abstract:
Non‑contrast chest CTs offer a rich opportunity for both conventional pulmonary and opportunistic extra‑pulmonary screening. While Multi‑Task Learning (MTL) can unify these diverse tasks, standard hard‑parameter sharing approaches are often suboptimal for modeling distinct pathologies. We propose HyperCT, a framework that dynamically adapts a Vision Transformer backbone via a Hypernetwork. To ensure computational efficiency, we integrate Low‑Rank Adaptation (LoRA), allowing the model to regress task‑specific low‑rank weight updates rather than full parameters. Validated on a large‑scale dataset of radiological and cardiological tasks, \method outperforms various strong baselines, offering a unified, parameter‑efficient solution for holistic patient assessment. Our code is available at https://github.com/lfb‑1/HyperCT.
Authors:Jiekai Wu, Rong Fu, Chuangqi Li, Zijian Zhang, Guangxin Wu, Hao Zhang, Shiyin Lin, Jianyuan Ni, Yang Li, Dongxu Zhang, Amir H. Gandomi, Simon Fong, Pengbin Feng
Abstract:
Remote sensing segmentation in real deployment is inherently continual: new semantic categories emerge, and acquisition conditions shift across seasons, cities, and sensors. Despite recent progress, many incremental approaches still treat training steps as isolated updates, which leaves representation drift and forgetting insufficiently controlled. We present ProtoFlow, a time‑aware prototype dynamics framework that models class prototypes as trajectories and learns their evolution with an explicit temporal vector field. By jointly enforcing low‑curvature motion and inter‑class separation, ProtoFlow stabilizes prototype geometry throughout incremental learning. Experiments on standard class‑ and domain‑incremental remote sensing benchmarks show consistent gains over strong baselines, including up to 1.5‑2.0 points improvement in mIoUall, together with reduced forgetting. These results suggest that explicitly modeling temporal prototype evolution is a practical and interpretable strategy for robust continual remote sensing segmentation. Open‑source code:https://github.com/dudududke/protoflow.
Authors:Peiyan Li, Yixiang Chen, Yuan Xu, Jiabing Yang, Xiangnan Wu, Jun Guo, Nan Sun, Long Qian, Xinghang Li, Xin Xiao, Jing Liu, Nianfeng Liu, Tao Kong, Yan Huang, Liang Wang, Tieniu Tan
Abstract:
Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies overlook one or both. They typically rely on 2D visual observations and backbones pretrained on static image‑‑text pairs, resulting in high data requirements and limited understanding of environment dynamics. To address this, we introduce MV‑VDP, a multi‑view video diffusion policy that jointly models the 3D spatio‑temporal state of the environment. The core idea is to simultaneously predict multi‑view heatmap videos and RGB videos, which 1) align the representation format of video pretraining with action finetuning, and 2) specify not only what actions the robot should take, but also how the environment is expected to evolve in response to those actions. Extensive experiments show that MV‑VDP enables data‑efficient, robust, generalizable, and interpretable manipulation. With only ten demonstration trajectories and without additional pretraining, MV‑VDP successfully performs complex real‑world tasks, demonstrates strong robustness across a range of model hyperparameters, generalizes to out‑of‑distribution settings, and predicts realistic future videos. Experiments on Meta‑World and real‑world robotic platforms demonstrate that MV‑VDP consistently outperforms video‑prediction‑‑based, 3D‑based, and vision‑‑language‑‑action models, establishing a new state of the art in data‑efficient multi‑task manipulation.
Authors:Wenfeng Zhang, Jun Ni, Yue Meng, Xiaodong Pei, Wei Hu, Qibing Qin, Lei Huang
Abstract:
Object detection in unmanned aerial vehicle (UAV) images remains a highly challenging task, primarily caused by the complexity of background noise and the imbalance of target scales. Traditional methods easily struggle to effectively separate objects from intricate backgrounds and fail to fully leverage the rich multi‑scale information contained within images. To address these issues, we have developed a synergistic feature fusion network (SFFNet) with dual‑domain edge enhancement specifically tailored for object detection in UAV images. Firstly, the multi‑scale dynamic dual‑domain coupling (MDDC) module is designed. This component introduces a dual‑driven edge extraction architecture that operates in both the frequency and spatial domains, enabling effective decoupling of multi‑scale object edges from background noise. Secondly, to further enhance the representation capability of the model's neck in terms of both geometric and semantic information, a synergistic feature pyramid network (SFPN) is proposed. SFPN leverages linear deformable convolutions to adaptively capture irregular object shapes and establishes long‑range contextual associations around targets through the designed wide‑area perception module (WPM). Moreover, to adapt to the various applications or resource‑constrained scenarios, six detectors of different scales (N/S/M/B/L/X) are designed. Experiments on two challenging aerial datasets (VisDrone and UAVDT) demonstrate the outstanding performance of SFFNet‑X, achieving 36.8 AP and 20.6 AP, respectively. The lightweight models (N/S) also maintain a balance between detection accuracy and parameter efficiency. The code will be available at https://github.com/CQNU‑ZhangLab/SFFNet.
Authors:Xiaoran Zhang, Yu Liu, Jinyu Liang, Kangqiushi Li, Zhiwei Huang, Huaxin Xiao
Abstract:
Cross‑modal Thermal Geo‑localization (TG) provides a robust, all‑weather solution for Unmanned Aerial Vehicles (UAVs) in Global Navigation Satellite System (GNSS)‑denied environments. However, profound thermal‑visible modality gaps introduce severe feature ambiguity, systematically corrupting conventional coarse‑to‑fine registration. To dismantle this bottleneck, we propose SCC‑Loc, a unified Semantic‑Cascade‑Consensus localization framework. By sharing a single DINOv2 backbone across global retrieval and MINIMA_\textRoMa matching, it minimizes memory footprint and achieves zero‑shot, highly accurate absolute position estimation. Specifically, we tackle modality ambiguity by introducing three cohesive components. First, we design the Semantic‑Guided Viewport Alignment (SGVA) module to adaptively optimize satellite crop regions, effectively correcting initial spatial deviations. Second, we develop the Cascaded Spatial‑Adaptive Texture‑Structure Filtering (C‑SATSF) mechanism to explicitly enforce geometric consistency, thereby eradicating dense cross‑modal outliers. Finally, we propose the Consensus‑Driven Reliability‑Aware Position Selection (CD‑RAPS) strategy to derive the optimal solution through a synergy of physically constrained pose optimization. To address data scarcity, we construct Thermal‑UAV, a comprehensive dataset providing 11,890 diverse thermal queries referenced against a large‑scale satellite ortho‑photo and corresponding spatially aligned Digital Surface Model (DSM). Extensive experiments demonstrate that SCC‑Loc establishes a new state‑of‑the‑art, suppressing the mean localization error to 9.37 m and providing a 7.6‑fold accuracy improvement within a strict 5‑m threshold over the strongest baseline. Code and dataset are available at https://github.com/FloralHercules/SCC‑Loc.
Authors:Xingtong Ge, Yi Zhang, Yushi Huang, Dailan He, Xiahong Wang, Bingqi Ma, Guanglu Song, Yu Liu, Jun Zhang
Abstract:
Distilling video generation models to extremely low inference budgets (e.g., 2‑‑4 NFEs) is crucial for real‑time deployment, yet remains challenging. Trajectory‑style consistency distillation often becomes conservative under complex video dynamics, yielding an over‑smoothed appearance and weak motion. Distribution matching distillation (DMD) can recover sharp, mode‑seeking samples, but its local training signals do not explicitly regularize how denoising updates compose across timesteps, making composed rollouts prone to drift. To overcome this challenge, we propose Self‑Consistent Distribution Matching Distillation (SC‑DMD), which explicitly regularizes the endpoint‑consistent composition of consecutive denoising updates. For real‑time autoregressive video generation, we further treat the KV cache as a quality parameterized condition and propose Cache‑Distribution‑Aware training. This training scheme applies SC‑DMD over multi‑step rollouts and introduces a cache‑conditioned feature alignment objective that steers low‑quality outputs toward high‑quality references. Across extensive experiments on both non‑autoregressive backbones (e.g., Wan~2.1) and autoregressive real‑time paradigms (e.g., Self Forcing), our method, dubbed Salt, consistently improves low‑NFE video generation quality while remaining compatible with diverse KV‑cache memory mechanisms. Source code will be released at \hrefhttps://github.com/XingtongGe/Salthttps://github.com/XingtongGe/Salt.
Authors:Zicheng Zhang, Xiangting Meng, Ke Wu, Wenchao Ding
Abstract:
Recent progress in feed‑forward 3D Gaussian Splatting (3DGS) has notably improved rendering quality. However, the spatially uniform and highly redundant 3DGS map generated by previous feed‑forward 3DGS methods limits their integration into downstream reconstruction tasks. We propose SparseSplat, the first feed‑forward 3DGS model that adaptively adjusts Gaussian density according to scene structure and information richness of local regions, yielding highly compact 3DGS maps. To achieve this, we propose entropy‑based probabilistic sampling, generating large, sparse Gaussians in textureless areas and assigning small, dense Gaussians to regions with rich information. Additionally, we designed a specialized point cloud network that efficiently encodes local context and decodes it into 3DGS attributes, addressing the receptive field mismatch between the general 3DGS optimization pipeline and feed‑forward models. Extensive experimental results demonstrate that SparseSplat can achieve state‑of‑the‑art rendering quality with only 22% of the Gaussians and maintain reasonable rendering quality with only 1.5% of the Gaussians. Project page: https://victkk.github.io/SparseSplat‑page/.
Authors:Weixiong Sun, Xiang Yin, Chao Dong
Abstract:
Recent advances in generative AI raise the question of whether general‑purpose image editing models can serve as unified solutions for image restoration. We conduct a systematic evaluation of Nano Banana 2 across diverse scenes and degradations. Our results show that prompt design is critical, with concise prompts and explicit fidelity constraints achieving a better balance between reconstruction and perceptual quality. Nano Banana 2 achieves competitive full‑reference performance and is consistently preferred in user studies, while showing strong generalization in challenging scenarios. However, we observe a gap between perceptual quality and restoration fidelity, as the model tends to produce visually rich results with over‑enhanced details and inconsistencies. This issue is not well captured by existing IQA metrics or user studies. Overall, general‑purpose models show promise as unified IR solvers from a perceptual perspective, but require improved controllability and fidelity‑aware evaluation. Further comparisons and detailed analyses are available in our project repository: https://github.com/yxyuanxiao/NanoBanana2TestOnIR.
Authors:Qida Cao, Xinyuan Hu, Changyue Shi, Jiajun Ding, Zhou Yu, Jun Yu
Abstract:
This paper describes our method for Track 2 of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge on smoke‑degraded images. In this task, smoke reduces image visibility and weakens the cross‑view consistency required by scene optimization and rendering. We address this problem with a multi‑stage pipeline consisting of image restoration, dehazing, MLLM‑based enhancement, 3DGS‑MCMC optimization, and averaging over repeated runs. The main purpose of the pipeline is to improve visibility before rendering while limiting scene‑content changes across input views. Experimental results on the challenge benchmark show improved quantitative performance and better visual quality than the provided baselines. The code is available at https://github.com/plbbl/GenSmoke‑GS. Our method achieved a ranking of 1 out of 14 participants in Track 2 of the NTIRE 3DRR Challenge, as reported on the official competition website: https://www.codabench.org/competitions/13993/#/results‑tab.
Authors:Wenhao Li, Zimeng Wu, Yu Wu, Zehua Fu, Jiaxin Chen
Abstract:
Unmanned aerial vehicle (UAV) based object detection is a critical but challenging task, when applied in dynamically changing scenarios with limited annotated training data. Layout‑to‑image generation approaches have proved effective in promoting detection accuracy by synthesizing labeled images based on diffusion models. However, they suffer from frequently producing artifacts, especially near layout boundaries of tiny objects, thus substantially limiting their performance. To address these issues, we propose UAVGen, a novel layout‑to‑image generation framework tailored for UAV‑based object detection. Specifically, UAVGen designs a Visual Prototype Conditioned Diffusion Model (VPC‑DM) that constructs representative instances for each class and integrates them into latent embeddings for high‑fidelity object generation. Moreover, a Focal Region Enhanced Data Pipeline (FRE‑DP) is introduced to emphasize object‑concentrated foreground regions in synthesis, combined with a label refinement to correct missing, extra and misaligned generations. Extensive experimental results demonstrate that our method significantly outperforms state‑of‑the‑art approaches, and consistently promotes accuracy when integrated with distinct detectors. The source code is available at https://github.com/Sirius‑Li/UAVGen.
Authors:Zimeng Wu, Yunhong Wang, Donghao Wang, Jiaxin Chen
Abstract:
Vision‑Language Models (VLMs) have advanced rapidly within the unified Transformer architecture, yet their deployment on resource‑constrained devices remains challenging due to high computational complexity. While pruning has emerged as an effective technique for compressing VLMs, existing approaches predominantly focus on a single mode by pruning either parameters or tokens, neglecting fully exploring the inherent redundancy in each mode, which leads to substantial performance degradation at high pruning ratios. To address the above limitations, we propose Collaborative Multi‑Mode Pruning (CoMP), a novel framework tailored for VLMs by performing joint parameter and token pruning. Specifically, we first design a Collaborative Importance Metric (CIM) that investigates the mutual interference between the coupled parameters and tokens. It incorporates distinct significance of tokens into the computation of parameter importance scores, while simultaneously mitigating the affect of pruned parameters on token importance scores. Moreover, we develop a Multi‑Mode Pruning Strategy (MPS) that decomposes the overall pruning process into a sequence of pruning stages, while in each stage we estimate the priory of different pruning modes based on their pruning cost and adaptively shift to the optimal one. Additionally, MPS integrates the historical cost and random exploration, in order to achieve a stable pruning process and avoid local optimum. Extensive experiments across various vision‑language tasks and models demonstrate that our method effectively promotes the performance under high pruning ratios by comparing to the state‑of‑the‑art approaches. The source code is available at https://github.com/Wuzimeng/CoMP.git.
Authors:Yuzhen Niu, Yangqing Wang, Ri Cheng, Fusheng Li, Rongshen Wang, Zhichen Yang
Abstract:
Camouflaged object detection (COD) is challenging due to high target‑background similarity, and recent methods address this by complementarily using RGB‑D texture and geometry cues. However, RGB‑D COD methods still underutilize modality‑specific cues, which limits fusion quality. We believe this is because RGB and depth features are fused directly after backbone extraction without modality‑specific enhancement. To address this limitation, we propose MHENet, an RGB‑D COD framework that performs modality‑specific hierarchical enhancement and adaptive fusion of RGB and depth features. Specifically, we introduce a Texture Hierarchical Enhancement Module (THEM) to amplify subtle texture variations by extracting high‑frequency information and a Geometry Hierarchical Enhancement Module (GHEM) to enhance geometric structures via learnable gradient extraction, while preserving cross‑scale semantic consistency. Finally, an Adaptive Dynamic Fusion Module (ADFM) adaptively fuses the enhanced texture and geometry features with spatially varying weights. Experiments on four benchmarks demonstrate that MHENet surpasses 16 state‑of‑the‑art methods qualitatively and quantitatively. Code is available at https://github.com/afdsgh/MHENet.
Authors:Chunyang Cheng, Tianyang Xu, Xiao-Jun Wu, Tao Zhou, Hui Li, Zhangyong Tang, Josef Kittler
Abstract:
Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide‑and‑conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no‑reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning‑based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks. Our code will be publicly available at https://github.com/AWCXV/EvaNet.
Authors:Chengxing Lin, Jinhong Deng, Yinjie Lei, Wen Li
Abstract:
Recent advances in point cloud In‑Context Learning (ICL) have demonstrated strong multitask capabilities. Existing approaches typically adopt a Masked Point Modeling (MPM)‑based paradigm for point cloud ICL. However, MPM‑based methods directly predict the target point cloud from masked tokens without leveraging geometric priors, requiring the model to infer spatial structure and geometric details solely from token‑level correlations via transformers. Additionally, these methods suffer from a training‑inference objective mismatch, as the model learns to predict the target point cloud using target‑side information that is unavailable at inference time. To address these challenges, we propose DeformPIC, a deformation‑based framework for point cloud ICL. Unlike existing approaches that rely on masked reconstruction, DeformPIC learns to deform the query point cloud under task‑specific guidance from prompts, enabling explicit geometric reasoning and consistent objectives. Extensive experiments demonstrate that DeformPIC consistently outperforms previous state‑of‑the‑art methods, achieving reductions of 1.6, 1.8, and 4.7 points in average Chamfer Distance on reconstruction, denoising, and registration tasks, respectively. Furthermore, we introduce a new out‑of‑domain benchmark to evaluate generalization across unseen data distributions, where DeformPIC achieves state‑of‑the‑art performance.
Authors:Hao Ren, Zetong Bi, Yiming Zeng, Zhaoliang Wan, Lu Qi, Hui Cheng
Abstract:
Visual navigation requires the robot to reach a specified goal such as an image, based on a sequence of first‑person visual observations. While recent learning‑based approaches have made significant progress, they often focus on improving policy heads or decision strategies while relying on simplistic feature encoders and temporal pooling to represent visual input. This leads to the loss of fine‑grained spatial and temporal structure, ultimately limiting accurate action prediction and progress estimation. In this paper, we propose a unified spatio‑temporal representation framework that enhances visual encoding for robotic navigation. Our approach extracts features from both image sequences and goal observations, and fuses them using the designed spatio‑temporal fusion module. This module performs spatial graph reasoning within each frame and models temporal dynamics using a hybrid temporal shift module combined with multi‑resolution difference‑aware convolution. Experimental results demonstrate that our approach consistently improves navigation performance and offers a generalizable visual backbone for goal‑conditioned control. Code is available at \hrefhttps://github.com/hren20/STRNethttps://github.com/hren20/STRNet.
Authors:Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao
Abstract:
Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel‑only reconstruction. To address this, we propose MMPhysVideo, the first framework to scale physical plausibility in video generation through joint multimodal modeling. We recast perceptual cues, specifically semantics, geometry, and spatio‑temporal trajectory, into a unified pseudo‑RGB format, enabling VDMs to directly capture complex physical dynamics. To mitigate cross‑modal interference, we propose a Bidirectionally Controlled Teacher architecture, which utilizes parallel branches to fully decouple RGB and perception processing and adopts two zero‑initialized control links to gradually learn pixel‑wise consistency. For inference efficiency, the teacher's physical prior is distilled into a single‑stream student model via representation alignment. Furthermore, we present MMPhysPipe, a scalable data curation and annotation pipeline tailored for constructing physics‑rich multimodal datasets. MMPhysPipe employs a vision‑language model (VLM) guided by a chain‑of‑visual‑evidence rule to pinpoint physical subjects, enabling expert models to extract multi‑granular perceptual information. Without additional inference costs, MMPhysVideo consistently improves physical plausibility and visual quality over advanced models across various benchmarks and achieves state‑of‑the‑art performance compared to existing methods.
Authors:Jiahe Zhu, Xinyao Wang, Yiyu Zhuang, Yanwen Wang, Jing Tian, Yao Yao, Hao Zhu
Abstract:
Controllable 3D human avatars have found widespread applications in 3D games, the metaverse, and AR/VR scenarios. The conventional approach to creating such a 3D avatar requires a lengthy, intricate pipeline encompassing appearance modeling, motion planning, rigging, and physical simulation. In this paper, we introduce UNICA (UNIfied neural Controllable Avatar), a skeleton‑free generative model that unifies all avatar control components into a single neural framework. Given keyboard inputs akin to video game controls, UNICA generates the next frame of a 3D avatar's geometry through an action‑conditioned diffusion model operating on 2D position maps. A point transformer then maps the resulting geometry to 3D Gaussian Splatting for high‑fidelity free‑view rendering. Our approach naturally captures hair and loose clothing dynamics without manually designed physical simulation, and supports extra‑long autoregressive generation. To the best of our knowledge, UNICA is the first model to unify the workflow of "motion planning, rigging, physical simulation, and rendering". Code is released at https://github.com/zjh21/UNICA.
Authors:Rong-Lin Jian, Ting-Yao Chen, Yu-Fan Lin, Chia-Ming Lee, Fu-En Yang, Yu-Chiang Frank Wang, Chih-Chung Hsu
Abstract:
Color ambient lighting normalization under multi‑colored illumination is challenging due to severe chromatic shifts, highlight saturation, and material‑dependent reflectance. Existing geometric and low‑level priors are insufficient for recovering object‑intrinsic color when illumination‑induced chromatic bias dominates. We observe that DINOv3's self‑supervised features remain highly consistent between colored‑light inputs and ambient‑lit ground truth, motivating their use as illumination‑robust semantic priors. We propose CANDLE (Color Ambient Normalization with DINO Layer Enhancement), which introduces DINO Omni‑layer Guidance (D.O.G.) to adaptively inject multi‑layer DINOv3 features into successive encoder stages, and a color‑frequency refinement design (BFACG + SFFB) to suppress decoder‑side chromatic collapse and detail contamination. Experiments on CL3AN show a +1.22 dB PSNR gain over the strongest prior method. CANDLE achieves 3rd place on the NTIRE 2026 ALN Color Lighting Challenge and 2nd place in fidelity on the White Lighting track with the lowest FID, confirming strong generalization across both chromatic and luminance‑dominant illumination conditions. Code is available at https://github.com/ron941/CANDLE.
Authors:Haoran Zhu, Wen Yang, Guangyou Yang, Chang Xu, Ruixiang Zhang, Fang Xu, Haijian Zhang, Gui-Song Xia
Abstract:
Small object detection (SOD) remains challenging due to extremely limited pixels and ambiguous object boundaries. These characteristics lead to challenging annotation, limited availability of large‑scale high‑quality datasets, and inherently weak semantic representations for small objects. In this work, we first address the data limitation by introducing TinySet‑9M, the first large‑scale, multi‑domain dataset for small object detection. Beyond filling the gap in large‑scale datasets, we establish a benchmark to evaluate the effectiveness of existing label‑efficient detection methods for small objects. Our evaluation reveals that weak visual cues further exacerbate the performance degradation of label‑efficient methods in small object detection, highlighting a critical challenge in label‑efficient SOD. Secondly, to tackle the limitation of insufficient semantic representation, we move beyond training‑time feature enhancement and propose a new paradigm termed Point‑Prompt Small Object Detection (P2SOD). This paradigm introduces sparse point prompts at inference time as an efficient information bridge for category‑level localization, enabling semantic augmentation. Building upon the P2SOD paradigm and the large‑scale TinySet‑9M dataset, we further develop DEAL (DEtect Any smalL object), a scalable and transferable point‑prompted detection framework that learns robust, prompt‑conditioned representations from large‑scale data. With only a single click at inference time, DEAL achieves a 31.4% relative improvement over fully supervised baselines under strict localization metrics (e.g., AP75) on TinySet‑9M, while generalizing effectively to unseen categories and unseen datasets. Our project is available at https://zhuhaoraneis.github.io/TinySet‑9M/.
Authors:Wenli Huang, Yang Wu, Xiaomeng Xin, Zhihong Liu, Jinjun Wang, Ye Deng
Abstract:
Remote sensing image restoration (RSIR) is essential for recovering high‑fidelity imagery from degraded observations, enabling accurate downstream analysis. However, most existing methods focus on single degradation types within homogeneous data, restricting their practicality in real‑world scenarios where multiple degradations often across diverse spectral bands or sensor modalities, creating a significant operational bottleneck. To address this fundamental gap, we propose TGPNet, a unified framework capable of handling denoising, cloud removal, shadow removal, deblurring, and SAR despeckling within a single, unified architecture. The core of our framework is a novel Task‑Guided Prompting (TGP) strategy. TGP leverages learnable, task‑specific embeddings to generate degradation‑aware cues, which then hierarchically modulate features throughout the decoder. This task‑adaptive mechanism allows the network to precisely tailor its restoration process for distinct degradation patterns while maintaining a single set of shared weights. To validate our framework, we construct a unified RSIR benchmark covering RGB, multispectral, SAR, and thermal infrared modalities for five aforementioned restoration tasks. Experimental results demonstrate that TGPNet achieves state‑of‑the‑art performance on both unified multi‑task scenarios and unseen composite degradations, surpassing even specialized models in individual domains such as cloud removal. By successfully unifying heterogeneous degradation removal within a single adaptive framework, this work presents a significant advancement for multi‑task RSIR, offering a practical and scalable solution for operational pipelines. The code and benchmark will be released at https://github.com/huangwenwenlili/TGPNet.
Authors:Mirali Purohit, Bimal Gajera, Irish Mehta, Bhanu Tokas, Jacob Adler, Steven Lu, Scott Dickenshied, Serina Diniega, Brian Bue, Umaa Rebbapragada, Hannah Kerner
Abstract:
We introduce MOMO, the first multi‑sensor foundation model for Mars remote sensing. MOMO uses model merge to integrate representations learned independently from three key Martian sensors (HiRISE, CTX, and THEMIS), spanning resolutions from 0.25 m/pixel to 100 m/pixel. Central to our method is our novel Equal Validation Loss (EVL) strategy, which aligns checkpoints across sensors based on validation loss similarity before fusion via task arithmetic. This ensures models are merged at compatible convergence stages, leading to improved stability and generalization. We train MOMO on a large‑scale, high‑quality corpus of ~ 12 million samples curated from Mars orbital data and evaluate it on 9 downstream tasks from Mars‑Bench. MOMO achieves better overall performance compared to ImageNet pre‑trained, earth observation foundation model, sensor‑specific pre‑training, and fully‑supervised baselines. Particularly on segmentation tasks, MOMO shows consistent and significant performance improvement. Our results demonstrate that model merging through an optimal checkpoint selection strategy provides an effective approach for building foundation models for multi‑resolution data. The model weights, pretraining code, pretraining data, and evaluation code are available at: https://github.com/kerner‑lab/MOMO.
Authors:Zihao Sheng, Xin Ye, Jingru Luo, Sikai Chen, Liu Ren
Abstract:
End‑to‑end autonomous driving models based on Vision‑Language‑Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on expert demonstrations. However, imitation learning inherently limits the model to replicating observed behaviors without exploring diverse driving strategies, leaving it brittle in novel or out‑of‑distribution scenarios. Reinforcement learning (RL) offers a natural remedy by enabling policy exploration beyond the expert distribution. Yet VLA models, typically trained on offline datasets, lack directly observable state transitions, necessitating a learned world model to anticipate action consequences. In this work, we propose a unified understanding‑and‑generation framework that leverages world modeling to simultaneously enable meaningful exploration and provide dense supervision. Specifically, we augment trajectory prediction with future RGB and depth image generation as dense world modeling objectives, requiring the model to learn fine‑grained visual and geometric representations that substantially enrich the planning backbone. Beyond serving as a supervisory signal, the world model further acts as a source of intrinsic reward for policy exploration: its image prediction uncertainty naturally measures a trajectory's novelty relative to the training distribution, where high uncertainty indicates out‑of‑distribution scenarios that, if safe, represent valuable learning opportunities. We incorporate this exploration signal into a safety‑gated reward and optimize the policy via Group Relative Policy Optimization (GRPO). Experiments on the NAVSIM and nuScenes benchmarks demonstrate the effectiveness of our approach, achieving a state‑of‑the‑art PDMS score of 93.7 and an EPDMS of 88.8 on NAVSIM. The code and demo will be publicly available at https://zihaosheng.github.io/ExploreVLA/.
Authors:Junwei You, Pei Li, Zhuoyu Jiang, Weizhe Tang, Zilin Huang, Rui Gan, Jiaxi Liu, Yan Zhao, Sikai Chen, Bin Ran
Abstract:
Multimodal large language models (MLLMs) have shown strong potential for autonomous driving, yet existing benchmarks remain largely ego‑centric and therefore cannot systematically assess model performance in infrastructure‑centric and cooperative driving conditions. In this work, we introduce V2X‑QA, a real‑world dataset and benchmark for evaluating MLLMs across vehicle‑side, infrastructure‑side, and cooperative viewpoints. V2X‑QA is built around a view‑decoupled evaluation protocol that enables controlled comparison under vehicle‑only, infrastructure‑only, and cooperative driving conditions within a unified multiple‑choice question answering (MCQA) framework. The benchmark is organized into a twelve‑task taxonomy spanning perception, prediction, and reasoning and planning, and is constructed through expert‑verified MCQA annotation to enable fine‑grained diagnosis of viewpoint‑dependent capabilities. Benchmark results across ten representative state‑of‑the‑art proprietary and open‑source models show that viewpoint accessibility substantially affects performance, and infrastructure‑side reasoning supports meaningful macroscopic traffic understanding. Results also indicate that cooperative reasoning remains challenging since it requires cross‑view alignment and evidence integration rather than simply additional visual input. To address these challenges, we introduce V2X‑MoE, a benchmark‑aligned baseline with explicit view routing and viewpoint‑specific LoRA experts. The strong performance of V2X‑MoE further suggests that explicit viewpoint specialization is a promising direction for multi‑view reasoning in autonomous driving. Overall, V2X‑QA provides a foundation for studying multi‑perspective reasoning, reliability, and cooperative physical intelligence in connected autonomous driving. The dataset and V2X‑MoE resources are publicly available at: https://github.com/junwei0001/V2X‑QA.
Authors:Yuhui Lin, Siyue Yu, Yuxing Yang, Guangliang Cheng, Jimin Xiao
Abstract:
Recent advances in Multimodal Large Language Models (MLLMs) have expanded reasoning capabilities into 3D domains, enabling fine‑grained spatial understanding. However, the substantial size of 3D MLLMs and the high dimensionality of input features introduce considerable inference overhead, which limits practical deployment on resource constrained platforms. To overcome this limitation, this paper presents Efficient3D, a unified framework for visual token pruning that accelerates 3D MLLMs while maintaining competitive accuracy. The proposed framework introduces a Debiased Visual Token Importance Estimator (DVTIE) module, which considers the influence of shallow initial layers during attention aggregation, thereby producing more reliable importance predictions for visual tokens. In addition, an Adaptive Token Rebalancing (ATR) strategy is developed to dynamically adjust pruning strength based on scene complexity, preserving semantic completeness and maintaining balanced attention across layers. Together, they enable context‑aware token reduction that maintains essential semantics with lower computation. Comprehensive experiments conducted on five representative 3D vision and language benchmarks, including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D, demonstrate that Efficient3D achieves superior performance compared with unpruned baselines, with a +2.57% CIDEr improvement on the Scan2Cap dataset. Therefore, Efficient3D provides a scalable and effective solution for efficient inference in 3D MLLMs. The code is released at: https://github.com/sol924/Efficient3D
Authors:Hao Li, Liwei Zou, Wenping Yin, Gulsen Taskin, Naoto Yokoya, Danfeng Hong, Wufan Zhao
Abstract:
Living in a changing climate, human society now faces more frequent and severe natural disasters than ever before. As a consequence, rapid disaster response during the "Golden 72 Hours" of search and rescue becomes a vital humanitarian necessity and community concern. However, traditional disaster damage surveys routinely fail to generalize across distinct urban morphologies and new disaster events. Effective damage mapping typically requires exhaustive and time‑consuming manual data annotation. To address this issue, we introduce Smart Transfer, a novel Geospatial Artificial Intelligence (GeoAI) framework, leveraging state‑of‑the‑art vision Foundation Models (FMs) for rapid building damage mapping with post‑earthquake Very High Resolution (VHR) imagery. Specifically, we design two novel model transfer strategies: first, Pixel‑wise Clustering (PC), ensuring robust prototype‑level global feature alignment; second, a Distance‑Penalized Triplet (DPT), integrating patch‑level spatial autocorrelation patterns by assigning stronger penalties to semantically inconsistent yet spatially adjacent patches. Extensive experiments and ablations from the recent 2023 Turkiye‑Syria earthquake show promising performance in multiple cross‑region transfer settings, namely Leave One Domain Out (LODO) and Specific Source Domain Combination (SSDC). Moreover, Smart Transfer provides a scalable, automated GeoAI solution to accelerate building damage mapping and support rapid disaster response, offering new opportunities to enhance disaster resilience in climate‑vulnerable regions and communities. The data and code are publicly available at https://github.com/ai4city‑hkust/SmartTransfer.
Authors:Daheng Yin, Isaac Ding, Yili Jin, Jianxin Shi, Jiangchuan Liu
Abstract:
Recent advancements in 3D Gaussian Splatting (3DGS) have demonstrated its potential for efficient and photorealistic 3D reconstructions, which is crucial for diverse applications such as robotics and immersive media. However, current Gaussian‑based methods for dynamic scene reconstruction struggle with large inter‑frame displacements, leading to artifacts and temporal inconsistencies under fast object motions. To address this, we introduce TrackerSplat, a novel method that integrates advanced point tracking methods to enhance the robustness and scalability of 3DGS for dynamic scene reconstruction. TrackerSplat utilizes off‑the‑shelf point tracking models to extract pixel trajectories and triangulate per‑view pixel trajectories onto 3D Gaussians to guide the relocation, rotation, and scaling of Gaussians before training. This strategy effectively handles large displacements between frames, dramatically reducing the fading and recoloring artifacts prevalent in prior methods. By accurately positioning Gaussians prior to gradient‑based optimization, TrackerSplat overcomes the quality degradation associated with large frame gaps when processing multiple adjacent frames in parallel across multiple devices, thereby boosting reconstruction throughput while preserving rendering quality. Experiments on real‑world datasets confirm the robustness of TrackerSplat in challenging scenarios with significant displacements, achieving superior throughput under parallel settings and maintaining visual quality compared to baselines. The code is available at https://github.com/yindaheng98/TrackerSplat.
Authors:Haiyu Wang, Yutong Wang, Jack Jiang, Sai Qian Zhang
Abstract:
Singular Value Decomposition (SVD) has become an important technique for reducing the computational burden of Vision Language Models (VLMs), which play a central role in tasks such as image captioning and visual question answering. Although multiple prior works have proposed efficient SVD variants to enable low‑rank operations, we find that in practice it remains difficult to achieve substantial latency reduction during model execution. To address this limitation, we introduce a new computational pattern and apply SVD at a finer granularity, enabling real and measurable improvements in execution latency. Furthermore, recognizing that weight elements differ in their relative importance, we adaptively allocate relative importance to each element during SVD process to better preserve accuracy, then extend this framework with quantization applied to both weights and activations, resulting in a highly efficient VLM. Collectively, we introduce~Weighted SVD (WSVD), which outperforms other approaches by achieving over 1.8× decoding speedup while preserving accuracy. We open source our code at: \hrefhttps://github.com/SAI‑Lab‑NYU/WSVD\texttthttps://github.com/SAI‑Lab‑NYU/WSVD
Authors:Sebo Diaz, Polina Golland, Elfar Adalsteinsson, Neel Dey
Abstract:
We present MaskGen, a theoretically grounded and deliberately simple approach for domain generalization in 3D biomedical image segmentation. Modern segmentation models degrade sharply under shifts in modality, disease severity, clinical sites, and more, limiting their reliable adoption. Existing generalization methods address this using extreme augmentations, hand‑engineered domain statistics mixing, or architectural redesigns that add significant implementation overhead while yielding inconsistent performance across biomedical settings. MaskGen instead presents a principled learning strategy with marginal overhead that utilizes both source‑domain image intensities and domain‑stable foundation model representations to train robust segmentation models. As a result, MaskGen achieves strong gains in both fully supervised and few‑shot segmentation across broad clinical shifts in biomedical studies. Unlike prior approaches, MaskGen is architecture‑ and loss‑agnostic, compatible with standard augmentation pipelines, easy to implement, and tackles arbitrary anatomical regions. Its implementation is freely available at https://github.com/sebodiaz/MaskGen.
Authors:Ye Mao, Weixun Luo, Ranran Huang, Junpeng Jing, Krystian Mikolajczyk
Abstract:
Pretraining 3D encoders by aligning with Contrastive Language Image Pretraining (CLIP) has emerged as a promising direction to learn generalizable representations for 3D scene understanding. In this paper, we propose UniScene3D, a transformer‑based encoder that learns unified scene representations from multi‑view colored pointmaps, jointly modeling image appearance and geometry. For robust colored pointmap representation learning, we introduce novel cross‑view geometric alignment and grounded view alignment to enforce cross‑view geometry and semantic consistency. Extensive low‑shot and task‑specific fine‑tuning evaluations on viewpoint grounding, scene retrieval, scene type classification, and 3D VQA demonstrate our state‑of‑the‑art performance. These results highlight the effectiveness of our approach for unified 3D scene understanding. https://yebulabula.github.io/UniScene3D/
Authors:Luca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano Mattoccia, Guillermo Gallego
Abstract:
We propose EventHub, a novel framework for training deep‑event stereo networks without ground truth annotations from costly active sensors, relying instead on standard color images. From these images, we derive either proxy annotations and proxy events through state‑of‑the‑art novel view synthesis techniques, or simply proxy annotations when images are already paired with event data. Using the training set generated by our data factory, we repurpose state‑of‑the‑art stereo models from RGB literature to process event data, obtaining new event stereo models with unprecedented generalization capabilities. Experiments on widely used event stereo datasets support the effectiveness of EventHub and show how the same data distillation mechanism can improve the accuracy of RGB stereo foundation models in challenging conditions such as nighttime scenes.
Authors:Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan, Ruihan Yu, Yidan Zhang, Bo Zheng, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang
Abstract:
Scaling generative inverse and forward rendering to real‑world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge this persistent domain gap, we introduce a large‑scale, dynamic dataset curated from visually complex AAA games. Using a novel dual‑screen stitched capture method, we extracted 4M continuous frames (720p/30 FPS) of synchronized RGB and five G‑buffer channels across diverse scenes, visual effects, and environments, including adverse weather and motion‑blur variants. This dataset uniquely advances bidirectional rendering: enabling robust in‑the‑wild geometry and material decomposition, and facilitating high‑fidelity G‑buffer‑guided video generation. Furthermore, to evaluate the real‑world performance of inverse rendering without ground truth, we propose a novel VLM‑based assessment protocol measuring semantic, spatial, and temporal consistency. Experiments demonstrate that inverse renderers fine‑tuned on our data achieve superior cross‑dataset generalization and controllable generation, while our VLM evaluation strongly correlates with human judgment. Combined with our toolkit, our forward renderer enables users to edit styles of AAA games from G‑buffers using text prompts.
Authors:Ruozhen He, Nisarg A. Shah, Qihua Dong, Zilin Xiao, Jaywon Koo, Vicente Ordonez
Abstract:
Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent named category. We explore a complementary and more challenging setting of scenario‑based visual grounding, where the target must be inferred from roles, intentions, and relational context rather than explicit naming. We introduce Referring Scenario Comprehension (RSC), a benchmark designed for this setting. The queries in this benchmark are paragraph‑length texts that describe object roles, user goals, and contextual cues, including deliberate references to distractor objects that often require deep understanding to resolve. Each instance is annotated with interpretable difficulty tags for uniqueness, clutter, size, overlap, and position which expose distinct failure modes and support fine‑grained analysis. RSC contains approximately 31k training examples, 4k in‑domain test examples, and a 3k out‑of‑distribution split with unseen object categories. We further propose ScenGround, a curriculum reasoning method serving as a reference point for this setting, combining supervised warm‑starting with difficulty‑aware reinforcement learning. Experiments show that scenario‑based queries expose systematic failures in current models that standard benchmarks do not reveal, and that curriculum training improves performance on challenging slices and transfers to standard benchmarks.
Authors:Junxuan Li, Rawal Khirodkar, Chengan He, Zhongshi Jiang, Giljoo Nam, Lingchen Yang, Jihyun Lee, Egor Zakharov, Zhaoen Su, Rinat Abdrashitov, Yuan Dong, Julieta Martinez, Kai Li, Qingyang Tan, Takaaki Shiratori, Matthew Hu, Peihong Guo, Xuhua Huang, Ariyan Zarei, Marco Pesavento, Yichen Xu, He Wen, Teng Deng, Wyatt Borsos, Anjali Thakrar, Jean-Charles Bazin, Carsten Stoll, Ginés Hidalgo, James Booth, Lucy Wang, Xiaowen Ma, Yu Rong, Sairanjith Thalanki, Chen Cao, Christian Häne, Abhishek Kar, Sofien Bouaziz, Jason Saragih, Yaser Sheikh, Shunsuke Saito
Abstract:
High‑quality 3D avatar modeling faces a critical trade‑off between fidelity and generalization. On the one hand, multi‑view studio data enables high‑fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real‑world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large‑scale avatar models trained on millions of in‑the‑wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low‑quality due to inherent 3D ambiguities. To address this, we present Large‑Scale Codec Avatars (LCA), a high‑fidelity, full‑body 3D avatar model that generalizes to world‑scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, we present, for the first time, a pre/post‑training paradigm for 3D avatar modeling at scale: we pretrain on 1M in‑the‑wild videos to learn broad priors over appearance and geometry, then post‑train on high‑quality curated data to enhance expressivity and fidelity. LCA generalizes across hair styles, clothing, and demographics while providing precise, fine‑grained facial expressions and finger‑level articulation control, with strong identity preservation. Notably, we observe emergent generalization to relightability and loose garment support to unconstrained inputs, and zero‑shot robustness to stylized imagery, despite the absence of direct supervision.
Authors:Naomi Kombol, Ivan Martinović, Siniša Šegvić, Giorgos Tolias
Abstract:
Foundational Vision Transformers (ViTs) have limited effectiveness in tasks requiring fine‑grained spatial understanding, due to their fixed pre‑training resolution and inherently coarse patch‑level representations. These challenges are especially pronounced in dense prediction scenarios, such as open‑vocabulary segmentation with ViT‑based vision‑language models, where high‑resolution inputs are essential for accurate pixel‑level reasoning. Existing approaches typically process large‑resolution images using a sliding‑window strategy at the pre‑training resolution. While this improves accuracy through finer strides, it comes at a significant computational cost. We introduce SPAR: Single‑Pass Any‑Resolution ViT, a resolution‑agnostic dense feature extractor designed for efficient high‑resolution inference. We distill the spatial reasoning capabilities of a finely‑strided, sliding‑window teacher into a single‑pass student using a feature regression loss, without requiring architectural changes or pixel‑level supervision. Applied to open‑vocabulary segmentation, SPAR improves single‑pass baselines by up to 10.5 mIoU and even surpasses the teacher, demonstrating effectiveness in efficient, high‑resolution reasoning. Code: https://github.com/naomikombol/SPAR
Authors:Qiyao Zhang, Shuhua Zheng, Jianli Sun, Chengxiang Li, Xianke Wu, Zihan Song, Zhiyong Cui, Yisheng Lv, Yonglin Tian
Abstract:
Embodied visual tracking is crucial for Unmanned Aerial Vehicles (UAVs) executing complex real‑world tasks. In dynamic urban scenarios with complex semantic requirements, Vision‑Language‑Action (VLA) models show great promise due to their cross‑modal fusion and continuous action generation capabilities. To benchmark multimodal tracking in such environments, we construct a dedicated evaluation benchmark and a large‑scale dataset encompassing over 890K frames, 176 tasks, and 85 diverse objects. Furthermore, to address temporal feature redundancy and the lack of spatial geometric priors in existing VLA models, we propose an improved VLA tracking model, UAV‑Track VLA. Built upon the π_0.5 architecture, our model introduces a temporal compression net to efficiently capture inter‑frame dynamics. Additionally, a parallel dual‑branch decoder comprising a spatial‑aware auxiliary grounding head and a flow matching action expert is designed to decouple cross‑modal features and generate fine‑grained continuous actions. Systematic experiments in the CARLA simulator validate the superior end‑to‑end performance of our method. Notably, in challenging long‑distance pedestrian tracking tasks, UAV‑Track VLA achieves a 61.76% success rate and 269.65 average tracking frames, significantly outperforming existing baselines. Furthermore, it demonstrates robust zero‑shot generalization in unseen environments and reduces single‑step inference latency by 33.4% (to 0.0571s) compared to the original π_0.5, enabling highly efficient, real‑time UAV control. Data samples and demonstration videos are available at: https://github.com/Hub‑Tian/UAV‑Track_VLA.
Authors:Yongkang Li, Lijun Zhou, Sixu Yan, Bencheng Liao, Tianyi Yan, Kaixin Xiong, Long Chen, Hongwei Xie, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Haiyang Sun, Xinggang Wang
Abstract:
Vision‑Language‑Action (VLA) models have recently emerged in autonomous driving, with the promise of leveraging rich world knowledge to improve the cognitive capabilities of driving systems. However, adapting such models for driving tasks currently faces a critical dilemma between spatial perception and semantic reasoning. Consequently, existing VLA systems are forced into suboptimal compromises: directly adopting 2D Vision‑Language Models yields limited spatial perception, whereas enhancing them with 3D spatial representations often impairs the native reasoning capacity of VLMs. We argue that this dilemma largely stems from the coupled optimization of spatial perception and semantic reasoning within shared model parameters. To overcome this, we propose UniDriveVLA, a Unified Driving Vision‑Language‑Action model based on Mixture‑of‑Transformers that addresses the perception‑reasoning conflict via expert decoupling. Specifically, it comprises three experts for driving understanding, scene perception, and action planning, which are coordinated through masked joint attention. In addition, we combine a sparse perception paradigm with a three‑stage progressive training strategy to improve spatial perception while maintaining semantic reasoning capability. Extensive experiments show that UniDriveVLA achieves state‑of‑the‑art performance in open‑loop evaluation on nuScenes and closed‑loop evaluation on Bench2Drive. Moreover, it demonstrates strong performance across a broad range of perception, prediction, and understanding tasks, including 3D detection, online mapping, motion forecasting, and driving‑oriented VQA, highlighting its broad applicability as a unified model for autonomous driving. Code and model have been released at https://github.com/xiaomi‑research/unidrivevla
Authors:Rong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang, Kai Dai, Zhao Yang
Abstract:
Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid‑LLMs) to broader applications. However, existing Vid‑LLMs rely on uniform frame sampling to extract video information, resulting in a sparse distribution of key frames and the loss of crucial temporal cues. To address this limitation, we propose Grounded Visual Token Sampling (GroundVTS), a Vid‑LLM architecture that focuses on the most informative temporal segments. GroundVTS employs a fine‑grained, query‑guided mechanism to filter visual tokens before feeding them into the LLM, thereby preserving essential spatio‑temporal information and maintaining temporal coherence. Futhermore, we introduce a progressive optimization strategy that enables the LLM to effectively adapt to the non‑uniform distribution of visual features, enhancing its ability to model temporal dependencies and achieve precise video localization. We comprehensively evaluate GroundVTS on three standard VTG benchmarks, where it outperforms existing methods, achieving a 7.7‑point improvement in mIoU for moment retrieval and 12.0‑point improvement in mAP for highlight detection. Code is available at https://github.com/Florence365/GroundVTS.
Authors:Yan Kong, Yuan Yin, Hongan Chen, Yuqi Fang, Caifeng Shan
Abstract:
Automated analysis of Pap smear images is critical for cervical cancer screening but remains challenging due to dense cell distribution and complex morphology. In this paper, we present our winning solution for the RIVA Cervical Cytology Challenge, achieving 1st place in Track B and 2nd place in Track A. Our approach leverages a powerful baseline, integrating the Co‑DINO framework with a Swin‑Large backbone for robust multi‑scale feature extraction. To address the dataset's unique fixed‑size bounding box annotations, we formulate the detection task as a center‑point prediction problem. Tailoring our approach to this formulation, we introduce a center‑preserving data augmentation strategy and an analytical geometric box optimization to effectively absorb localization jitter. Finally, we apply track‑specific loss tuning to adapt the loss weights for each task. Experiments demonstrate that our targeted optimizations improve detection performance, providing an effective pipeline for cytology image analysis. Our code is available at https://github.com/YanKong0408/Center‑DETR.
Authors:Soo Won Seo, KyungChae Lee, Hyungchan Cho, Taein Son, Nam Ik Cho, Jun Won Choi
Abstract:
Human‑Object Interaction (HOI) detection aims to localize human‑object pairs and classify their interactions from a single image, a task that demands strong visual understanding and nuanced contextual reasoning. Recent approaches have leveraged Vision‑Language Models (VLMs) to introduce semantic priors, significantly improving HOI detection performance. However, existing methods often fail to fully capitalize on the diverse contextual cues distributed across the entire scene. To overcome these limitations, we propose the Instance‑centric Context Mining Network (InCoM‑Net)‑a novel framework that effectively integrates rich semantic knowledge extracted from VLMs with instance‑specific features produced by an object detector. This design enables deeper interaction reasoning by modeling relationships not only within each detected instance but also across instances and their surrounding scene context. InCoM‑Net comprises two core components: Instancecentric Context Refinement (ICR), which separately extracts intra‑instance, inter‑instance, and global contextual cues from VLM‑derived features, and Progressive Context Aggregation (ProCA), which iteratively fuses these multicontext features with instance‑level detector features to support high‑level HOI reasoning. Extensive experiments on the HICO‑DET and V‑COCO benchmarks show that InCoM‑Net achieves state‑of‑the‑art performance, surpassing previous HOI detection methods. Code is available at https://github.com/nowuss/InCoM‑Net.
Authors:Qing Zhou, Shiyu Zhang, Yuyu Jia, Junyu Gao, Weiping Ni, Junzheng Wu, Qi Wang
Abstract:
Chain‑of‑thought (CoT) reasoning has significantly improved the performance of large multimodal models in language‑guided segmentation, yet its prohibitive computational cost, stemming from generating verbose rationales, limits real‑world applicability. We introduce WISE (Wisdom from Internal Self‑Exploration), a novel paradigm for efficient reasoning guided by the principle of thinking twice ‑‑ once for learning, once for speed. WISE trains a model to generate a structured sequence: a concise rationale, the final answer, and then a detailed explanation. By placing the concise rationale first, our method leverages autoregressive conditioning to enforce that the concise rationale acts as a sufficient summary for generating the detailed explanation. This structure is reinforced by a self‑distillation objective that jointly rewards semantic fidelity and conciseness, compelling the model to internalize its detailed reasoning into a compact form. At inference, the detailed explanation is omitted. To address the resulting conditional distribution shift, our inference strategy, WISE‑S, employs a simple prompting technique that injects a brevity‑focused instruction into the user's query. This final adjustment facilitates the robust activation of the learned concise policy, unlocking the full benefits of our framework. Extensive experiments show that WISE‑S achieves state‑of‑the‑art zero‑shot performance on the ReasonSeg benchmark with 58.3 cIoU, while reducing the average reasoning length by nearly 5× ‑‑ from 112 to just 23 tokens. Code is available at \hrefhttps://github.com/mrazhou/WISEWISE.
Authors:Sebastian-Ion Nae, Radu Moldoveanu, Alexandra Stefania Ghita, Adina Magda Florea
Abstract:
Understanding human behaviour in crowded indoor environments is central to surveillance, smart buildings, and human‑robot interaction, yet existing datasets rarely capture real‑world indoor complexity at scale. We introduce IndoorCrowd, a multi‑scene dataset for indoor human detection, instance segmentation, and multi‑object tracking, collected across four campus locations (ACS‑EC, ACS‑EG, IE‑Central, R‑Central). It comprises 31 videos (9,913 frames at 5fps) with human‑verified, per‑instance segmentation masks. A 620‑frame control subset benchmarks three foundation‑model auto‑annotators: SAM3, GroundingSAM, and EfficientGroundingSAM, against human labels using Cohen's κ, AP, precision, recall, and mask IoU. A further 2,552‑frame subset supports multi‑object tracking with continuous identity tracks in MOTChallenge format. We establish detection, segmentation, and tracking baselines using YOLOv8n, YOLOv26n, and RT‑DETR‑L paired with ByteTrack, BoT‑SORT, and OC‑SORT. Per‑scene analysis reveals substantial difficulty variation driven by crowd density, scale, and occlusion: ACS‑EC, with 79.3% dense frames and a mean instance scale of 60.8px, is the most challenging scene. The project page is available at https://sheepseb.github.io/IndoorCrowd/.
Authors:Sirshapan Mitra, Yogesh S. Rawat
Abstract:
Generating ground‑level views and coherent 3D site models from aerial‑only imagery is challenging due to extreme viewpoint changes, missing intermediate observations, and large scale variations. Existing methods either refine renderings post‑hoc, often producing geometrically inconsistent results, or rely on multi‑altitude ground‑truth, which is rarely available. Gaussian Splatting and diffusion‑based refinements improve fidelity under small variations but fail under wide aerial‑toground gaps. To address these limitations, we introduce ProDiG (Progressive Diffusion‑Guided Gaussian Splatting for Aerial to Ground Reconstruction), a diffusionguided framework that progressively transforms aerial 3D representations toward ground‑level fidelity. ProDiG synthesizes intermediate‑altitude views and refines the Gaussian representation at each stage using a geometry‑aware causal attention module that injects epipolar structure into reference‑view diffusion. A distance‑adaptive Gaussian module dynamically adjusts Gaussian scale and opacity based on camera distance, ensuring stable reconstruction across large viewpoint gaps. Together, these components enable progressive, geometrically grounded refinement without requiring additional ground‑truth viewpoints. Extensive experiments on synthetic and real‑world datasets demonstrate that ProDiG produces visually realistic ground‑level renderings and coherent 3D geometry, significantly outperforming existing approaches in terms of visual quality, geometric consistency, and robustness to extreme viewpoint changes. Project Page: https://sirsh07.github.io/research/prodig
Authors:Yuqing Huang, Guotian Zeng, Zhenqiao Yuan, Zhenyu He, Xin Li, Yaowei Wang, Ming-Hsuan Yang
Abstract:
Existing visual trackers mainly operate in a non‑interactive, fire‑and‑forget manner, making them impractical for real‑world scenarios that require human‑in‑the‑loop adaptation. To overcome this limitation, we introduce Interactive Tracking, a new paradigm that allows users to guide the tracker at any time using natural language commands. To support research in this direction, we make three main contributions. First, we present InteractTrack, the first large‑scale benchmark for interactive tracking, containing 150 videos with dense bounding box annotations and timestamped language instructions. Second, we propose a comprehensive evaluation protocol and evaluate 25 representative trackers, showing that state‑of‑the‑art methods fail in interactive scenarios; strong performance on conventional benchmarks does not transfer. Third, we introduce Interactive Memory‑Augmented Tracking (IMAT), a new baseline that employs a dynamic memory mechanism to learn from user feedback and update tracking behavior accordingly. Our benchmark, protocol, and baseline establish a foundation for developing more intelligent, adaptive, and collaborative tracking systems, bridging the gap between automated perception and human guidance. The full benchmark, tracking results, and analysis are available at https://github.com/NorahGreen/InteractTrack.git.
Authors:Aleksandar Cvejic, Rameen Abdal, Abdelrahman Eldesokey, Bernard Ghanem, Peter Wonka
Abstract:
When evaluating identity‑focused tasks such as personalized generation and image editing, existing vision encoders entangle object identity with background context, leading to unreliable representations and metrics. We introduce the first principled framework to address this vulnerability using Near‑identity (NearID) distractors, where semantically similar but distinct instances are placed on the exact same background as a reference image, eliminating contextual shortcuts and isolating identity as the sole discriminative signal. Based on this principle, we present the NearID dataset (19K identities, 316K matched‑context distractors) together with a strict margin‑based evaluation protocol. Under this setting, pre‑trained encoders perform poorly, achieving Sample Success Rates (SSR), a strict margin‑based identity discrimination metric, as low as 30.7% and often ranking distractors above true cross‑view matches. We address this by learning identity‑aware representations on a frozen backbone using a two‑tier contrastive objective enforcing the hierarchy: same identity > NearID distractor > random negative. This improves SSR to 99.2%, enhances part‑level discrimination by 28.0%, and yields stronger alignment with human judgments on DreamBench++, a human‑aligned benchmark for personalization. Project page: https://gorluxor.github.io/NearID/
Authors:Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela Yao
Abstract:
We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question‑answering requiring ego‑grounding ‑ the ability to understand the camera‑wearer in egocentric videos. To this end, we introduce MyEgo, the first egocentric VideoQA dataset designed to evaluate MLLMs' ability to understand, remember, and reason about the camera wearer. MyEgo comprises 541 long videos and 5K personalized questions asking about "my things", "my activities", and "my past". Benchmarking reveals that competitive MLLMs across variants, including open‑source vs. proprietary, thinking vs. non‑thinking, small vs. large scales all struggle on MyEgo. Top closed‑ and open‑source models (e.g., GPT‑5 and Qwen3‑VL) achieve only~46% and 36% accuracy, trailing human performance by near 40% and 50% respectively. Surprisingly, neither explicit reasoning nor model scaling yield consistent improvements. Models improve when relevant evidence is explicitly provided, but gains drop over time, indicating limitations in tracking and remembering "me" and "my past". These findings collectively highlight the crucial role of ego‑grounding and long‑range memory in enabling personalized QA in egocentric videos. We hope MyEgo and our analyses catalyze further progress in these areas for egocentric personalized assistance. Data and code are available at https://github.com/Ryougetsu3606/MyEgo
Authors:Xilai Li, Weijun Jiang, Xiaosong Li, Yang Liu, Hongbin Wang, Tao Ye, Huafeng Li, Haishu Tan
Abstract:
Infrared and visible video fusion combines the object saliency from infrared images with the texture details from visible images to produce semantically rich fusion results. However, most existing methods are designed for static image fusion and cannot effectively handle frame‑to‑frame motion in videos. Current video fusion methods improve temporal consistency by introducing interactions across frames, but they often require high computational cost. To mitigate these challenges, we propose MAVFusion, an end‑to‑end video fusion framework featuring a motion‑aware sparse interaction mechanism that enhances efficiency while maintaining superior fusion quality. Specifically, we leverage optical flow to identify dynamic regions in multi‑modal sequences, adaptively allocating computationally intensive cross‑modal attention to these sparse areas to capture salient transitions and facilitate inter‑modal information exchange. For static background regions, a lightweight weak interaction module is employed to maintain structural and appearance integrity. By decoupling the processing of dynamic and static regions, MAVFusion simultaneously preserves temporal consistency and fine‑grained details while significantly accelerating inference. Extensive experiments demonstrate that MAVFusion achieves state‑of‑the‑art performance on multiple infrared and visible video benchmarks, achieving a speed of 14.16\,FPS at 640 × 480 resolution. The source code will be available at https://github.com/ixilai/MAVFusion.
Authors:Yimin Fu, Songbo Wang, Feiyan Wu, Jialin Lyu, Zhunga Liu, Michael K. Ng
Abstract:
The accurate target‑background separation in infrared small target detection (IRSTD) highly depends on the discriminability of extracted representations. However, most existing methods are confined to domain‑consistent settings, while overlooking whether such discriminability can generalize to unseen domains. In practice, distribution shifts between training and testing data are inevitable due to variations in observational conditions and environmental factors. Meanwhile, the intrinsic indistinctiveness of infrared small targets aggravates overfitting to domain‑specific patterns. Consequently, the detection performance of models trained on source domains can be severely degraded when deployed in unseen domains. To address this challenge, we propose a spatial‑spectral collaborative perception network (S^2CPNet) for cross‑domain IRSTD. Moving beyond conventional spatial learning pipelines, we rethink IRSTD representations from a frequency perspective and reveal inconsistencies in spectral phase as the primary manifestation of domain discrepancies. Based on this insight, we develop a phase rectification module (PRM) to derive generalizable target awareness. Then, we employ an orthogonal attention mechanism (OAM) in skip connections to preserve positional information while refining informative representations. Moreover, the bias toward domain‑specific patterns is further mitigated through selective style recomposition (SSR). Extensive experiments have been conducted on three IRSTD datasets, and the proposed method consistently achieves state‑of‑the‑art performance under diverse cross‑domain settings.
Authors:Xilai Li, Chusheng Fang, Xiaosong Li
Abstract:
Infrared and visible video fusion plays a critical role in intelligent surveillance and low‑light monitoring. However, maintaining temporal stability while preserving spatial detail remains a fundamental challenge. Existing methods either focus on frame‑wise enhancement with limited temporal modeling or rely on heavy spatio‑temporal aggregation that often sacrifices high‑frequency details. In this paper, we propose FTPFusion, a frequency‑aware infrared and visible video fusion method based on temporal perturbation and sparse cross‑modal interaction. Specifically, FTPFusion decomposes the feature representations into high‑frequency and low‑frequency components for collaborative modeling. The high‑frequency branch performs sparse cross‑modal spatio‑temporal interaction to capture motion‑related context and complementary details. The low‑frequency branch introduces a temporal perturbation strategy to enhance robustness against complex video variations, such as flickering, jitter, and local misalignment. Furthermore, we design an offset‑aware temporal consistency constraint to explicitly stabilize cross‑frame representations under temporal disturbances. Extensive experiments on multiple public benchmarks demonstrate that FTPFusion consistently outperforms state‑of‑the‑art methods across multiple metrics in both spatial fidelity and temporal consistency. The source code will be available at https://github.com/ixilai/FTPFusion.
Authors:Panagiotis Sapoutzoglou, George Terzakis, Maria Pateraki
Abstract:
We propose SHARC, a novel framework that synthesizes arbitrary, genus‑agnostic shapes by means of a collection of Spherical Harmonic (SH) representations of distance fields. These distance fields are anchored at optimally placed reference points in the interior volume of the surface in a way that maximizes learning of the finer details of the surface. To achieve this, we employ a cost function that jointly maximizes sparsity and centrality in terms of positioning, as well as visibility of the surface from their location. For each selected reference point, we sample the visible distance field to the surface geometry via ray‑casting and compute the SH coefficients using the Fast Spherical Harmonic Transform (FSHT). To enhance geometric fidelity, we apply a configurable low‑pass filter to the coefficients and refine the output using a local consistency constraint based on proximity. Evaluation of SHARC against state‑of‑the‑art methods demonstrates that the proposed method outperforms existing approaches in both reconstruction accuracy and time efficiency without sacrificing model parsimony. The source code is available at https://github.com/POSE‑Lab/SHARC.
Authors:Pawel Tomasz Pieta, Rasmus Juul Pedersen, Sina Borgi, Jakob Sauer Jørgensen, Jens Wenzel Andreasen, Vedrana Andersen Dahl
Abstract:
Gaussian Splatting (GS) has emerged as a dominating technique for image rendering and has quickly been adapted for the X‑ray Computed Tomography (CT) reconstruction task. However, despite being on par or better than many of its predecessors, the benefits of GS are typically not substantial enough to motivate a transition from well‑established reconstruction algorithms. This paper addresses the most significant remaining limitations of the GS‑based approach by introducing FaCT‑GS, a framework for fast and flexible CT reconstruction. Enabled by an in‑depth optimization of the voxelization and rasterization pipelines, our new method is significantly faster than its predecessors and scales well with projection and output volume size. Furthermore, the improved voxelization enables rapid fitting of Gaussians to pre‑existing volumes, which can serve as a prior for warm‑starting the reconstruction, or simply as an alternative, compressed representation. FaCT‑GS is over 4X faster than the State of the Art GS CT reconstruction on standard 512x512 projections, and over 13X faster on 2k projections. Implementation available at: https://github.com/PaPieta/fact‑gs.
Authors:Xiang Yang, Feifei Li, Mi Zhang, Geng Hong, Xiaoyu You, Min Yang
Abstract:
Recent Text‑to‑Image (T2I) models based on rectified‑flow transformers (e.g., SD3, FLUX) achieve high generative fidelity but remain vulnerable to unsafe semantics, especially when triggered by multi‑token interactions. Existing mitigation methods largely rely on fine‑tuning or attention modulation for concept unlearning; however, their expensive computational overhead and design tailored to U‑Net‑based denoisers hinder direct adaptation to transformer‑based diffusion models (e.g., MMDiT). In this paper, we conduct an in‑depth analysis of the attention mechanism in MMDiT and find that unsafe semantics concentrate within interpretable, low‑dimensional subspaces at head level, where a finite set of safety‑critical heads is responsible for unsafe feature extraction. We further observe that perturbing the Rotary Positional Embedding (RoPE) applied to the query and key vectors can effectively modify some specific concepts in the generated images. Motivated by these insights, we propose SafeRoPE, a lightweight and fine‑grained safe generation framework for MMDiT. Specifically, SafeRoPE first constructs head‑wise unsafe subspaces by decomposing unsafe embeddings within safety‑critical heads, and computes a Latent Risk Score (LRS) for each input vector via projection onto these subspaces. We then introduce head‑wise RoPE perturbations that can suppress unsafe semantics without degrading benign content or image quality. SafeRoPE combines both head‑wise LRS and RoPE perturbations to perform risk‑specific head‑wise rotation on query and key vector embeddings, enabling precise suppression of unsafe outputs while maintaining generation fidelity. Extensive experiments demonstrate that SafeRoPE achieves SOTA performance in balancing effective harmful content mitigation and utility preservation for safe generation of MMDiT. Codes are available at https://github.com/deng12yx/SafeRoPE.
Authors:Mengtian Li, Fan Yang, Ruixue Xiong, Yiyan Fan, Zhifeng Xie, Zeyu Wang
Abstract:
Jiangnan gardens, a prominent style of Chinese classical gardens, hold great potential as digital assets for film and game production and digital tourism. However, manual modeling of Jiangnan gardens heavily relies on expert experience for layout design and asset creation, making the process time‑consuming. To address this gap, we propose GardenDesigner, a novel framework that encodes aesthetic principles for Jiangnan garden construction and integrates a chain of agents based on procedural modeling. The water‑centric terrain and explorative pathway rules are applied by terrain distribution and road generation agents. Selection and spatial layout of garden assets follow the aesthetic and cultural constraints. Consequently, we propose asset selection and layout optimization agents to select and arrange objects for each area in the garden. Additionally, we introduce GardenVerse for Jiangnan garden construction, including expert‑annotated garden knowledge to enhance the asset arrangement process. To enable interaction and editing, we develop an interactive interface and tools in Unity, in which non‑expert users can construct Jiangnan gardens via text input within one minute. Experiments and human evaluations demonstrate that GardenDesigner can generate diverse and aesthetically pleasing Jiangnan gardens. Project page is available at https://monad‑cube.github.io/GardenDesigner.
Authors:Edoardo A. Dominici, Thomas Deixelberger, Konstantinos Vardis, Markus Steinberger
Abstract:
Video models have recently been applied with success to problems in content generation, novel view synthesis, and, more broadly, world simulation. Many applications in generation and transfer rely on conditioning these models, typically through perceptual, geometric, or simple semantic signals, fundamentally using them as generative renderers. At the same time, high‑dimensional features obtained from large‑scale self‑supervised learning on images or point clouds are increasingly used as a general‑purpose interface for vision models. The connection between the two has been explored for subject specific editing, aligning and training video diffusion models, but not in the role of a more general conditioning signal for pretrained video diffusion models. Features obtained through self‑supervised learning like DINO, contain a lot of entangled information about style, lighting and semantics of the scene. This makes them great at reconstruction tasks but limits their generative capabilities. In this paper, we show how we can use the features for tasks such as video domain transfer and video‑from‑3D generation. We introduce a lightweight architecture and training strategy that decouples appearance from other features that we wish to preserve, enabling robust control for appearance changes such as stylization and relighting. Furthermore, we show that low spatial resolution can be compensated by higher feature dimensionality, improving controllability in generative rendering from explicit spatial representations.
Authors:Haibo Li, Qingyue Deng, Jijiang Li, Haibin Ling, Bingyao Huang
Abstract:
Projector compensation seeks to correct geometric and photometric distortions that occur when images are projected onto nonplanar or textured surfaces. However, most existing methods are highly setup‑dependent, requiring fine‑tuning or retraining whenever the surface, lighting, or projector‑camera pose changes. Progress has been limited by two key challenges: (1) the absence of large, diverse training datasets and (2) existing geometric correction models are typically constrained by specific spatial setups; without further retraining or fine‑tuning, they often fail to generalize directly to novel geometric configurations. We introduce SIComp, the first Setup‑Independent framework for full projector Compensation, capable of generalizing to unseen setups without fine‑tuning or retraining. To enable this, we construct a large‑scale real‑world dataset spanning 277 distinct projector‑camera setups. SIComp adopts a co‑adaptive design that decouples geometry and photometry: A carefully tailored optical flow module performs online geometric correction, while a novel photometric network handles photometric compensation. To further enhance robustness under varying illumination, we integrate intensity‑varying surface priors into the network design. Extensive experiments demonstrate that SIComp consistently produces high‑quality compensation across diverse unseen setups, substantially outperforming existing methods in terms of generalization ability and establishing the first generalizable solution to projector compensation. The code and dataset are available on our project page: https://hai‑bo‑li.github.io/SIComp/
Authors:Chihiro Nakatani, Norimichi Ukita, Jean-Marc Odobez
Abstract:
This paper proposes an end‑to‑end shared attention estimation method via group detection. Most previous methods estimate shared attention (SA) without detecting the actual group of people focusing on it, or assume that there is a single SA point in a given image. These issues limit the applicability of SA detection in practice and impact performance. To address them, we propose to simultaneously achieve group detection and shared attention estimation using a two step process: (i) the generation of SA heatmaps relying on individual gaze attention heatmaps and group membership scalars estimated in a group inference; (ii) a refinement of the initial group memberships allowing to account for the initial SA heatmaps, and the final prediction of the SA heatmap. Experiments demonstrate that our method outperforms other methods in group detection and shared attention estimation. Additional analyses validate the effectiveness of the proposed components. Code: https://github.com/chihina/sagd‑CVPRW2026.
Authors:Meng Yu, Kun Zhan
Abstract:
Most existing graph diffusion models have significant bias problems. We observe that the forward diffusion's maximum perturbation distribution in most models deviates from the standard Gaussian distribution, while reverse sampling consistently starts from a standard Gaussian distribution, which results in a reverse‑starting bias. Together with the inherent exposure bias of diffusion models, this results in degraded generation quality. This paper proposes a comprehensive approach to mitigate both biases. To mitigate reverse‑starting bias, we employ a newly designed Langevin sampling algorithm to align with the forward maximum perturbation distribution, establishing a new reverse‑starting point. To address the exposure bias, we introduce a score correction mechanism based on a newly defined score difference. Our approach, which requires no network modifications, is validated across multiple models, datasets, and tasks, achieving state‑of‑the‑art results.Code is at https://github.com/kunzhan/spp
Authors:Dingming Liu, Wenjing Wang, Chen Li, Jing Lyu
Abstract:
Video object removal aims to eliminate target objects from videos while plausibly completing missing regions and preserving spatio‑temporal consistency. Although diffusion models have recently advanced this task, it remains challenging to remove object‑induced side effects (e.g., shadows, reflections, and illumination changes) without compromising overall coherence. This limitation stems from the insufficient physical and semantic understanding of the target object and its interactions with the scene. In this paper, we propose to introduce understanding into erasing from two complementary perspectives. Externally, we introduce a distillation scheme that transfers the relationships between objects and their induced effects from vision foundation models to video diffusion models. Internally, we propose a framewise context cross‑attention mechanism that grounds each denoising block in informative, unmasked context surrounding the target region. External and internal guidance jointly enable our model to understand the target object, its induced effects, and the global background context, resulting in clear and coherent object removal. Extensive experiments demonstrate our state‑of‑the‑art performance, and we establish the first real‑world benchmark for video object removal to facilitate future research and community progress. Our code, data, and models are available at: https://github.com/WeChatCV/UnderEraser.
Authors:Yuheng Jiang, Yiwen Cai, Zihao Wang, Yize Wu, Sicheng Li, Zhuo Su, Shaohui Jiao, Lan Xu
Abstract:
Volumetric video seeks to model dynamic scenes as temporally coherent 4D representations. While recent Gaussian‑based approaches achieve impressive rendering fidelity, they primarily emphasize appearance but are largely agnostic to instance‑level structure, limiting stable tracking and semantic reasoning in highly dynamic scenarios. In this paper, we present Director, a unified spatio‑temporal Gaussian representation that jointly models human performance, high‑fidelity rendering, and instance‑level semantics. Our key insight is that embedding instance‑consistent semantics naturally complements 4D modeling, enabling more accurate scene decomposition while supporting robust dynamic scene understanding. To this end, we leverage temporally aligned instance masks and sentence embeddings derived from Multimodal Large Language Models to supervise the learnable semantic features of each Gaussian via two MLP decoders, enabling language‑aligned 4D representations and enforcing identity consistency over time. To enhance temporal stability, we bridge 2D optical flow with 4D Gaussians and finetune their motions, yielding reliable initialization and reducing drift. For the training, we further introduce a geometry‑aware SDF constraints, along with regularization terms that enforces surface continuity, enhancing temporal coherence in dynamic foreground modeling. Experiments demonstrate that Director achieves temporally coherent 4D reconstructions while simultaneously enabling instance segmentation and open‑vocabulary querying.
Authors:Wonjoon Jin, Jiyun Won, Janghyeok Han, Qi Dai, Chong Luo, Seung-Hwan Baek, Sunghyun Cho
Abstract:
Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine‑grained motion controllability. A central limitation lies in the scarcity of such examples in commonly used training datasets. To address this, we introduce DynaVid, a video synthesis framework that leverages synthetic motion data in training, which is represented as optical flow and rendered using computer graphics pipelines. This approach offers two key advantages. First, synthetic motion offers diverse motion patterns and precise control signals that are difficult to obtain from real data. Second, unlike rendered videos with artificial appearances, rendered optical flow encodes only motion and is decoupled from appearance, thereby preventing models from reproducing the unnatural look of synthetic videos. Building on this idea, DynaVid adopts a two‑stage generation framework: a motion generator first synthesizes motion, and then a motion‑guided video generator produces video frames conditioned on that motion. This decoupled formulation enables the model to learn dynamic motion patterns from synthetic data while preserving visual realism from real‑world videos. We validate our framework on two challenging scenarios, vigorous human motion generation and extreme camera motion control, where existing datasets are particularly limited. Extensive experiments demonstrate that DynaVid improves the realism and controllability in dynamic motion generation and camera motion control.
Authors:Junyoung Jung, Seokwon Kim, Jung Uk Kim
Abstract:
Monocular 3D object detection has achieved impressive performance on densely annotated datasets. However, it struggles when only a fraction of objects are labeled due to the high cost of 3D annotation. This sparsely annotated setting is common in real‑world scenarios where annotating every object is impractical. To address this, we propose a novel framework for sparsely annotated monocular 3D object detection with two key modules. First, we propose Road‑Aware Patch Augmentation (RAPA), which leverages sparse annotations by augmenting segmented object patches onto road regions while preserving 3D geometric consistency. Second, we propose Prototype‑Based Filtering (PBF), which generates high‑quality pseudo‑labels by filtering predictions through prototype similarity and depth uncertainty. It maintains global 2D RoI feature prototypes and selects pseudo‑labels that are both feature‑consistent with learned prototypes and have reliable depth estimates. Our training strategy combines geometry‑preserving augmentation with prototype‑guided pseudo‑labeling to achieve robust detection under sparse supervision. Extensive experiments demonstrate the effectiveness of the proposed method. The source code is available at https://github.com/VisualAIKHU/MonoSAOD .
Authors:Youqi Liao, Shuhao Kang, Jingyu Xu, Olaf Wysocki, Yan Xia, Jianping Li, Zhen Dong, Bisheng Yang, Xieyuanli Chen
Abstract:
Natural language provides an intuitive way to express spatial intent in geospatial applications. While existing localization methods often rely on dense point cloud maps or high‑resolution imagery, OpenStreetMap (OSM) offers a compact and freely available map representation that encodes rich semantic and structural information, making it well‑suited for large‑scale localization. However, text‑to‑OSM (T2O) localization remains largely unexplored. In this paper, we formulate the T2O localization task, which aims to estimate accurate 2D positions in urban environments from textual scene descriptions without relying on geometric observations or GNSS‑based initial location. To support the proposed task, we introduce TOL, a large‑scale benchmark spanning multiple continents and diverse urban environments. TOL contains approximately 121K textual queries paired with OSM map tiles and covers about 316 km of road trajectories across Boston, Karlsruhe, and Singapore. We further propose TOLoc, a coarse‑to‑fine localization framework that explicitly models the semantics of surrounding objects and their directional information. In the coarse stage, direction‑aware features are extracted from both textual descriptions and OSM tiles to construct global descriptors, which are used to retrieve candidate locations for the query. In the fine stage, the query text and top‑1 retrieved tile are jointly processed, where a dedicated alignment module fuses the textual descriptor and local map features to regress the 2‑DoF pose. Experimental results demonstrate that TOLoc achieves strong localization performance, outperforming the best existing method by 6.53%, 9.93%, and 8.32% at 5 m, 10 m, and 25 m thresholds, respectively, and shows strong generalization to unseen environments. Dataset, code and models will be publicly available at: https://github.com/WHU‑USI3DV/TOL.
Authors:Longfei Huang, Yang Yang
Abstract:
Multimodal tabular‑image fusion is an emerging task that has received increasing attention in various domains. However, existing methods may be hindered by gradient conflicts between modalities, misleading the optimization of the unimodal learner. In this paper, we propose a novel Gradient‑Aligned Alternating Learning (GAAL) paradigm to address this issue by aligning modality gradients. Specifically, GAAL adopts an alternating unimodal learning and shared classifier to decouple the multimodal gradient and facilitate interaction. Furthermore, we design uncertainty‑based cross‑modal gradient surgery to selectively align cross‑modal gradients, thereby steering the shared parameters to benefit all modalities. As a result, GAAL can provide effective unimodal assistance and help boost the overall fusion performance. Empirical experiments on widely used datasets reveal the superiority of our method through comparison with various state‑of‑the‑art (SoTA) tabular‑image fusion baselines and test‑time tabular missing baselines. The source code is available at https://github.com/njustkmg/ICME26‑GAAL.
Authors:Yanzhe Liang, Ruijie Zhu, Hanzhi Chang, Zhuoyuan Li, Jiahao Lu, Tianzhu Zhang
Abstract:
We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self‑correction manner from raw video. Existing methods often suffer from incomplete scene initialization for dynamic regions, leading to unstable reconstruction and motion estimation, which often resorts to external dense motion guidance such as pre‑computed optical flow to further stabilize and constrain the reconstruction of dynamic components. However, this introduces additional complexity and potential error propagation. To address these issues, ReFlow integrates a Complete Canonical Space Construction module for enhanced initialization of both static and dynamic regions, and a Separation‑Based Dynamic Scene Modeling module that decouples static and dynamic components for targeted motion supervision. The core of ReFlow is a novel self‑correction flow matching mechanism, consisting of Full Flow Matching to align 3D scene flow with time‑varying 2D observations, and Camera Flow Matching to enforce multi‑view consistency for static objects. Together, these modules enable robust and accurate dynamic scene reconstruction. Extensive experiments across diverse scenarios demonstrate that ReFlow achieves superior reconstruction quality and robustness, establishing a novel self‑correction paradigm for monocular 4D reconstruction.
Authors:Da Zhang, Gao Junyu, Zhao Zhiyuan
Abstract:
Semantic segmentation of low‑altitude UAV imagery presents unique challenges due to extreme scale variations, complex object boundaries, and limited computational resources on edge devices. Existing transformer‑based segmentation methods achieve remarkable performance but incur high computational overhead, while lightweight approaches struggle to capture fine‑grained details in high‑resolution aerial scenes. To address these limitations, we propose PBSeg, an efficient prototype‑based segmentation framework tailored for UAV applications. PBSeg introduces a novel prototype‑based cross‑attention (PBCA) that exploits feature redundancy to reduce computational complexity while maintaining segmentation quality. The framework incorporates an efficient multi‑scale feature extraction module that combines deformable convolutions (DConv) with context‑aware modulation (CAM) to capture both local details and global semantics. Experiments on two challenging UAV datasets demonstrate the effectiveness of the proposed approach. PBSeg achieves 71.86% mIoU on UAVid and 80.92% mIoU on UDD6, establishing competitive performance while maintaining computational efficiency. Code is available at https://github.com/zhangda1018/PBSeg.
Authors:Zhisheng Huang, Jiahao Chen, Cheng Lin, Chenyu Hu, Hanzhuo Huang, Zhengming Yu, Mengfei Li, Yuheng Liu, Zekai Gu, Zibo Zhao, Yuan Liu, Xin Li, Wenping Wang
Abstract:
Sparse‑view 3D modeling represents a fundamental tension between reconstruction fidelity and generative plausibility. While feed‑forward reconstruction excels in efficiency and input alignment, it often lacks the global priors needed for structural completeness. Conversely, diffusion‑based generation provides rich geometric details but struggles with multi‑view consistency. We present UniRecGen, a unified framework that integrates these two paradigms into a single cooperative system. To overcome inherent conflicts in coordinate spaces, 3D representations, and training objectives, we align both models within a shared canonical space. We employ disentangled cooperative learning, which maintains stable training while enabling seamless collaboration during inference. Specifically, the reconstruction module is adapted to provide canonical geometric anchors, while the diffusion generator leverages latent‑augmented conditioning to refine and complete the geometric structure. Experimental results demonstrate that UniRecGen achieves superior fidelity and robustness, outperforming existing methods in creating complete and consistent 3D models from sparse observations. Code is available at https://github.com/zsh523/UniRecGen.
Authors:Abhishek Saroha, Huajian Zeng, Xingxing Zuo, Daniel Cremers, Xi Wang
Abstract:
Understanding and predicting object motion from egocentric video is fundamental to embodied perception and interaction. However, generating physically consistent 6DoF trajectories remains challenging due to occlusions, fast motion, and the lack of explicit physical reasoning in existing generative models. We present EgoFlow, a flow‑matching framework that synthesizes realistic and physically plausible trajectories conditioned on multimodal egocentric observations. EgoFlow employs a hybrid Mamba‑Transformer‑Perceiver architecture to jointly model temporal dynamics, scene geometry, and semantic intent, while a gradient‑guided inference process enforces differentiable physical constraints such as collision avoidance and motion smoothness. This combination yields coherent and controllable motion generation without post‑hoc filtering or additional supervision. Experiments on real‑world datasets HD‑EPIC, EgoExo4D, and HOT3D show that EgoFlow outperforms diffusion‑based and transformer baselines in accuracy, generalization, and physical realism, reducing collision rates by up to 79%, and strong generalization to unseen scenes. Our results highlight the promise of flow‑based generative modeling for scalable and physically grounded egocentric motion understanding.
Authors:Syed Ahsan Masud Zaidi, Lior Shamir, William Hsu, Scott Dietrich, Talha Zaidi
Abstract:
American football practice generates video at scale, yet the interaction of interest occupies only a brief window of each long, untrimmed clip. Reliable biomechanical analysis, therefore, depends on spatiotemporal localization that identifies both the interacting entities and the onset of contact. We study First Point of Contact (FPOC), defined as the first frame in which a player physically touches a tackle dummy, in unconstrained practice footage with camera motion, clutter, multiple similarly equipped athletes, and rapid pose changes around impact. We present GRAZE, a training‑free pipeline for FPOC localization that requires no labeled tackle‑contact examples. GRAZE uses Grounding DINO to discover candidate player‑dummy interactions, refines them with motion‑aware temporal reasoning, and uses SAM2 as an explicit pixel‑level verifier of contact rather than relying on detection confidence alone. This separation between candidate discovery and contact confirmation makes the approach robust to cluttered scenes and unstable grounding near impact. On 738 tackle‑practice videos, GRAZE produces valid outputs for 97.4% of clips and localizes FPOC within \pm 10 frames on 77.5% of all clips and within \pm 20 frames on 82.7% of all clips. These results show that frame‑accurate contact onset localization in real‑world practice footage is feasible without task‑specific training.
Authors:Nermin Samet, Gilles Puy, Renaud Marlet
Abstract:
This paper presents a new method for the zero‑shot open‑vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image‑text modality gap that is intrinsic to approaches based on Vision Language Models (VLMs) such as CLIP, our method relies instead on image generation from text, to create prototype images. Given a 3D network distilled from a 2D Vision Foundation Model (VFM), we then label a point cloud by matching 3D point features with 2D image features of these prototypes. Our method is state‑of‑the‑art for OVSS on nuScenes and SemanticKITTI. Code, pre‑trained models, and generated images are available at https://github.com/valeoai/IGLOSS.
Authors:Neo Christopher Chung, Maxim Laletin
Abstract:
Vision transformers (ViT) rely on attention mechanism to weigh input features, and therefore attention scores have naturally been considered as explanations for its decision‑making process. However, attention scores are almost always non‑zero, resulting in noisy and diffused attention maps and limiting interpretability. Can we quantify uncertainty measures of attention scores and obtain regularized attention scores? To this end, we consider attention scores of ViT in a statistical framework where independent noise would lead to insignificant yet non‑zero scores. Leveraging statistical learning techniques, we introduce the bootstrapping for attention scores which generates a baseline distribution of attention scores by resampling input features. Such a bootstrap distribution is then used to estimate significances and posterior probabilities of attention scores. In natural and medical images, the proposed \emphAttention Regularization approach demonstrates a straightforward removal of spurious attention arising from noise, drastically improving shrinkage and sparsity. Quantitative evaluations are conducted using both simulation and real‑world datasets. Our study highlights bootstrapping as a practical regularization tool when using attention scores as explanations for ViT.
Code available: https://github.com/ncchung/AttentionRegularization
Authors:Marco Morini, Sara Sarto, Marcella Cornia, Lorenzo Baraldi
Abstract:
Answering questions about images often requires combining visual understanding with external knowledge. Multimodal Large Language Models (MLLMs) provide a natural framework for this setting, but they often struggle to identify the most relevant visual and textual evidence when answering knowledge‑intensive queries. In such scenarios, models must integrate visual cues with retrieved textual evidence that is often noisy or only partially relevant, while also localizing fine‑grained visual information in the image. In this work, we introduce Look Twice (LoT), a training‑free inference‑time framework that improves how pretrained MLLMs utilize multimodal evidence. Specifically, we exploit the model attention patterns to estimate which visual regions and retrieved textual elements are relevant to a query, and then generate the answer conditioned on this highlighted evidence. The selected cues are highlighted through lightweight prompt‑level markers that encourage the model to re‑attend to the relevant evidence during generation. Experiments across multiple knowledge‑based VQA benchmarks show consistent improvements over zero‑shot MLLMs. Additional evaluations on vision‑centric and hallucination‑oriented benchmarks further demonstrate that visual evidence highlighting alone improves model performance in settings without textual context, all without additional training or architectural modifications. Source code will be publicly released.
Authors:Yao Jiang, Zhongkuan Mao, Xuan Wu, Keren Fu, Qijun Zhao
Abstract:
Camouflaged scene understanding (CSU) has attracted significant attention due to its broad practical implications. However, in this field, robust image‑text cross‑modal alignment remains under‑explored, hindering deeper understanding of camouflaged scenarios and their related applications. To this end, we focus on the typical image‑text retrieval task, and formulate a new task dubbed ``camouflage‑aware image‑text retrieval'' (CA‑ITR). We first construct a dedicated camouflage image‑text retrieval dataset (CamoIT), comprising ~10.5K samples with multi‑granularity textual annotations. Benchmark results conducted on CamoIT reveal the underlying challenges of CA‑ITR for existing cutting‑edge retrieval techniques, which are mainly caused by objects' camouflage properties as well as those complex image contents. As a solution, we propose a camouflage‑expert collaborative network (CECNet), which features a dual‑branch visual encoder: one branch captures holistic image representations, while the other incorporates a dedicated model to inject representations of camouflaged objects. A novel confidence‑conditioned graph attention (C\textsuperscript2GA) mechanism is incorporated to exploit the complementarity across branches. Comparative experiments show that CECNet achieves ~29% overall CA‑ITR accuracy boost, surpassing seven representative retrieval models. The dataset and code will be available at https://github.com/jiangyao‑scu/CA‑ITR.
Authors:J. E. Domínguez-Vidal
Abstract:
Foundation vision‑language models are becoming increasingly relevant to robotics because they can provide richer semantic perception than narrow task‑specific pipelines. However, their practical adoption in robot software stacks still depends on reproducible middleware integrations rather than on model quality alone. Florence‑2 is especially attractive in this regard because it unifies captioning, optical character recognition, open‑vocabulary detection, grounding and related vision‑language tasks within a comparatively manageable model size. This article presents a ROS 2 wrapper for Florence‑2 that exposes the model through three complementary interaction modes: continuous topic‑driven processing, synchronous service calls and asynchronous actions. The wrapper is designed for local execution and supports both native installation and Docker container deployment. It also combines generic JSON outputs with standard ROS 2 message bindings for detection‑oriented tasks. A functional validation is reported together with a throughput study on several GPUs, showing that local deployment is feasible with consumer grade hardware. The repository is publicly available here: https://github.com/JEDominguezVidal/florence2_ros2_wrapper
Authors:Hanzhe Liang, Luocheng Zhang, Junyang Xia, HanLiang Zhou, Bingyang Guo, Yingxi Xie, Can Gao, Ruiyun Yu, Jinbao Wang, Pan Li
Abstract:
Although self‑supervised 3D anomaly detection assumes that acquiring high‑precision point clouds is computationally expensive, in real manufacturing scenarios it is often feasible to collect a limited number of anomalous samples. Therefore, we study open‑set supervised 3D anomaly detection, where the model is trained with only normal samples and a small number of known anomalous samples, aiming to identify unknown anomalies at test time. We present Open‑Industry, a high‑quality industrial dataset containing 15 categories, each with five real anomaly types collected from production lines. We first adapt general open‑set anomaly detection methods to accommodate 3D point cloud inputs better. Building upon this, we propose Open3D‑AD, a point‑cloud‑oriented approach that leverages normal samples, simulated anomalies, and partially observed real anomalies to model the probability density distributions of normal and anomalous data. Then, we introduce a simple Correspondence Distributions Subsampling to reduce the overlap between normal and non‑normal distributions, enabling stronger dual distributions modeling. Based on these contributions, we establish a comprehensive benchmark and evaluate the proposed method extensively on Open‑Industry as well as established datasets including Real3D‑AD and Anomaly‑ShapeNet. Benchmark results and ablation studies demonstrate the effectiveness of Open3D‑AD and further reveal the potential of open‑set supervised 3D anomaly detection.
Authors:Prantik Deb, Srimanth Dhondy, N. Ramakrishna, Anu Kapoor, Raju S. Bapi, Tapabrata Chakraborti
Abstract:
Chest X‑ray (CXR) segmentation is an important step in computer‑aided diagnosis, yet deploying large foundation models in clinical settings remains challenging due to computational constraints. We propose AdaLoRA‑QAT, a two‑stage fine‑tuning framework that combines adaptive low‑rank encoder adaptation with full quantization‑aware training. Adaptive rank allocation improves parameter efficiency, while selective mixed‑precision INT8 quantization preserves structural fidelity crucial for clinical reliability. Evaluated across large‑scale CXR datasets, AdaLoRA‑QAT achieves 95.6% Dice, matching full‑precision SAM decoder fine‑tuning while reducing trainable parameters by 16.6× and yielding 2.24× model compression. A Wilcoxon signed‑rank test confirms that quantization does not significantly degrade segmentation accuracy. These results demonstrate that AdaLoRA‑QAT effectively balances accuracy, efficiency, and structural trust‑worthiness, enabling compact and deployable foundation models for medical image segmentation. Code and pretrained models are available at: https://prantik‑pdeb.github.io/adaloraqat.github.io/
Authors:Hao Zhang, Lue Fan, Weikang Bian, Zehuan Wu, Lewei Lu, Zhaoxiang Zhang, Hongsheng Li
Abstract:
We present ReinDriveGen, a framework that enables full controllability over dynamic driving scenes, allowing users to freely edit actor trajectories to simulate safety‑critical corner cases such as front‑vehicle collisions, drifting cars, vehicles spinning out of control, pedestrians jaywalking, and cyclists cutting across lanes. Our approach constructs a dynamic 3D point cloud scene from multi‑frame LiDAR data, introduces a vehicle completion module to reconstruct full 360° geometry from partial observations, and renders the edited scene into 2D condition images that guide a video diffusion model to synthesize realistic driving videos. Since such edited scenarios inevitably fall outside the training distribution, we further propose an RL‑based post‑training strategy with a pairwise preference model and a pairwise reward mechanism, enabling robust quality improvement under out‑of‑distribution conditions without ground‑truth supervision. Extensive experiments demonstrate that ReinDriveGen outperforms existing approaches on edited driving scenarios and achieves state‑of‑the‑art results on novel ego viewpoint synthesis.
Authors:Yaoqin Ye, Yiteng Xu, Qin Sun, Xinge Zhu, Yujing Sun, Yuexin Ma
Abstract:
Human behaviors in real‑world environments are inherently interactive, with an individual's motion shaped by surrounding agents and the scene. Such capabilities are essential for applications in virtual avatars, interactive animation, and human‑robot collaboration. We target real‑time human interaction‑to‑reaction generation, which generates the ego's future motion from dynamic multi‑source cues, including others' actions, scene geometry, and optional high‑level semantic inputs. This task is fundamentally challenging due to (i) limited and fragmented interaction data distributed across heterogeneous single‑person, human‑human, and human‑scene domains, and (ii) the need to produce low‑latency yet high‑fidelity motion responses during continuous online interaction. To address these challenges, we propose ReMoGen (Reaction Motion Generation), a modular learning framework for real‑time interaction‑to‑reaction generation. ReMoGen leverages a universal motion prior learned from large‑scale single‑person motion datasets and adapts it to target interaction domains through independently trained Meta‑Interaction modules, enabling robust generalization under data‑scarce and heterogeneous supervision. To support responsive online interaction, ReMoGen performs segment‑level generation together with a lightweight Frame‑wise Segment Refinement module that incorporates newly observed cues at the frame level, improving both responsiveness and temporal coherence without expensive full‑sequence inference. Extensive experiments across human‑human, human‑scene, and mixed‑modality interaction settings show that ReMoGen produces high‑quality, coherent, and responsive reactions, while generalizing effectively across diverse interaction scenarios.
Authors:Yuheng Zhang, Mengfei Duan, Kunyu Peng, Yuhang Wang, Di Wen, Danda Pani Paudel, Luc Van Gool, Kailun Yang
Abstract:
3D semantic occupancy prediction is central to autonomous driving, yet current methods are vulnerable to long‑tailed class bias and out‑of‑distribution (OOD) inputs, often overconfidently assigning anomalies to rare classes. We present ProOOD, a lightweight, plug‑and‑play method that couples prototype‑guided refinement with training‑free OOD scoring. ProOOD comprises (i) prototype‑guided semantic imputation that fills occluded regions with class‑consistent features, (ii) prototype‑guided tail mining that strengthens rare‑class representations to curb OOD absorption, and (iii) EchoOOD, which fuses local logit coherence with local and global prototype matching to produce reliable voxel‑level OOD scores. Extensive experiments on five datasets demonstrate that ProOOD achieves state‑of‑the‑art performance on both in‑distribution 3D occupancy prediction and OOD detection. On SemanticKITTI, it surpasses baselines by +3.57% mIoU overall and +24.80% tail‑class mIoU; on VAA‑KITTI, it improves AuPRCr by +19.34 points, with consistent gains across benchmarks. These improvements yield more calibrated occupancy estimates and more reliable OOD detection in safety‑critical urban driving. The source code is publicly available at https://github.com/7uHeng/ProOOD.
Authors:Fengyuan Yang, Luying Huang, Jiazhi Guan, Quanwei Yang, Dongwei Pan, Jianglin Fu, Haocheng Feng, Wei He, Kaisiyuan Wang, Hang Zhou, Angela Yao
Abstract:
Recent advances in Video Foundation Models (VFMs) have revolutionized human‑centric video synthesis, yet fine‑grained and independent editing of subjects and scenes remains a critical challenge. Recent attempts to incorporate richer environment control through rigid 3D geometric compositions often encounter a stark trade‑off between precise control and generative flexibility. Furthermore, the heavy 3D pre‑processing still limits practical scalability. In this paper, we propose ONE‑SHOT, a parameter‑efficient framework for compositional human‑environment video generation. Our key insight is to factorize the generative process into disentangled signals. Specifically, we introduce a canonical‑space injection mechanism that decouples human dynamics from environmental cues via cross‑attention. We also propose Dynamic‑Grounded‑RoPE, a novel positional embedding strategy that establishes spatial correspondences between disparate spatial domains without any heuristic 3D alignments. To support long‑horizon synthesis, we introduce a Hybrid Context Integration mechanism to maintain subject and scene consistency across minute‑level generations. Experiments demonstrate that our method significantly outperforms state‑of‑the‑art methods, offering superior structural control and creative diversity for video synthesis. Our project has been available on: https://martayang.github.io/ONE‑SHOT/.
Authors:Yueh-Cheng Liu, Jozef Hladký, Matthias Nießner, Angela Dai
Abstract:
Recent advances in 3D Gaussian Splatting (3DGS) present two main directions: feed‑forward models offer fast inference in sparse‑view settings, while per‑scene optimization yields high‑quality renderings but is computationally expensive. To combine the benefits of both, we introduce Diff3R, a novel framework that explicitly bridges feed‑forward prediction and test‑time optimization. By incorporating a differentiable 3DGS optimization layer directly into the training loop, our network learns to predict an optimal initialization for test‑time optimization rather than a conventional zero‑shot result. To overcome the computational cost of backpropagating through the optimization steps, we propose computing gradients via the Implicit Function Theorem and a scalable, matrix‑free PCG solver tailored for 3DGS optimization. Additionally, we incorporate a data‑driven uncertainty model into the optimization process by adaptively controlling how much the parameters are allowed to change during optimization. This approach effectively mitigates overfitting in under‑constrained regions and increases robustness against input outliers. Since our proposed optimization layer is model‑agnostic, we show that it can be seamlessly integrated into existing feed‑forward 3DGS architectures for both pose‑given and pose‑free methods, providing improvements for test‑time optimization.
Authors:Miro Miranda, Deepak Pathak, Patrick Helber, Benjamin Bischke, Hiba Najjar, Francisco Mena, Cristhian Sanchez, Akshay Pai, Diego Arenas, Matias Valdenegro-Toro, Marcela Charfuelan, Marlon Nuske, Andreas Dengel
Abstract:
Crop yield prediction requires substantial data to train scalable models. However, creating yield prediction datasets is constrained by high acquisition costs, heterogeneous data quality, and data privacy regulations. Consequently, existing datasets are scarce, low in quality, or limited to regional levels or single crop types, hindering the development of scalable data‑driven solutions. In this work, we release YieldSAT, a large, high‑quality, and multimodal dataset for high‑resolution crop yield prediction. YieldSAT spans various climate zones across multiple countries, including Argentina, Brazil, Uruguay, and Germany, and includes major crop types, including corn, rapeseed, soybeans, and wheat, across 2,173 expert‑curated fields. In total, over 12.2 million yield samples are available, each with a spatial resolution of 10 m. Each field is paired with multispectral satellite imagery, resulting in 113,555 labeled satellite images, complemented by auxiliary environmental data. We demonstrate the potential of large‑scale and high‑resolution crop yield prediction as a pixel regression task by comparing various deep learning models and data fusion architectures. Furthermore, we highlight open challenges arising from severe distribution shifts in the ground truth data under real‑world conditions. To mitigate this, we explore a domain‑informed Deep Ensemble approach that exhibits significant performance gains. The dataset is available at https://yieldsat.github.io/.
Authors:Michael Steiner, Zhang Chen, Alexander Richard, Vasu Agrawal, Markus Steinberger, Michael Zollhöfer
Abstract:
A photorealistic and immersive human avatar experience demands capturing fine, person‑specific details such as cloth and hair dynamics, subtle facial expressions, and characteristic motion patterns. Achieving this requires large, high‑quality datasets, which often introduce ambiguities and spurious correlations when very similar poses correspond to different appearances. Models that fit these details during training can overfit and produce unstable, abrupt appearance changes for novel poses. We propose a 3D Gaussian Splatting avatar model with a spatial MLP backbone that is conditioned on both pose and an appearance latent. The latent is learned during training by an encoder, yielding a compact representation that improves reconstruction quality and helps disambiguate pose‑driven renderings. At driving time, our predictor autoregressively infers the latent, producing temporally smooth appearance evolution and improved stability. Overall, our method delivers a robust and practical path to high‑fidelity, stable avatar driving.
Authors:Zhuchenyang Liu, Yao Zhang, Yu Xiao
Abstract:
2D assembly diagrams are often abstract and hard to follow, creating a need for intelligent assistants that can monitor progress, detect errors, and provide step‑by‑step guidance. In mixed reality settings, such systems must recognize completed and ongoing steps from the camera feed and align them with the diagram instructions. Vision Language Models (VLMs) show promise for this task, but face a depiction gap because assembly diagrams and video frames share few visual features. To systematically assess this gap, we construct IKEA‑Bench, a benchmark of 1,623 questions across 6 task types on 29 IKEA furniture products, and evaluate 19 VLMs (2B‑38B) under three alignment strategies. Our key findings: (1) assembly instruction understanding is recoverable via text, but text simultaneously degrades diagram‑to‑video alignment; (2) architecture family predicts alignment accuracy more strongly than parameter count; (3) video understanding remains a hard bottleneck unaffected by strategy. A three‑level mechanistic analysis further reveals that diagrams and video occupy disjoint ViT subspaces, and that adding text shifts models from visual to text‑driven reasoning. These results identify visual encoding as the primary target for improving cross‑depiction robustness. Project page: https://ryenhails.github.io/IKEA‑Bench/
Authors:Zimo Cao, Yuchen Deng, Haibin Ling, Bingyao Huang
Abstract:
Spatial augmented reality (SAR) directly projects digital content onto physical scenes using projectors, creating immersive experience without head‑mounted displays. However, for SAR to support intelligent interaction, such as reasoning about the scene or answering user queries, it must semantically distinguish between the physical scene and the projected content. Standard Vision Language Models (VLMs) struggle with this virtual‑physical ambiguity, often confusing the two contexts. To address this issue, we introduce ProCap, a novel framework that explicitly decouples projected content from physical scenes. ProCap employs a two‑stage pipeline: first it visually isolates virtual and physical layers via automated segmentation; then it uses region‑aware retrieval to avoid ambiguous semantic context due to projection distortion. To support this, we present RGBP (RGB + Projections), the first large‑scale SAR semantic benchmark dataset, featuring 65 diverse physical scenes and over 180,000 projections with dense, decoupled annotations. Finally, we establish a dual‑captioning evaluation protocol using task‑specific tokens to assess physical scene and projection descriptions independently. Our experiments show that ProCap provides a robust semantic foundation for future SAR research. The source code, pre‑trained models and the RGBP dataset are available on the project page: https://ZimoCao.github.io/ProCap/.
Authors:Yiming Zhang, Weibo Qin, Feng Wang
Abstract:
Deep neural networks have demonstrated excellent performance in SAR target detection tasks but remain susceptible to adversarial attacks. Existing SAR‑specific attack methods can effectively deceive detectors; however, they often introduce noticeable perturbations and are largely confined to digital domain, neglecting physical implementation constrains for attacking SAR systems. In this paper, a novel Adversarial Attenuation Patch (AAP) method is proposed that employs energy‑constrained optimization strategy coupled with an attenuation‑based deployment framework to achieve a seamless balance between attack effectiveness and stealthiness. More importantly, AAP exhibits strong potential for physical realization by aligning with signal‑level electronic jamming mechanisms. Experimental results show that AAP effectively degrades detection performance while preserving high imperceptibility, and shows favorable transferability across different models. This study provides a physical grounded perspective for adversarial attacks on SAR target detection systems and facilitates the design of more covert and practically deployable attack strategies. The source code is made available at https://github.com/boremycin/SAAP.
Authors:Nan Wang, Zhiwei Jin, Chen Chen, Haonan Lu
Abstract:
Document understanding and GUI interaction are among the highest‑value applications of Vision‑Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine‑grained text and small UI elements demand high‑resolution inputs that produce tens of thousands of visual tokens. We observe that this cost is largely wasteful ‑‑ across document and GUI benchmarks, only 22‑‑71% of image patches are pixel‑unique, the rest being exact duplicates of another patch in the same image. We propose PixelPrune, which exploits this pixel‑level redundancy through predictive‑coding‑based compression, pruning redundant patches \emphbefore the Vision Transformer (ViT) encoder. Because it operates in pixel space prior to any neural computation, PixelPrune accelerates both the ViT encoder and the downstream LLM, covering the full inference pipeline. The method is training‑free, requires no learnable parameters, and supports pixel‑lossless compression (τ=0) as well as controlled lossy compression (τ>0). Experiments across three model scales and document and GUI benchmarks show that PixelPrune maintains competitive task accuracy while delivering up to 4.2× inference speedup and 1.9× training acceleration. Code is available at https://github.com/OPPO‑Mente‑Lab/PixelPrune.
Authors:Maximilian Fehrentz, Nicolas Stellwag, Robert Wiebe, Nicole Thorisch, Fabian Grob, Patrick Remerscheid, Ken-Joel Simmoteit, Benjamin D. Killeen, Christian Heiliger, Nassir Navab
Abstract:
Spatiotemporal reasoning is a fundamental capability for artificial intelligence (AI) in soft tissue surgery, paving the way for intelligent assistive systems and autonomous robotics. While 2D vision‑language models show increasing promise at understanding surgical video, the spatial complexity of surgical scenes suggests that reasoning systems may benefit from explicit 4D representations. Here, we propose a framework for equipping surgical agents with spatiotemporal tools based on an explicit 4D representation, enabling AI systems to ground their natural language reasoning in both time and 3D space. Leveraging models for point tracking, depth, and segmentation, we develop a coherent 4D model with spatiotemporally consistent tool and tissue semantics. A Multimodal Large Language Model (MLLM) then acts as an agent on tools derived from the explicit 4D representation (e.g., trajectories) without any fine‑tuning. We evaluate our method on a new dataset of 134 clinically relevant questions and find that the combination of a general purpose reasoning backbone and our 4D representation significantly improves spatiotemporal understanding and allows for 4D grounding. We demonstrate that spatiotemporal intelligence can be "assembled" from 2D MLLMs and 3D computer vision models without additional training. Code, data, and examples are available at https://tum‑ai.github.io/surg4d/
Authors:Samuel Teodoro, Yun Chen, Agus Gunawan, Soo Ye Kim, Jihyong Oh, Munchurl Kim
Abstract:
Motion transfer enables controllable video generation by transferring temporal dynamics from a reference video to synthesize a new video conditioned on a target caption. However, existing Diffusion Transformer (DiT)‑based methods are limited to single‑object videos, restricting fine‑grained control in real‑world scenes with multiple objects. In this work, we introduce MotionGrounder, a DiT‑based framework that firstly handles motion transfer with multi‑object controllability. Our Flow‑based Motion Signal (FMS) in MotionGrounder provides a stable motion prior for target video generation, while our Object‑Caption Alignment Loss (OCAL) grounds object captions to their corresponding spatial regions. We further propose a new Object Grounding Score (OGS), which jointly evaluates (i) spatial alignment between source video objects and their generated counterparts and (ii) semantic consistency between each generated object and its target caption. Our experiments show that MotionGrounder consistently outperforms recent baselines across quantitative, qualitative, and human evaluations.
Authors:Sicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu, Fang Li, Hanbing Li, Long Chen, Zhi-Xin Yang, Jiwen Lu
Abstract:
End‑to‑end autonomous driving has evolved from the conventional paradigm based on sparse perception into vision‑language‑action (VLA) models, which focus on learning language descriptions as an auxiliary task to facilitate planning. In this paper, we propose an alternative Vision‑Geometry‑Action (VGA) paradigm that advocates dense 3D geometry as the critical cue for autonomous driving. As vehicles operate in a 3D world, we think dense 3D geometry provides the most comprehensive information for decision‑making. However, most existing geometry reconstruction methods (e.g., DVGT) rely on computationally expensive batch processing of multi‑frame inputs and cannot be applied to online planning. To address this, we introduce a streaming Driving Visual Geometry Transformer (DVGT‑2), which processes inputs in an online manner and jointly outputs dense geometry and trajectory planning for the current frame. We employ temporal causal attention and cache historical features to support on‑the‑fly inference. To further enhance efficiency, we propose a sliding‑window streaming strategy and use historical caches within a certain interval to avoid repetitive computations. Despite the faster speed, DVGT‑2 achieves superior geometry reconstruction performance on various datasets. The same trained DVGT‑2 can be directly applied to planning across diverse camera configurations without fine‑tuning, including closed‑loop NAVSIM and open‑loop nuScenes benchmarks.
Authors:Monica M. Q. Li, Pierre-Yves Lajoie, Jialiang Liu, Giovanni Beltrame
Abstract:
Efficient multi‑agent 3D mapping is essential for robotic teams operating in unknown environments, but dense representations hinder real‑time exchange over constrained communication links. In multi‑agent Simultaneous Localization and Mapping (SLAM), systems typically rely on a centralized server to merge and optimize the local maps produced by individual agents. However, sharing these large map representations, particularly those generated by recent methods such as Gaussian Splatting, becomes a bottleneck in real‑world scenarios with limited bandwidth. We present an improved multi‑agent RGB‑D Gaussian Splatting SLAM framework that reduces communication load while preserving map fidelity. First, we incorporate a compaction step into our SLAM system to remove redundant 3D Gaussians, without degrading the rendering quality. Second, our approach performs centralized loop closure computation without initial guess, operating in two modes: a pure rendered‑depth mode that requires no data beyond the 3D Gaussians, and a camera‑depth mode that includes lightweight depth images for improved registration accuracy and additional Gaussian pruning. Evaluation on both synthetic and real‑world datasets shows up to 85‑95% reduction in transmitted data compared to state‑of‑the‑art approaches in both modes, bringing 3D Gaussian multi‑agent SLAM closer to practical deployment in real‑world scenarios. Code: https://github.com/lemonci/coko‑slam
Authors:Clémentine Grethen, Yuang Shi, Simone Gasparini, Géraldine Morin
Abstract:
Accurate perception of lunar surfaces is critical for modern lunar exploration missions. However, developing robust learning‑based perception systems is hindered by the lack of datasets that provide both geometric and photometric supervision. Existing lunar datasets typically lack either geometric ground truth, photometric realism, illumination diversity, or large‑scale coverage. In this paper, we introduce MoonAnything, a unified benchmark built on real lunar topography with physically‑based rendering, providing the first comprehensive geometric and photometric supervision under diverse illumination with large scale. The benchmark comprises two complementary sub‑datasets : i) LunarGeo provides stereo images with corresponding dense depth maps and camera calibration enabling 3D reconstruction and pose estimation; ii) LunarPhoto provides photorealistic images using a spatially‑varying BRDF model, along with multi‑illumination renderings under real solar configurations, enabling reflectance estimation and illumination‑robust perception. Together, these datasets offer over 130K samples with comprehensive supervision. Beyond lunar applications, MoonAnything offers a unique setting and challenging testbed for algorithms under low‑textured, high‑contrast conditions and applies to other airless celestial bodies and could generalize beyond. We establish baselines using state‑of‑the‑art methods and release the complete dataset along with generation tools to support community extension: https://github.com/clementinegrethen/MoonAnything.
Authors:Zhengxian Yang, Fei Xie, Xutao Xue, Rui Zhang, Taicheng Huang, Yang Liu, Mengqi Ji, Tao Yu
Abstract:
3D Gaussian Splatting (3DGS) has enabled efficient 3D scene reconstruction from everyday images with real‑time, high‑fidelity rendering, greatly advancing VR/AR applications. Fisheye cameras, with their wider field of view (FOV), promise high‑quality reconstructions from fewer inputs and have recently attracted much attention. However, since 3DGS relies on rasterization, most subsequent works involving fisheye camera inputs first undistort images before training, which introduces two problems: 1) Black borders at image edges cause information loss and negate the fisheye's large FOV advantage; 2) Undistortion's stretch‑and‑interpolate resampling spreads each pixel's value over a larger area, diluting detail density ‑‑ causes 3DGS overfitting these low‑frequency zones, producing blur and floating artifacts. In this work, we integrate fisheye camera model into the original 3DGS framework, enabling native fisheye image input for training without preprocessing. Despite correct modeling, we observed that the reconstructed scenes still exhibit floaters at image edges: Distortion increases toward the periphery, and 3DGS's original per‑iteration random‑selecting‑view optimization ignores the cross‑view correlations of a Gaussian, leading to extreme shapes (e.g., oversized or elongated) that degrade reconstruction quality. To address this, we introduce a feature‑overlap‑driven cross‑view joint optimization strategy that establishes consistent geometric and photometric constraints across views‑a technique equally applicable to existing pinhole‑camera‑based pipelines. Our DirectFisheye‑GS matches or surpasses state‑of‑the‑art performance on public datasets. Project Page: https://yzxqh.github.io/DirectFisheye‑GS/ .
Authors:Shuo Jin, Siyue Yu, Bingfeng Zhang, Chao Yao, Meiqin Liu, Jimin Xiao
Abstract:
Referring image segmentation aims to segment specific targets based on a natural text expression. Recently, parameter‑efficient tuning (PET) has emerged as a promising paradigm. However, existing PET‑based methods often suffer from the fact that visual features can't emphasize the text‑referred target instance but activate co‑category yet unrelated objects. We analyze and quantify this problem, terming it the `non‑target activation' (NTA) issue. To address this, we propose a novel framework, TALENT, which utilizes target‑aware efficient tuning for PET‑based RIS. Specifically, we first propose a Rectified Cost Aggregator (RCA) to efficiently aggregate text‑referred features. Then, to calibrate `NTA' into accurate target activation, we adopt a Target‑aware Learning Mechanism (TLM), including contextual pairwise consistency learning and target‑centric contrastive learning. The former uses the sentence‑level text feature to achieve a holistic understanding of the referent and constructs a text‑referred affinity map to optimize the semantic association of visual features. The latter further enhances target localization to discover the distinct instance while suppressing associations with other unrelated ones. The two objectives work in concert and address `NTA' effectively. Extensive evaluations show that TALENT outperforms existing methods across various metrics (e.g., 2.5% mIoU gains on G‑Ref val set). Our codes will be released at: https://github.com/Kimsure/TALENT.
Authors:Yichen Xie, Yixiao Wang, Shuqi Zhao, Cheng-En Wu, Masayoshi Tomizuka, Jianwen Xie, Hao-Shu Fang
Abstract:
The generalization ability of imitation learning policies for robotic manipulation is fundamentally constrained by the diversity of expert demonstrations, while collecting demonstrations across varied environments is costly and difficult in practice. In this paper, we propose a practical framework that exploits inherent scene diversity without additional human effort by scaling camera views during demonstration collection. Instead of acquiring more trajectories, multiple synchronized camera perspectives are used to generate pseudo‑demonstrations from each expert trajectory, which enriches the training distribution and improves viewpoint invariance in visual representations. We analyze how different action spaces interact with view scaling and show that camera‑space representations further enhance diversity. In addition, we introduce a multiview action aggregation method that allows single‑view policies to benefit from multiple cameras during deployment. Extensive experiments in simulation and real‑world manipulation tasks demonstrate significant gains in data efficiency and generalization compared to single‑view baselines. Our results suggest that scaling camera views provides a practical and scalable solution for imitation learning, which requires minimal additional hardware setup and integrates seamlessly with existing imitation learning algorithms. The website of our project is https://yichen928.github.io/robot_multiview.
Authors:Zhijin He, Shuo Jin, Siyue Yu, Shuwei Wu, Bingfeng Zhang, Li Yu, Jimin Xiao
Abstract:
Co‑salient Object Detection (CoSOD) aims to segment salient objects that consistently appear across a group of related images. Despite the notable progress achieved by recent training‑based approaches, they still remain constrained by the closed‑set datasets and exhibit limited generalization. However, few studies explore the potential of Vision Foundation Models (VFMs) to address CoSOD, which demonstrate a strong generalized ability and robust saliency understanding. In this paper, we investigate and leverage VFMs for CoSOD, and further propose a novel training‑free method, TF‑SSD, through the synergy between SAM and DINO. Specifically, we first utilize SAM to generate comprehensive raw proposals, which serve as a candidate mask pool. Then, we introduce a quality mask generator to filter out redundant masks, thereby acquiring a refined mask set. Since this generator is built upon SAM, it inherently lacks semantic understanding of saliency. To this end, we adopt an intra‑image saliency filter that employs DINO's attention maps to identify visually salient masks within individual images. Moreover, to extend saliency understanding across group images, we propose an inter‑image prototype selector, which computes similarity scores among cross‑image prototypes to select masks with the highest score. These selected masks serve as final predictions for CoSOD. Extensive experiments show that our TF‑SSD outperforms existing methods (e.g., 13.7% gains over the recent training‑free method). Codes are available at https://github.com/hzz‑yy/TF‑SSD.
Authors:Suwoong Yeom, Joonsik Nam, Seunggyu Choi, Lucas Yunkyu Lee, Sangmin Kim, Jaesik Park, Joonsoo Kim, Kugjin Yun, Kyeongbo Kong, Sukju Kang
Abstract:
Recent 4D Gaussian Splatting (4DGS) methods achieve impressive dynamic scene reconstruction but often rely on piecewise linear velocity approximations and short temporal windows. This disjointed modeling leads to severe temporal fragmentation, forcing primitives to be repeatedly eliminated and regenerated to track complex nonlinear dynamics. This makeshift approximation eliminates the long‑term temporal identity of objects and causes an inevitable proliferation of Gaussians, hindering scalability to extended video sequences. To address this, we propose TRiGS, a novel 4D representation that utilizes unified, continuous geometric transformations. By integrating SE(3) transformations, hierarchical Bezier residuals, and learnable local anchors, TRiGS models geometrically consistent rigid motions for individual primitives. This continuous formulation preserves temporal identity and effectively mitigates unbounded memory growth. Extensive experiments demonstrate that TRiGS achieves high fidelity rendering on standard benchmarks while uniquely scaling to extended video sequences (e.g., 600 to 1200 frames) without severe memory bottlenecks, significantly outperforming prior works in temporal stability.
Authors:Jeffrey A. Chan-Santiago, Mubarak Shah
Abstract:
Training machine learning models on massive datasets is expensive and time‑consuming. Dataset distillation addresses this by creating a small synthetic dataset that achieves the same performance as the full dataset. Recent methods use diffusion models to generate distilled data, either by promoting diversity or matching training gradients. However, existing approaches produce redundant training signals, where samples convey overlapping information. Empirically, disjoint subsets of distilled datasets capture 80‑90% overlapping signals. This redundancy stems from optimizing visual diversity or average training dynamics without accounting for similarity across samples, leading to datasets where multiple samples share similar information rather than complementary knowledge. We propose learnability‑driven dataset distillation, which constructs synthetic datasets incrementally through successive stages. Starting from a small set, we train a model and generate new samples guided by learnability scores that identify what the current model can learn from, creating an adaptive curriculum. We introduce Learnability‑Guided Diffusion (LGD), which balances training utility for the current model with validity under a reference model to generate curriculum‑aligned samples. Our approach reduces redundancy by 39.1%, promotes specialization across training stages, and achieves state‑of‑the‑art results on ImageNet‑1K (60.1%), ImageNette (87.2%), and ImageWoof (72.9%). Our code is available on our project page https://jachansantiago.github.io/learnability‑guided‑distillation/.
Authors:Jihwan Park, Chanhyeong Yang, Jinyoung Park, Taehoon Song, Hyunwoo J. Kim
Abstract:
Weakly‑supervised Human‑Object Interaction (HOI) detection is essential for scalable scene understanding, as it learns interactions from only image‑level annotations. Due to the lack of localization signals, prior works typically rely on an external object detector to generate candidate pairs and then infer their interactions through pairwise reasoning. However, this framework often struggles to scale due to the substantial computational cost incurred by enumerating numerous instance pairs. In addition, it suffers from false positives arising from non‑interactive combinations, which hinder accurate instance‑level HOI reasoning. To address these issues, we introduce Relational Grounding Transformer (RegFormer), a versatile interaction recognition module for efficient and accurate HOI reasoning. Under image‑level supervision, RegFormer leverages spatially grounded signals as guidance for the reasoning process and promotes locality‑aware interaction learning. By learning localized interaction cues, our module distinguishes humans, objects, and their interactions, enabling direct transfer from image‑level interaction reasoning to precise and efficient instance‑level reasoning without additional training. Our extensive experiments and analyses demonstrate that RegFormer effectively learns spatial cues for instance‑level interaction reasoning, operates with high efficiency, and even achieves performance comparable to fully supervised models. Our code is available at https://github.com/mlvlab/RegFormer.
Authors:Weifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li, Yuhuan Lin, Hanqiu Deng, Wenbing Tao, Yong Liu, Chengjie Wang
Abstract:
Open‑Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image‑text pairs for rare categories. This results in suboptimal performance in specialized domains or with complex objects. Recent visual‑prompted methods partially address these issues but often involve complex multi‑modal designs and multi‑stage optimizations, prolonging the development cycle. Additionally, effective training strategies for data‑driven OSOD models remain largely unexplored. To address these challenges, we propose PET‑DINO, a universal detector supporting both text and visual prompts. Our Alignment‑Friendly Visual Prompt Generation (AFVPG) module builds upon an advanced text‑prompted detector, addressing the limitations of text representation guidance and reducing the development cycle. We introduce two prompt‑enriched training strategies: Intra‑Batch Parallel Prompting (IBP) at the iteration level and Dynamic Memory‑Driven Prompting (DMD) at the overall training level. These strategies enable simultaneous modeling of multiple prompt routes, facilitating parallel alignment with diverse real‑world usage scenarios. Comprehensive experiments demonstrate that PET‑DINO exhibits competitive zero‑shot object detection capabilities across various prompt‑based detection protocols. These strengths can be attributed to inheritance‑based philosophy and prompt‑enriched training strategies, which play a critical role in building an effective generic object detector. Project page: https://fuweifuvtoo.github.io/pet‑dino.
Authors:Chengcheng Lv, Rushi Li, Mincheng Wu, Xiufang Shi, Zhenyu Wen, Shibo He
Abstract:
Road masks obtained from remote sensing images effectively support a wide range of downstream tasks. In recent years, most studies have focused on improving the performance of fully automatic segmentation models for this task, achieving significant gains. However, current fully automatic methods are still insufficient for identifying certain challenging road segments and often produce false positive and false negative regions. Moreover, fully automatic segmentation does not support local segmentation of regions of interest or refinement of existing masks. Although the SAM model is widely used as an interactive segmentation model and performs well on natural images, it shows poor performance in remote sensing road segmentation and cannot support fine‑grained local refinement. To address these limitations, we propose PC‑SAM, which integrates fully automatic road segmentation and interactive segmentation within a unified framework. By carefully designing a fine‑tuning strategy, the influence of point prompts is constrained to their corresponding patches, overcoming the inability of the original SAM to perform fine local corrections and enabling fine‑grained interactive mask refinement. Extensive experiments on several representative remote sensing road segmentation datasets demonstrate that, when combined with point prompts, PC‑SAM significantly outperforms state‑of‑the‑art fully automatic models in road mask segmentation, while also providing flexible local mask refinement and local road segmentation. The code will be available at https://github.com/Cyber‑CCOrange/PC‑SAM.
Authors:Yabin Zhang, Chong Wang, Yunhe Gao, Jiaming Liu, Maya Varma, Justin Xu, Sophie Ostmeier, Jin Long, Sergios Gatidis, Seena Dehkharghani, Arne Michalson, Eun Kyoung Hong, Christian Bluethgen, Haiwei Henry Guo, Alexander Victor Ortiz, Stephan Altmayer, Sandhya Bodapati, Joseph David Janizek, Ken Chang, Jean-Benoit Delbrouck, Akshay S. Chaudhari, Curtis P. Langlotz
Abstract:
Chest X‑rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic errors. Although artificial intelligence (AI) systems have shown promise for CXR interpretation, most generate only final predictions, without making explicit how visual evidence is translated into radiographic findings and diagnostic predictions. We present CheXOne, a reasoning‑enabled vision‑language model for CXR interpretation. CheXOne jointly generates diagnostic predictions and explicit, clinically grounded reasoning traces that connect visual evidence, radiographic findings, and these predictions. The model is trained on 14.7 million instruction and reasoning samples curated from 30 public datasets spanning 36 CXR interpretation tasks, using a two‑stage framework that combines instruction tuning with reinforcement learning to improve reasoning quality. We evaluate CheXOne in zero‑shot settings across visual question answering, report generation, visual grounding and reasoning assessment, covering 17 evaluation settings. CheXOne outperforms existing medical and general‑domain foundation models and achieves strong performance on independent public benchmarks. A clinical reader study demonstrates that CheXOne‑drafted reports are comparable to or better than resident‑written reports in 55% of cases, while effectively addressing clinical indications and enhancing both report writing and CXR interpretation efficiency. Further analyses involving radiologists reveal that the generated reasoning traces show high clinical factuality and provide causal support for the final predictions, offering a plausible explanation for the performance gains. These results suggest that explicit reasoning can improve model performance, interpretability and clinical utility in AI‑assisted CXR interpretation.
Authors:Xinyu Tian, Shu Zou, Zhaoyuan Yang, Mengqi He, Peter Tu, Jing Zhang
Abstract:
Recent studies have demonstrated that Reinforcement Learning (RL), notably Group Relative Policy Optimization (GRPO), can intrinsically elicit and enhance the reasoning capabilities of Vision‑Language Models (VLMs). However, despite the promise, the underlying mechanisms that drive the effectiveness of RL models as well as their limitations remain underexplored. In this paper, we highlight a fundamental behavioral distinction between RL and base models, where the former engages in deeper yet narrow reasoning, while base models, despite less refined along individual path, exhibit broader and more diverse thinking patterns. Through further analysis of training dynamics, we show that GRPO is prone to diversity collapse, causing models to prematurely converge to a limited subset of reasoning strategies while discarding the majority of potential alternatives, leading to local optima and poor scalability. To address this, we propose Multi‑Group Policy Optimization (MUPO), a simple yet effective approach designed to incentivize divergent thinking across multiple solutions, and demonstrate its effectiveness on established benchmarks. Project page: https://xytian1008.github.io/MUPO/
Authors:Michael Maynord, Minghui Liu, Cornelia Fermüller, Seongjin Choi, Yuxin Zeng, Shishir Dahal, Daniel M. Harrison
Abstract:
Ultra‑high field 7‑tesla (7T) MRI improves visualization of multiple sclerosis (MS) white matter lesions (WML) but differs sufficiently in contrast and artifacts from 1.5‑3T imaging ‑ suggesting that widely used automated segmentation tools may not translate directly. We analyzed 7T FLAIR scans and generated reference WML masks from Lesion Segmentation Tool (LST) outputs followed by expert manual revision. As external comparators, we applied LST‑LPA and the more recent LST‑AI ensemble, both originally developed on lower‑field data. We then trained 3D UNETR and SegFormer transformer‑based models on 7T FLAIR at multiple resolutions (0.5x0.5x0.5^3, 1.0x1.0x1.0^3, and 1.5x1.5x2.0^3) and evaluated all methods using voxel‑wise and lesion‑wise metrics from the BraTS 2023 framework. On the held‑out test set at native 0.5x0.5x0.5^3 resolution, 7T‑trained transformers achieved competitive overlap with LST‑AI while recovering additional small lesions that were missed by classical methods, at the cost of some boundary variability and occasional artifact‑related false positives. On a held‑out 7 T test set, our best transformer model (SegFormer) achieved a voxel‑wise Dice of 0.61 and lesion‑wise Dice of 0.20, improving on the classical LST‑LPA tool (Dice 0.39, lesion‑wise Dice 0.02). Performance decreased for models trained on downsampled images, underscoring the value of native 7T resolution for small‑lesion detection. By releasing our 7T‑trained models, we aim to provide a reproducible, ready‑to‑use resource for automated lesion quantification in ultra‑high field MS research (https://github.com/maynord/7T‑MS‑lesion‑segmentation).
Authors:Jiwoo Ha, Jongwoo Baek, Jinhyun So
Abstract:
Recent Large Vision‑Language Models (LVLMs) have demonstrated remarkable performance across various multimodal tasks that require understanding both visual and linguistic inputs. However, object hallucination ‑‑ the generation of nonexistent objects in answers ‑‑ remains a persistent challenge. Although several approaches such as retraining and external grounding methods have been proposed to mitigate this issue, they still suffer from high data costs or structural complexity. Training‑free methods such as Contrastive Decoding (CD) are more cost‑effective, avoiding additional training or external models, but still suffer from long‑term decay, where visual grounding weakens and language priors dominate as the generation progresses. In this paper, we propose First Logit Boosting (FLB), a simple yet effective training‑free technique designed to alleviate long‑term decay in LVLMs. FLB stores the logit of the first generated token and adds it to subsequent token predictions, effectively mitigating long‑term decay of visual information. We observe that FLB (1) sustains the visual information embedded in the first token throughout generation, and (2) suppresses hallucinated words through the stabilizing effect of the ``The'' token. Experimental results show that FLB significantly reduces object hallucination across various tasks, benchmarks, and backbone models. Notably, it causes negligible inference overhead, making it highly applicable to real‑time multimodal systems. Code is available at https://github.com/jiwooha20/FLB
Authors:Xusheng He, Canyang Wu, Jinrong Zhang, Weili Guan, Jianlong Wu, Liqiang Nie
Abstract:
This report presents our winning solution to the 5th PVUW MeViS‑Text Challenge. The track studies referring video object segmentation under motion‑centric language expressions, where the model must jointly understand appearance, temporal behavior, and object interactions. To address this problem, we build a fully training‑free pipeline that combines strong multimodal large language models with SAM3. Our method contains three stages. First, Gemini‑3.1 Pro decomposes each target event into instance‑level grounding targets, selects the frame where the target is most clearly visible, and generates a discriminative description. Second, SAM3‑agent produces a precise seed mask on the selected frame, and the official SAM3 tracker propagates the mask through the whole video. Third, a refinement stage uses Qwen3.5‑Plus and behavior‑level verification to correct ambiguous or semantically inconsistent predictions. Without task‑specific fine‑tuning, our method ranks first on the PVUW 2026 MeViS‑Text test set, achieving a Final score of 0.909064 and a J&F score of 0.7897. The code is available at https://github.com/Moujuruo/MeViSv2_Track_Solution_2026.
Authors:Daehyun Kim, Youngmin Kim, Yoon Ju Oh, Tae Hyun Kim
Abstract:
Under‑display cameras (UDCs) allow for full‑screen designs by positioning the imaging sensor underneath the display. Nonetheless, light diffraction and scattering through the various display layers result in spatially varying and complex degradations, which significantly reduce high‑frequency details. Current PSF‑based physical modeling techniques and frequency‑separation networks are effective at reconstructing low‑frequency structures and maintaining overall color consistency. However, they still face challenges in recovering fine details when dealing with complex, spatially varying degradation. To solve this problem, we propose a lightweight Uncertainty‑aware Context‑Memory Network (UCMNet), for UDC image restoration. Unlike previous methods that apply uniform restoration, UCMNet performs uncertainty‑aware adaptive processing to restore high‑frequency details in regions with varying degradations. The estimated uncertainty maps, learned through an uncertainty‑driven loss, quantify spatial uncertainty induced by diffraction and scattering, and guide the Memory Bank to retrieve region‑adaptive context from the Context Bank. This process enables effective modeling of the non‑uniform degradation characteristics inherent to UDC imaging. Leveraging this uncertainty as a prior, UCMNet achieves state‑of‑the‑art performance on multiple benchmarks with 30% fewer parameters than previous models. Project page: \hrefhttps://kdhrick2222.github.io/projects/UCMNet/https://kdhrick2222.github.io/projects/UCMNet.
Authors:Utkarsh Pratiush, Huaixun Huyan, Maryam Zahiri Azar, Esmeralda Yitamben, Allen Bourez, Sergei V Kalinin, Vasfi Burak Ozdol
Abstract:
Scanning transmission electron microscopy (STEM) has become a cornerstone instrument for semiconductor materials metrology, enabling nanoscale analysis of complex multilayer structures that define device performance. Developing effective metrology workflows for such systems requires balancing automation with flexibility; rigid pipelines are brittle to sample variability, while purely manual approaches are slow and subjective. Here, we present a tunable human‑AI‑assisted workflow framework that enables modular and adaptive analysis of STEM images for device characterization. As an illustrative example, we demonstrate a workflow for automated layer thickness and interface roughness quantification in multilayer thin films. The system integrates gradient‑based peak detection with interactive correction modules, allowing human input at the design stage while maintaining fully automated execution across samples. Implemented as a web‑based interface, it processes TEM/EMD files directly, applies noise reduction and interface tracking algorithms, and outputs statistical roughness and thickness metrics with nanometer precision. This architecture exemplifies a general approach toward adaptive, reusable metrology workflows ‑ bridging human insight and machine precision for scalable, standardized analysis in semiconductor manufacturing. The code is made available at https://github.com/utkarshp1161/thickness‑mapping‑webapp
Authors:Xinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng, Cihang Xie, Yuyin Zhou, James Matthew Rehg, Yapeng Tian
Abstract:
We introduce Omni‑MMSI, a new task that requires comprehensive social interaction understanding from raw audio, vision, and speech input. The task involves perceiving identity‑attributed social cues (e.g., who is speaking what) and reasoning about the social interaction (e.g., whom the speaker refers to). This task is essential for developing AI assistants that can perceive and respond to human interactions. Unlike prior studies that operate on oracle‑preprocessed social cues, Omni‑MMSI reflects realistic scenarios where AI assistants must perceive and reason from raw data. However, existing pipelines and multi‑modal LLMs perform poorly on Omni‑MMSI because they lack reliable identity attribution capabilities, which leads to inaccurate social interaction understanding. To address this challenge, we propose Omni‑MMSI‑R, a reference‑guided pipeline that produces identity‑attributed social cues with tools and conducts chain‑of‑thought social reasoning. To facilitate this pipeline, we construct participant‑level reference pairs and curate reasoning annotations on top of the existing datasets. Experiments demonstrate that Omni‑MMSI‑R outperforms advanced LLMs and counterparts on Omni‑MMSI. Project page: https://sampson‑lee.github.io/omni‑mmsi‑project‑page.
Authors:Nicholas Kuang, Vanessa Scalon, Ji Yu
Abstract:
The modern deep learning field is a scale‑centric one. Larger models have been shown to consistently perform better than smaller models of similar architecture. In many sub‑domains of biomedical research, however, the model scaling is bottlenecked by the amount of available training data, and the high cost associated with generating and validating additional high quality data. Despite the practical hurdle, the majority of the ongoing research still focuses on building bigger foundation models, whereas the alternative of improving the ability of small models has been under‑explored. Here we experiment with building models with 10‑30M parameters, tiny by modern standards, to perform the single‑cell segmentation task. An important design choice is the incorporation of a recursive structure into the model's forward computation graph, leading to a more parameter‑efficient architecture. We found that for the single‑cell segmentation, on multiple benchmarks, our small model, UCell, matches the performance of models 10‑20 times its size, and with a similar generalizability to unseen out‑of‑domain data. More importantly, we found that ucell can be trained from scratch using only a set of microscopy imaging data, without relying on massive pretraining on natural images, and therefore decouples the model building from any external commercial interests. Finally, we examined and confirmed the adaptability of ucell by performing a wide range of one‑shot and few‑shot fine tuning experiments on a diverse set of small datasets. Implementation is available at https://github.com/jiyuuchc/ucell
Authors:Silong Yong, Stephen Sheng, Carl Qi, Xiaojie Wang, Evan Sheehan, Anurag Shivaprasad, Yaqi Xie, Katia Sycara, Yesh Dattatreya
Abstract:
Existing robotic foundation policies are trained primarily via large‑scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long‑horizon tasks due to distribution shift and error accumulation. While reinforcement learning (RL) can finetune these models, it cannot work well across diverse tasks without manual reward engineering. We propose VLLR, a dense reward framework combining (1) an extrinsic reward from Large Language Models (LLMs) and Vision‑Language Models (VLMs) for task progress recognition, and (2) an intrinsic reward based on policy self‑certainty. VLLR uses LLMs to decompose tasks into verifiable subtasks and then VLMs to estimate progress to initialize the value function for a brief warm‑up phase, avoiding prohibitive inference cost during full training; and self‑certainty provides per‑step intrinsic guidance throughout PPO finetuning. Ablation studies reveal complementary benefits: VLM‑based value initialization primarily improves task completion efficiency, while self‑certainty primarily enhances success rates, particularly on out‑of‑distribution tasks. On the CHORES benchmark covering mobile manipulation and navigation, VLLR achieves up to 56% absolute success rate gains over the pretrained policy, up to 5% gains over state‑of‑the‑art RL finetuning methods on in‑distribution tasks, and up to 10% gains on out‑of‑distribution tasks, all without manual reward engineering. Additional visualizations can be found in https://silongyong.github.io/vllr_project_page/
Authors:Yuheng Liu, Xin Lin, Xinke Li, Baihan Yang, Chen Wang, Kalyan Sunkavalli, Yannick Hold-Geoffroy, Hao Tan, Kai Zhang, Xiaohui Xie, Zifan Shi, Yiwei Hu
Abstract:
Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthesize only limited observations of a scene, leading to issues of completeness and global consistency. We propose OmniRoam, a controllable panoramic video generation framework that exploits the rich per‑frame scene coverage and inherent long‑term spatial and temporal consistency of panoramic representation, enabling long‑horizon scene wandering. Our framework begins with a preview stage, where a trajectory‑controlled video generation model creates a quick overview of the scene from a given input image or video. Then, in the refine stage, this video is temporally extended and spatially upsampled to produce long‑range, high‑resolution videos, thus enabling high‑fidelity world wandering. To train our model, we introduce two panoramic video datasets that incorporate both synthetic and real‑world captured videos. Experiments show that our framework consistently outperforms state‑of‑the‑art methods in terms of visual quality, controllability, and long‑term scene consistency, both qualitatively and quantitatively. We further showcase several extensions of this framework, including real‑time video generation and 3D reconstruction. Code is available at https://github.com/yuhengliu02/OmniRoam.
Authors:Abdullah Thabit, Mohamed Benmahdjoub, Rafiuddin Jinabade, Hizirwan S. Salim, Marie-Lise C. van Veelen, Mark G. van Vledder, Eppo B. Wolvius, Theo van Walsum
Abstract:
Augmented reality (AR) devices with head mounted displays (HMDs) facilitate the direct superimposition of 3D preoperative imaging data onto the patient during surgery. To use an HMD‑AR device as a stand‑alone surgical navigation system, the device should be able to locate the patient and surgical instruments, align preoperative imaging data with the patient, and visualize navigation data in real time during surgery. Whereas some of the technologies required for this are known, integration in such devices is cumbersome and requires specific knowledge and expertise, hampering scientific progress in this field. This work therefore aims to present and evaluate an integrated HMD‑based AR surgical navigation framework that is adaptable to diverse surgical applications. The framework tracks 2D patterns as reference markers attached to the patient and surgical instruments. It allows for the calibration of surgical tools using pivot and reference‑based calibration techniques. It enables image‑to‑patient registration using point‑based matching and manual positioning. The integrated functionalities of the framework are evaluated on two HMD devices, the HoloLens 2 and Magic Leap 2, with two surgical use cases being evaluated in a phantom setup: AR‑guided needle insertion and rib fracture localization. The framework was able to achieve a mean tooltip calibration accuracy of 1 mm, a registration accuracy of 3 mm, and a targeting accuracy below 5 mm on the two surgical use cases. The framework presents an easy‑to‑use configurable tool for HMD‑based AR surgical navigation, which can be extended and adapted to many surgical applications. The framework is publicly available at https://github.com/abdullahthabit/SurgNavAR.
Authors:Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
Abstract:
Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes. Computer‑assisted systems such as surgical visual question answering (VQA) offer promises for education and intraoperative support. Current surgical VQA research largely focuses on static frame analysis, overlooking rich temporal semantics. Surgical video question answering is further challenged by low visual contrast, its highly knowledge‑driven nature, diverse analytical needs spanning scattered temporal windows, and the hierarchy from basic perception to high‑level intraoperative assessment. To address these challenges, we propose SurgTEMP, a multimodal LLM framework featuring (i) a query‑guided token selection module that builds hierarchical visual memory (spatial and temporal memory banks) and (ii) a Surgical Competency Progression (SCP) training scheme. Together, they enable effective modeling of variable‑length surgical videos while preserving procedure‑relevant cues and temporal coherence, and better support diverse downstream assessment tasks. To support model development, we introduce CholeVidQA‑32K, a surgical video question answering dataset comprising 32K open‑ended QA pairs and 3,855 video segments (approximately 128 h total) from laparoscopic cholecystectomy. The dataset is organized into a three‑level hierarchy ‑‑ Perception, Assessment, and Reasoning ‑‑ spanning 11 tasks from instrument/action/anatomy perception to Critical View of Safety (CVS), intraoperative difficulty, skill proficiency, and adverse event assessment. In comprehensive evaluations against state‑of‑the‑art open‑source multimodal and video LLMs (fine‑tuned and zero‑shot), SurgTEMP achieves substantial performance improvements, advancing the state of video‑based surgical VQA. The project page is available at: https://camma‑public.github.io/SurgTEMP/
Authors:Fumihiko Tsuchiya, Taiki Miyanishi, Mahiro Ukai, Nakamasa Inoue, Shuhei Kurita, Yusuke Iwasawa, Yutaka Matsuo
Abstract:
Counting in long videos remains a fundamental yet underexplored challenge in computer vision. Real‑world recordings often span tens of minutes or longer and contain sparse, diverse events, making long‑range temporal reasoning particularly difficult. However, most existing video counting benchmarks focus on short clips and evaluate only the final numerical answer, providing little insight into what should be counted or whether models consistently identify relevant instances across time. We introduce EC‑Bench, a benchmark that jointly evaluates enumeration, counting, and temporal evidence grounding in long‑form videos. EC‑Bench contains 152 videos longer than 30 minutes and 1,699 queries paired with explicit evidence spans. Across 22 multimodal large language models (MLLMs), the best model achieves only 29.98% accuracy on Enumeration and 23.74% on Counting, while human performance reaches 78.57% and 82.97%, respectively. Our analysis reveals strong relationships between enumeration accuracy, temporal grounding, and counting performance. These results highlight fundamental limitations of current MLLMs and establish EC‑Bench as a challenging benchmark for long‑form quantitative video reasoning.
Authors:Yuhang Yang, Fan Zhang, Huaijin Pi, Shuai Guo, Guowei Xu, Wei Zhai, Yang Cao, Zheng-Jun Zha
Abstract:
Digital characters are central to modern media, yet generating character videos with long‑duration, consistent multi‑view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non‑character‑centric information as the memory, leading to suboptimal consistency. Recognizing that character video generation inherently resembles an outside‑looking‑in scenario. In this work, we propose representing the character visual attributes through a compact set of anchor frames. This design provides stable references for consistency, while reference‑based video generation inherently faces challenges of copy‑pasting and multi‑reference conflicts. To address these, we introduce two mechanisms: Superset Content Anchoring, providing intra‑ and extra‑training clip cues to prevent duplication, and RoPE as Weak Condition, encoding positional offsets to distinguish multiple anchors. Furthermore, we construct a scalable pipeline to extract these anchors from massive videos. Experiments show our method generates high‑quality character videos exceeding 10 minutes, and achieves expressive identity and appearance consistency across views, surpassing existing methods.
Authors:Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, Xihui Liu
Abstract:
The development of Vision‑Language‑Action (VLA) models has been significantly accelerated by pre‑trained Vision‑Language Models (VLMs). However, most existing end‑to‑end VLAs treat the VLM primarily as a multimodal encoder, directly mapping vision‑language features to low‑level actions. This paradigm underutilizes the VLM's potential in high‑level decision making and introduces training instability, frequently degrading its rich semantic representations. To address these limitations, we introduce DIAL, a framework bridging high‑level decision making and low‑level motor execution through a differentiable latent intent bottleneck. Specifically, a VLM‑based System‑2 performs latent world modeling by synthesizing latent visual foresight within the VLM's native feature space; this foresight explicitly encodes intent and serves as the structural bottleneck. A lightweight System‑1 policy then decodes this predicted intent together with the current observation into precise robot actions via latent inverse dynamics. To ensure optimization stability, we employ a two‑stage training paradigm: a decoupled warmup phase where System‑2 learns to predict latent futures while System‑1 learns motor control under ground‑truth future guidance within a unified feature space, followed by seamless end‑to‑end joint optimization. This enables action‑aware gradients to refine the VLM backbone in a controlled manner, preserving pre‑trained knowledge. Extensive experiments on the RoboCasa GR1 Tabletop benchmark show that DIAL establishes a new state‑of‑the‑art, achieving superior performance with 10x fewer demonstrations than prior methods. Furthermore, by leveraging heterogeneous human demonstrations, DIAL learns physically grounded manipulation priors and exhibits robust zero‑shot generalization to unseen objects and novel configurations during real‑world deployment on a humanoid robot.
Authors:Rosario Leonardi, Antonino Furnari, Francesco Ragusa, Giovanni Maria Farinella
Abstract:
In this work, we explore the role of synthetic data in improving the detection of Hand‑Object Interactions from egocentric images. Through extensive experimentation and comparative analysis on VISOR, EgoHOS, and ENIGMA‑51 datasets, our findings demonstrate the potential of synthetic data to significantly improve HOI detection, particularly when real labeled data are scarce or unavailable. By using synthetic data and only 10% of the real labeled data, we achieve improvements in Overall AP over models trained exclusively on real data, with gains of +5.67% on VISOR, +8.24% on EgoHOS, and +11.69% on ENIGMA‑51. Furthermore, we systematically study how aligning synthetic data to specific real‑world benchmarks with respect to objects, grasps, and environments, showing that the effectiveness of synthetic data consistently improves with better synthetic‑real alignment. As a result of this work, we release a new data generation pipeline and the new HOI‑Synth benchmark, which augments existing datasets with synthetic images of hand‑object interaction. These data are automatically annotated with hand‑object contact states, bounding boxes, and pixel‑wise segmentation masks. All data, code, and tools for synthetic data generation are available at: https://fpv‑iplab.github.io/HOI‑Synth/.
Authors:Lixin Xiu, Xufang Luo, Hideki Nakayama
Abstract:
Large vision‑language models (LVLMs) achieve impressive performance, yet their internal decision‑making processes remain opaque, making it difficult to determine if the success stems from true multimodal fusion or from reliance on unimodal priors. To address this attribution gap, we introduce a novel framework using partial information decomposition (PID) to quantitatively measure the "information spectrum" of LVLMs ‑‑ decomposing a model's decision‑relevant information into redundant, unique, and synergistic components. By adapting a scalable estimator to modern LVLM outputs, our model‑agnostic pipeline profiles 26 LVLMs on four datasets across three dimensions ‑‑ breadth (cross‑model & cross‑task), depth (layer‑wise information dynamics), and time (learning dynamics across training). Our analysis reveals two key results: (i) two task regimes (synergy‑driven vs. knowledge‑driven) and (ii) two stable, contrasting family‑level strategies (fusion‑centric vs. language‑centric). We also uncover a consistent three‑phase pattern in layer‑wise processing and identify visual instruction tuning as the key stage where fusion is learned. Together, these contributions provide a quantitative lens beyond accuracy‑only evaluation and offer insights for analyzing and designing the next generation of LVLMs. Code and data are available at https://github.com/RiiShin/pid‑lvlm‑analysis .
Authors:Dimitrios Anastasiou, Razvan Caramalau, Jialang Xu, Runlong He, Freweini Tesfai, Matthew Boal, Nader Francis, Danail Stoyanov, Evangelos B. Mazomenos
Abstract:
Vision‑based surgical skill assessment (SSA) enables objective and scalable evaluation of operative performance. Progress in this field is constrained by the high cost and time demands for manual annotation of quantitative skill scores, as well as the poor generalization of existing regression models to new surgical tasks and environments. Meanwhile, appreciable volumes of unlabeled video data are now available, motivating the development of unsupervised domain adaptation (UDA) methods for SSA. We introduce the first benchmark for UDA in SSA regression, spanning four datasets across dry‑lab and clinical settings as well as open and robotic surgery. We evaluate eight representative models under challenging domain shifts and propose CoRe‑DA, a novel contrastive regression‑based adaptation framework. Our method learns domain‑invariant representations through relative‑score supervision and target‑domain self‑training. Comprehensive experiments across two UDA settings show that CoRe‑DA is superior to state‑of‑the‑art methods, achieving Spearman Correlation Coefficients of 0.46 and 0.41 on dry‑lab and clinical target datasets, respectively, without using any labeled target data for training. Overall, CoRe‑DA enables scalable SSA with reliable cross‑domain generalization, where existing methods underperform. Our code and datasets will be released at https://github.com/anastadimi/CoRe‑DA.
Authors:Shifang Zhao, Yihan Hu, Ying Shan, Yunchao Wei, Xiaodong Cun
Abstract:
Editing the video content with audio alignment forms a digital human‑made art in current social media. However, the time‑consuming and repetitive nature of manual video editing has long been a challenge for filmmakers and professional content creators alike. In this paper, we introduce CutClaw, an autonomous multi‑agent framework designed to edit hours‑long raw footage into meaningful short videos that leverages the capabilities of multiple Multimodal Language Models~(MLLMs) as an agent system. It produces videos with synchronized music, followed by instructions, and a visually appealing appearance. In detail, our approach begins by employing a hierarchical multimodal decomposition that captures both fine‑grained details and global structures across visual and audio footage. Then, to ensure narrative consistency, a Playwriter Agent orchestrates the whole storytelling flow and structures the long‑term narrative, anchoring visual scenes to musical shifts. Finally, to construct a short edited video, Editor and Reviewer Agents collaboratively optimize the final cut via selecting fine‑grained visual content based on rigorous aesthetic and semantic criteria. We conduct detailed experiments to demonstrate that CutClaw significantly outperforms state‑of‑the‑art baselines in generating high‑quality, rhythm‑aligned videos. The code is available at: https://github.com/GVCLab/CutClaw.
Authors:Pengfei Zhou, Xiangyue Zhang, Xukun Shen, Yong Hu
Abstract:
Masked generative models have become a strong paradigm for text‑to‑motion synthesis, but they still treat motion frames too uniformly during masking, attention, and decoding. This is a poor match for motion, where local dynamic complexity varies sharply over time. We show that current masked motion generators degrade disproportionately on dynamically complex motions, and that frame‑wise generation error is strongly correlated with motion dynamics. Motivated by this mismatch, we introduce the Motion Spectral Descriptor (MSD), a simple and parameter‑free measure of local dynamic complexity computed from the short‑time spectrum of motion velocity. Unlike learned difficulty predictors, MSD is deterministic, interpretable, and derived directly from the motion signal itself. We use MSD to make masked motion generation complexity‑aware. In particular, MSD guides content‑focused masking during training, provides a spectral similarity prior for self‑attention, and can additionally modulate token‑level sampling during iterative decoding. Built on top of masked motion generators, our method, DynMask, improves motion generation most clearly on dynamically complex motions while also yielding stronger overall FID on HumanML3D and KIT‑ML. These results suggest that respecting local motion complexity is a useful design principle for masked motion generation. Project page: https://xiangyue‑zhang.github.io/DynMask
Authors:Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou, Kaituo Feng, Wenbo Hu, Yi-Fan Zhang, Yunlong Lin, Wenxuan Huang, Mingyang Song, Dasen Dai, Bolin Jiang, Manyuan Zhang, Shi-Xue Zhang, Zhengkai Jiang, Lucas Wang, Zhao Zhong, Yu Cheng, Nanyun Peng
Abstract:
Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real‑world knowledge while generating high‑quality images. However, they still rely primarily on frozen parametric knowledge, which makes them struggle with real‑world image generation involving long‑tail and knowledge‑intensive concepts. Inspired by the broad success of agents on real‑world tasks, we explore agentic modeling to address this limitation. Specifically, we present Unify‑Agent, a unified multimodal agent for world‑grounded image synthesis, which reframes image generation as an agentic pipeline consisting of prompt understanding, multimodal evidence searching, grounded recaptioning, and final synthesis. To train our model, we construct a tailored multimodal data pipeline and curate 143K high‑quality agent trajectories for world‑grounded image synthesis, enabling effective supervision over the full agentic generation process. We further introduce FactIP, a benchmark covering 12 categories of culturally significant and long‑tail factual concepts that explicitly requires external knowledge grounding. Extensive experiments show that our proposed Unify‑Agent substantially improves over its base unified model across diverse benchmarks and real world generation tasks, while approaching the world knowledge capabilities of the strongest closed‑source models. As an early exploration of agent‑based modeling for world‑grounded image synthesis, our work highlights the value of tightly coupling reasoning, searching, and generation for reliable open‑world agentic image synthesis.
Authors:Geuntaek Lim, Minho Shim, Sungjune Park, Jaeyun Lee, Inwoong Lee, Taeoh Kim, Dongyoon Wee, Yukyung Choi
Abstract:
The inherent complexity of video understanding makes it difficult to attribute whether performance gains stem from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high‑level reasoning, the essential criteria that constitute video understanding remain largely overlooked. Instead of introducing yet another benchmark, we take a step back to re‑examine the current landscape of video understanding. In this work, we provide Video‑Oasis, a sustainable diagnostic suite designed to systematically evaluate existing evaluations and distill spatio‑temporal challenges for video understanding. Our analysis reveals two critical findings: (1) 54% of existing benchmark samples are solvable without visual input or temporal context, and (2) on the remaining samples, state‑of‑the‑art models exhibit performance barely exceeding random guessing. To bridge this gap, we investigate which algorithmic design choices contribute to robust video understanding, providing practical guidelines for future research. We hope our work serves as a standard guideline for benchmark construction and the rigorous evaluation of architecture development. Code is available at https://github.com/sejong‑rcv/Video‑Oasis.
Authors:Fei Shen, Chengyu Xie, Lihong Wang, Zhanyi Zhang, Xin Jiang, Xiaoyu Du, Jinhui Tang
Abstract:
Existing multi‑turn image editing paradigms are often confined to isolated single‑step execution. Due to a lack of context‑awareness and closed‑loop feedback mechanisms, they are prone to error accumulation and semantic drift during multi‑turn interactions, ultimately resulting in severe structural distortion of the generated images. For that, we propose IMAGAgent, a multi‑turn image editing agent framework based on a "plan‑execute‑reflect" closed‑loop mechanism that achieves deep synergy among instruction parsing, tool scheduling, and adaptive correction within a unified pipeline. Specifically, we first present a constraint‑aware planning module that leverages a vision‑language model (VLM) to precisely decompose complex natural language instructions into a series of executable sub‑tasks, governed by target singularity, semantic atomicity, and visual perceptibility. Then, the tool‑chain orchestration module dynamically constructs execution paths based on the current image, the current sub‑task, and the historical context, enabling adaptive scheduling and collaborative operation among heterogeneous operation models covering image retrieval, segmentation, detection, and editing. Finally, we devise a multi‑expert collaborative reflection mechanism where a central large language model (LLM) receives the image to be edited and synthesizes VLM critiques into holistic feedback, simultaneously triggering fine‑grained self‑correction and recording feedback outcomes to optimize future decisions. Extensive experiments on our constructed MTEditBench and the MagicBrush dataset demonstrate that IMAGAgent achieves performance significantly superior to existing methods in terms of instruction consistency, editing precision, and overall quality. The code is available at https://github.com/hackermmzz/IMAGAgent.git.
Authors:Jagadish Kashinath Kamble, Jayanta Mukhopadhyay, Debaditya Roy, Partha Pratim Das
Abstract:
Preserving intangible cultural dances rooted in centuries of tradition and governed by strict structural and symbolic rules presents unique challenges in the digital era. Among these, Bharatanatyam, a classical Indian dance form, stands out for its emphasis on codified adavus and precise key postures. Accurately generating these postures is crucial not only for maintaining anatomical and stylistic integrity, but also for enabling effective documentation, analysis, and transmission to broader global audiences through digital means. We propose a pose‑aware generative framework integrated with a pose estimation module, guided by keypoint‑based loss and pose consistency constraints. These supervisory signals ensure anatomical accuracy and stylistic integrity in the synthesized outputs. We evaluate four configurations: standard conditional generative adversarial network (cGAN), cGAN with pose supervision, conditional diffusion, and conditional diffusion with pose supervision. Each model is conditioned on key posture class labels and optimized to maintain geometric structure. In both cGAN and conditional diffusion settings, the integrated pose guidance aligns generated poses with ground‑truth keypoint structures, promoting cultural fidelity. Our results demonstrate that incorporating pose supervision significantly enhances the quality, realism, and authenticity of generated Bharatanatyam postures. This framework provides a scalable approach for the digital preservation, education, and dissemination of traditional dance forms, enabling high‑fidelity generation without compromising cultural precision. Code is available at https://github.com/jagidsh/Generating‑Key‑Postures‑of‑Bharatanatyam‑Adavus‑with‑Pose‑Estimation.
Authors:Anmin Liu, Ruixuan Yang, Huiqiang Jiang, Bin Lin, Minmin Sun, Yong Li, Chen Zhang, Tao Xie
Abstract:
Long‑context video understanding and generation pose a significant computational challenge for Transformer‑based video models due to the quadratic complexity of self‑attention. While existing sparse attention methods employ coarse‑grained patterns to improve efficiency, they typically incur redundant computation and suboptimal performance. To address this issue, in this paper, we propose VecAttention, a novel framework of vector‑wise sparse attention that achieves superior accuracy‑efficiency trade‑offs for video models. We observe that video attention maps exhibit a strong vertical‑vector sparse pattern, and further demonstrate that this vertical‑vector pattern offers consistently better accuracy‑sparsity trade‑offs compared with existing coarse‑grained sparse patterns. Based on this observation, VecAttention dynamically selects and processes only informative vertical vectors through a lightweight important‑vector selection that minimizes memory access overhead and an optimized kernel of vector sparse attention. Comprehensive evaluations on video understanding (VideoMME, LongVideoBench, and VCRBench) and generation (VBench) tasks show that VecAttention delivers a 2.65× speedup over full attention and a 1.83× speedup over state‑of‑the‑art sparse attention methods, with comparable accuracy to full attention. Our code is available at https://github.com/anminliu/VecAttention.
Authors:Antoine Bottenmuller, Etienne Decencière, Petr Dokládal
Abstract:
Semantic segmentation and hyperspectral unmixing are two central problems in spectral image analysis. The former assigns each pixel a discrete label corresponding to its material class, whereas the latter estimates pure material spectra, called endmembers, and, for each pixel, a vector representing material abundances in the observed scene. Despite their complementarity, these two problems are usually addressed independently. This paper aims to bridge these two lines of work by formally showing that, under the linear mixing model, pixel classification by dominant materials induces polyhedral‑cone regions in the spectral space. We leverage this fundamental property to propose a direct segmentation‑to‑unmixing pipeline that performs blind hyperspectral unmixing from any semantic segmentation by constructing a polyhedral‑cone partition of the space that best fits the labeled pixels. Signed distances from pixels to the estimated regions are then computed, linearly transformed via a change of basis in the distance space, and projected onto the probability simplex, yielding an initial abundance estimate. This estimate is used to extract endmembers and recover final abundances via matrix pseudo‑inversion. Because the segmentation method can be freely chosen, the user gains explicit control over the unmixing process, while the rest of the pipeline remains essentially deterministic and lightweight. Beyond improving interpretability, experiments on three real datasets demonstrate the effectiveness of the proposed approach when associated with appropriate clustering algorithms, and show consistent improvements over recent deep and non‑deep state‑of‑the‑art methods. The code is available at: https://github.com/antoine‑bottenmuller/polyhedral‑unmixing
Authors:Xuesong Wang, Harry Wang
Abstract:
Vision‑language models (VLMs) exhibit a systematic bias when confronted with classic optical illusions: they overwhelmingly predict the illusion as "real" regardless of whether the image has been counterfactually modified. We present a tool‑guided inference framework for the DataCV 2026 Challenge (Tasks I and II) that addresses this failure mode without any model training. An off‑the‑shelf vision‑language model is given access to a small set of generic image manipulation tools: line drawing, region cropping, side‑by‑side comparison, and channel isolation, together with an illusion‑type‑routing system prompt that prescribes which tools to invoke for each perceptual question category. Critically, every tool call produces a new, immutable image resource appended to a persistent registry, so the model can reference and compose any prior annotated view throughout its reasoning chain. Rather than hard‑coding illusion‑specific modules, this generic‑tool‑plus‑routing design yields strong cross‑structural generalization: performance remained consistent from the validation set to a test set containing structurally unfamiliar illusion variants (e.g., Mach Bands rotated from vertical to horizontal stacking). We further report three empirical observations that we believe warrant additional investigation: (i) a strong positive‑detection bias likely rooted in imbalanced illusion training data, (ii) a striking dissociation between pixel‑accurate spatial reasoning and logical inference over self‑generated annotations, and (iii) pronounced sensitivity to image compression artifacts that compounds false positives.
Authors:Qiyuan Zhuang, He-Yang Xu, Yijun Wang, Xin-Yang Zhao, Yang-Yang Li, Xiu-Shen Wei
Abstract:
Understanding object affordances is essential for enabling robots to perform purposeful and fine‑grained interactions in diverse and unstructured environments. However, existing approaches either rely on retrieval, which is fragile due to sparsity and coverage gaps, or on large‑scale models, which frequently mislocalize contact points and mispredict post‑contact actions when applied to unseen categories, thereby hindering robust generalization. We introduce Retrieval‑Augmented Affordance Prediction (RAAP), a framework that unifies affordance retrieval with alignment‑based learning. By decoupling static contact localization and dynamic action direction, RAAP transfers contact points via dense correspondence and predicts action directions through a retrieval‑augmented alignment model that consolidates multiple references with dual‑weighted attention. Trained on compact subsets of DROID and HOI4D with as few as tens of samples per task, RAAP achieves consistent performance across unseen objects and categories, and enables zero‑shot robotic manipulation in both simulation and the real world. Project website: https://github.com/SEU‑VIPGroup/RAAP.
Authors:Ni Ou, Zhuo Chen, Xinru Zhang, Junzheng Wang
Abstract:
Accurate camera‑LiDAR fusion relies on precise extrinsic calibration, which fundamentally depends on establishing reliable cross‑modal correspondences under potentially large misalignments. Existing learning‑based methods typically project LiDAR points into depth maps for feature fusion, which distorts 3D geometry and degrades performance when the extrinsic initialization is far from the ground truth. To address this issue, we propose an extrinsic‑aware cross‑attention framework that directly aligns image patches and LiDAR point groups in their native domains. The proposed attention mechanism explicitly injects extrinsic parameter hypotheses into the correspondence modeling process, enabling geometry‑consistent cross‑modal interaction without relying on projected 2D depth maps. Extensive experiments on the KITTI and nuScenes benchmarks demonstrate that our method consistently outperforms state‑of‑the‑art approaches in both accuracy and robustness. Under large extrinsic perturbations, our approach achieves accurate calibration in 88% of KITTI cases and 99% of nuScenes cases, substantially surpassing the second‑best baseline. We have open sourced our code on https://github.com/gitouni/ProjFusion to benefit the community.
Authors:Yubo Cui, Xianchao Guan, Zijun Xiong, Zheng Zhang
Abstract:
Pre‑trained vision‑language models (VLMs) exhibit strong zero‑shot generalization but remain vulnerable to adversarial perturbations. Existing classification‑guided adversarial fine‑tuning methods often disrupt pre‑trained cross‑modal alignment, weakening visual‑textual correspondence and degrading zero‑shot performance. In this paper, we propose an Alignment‑Guided Fine‑Tuning (AGFT) framework that enhances zero‑shot adversarial robustness while preserving the cross‑modal semantic structure. Unlike label‑based methods that rely on hard labels and fail to maintain the relative relationships between image and text, AGFT leverages the probabilistic predictions of the original model for text‑guided adversarial training, which aligns adversarial visual features with textual embeddings via soft alignment distributions, improving zero‑shot adversarial robustness. To address structural discrepancies introduced by fine‑tuning, we introduce a distribution consistency calibration mechanism that adjusts the robust model output to match a temperature‑scaled version of the pre‑trained model predictions. Extensive experiments across multiple zero‑shot benchmarks demonstrate that AGFT outperforms state‑of‑the‑art methods while significantly improving zero‑shot adversarial robustness.
Authors:Wei Suo, Hanzu Zhang, Lijun Zhang, Ji Ma, Peng Wang, Yanning Zhang
Abstract:
Large Vision‑Language Models have demonstrated exceptional performance in multimodal reasoning and complex scene understanding. However, these models still face significant hallucination issues, where outputs contradict visual facts. Recent research on hallucination mitigation has focused on retraining methods and Contrastive Decoding (CD) methods. While both methods perform well, retraining methods require substantial training resources, and CD methods introduce dual inference overhead. These factors hinder their practical applicability. To address the above issue, we propose a framework for dynamically detecting hallucination representations and performing hallucination‑eliminating edits on these representations. With minimal additional computational cost, we achieve state‑of‑the‑art performance on existing benchmarks. Extensive experiments demonstrate the effectiveness of our approach, highlighting its efficient and robust hallucination elimination capability and its powerful controllability over hallucinations. Code is available at https://github.com/ASGO‑MM/HIRE
Authors:Taewoo Suh, Sungpyo Kim, Jongmin Park, Munchurl Kim
Abstract:
Feed‑forward 3D Gaussian Splatting (FF‑3DGS) emerges as a fast and robust solution for sparse‑view 3D reconstruction and novel view synthesis (NVS). However, existing FF‑3DGS methods are built on incorrect screen‑space dilation filters, causing severe rendering artifacts when rendering at out‑of‑distribution sampling rates. We firstly propose an FF‑3DGS model, called AA‑Splat, to enable robust anti‑aliased rendering at any resolution. AA‑Splat utilizes an opacity‑balanced band‑limiting (OBBL) design, which combines two components: a 3D band‑limiting post‑filter integrates multi‑view maximal frequency bounds into the feed‑forward reconstruction pipeline, effectively band‑limiting the resulting 3D scene representations and eliminating degenerate Gaussians; an Opacity Balancing (OB) to seamlessly integrate all pixel‑aligned Gaussian primitives into the rendering process, compensating for the increased overlap between expanded Gaussian primitives. AA‑Splat demonstrates drastic improvements with average 5.4~7.5dB PSNR gains on NVS performance over a state‑of‑the‑art (SOTA) baseline, DepthSplat, at all resolutions, between 4× and 1/4×. Code will be made available.
Authors:Seungwoo Yoon, Jinmo Kim, Jaesik Park
Abstract:
In this paper, we propose Extend3D, a training‑free pipeline for 3D scene generation from a single image, built upon an object‑centric 3D generative model. To overcome the limitations of fixed‑size latent spaces in object‑centric models for representing wide scenes, we extend the latent space in the x and y directions. Then, by dividing the extended latent space into overlapping patches, we apply the object‑centric 3D generative model to each patch and couple them at each time step. Since patch‑wise 3D generation with image conditioning requires strict spatial alignment between image and latent patches, we initialize the scene using a point cloud prior from a monocular depth estimator and iteratively refine occluded regions through SDEdit. We discovered that treating the incompleteness of 3D structure as noise during 3D refinement enables 3D completion via a concept, which we term under‑noising. Furthermore, to address the sub‑optimality of object‑centric models for sub‑scene generation, we optimize the extended latent during denoising, ensuring that the denoising trajectories remain consistent with the sub‑scene dynamics. To this end, we introduce 3D‑aware optimization objectives for improved geometric structure and texture fidelity. We demonstrate that our method yields better results than prior methods, as evidenced by human preference and quantitative experiments.
Authors:Jintao Sun, Hu Zhang, Gangyi Ding, Zhedong Zheng
Abstract:
Trajectory prediction seeks to forecast the future motion of dynamic entities, such as vehicles and pedestrians, given a temporal horizon of historical movement data and environmental context. A central challenge in this domain is the inherent uncertainty in real‑time maps, arising from two primary sources: (1) positional inaccuracies due to sensor limitations or environmental occlusions, and (2) semantic errors stemming from misinterpretations of scene context. To address these challenges, we propose a novel unified framework that jointly models positional and semantic uncertainties and explicitly integrates them into the trajectory prediction pipeline. Our approach employs a dual‑head architecture to independently estimate semantic and positional predictions in a dual‑pass manner, deriving prediction variances as uncertainty indicators in an end‑to‑end fashion. These uncertainties are subsequently fused with the semantic and positional predictions to enhance the robustness of trajectory forecasts. We evaluate our uncertainty‑aware framework on the nuScenes real‑world driving dataset, conducting extensive experiments across four map estimation methods and two trajectory prediction baselines. Results verify that our method (1) effectively quantifies map uncertainties through both positional and semantic dimensions, and (2) consistently improves the performance of existing trajectory prediction models across multiple metrics, including minimum Average Displacement Error (minADE), minimum Final Displacement Error (minFDE), and Miss Rate (MR). Code will available at https://github.com/JT‑Sun/UATP.
Authors:Haoran Zhou, Gim Hee Lee
Abstract:
Realistic reconstruction of dynamic 4D scenes from monocular videos is essential for understanding the physical world. Despite recent progress in neural rendering, existing methods often struggle to recover accurate 3D geometry and temporally consistent motion in complex environments. To address these challenges, we propose MotionScale, a 4D Gaussian Splatting framework that scales efficiently to large scenes and extended sequences while maintaining high‑fidelity structural and motion coherence. At the core of our approach is a scalable motion field parameterized by cluster‑centric basis transformations that adaptively expand to capture diverse and evolving motion patterns. To ensure robust reconstruction over long durations, we introduce a progressive optimization strategy comprising two decoupled propagation stages: 1) A background extension stage that adapts to newly visible regions, refines camera poses, and explicitly models transient shadows; 2) A foreground propagation stage that enforces motion consistency through a specialized three‑stage refinement process. Extensive experiments on challenging real‑world benchmarks demonstrate that MotionScale significantly outperforms state‑of‑the‑art methods in both reconstruction quality and temporal stability. Project page: https://hrzhou2.github.io/motion‑scale‑web/.
Authors:Guozhi Qiu, Zhiwei Chen, Zixu Li, Qinlei Huang, Zhiheng Fu, Xuemeng Song, Yupeng Hu
Abstract:
Composed Image Retrieval (CIR) uses a reference image and a modification text as a query to retrieve a target image satisfying the requirement of ``modifying the reference image according to the text instructions''. However, existing CIR methods face two limitations: (1) frequency bias leading to ``Rare Sample Neglect'', and (2) susceptibility of similarity scores to interference from hard negative samples and noise. To address these limitations, we confront two key challenges: asymmetric rare semantic localization and robust similarity estimation under hard negative samples. To solve these challenges, we propose the Modification frEquentation‑rarity baLance neTwork MELT. MELT assigns increased attention to rare modification semantics in multimodal contexts while applying diffusion‑based denoising to hard negative samples with high similarity scores, enhancing multimodal fusion and matching. Extensive experiments on two CIR benchmarks validate the superior performance of MELT. Codes are available at https://github.com/luckylittlezhi/MELT.
Authors:Wenyang Chen, Zhanxuan Hu, Yaping Zhang, Hailong Ning, Yonghang Tai
Abstract:
Training‑free open‑vocabulary remote sensing segmentation (OVRSS), empowered by vision‑language models, has emerged as a promising paradigm for achieving category‑agnostic semantic understanding in remote sensing imagery. Existing approaches mainly focus on enhancing feature representations or mitigating modality discrepancies to improve patch‑level prediction accuracy. However, such independent prediction schemes are fundamentally misaligned with the intrinsic characteristics of remote sensing data. In real‑world applications, remote sensing scenes are typically large‑scale and exhibit strong spatial as well as semantic correlations, making isolated patch‑wise predictions insufficient for accurate segmentation. To address this limitation, we propose ConInfer, a context‑aware inference framework for OVRSS that performs joint prediction across multiple spatial units while explicitly modeling their inter‑unit semantic dependencies. By incorporating global contextual cues, our method significantly enhances segmentation consistency, robustness, and generalization in complex remote sensing environments. Extensive experiments on multiple benchmark datasets demonstrate that our approach consistently surpasses state‑of‑the‑art per‑pixel VLM‑based baselines such as SegEarth‑OV, achieving average improvements of 2.80% and 6.13% on open‑vocabulary semantic segmentation and object extraction tasks, respectively. The implementation code is available at: https://github.com/Dog‑Yang/ConInfer
Authors:Wenchao Sun, Xuewu Lin, Keyu Chen, Zixiang Pei, Xiang Li, Yining Shi, Sifa Zheng
Abstract:
End‑to‑end multi‑modal planning has been widely adopted to model the uncertainty of driving behavior, typically by scoring candidate trajectories and selecting the optimal one. Existing approaches generally fall into two categories: scoring a large static trajectory vocabulary, or scoring a small set of dynamically generated proposals. While static vocabularies often suffer from coarse discretization of the action space, dynamic proposals provide finer‑grained precision and have shown stronger empirical performance on existing benchmarks. However, it remains unclear whether dynamic generation is fundamentally necessary, or whether static vocabularies can already achieve comparable performance when they are sufficiently dense to cover the action space. In this work, we start with a systematic scaling study of Hydra‑MDP, a representative scoring‑based method, revealing that performance consistently improves as trajectory anchors become denser, without exhibiting saturation before computational constraints are reached. Motivated by this observation, we propose SparseDriveV2 to push the performance boundary of scoring‑based planning through two complementary innovations: (1) a scalable vocabulary representation with a factorized structure that decomposes trajectories into geometric paths and velocity profiles, enabling combinatorial coverage of the action space, and (2) a scalable scoring strategy with coarse factorized scoring over paths and velocity profiles followed by fine‑grained scoring on a small set of composed trajectories. By combining these two techniques, SparseDriveV2 achieves 92.0 PDMS and 90.1 EPDMS on NAVSIM, with 89.15 Driving Score and 70.00 Success Rate on Bench2Drive with a lightweight ResNet‑34 as backbone. Code and model are released at https://github.com/swc‑17/SparseDriveV2.
Authors:Xiaoyan Zhang, Jiangpeng He
Abstract:
Visual food recognition in real‑world dietary logging scenarios naturally exhibits severe data imbalance, where a small number of food categories appear frequently while many others occur rarely, resulting in long‑tailed class distributions. In practice, food recognition systems often operate in a continual learning setting, where new categories are introduced sequentially over time. However, existing studies typically assume that each incremental step introduces a similar number of new food classes, which rarely happens in real world where the number of newly observed categories can vary significantly across steps, leading to highly uneven learning dynamics. As a result, continual food recognition exhibits a dual imbalance: imbalanced samples within each food class and imbalanced numbers of new food classes to learn at each incremental learning step. In this work, we introduce DIME, a Dual‑Imbalance‑aware Adapter Merging framework for continual food recognition. DIME learns lightweight adapters for each task using parameter‑efficient fine‑tuning and progressively integrates them through a class‑count guided spectral merging strategy. A rank‑wise threshold modulation mechanism further stabilizes the merging process by preserving dominant knowledge while allowing adaptive updates. The resulting model maintains a single merged adapter for inference, enabling efficient deployment without accumulating task‑specific modules. Experiments on realistic long‑tailed food benchmarks under our step‑imbalanced setup show that the proposed method consistently improves by more than 3% over the strongest existing continual learning baselines. Code is available at https://github.com/xiaoyanzhang1/DIME.
Authors:Kiran Chhatre, Hyeonho Jeong, Yulia Gryaditskaya, Christopher E. Peters, Chun-Hao Paul Huang, Paul Guerrero
Abstract:
Generative video editing has enabled several intuitive editing operations for short video clips that would previously have been difficult to achieve, especially for non‑expert editors. Existing methods focus on prescribing an object's 3D or 2D motion trajectory in a video, or on altering the appearance of an object or a scene, while preserving both the video's plausibility and identity. Yet a method to move an object's 3D motion trajectory in a video, i.e., moving an object while preserving its relative 3D motion, is currently still missing. The main challenge lies in obtaining paired video data for this scenario. Previous methods typically rely on clever data generation approaches to construct plausible paired data from unpaired videos, but this approach fails if one of the videos in a pair can not easily be constructed from the other. Instead, we introduce TrajectoryAtlas, a new data generation pipeline for large‑scale synthetic paired video data and a video generator TrajectoryMover fine‑tuned with this data. We show that this successfully enables generative movement of object trajectories. Project page: https://chhatrekiran.github.io/trajectorymover
Authors:Jaber Jaber, Osama Jaber
Abstract:
World models that predict future states from video remain limited by flat latent representations that entangle objects, ignore causal structure, and collapse temporal dynamics into a single scale. We present HCLSM, a world model architecture that operates on three interconnected principles: object‑centric decomposition via slot attention with spatial broadcast decoding, hierarchical temporal dynamics through a three‑level engine combining selective state space models for continuous physics, sparse transformers for discrete events, and compressed transformers for abstract goals, and causal structure learning through graph neural network interaction patterns. HCLSM introduces a two‑stage training protocol where spatial reconstruction forces slot specialization before dynamics prediction begins. We train a 68M‑parameter model on the PushT robotic manipulation benchmark from the Open X‑Embodiment dataset, achieving 0.008 MSE next‑state prediction loss with emerging spatial decomposition (SBD loss: 0.0075) and learned event boundaries. A custom Triton kernel for the SSM scan delivers 38x speedup over sequential PyTorch. The full system spans 8,478 lines of Python across 51 modules with 171 unit tests. Code: https://github.com/rightnow‑ai/hclsm
Authors:Kushal Vyas, Alper Kayabasi, Daniel Kim, Vishwanath Saragadam, Ashok Veeraraghavan, Guha Balakrishnan
Abstract:
The approximation and convergence properties of implicit neural representations (INRs) are known to be highly sensitive to parameter initialization strategies. While several data‑driven initialization methods demonstrate significant improvements over standard random sampling, the reasons for their success ‑‑ specifically, whether they encode classical statistical signal priors or more complex features ‑‑ remain poorly understood. In this study, we explore this phenomenon through a series of experimental analyses leveraging noise pretraining. We pretrain INRs on diverse noise classes (e.g., Gaussian, Dead Leaves, Spectral) and measure their ability to both fit unseen signals and encode priors for an inverse imaging task (denoising). Our analyses on image and video data reveal a surprising finding: simply pretraining on unstructured noise (Uniform, Gaussian) dramatically improves signal fitting capacity compared to all other baselines. However, unstructured noise also yields poor deep image priors for denoising. In contrast, we also find that noise with the classic 1/|f^α| spectral structure of natural images achieves an excellent balance of signal fitting and inverse imaging capabilities, performing on par with the best data‑driven initialization methods. This finding enables more efficient INR training in applications lacking sufficient prior domain‑specific data. For more details, visit project page at https://kushalvyas.github.io/noisepretraining.html
Authors:Bharath Krishnamurthy, Ajita Rattani
Abstract:
Recent multimodal face generation models address the spatial control limitations of text‑to‑image diffusion models by augmenting text‑based conditioning with spatial priors such as segmentation masks, sketches, or edge maps. This multimodal fusion enables controllable synthesis aligned with both high‑level semantic intent and low‑level structural layout. However, most existing approaches typically extend pre‑trained text‑to‑image pipelines by appending auxiliary control modules or stitching together separate uni‑modal networks. These ad hoc designs inherit architectural constraints, duplicate parameters, and often fail under conflicting modalities or mismatched latent spaces, limiting their ability to perform synergistic fusion across semantic and spatial domains. We introduce MMFace‑DiT, a unified dual‑stream diffusion transformer engineered for synergistic multimodal face synthesis. Its core novelty lies in a dual‑stream transformer block that processes spatial (mask/sketch) and semantic (text) tokens in parallel, deeply fusing them through a shared Rotary Position‑Embedded (RoPE) Attention mechanism. This design prevents modal dominance and ensures strong adherence to both text and structural priors to achieve unprecedented spatial‑semantic consistency for controllable face generation. Furthermore, a novel Modality Embedder enables a single cohesive model to dynamically adapt to varying spatial conditions without retraining. MMFace‑DiT achieves a 40% improvement in visual fidelity and prompt alignment over six state‑of‑the‑art multimodal face generation models, establishing a flexible new paradigm for end‑to‑end controllable generative modeling. The code and dataset are available on our project page: https://vcbsl.github.io/MMFace‑DiT/
Authors:Felix Wimbauer, Fabian Manhardt, Michael Oechsle, Nikolai Kalischek, Christian Rupprecht, Daniel Cremers, Federico Tombari
Abstract:
The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed‑forward 3D reconstruction, with vast potential in AR/VR and world modeling. While panoramic images have proven effective for scene initialization, existing approaches suffer from a trade‑off between visual fidelity and explorability: autoregressive expansion suffers from context drift, while panoramic video generation is limited to low resolution. We present Stepper, a unified framework for text‑driven immersive 3D scene synthesis that circumvents these limitations via stepwise panoramic scene expansion. Stepper leverages a novel multi‑view 360° diffusion model that enables consistent, high‑resolution expansion, coupled with a geometry reconstruction pipeline that enforces geometric coherence. Trained on a new large‑scale, multi‑view panorama dataset, Stepper achieves state‑of‑the‑art fidelity and structural consistency, outperforming prior approaches, thereby setting a new standard for immersive scene generation.
Authors:Kaituo Feng, Manyuan Zhang, Shuang Chen, Yunlong Lin, Kaixuan Fan, Yilei Jiang, Hongyu Li, Dian Zheng, Chenyang Wang, Xiangyu Yue
Abstract:
Recent image generation models have shown strong capabilities in generating high‑fidelity and photorealistic images. However, they are fundamentally constrained by frozen internal knowledge, thus often failing on real‑world scenarios that are knowledge‑intensive or require up‑to‑date information. In this paper, we present Gen‑Searcher, as the first attempt to train a search‑augmented image generation agent, which performs multi‑hop reasoning and search to collect the textual knowledge and reference images needed for grounded generation. To achieve this, we construct a tailored data pipeline and curate two high‑quality datasets, Gen‑Searcher‑SFT‑10k and Gen‑Searcher‑RL‑6k, containing diverse search‑intensive prompts and corresponding ground‑truth synthesis images. We further introduce KnowGen, a comprehensive benchmark that explicitly requires search‑grounded external knowledge for image generation and evaluates models from multiple dimensions. Based on these resources, we train Gen‑Searcher with SFT followed by agentic reinforcement learning with dual reward feedback, which combines text‑based and image‑based rewards to provide more stable and informative learning signals for GRPO training. Experiments show that Gen‑Searcher brings substantial gains, improving Qwen‑Image by around 16 points on KnowGen and 15 points on WISE. We hope this work can serve as an open foundation for search agents in image generation, and we fully open‑source our data, models, and code.
Authors:Zimu Zhang, Yucheng Zhang, Xiyan Xu, Ziyin Wang, Sirui Xu, Kai Zhou, Bing Zhou, Chuan Guo, Jian Wang, Yu-Xiong Wang, Liang-Yan Gui
Abstract:
Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole‑body models often miss the fine‑grained cues that drive dexterous behavior, finger articulation, contact timing, and inter‑hand coordination, and existing resources lack high‑fidelity bimanual sequences that capture nuanced finger dynamics and collaboration. To fill this gap, we present HandX, a unified foundation spanning data, annotation, and evaluation. We consolidate and filter existing datasets for quality, and collect a new motion‑capture dataset targeting underrepresented bimanual interactions with detailed finger dynamics. For scalable annotation, we introduce a decoupled strategy that extracts representative motion features, e.g., contact events and finger flexion, and then leverages reasoning from large language models to produce fine‑grained, semantically rich descriptions aligned with these features. Building on the resulting data and annotations, we benchmark diffusion and autoregressive models with versatile conditioning modes. Experiments demonstrate high‑quality dexterous motion generation, supported by our newly proposed hand‑focused metrics. We further observe clear scaling trends: larger models trained on larger, higher‑quality datasets produce more semantically coherent bimanual motion. Our dataset is released to support future research.
Authors:Derong Jin, Xiyi Chen, Ming C. Lin, Ruohan Gao
Abstract:
Tremendous progress in visual scene generation now turns a single image into an explorable 3D world, yet immersion remains incomplete without sound. We introduce Image2AVScene, the task of generating a 3D audio‑visual scene from a single image, and present SonoWorld, the first framework to tackle this challenge. From one image, our pipeline outpaints a 360° panorama, lifts it into a navigable 3D scene, places language‑guided sound anchors, and renders ambisonics for point, areal, and ambient sources, yielding spatial audio aligned with scene geometry and semantics. Quantitative evaluations on a newly curated real‑world dataset and a controlled user study confirm the effectiveness of our approach. Beyond free‑viewpoint audio‑visual rendering, we also demonstrate applications to one‑shot acoustic learning and audio‑visual spatial source separation. Project website: https://humathe.github.io/sonoworld/
Authors:Philip Schroeder, Thomas Weng, Karl Schmeckpeper, Eric Rosen, Stephen Hart, Ondrej Biza
Abstract:
Vision‑language models (VLMs) have shown impressive capabilities across diverse tasks, motivating efforts to leverage these models to supervise robot learning. However, when used as evaluators in reinforcement learning (RL), today's strongest models often fail under partial observability and distribution shift, enabling policies to exploit perceptual errors rather than solve the task. We introduce SOLE‑R1 (Self‑Observing LEarner), a video‑language reasoning model explicitly designed to serve as the sole reward signal for online RL. Given only raw video observations and a natural‑language goal, SOLE‑R1 performs per‑timestep spatiotemporal chain‑of‑thought (CoT) reasoning and produces dense estimates of task progress that can be used directly as rewards. To train SOLE‑R1, we develop a large‑scale video trajectory and reasoning synthesis pipeline that generates temporally grounded CoT traces aligned with continuous progress supervision. This data is combined with foundational spatial and multi‑frame temporal reasoning, and used to train the model with a hybrid framework that couples supervised fine‑tuning with RL from verifiable rewards. Across four different simulation environments and a real‑robot setting, SOLE‑R1 enables zero‑shot online RL from random initialization: robots learn previously unseen manipulation tasks without ground‑truth rewards, success indicators, demonstrations, or task‑specific tuning. SOLE‑R1 succeeds on 24 unseen tasks and substantially outperforms strong vision‑language rewarders, including Robometer, RoboReward, ReWiND, GPT‑5, and Gemini‑3‑Pro, while exhibiting markedly greater robustness to reward hacking. We release all models, data, code, and demos at the anonymous page: https://philip‑mit.github.io/sole‑r1/
Authors:Kailai Feng, Yuxiang Wei, Bo Chen, Yang Pan, Hu Ye, Songwei Liu, Chenqian Yan, Yuan Gao
Abstract:
Diffusion models have made significant progress in both text‑to‑image (T2I) generation and text‑guided image editing. However, these models are typically built with billions of parameters, leading to high latency and increased deployment challenges. While on‑device diffusion models improve efficiency, they largely focus on T2I generation and lack support for image editing. In this paper, we propose DreamLite, a compact unified on‑device diffusion model (0.39B) that supports both T2I generation and text‑guided image editing within a single network. DreamLite is built on a pruned mobile U‑Net backbone and unifies conditioning through in‑context spatial concatenation in the latent space. It concatenates images horizontally as input, using a (target | blank) configuration for generation tasks and (target | source) for editing tasks. To stabilize the training of this compact model, we introduce a task‑progressive joint pretraining strategy that sequentially targets T2I, editing, and joint tasks. After high‑quality SFT and reinforcement learning, DreamLite achieves GenEval (0.72) for image generation and ImgEdit (4.11) for image editing, outperforming existing on‑device models and remaining competitive with several server‑side models. By employing step distillation, we further reduce denoising processing to just 4 steps, enabling our DreamLite could generate or edit a 1024 x 1024 image in less than 1s on a Xiaomi 14 smartphone. To the best of our knowledge, DreamLite is the first unified on‑device diffusion model that supports both image generation and image editing.
Authors:Haozhe Qi, Kevin Qu, Mahdi Rad, Rui Wang, Alexander Mathis, Marc Pollefeys
Abstract:
Long video understanding remains challenging for Multi‑modal Large Language Models (MLLMs) due to high memory costs and context‑length limits. Prior approaches mitigate this by scoring and selecting frames/tokens within short clips, but they lack a principled mechanism to (i) compare relevance across distant video clips and (ii) stop processing once sufficient evidence has been gathered. We propose AdaptToken, a training‑free framework that turns an MLLM's self‑uncertainty into a global control signal for long‑video token selection. AdaptToken splits a video into groups, extracts cross‑modal attention to rank tokens within each group, and uses the model's response entropy to estimate each group's prompt relevance. This entropy signal enables a global token budget allocation across groups and further supports early stopping (AdaptToken‑Lite), skipping the remaining groups when the model becomes sufficiently certain. Across four long‑video benchmarks (VideoMME, LongVideoBench, LVBench, and MLVU) and multiple base MLLMs (7B‑72B), AdaptToken consistently improves accuracy (e.g., +6.7 on average over Qwen2.5‑VL 7B) and continues to benefit from extremely long inputs (up to 10K frames), while AdaptToken‑Lite reduces inference time by about half with comparable performance. Project page: https://haozheqi.github.io/adapt‑token
Authors:Chao Yin, Hongzhe Yue, Qing Han, Difeng Hu, Zhenyu Liang, Fangzhou Lin, Bing Sun, Boyu Wang, Mingkai Li, Wei Yao, Jack C. P. Cheng
Abstract:
Automated semantic understanding of dense point clouds is a prerequisite for Scan‑to‑BIM pipelines, digital twin construction, and as‑built verification‑‑core tasks in the digital transformation of the construction industry. Yet for industrial mechanical, electrical, and plumbing (MEP) facilities, this challenge remains largely unsolved: TLS acquisitions of water treatment plants, chiller halls, and pumping stations exhibit extreme geometric ambiguity, severe occlusion, and extreme class imbalance that architectural benchmarks (e.g., S3DIS or ScanNet) cannot adequately represent. We present Industrial3D, a terrestrial LiDAR dataset comprising 612 million expertly labelled points at 6 mm resolution from 13 water treatment facilities. At 6.6x the scale of the closest comparable MEP dataset, Industrial3D provides the largest and most demanding testbed for industrial 3D scene understanding to date. We further establish the first industrial cross‑paradigm benchmark, evaluating nine representative methods across fully supervised, weakly supervised, unsupervised, and foundation model settings under a unified benchmark protocol. The best supervised method achieves 55.74% mIoU, whereas zero‑shot Point‑SAM reaches only 15.79%‑‑a 39.95 percentage‑point gap that quantifies the unresolved domain‑transfer challenge for industrial TLS data. Systematic analysis reveals that this gap originates from a dual crisis: statistical rarity (215:1 imbalance, 3.5x more severe than S3DIS) and geometric ambiguity (tail‑class points share cylindrical primitives with head‑class pipes) that frequency‑based re‑weighting alone cannot resolve. Industrial3D, along with benchmark code and pre‑trained models, will be publicly available at https://github.com/pointcloudyc/Industrial3D.
Authors:Hannes Mareen, Dimitrios Karageorgiou, Paschalis Giakoumoglou, Peter Lambert, Symeon Papadopoulos, Glenn Van Wallendael
Abstract:
Generative AI has made text‑guided inpainting a powerful image editing tool, but at the same time a growing challenge for media forensics. Existing benchmarks, including our text‑guided inpainting forgery (TGIF) dataset, show that image forgery localization (IFL) methods can localize manipulations in spliced images but struggle not in fully regenerated (FR) images, while synthetic image detection (SID) methods can detect fully regenerated images but cannot perform localization. With new generative inpainting models emerging and the open problem of localization in FR images remaining, updated datasets and benchmarks are needed. We introduce TGIF2, an extended version of TGIF, that captures recent advances in text‑guided inpainting and enables a deeper analysis of forensic robustness. TGIF2 augments the original dataset with edits generated by FLUX.1 models, as well as with random non‑semantic masks. Using the TGIF2 dataset, we conduct a forensic evaluation spanning IFL and SID, including fine‑tuning IFL methods on FR images and generative super‑resolution attacks. Our experiments show that both IFL and SID methods degrade on FLUX.1 manipulations, highlighting limited generalization. Additionally, while fine‑tuning improves localization on FR images, evaluation with random non‑semantic masks reveals object bias. Furthermore, generative super‑resolution significantly weakens forensic traces, demonstrating that common image enhancement operations can undermine current forensic pipelines. In summary, TGIF2 provides an updated dataset and benchmark, which enables new insights into the challenges posed by modern inpainting and AI‑based image enhancements. TGIF2 is available at https://github.com/IDLabMedia/tgif‑dataset.
Authors:Huanxuan Liao, Zhongtao Jiang, Yupu Hao, Yuqiao Tan, Shizhu He, Ben Wang, Jun Zhao, Kun Xu, Kang Liu
Abstract:
Multimodal Large Language Models (MLLMs) achieve stronger visual understanding by scaling input fidelity, yet the resulting visual token growth makes jointly sustaining high spatial resolution and long temporal context prohibitive. We argue that the bottleneck lies not in how post‑encoding representations are compressed but in the volume of pixels the encoder receives, and address it with ResAdapt, an Input‑side adaptation framework that learns how much visual budget each frame should receive before encoding. ResAdapt couples a lightweight Allocator with an unchanged MLLM backbone, so the backbone retains its native visual‑token interface while receiving an operator‑transformed input. We formulate allocation as a contextual bandit and train the Allocator with Cost‑Aware Policy Optimization (CAPO), which converts sparse rollout feedback into a stable accuracy‑cost learning signal. Across budget‑controlled video QA, temporal grounding, and image reasoning tasks, ResAdapt improves low‑budget operating points and often lies on or near the efficiency‑accuracy frontier, with the clearest gains on reasoning‑intensive benchmarks under aggressive compression. Notably, ResAdapt supports up to 16x more frames at the same visual budget while delivering over 15% performance gain. Code is available at https://github.com/Xnhyacinth/ResAdapt.
Authors:Pavel Suma, Giorgos Kordopatis-Zilos, Yannis Kalantidis, Giorgos Tolias
Abstract:
Large‑scale instance‑level training data is scarce, so models are typically trained on domain‑specific datasets. Yet in real‑world retrieval, they must handle diverse domains, making generalization to unseen data critical. We introduce ELViS, an image‑to‑image similarity model that generalizes effectively to unseen domains. Unlike conventional approaches, our model operates in similarity space rather than representation space, promoting cross‑domain transfer. It leverages local descriptor correspondences, refines their similarities through an optimal transport step with data‑dependent gains that suppress uninformative descriptors, and aggregates strong correspondences via a voting process into an image‑level similarity. This design injects strong inductive biases, yielding a simple, efficient, and interpretable model. To assess generalization, we compile a benchmark of eight datasets spanning landmarks, artworks, products, and multi‑domain collections, and evaluate ELViS as a re‑ranking method. Our experiments show that ELViS outperforms competing methods by a large margin in out‑of‑domain scenarios and on average, while requiring only a fraction of their computational cost. Code available at: https://github.com/pavelsuma/ELViS/
Authors:Quan Meng, Yujin Chen, Lei Li, Matthias Nießner, Angela Dai
Abstract:
We present Seen2Scene, the first flow matching‑based approach that trains directly on incomplete, real‑world 3D scans for scene completion and generation. Unlike prior methods that rely on complete and hence synthetic 3D data, our approach introduces visibility‑guided flow matching, which explicitly masks out unknown regions in real scans, enabling effective learning from real‑world, partial observations. We represent 3D scenes using truncated signed distance field (TSDF) volumes encoded in sparse grids and employ a sparse transformer to efficiently model complex scene structures while masking unknown regions. We employ 3D layout boxes as an input conditioning signal, and our approach is flexibly adapted to various other inputs such as text or partial scans. By learning directly from real‑world, incomplete 3D scans, Seen2Scene enables realistic 3D scene completion for complex, cluttered real environments. Experiments demonstrate that our model produces coherent, complete, and realistic 3D scenes, outperforming baselines in completion accuracy and generation quality.
Authors:Claudia Cuttano, Gabriele Trivigno, Christoph Reich, Daniel Cremers, Carlo Masone, Stefan Roth
Abstract:
In‑context segmentation (ICS) aims to segment arbitrary concepts, e.g., objects, parts, or personalized instances, given one annotated visual examples. Existing work relies on (i) fine‑tuning vision foundation models (VFMs), which improves in‑domain results but harms generalization, or (ii) combines multiple frozen VFMs, which preserves generalization but yields architectural complexity and fixed segmentation granularities. We revisit ICS from a minimalist perspective and ask: Can a single self‑supervised backbone support both semantic matching and segmentation, without any supervision or auxiliary models? We show that scaled‑up dense self‑supervised features from DINOv3 exhibit strong spatial structure and semantic correspondence. We introduce INSID3, a training‑free approach that segments concepts at varying granularities only from frozen DINOv3 features, given an in‑context example. INSID3 achieves state‑of‑the‑art results across one‑shot semantic, part, and personalized segmentation, outperforming previous work by +7.5 % mIoU, while using 3x fewer parameters and without any mask or category‑level supervision. Code is available at https://github.com/visinf/INSID3 .
Authors:Yangmei Chen, Zhongyuan Zhang, Xikun Zhang, Xinyu Hao, Mingliang Hou, Renqiang Luo, Ziqi Xu
Abstract:
Thyroid nodule classification using ultrasound imaging is essential for early diagnosis and clinical decision‑making; however, despite promising performance on in‑distribution data, existing deep learning methods often exhibit limited robustness and generalisation when deployed across different ultrasound devices or clinical environments. This limitation is mainly attributed to the pronounced heterogeneity of thyroid ultrasound images, which can lead models to capture spurious correlations rather than reliable diagnostic cues. To address this challenge, we propose PEMV‑thyroid, a Prototype‑Enhanced Multi‑View learning framework that accounts for data heterogeneity by learning complementary representations from multiple feature perspectives and refining decision boundaries through a prototype‑based correction mechanism with mixed prototype information. By integrating multi‑view representations with prototype‑level guidance, the proposed approach enables more stable representation learning under heterogeneous imaging conditions. Extensive experiments on multiple thyroid ultrasound datasets demonstrate that PEMV‑thyroid consistently outperforms state‑of‑the‑art methods, particularly in cross‑device and cross‑domain evaluation scenarios, leading to improved diagnostic accuracy and generalisation performance in real‑world clinical settings. The source code is available at https://github.com/chenyangmeii/Prototype‑Enhanced‑Multi‑View‑Learning.
Authors:Minh-Khoi Do, Huy Che, Dinh-Duy Phan, Duc-Khai Lam, Duc-Lung Vu
Abstract:
Accurate and efficient perception is essential for autonomous driving, where segmentation tasks such as drivable‑area and lane segmentation provide critical cues for motion planning and control. However, achieving high segmentation accuracy while maintaining real‑time performance on low‑cost hardware remains a challenging problem. To address this issue, we introduce TwinMixing, a lightweight multi‑task segmentation model designed explicitly for drivable‑area and lane segmentation. The proposed network features a shared encoder and task‑specific decoders, enabling both feature sharing and task specialization. Within the encoder, we propose an Efficient Pyramid Mixing (EPM) module that enhances multi‑scale feature extraction through a combination of grouped convolutions, depthwise dilated convolutions and channel shuffle operations, effectively expanding the receptive field while minimizing computational cost. Each decoder adopts a Dual‑Branch Upsampling (DBU) Block composed of a learnable transposed convolution‑based Fine detailed branch and a parameter‑free bilinear interpolation‑based Coarse grained branch, achieving detailed yet spatially consistent feature reconstruction. Extensive experiments on the BDD100K dataset validate the effectiveness of TwinMixing across three configurations ‑ tiny, base, and large. Among them, the base configuration achieves the best trade‑off between accuracy and computational efficiency, reaching 92.0% mIoU for drivable‑area segmentation and 32.3% IoU for lane segmentation with only 0.43M parameters and 3.95 GFLOPs. Moreover, TwinMixing consistently outperforms existing segmentation models on the same tasks, as illustrated in Fig. 1. Thanks to its compact and modular design, TwinMixing demonstrates strong potential for real‑time deployment in autonomous driving and embedded perception systems. The source code: https://github.com/Jun0se7en/TwinMixing.
Authors:Kazuma Ikeda, Ryosei Hara, Rokuto Nagata, Ozora Sako. Zihao Ding, Takahiro Kado, Ibuki Fujioka, Taro Beppu, Mariko Isogawa, Kentaro Yoshioka
Abstract:
LiDAR has become an essential sensing modality in autonomous driving, robotics, and smart‑city applications. However, ghost points (or ghosts), which are false reflections caused by multi‑path laser returns from glass and reflective surfaces, severely degrade 3D mapping and localization accuracy. Prior ghost removal relies on geometric consistency in dense point clouds, failing on mobile LiDAR's sparse, dynamic data. We address this by exploiting full‑waveform LiDAR (FWL), which captures complete temporal intensity profiles rather than just peak distances, providing crucial cues for distinguishing ghosts from genuine reflections in mobile scenarios. As this is a new task, we present Ghost‑FWL, the first and largest annotated mobile FWL dataset for ghost detection and removal. Ghost‑FWL comprises 24K frames across 10 diverse scenes with 7.5 billion peak‑level annotations, which is 100x larger than existing annotated FWL datasets. Benefiting from this large‑scale dataset, we establish a FWL‑based baseline model for ghost detection and propose FWL‑MAE, a masked autoencoder for efficient self‑supervised representation learning on FWL data. Experiments show that our baseline outperforms existing methods in ghost removal accuracy, and our ghost removal further enhances downstream tasks such as LiDAR‑based SLAM (66% trajectory error reduction) and 3D object detection (50x false positive reduction). The dataset and code is publicly available and can be accessed via the project page: https://keio‑csg.github.io/Ghost‑FWL
Authors:Onat Ozdemir, Anders Christensen, Stephan Alaniz, Zeynep Akata, Emre Akbas
Abstract:
Large‑scale vision‑language models such as CLIP have achieved remarkable success in zero‑shot image recognition, yet their predictions remain largely opaque to human understanding. In contrast, Concept Bottleneck Models provide interpretable intermediate representations by reasoning through human‑defined concepts, but they rely on concept supervision and lack the ability to generalize to unseen classes. We introduce EZPC that bridges these two paradigms by explaining CLIP's zero‑shot predictions through human‑understandable concepts. Our method projects CLIP's joint image‑text embeddings into a concept space learned from language descriptions, enabling faithful and transparent explanations without additional supervision. The model learns this projection via a combination of alignment and reconstruction objectives, ensuring that concept activations preserve CLIP's semantic structure while remaining interpretable. Extensive experiments on five benchmark datasets, CIFAR‑100, CUB‑200‑2011, Places365, ImageNet‑100, and ImageNet‑1k, demonstrate that our approach maintains CLIP's strong zero‑shot classification accuracy while providing meaningful concept‑level explanations. By grounding open‑vocabulary predictions in explicit semantic concepts, our method offers a principled step toward interpretable and trustworthy vision‑language models. Code is available at https://github.com/oonat/ezpc.
Authors:Xuanlong Yu, Youyang Sha, Longfei Liu, Xi Shen, Di Yang
Abstract:
Few‑shot object detection (FSOD) is challenging due to unstable optimization and limited generalization arising from the scarcity of training samples. To address these issues, we propose a hybrid ensemble decoder that enhances generalization during fine‑tuning. Inspired by ensemble learning, the decoder comprises a shared hierarchical layer followed by multiple parallel decoder branches, where each branch employs denoising queries either inherited from the shared layer or newly initialized to encourage prediction diversity. This design fully exploits pretrained weights without introducing additional parameters, and the resulting diverse predictions can be effectively ensembled to improve generalization. We further leverage a unified progressive fine‑tuning framework with a plateau‑aware learning rate schedule, which stabilizes optimization and achieves strong few‑shot adaptation without complex data augmentations or extensive hyperparameter tuning. Extensive experiments on CD‑FSOD, ODinW‑13, and RF100‑VL validate the effectiveness of our approach. Notably, on RF100‑VL, which includes 100 datasets across diverse domains, our method achieves an average performance of 41.9 in the 10‑shot setting, significantly outperforming the recent approach SAM3, which obtains 35.7. We further construct a mixed‑domain test set from CD‑FSOD to evaluate robustness to out‑of‑distribution (OOD) samples, showing that our proposed modules lead to clear improvement gains. These results highlight the effectiveness, generalization, and robustness of the proposed method. Code is available at: https://github.com/Intellindust‑AI‑Lab/FT‑FSOD.
Authors:Qiya Song, Yiqiang Xie, Yuan Sun, Renwei Dian, Xudong Kang
Abstract:
As a pivotal task that bridges remote visual and linguistic understanding, Remote Sensing Image‑Text Retrieval (RSITR) has attracted considerable research interest in recent years. However, almost all RSITR methods implicitly assume that image‑text pairs are matched perfectly. In practice, acquiring a large set of well‑aligned data pairs is often prohibitively expensive or even infeasible. In addition, we also notice that the remote sensing datasets (e.g., RSITMD) truly contain some inaccurate or mismatched image text descriptions. Based on the above observations, we reveal an important but untouched problem in RSITR, i.e., Noisy Correspondence (NC). To overcome these challenges, we propose a novel Robust Remote Sensing Image‑Text Retrieval (RRSITR) paradigm that designs a self‑paced learning strategy to mimic human cognitive learning patterns, thereby learning from easy to hard from multi‑modal data with NC. Specifically, we first divide all training sample pairs into three categories based on the loss magnitude of each pair, i.e., clean sample pairs, ambiguous sample pairs, and noisy sample pairs. Then, we respectively estimate the reliability of each training pair by assigning a weight to each pair based on the values of the loss. Further, we respectively design a new multi‑modal self‑paced function to dynamically regulate the training sequence and weights of the samples, thus establishing a progressive learning process. Finally, for noisy sample pairs, we present a robust triplet loss to dynamically adjust the soft margin based on semantic similarity, thereby enhancing the robustness against noise. Extensive experiments on three popular benchmark datasets demonstrate that the proposed RRSITR significantly outperforms the state‑of‑the‑art methods, especially in high noise rates. The code is available at: https://github.com/MSFLabX/RRSITR
Authors:Zhang Li, Zhibo Lin, Qiang Liu, Ziyang Zhang, Shuo Zhang, Zidun Guo, Jiajun Song, Jiarui Zhang, Xiang Bai, Yuliang Liu
Abstract:
We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, yet almost exclusively on clean, digital, well‑formatted pages in a handful of dominant languages. No systematic benchmark exists to evaluate how models perform on digital and photographed documents across diverse scripts and low‑resource languages. MDPBench comprises 3,400 document images spanning 17 languages, diverse scripts, and varied photographic conditions, with high‑quality annotations produced through a rigorous pipeline of expert model labeling, manual correction, and human verification. To ensure fair comparison and prevent data leakage, we maintain separate public and private evaluation splits. Our comprehensive evaluation of both open‑source and closed‑source models uncovers a striking finding: while closed‑source models (notably Gemini3‑Pro) prove relatively robust, open‑source alternatives suffer dramatic performance collapse, particularly on non‑Latin scripts and real‑world photographed documents, with an average drop of 17.8% on photographed documents and 14.0% on non‑Latin scripts. These results reveal significant performance imbalances across languages and conditions, and point to concrete directions for building more inclusive, deployment‑ready parsing systems. Source available at https://github.com/Yuliang‑Liu/MultimodalOCR.
Authors:Pengcheng Xue, Yan Tian, Qiutao Song, Ziyi Wang, Linyang He, Weiping Ding, Mahmoud Hassaballah, Karen Egiazarian, Wei-Fa Yang, Leszek Rutkowski
Abstract:
Text‑driven 3D scene editing has attracted considerable interest due to its convenience and user‑friendliness. However, methods that rely on implicit 3D representations, such as Neural Radiance Fields (NeRF), while effective in rendering complex scenes, are hindered by slow processing speeds and limited control over specific regions of the scene. Moreover, existing approaches, including Instruct‑NeRF2NeRF and GaussianEditor, which utilize multi‑view editing strategies, frequently produce inconsistent results across different views when executing text instructions. This inconsistency can adversely affect the overall performance of the model, complicating the task of balancing the consistency of editing results with editing efficiency. To address these challenges, we propose a novel method termed Single‑View to 3D Object Editing via Gaussian Splatting (SVGS), which is a single‑view text‑driven editing technique based on 3D Gaussian Splatting (3DGS). Specifically, in response to text instructions, we introduce a single‑view editing strategy grounded in multi‑view diffusion models, which reconstructs 3D scenes by leveraging only those views that yield consistent editing results. Additionally, we employ sparse 3D Gaussian Splatting as the 3D representation, which significantly enhances editing efficiency. We conducted a comparative analysis of SVGS against existing baseline methods across various scene settings, and the results indicate that SVGS outperforms its counterparts in both editing capability and processing speed, representing a significant advancement in 3D editing technology. For further details, please visit our project page at: https://amateurc.github.io/svgs.github.io.
Authors:Guangjing Yang, Ziyuan Qin, Chaoran Zhang, Chenlin Du, Jinlin Wang, Wanran Sun, Zhenyu Zhang, Bing Ji, Qicheng Lao
Abstract:
Medical visual grounding serves as a crucial foundation for fine‑grained multimodal reasoning and interpretable clinical decision support. Despite recent advances in reinforcement learning (RL) for grounding tasks, existing approaches such as Group Relative Policy Optimization~(GRPO) suffer from severe reward sparsity when directly applied to medical images, primarily due to the inherent difficulty of localizing small or ambiguous regions of interest, which is further exacerbated by the rigid and suboptimal nature of fixed IoU‑based reward schemes in RL. This leads to vanishing policy gradients and stagnated optimization, particularly during early training. To address this challenge, we propose MedLoc‑R1, a performance‑aware reward scheduling framework that progressively tightens the reward criterion in accordance with model readiness. MedLoc‑R1 introduces a sliding‑window performance tracker and a multi‑condition update rule that automatically adjust the reward schedule from dense, easily obtainable signals to stricter, fine‑grained localization requirements, while preserving the favorable properties of GRPO without introducing auxiliary networks or additional gradient paths. Experiments on three medical visual grounding benchmarks demonstrate that MedLoc‑R1 consistently improves both localization accuracy and training stability over GRPO‑based baselines. Our framework offers a general, lightweight, and effective solution for RL‑based grounding in high‑stakes medical applications. Code \& checkpoints are available at \hyperlinkhttps://github.com/MembrAI/MedLoc‑R1.
Authors:Yuqi Ye, Zijian Zhang, Junhong Lin, Shangkun Sun, Changhao Peng, Wei Gao
Abstract:
Vision‑language models (VLMs) are increasingly being adopted for end‑to‑end autonomous driving systems due to their exceptional performance in handling long‑tail scenarios. However, current VLM‑based approaches suffer from two major limitations: 1) Some VLMs directly output planning results without chain‑of‑thought (CoT) reasoning, bypassing crucial perception and prediction stages which creates a significant domain gap and compromises decision‑making capability; 2) Other VLMs can generate outputs for perception, prediction, and planning tasks but employ a fragmented decision‑making approach where these modules operate separately, leading to a significant lack of synergy that undermines true planning performance. To address these limitations, we propose AutoDrive\text‑P^3, a novel framework that seamlessly integrates Perception, Prediction, and Planning through structured reasoning. We introduce the P^3\text‑CoT dataset to facilitate coherent reasoning and propose P^3\text‑GRPO, a hierarchical reinforcement learning algorithm that provides progressive supervision across all three tasks. Specifically, AutoDrive\text‑P^3 progressively generates CoT reasoning and answers for perception, prediction, and planning, where perception provides essential information for subsequent prediction and planning, while both perception and prediction collectively contribute to the final planning decisions, enabling safer and more interpretable autonomous driving. Additionally, to balance inference efficiency with performance, we introduce dual thinking modes: detailed thinking and fast thinking. Extensive experiments on both open‑loop (nuScenes) and closed‑loop (NAVSIMv1/v2) benchmarks demonstrate that our approach achieves state‑of‑the‑art performance in planning tasks. Code is available at https://github.com/haha‑yuki‑haha/AutoDrive‑P3.
Authors:Chunhang Zheng, Tongda Xu, Mingli Xie, Yan Wang, Dou Li
Abstract:
Raw images preserve linear sensor measurements and high bit‑depth information crucial for advanced vision tasks and photography applications, yet their storage remains challenging due to large file sizes, varying bit depths, and sensor‑dependent characteristics. Existing learned lossless compression methods mainly target 8‑bit sRGB images, while raw reconstruction approaches are inherently lossy and rely on camera‑specific assumptions. To address these challenges, we introduce RAWIC, a bit‑depth‑adaptive learned lossless compression framework for Bayer‑pattern raw images. We first convert single‑channel Bayer data into a four‑channel RGGB format and partition it into patches. For each patch, we compute its bit depth and use it as auxiliary input to guide compression. A bit‑depth‑adaptive entropy model is then designed to estimate patch distributions conditioned on their bit depths. This architecture enables a single model to handle raw images from diverse cameras and bit depths. Experiments show that RAWIC consistently surpasses traditional lossless codecs, achieving an average 7.7% bitrate reduction over JPEG‑XL. Our code is available at https://github.com/chunbaobao/RAWIC.
Authors:Alexander Prutsch, Christian Fruhwirth-Reisinger, David Schinagl, Horst Possegger
Abstract:
In dynamic traffic environments, motion forecasting models must be able to accurately estimate future trajectories continuously. Streaming‑based methods are a promising solution, but despite recent advances, their performance often degrades when exposed to heterogeneous observation lengths. To address this, we propose a novel streaming‑based motion forecasting framework that explicitly focuses on evolving scenes. Our method incrementally processes incoming observation windows and leverages an instance‑aware context streaming to maintain and update latent agent representations across inference steps. A dual training objective further enables consistent forecasting accuracy across diverse observation horizons. Extensive experiments on Argoverse 2, nuScenes, and Argoverse 1 demonstrate the robustness of our approach under evolving scene conditions and also on the single‑agent benchmarks. Our model achieves state‑of‑the‑art performance in streaming inference on the Argoverse 2 multi‑agent benchmark, while maintaining minimal latency, highlighting its suitability for real‑world deployment.
Authors:Zhen Zou, Xiaoxiao Ma, Mingde Yao, Jie Huang, LinJiang Huang, Feng Zhao
Abstract:
Autoregressive (AR)‑Diffusion hybrid paradigms combine AR's structured semantic modeling with diffusion's high‑fidelity synthesis, yet suffer from a dual speed bottleneck: the sequential AR stage and the iterative multi‑step denoising of the diffusion vision decode stage. Existing methods address each in isolation without a unified principle design. We observe that the per‑position \emphprediction entropy of continuous‑space AR models naturally encodes spatially varying generation uncertainty, which simultaneously governing draft prediction quality in the AR stage and reflecting the corrective effort required by vision decoding stage, which is not fully explored before. Since entropy is inherently tied to both bottlenecks, it serves as a natural unifying signal for joint acceleration. In this work, we propose Drift‑AR, which leverages entropy signal to accelerate both stages: 1) for AR acceleration, we introduce Entropy‑Informed Speculative Decoding that align draft‑target entropy distributions via a causal‑normalized entropy loss, resolving the entropy mismatch that causes excessive draft rejection; 2) for visual decoder acceleration, we reinterpret entropy as the \emphphysical variance of the initial state for an anti‑symmetric drifting field ‑‑ high‑entropy positions activate stronger drift toward the data manifold while low‑entropy positions yield vanishing drift ‑‑ enabling single‑step (1‑NFE) decoding without iterative denoising or distillation. Moreover, both stages share the same entropy signal, which is computed once with no extra cost. Experiments on MAR, TransDiff, and NextStep‑1 demonstrate 3.8‑5.5× speedup with genuine 1‑NFE decoding, matching or surpassing original quality. Code will be available at https://github.com/aSleepyTree/Drift‑AR.
Authors:Jae-Young Kang, Hoonhee Cho, Taeyeop Lee, Minjun Kang, Bowen Wen, Youngho Kim, Kuk-Jin Yoon
Abstract:
Event cameras provide microsecond latency, making them suitable for 6D object pose tracking in fast, dynamic scenes where conventional RGB and depth pipelines suffer from motion blur and large pixel displacements. We introduce EventTrack6D, an event‑depth tracking framework that generalizes to novel objects without object‑specific training by reconstructing both intensity and depth at arbitrary timestamps between depth frames. Conditioned on the most recent depth measurement, our dual reconstruction recovers dense photometric and geometric cues from sparse event streams. Our EventTrack6D operates at over 120 FPS and maintains temporal consistency under rapid motion. To support training and evaluation, we introduce a comprehensive benchmark suite: a large‑scale synthetic dataset for training and two complementary evaluation sets, including real and simulated event datasets. Trained exclusively on synthetic data, EventTrack6D generalizes effectively to real‑world scenarios without fine‑tuning, maintaining accurate tracking across diverse objects and motion patterns. Our method and datasets validate the effectiveness of event cameras for event‑based 6D pose tracking of novel objects. Code and datasets are publicly available at https://chohoonhee.github.io/Event6D.
Authors:Tianle Zeng, Yanci Wen, Hong Zhang
Abstract:
The convergence of low‑altitude economies, embodied intelligence, and air‑ground cooperative systems creates growing demand for simulation infrastructure capable of jointly modeling aerial and ground agents within a single physically coherent environment. Existing open‑source platforms remain domain‑segregated: driving simulators lack aerial dynamics, while multirotor simulators lack realistic ground scenes. Bridge‑based co‑simulation introduces synchronization overhead and cannot guarantee strict spatial‑temporal consistency.
We present CARLA‑Air, an open‑source infrastructure that unifies high‑fidelity urban driving and physics‑accurate multirotor flight within a single Unreal Engine process. The platform preserves both CARLA and AirSim native Python APIs and ROS 2 interfaces, enabling zero‑modification code reuse. Within a shared physics tick and rendering pipeline, CARLA‑Air delivers photorealistic environments with rule‑compliant traffic, socially‑aware pedestrians, and aerodynamically consistent UAV dynamics, synchronously capturing up to 18 sensor modalities across all platforms at each tick. The platform supports representative air‑ground embodied intelligence workloads spanning cooperation, embodied navigation and vision‑language action, multi‑modal perception and dataset construction, and reinforcement‑learning‑based policy training. An extensible asset pipeline allows integration of custom robot platforms into the shared world. By inheriting AirSim's aerial capabilities ‑‑ whose upstream development has been archived ‑‑ CARLA‑Air ensures this widely adopted flight stack continues to evolve within a modern infrastructure.
Released with prebuilt binaries and full source: https://github.com/louiszengCN/CarlaAir
Authors:Huimin Zeng, Yue Bai, Hailing Wang, Yun Fu
Abstract:
High dynamic range novel view synthesis (HDR‑NVS) reconstructs scenes with dynamic details by fusing multi‑exposure low dynamic range (LDR) views, yet it struggles to capture ambient illumination‑dependent appearance. Implicitly supervising HDR content by constraining tone‑mapped results fails in correcting abnormal HDR values, and results in limited gradients for Gaussians in under/over‑exposed regions. To this end, we introduce PhysHDR‑GS, a physically inspired HDR‑NVS framework that models scene appearance via intrinsic reflectance and adjustable ambient illumination. PhysHDR‑GS employs a complementary image‑exposure (IE) branch and Gaussian‑illumination (GI) branch to faithfully reproduce standard camera observations and capture illumination‑dependent appearance changes, respectively. During training, the proposed cross‑branch HDR consistency loss provides explicit supervision for HDR content, while an illumination‑guided gradient scaling strategy mitigates exposure‑biased gradient starvation and reduces under‑densified representations. Experimental results across realistic and synthetic datasets demonstrate our superiority in reconstructing HDR details (e.g., a PSNR gain of 2.04 dB over HDR‑GS), while maintaining real‑time rendering speed (up to 76 FPS). Code and models are available at https://huimin‑zeng.github.io/PhysHDR‑GS/.
Authors:Muhammad Osama Zeeshan, Masoumeh Sharafi, Benoît Savary, Alessandro Lameiras Koerich, Marco Pedersoli, Eric Granger
Abstract:
Personalization in emotion recognition (ER) is essential for an accurate interpretation of subtle and subject‑specific expressive patterns. Recent advances in vision‑language models (VLMs) such as CLIP demonstrate strong potential for leveraging joint image‑text representations in ER. However, CLIP‑based methods either depend on CLIP's contrastive pretraining or on LLMs to generate descriptive text prompts, which are noisy, computationally expensive, and fail to capture fine‑grained expressions, leading to degraded performance. In this work, we leverage Action Units (AUs) as structured textual prompts within CLIP to model fine‑grained facial expressions. AUs encode the subtle muscle activations underlying expressions, providing localized and interpretable semantic cues for more robust ER. We introduce CLIP‑AU, a lightweight AU‑guided temporal learning method that integrates interpretable AU semantics into CLIP. It learns generic, subject‑agnostic representations by aligning AU prompts with facial dynamics, enabling fine‑grained ER without CLIP fine‑tuning or LLM‑generated text supervision. Although CLIP‑AU models fine‑grained AU semantics, it does not adapt to subject‑specific variability in subtle expressions. To address this limitation, we propose CLIP‑AUTT, a video‑based test‑time personalization method that dynamically adapts AU prompts to videos from unseen subjects. By combining entropy‑guided temporal window selection with prompt tuning, CLIP‑AUTT enables subject‑specific adaptation while preserving temporal consistency. Our extensive experiments on three challenging video‑based subtle ER datasets, BioVid, StressID, and BAH, indicate that CLIP‑AU and CLIP‑AUTT outperform state‑of‑the‑art CLIP‑based FER and TTA methods, achieving robust and personalized subtle ER. Our code is publicly available at: https://github.com/osamazeeshan/CLIP‑AUTT.
Authors:Mohab Kishawy, Jun Chen
Abstract:
We propose RetinexDualV2, a unified, physically grounded dual‑branch framework for diverse Ultra‑High‑Definition (UHD) image restoration. Unlike generic models, our method employs a Task‑Specific Physical Grounding Module (TS‑PGM) to extract degradation‑aware priors (e.g., rain masks and dark channels). These explicitly guide a Retinex decomposition network via a novel Physical‑Conditioned Multi‑head Self‑Attention (PC‑MSA) mechanism, enabling robust reflection and illumination correction. This physical conditioning allows a single architecture to handle various complex degradations seamlessly, without task‑specific structural modifications. RetinexDualV2 demonstrates exceptional generalizability, securing 4th place in the NTIRE 2026 Day and Night Raindrop Removal Challenge and 5th place in the Joint Noise Low‑light Enhancement (JNLLIE) Challenge. Extensive experiments confirm the state‑of‑the‑art performance and efficiency of our physically motivated approach. Code is available at https://github.com/ErrorLogic1211/RetinexDual/tree/master/RetinexDualV2
Authors:Pei An, Junfeng Ding, Jiaqi Yang, Yulong Wang, Jie Ma, Liangliang Nan
Abstract:
Image‑to‑point‑cloud (I2P) registration aims to align 2D images with 3D point clouds by establishing reliable 2D‑3D correspondences. The drastic modality gap between images and point clouds makes it challenging to learn features that are both discriminative and generalizable, leading to severe performance drops in unseen scenarios.
We address this challenge by introducing a heterogeneous graph that enables refining both cross‑modal features and correspondences within a unified architecture. The proposed graph represents a mapping between segmented 2D and 3D regions, which enhances cross‑modal feature interaction and thus improves feature discriminability. In addition, modeling the consistency among vertices and edges within the graph enables pruning of unreliable correspondences. Building on these insights, we propose a heterogeneous graph embedded I2P registration method, termed Hg‑I2P. It learns a heterogeneous graph by mining multi‑path feature relationships, adapts features under the guidance of heterogeneous edges, and prunes correspondences using graph‑based projection consistency. Experiments on six indoor and outdoor benchmarks under cross‑domain setups demonstrate that Hg‑I2P significantly outperforms existing methods in both generalization and accuracy. Code is released on https://github.com/anpei96/hg‑i2p‑demo.
Authors:Pragat Wagle, Zheng Chen, Lantao Liu
Abstract:
Robust scene understanding is essential for intelligent vehicles operating in natural, unstructured environments. While semantic segmentation datasets for structured urban driving are abundant, the datasets for extremely unstructured wild environments remain scarce due to the difficulty and cost of generating pixel‑accurate annotations. These limitations hinder the development of perception systems needed for intelligent ground vehicles tasked with forestry automation, agricultural robotics, disaster response, and all‑terrain mobility. To address this gap, we present ForestSim, a high‑fidelity synthetic dataset designed for training and evaluating semantic segmentation models for intelligent vehicles in forested off‑road and no‑road environments. ForestSim contains 2094 photorealistic images across 25 diverse environments, covering multiple seasons, terrain types, and foliage densities. Using Unreal Engine environments integrated with Microsoft AirSim, we generate consistent, pixel‑accurate labels across 20 classes relevant to autonomous navigation. We benchmark ForestSim using state‑of‑the‑art architectures and report strong performance despite the inherent challenges of unstructured scenes. ForestSim provides a scalable and accessible foundation for perception research supporting the next generation of intelligent off‑road vehicles. The dataset and code are publicly available: Dataset: https://vailforestsim.github.io Code: https://github.com/pragatwagle/ForestSim
Authors:Liuzhou Zhang, Zeyu Zhang, Biao Wu, Luyao Tang, Zirui Song, Hongyang He, Renda Han, Guangzhen Yao, Huacan Wang, Ronghao Chen, Xiuying Chen, Guan Huang, Zheng Zhu
Abstract:
Sign language plays a crucial role in bridging communication gaps between the deaf and hard‑of‑hearing communities. However, existing sign language video generation models often rely on complex intermediate representations, which limits their flexibility and efficiency. In this work, we propose a novel pose‑free framework for real‑time sign language video generation. Our method eliminates the need for intermediate pose representations by directly mapping natural language text to sign language videos using a diffusion‑based approach. We introduce two key innovations: (1) a pose‑free generative model based on the a state‑of‑the‑art diffusion backbone, which learns implicit text‑to‑gesture alignments without pose estimation, and (2) a Trainable Sliding Tile Attention (T‑STA) mechanism that accelerates inference by exploiting spatio‑temporal locality patterns. Unlike previous training‑free sparsity approaches, T‑STA integrates trainable sparsity into both training and inference, ensuring consistency and eliminating the train‑test gap. This approach significantly reduces computational overhead while maintaining high generation quality, making real‑time deployment feasible. Our method increases video generation speed by 3.07x without compromising video quality. Our contributions open new avenues for real‑time, high‑quality, pose‑free sign language synthesis, with potential applications in inclusive communication tools for diverse communities. Code: https://github.com/AIGeeksGroup/FlashSign.
Authors:Dexing Huang, Shiao Wang, Fan Zhang, Xiao Wang
Abstract:
Robust visual object tracking (VOT) remains challenging in high‑speed motion scenarios, where conventional RGB sensors suffer from severe motion blur and performance degradation. Event cameras, with microsecond temporal resolution and high dynamic range, provide complementary structural cues that can potentially compensate for these limitations. However, existing RGB‑Event fusion methods typically treat event data as dense intensity representations and adopt black‑box fusion strategies, failing to explicitly leverage the directional geometric priors inherently encoded in event streams to rectify degraded RGB features. To address this limitation, we propose SOR‑Track, a streamlined framework for robust RGB‑Event tracking based on Spatial Orthogonal Refinement (SOR). The core SOR module employs a set of orthogonal directional filters that are dynamically guided by local motion orientations to extract sharp and motion‑consistent structural responses from event streams. These responses serve as geometric anchors to modulate and refine aliased RGB textures through an asymmetric structural modulation mechanism, thereby explicitly bridging structural discrepancies between two modalities. Extensive experiments on the large‑scale FE108 benchmark demonstrate that SOR‑Track consistently outperforms existing fusion‑based trackers, particularly under motion blur and low‑light conditions. Despite its simplicity, the proposed method offers a principled and physics‑grounded approach to multi‑modal feature alignment and texture rectification. The source code of this paper will be released on https://github.com/Event‑AHU/OpenEvTracking
Authors:Irene Kim, Sai Tanmay Reddy Chakkera, Alexandros Graikos, Dimitris Samaras, Akshat Dave
Abstract:
Monocular surface normal estimators trained on large‑scale RGB‑normal data often perform poorly in the edge cases of reflective, textureless, and dark surfaces. Polarization encodes surface orientation independently of texture and albedo, offering a physics‑based complement for these cases. Existing polarization methods, however, require multi‑view capture or specialized training data, limiting generalization. We introduce Poppy, a training‑free framework that refines normals from any frozen RGB backbone using single‑shot polarization measurements at test time. Keeping backbone weights frozen, Poppy optimizes per‑pixel offsets to the input RGB and output normal along with a learned reflectance decomposition. A differentiable rendering layer converts the refined normals into polarization predictions and penalizes mismatches with the observed signal. Across seven benchmarks and three backbone architectures (diffusion, flow, and feed‑forward), Poppy reduces mean angular error by 23‑26% on synthetic data and 6‑16% on real data. These results show that guiding learned RGB‑based normal estimators with polarization cues at test time refines normals on challenging surfaces without retraining.
Authors:Xiangzhong Liu, Hao Shen
Abstract:
Modern autonomous driving systems increasingly rely on mixed camera configurations with pinhole and fisheye cameras for full view perception. However, Bird's‑Eye View (BEV) 3D object detection models are predominantly designed for pinhole cameras, leading to performance degradation under fisheye distortion. To bridge this gap, we introduce a multi‑view BEV detection benchmark with mixed cameras by converting KITTI‑360 into nuScenes format. Our study encompasses three adaptations: rectification for zero‑shot evaluation and fine‑tuning of nuScenes‑trained models, distortion‑aware view transformation modules (VTMs) via the MEI camera model, and polar coordinate representations to better align with radial distortion. We systematically evaluate three representative BEV architectures, BEVFormer, BEVDet and PETR, across these strategies. We demonstrate that projection‑free architectures are inherently more robust and effective against fisheye distortion than other VTMs. This work establishes the first real‑data 3D detection benchmark with fisheye and pinhole images and provides systematic adaptation and practical guidelines for designing robust and cost‑effective 3D perception systems. The code is available at https://github.com/CesarLiu/FishBEVOD.git.
Authors:Linfei Li, Lin Zhang, Zhong Wang, Ying Shen
Abstract:
Recently, the multi‑modal fusion of RGB, depth, and semantics has shown great potential in dense Simultaneous Localization and Mapping (SLAM). However, a prerequisite for generating consistent semantic maps is the availability of dense, efficient, and scalable scene representations. Existing semantic SLAM systems based on explicit representations are often limited by resolution and an inability to predict unknown areas. Conversely, implicit representations typically rely on time‑consuming ray tracing, failing to meet real‑time requirements. Fortunately, 3D Gaussian Splatting (3DGS) has emerged as a promising representation that combines the efficiency of point‑based methods with the continuity of geometric structures. To this end, we propose GS3LAM, a Gaussian Semantic Splatting SLAM framework that processes multimodal data to render consistent, dense semantic maps in real‑time. GS3LAM models the scene as a Semantic Gaussian Field (SG‑Field) and jointly optimizes camera poses and the field via multimodal error constraints. Furthermore, a Depth‑adaptive Scale Regularization (DSR) scheme is introduced to resolve misalignments between scale‑invariant Gaussians and geometric surfaces. To mitigate catastrophic forgetting, we propose a Random Sampling‑based Keyframe Mapping (RSKM) strategy, which demonstrates superior performance over common local covisibility optimization methods. Extensive experiments on benchmark datasets show that GS3LAM achieves increased tracking robustness, superior rendering quality, and enhanced semantic precision compared to state‑of‑the‑art methods. Source code is available at https://github.com/lif314/GS3LAM.
Authors:Junwei Zheng, Ruize Dai, Ruiping Liu, Zichao Zeng, Yufan Chen, Fangjinhua Wang, Kunyu Peng, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen
Abstract:
Metric Cross‑View Geo‑Localization (MCVGL) aims to estimate the 3‑DoF camera pose (position and heading) by matching ground and satellite images. In this work, instead of pinhole and satellite images, we study robust MCVGL using holistic panoramas and OpenStreetMap (OSM). To this end, we establish a large‑scale MCVGL benchmark dataset, CV‑RHO, with over 2.7M images under different weather and lighting conditions, as well as sensor noise. Furthermore, we propose a model termed RHO with a two‑branch Pin‑Pan architecture for accurate visual localization. A Split‑Undistort‑Merge (SUM) module is introduced to address the panoramic distortion, and a Position‑Orientation Fusion (POF) mechanism is designed to enhance the localization accuracy. Extensive experiments prove the value of our CV‑RHO dataset and the effectiveness of the RHO model, with a significant performance gain up to 20% compared with the state‑of‑the‑art baselines. Project page: https://github.com/InSAI‑Lab/RHO.
Authors:Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng
Abstract:
This work introduces a new approach to automatic oil painting that emphasizes the creation of dynamic and expressive brushstrokes. A pivotal challenge lies in mitigating the duplicate and common‑place strokes, which often lead to less aesthetic outcomes. Inspired by the human painting process, \ie, observing, comparing, and drawing, we incorporate differential image analysis into a neural oil painting model, allowing the model to effectively concentrate on the incremental impact of successive brushstrokes. To operationalize this concept, we propose the Differential Query Transformer (DQ‑Transformer), a new architecture that leverages differentially derived image representations enriched with positional encoding to guide the stroke prediction process. This integration enables the model to maintain heightened sensitivity to local details, resulting in more refined and nuanced stroke generation. Furthermore, we incorporate adversarial training into our framework, enhancing the accuracy of stroke prediction and thereby improving the overall realism and fidelity of the synthesized paintings. Extensive qualitative evaluations, complemented by a controlled user study, validate that our DQ‑Transformer surpasses existing methods in both visual realism and artistic authenticity, typically achieving these results with fewer strokes. The stroke‑by‑stroke painting animations are available on our project website.
Authors:Xinying Lin, Xuyang Liu, Yiyu Wang, Teng Ma, Wenqi Ren
Abstract:
Video large language models (VideoLLMs) show strong capability in video understanding, yet long‑context inference is still dominated by massive redundant visual tokens in the prefill stage. We revisit token compression for VideoLLMs under a tight budget and identify a key bottleneck, namely insufficient spatio‑temporal information coverage. Existing methods often introduce discontinuous coverage through coarse per‑frame allocation or scene segmentation, and token merging can further misalign spatio‑temporal coordinates under MRoPE‑style discrete (t,h,w) bindings. To address these issues, we propose V‑CAST (Video Curvature‑Aware Spatio‑Temporal Pruning), a training‑free, plug‑and‑play pruning policy for long‑context video inference. V‑CAST casts token compression as a trajectory approximation problem and introduces a curvature‑guided temporal allocation module that routes per‑frame token budgets to semantic turns and event boundaries. It further adopts a dual‑anchor spatial selection mechanism that preserves high‑entropy visual evidence without attention intervention, while keeping retained tokens at their original coordinates to maintain positional alignment. Extensive experiments across multiple VideoLLMs of different architectures and scales demonstrate that V‑CAST achieves 98.6% of the original performance, outperforms the second‑best method by +1.1% on average, and reduces peak memory and total latency to 86.7% and 86.4% of vanilla Qwen3‑VL‑8B‑Instruct.
Authors:Yixing Zhu, Qing Zhang, Wenju Xu, Wei-Shi Zheng
Abstract:
We present YOEO, an approach for object erasure. Unlike recent diffusion‑based methods which struggle to erase target objects without generating unexpected content within the masked regions due to lack of sufficient paired training data and explicit constraint on content generation, our method allows to produce high‑quality object erasure results free of unwanted objects or artifacts while faithfully preserving the overall context coherence to the surrounding content. We achieve this goal by training an object erasure diffusion model on unpaired data containing only large‑scale real‑world images, under the supervision of a sundries detector and a context coherence loss that are built upon an entity segmentation model. To enable more efficient training and inference, a diffusion distillation strategy is employed to train for a few‑step erasure diffusion model. Extensive experiments show that our method outperforms the state‑of‑the‑art object erasure methods. Code will be available at https://zyxunh.github.io/YOEO‑ProjectPage/.
Authors:Junho Kim, Hosu Lee, James M. Rehg, Minsu Kim, Yong Man Ro
Abstract:
Recent progress in video large language models (Video‑LLMs) has enabled strong offline reasoning over long and complex videos. However, real‑world deployments increasingly require streaming perception and proactive interaction, where video frames arrive online and the system must decide not only what to respond, but also when to respond. In this work, we revisit proactive activation in streaming video as a structured sequence modeling problem, motivated by the observation that temporal transitions in streaming video naturally form span‑structured activation patterns. To capture this span‑level structure, we model activation signals jointly over a sliding temporal window and update them iteratively as new frames arrive. We propose STRIDE (Structured Temporal Refinement with Iterative DEnoising), which employs a lightweight masked diffusion module at the activation interface to jointly predict and progressively refine activation signals across the window. Extensive experiments on diverse streaming benchmarks and downstream models demonstrate that STRIDE shows more reliable and temporally coherent proactive responses, significantly improving when‑to‑speak decision quality in online streaming scenarios.
Authors:Dinh-Khoi Vo, Van-Loc Nguyen, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
Abstract:
Removing objects from natural images is challenging due to difficulty of synthesizing semantically coherent content while preserving background integrity. Existing methods often rely on fine‑tuning, prompt engineering, or inference‑time optimization, yet still suffer from texture inconsistency, rigid artifacts, weak foreground‑background disentanglement, and poor scalability for multi‑object removal. We propose a novel zero‑shot object removal framework, namely PANDORA, that operates directly on pre‑trained text‑to‑image diffusion models, requiring no fine‑tuning, prompts, or optimization. We propose Pixel‑wise Attention Dissolution to remove object by nullifying the most correlated attention keys for masked pixels, effectively eliminating the object from self‑attention flow and allowing background context to dominate reconstruction. We further introduce Localized Attentional Disentanglement Guidance to steer denoising toward latent manifolds favorable to clean object removal. Together, these components enable precise, non‑rigid, prompt‑free, and scalable multi‑object erasure in a single pass. Experiments demonstrate superior visual fidelity and semantic plausibility compared to state‑of‑the‑art methods. The project page is available at https://vdkhoi20.github.io/PANDORA.
Authors:Jongmin Lee, Seungyeop Kang, Sungjoo Yoo
Abstract:
Establishing consistent correspondences across images is essential for 3D vision tasks such as structure‑from‑motion (SfM), yet most existing matchers operate in a pairwise manner, often producing fragmented and geometrically inconsistent tracks when their predictions are chained across views. We propose MV‑RoMa, a multi‑view dense matching model that jointly estimates dense correspondences from a source image to multiple co‑visible targets. Specifically, we design an efficient model architecture which avoids high computational cost of full cross‑attention for multi‑view feature interaction: (i) multi‑view encoder that leverages pair‑wise matching results as a geometric prior, and (ii) multi‑view matching refiner that refines correspondences using pixel‑wise attention. Additionally, we propose a post‑processing strategy that integrates our model's consistent multi‑view correspondences as high‑quality tracks for SfM. Across diverse and challenging benchmarks, MV‑RoMa produces more reliable correspondences and substantially denser, more accurate 3D reconstructions than existing sparse and dense matching methods. Project page: https://icetea‑cv.github.io/mv‑roma/.
Authors:Meituan LongCat Team, Bin Xiao, Chao Wang, Chengjiang Li, Chi Zhang, Chong Peng, Hang Yu, Hao Yang, Haonan Yan, Haoze Sun, Haozhe Zhao, Hong Liu, Hui Su, Jiaqi Zhang, Jiawei Wang, Jing Li, Kefeng Zhang, Manyuan Zhang, Minhao Jing, Peng Pei, Quan Chen, Taofeng Xue, Tongxin Pan, Xiaotong Li, Xiaoyang Li, Xiaoyu Zhao, Xing Hu, Xinyang Lin, Xunliang Cai, Yan Bai, Yan Feng, Yanjie Li, Yao Qiu, Yerui Sun, Yifan Lu, Ying Luo, Yipeng Mei, Yitian Chen, Yuchen Xie, Yufang Liu, Yufei Chen, Yulei Qian, Yuqi Peng, Zhihang Yu, Zhixiong Han, Changran Wang, Chen Chen, Dian Zheng, Fengjiao Chen, Ge Yang, Haowei Guo, Haozhe Wang, Hongyu Li, Huicheng Jiang, Jiale Hong, Jialv Zou, Jiamu Li, Jianping Lin, Jiaxing Liu, Jie Yang, Jing Jin, Jun Kuang, Juncheng She, Kunming Luo, Kuofeng Gao, Lin Qiu, Linsen Guo, Mianqiu Huang, Qi Li, Qian Wang, Rumei Li, Siyu Ren, Wei Wang, Wenlong He, Xi Chen, Xiao Liu, Xiaoyu Li, Xu Huang, Xuanyu Zhu, Xuezhi Cao, Yaoming Zhu, Yifei Cao, Yimeng Jia, Yizhen Jiang, Yufei Gao, Zeyang Hu, Zhenlong Yuan, Zijian Zhang, Ziwen Wang
Abstract:
The prevailing Next‑Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal systems remain language‑centric, often treating non‑linguistic modalities as external attachments, leading to fragmented architectures and suboptimal integration. To transcend this limitation, we introduce Discrete Native Autoregressive (DiNA), a unified framework that represents multimodal information within a shared discrete space, enabling a consistent and principled autoregressive modeling across modalities. A key innovation is the Discrete Native Any‑resolution Visual Transformer (dNaViT), which performs tokenization and de‑tokenization at arbitrary resolutions, transforming continuous visual signals into hierarchical discrete tokens. Building on this foundation, we develop LongCat‑Next, a native multimodal model that processes text, vision, and audio under a single autoregressive objective with minimal modality‑specific design. As an industrial‑strength foundation model, it excels at seeing, painting, and talking within a single framework, achieving strong performance across a wide range of multimodal benchmarks. In particular, LongCat‑Next addresses the long‑standing performance ceiling of discrete vision modeling on understanding tasks and provides a unified approach to effectively reconcile the conflict between understanding and generation. As an attempt toward native multimodality, we open‑source the LongCat‑Next and its tokenizers, hoping to foster further research and development in the community. GitHub: https://github.com/meituan‑longcat/LongCat‑Next
Authors:Xulu Zhang, Haoqian Du, Xiaoyong Wei, Qing Li
Abstract:
Lineart colorization is a critical stage in professional content creation, yet achieving precise and flexible results under diverse user constraints remains a significant challenge. To address this, we propose OmniColor, a unified framework for multi‑modal lineart colorization that supports arbitrary combinations of control signals. Specifically, we systematically categorize guidance signals into two types: spatially‑aligned conditions and semantic‑reference conditions. For spatially‑aligned inputs, we employ a dual‑path encoding strategy paired with a Dense Feature Alignment loss to ensure rigorous boundary preservation and precise color restoration. For semantic‑reference inputs, we utilize a VLM‑only encoding scheme integrated with a Temporal Redundancy Elimination mechanism to filter repetitive information and enhance inference efficiency. To resolve potential input conflicts, we introduce an Adaptive Spatial‑Semantic Gating module that dynamically balances multi‑modal constraints. Experimental results demonstrate that OmniColor achieves superior controllability, visual quality, and temporal stability, providing a robust and practical solution for lineart colorization. The source code and dataset will be open at https://github.com/zhangxulu1996/OmniColor.
Authors:Shuai Xiang, Wei Guo, James Burridge, Shouyang Liu, Hao Lu, Tokihiro Fukatsu
Abstract:
Vision Foundation Models (VFM) pre‑trained on large‑scale unlabeled data have achieved remarkable success on general computer vision tasks, yet typically suffer from significant domain gaps when applied to agriculture. In this context, we introduce SPROUT (Scalable Plant Representation model via Open‑field Unsupervised Training), a multi‑crop, multi‑task agricultural foundation model trained via diffusion denoising. SPROUT leverages a VAE‑free Pixel‑space Diffusion Transformer to learn rich, structure‑aware representations through denoising and enabling efficient end‑to‑end training. We pre‑train SPROUT on a curated dataset of 2.6 million high‑quality agricultural images spanning diverse crops, growth stages, and environments. Extensive experiments demonstrate that SPROUT consistently outperforms state‑of‑the‑art web‑pretrained and agricultural foundation models across a wide range of downstream tasks, while requiring substantially lower pre‑training cost. The code and model are available at https://github.com/UTokyo‑FieldPhenomics‑Lab/SPROUT.
Authors:Jiahao Niu, Rongjia Zheng, Wenju Xu, Wei-Shi Zheng, Qing Zhang
Abstract:
We present SGS‑Intrinsic, an indoor inverse rendering framework that works well for sparse‑view images. Unlike existing 3D Gaussian Splatting (3DGS) based methods that focus on object‑centric reconstruction and fail to work under sparse view settings, our method allows to achieve high‑quality geometry reconstruction and accurate disentanglement of material and illumination. The core idea is to construct a dense and geometry‑consistent Gaussian semantic field guided by semantic and geometric priors, providing a reliable foundation for subsequent inverse rendering. Building upon this, we perform material‑illumination disentanglement by combining a hybrid illumination model and material prior to effectively capture illumination‑material interactions. To mitigate the impact of cast shadows and enhance the robustness of material recovery, we introduce illumination‑invariant material constraint together with a deshadowing model. Extensive experiments on benchmark datasets show that our method consistently improves both reconstruction fidelity and inverse rendering quality over existing 3DGS‑based inverse rendering approaches. Our code is available at https://github.com/GrumpySloths/SGS_Intrinsic.github.io.
Authors:Chang Sun, Dongliang Liao, Changxing Ding
Abstract:
Open‑vocabulary human‑object interaction (HOI) detection aims to localize and recognize all human‑object interactions in an image, including those unseen during training. Existing approaches usually rely on the collaboration between a conventional HOI detector and a Vision‑Language Model (VLM) to recognize unseen HOI categories. However, feature fusion in this paradigm is challenging due to significant gaps in cross‑model representations. To address this issue, we introduce SL‑HOI, a StreamLined open‑vocabulary HOI detection framework based solely on the powerful DINOv3 model. Our design leverages the complementary strengths of DINOv3's components: its backbone for fine‑grained localization and its text‑aligned vision head for open‑vocabulary interaction classification. Moreover, to facilitate smooth cross‑attention between the interaction queries and the vision head's output, we propose first feeding both the interaction queries and the backbone image tokens into the vision head, effectively bridging their representation gaps. All DINOv3 parameters in our approach are frozen, with only a small number of learnable parameters added, allowing a fast adaptation to the HOI detection task. Extensive experiments show that SL‑HOI achieves state‑of‑the‑art performance on both the SWiG‑HOI and HICO‑DET benchmarks, demonstrating the effectiveness of our streamlined model architecture. Code is available at https://github.com/MPI‑Lab/SL‑HOI.
Authors:Xuanpu Zhao, Zhentao Tan, Dianmo Sheng, Tianxiang Chen, Yao Liu, Yue Wu, Tao Gong, Qi Chu, Nenghai Yu
Abstract:
To enhance the perception and reasoning capabilities of multimodal large language models in complex visual scenes, recent research has introduced agent‑based workflows. In these works, MLLMs autonomously utilize image cropping tool to analyze regions of interest for question answering. While existing training strategies, such as those employing supervised fine‑tuning and reinforcement learning, have made significant progress, our empirical analysis reveals a key limitation. We demonstrate the model's strong reliance on global input and its weak dependence on the details within the cropped region. To address this issue, we propose a novel two‑stage reinforcement learning framework that does not require trajectory supervision. In the first stage, we introduce the ``Information Gap" mechanism by adjusting the granularity of the global image. This mechanism trains the model to answer questions by focusing on cropped key regions, driven by the information gain these regions provide. The second stage further enhances cropping precision by incorporating a grounding loss, using a small number of bounding box annotations. Experiments show that our method significantly enhances the model's attention to cropped regions, enabling it to achieve state‑of‑the‑art performance on high‑resolution visual question‑answering benchmarks. Our method provides a more efficient approach for perceiving and reasoning fine‑grained details in MLLMs. Code is available at: https://github.com/XuanPu‑Z/LFPC.
Authors:Zhongying Deng, Cheng Tang, Ziyan Huang, Jiashi Lin, Ying Chen, Junzhi Ning, Chenglong Ma, Jiyao Liu, Wei Li, Yinghao Zhu, Shujian Gao, Yanyan Huang, Sibo Ju, Yanzhou Su, Pengcheng Chen, Wenhao Tang, Tianbin Li, Haoyu Wang, Yuanfeng Ji, Hui Sun, Shaobo Min, Liang Peng, Feilong Tang, Haochen Xue, Rulin Zhou, Chaoyang Zhang, Wenjie Li, Shaohao Rui, Weijie Ma, Xingyue Zhao, Yibin Wang, Kun Yuan, Zhaohui Lu, Shujun Wang, Jinjie Wei, Lihao Liu, Dingkang Yang, Lin Wang, Yulong Li, Haolin Yang, Yiqing Shen, Lequan Yu, Xiaowei Hu, Yun Gu, Yicheng Wu, Benyou Wang, Minghui Zhang, Angelica I. Aviles-Rivero, Qi Gao, Hongming Shan, Xiaoyu Ren, Fang Yan, Hongyu Zhou, Haodong Duan, Maosong Cao, Shanshan Wang, Bin Fu, Xiaomeng Li, Zhi Hou, Chunfeng Song, Lei Bai, Yuan Cheng, Yuandong Pu, Xiang Li, Wenhai Wang, Hao Chen, Jiaxin Zhuang, Songyang Zhang, Huiguang He, Mengzhang Li, Bohan Zhuang, Zhian Bai, Rongshan Yu, Liansheng Wang, Yukun Zhou, Xiaosong Wang, Xin Guo, Guanbin Li, Xiangru Lin, Dakai Jin, Mianxin Liu, Wenlong Zhang, Qi Qin, Conghui He, Yuqiang Li, Ye Luo, Nanqing Dong, Jie Xu, Wenqi Shao, Bo Zhang, Qiujuan Yan, Yihao Liu, Jun Ma, Zhi Lu, Yuewen Cao, Zongwei Zhou, Jianming Liang, Shixiang Tang, Qi Duan, Dongzhan Zhou, Chen Jiang, Yuyin Zhou, Yanwu Xu, Jiancheng Yang, Shaoting Zhang, Xiaohong Liu, Siqi Luo, Yi Xin, Chaoyu Liu, Haochen Wen, Xin Chen, Alejandro Lozano, Min Woo Sun, Yuhui Zhang, Yue Yao, Xiaoxiao Sun, Serena Yeung-Levy, Xia Li, Jing Ke, Chunhui Zhang, Zongyuan Ge, Ming Hu, Jin Ye, Zhifeng Li, Yirong Chen, Yu Qiao, Junjun He
Abstract:
Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large‑scale, diverse, and high‑quality datasets. However, in the field of medical imaging, the curation and assembling of such medical datasets are highly challenging due to the reliance on clinical expertise and strict ethical and privacy constraints, resulting in a scarcity of large‑scale unified medical datasets and hindering the development of powerful medical foundation models. In this work, we present the largest survey to date of medical image datasets, covering over 1,000 open‑access datasets with a systematic catalog of their modalities, tasks, anatomies, annotations, limitations, and potential for integration. Our analysis exposes a landscape that is modest in scale, fragmented across narrowly scoped tasks, and unevenly distributed across organs and modalities, which in turn limits the utility of existing medical image datasets for developing versatile and robust medical foundation models. To turn fragmentation into scale, we propose a metadata‑driven fusion paradigm (MDFP) that integrates public datasets with shared modalities or tasks, thereby transforming multiple small data silos into larger, more coherent resources. Building on MDFP, we release an interactive discovery portal that enables end‑to‑end, automated medical image dataset integration, and compile all surveyed datasets into a unified, structured table that clearly summarizes their key characteristics and provides reference links, offering the community an accessible and comprehensive repository. By charting the current terrain and offering a principled path to dataset consolidation, our survey provides a practical roadmap for scaling medical imaging corpora, supporting faster data discovery, more principled dataset creation, and more capable medical foundation models.
Authors:Nazia Tasnim, Shrimai Prabhumoye, Bryan A. Plummer
Abstract:
Parameter Recombination (PR) methods aim to efficiently compose the weights of a neural network for applications like Parameter‑Efficient FineTuning (PEFT) and Model Compression (MC), among others. Most methods typically focus on one application of PR, which can make composing them challenging. For example, when deploying a large model you may wish to compress the model and also quickly adapt to new settings. However, PEFT methods often can still contain millions of parameters. This may be small compared to the original model size, but can be problematic in resource constrained deployments like edge devices, where they take a larger portion of the compressed model's parameters. To address this, we present Coefficient‑gated weight Recombination by Interpolated Shared basis Projections (CRISP), a general approach that seamlessly integrates multiple PR tasks within the same framework. CRISP accomplishes this by factorizing pretrained weights into basis matrices and their component mixing projections. Sharing basis matrices across layers and adjusting its size enables us to perform MC, whereas the mixer weight's small size (fewer than 200 in some experiments) enables CRISP to support PEFT. Experiments show CRISP outperforms methods from prior work capable of dual‑task applications by 4‑5% while also outperforming the state‑of‑the‑art in PEFT by 1.5% and PEFT+MC combinations by 1%. Our code is available on the repository: https://github.com/appledora/CRISP‑CVPR26.
Authors:Yuhang Han, Yuyang Wu, Zhengbo Jiao, Yiyu Wang, Xuyang Liu, Shaobo Wang, Hanlin Xu, Xuming Hu, Linfeng Zhang
Abstract:
Reinforcement Learning from Verifiable Rewards (RLVR) has substantially enhanced the reasoning capabilities of large language models in abstract reasoning tasks. However, its application to Large Vision‑Language Models (LVLMs) remains constrained by a structural representational bottleneck. Existing approaches generally lack explicit modeling and effective utilization of visual information, preventing visual representations from being tightly coupled with the reinforcement learning optimization process and thereby limiting further improvements in multimodal reasoning performance. To address this limitation, we propose KAWHI (Key‑Region Aligned Weighted Harmonic Incentive), a plug‑and‑play reward reweighting mechanism that explicitly incorporates structured visual information into uniform reward policy optimization methods (e.g., GRPO and GSPO). The method adaptively localizes semantically salient regions through hierarchical geometric aggregation, identifies vision‑critical attention heads via structured attribution, and performs paragraph‑level credit reallocation to align spatial visual evidence with semantically decisive reasoning steps. Extensive empirical evaluations on diverse reasoning benchmarks substantiate KAWHI as a general‑purpose enhancement module, consistently improving the performance of various uniform reward optimization methods. Project page: KAWHI (https://kawhiiiileo.github.io/KAWHI_PAGE/)
Authors:Ke Li, Tianjia Yang, Kaidi Liang, Xianbiao Hu, Ruwen Qin
Abstract:
Video prediction is a useful function for autonomous driving, enabling intelligent vehicles to reliably anticipate how driving scenes will evolve and thereby supporting reasoning and safer planning. However, existing models are constrained by multi‑stage training pipelines and remain insufficient in modeling the diverse motion patterns in real driving scenes, leading to degraded temporal consistency and visual quality. To address these challenges, this paper introduces the historical motion priors‑informed diffusion model (HMPDM), a video prediction model that leverages historical motion priors to enhance motion understanding and temporal coherence. The proposed deep learning system introduces three key designs: (i) a Temporal‑aware Latent Conditioning (TaLC) module for implicit historical motion injection; (ii) a Motion‑aware Pyramid Encoder (MaPE) for multi‑scale motion representation; (iii) a Self‑Conditioning (SC) strategy for stable iterative denoising. Extensive experiments on the Cityscapes and KITTI benchmarks demonstrate that HMPDM outperforms state‑of‑the‑art video prediction methods with efficiency, achieving a 28.2% improvement in FVD on Cityscapes under the same monocular RGB input configuration setting. The implementation codes are publicly available at https://github.com/KELISBU/HMPDM.
Authors:Amartya Bhattacharya
Abstract:
Vision‑language models (VLMs) excel at image‑text retrieval yet persistently fail at compositional reasoning, distinguishing captions that share the same words but differ in relational structure. We present, a unified evaluation and augmentation framework benchmarking four architecturally diverse VLMs,CLIP, BLIP, LLaVA, and Qwen3‑VL‑8B‑Thinking,on the Winoground benchmark under plain and scene‑graph‑augmented regimes. We introduce a dependency‑based TextSceneGraphParser (spaCy) extracting subject‑relation‑object triples, and a Graph Asymmetry Scorer using optimal bipartite matching to inject structural relational priors. Caption ablation experiments (subject‑object masking and swapping) reveal that Qwen3‑VL‑8B‑Thinking achieves a group score of 62.75, far above all encoder‑based models, while a proposed multi‑turn SG filtering strategy further lifts it to 66.0, surpassing prior open‑source state‑of‑the‑art. We analyze the capability augmentation tradeoff and find that SG augmentation benefits already capable models while providing negligible or negative gains for weaker baselines. Code: https://github.com/amartyacodes/Inference‑Time‑Structural‑Reasoning‑for‑Compositional‑Vision‑Language‑Understanding
Authors:Ji-Xuan He, Jia-Cheng Zhao, Feng-Qi Cui, Jinyang Huang, Yang Liu, Sirui Zhao, Meng Li, Zhi Liu
Abstract:
Low‑light image super‑resolution (LLISR) is essential for restoring fine visual details and perceptual quality under insufficient illumination conditions with ubiquitous low‑resolution devices. Although pioneer methods achieve high performance on single tasks, they solve both tasks in a serial manner, which inevitably leads to artifact amplification, texture suppression, and structural degradation. To address this, we propose Decoupling then Perceive (DTP), a novel frequency‑aware framework that explicitly separates luminance and texture into semantically independent components, enabling specialized modeling and coherent reconstruction. Specifically, to adaptively separate the input into low‑frequency luminance and high‑frequency texture subspaces, we propose a Frequency‑aware Structural Decoupling (FSD) mechanism, which lays a solid foundation for targeted representation learning and reconstruction. Based on the decoupled representation, a Semantics‑specific Dual‑path Representation (SDR) learning strategy that performs targeted enhancement and reconstruction for each frequency component is further designed, facilitating robust luminance adjustment and fine‑grained texture recovery. To promote structural consistency and perceptual alignment in the reconstructed output, building upon this dual‑path modeling, we further introduce a Cross‑frequency Semantic Recomposition (CSR) module that selectively integrates the decoupled representations. Extensive experiments on the most widely used LLISR benchmarks demonstrate the superiority of our DTP framework, improving +1.6% PSNR, +9.6% SSIM, and ‑48% LPIPS compared to the most state‑of‑the‑art (SOTA) algorithm. Codes are released at https://github.com/JXVision/DTP.
Authors:Renaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard Ghanem
Abstract:
Masked video modeling (MVM) has emerged as a simple and scalable self‑supervised pretraining paradigm, but only encodes motion information implicitly, limiting the encoding of temporal dynamics in the learned representations. As a result, such models struggle on motion‑centric tasks that require fine‑grained motion awareness. To address this, we propose TrackMAE, a simple masked video modeling paradigm that explicitly uses motion information as a reconstruction signal. In TrackMAE, we use an off‑the‑shelf point tracker to sparsely track points in the input videos, generating motion trajectories. Furthermore, we exploit the extracted trajectories to improve random tube masking with a motion‑aware masking strategy. We enhance video representations learned in both pixel and feature semantic reconstruction spaces by providing a complementary supervision signal in the form of motion targets. We evaluate on six datasets across diverse downstream settings and find that TrackMAE consistently outperforms state‑of‑the‑art video self‑supervised learning baselines, learning more discriminative and generalizable representations. Code available at https://github.com/rvandeghen/TrackMAE
Authors:Yanying Li, Jinyang Li, Shengfeng He, Yangyang Xu, Junyu Dong, Yong Du
Abstract:
We present NimbusGS, a unified framework for reconstructing high‑quality 3D scenes from degraded multi‑view inputs captured under diverse and mixed adverse weather conditions. Unlike existing methods that target specific weather types, NimbusGS addresses the broader challenge of generalization by modeling the dual nature of weather: a continuous, view‑consistent medium that attenuates light, and dynamic, view‑dependent particles that cause scattering and occlusion. To capture this structure, we decompose degradations into a global transmission field and per‑view particulate residuals. The transmission field represents static atmospheric effects shared across views, while the residuals model transient disturbances unique to each input. To enable stable geometry learning under severe visibility degradation, we introduce a geometry‑guided gradient scaling mechanism that mitigates gradient imbalance during the self‑supervised optimization of 3D Gaussian representations. This physically grounded formulation allows NimbusGS to disentangle complex degradations while preserving scene structure, yielding superior geometry reconstruction and outperforming task‑specific methods across diverse and challenging weather conditions. Code is available at https://github.com/lyy‑ovo/NimbusGS.
Authors:Ji Ma, Wei Suo, Peng Wang, Yanning Zhang
Abstract:
Multimodal Chain‑of‑Thought (MCoT) models have demonstrated impressive capability in complex visual reasoning tasks. Unfortunately, recent studies reveal that they suffer from severe hallucination problems due to diminished visual attention during the generation process. However, visual attention decay is a well‑studied problem in Large Vision‑Language Models (LVLMs). Considering the fundamental differences in reasoning processes between MCoT models and traditional LVLMs, we raise a basic question: Whether MCoT models have unique causes of hallucinations? To answer this question, we systematically investigate the hallucination patterns of MCoT models and find that fabricated texts are primarily generated in associative reasoning steps, which we term divergent thinking. Leveraging these insights, we introduce a simple yet effective strategy that can effectively localize divergent thinking steps and intervene in the decoding process to mitigate hallucinations. Extensive experiments show that our method outperforms existing methods by a large margin. More importantly, our proposed method can be conveniently integrated with other hallucination mitigation methods and further boost their performance. The code is publicly available at https://github.com/ASGO‑MM/MCoT‑hallucination.
Authors:Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil, Ponnurangam Kumaraguru, Aishwarya Agrawal
Abstract:
Humans build shared spatial understanding by communicating partial, viewpoint‑dependent observations. We ask whether Multimodal Large Language Models (MLLMs) can do the same, aligning distinct egocentric views through dialogue to form a coherent, allocentric mental model of a shared environment. To study this systematically, we introduce COSMIC, a benchmark for Collaborative Spatial Communication. In this setting, two static MLLM agents observe a 3D indoor environment from different viewpoints and exchange natural‑language messages to solve spatial queries. COSMIC contains 899 diverse scenes and 1250 question‑answer pairs spanning five tasks. We find a capability hierarchy, MLLMs are most reliable at identifying shared anchor objects across views, perform worse on relational reasoning, and largely fail at building globally consistent maps, performing near chance, even for frontier models. Moreover, we find thinking capability yields gains in anchor grounding, but is insufficient for higher‑level spatial communication. To contextualize model behavior, we collect 250 human‑human dialogues. Humans achieve 95% aggregate accuracy, while the best model, Gemini‑3‑Pro‑Thinking, reaches 72%, leaving substantial room for improvement. Moreover, human conversations grow more precise as partners align on a shared spatial understanding, whereas MLLMs keep exploring without converging, suggesting limited capacity to form and sustain a robust shared mental model throughout the dialogue. Our code and data is available at https://github.com/ankursikarwar/Cosmic.
Authors:Yizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang, Qingjie Liu, Yunhong Wang
Abstract:
Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to image‑level anomaly detection and textual reasoning, while pixel‑level localization still relies on external vision modules and dense annotations. In this work, we activate the intrinsic reasoning potential of MLLMs to perform anomaly detection, pixel‑level localization, and interpretable reasoning solely from image‑level supervision, without any auxiliary components or pixel‑wise labels. Specifically, we propose Reasoning‑Driven Anomaly Localization (ReAL), which extracts anomaly‑related tokens from the autoregressive reasoning process and aggregates their attention responses to produce pixel‑level anomaly maps. We further introduce a Consistency‑Guided Reasoning Optimization (CGRO) module that leverages reinforcement learning to align reasoning tokens with visual attentions, resulting in more coherent reasoning and accurate anomaly localization. Extensive experiments on four public benchmarks demonstrate that our method significantly improves anomaly detection, localization, and interpretability. Remarkably, despite relying solely on image‑level supervision, our approach achieves performance competitive with MLLM‑based methods trained under dense pixel‑level supervision. Code is available at https://github.com/YizhouJin313/ReADL.
Authors:Kenji Tojo, Bernd Bickel, Nobuyuki Umetani
Abstract:
Radiance field reconstruction aims to recover high‑quality 3D representations from multi‑view RGB images. Recent advances, such as 3D Gaussian splatting, enable real‑time rendering with high visual fidelity on sufficiently powerful graphics hardware. However, efficient online transmission and rendering across diverse platforms requires drastic model simplification, reducing the number of primitives by several orders of magnitude. We introduce DiffSoup, a radiance field representation that employs a soup (i.e., a highly unstructured set) of a small number of triangles with neural textures and binary opacity. We show that this binary opacity representation is directly differentiable via stochastic opacity masking, enabling stable training without a mollifier (i.e., smooth rasterization). DiffSoup can be rasterized using standard depth testing, enabling seamless integration into traditional graphics pipelines and interactive rendering on consumer‑grade laptops and mobile devices. Code is available at https://github.com/kenji‑tojo/diffsoup.
Authors:Sen Zhang, Runmei Li, Shizhuang Deng, Zhichao Zheng, Yuhe Zhang, Jiani Li, Kailun Zhang, Tao Zhang, Wenjun Wu, Qunbo Wang
Abstract:
As Automatic Train Operation (ATO) advances toward GoA4 and beyond, it increasingly depends on efficient, reliable cab‑view visual perception and decision‑oriented inference to ensure safe operation in complex and dynamic railway environments. However, existing approaches focus primarily on basic perception and often generalize poorly to rare yet safety‑critical corner cases. They also lack the high‑level reasoning and planning capabilities required for operational decision‑making. Although recent Large Multi‑modal Models (LMMs) show strong generalization and cognitive capabilities, their use in safety‑critical ATO is hindered by high computational cost and hallucination risk. Meanwhile, reliable domain‑specific benchmarks for systematically evaluating cognitive capabilities are still lacking. To address these gaps, we introduce RailVQA‑bench, the first VQA benchmark for cab‑view visual cognition in ATO, comprising 20,000 single‑frame and 1,168 video based QA pairs to evaluate cognitive generalization and interpretability in both static and dynamic scenarios. Furthermore, we propose RailVQA‑CoM, a collaborative large‑small model framework that combines small‑model efficiency with large‑model cognition via a transparent three‑module architecture and adaptive temporal sampling, improving perceptual generalization and enabling more efficient reasoning and planning. Experiments demonstrate that the proposed approach substantially improves performance, enhances interpretability, improves efficiency, and strengthens cross‑domain generalization in autonomous driving systems. Code and datasets will be available at https://cybereye‑bjtu.github.io/RailVQA.html.
Authors:Yizuo Peng, Xuelin Chen, Kai Zhang, Xiaodong Cun
Abstract:
Recent diffusion models have achieved remarkable success in image relighting, and this success has quickly been extended to video relighting. However, existing methods offer limited explicit control over illumination in the relighted output. We present LightCtrl, the first controllable video relighting method that enables explicit control of video illumination through a user‑supplied light trajectory in a training‑free manner. Our approach combines pre‑trained diffusion models: an image relighting model processes each frame individually, followed by a video diffusion prior to enhance temporal consistency. To achieve explicit control over dynamically varying lighting, we introduce two key components. First, a Light Map Injection module samples light trajectory‑specific noise and injects it into the latent representation of the source video, improving illumination coherence with the conditional light trajectory. Second, a Geometry‑Aware Relighting module dynamically combines RGB and normal map latents in the frequency domain to suppress the influence of the original lighting, further enhancing adherence to the input light trajectory. Experiments show that LightCtrl produces high‑quality videos with diverse illumination changes that closely follow the specified light trajectory, demonstrating improved controllability over baseline methods. Code is available at: https://github.com/GVCLab/LightCtrl.
Authors:Haoyu He, Yue Zhuo, Yu Zheng, Qi R. Wang
Abstract:
Vision‑language models (VLMs) achieve strong multimodal performance, yet how computation is organized across populations of neurons remains poorly understood. In this work, we study VLMs through the lens of neural topology, representing each layer as a within‑layer correlation graph derived from neuron‑neuron co‑activations. This view allows us to ask whether population‑level structure is behaviorally meaningful, how it changes across modalities and depth, and whether it identifies causally influential internal components under intervention. We show that correlation topology carries recoverable behavioral signal; moreover, cross‑modal structure progressively consolidates with depth around a compact set of recurrent hub neurons, whose targeted perturbation substantially alters model output. Neural topology thus emerges as a meaningful intermediate scale for VLM interpretability: richer than local attribution, more tractable than full circuit recovery, and empirically tied to multimodal behavior. Code is publicly available at https://github.com/he‑h/vlm‑graph‑probing.
Authors:Jihwan Hong, Jaeyoung Do
Abstract:
Referring Video Object Segmentation (RVOS) aims to segment target objects in videos based on natural language descriptions. However, fixed keyframe‑based approaches that couple a vision language model with a separate propagation module often fail to capture rapidly changing spatiotemporal dynamics and to handle queries requiring multi‑step reasoning, leading to sharp performance drops on motion‑intensive and reasoning‑oriented videos beyond static RVOS benchmarks. To address these limitations, we propose VIRST (Video‑Instructed Reasoning Assistant for Spatio‑Temporal Segmentation), an end‑to‑end framework that unifies global video reasoning and pixel‑level mask prediction within a single model. VIRST bridges semantic and segmentation representations through the Spatio‑Temporal Fusion (STF), which fuses segmentation‑aware video features into the vision‑language backbone, and employs the Temporal Dynamic Anchor Updater to maintain temporally adjacent anchor frames that provide stable temporal cues under large motion, occlusion, and reappearance. This unified design achieves state‑of‑the‑art results across diverse RVOS benchmarks under realistic and challenging conditions, demonstrating strong generalization to both referring and reasoning oriented settings. The code and checkpoints are available at https://github.com/AIDASLab/VIRST.
Authors:Guanhe Huang, Oya Celiktutan
Abstract:
Generative models excel at motion synthesis for a fixed number of agents but struggle to generalize with variable agents. Based on limited, domain‑specific data, existing methods employ autoregressive models to generate motion recursively, which suffer from inefficiency and error accumulation. We propose Unified Motion Flow (UMF), which consists of Pyramid Motion Flow (P‑Flow) and Semi‑Noise Motion Flow (S‑Flow). UMF decomposes the number‑free motion generation into a single‑pass motion prior generation stage and multi‑pass reaction generation stages. Specifically, UMF utilizes a unified latent space to bridge the distribution gap between heterogeneous motion datasets, enabling effective unified training. For motion prior generation, P‑Flow operates on hierarchical resolutions conditioned on different noise levels, thereby mitigating computational overheads. For reaction generation, S‑Flow learns a joint probabilistic path that adaptively performs reaction transformation and context reconstruction, alleviating error accumulation. Extensive results and user studies demonstrate UMF' s effectiveness as a generalist model for multi‑person motion generation from text. Project page: https://githubhgh.github.io/umf/.
Authors:Jiaming Li, Zhijia Liang, Weikai Chen, Lin Ma, Guanbin Li
Abstract:
Fine‑grained open‑vocabulary object detection (FG‑OVD) aims to detect novel object categories described by attribute‑rich texts. While existing open‑vocabulary detectors show promise at the base‑category level, they underperform in fine‑grained settings due to the semantic entanglement of subjects and attributes in pretrained vision‑language model (VLM) embeddings ‑‑ leading to over‑representation of attributes, mislocalization, and semantic drift in embedding space. We propose GUIDED, a decomposition framework specifically designed to address the semantic entanglement between subjects and attributes in fine‑grained prompts. By separating object localization and fine‑grained recognition into distinct pathways, HUIDED aligns each subtask with the module best suited for its respective roles. Specifically, given a fine‑grained class name, we first use a language model to extract a coarse‑grained subject and its descriptive attributes. Then the detector is guided solely by the subject embedding, ensuring stable localization unaffected by irrelevant or overrepresented attributes. To selectively retain helpful attributes, we introduce an attribute embedding fusion module that incorporates attribute information into detection queries in an attention‑based manner. This mitigates over‑representation while preserving discriminative power. Finally, a region‑level attribute discrimination module compares each detected region against full fine‑grained class names using a refined vision‑language model with a projection head for improved alignment. Extensive experiments on FG‑OVD and 3F‑OVD benchmarks show that GUIDED achieves new state‑of‑the‑art results, demonstrating the benefits of disentangled modeling and modular optimization. Our code will be released at https://github.com/lijm48/GUIDED.
Authors:Xinyu Yang, Haozheng Yu, Yihong Sun, Bharath Hariharan, Jennifer J. Sun
Abstract:
Interactive video segmentation often requires many user interventions for robust performance in challenging scenarios (e.g., occlusions, object separations, camouflage, etc.). Yet, even state‑of‑the‑art models like SAM2 use corrections only for immediate fixes without learning from this feedback, leading to inefficient, repetitive user effort. To address this, we introduce Live Interactive Training (LIT), a novel framework for prompt‑based visual systems where models also learn online from human corrections at inference time. Our primary instantiation, LIT‑LoRA, implements this by continually updating a lightweight LoRA module on‑the‑fly. When a user provides a correction, this module is rapidly trained on that feedback, allowing the vision system to improve performance on subsequent frames of the same video. Leveraging the core principles of LIT, our LIT‑LoRA implementation achieves an average 18‑34% reduction in total corrections on challenging video segmentation benchmarks, with a negligible training overhead of ~0.5s per correction. We further demonstrate its generality by successfully adapting it to other segmentation models and extending it to CLIP‑based fine‑grained image classification. Our work highlights the promise of live adaptation to transform interactive tools and significantly reduce redundant human effort in complex visual tasks. Project: https://youngxinyu1802.github.io/projects/LIT/.
Authors:Jie Zhu, Xiao Guo, Yiyang Su, Anil Jain, Xiaoming Liu
Abstract:
Model fusion is a key strategy for robust recognition in unconstrained scenarios, as different models provide complementary strengths. This is especially important for whole‑body human recognition, where biometric cues such as face, gait, and body shape vary across samples and are typically integrated via score‑fusion. However, existing score‑fusion strategies are usually static, invoking all models for every test sample regardless of sample quality or modality reliability. To overcome these limitations, we propose FusionAgent, a novel agentic framework that leverages a Multimodal Large Language Model (MLLM) to perform dynamic, sample‑specific model selection. Each expert model is treated as a tool, and through Reinforcement Fine‑Tuning (RFT) with a metric‑based reward, the agent learns to adaptively determine the optimal model combination for each test input. To address the model score misalignment and embedding heterogeneity, we introduce Anchor‑based Confidence Top‑k (ACT) score‑fusion, which anchors on the most confident model and integrates complementary predictions in a confidence‑aware manner. Extensive experiments on multiple whole‑body biometric benchmarks demonstrate that FusionAgent significantly outperforms SoTA methods while achieving higher efficiency through fewer model invocations, underscoring the critical role of dynamic, explainable, and robust model fusion in real‑world recognition systems. Project page: \hrefhttps://fusionagent.github.io/FusionAgent.
Authors:Dongsheng Yang, Yinfeng Yu, Liejun Wang
Abstract:
Vision‑and‑Language Navigation (VLN) requires an agent to navigate through complex unseen environments based on natural language instructions. However, existing methods often struggle to effectively capture key semantic cues and accurately align them with visual observations. To address this limitation, we propose Beyond Textual Knowledge (BTK), a VLN framework that synergistically integrates environment‑specific textual knowledge with generative image knowledge bases. BTK employs Qwen3‑4B to extract goal‑related phrases and utilizes Flux‑Schnell to construct two large‑scale image knowledge bases: R2R‑GP and REVERIE‑GP. Additionally, we leverage BLIP‑2 to construct a large‑scale textual knowledge base derived from panoramic views, providing environment‑specific semantic cues. These multimodal knowledge bases are effectively integrated via the Goal‑Aware Augmentor and Knowledge Augmentor, significantly enhancing semantic grounding and cross‑modal alignment. Extensive experiments on the R2R dataset with 7,189 trajectories and the REVERIE dataset with 21,702 instructions demonstrate that BTK significantly outperforms existing baselines. On the test unseen splits of R2R and REVERIE, SR increased by 5% and 2.07% respectively, and SPL increased by 4% and 3.69% respectively. The source code is available at https://github.com/yds3/IPM‑BTK/.
Authors:PengYu Chen, Shang Wan, Xiaohou Shi, Yuan Chang, Yan Sun, Sajal K. Das
Abstract:
Time series anomaly detection (TSAD) is essential for maintaining the reliability and security of IoT‑enabled service systems. Existing methods require training one specific model for each dataset, which exhibits limited generalization capability across different target datasets, hindering anomaly detection performance in various scenarios with scarce training data. To address this limitation, foundation models have emerged as a promising direction. However, existing approaches either repurpose large language models (LLMs) or construct largescale time series datasets to develop general anomaly detection foundation models, and still face challenges caused by severe cross‑modal gaps or in‑domain heterogeneity. In this paper, we investigate the applicability of large‑scale vision models to TSAD. Specifically, we adapt a visual Masked Autoencoder (MAE) pretrained on ImageNet to the TSAD task. However, directly transferring MAE to TSAD introduces two key challenges: overgeneralization and limited local perception. To address these challenges, we propose VAN‑AD, a novel MAE‑based framework for TSAD. To alleviate the over‑generalization issue, we design an Adaptive Distribution Mapping Module (ADMM), which maps the reconstruction results before and after MAE into a unified statistical space to amplify discrepancies caused by abnormal patterns. To overcome the limitation of local perception, we further develop a Normalizing Flow Module (NFM), which combines MAE with normalizing flow to estimate the probability density of the current window under the global distribution. Extensive experiments on nine real‑world datasets demonstrate that VAN‑AD consistently outperforms existing state‑of‑the‑art methods across multiple evaluation metrics.We make our code and datasets available at https://github.com/PenyChen/VAN‑AD.
Authors:Alberto G. Rodriguez Salgado
Abstract:
How do multimodal models solve visual spatial tasks ‑‑ through genuine planning, or through brute‑force search in token space? We introduce \textscMazeBench, a benchmark of 110 procedurally generated maze images across nine controlled groups, and evaluate 16 model configurations from OpenAI, Anthropic, Google, and Alibaba. GPT‑5.4 solves 91% and Gemini 3.1 Pro 79%, but these scores are misleading: models typically translate images into text grids and then enumerate paths step by step, consuming 1,710‑‑22,818 tokens per solve for a task humans do quickly. Without added reasoning budgets, all configurations score only 2‑‑12%; on 20×20 ultra‑hard mazes, they hit token limits and fail. Qualitative traces reveal a common two‑stage strategy: image‑to‑grid translation followed by token‑level search, effectively BFS in prose. A text‑grid ablation shows Claude Sonnet 4.6 rising from 6% on images to 80% when given the correct grid, isolating weak visual extraction from downstream search. When explicitly instructed not to construct a grid or perform graph search, models still revert to the same enumeration strategy. \textscMazeBench therefore shows that high accuracy on visual planning tasks does not imply human‑like spatial understanding.
Authors:Jiwen Zhang, Xiangyu Shi, Siyuan Wang, Zerui Li, Zhongyu Wei, Qi Wu
Abstract:
Vision‑and‑Language Navigation (VLN) has recently benefited from Multimodal Large Language Models (MLLMs), enabling zero‑shot navigation. While recent exploration‑based zero‑shot methods have shown promising results by leveraging global scene priors, they rely on high‑quality human‑crafted scene reconstructions, which are impractical for real‑world robot deployment. When encountering an unseen environment, a robot should build its own priors through pre‑exploration. However, these self‑built reconstructions are inevitably incomplete and noisy, which severely degrade methods that depend on high‑quality scene reconstructions. To address these issues, we propose SpatialAnt, a zero‑shot navigation framework designed to bridge the gap between imperfect self‑reconstructions and robust execution. SpatialAnt introduces a physical grounding strategy to recover the absolute metric scale for monocular‑based reconstructions. Furthermore, rather than treating the noisy self‑reconstructed scenes as absolute spatial references, we propose a novel visual anticipation mechanism. This mechanism leverages the noisy point clouds to render future observations, enabling the agent to perform counterfactual reasoning and prune paths that contradict human instructions. Extensive experiments in both simulated and real‑world environments demonstrate that SpatialAnt significantly outperforms existing zero‑shot methods. We achieve a 66% Success Rate (SR) on R2R‑CE and 50.8% SR on RxR‑CE benchmarks. Physical deployment on a Hello Robot further confirms the efficiency and efficacy of our framework, achieving a 52% SR in challenging real‑world settings.
Authors:Ling Zhang, Boxiang Yun, Ting Jin, Qingli Li, Xinxing Li, Yan Wang
Abstract:
Prediction of genetic biomarkers, e.g., microsatellite instability in colorectal cancer is crucial for clinical decision making. But, two primary challenges hamper accurate prediction: (1) It is difficult to construct a pathology‑aware representation involving the complex interconnections among pathological components. (2) WSIs contain a large proportion of areas unrelated to genetic biomarkers, which make the model easily overfit simple but irrelative instances. We hereby propose a Dictionary‑based hierarchical pathology mining with hard‑instance‑assisted classifier Debiasing framework to address these challenges, dubbed as D2Bio. Our first module, dictionary‑based hierarchical pathology mining, is able to mine diverse and very fine‑grained pathological contextual interaction without the limit to the distances between patches. The second module, hard‑instance‑assisted classfier debiasing, learns a debiased classifier via focusing on hard but task‑related features, without any additional annotations. Experimental results on five cohorts show the superiority of our method, with over 4% improvement in AUROC compared with the second best on the TCGA‑CRC‑MSI cohort. Our analysis further shows the clinical interpretability of D2Bio in genetic biomarker diagnosis and potential clinical utility in survival analysis. Code will be available at https://github.com/DeepMed‑Lab‑ECNU/D2Bio.
Authors:Charles Jones, Emmanuel Noutahi, Jason Hartford, Cian Eastwood
Abstract:
Flow‑matching generative models are increasingly used to simulate cell responses to biological perturbations. However, the design space for building such models is large and underexplored. We systematically analyse the design space of flow matching models for cell‑microscopy images, finding that many popular techniques are unnecessary and can even hurt performance. We develop a simple, stable, and scalable recipe which we use to train our foundation model. We scale our model to two orders of magnitude larger than prior methods, achieving a two‑fold FID and ten‑fold KID improvement over prior methods. We then fine‑tune our model with pre‑trained molecular embeddings to achieve state‑of‑the‑art performance simulating responses to unseen molecules.
Code is available at https://github.com/valence‑labs/microscopy‑flow‑matching
Authors:Xintao Zong, Xian Zhong, Wenxuan Liu, Jianhao Ding, Zhaofei Yu, Tiejun Huang
Abstract:
Spiking neural networks (SNNs) have recently shown strong potential in unimodal visual and textual tasks, yet building a directly trained, low‑energy, and high‑performance SNN for multimodal applications such as image‑text retrieval (ITR) remains highly challenging. Existing artificial neural network (ANN)‑based methods often pursue richer unimodal semantics using deeper and more complex architectures, while overlooking cross‑modal interaction, retrieval latency, and energy efficiency. To address these limitations, we present a brain‑inspired Cross‑Modal Spike Fusion network (CMSF) and apply it to ITR for the first time. The proposed spike fusion mechanism integrates unimodal features at the spike level, generating enhanced multimodal representations that act as soft supervisory signals to refine unimodal spike embeddings, effectively mitigating semantic loss within CMSF. Despite requiring only two time steps, CMSF achieves top‑tier retrieval accuracy, surpassing state‑of‑the‑art ANN counterparts while maintaining exceptionally low energy consumption and high retrieval speed. This work marks a significant step toward multimodal SNNs, offering a brain‑inspired framework that unifies temporal dynamics with cross‑modal alignment and provides new insights for future spiking‑based multimodal research. The code is available at https://github.com/zxt6174/CMSF.
Authors:Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Xu Zhu, Qiyu Hu, Tianyu Wang, Johnalbert Garnica, Feng Liu, Siyu Huang, Qi Dai, Zhi-Qi Cheng
Abstract:
We study language‑conditioned visual navigation (LCVN), in which an embodied agent is asked to follow a natural language instruction based only on an initial egocentric observation. Without access to goal images, the agent must rely on language to shape its perception and continuous control, making the grounding problem particularly challenging. We formulate this problem as open‑loop trajectory prediction conditioned on linguistic instructions and introduce the LCVN Dataset, a benchmark of 39,016 trajectories and 117,048 human‑verified instructions that supports reproducible research across a range of environments and instruction styles. Using this dataset, we develop LCVN frameworks that link language grounding, future‑state prediction, and action generation through two complementary model families. The first family combines LCVN‑WM, a diffusion‑based world model, with LCVN‑AC, an actor‑critic agent trained in the latent space of the world model. The second family, LCVN‑Uni, adopts an autoregressive multimodal architecture that predicts both actions and future observations. Experiments show that these families offer different advantages: the former provides more temporally coherent rollouts, whereas the latter generalizes better to unseen environments. Taken together, these observations point to the value of jointly studying language grounding, imagination, and policy learning in a unified task setting, and LCVN provides a concrete basis for further investigation of language‑conditioned world models. The code is available at https://github.com/F1y1113/LCVN.
Authors:Nicolas von Lützow, Barbara Rössle, Katharina Schmid, Matthias Nießner
Abstract:
Most recent advances in 3D generative modeling rely on diffusion or flow‑matching formulations. We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer‑based model that directly generates 3D Gaussians via next‑token prediction, thus facilitating full 3D scene generation. We first compress Gaussian primitives into a discrete latent grid using a sparse 3D convolutional autoencoder with vector quantization. The resulting tokens are serialized and modeled using a causal transformer with 3D rotary positional embedding, enabling sequential generation of spatial structure and appearance. Unlike diffusion‑based methods that refine scenes holistically, our formulation constructs scenes step‑by‑step, naturally supporting completion, outpainting, controllable sampling via temperature, and flexible generation horizons. This formulation leverages the compositional inductive biases and scalability of autoregressive modeling while operating on explicit representations compatible with modern neural rendering pipelines, positioning autoregressive transformers as a complementary paradigm for controllable and context‑aware 3D generation.
Authors:Yiming Zuo, Hongyu Wen, Venkat Subramanian, Patrick Chen, Karhan Kayan, Mario Bijelic, Felix Heide, Jia Deng
Abstract:
Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain dataset, this paper focuses on the challenging and practical setting of zero‑shot generalization. We first propose a new real‑world DfD benchmark ZEDD, which contains 8.3x more scenes and significantly higher quality images and ground‑truth depth maps compared to previous benchmarks. We also design a novel network architecture named FOSSA. FOSSA is a Transformer‑based architecture with novel designs tailored to the DfD task. The key contribution is a stack attention layer with a focus distance embedding, allowing efficient information exchange across the focus stack. Finally, we develop a new training data pipeline allowing us to utilize existing large‑scale RGBD datasets to generate synthetic focus stacks. Experiment results on ZEDD and other benchmarks show a significant improvement over the baselines, reducing errors by up to 55.7%. The ZEDD benchmark is released at https://zedd.cs.princeton.edu. The code and checkpoints are released at https://github.com/princeton‑vl/FOSSA.
Authors:Shihua Zhang, Qiuhong Shen, Shizun Wang, Tianbo Pan, Xinchao Wang
Abstract:
Empowered by large‑scale training, vision‑language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to handle this limitation by injecting geometry tokens from pretrained 3D foundation models into VLMs. Nevertheless, we observe that naive token fusion followed by standard fine‑tuning in this line of work often leaves such geometric cues underutilized for spatial reasoning, as VLMs tend to rely heavily on 2D visual cues. In this paper, we propose GeoSR, a framework designed to make geometry matter by encouraging VLMs to actively reason with geometry tokens. GeoSR introduces two key components: (1) Geometry‑Unleashing Masking, which strategically masks portions of 2D vision tokens during training to weaken non‑geometric shortcuts and force the model to consult geometry tokens for spatial reasoning; and (2) Geometry‑Guided Fusion, a gated routing mechanism that adaptively amplifies geometry token contributions in regions where geometric evidence is critical. Together, these designs unleash the potential of geometry tokens for spatial reasoning tasks. Extensive experiments on both static and dynamic spatial reasoning benchmarks demonstrate that GeoSR consistently outperforms prior methods and establishes new state‑of‑the‑art performance by effectively leveraging geometric information. The project page is available at https://suhzhang.github.io/GeoSR/.
Authors:Zhaochong An, Orest Kupyn, Théo Uscidda, Andrea Colaco, Karan Ahuja, Serge Belongie, Mar Gonzalez-Franco, Marta Tintore Gazulla
Abstract:
Large‑scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting the generator with additional modules or applying geometry‑aware alignment. However, architectural modifications can compromise the generalization of internet‑scale pretrained models, while existing alignment methods are limited to static scenes and rely on RGB‑space rewards that require repeated VAE decoding, incurring substantial compute overhead and failing to generalize to highly dynamic real‑world scenes. To preserve the pretrained capacity while improving geometric consistency, we propose VGGRPO (Visual Geometry GRPO), a latent geometry‑guided framework for geometry‑aware video post‑training. VGGRPO introduces a Latent Geometry Model (LGM) that stitches video diffusion latents to geometry foundation models, enabling direct decoding of scene geometry from the latent space. By constructing LGM from a geometry model with 4D reconstruction capability, VGGRPO naturally extends to dynamic scenes, overcoming the static‑scene limitations of prior methods. Building on this, we perform latent‑space Group Relative Policy Optimization with two complementary rewards: a camera motion smoothness reward that penalizes jittery trajectories, and a geometry reprojection consistency reward that enforces cross‑view geometric coherence. Experiments on both static and dynamic benchmarks show that VGGRPO improves camera stability, geometry consistency, and overall quality while eliminating costly VAE decoding, making latent‑space geometry‑guided reinforcement an efficient and flexible approach to world‑consistent video generation.
Authors:Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Xilin Zhao, Qingming Huang
Abstract:
Recent studies have made notable progress in video representation learning by transferring image‑pretrained models to video tasks, typically with complex temporal modules and video fine‑tuning. However, fine‑tuning heavy modules may compromise inter‑video semantic separability, i.e., the essential ability to distinguish objects across videos. While reducing the tunable parameters hinders their intra‑video temporal consistency, which is required for stable representations of the same object within a video. This dilemma indicates a potential trade‑off between the intra‑video temporal consistency and inter‑video semantic separability during image‑to‑video transfer. To this end, we propose the Consistency‑Separability Trade‑off Transfer Learning (Co‑Settle) framework, which applies a lightweight projection layer on top of the frozen image‑pretrained encoder to adjust representation space with a temporal cycle consistency objective and a semantic separability constraint. We further provide a theoretical support showing that the optimized projection yields a better trade‑off between the two properties under appropriate conditions. Experiments on eight image‑pretrained models demonstrate consistent improvements across multiple levels of video tasks with only five epochs of self‑supervised training. The code is available at https://github.com/yafeng19/Co‑Settle.
Authors:Dávid Pukanec, Tibor Kubík, Michal Španěl
Abstract:
We present ToothCraft, a diffusion‑based model for the contextual generation of tooth crowns, trained on artificially created incomplete teeth. Building upon recent advancements in conditioned diffusion models for 3D shapes, we developed a model capable of an automated tooth crown completion conditioned on local anatomical context. To address the lack of training data for this task, we designed an augmentation pipeline that generates incomplete tooth geometries from a publicly available dataset of complete dental arches (3DS, ODD). By synthesising a diverse set of training examples, our approach enables robust learning across a wide spectrum of tooth defects. Experimental results demonstrate the strong capability of our model to reconstruct complete tooth crowns, achieving an intersection over union (IoU) of 81.8% and a Chamfer Distance (CD) of 0.00034 on synthetically damaged testing restorations. Our experiments demonstrate that the model can be applied directly to real‑world cases, effectively filling in incomplete teeth, while generated crowns show minimal intersection with the opposing dentition, thus reducing the risk of occlusal interference. Access to the code, model weights, and dataset information will be available at: https://github.com/ikarus1211/VISAPP_ToothCraft
Authors:Tamir Cohen, Leo Segre, Shay Shomer-Chai, Shai Avidan, Hadar Averbuch-Elor
Abstract:
Reconstructing accurate 3D models of large‑scale real‑world scenes from unstructured, in‑the‑wild imagery remains a core challenge in computer vision, especially when the input views have little or no overlap. In such cases, existing reconstruction pipelines often produce multiple disconnected partial reconstructions or erroneously merge non‑overlapping regions into overlapping geometry. In this work, we propose a framework that grounds each partial reconstruction to a complete reference model of the scene, enabling globally consistent alignment even in the absence of visual overlap. We obtain reference models from dense, geospatially accurate pseudo‑synthetic renderings derived from Google Earth Studio. These renderings provide full scene coverage but differ substantially in appearance from real‑world photographs. Our key insight is that, despite this significant domain gap, both domains share the same underlying scene semantics. We represent the reference model using 3D Gaussian Splatting, augmenting each Gaussian with semantic features, and formulate alignment as an inverse feature‑based optimization scheme that estimates a global 6DoF pose and scale while keeping the reference model fixed. Furthermore, we introduce the WikiEarth dataset, which registers existing partial 3D reconstructions with pseudo‑synthetic reference models. We demonstrate that our approach consistently improves global alignment when initialized with various classical and learning‑based pipelines, while mitigating failure modes of state‑of‑the‑art end‑to‑end models.
Authors:Moritz Nottebaum, Matteo Dunnhofer, Christian Micheloni
Abstract:
Vision backbone networks play a central role in modern computer vision. Enhancing their efficiency directly benefits a wide range of downstream applications. To measure efficiency, many publications rely on MACs (Multiply Accumulate operations) as a predictor of execution time. In this paper, we experimentally demonstrate the shortcomings of such a metric, especially in the context of edge devices. By contrasting the MAC count and execution time of common architectural design elements, we identify key factors for efficient execution and provide insights to optimize backbone design. Based on these insights, we present LowFormer, a novel vision backbone family. LowFormer features a streamlined macro and micro design that includes Lowtention, a lightweight alternative to Multi‑Head Self‑Attention. Lowtention not only proves more efficient, but also enables superior results on ImageNet. Additionally, we present an edge GPU version of LowFormer, that can further improve upon its baseline's speed on edge GPU and desktop GPU. We demonstrate LowFormer's wide applicability by evaluating it on smaller image classification datasets, as well as adapting it to several downstream tasks, such as object detection, semantic segmentation, image retrieval, and visual object tracking. LowFormer models consistently achieve remarkable speed‑ups across various hardware platforms compared to recent state‑of‑the‑art backbones. Code and models are available at https://github.com/altair199797/LowFormer/blob/main/Beyond_MACs.md.
Authors:Tianyu Liu, Weitao Xiong, Kunming Luo, Manyuan Zhang, Peng Li, Yuan Liu, Ping Tan
Abstract:
Generative video models have significantly advanced the photorealistic synthesis of adverse weather for autonomous driving; however, they consistently demand massive datasets to learn rare weather scenarios. While 3D‑aware editing methods alleviate these data constraints by augmenting existing video footage, they are fundamentally bottlenecked by costly per‑scene optimization and suffer from inherent geometric and illumination entanglement. In this work, we introduce AutoWeather4D, a feed‑forward 3D‑aware weather editing framework designed to explicitly decouple geometry and illumination. At the core of our approach is a G‑buffer Dual‑pass Editing mechanism. The Geometry Pass leverages explicit structural foundations to enable surface‑anchored physical interactions, while the Light Pass analytically resolves light transport, accumulating the contributions of local illuminants into the global illumination to enable dynamic 3D local relighting. Extensive experiments demonstrate that AutoWeather4D achieves comparable photorealism and structural consistency to generative baselines while enabling fine‑grained parametric physical control, serving as a practical data engine for autonomous driving.
Authors:Martin Rath, Morteza Ghahremani, Yitong Li, Ashkan Taghipour, Marcus Makowski, Christian Wachinger
Abstract:
Computed tomography (CT) provides rich 3D anatomical details but is often constrained by high radiation exposure, substantial costs, and limited availability. While standard chest X‑rays are cost‑effective and widely accessible, they only provide 2D projections with limited pathological information. Reconstructing 3D CT volumes from 2D X‑rays offers a transformative solution to increase diagnostic accessibility, yet existing methods predominantly rely on synthetic X‑ray projections, limiting clinical generalization. In this work, we propose AXON, a multi‑stage diffusion‑based framework that reconstructs high‑fidelity 3D CT volumes directly from real X‑rays. AXON employs a coarse‑to‑fine strategy, with a Brownian Bridge diffusion model‑based initial stage for global structural synthesis, followed by a ControlNet‑based refinement stage for local intensity optimization. It also supports bi‑planar X‑ray input to mitigate depth ambiguities inherent in 2D‑to‑3D reconstruction. A super‑resolution network is integrated to upscale the generated volumes to achieve diagnostic‑grade resolution. Evaluations on both public and external datasets demonstrate that AXON significantly outperforms state‑of‑the‑art baselines, achieving a 11.9% improvement in PSNR and a 11.0% increase in SSIM with robust generalizability across disparate clinical distributions. Our code is available at https://github.com/ai‑med/AXON.
Authors:Weihong Pan, Xiaoyu Zhang, Zhuang Zhang, Zhichao Ye, Nan Wang, Haomin Liu, Guofeng Zhang
Abstract:
High‑quality 4D reconstruction enables photorealistic and immersive rendering of the dynamic real world. However, unlike static scenes that can be fully captured with a single camera, high‑quality dynamic scenes typically require dense arrays of tens or even hundreds of synchronized cameras. Dependence on such costly lab setups severely limits practical scalability. To this end, we propose a sparse‑camera dynamic reconstruction framework that exploits abundant yet inconsistent generative observations. Our key innovation is the Spatio‑Temporal Distortion Field, which provides a unified mechanism for modeling inconsistencies in generative observations across both spatial and temporal dimensions. Building on this, we develop a complete pipeline that enables 4D reconstruction from sparse and uncalibrated camera inputs. We evaluate our method on multi‑camera dynamic scene benchmarks, achieving spatio‑temporally consistent high‑fidelity renderings and significantly outperforming existing approaches. Project page available at https://inspatio.github.io/sparse‑cam4d/
Authors:Moritz Nottebaum, Matteo Dunnhofer, Christian Micheloni
Abstract:
Recent research on vision backbone architectures has predominantly focused on optimizing efficiency for hardware platforms with high parallel processing capabilities. This category increasingly includes embedded systems such as mobile phones and embedded AI accelerator modules. In contrast, CPUs do not have the possibility to parallelize operations in the same manner, wherefore models benefit from a specific design philosophy that balances amount of operations (MACs) and hardware‑efficient execution by having high MACs per second (MACpS). In pursuit of this, we investigate two modifications to standard convolutions, aimed at reducing computational cost: grouping convolutions and reducing kernel sizes. While both adaptations substantially decrease the total number of MACs required for inference, sustaining low latency necessitates preserving hardware‑efficiency. Our experiments across diverse CPU devices confirm that these adaptations successfully retain high hardware‑efficiency on CPUs. Based on these insights, we introduce CPUBone, a new family of vision backbone models optimized for CPU‑based inference. CPUBone achieves state‑of‑the‑art Speed‑Accuracy Trade‑offs (SATs) across a wide range of CPU devices and effectively transfers its efficiency to downstream tasks such as object detection and semantic segmentation. Models and code are available at https://github.com/altair199797/CPUBone.
Authors:MD Khalequzzaman Chowdhury Sayem, Mubarrat Tajoar Chowdhury, Yihalem Yimolal Tiruneh, Muneeb A. Khan, Muhammad Salman Ali, Binod Bhattarai, Seungryul Baek
Abstract:
Understanding the fine‑grained articulation of human hands is critical in high‑stakes settings such as robot‑assisted surgery, chip manufacturing, and AR/VR‑based human‑AI interaction. Despite achieving near‑human performance on general vision‑language benchmarks, current vision‑language models (VLMs) struggle with fine‑grained spatial reasoning, especially in interpreting complex and articulated hand poses. We introduce HandVQA, a large‑scale diagnostic benchmark designed to evaluate VLMs' understanding of detailed hand anatomy through visual question answering. Built upon high‑quality 3D hand datasets (FreiHAND, InterHand2.6M, FPHA), our benchmark includes over 1.6M controlled multiple‑choice questions that probe spatial relationships between hand joints, such as angles, distances, and relative positions. We evaluate several state‑of‑the‑art VLMs (LLaVA, DeepSeek and Qwen‑VL) in both base and fine‑tuned settings, using lightweight fine‑tuning via LoRA. Our findings reveal systematic limitations in current models, including hallucinated finger parts, incorrect geometric interpretations, and poor generalization. HandVQA not only exposes these critical reasoning gaps but provides a validated path to improvement. We demonstrate that the 3D‑grounded spatial knowledge learned from our benchmark transfers in a zero‑shot setting, significantly improving accuracy of model on novel downstream tasks like hand gesture recognition (+10.33%) and hand‑object interaction (+2.63%).
Authors:Quan Dao, Dimitris Metaxas
Abstract:
Transformer architectures, particularly Diffusion Transformers (DiTs), have become widely used in diffusion and flow‑matching models due to their strong performance compared to convolutional UNets. However, the isotropic design of DiTs processes the same number of patchified tokens in every block, leading to relatively heavy computation during training process. In this work, we introduce a multi‑patch transformer design in which early blocks operate on larger patches to capture coarse global context, while later blocks use smaller patches to refine local details. This hierarchical design could reduces computational cost by up to 50% in GFLOPs while achieving good generative performance. In addition, we also propose improved designs for time and class embeddings that accelerate training convergence. Extensive experiments on the ImageNet dataset demonstrate the effectiveness of our architectural choices. Code is released at: https://github.com/quandao10/MPDiT
Authors:Shuai Lv, Chang Liu, Feng Tang, Yujie Yuan, Aojun Zhou, Kui Zhang, Xi Yang, Yangqiu Song
Abstract:
Multimodal Large Language Models (MLLMs) achieve strong multimodal reasoning performance, yet we identify a recurring failure mode in long‑form generation: as outputs grow longer, models progressively drift away from image evidence and fall back on textual priors, resulting in ungrounded reasoning and hallucinations. Interestingly, Based on attention analysis, we find that MLLMs have a latent capability for late‑stage visual verification that is present but not consistently activated. Motivated by this observation, we propose Visual Re‑Examination (VRE), a self‑evolving training framework that enables MLLMs to autonomously perform visual introspection during reasoning without additional visual inputs. Rather than distilling visual capabilities from a stronger teacher, VRE promotes iterative self‑improvement by leveraging the model itself to generate reflection traces, making visual information actionable through information gain. Extensive experiments across diverse multimodal benchmarks demonstrate that VRE consistently improves reasoning accuracy and perceptual reliability, while substantially reducing hallucinations, especially in long‑chain settings. Code is available at https://github.com/Xiaobu‑USTC/VRE.
Authors:Mingyu Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Xiaowei Zhu, Jiajia Nie, Yinwei Wei, Yupeng Hu
Abstract:
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm. It aims to retrieve target images from large‑scale image databases that are consistent with the modification semantics, based on a multimodal query composed of a reference image and modification text. Although existing methods have made significant progress in cross‑modal alignment and feature fusion, a key flaw remains: the neglect of contextual information in discriminating matching samples. However, addressing this limitation is not an easy task due to two challenges: 1) implicit dependencies and 2) the lack of a differential amplification mechanism. To address these challenges, we propose a dual‑patH composItional coNtextualized neTwork (HINT), which can perform contextualized encoding and amplify the similarity differences between matching and non‑matching samples, thus improving the upper performance of CIR models in complex scenarios. Our HINT model achieves optimal performance on all metrics across two CIR benchmark datasets, demonstrating the superiority of our HINT model. Codes are available at https://github.com/zh‑mingyu/HINT.
Authors:Jiayi Chen, Wenxuan Song, Shuai Chen, Jingbo Wang, Zhijun Li, Haoang Li
Abstract:
Vision‑‑Language‑‑Action (VLA) models that encode actions using a discrete tokenization scheme are increasingly adopted for robotic manipulation, but existing decoding paradigms remain fundamentally limited. Whether actions are decoded sequentially by autoregressive VLAs or in parallel by discrete diffusion VLAs, once a token is generated, it is typically fixed and cannot be revised in subsequent iterations, so early token errors cannot be effectively corrected later. We propose DFM‑VLA, a discrete flow matching VLA for iterative refinement of action tokens. DFM‑VLA~models a token‑level probability velocity field that dynamically updates the full action sequence across refinement iterations. We investigate two ways to construct the velocity field: an auxiliary velocity‑head formulation and an action‑embedding‑guided formulation. Our framework further adopts a two‑stage decoding strategy with an iterative refinement stage followed by deterministic validation for stable convergence. Extensive experiments on CALVIN, LIBERO, and real‑world manipulation tasks show that DFM‑VLA consistently outperforms strong autoregressive, discrete diffusion, and continuous diffusion baselines in manipulation performance while retaining high inference efficiency. In particular, DFM‑VLA achieves an average success length of 4.44 on CALVIN and an average success rate of 95.7% on LIBERO, highlighting the value of action refinement via discrete flow matching for robotic manipulation. Our project is available https://chris1220313648.github.io/DFM‑VLA/
Authors:Cai Selvas-Sala, Lei Kang, Lluis Gomez
Abstract:
As multimodal models like CLIP become integral to downstream systems, the need to remove sensitive information is critical. However, machine unlearning for contrastively‑trained encoders remains underexplored, and existing evaluations fail to diagnose fine‑grained, association‑level forgetting. We introduce SALMUBench (Sensitive Association‑Level Multimodal Unlearning), a benchmark built upon a synthetic dataset of 60K persona‑attribute associations and two foundational models: a Compromised model polluted with this data, and a Clean model without it. To isolate unlearning effects, both are trained from scratch on the same 400M‑pair retain base, with the Compromised model additionally trained on the sensitive set. We propose a novel evaluation protocol with structured holdout sets (holdout identity, holdout association) to precisely measure unlearning efficacy and collateral damage. Our benchmark reveals that while utility‑efficient deletion is feasible, current methods exhibit distinct failure modes: they either fail to forget effectively or over‑generalize by erasing more than intended. SALMUBench sets a new standard for comprehensive unlearning evaluation, and we publicly release our dataset, models, evaluation scripts, and leaderboards to foster future research.
Authors:Tomoya Miyawaki, Kazuto Nakashima, Yumi Iwashita, Ryo Kurazume
Abstract:
LiDAR‑based semantic segmentation is a key component for autonomous mobile robots, yet large‑scale annotation of LiDAR point clouds is prohibitively expensive and time‑consuming. Although simulators can provide labeled synthetic data, models trained on synthetic data often underperform on real‑world data due to a data‑level domain gap. To address this issue, we propose DRUM, a novel Sim2Real translation framework. We leverage a diffusion model pre‑trained on unlabeled real‑world data as a generative prior and translate synthetic data by reproducing two key measurement characteristics: reflectance intensity and raydrop noise. To improve sample fidelity, we introduce a raydrop‑aware masked guidance mechanism that selectively enforces consistency with the input synthetic data while preserving realistic raydrop noise induced by the diffusion prior. Experimental results demonstrate that DRUM consistently improves Sim2Real performance across multiple representations of LiDAR data. The project page is available at https://miya‑tomoya.github.io/drum.
Authors:Rui Wang, Huisi Wu, Jing Qin
Abstract:
Accurate and temporally consistent segmentation of the left ventricle from echocardiography videos is essential for estimating the ejection fraction and assessing cardiac function. However, modeling spatiotemporal dynamics remains difficult due to severe speckle noise and rapid non‑rigid deformations. Existing linear recurrent models offer efficient in‑context associative recall for temporal tracking, but rely on unconstrained state updates, which cause progressive singular value decay in the state matrix, a phenomenon known as rank collapse, resulting in anatomical details being overwhelmed by noise. To address this, we propose OSA, a framework that constrains the state evolution on the Stiefel manifold. We introduce the Orthogonalized State Update (OSU) mechanism, which formulates the memory evolution as Euclidean projected gradient descent on the Stiefel manifold to prevent rank collapse and maintain stable temporal transitions. Furthermore, an Anatomical Prior‑aware Feature Enhancement module explicitly separates anatomical structures from speckle noise through a physics‑driven process, providing the temporal tracker with noise‑resilient structural cues. Comprehensive experiments on the CAMUS and EchoNet‑Dynamic datasets show that OSA achieves state‑of‑the‑art segmentation accuracy and temporal stability, while maintaining real‑time inference efficiency for clinical deployment. Codes are available at https://github.com/wangrui2025/OSA.
Authors:Pan Zhao, Hui Yuan, Chang Sun, Chongzhen Tian, Raouf Hamzaoui, Sam Kwong
Abstract:
Existing post‑decoding quality enhancement methods for point clouds are designed for static data and typically process each frame independently. As a result, they cannot effectively exploit the spatiotemporal correlations present in point cloud sequences.We propose a unified geometry and attribute enhancement framework (DUGAE) for G‑PCC compressed dynamic point clouds that explicitly exploits inter‑frame spatiotemporal correlations in both geometry and attributes. First, a dynamic geometry enhancement network (DGE‑Net) based on sparse convolution (SPConv) and feature‑domain geometry motion compensation (GMC) aligns and aggregates spatiotemporal information. Then, a detail‑aware k‑nearest neighbors (DA‑KNN) recoloring module maps the original attributes onto the enhanced geometry at the encoder side, improving mapping completeness and preserving attribute details. Finally, a dynamic attribute enhancement network (DAE‑Net) with dedicated temporal feature extraction and feature‑domain attribute motion compensation (AMC) refines attributes by modeling complex spatiotemporal correlations. On seven dynamic point clouds from the 8iVFB v2, Owlii, and MVUB datasets, DUGAE significantly enhanced the performance of the latest G‑PCC geometry‑based solid content test model (GeS‑TM v10). For geometry (D1), it achieved an average BD‑PSNR gain of 11.03 dB and a 93.95% BD‑bitrate reduction. For the luma component, it achieved a 4.23 dB BD‑PSNR gain with a 66.61% BD‑bitrate reduction. DUGAE also improved perceptual quality (as measured by PCQM) and outperformed V‑PCC. Our source code will be released on GitHub at: https://github.com/yuanhui0325/DUGAE
Authors:Youngju Na, Jaeseong Yun, Soohyun Ryu, Hyunsu Kim, Sung-Eui Yoon, Suyong Yeon
Abstract:
While 3D Gaussian splatting has emerged as a powerful paradigm, it fundamentally fails to model transparency such as glass panels. The core challenge lies in decoupling the intertwined radiance contributions from transparent interfaces and the transmitted geometry observed through the glass. We present GLINT, a framework that models scene‑scale transparency through explicit decomposed Gaussian representation. GLINT reconstructs the primary interface and models reflected and transmitted radiance separately, enabling consistent radiance transport. During optimization, GLINT bootstraps transparency localization from geometry‑separation cues induced by the decomposition, together with geometry and material priors from a pre‑trained video relighting model. Extensive experiments demonstrate consistent improvements over prior methods for reconstructing complex transparent scenes.
Authors:Bozhao Li, Shaocong Wu, Tong Shao, Senqiao Yang, Qiben Shan, Zhuotao Tian, Jingyong Su
Abstract:
Recent advances in open‑vocabulary object detection focus primarily on two aspects: scaling up datasets and leveraging contrastive learning to align language and vision modalities. However, these approaches often neglect internal consistency within a single modality, particularly when background or environmental changes occur. This lack of consistency leads to a performance drop because the model struggles to detect the same object in different scenes, which reveals a robustness gap. To address this issue, we introduce Contextual Consistency Learning (CCL), a novel framework that integrates two key strategies: Contextual Bootstrapped Data Generation (CBDG) and Contextual Consistency Loss (CCLoss). CBDG functions as a data generation mechanism, producing images that contain the same objects across diverse backgrounds. This is essential because existing datasets alone do not support our CCL framework. The CCLoss further enforces the invariance of object features despite environmental changes, thereby improving the model's robustness in different scenes. These strategies collectively form a unified framework for ensuring contextual consistency within the same modality. Our method achieves state‑of‑the‑art performance, surpassing previous approaches by +16.3 AP on OmniLabel and +14.9 AP on D3. These results demonstrate the importance of enforcing intra‑modal consistency, significantly enhancing model generalization in diverse environments. Our code is publicly available at: https://github.com/bozhao‑li/CCL.
Authors:Shubhi Shukla, Pravin Nair
Abstract:
Image restoration, the recovery of clean images from degraded measurements, has applications in various domains like surveillance, defense, and medical imaging. Despite achieving state‑of‑the‑art (SOTA) restoration performance, existing convolutional and attention‑based networks lack stability guarantees under minor shifts in input, exposing a robustness accuracy trade‑off. We develop provably contractive (global Lipschitz < 1) denoiser networks that considerably reduce this gap. Our design composes proximal layers obtained from unfolding techniques, with Lipschitz‑controlled convolutional refinements. By contractivity, our denoiser guarantees that input perturbations of strength \|δ\|\le\varepsilon induce at most \varepsilon change at the output, while strong baselines such as DnCNN and Restormer can exhibit larger deviations under the same perturbations. On image denoising, the proposed model is competitive with unconstrained SOTA denoisers, reporting the tightest gap for a provably 1‑Lipschitz model and establishing that such gaps are indeed achievable by contractive denoisers. Moreover, the proposed denoisers act as strong regularizers for image restoration that provably effect convergence in Plug‑and‑Play algorithms. Our results show that enforcing strict Lipschitz control does not inherently degrade output quality, challenging a common assumption in the literature and moving the field toward verifiable and stable vision models. Codes and pretrained models are available at https://github.com/SHUBHI1553/Contractive‑Denoisers
Authors:Yi Zhang, Hongbo Huang, Liang-Jie Zhang
Abstract:
Diffusion models generate high‑quality images but pose serious risks like copyright violation and disinformation. Watermarking is a key defense for tracing and authenticating AI‑generated content. However, existing methods rely on threshold‑based detection, which only supports fuzzy matching and cannot recover structured watermark data bit‑exactly, making them unsuitable for offline verification or applications requiring lossless metadata (e.g., licensing instructions). To address this problem, in this paper, we propose Gaussian Shannon, a watermarking framework that treats the diffusion process as a noisy communication channel and enables both robust tracing and exact bit recovery. Our method embeds watermarks in the initial Gaussian noise without fine‑tuning or quality loss. We identify two types of channel interference, namely local bit flips and global stochastic distortions, and design a cascaded defense combining error‑correcting codes and majority voting. This ensures reliable end‑to‑end transmission of semantic payloads. Experiments across three Stable Diffusion variants and seven perturbation types show that Gaussian Shannon achieves state‑of‑the‑art bit‑level accuracy while maintaining a high true positive rate, enabling trustworthy rights attribution in real‑world deployment. The source code have been made available at: https://github.com/Rambo‑Yi/Gaussian‑Shannon
Authors:Kang Liu, Zhuoqi Ma, Siyu Liang, Yunan Li, Xiyue Gao, Chao Liang, Kun Xie, Qiguang Miao
Abstract:
Despite recent advances in medical vision‑language pretraining, existing models still struggle to capture the diagnostic workflow: radiographs are typically treated as context‑agnostic images, while radiologists' gaze ‑‑ a crucial cue for visual reasoning ‑‑ remains largely underexplored by existing methods. These limitations hinder the modeling of disease‑specific patterns and weaken cross‑modal alignment. To bridge this gap, we introduce CoGaze, a Context‑ and Gaze‑guided vision‑language pretraining framework for chest X‑rays. We first propose a context‑infused vision encoder that models how radiologists integrate clinical context ‑‑ including patient history, symptoms, and diagnostic intent ‑‑ to guide diagnostic reasoning. We then present a multi‑level supervision paradigm that (1) enforces intra‑ and inter‑modal semantic alignment through hybrid‑positive contrastive learning, (2) injects diagnostic priors via disease‑aware cross‑modal representation learning, and (3) leverages radiologists' gaze as probabilistic priors to guide attention toward diagnostically salient regions. Extensive experiments demonstrate that CoGaze consistently outperforms state‑of‑the‑art methods across diverse tasks, achieving up to +2.0% CheXbertF1 and +1.2% BLEU2 for free‑text and structured report generation, +23.2% AUROC for zero‑shot classification, and +12.2% Precision@1 for image‑text retrieval. Code is available at https://github.com/mk‑runner/CoGaze.
Authors:Mahesh Bhosale, Abdul Wasi, Shantam Srivastava, Shifa Latif, Tianyu Luan, Mingchen Gao, David Doermann, Xuan Gong
Abstract:
While powerful in image‑conditioned generation, multimodal large language models (MLLMs) can display uneven performance across demographic groups, highlighting fairness risks. In safety‑critical clinical settings, such disparities risk producing unequal diagnostic narratives and eroding trust in AI‑assisted decision‑making. While fairness has been studied extensively in vision‑only and language‑only models, its impact on MLLMs remains largely underexplored. To address these biases, we introduce FairLLaVA, a parameter‑efficient fine‑tuning method that mitigates group disparities in visual instruction tuning without compromising overall performance. By minimizing the mutual information between target attributes, FairLLaVA regularizes the model's representations to be demographic‑invariant. The method can be incorporated as a lightweight plug‑in, maintaining efficiency with low‑rank adapter fine‑tuning, and provides an architecture‑agnostic approach to fair visual instruction following. Extensive experiments on large‑scale chest radiology report generation and dermoscopy visual question answering benchmarks show that FairLLaVA consistently reduces inter‑group disparities while improving both equity‑scaled clinical performance and natural language generation quality across diverse medical imaging modalities. Code can be accessed at https://github.com/bhosalems/FairLLaVA.
Authors:Zhuoli Zhuang, Yu-Cheng Chang, Yu-Kai Wang, Thomas Do, Chin-Teng Lin
Abstract:
Recent advancements in computer vision have accelerated the development of autonomous driving. Despite these advancements, training machines to drive in a way that aligns with human expectations remains a significant challenge. Human factors are still essential, as humans possess a sophisticated cognitive system capable of rapidly interpreting scene information and making accurate decisions. Aligning machine with human intent has been explored with Reinforcement Learning with Human Feedback (RLHF). Conventional RLHF methods rely on collecting human preference data by manually ranking generated outputs, which is time‑consuming and indirect. In this work, we propose an electroencephalography (EEG)‑guided decision‑making framework to incorporate human cognitive insights without behaviour response interruption into reinforcement learning (RL) for autonomous driving. We collected EEG signals from 20 participants in a realistic driving simulator and analyzed event‑related potentials (ERP) in response to sudden environmental changes. Our proposed framework employs a neural network to predict the strength of ERP based on the cognitive information from visual scene information. Moreover, we explore the integration of such cognitive information into the reward signal of the RL algorithm. Experimental results show that our framework can improve the collision avoidance ability of the RL algorithm, highlighting the potential of neuro‑cognitive feedback in enhancing autonomous driving systems. Our project page is: https://alex95gogo.github.io/Cognitive‑Reward/.
Authors:Shounak Sural, Ragunathan Rajkumar
Abstract:
Localization in GNSS‑denied and GNSS‑degraded environments is a challenge for the safe widespread deployment of autonomous vehicles. Such GNSS‑challenged environments require alternative methods for robust localization. In this work, we propose BEVMapMatch, a framework for robust vehicle re‑localization on a known map without the need for GNSS priors. BEVMapMatch uses a context‑aware lidar+camera fusion method to generate multimodal Bird's Eye View (BEV) segmentations around the ego vehicle in both good and adverse weather conditions. Leveraging a search mechanism based on cross‑attention, the generated BEV segmentation maps are then used for the retrieval of candidate map patches for map‑matching purposes. Finally, BEVMapMatch uses the top retrieved candidate for finer alignment against the generated BEV segmentation, achieving accurate global localization without the need for GNSS. Multiple frames of generated BEV segmentation further improve localization accuracy. Extensive evaluations show that BEVMapMatch outperforms existing methods for re‑localization in GNSS‑denied and adverse environments, with a Recall@1m of 39.8%, being nearly twice as much as the best performing re‑localization baseline. Our code and data will be made available at https://github.com/ssuralcmu/BEVMapMatch.git.
Authors:Julia Wolleb, Cristiana Baloescu, Alicia Durrer, Hemant D. Tagare, Xenophon Papademetris
Abstract:
Implicit neural representations (INRs) have emerged as a powerful framework for continuous image representation learning. In Functa‑based approaches, each image is encoded as a latent modulation vector that conditions a shared INR, enabling strong reconstruction performance. However, the structure and interpretability of the corresponding latent spaces remain largely unexplored. In this work, we investigate the latent space of Functa‑based models for ultrasound videos and propose Low‑Rank‑Modulated Functa (LRM‑Functa), a novel architecture that enforces a low‑rank adaptation of modulation vectors in the time‑resolved latent space. When applied to cardiac ultrasound, the resulting latent space exhibits clearly structured periodic trajectories, facilitating visualization and interpretability of temporal patterns. The latent space can be traversed to sample novel frames, revealing smooth transitions along the cardiac cycle, and enabling direct readout of end‑diastolic (ED) and end‑systolic (ES) frames without additional model training. We show that LRM‑Functa outperforms prior methods in unsupervised ED and ES frame detection, while compressing each video frame to as low as rank k=2 without sacrificing competitive downstream performance on ejection fraction prediction. Evaluations on out‑of‑distribution frame selection in a cardiac point‑of‑care dataset, as well as on lung ultrasound for B‑line classification, demonstrate the generalizability of our approach. Overall, LRM‑Functa provides a compact, interpretable, and generalizable framework for ultrasound video analysis. The code is available at https://github.com/JuliaWolleb/LRM_Functa.
Authors:Guoping Xu, Jayaram K. Udupa, Yubing Tong, Xin Long, Ying Zhang, Jie Deng, Weiguo Lu, You Zhang
Abstract:
Accurate lesion segmentation is essential in medical image analysis, yet most existing methods are designed for specific anatomical sites or imaging modalities, limiting their generalizability. Recent vision‑language foundation models enable concept‑driven segmentation in natural images, offering a promising direction for more flexible medical image analysis. However, concept‑prompt‑based lesion segmentation, particularly with the latest Segment Anything Model 3 (SAM3), remains underexplored.
In this work, we present a systematic evaluation of SAM3 for lesion segmentation. We assess its performance using geometric bounding boxes and concept‑based text and image prompts across multiple modalities, including multiparametric MRI, CT, ultrasound, dermoscopy, and endoscopy. To improve robustness, we incorporate additional prior knowledge, such as adjacent‑slice predictions, multiparametric information, and prior annotations. We further compare different fine‑tuning strategies, including partial module tuning, adapter‑based methods, and full‑model optimization.
Experiments on 13 datasets covering 11 lesion types demonstrate that SAM3 achieves strong cross‑modality generalization, reliable concept‑driven segmentation, and accurate lesion delineation. These results highlight the potential of concept‑based foundation models for scalable and practical medical image segmentation. Code and trained models will be released at: https://github.com/apple1986/lesion‑sam3
Authors:PAN Team, Qiyue Gao, Kun Zhou, Jiannan Xiang, Zihan Liu, Dequan Yang, Junrong Chen, Arif Ahmad, Cong Zeng, Ganesh Bannur, Xinqi Huang, Zheqi Liu, Yi Gu, Yichi Yang, Guangyi Liu, Zhiting Hu, Zhengzhong Liu, Eric Xing
Abstract:
World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM benchmarks remain narrowly focused on next‑state prediction and visual fidelity, overlooking the richer simulation capabilities required for intelligent behavior. To address this gap, we introduce WR‑Arena, a comprehensive benchmark for evaluating WMs along three fundamental dimensions of next world simulation: (i) Action Simulation Fidelity, the ability to interpret and follow semantically meaningful, multi‑step instructions and generate diverse counterfactual rollouts; (ii) Long‑horizon Forecast, the ability to sustain accurate, coherent, and physically plausible simulations across extended interactions; and (iii) Simulative Reasoning and Planning, the ability to support goal‑directed reasoning by simulating, comparing, and selecting among alternative futures in both structured and open‑ended environments. We build a task taxonomy and curate diverse datasets designed to probe these capabilities, moving beyond single‑turn and perceptual evaluations. Through extensive experiments with state‑of‑the‑art WMs, our results expose a substantial gap between current models and human‑level hypothetical reasoning, and establish WR‑Arena as both a diagnostic tool and a guideline for advancing next‑generation world models capable of robust understanding, forecasting, and purposeful action. The code is available at https://github.com/MBZUAI‑IFM/WR‑Arena.
Authors:Trong Thang Pham, Hien Nguyen, Ngan Le
Abstract:
Current multimodal large language models (MLLMs) cannot effectively utilize eye‑gaze information for video understanding, even when gaze cues are supplied via visual overlays or text descriptions. We introduce GazeQwen, a parameter efficient approach that equips an open‑source MLLM with gaze awareness through hidden‑state modulation. At its core is a compact gaze resampler (~1‑5 M trainable parameters) that encodes V‑JEPA 2.1 video features together with fixation‑derived positional encodings and produces additive residuals injected into selected LLM decoder layers via forward hooks. An optional second training stage adds low‑rank adapters (LoRA) to the LLM for tighter integration. Evaluated on all 10 tasks of the StreamGaze benchmark, GazeQwen reaches 63.9% accuracy, a +16.1 point gain over the same Qwen2.5‑VL‑7B backbone with gaze as visual prompts and +10.5 points over GPT‑4o, the highest score among all open‑source and proprietary models tested. These results suggest that learning where to inject gaze within an LLM is more effective than scaling model size or engineering better prompts. All code and checkpoints are available at https://github.com/phamtrongthang123/gazeqwen .
Authors:Laura Fink, Linus Franke, George Kopanas, Marc Stamminger, Peter Hedman
Abstract:
We propose a feed‑forward method for dense Signed Distance Field (SDF) regression from unstructured image collections in less than three seconds, without camera calibration or post‑hoc fusion. Our key insight is that the intermediate feature space of pretrained multi‑view feed‑forward geometry transformers already encodes a powerful joint world representation; yet, existing pipelines discard it, routing features through per‑view prediction heads before assembling 3D geometry post‑hoc, which discards valuable completeness information and accumulates inaccuracies.
We instead perform 3D extraction directly from geometry transformer features via learned volumetric extraction: voxelized canonical embeddings that progressively absorb multi‑view geometry information through interleaved cross‑ and self‑attention into a structured volumetric latent grid. A simple convolutional decoder then maps this grid to a dense SDF. We additionally propose a scalable, validity‑aware supervision scheme directly using SDFs derived from depth maps or 3D assets, tackling practical issues like non‑watertight meshes. Our approach yields complete and well‑defined distance values across sparse‑ and dense‑view settings and demonstrates geometrically plausible completions. Code and further material can be found at https://lorafib.github.io/fus3d.
Authors:Haonan Han, Jiancheng Huang, Xiaopeng Sun, Junyan He, Rui Yang, Jie Hu, Xiaojiang Peng, Lin Ma, Xiaoming Wei, Xiu Li
Abstract:
Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fail tasks that require physical, causal, or complex spatial reasoning. Current evaluations largely rely on superficial metrics or fragmented benchmarks, creating a ``performance mirage'' that overlooks the generative process. To address this, we introduce ViGoR Vision‑Gnerative Reasoning‑centric Benchmark), a unified framework designed to dismantle this mirage. ViGoR distinguishes itself through four key innovations: 1) holistic cross‑modal coverage bridging Image‑to‑Image and Video tasks; 2) a dual‑track mechanism evaluating both intermediate processes and final results; 3) an evidence‑grounded automated judge ensuring high human alignment; and 4) granular diagnostic analysis that decomposes performance into fine‑grained cognitive dimensions. Experiments on over 20 leading models reveal that even state‑of‑the‑art systems harbor significant reasoning deficits, establishing ViGoR as a critical ``stress test'' for the next generation of intelligent vision models. The demo have been available at https://vincenthancoder.github.io/ViGoR‑Bench/
Authors:Yuan Zhang, Sihao Dou, Kai Hu, Shuhua Deng, Chunhong Cao, Fen Xiao, Xieping Gao
Abstract:
Endoscopic video analysis is essential for early gastrointestinal screening but remains hindered by limited high‑quality annotations. While self‑supervised video pre‑training shows promise, existing methods developed for natural videos prioritize dense spatio‑temporal modeling and exhibit motion bias, overlooking the static, structured semantics critical to clinical decision‑making. To address this challenge, we propose Focus‑to‑Perceive Representation Learning (FPRL), a cognition‑inspired hierarchical framework that emulates clinical examination. FPRL first focuses on intra‑frame lesion‑centric regions to learn static semantics, and then perceives their evolution across frames to model contextual semantics. To achieve this, FPRL employs a hierarchical semantic modeling mechanism that explicitly distinguishes and collaboratively learns both types of semantics. Specifically, it begins by capturing static semantics via teacher‑prior adaptive masking (TPAM) combined with multi‑view sparse sampling. This approach mitigates redundant temporal dependencies and enables the model to concentrate on lesion‑related local semantics. Following this, contextual semantics are derived through cross‑view masked feature completion (CVMFC) and attention‑guided temporal prediction (AGTP). These processes establish cross‑view correspondences and effectively model structured inter‑frame evolution, thereby reinforcing temporal semantic continuity while preserving global contextual integrity. Extensive experiments on 11 endoscopic video datasets show that FPRL achieves superior performance across diverse downstream tasks, demonstrating its effectiveness in endoscopic video representation learning. The code is available at https://github.com/MLMIP/FPRL.
Authors:Yawen Luo, Xiaoyu Shi, Junhao Zhuang, Yutian Chen, Quande Liu, Xintao Wang, Pengfei Wan, Tianfan Xue
Abstract:
Multi‑shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi‑shot architecture that enables interactive storytelling and efficient on‑the‑fly frame generation. By reformulating the task as next‑shot generation conditioned on historical context, ShotStream allows users to dynamically instruct ongoing narratives via streaming prompts. We achieve this by first fine‑tuning a text‑to‑video model into a bidirectional next‑shot generator, which is then distilled into a causal student via Distribution Matching Distillation. To overcome the challenges of inter‑shot consistency and error accumulation inherent in autoregressive generation, we introduce two key innovations. First, a dual‑cache memory mechanism preserves visual coherence: a global context cache retains conditional frames for inter‑shot consistency, while a local context cache holds generated frames within the current shot for intra‑shot consistency. And a RoPE discontinuity indicator is employed to explicitly distinguish the two caches to eliminate ambiguity. Second, to mitigate error accumulation, we propose a two‑stage distillation strategy. This begins with intra‑shot self‑forcing conditioned on ground‑truth historical shots and progressively extends to inter‑shot self‑forcing using self‑generated histories, effectively bridging the train‑test gap. Extensive experiments demonstrate that ShotStream generates coherent multi‑shot videos with sub‑second latency, achieving 16 FPS on a single GPU. It matches or exceeds the quality of slower bidirectional models, paving the way for real‑time interactive storytelling. Training and inference code, as well as the models, are available on our
Authors:Yixing Lao, Xuyang Bai, Xiaoyang Wu, Nuoyuan Yan, Zixin Luo, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Shiwei Li, Hengshuang Zhao
Abstract:
Existing feed‑forward 3D Gaussian Splatting methods predict pixel‑aligned primitives, leading to a quadratic growth in primitive count as resolution increases. This fundamentally limits their scalability, making high‑resolution synthesis such as 4K intractable. We introduce LGTM (Less Gaussians, Texture More), a feed‑forward framework that overcomes this resolution scaling barrier. By predicting compact Gaussian primitives coupled with per‑primitive textures, LGTM decouples geometric complexity from rendering resolution. This approach enables high‑fidelity 4K novel view synthesis without per‑scene optimization, a capability previously out of reach for feed‑forward methods, all while using significantly fewer Gaussian primitives. Project page: https://yxlao.github.io/lgtm/
Authors:Sicheng Zuo, Yuxuan Li, Wenzhao Zheng, Zheng Zhu, Jie Zhou, Jiwen Lu
Abstract:
Vision‑language‑action models have reshaped autonomous driving to incorporate languages into the decision‑making process. However, most existing pipelines only utilize the language modality for scene descriptions or reasoning and lack the flexibility to follow diverse user instructions for personalized driving. To address this, we first construct a large‑scale driving dataset (InstructScene) containing around 100,000 scenes annotated with diverse driving instructions with the corresponding trajectories. We then propose a unified Vision‑Language‑World‑Action model, Vega, for instruction‑based generation and planning. We employ the autoregressive paradigm to process visual inputs (vision) and language instructions (language) and the diffusion paradigm to generate future predictions (world modeling) and trajectories (action). We perform joint attention to enable interactions between the modalities and use individual projection layers for different modalities for more capabilities. Extensive experiments demonstrate that our method not only achieves superior planning performance but also exhibits strong instruction‑following abilities, paving the way for more intelligent and personalized driving systems.
Authors:Zehao Wang, Huaide Jiang, Shuaiwu Dong, Yuping Wang, Hang Qiu, Jiachen Li
Abstract:
Human driving behavior is inherently personal, which is shaped by long‑term habits and influenced by short‑term intentions. Individuals differ in how they accelerate, brake, merge, yield, and overtake across diverse situations. However, existing end‑to‑end autonomous driving systems either optimize for generic objectives or rely on fixed driving modes, lacking the ability to adapt to individual preferences or interpret natural language intent. To address this gap, we propose Drive My Way (DMW), a personalized Vision‑Language‑Action (VLA) driving framework that aligns with users' long‑term driving habits and adapts to real‑time user instructions. DMW learns a user embedding from our personalized driving dataset collected across multiple real drivers and conditions the policy on this embedding during planning, while natural language instructions provide additional short‑term guidance. Closed‑loop evaluation on the Bench2Drive benchmark demonstrates that DMW improves style instruction adaptation, and user studies show that its generated behaviors are recognizable as each driver's own style, highlighting personalization as a key capability for human‑centered autonomous driving. Our data and code are available at https://dmw‑cvpr.github.io/.
Authors:Dingxi Zhang, Fangjinhua Wang, Marc Pollefeys, Haofei Xu
Abstract:
Accurate estimation of large displacement optical flow remains a critical challenge. Existing methods typically rely on iterative local search or/and domain‑specific fine‑tuning, which severely limits their performance in large displacement and zero‑shot generalization scenarios. To overcome this, we introduce MegaFlow, a simple yet powerful model for zero‑shot large displacement optical flow. Rather than relying on highly complex, task‑specific architectural designs, MegaFlow adapts powerful pre‑trained vision priors to produce temporally consistent motion fields. In particular, we formulate flow estimation as a global matching problem by leveraging pre‑trained global Vision Transformer features, which naturally capture large displacements. This is followed by a few lightweight iterative refinements to further improve the sub‑pixel accuracy. Extensive experiments demonstrate that MegaFlow achieves state‑of‑the‑art zero‑shot performance across multiple optical flow benchmarks. Moreover, our model also delivers highly competitive zero‑shot performance on long‑range point tracking benchmarks, demonstrating its robust transferability and suggesting a unified paradigm for generalizable motion estimation. Our project page is at: https://kristen‑z.github.io/projects/megaflow.
Authors:Ziyin Wang, Sirui Xu, Chuan Guo, Bing Zhou, Jiangshan Gong, Jian Wang, Yu-Xiong Wang, Liang-Yan Gui
Abstract:
Generating realistic human‑object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion‑based approaches often rely on hand‑crafted contact priors or human‑imposed kinematic constraints to improve contact quality. We propose LIGHT, a data‑driven alternative in which guidance emerges from the denoising pace itself, reducing dependence on manually designed priors. Building on diffusion forcing, we factor the representation into modality‑specific components and assign individualized noise levels with asynchronous denoising schedules. In this paradigm, cleaner components guide noisier ones through cross‑attention, yielding guidance without auxiliary classifiers. We find that this data‑driven guidance is inherently contact‑aware, and can be enhanced when training is augmented with a broad spectrum of synthetic object geometries, encouraging invariance of contact semantics to geometric diversity. Extensive experiments show that pace‑induced guidance more effectively mirrors the benefits of contact priors than conventional classifier‑free guidance, while achieving higher contact fidelity, more realistic HOI generation, and stronger generalization to unseen objects and tasks.
Authors:Xiaofeng Mao, Shaohao Rui, Kaining Ying, Bo Zheng, Chuanhao Li, Mingmin Chi, Kaipeng Zhang
Abstract:
Autoregressive video diffusion models have demonstrated remarkable progress, yet they remain bottlenecked by intractable linear KV‑cache growth, temporal repetition, and compounding errors during long‑video generation. To address these challenges, we present PackForcing, a unified framework that efficiently manages the generation history through a novel three‑partition KV‑cache strategy. Specifically, we categorize the historical context into three distinct types: (1) Sink tokens, which preserve early anchor frames at full resolution to maintain global semantics; (2) Mid tokens, which achieve a massive spatiotemporal compression (32x token reduction) via a dual‑branch network fusing progressive 3D convolutions with low‑resolution VAE re‑encoding; and (3) Recent tokens, kept at full resolution to ensure local temporal coherence. To strictly bound the memory footprint without sacrificing quality, we introduce a dynamic top‑k context selection mechanism for the mid tokens, coupled with a continuous Temporal RoPE Adjustment that seamlessly re‑aligns position gaps caused by dropped tokens with negligible overhead. Empowered by this principled hierarchical context compression, PackForcing can generate coherent 2‑minute, 832x480 videos at 16 FPS on a single H200 GPU. It achieves a bounded KV cache of just 4 GB and enables a remarkable 24x temporal extrapolation (5s to 120s), operating effectively either zero‑shot or trained on merely 5‑second clips. Extensive results on VBench demonstrate state‑of‑the‑art temporal consistency (26.07) and dynamic degree (56.25), proving that short‑video supervision is sufficient for high‑quality, long‑video synthesis. https://github.com/ShandaAI/PackForcing
Authors:Jiabin Hua, Hengyuan Xu, Aojie Li, Wei Cheng, Gang Yu, Xingjun Ma, Yu-Gang Jiang
Abstract:
Fine‑grained facial expression editing has long been limited by intrinsic semantic overlap. To address this, we construct the Flex Facial Expression (FFE) dataset with continuous affective annotations and establish FFE‑Bench to evaluate structural confusion, editing accuracy, linear controllability, and the trade‑off between expression editing and identity preservation. We propose PixelSmile, a diffusion framework that disentangles expression semantics via fully symmetric joint training. PixelSmile combines intensity supervision with contrastive learning to produce stronger and more distinguishable expressions, achieving precise and stable linear expression control through textual latent interpolation. Extensive experiments demonstrate that PixelSmile achieves superior disentanglement and robust identity preservation, confirming its effectiveness for continuous, controllable, and fine‑grained expression editing, while naturally supporting smooth expression blending.
Authors:Hai X. Pham, David T. Hoffmann, Ricardo Guerrero, Brais Martinez
Abstract:
Contrastive vision‑language (V&L) models remain a popular choice for various applications. However, several limitations have emerged, most notably the limited ability of V&L models to learn compositional representations. Prior methods often addressed this limitation by generating custom training data to obtain hard negative samples. Hard negatives have been shown to improve performance on compositionality tasks, but are often specific to a single benchmark, do not generalize, and can cause substantial degradation of basic V&L capabilities such as zero‑shot or retrieval performance, rendering them impractical. In this work we follow a different approach. We identify two root causes that limit compositionality performance of V&Ls: 1) Long training captions do not require a compositional representation; and 2) The final global pooling in the text and image encoders lead to a complete loss of the necessary information to learn binding in the first place. As a remedy, we propose two simple solutions: 1) We obtain short concept centric caption parts using standard NLP software and align those with the image; and 2) We introduce a parameter‑free cross‑modal attention‑pooling to obtain concept centric visual embeddings from the image encoder. With these two changes and simple auxiliary contrastive losses, we obtain SOTA performance on standard compositionality benchmarks, while maintaining or improving strong zero‑shot and retrieval capabilities. This is achieved without increasing inference cost. We release the code for this work at https://github.com/saic‑fi/concept_centric_clip.
Authors:Kaijin Chen, Dingkang Liang, Xin Zhou, Yikang Ding, Xiaoqiang Liu, Pengfei Wan, Xiang Bai
Abstract:
Video world models have shown immense potential in simulating the physical world, yet existing memory mechanisms primarily treat environments as static canvases. When dynamic subjects hide out of sight and later re‑emerge, current methods often struggle, leading to frozen, distorted, or vanishing subjects. To address this, we introduce Hybrid Memory, a novel paradigm requiring models to simultaneously act as precise archivists for static backgrounds and vigilant trackers for dynamic subjects, ensuring motion continuity during out‑of‑view intervals. To facilitate research in this direction, we construct HM‑World, the first large‑scale video dataset dedicated to hybrid memory. It features 59K high‑fidelity clips with decoupled camera and subject trajectories, encompassing 17 diverse scenes, 49 distinct subjects, and meticulously designed exit‑entry events to rigorously evaluate hybrid coherence. Furthermore, we propose HyDRA, a specialized memory architecture that compresses memory into tokens and utilizes a spatiotemporal relevance‑driven retrieval mechanism. By selectively attending to relevant motion cues, HyDRA effectively preserves the identity and motion of hidden subjects. Extensive experiments on HM‑World demonstrate that our method significantly outperforms state‑of‑the‑art approaches in both dynamic subject consistency and overall generation quality. Code is publicly available at https://github.com/H‑EmbodVis/HyDRA.
Authors:Jinbo Xing, Zeyinzi Jiang, Yuxiang Tuo, Chaojie Mao, Xiaotang Gai, Xi Chen, Jingfeng Zhang, Yulin Pan, Zhen Han, Jie Xiao, Keyu Yan, Chenwei Xie, Chongyang Zhong, Kai Zhu, Tong Shen, Lianghua Huang, Yu Liu, Yujiu Yang
Abstract:
Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi‑modal inputs, they typically produce only single‑modality outputs. This challenge of producing interleaved content is mainly due to training data scarcity and the difficulty of modeling long‑range cross‑modal context. To address this issue, we decompose interleaved generation into textual planning and visual consistency modeling, and introduce a framework consisting of a planner and a visualizer. The planner produces dense textual descriptions for visual content, while the visualizer synthesizes images accordingly. Under this guidance, we construct large‑scale textual‑proxy interleaved data (where visual content is represented in text) to train the planner, and curate reference‑guided image data to train the visualizer. These designs give rise to Wan‑Weaver, which exhibits emergent interleaved generation ability with long‑range textual coherence and visual consistency. Meanwhile, the integration of diverse understanding and generation data into planner training enables Wan‑Weaver to achieve robust task reasoning and generation proficiency. To assess the model's capability in interleaved generation, we further construct a benchmark that spans a wide range of use cases across multiple dimensions. Extensive experiments demonstrate that, even without access to any real interleaved data, Wan‑Weaver achieves superior performance over existing methods.
Authors:Yuqian Shao, Xiaosong Jia, Langechuan Liu, Junchi Yan
Abstract:
End‑to‑end autonomous driving (E2E‑AD) has achieved remarkable progress. However, one practical and useful function has been long overlooked: users may wish to customize the desired speed of the policy or specify whether to allow the autonomous vehicle to overtake. To bridge this gap, we present Bench2Drive‑Speed, a benchmark with metrics, dataset, and baselines for desired‑speed conditioned autonomous driving. We introduce explicit inputs of users' desired target‑speed and overtake/follow instructions to driving policy models. We design quantitative metrics, including Speed‑Adherence Score and Overtake Score, to measure how faithfully policies follow user specifications, while remaining compatible with standard autonomous driving metrics. To enable training of speed‑conditioned policies, one approach is to collect expert demonstrations that strictly follow speed requirements, an expensive and unscalable process in the real world. An alternative is to adapt existing regular driving data by treating the speed observed in future frames as the target speed for training. To investigate this, we construct CustomizedSpeedDataset, composed of 2,100 clips annotated with experts demonstrations, enabling systematic investigation of supervision strategies. Our experiments show that, under proper re‑annotation, models trained on regular driving data perform comparably to on expert demonstrations, suggesting that speed supervision can be introduced without additional complex real‑world data collection. Furthermore, we find that while target‑speed following can be achieved without degrading regular driving performance, executing overtaking commands remains challenging due to the inherent difficulty of interactive behaviors. All code, datasets and baselines are available at https://github.com/Thinklab‑SJTU/Bench2Drive‑Speed
Authors:Wenxuan Song, Jiayi Chen, Shuai Chen, Jingbo Wang, Pengxiang Ding, Han Zhao, Yikai Qin, Xinhu Zheng, Donglin Wang, Yan Wang, Haoang Li
Abstract:
This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard supervised finetuning (SFT). Some advanced finetuning methods with auxiliary training objectives can improve performance and reduce the number of convergence steps. However, they typically incur significant computational overhead due to the additional losses from auxiliary tasks. To simultaneously achieve the enhanced capabilities of auxiliary training with the simplicity of standard SFT, we decouple the two objectives of auxiliary task training within the parameter space, namely, enhancing general capabilities and fitting task‑specific action distributions. To deliver this goal, we only need to train the model to converge on a small‑scale task set using two distinct training strategies. The difference between the resulting model parameters can then be interpreted as capability vectors provided by auxiliary tasks. These vectors are then merged with pretrained parameters to form a capability‑enhanced meta model. Moreover, when standard SFT is augmented with a lightweight orthogonal regularization loss, the merged model attains performance comparable to auxiliary finetuned baselines with reduced computational overhead. Experimental results demonstrate that this approach is highly effective across diverse robot tasks. Project page: https://chris1220313648.github.io/Fast‑dVLA/
Authors:Chengfeng Zhao, Junbo Qi, Yulou Liu, Zhiyang Dou, Minchen Li, Taku Komura, Ziwei Liu, Wenping Wang, Yuan Liu
Abstract:
Simulating physically realistic garment deformations is an essential task for virtual immersive experience, which is often achieved by physics simulation methods. However, these methods are typically time‑consuming, computationally demanding, and require costly hardware, which is not suitable for real‑time applications. Recent learning‑based methods tried to resolve this problem by training graph neural networks to learn the garment deformation on vertices, which, however, fail to capture the intricate deformation of complex garment meshes with complex topologies. In this paper, we introduce a novel neural deformation field‑based method, named UNIC, to animate the garments of an avatar in real time, given the motion sequences. Our key idea is to learn the instance‑specific neural deformation field to animate the garment meshes. Such an instance‑specific learning scheme does not require UNIC to generalize to new garments but only to new motion sequences, which greatly reduces the difficulty in training and improves the deformation quality. Moreover, neural deformation fields map the 3D points to their deformation offsets, which not only avoids handling topologies of the complex garments but also injects a natural smoothness constraint in the deformation learning. Extensive experiments have been conducted on various kinds of garment meshes to demonstrate the effectiveness and efficiency of UNIC over baseline methods, making it potentially practical and useful in real‑world interactive applications like video games.
Authors:Yihao Wang, Yang Miao, Wenshuai Zhao, Wenyan Yang, Zihan Wang, Joni Pajarinen, Luc Van Gool, Danda Pani Paudel, Juho Kannala, Xi Wang, Arno Solin
Abstract:
Articulation perception aims to recover the motion and structure of articulated objects (e.g., drawers and cupboards), and is fundamental to 3D scene understanding in robotics, simulation, and animation. Existing learning‑based methods rely heavily on supervised training with high‑quality 3D data and manual annotations, limiting scalability and diversity. To address this limitation, we propose PAWS, a method that directly extracts object articulations from hand‑object interactions in large‑scale in‑the‑wild egocentric videos. We evaluate our method on the public data sets, including HD‑EPIC and Arti4D data sets, achieving significant improvements over baselines. We further demonstrate that the extracted articulations benefit downstream tasks, including fine‑tuning 3D articulation prediction models and enabling robot manipulation. See the project website at https://aaltoml.github.io/PAWS/.
Authors:Yufeng Yang, Xianfang Zeng, Zhangqi Jiang, Fukun Yin, Jianzhuang Liu, Wei Cheng, jinghong lan, Shiyu Liu, Yuqi Peng, Gang YU, Shifeng Chen
Abstract:
Image restoration under real‑world degradations is critical for downstream tasks such as autonomous driving and object detection. However, existing restoration models are often limited by the scale and distribution of their training data, resulting in poor generalization to real‑world scenarios. Recently, large‑scale image editing models have shown strong generalization ability in restoration tasks, especially for closed‑source models like Nano Banana Pro, which can restore images while preserving consistency. Nevertheless, achieving such performance with those large universal models requires substantial data and computational costs. To address this issue, we construct a large‑scale dataset covering nine common real‑world degradation types and train a state‑of‑the‑art open‑source model to narrow the gap with closed‑source alternatives. Furthermore, we introduce RealIR‑Bench, which contains 464 real‑world degraded images and tailored evaluation metrics focusing on degradation removal and consistency preservation. Extensive experiments demonstrate our model ranks first among open‑source methods, achieving state‑of‑the‑art performance.
Authors:Xuzhi Wang, Xinran Wu, Song Wang, Lingdong Kong, Ziping Zhao
Abstract:
Indoor monocular semantic scene completion (MSSC) is notably more challenging than its outdoor counterpart due to complex spatial layouts and severe occlusions. While transformers are well suited for modeling global dependencies, their high memory cost and difficulty in reconstructing fine‑grained details have limited their use in indoor MSSC. To address these limitations, we introduce AdaSFormer, a serialized transformer framework tailored for indoor MSSC. Our model features three key designs: (1) an Adaptive Serialized Transformer with learnable shifts that dynamically adjust receptive fields; (2) a Center‑Relative Positional Encoding that captures spatial information richness; and (3) a Convolution‑Modulated Layer Normalization that bridges heterogeneous representations between convolutional and transformer features. Extensive experiments on NYUv2 and Occ‑ScanNet demonstrate that AdaSFormer achieves state‑of‑the‑art performance. The code is publicly available at: https://github.com/alanWXZ/AdaSFormer.
Authors:Huizhi Liang, Yichao Shen, Yu Deng, Sicheng Xu, Zhiyuan Feng, Tong Zhang, Yaobo Liang, Jiaolong Yang
Abstract:
Achieving human‑like spatial intelligence for vision‑language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high‑level spatial reasoning. In this paper, we propose a principled hierarchical framework that decomposes the learning of 3D spatial understanding in VLMs into four progressively complex levels, from geometric perception to abstract spatial reasoning. Guided by this framework, we construct an automated pipeline that processes approximately 5M images with over 45M objects to generate 3D spatial VQA pairs across diverse tasks and scenes for VLM supervised fine‑tuning. We also develop an RGB‑D VLM incorporating metric‑scale point maps as auxiliary inputs to further enhance spatial understanding. Extensive experiments demonstrate that our approach achieves state‑of‑the‑art performance on multiple spatial understanding and reasoning benchmarks, surpassing specialized spatial models and large proprietary systems such as Gemini‑2.5‑pro and GPT‑5. Moreover, our analysis reveals clear dependencies among hierarchical task levels, offering new insights into how multi‑level task design facilitates the emergence of 3D spatial intelligence.
Authors:Xinkai Wang, Chenyi Wang, Yifu Xu, Mingzhe Ye, Fu-Cheng Zhang, Jialin Tian, Xinyu Zhan, Lifeng Zhu, Cewu Lu, Lixin Yang
Abstract:
We introduce LaMP, a dual‑expert Vision‑Language‑Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipulation. Existing VLA models regress actions directly from 2D semantic visual features, forcing them to learn complex 3D physical interactions implicitly. This implicit learning strategy degrades under unfamiliar spatial dynamics. LaMP addresses this limitation by aligning a flow‑matching \emphMotion Expert with a policy‑predicting \emphAction Expert through gated cross‑attention. Specifically, the Motion Expert generates a one‑step partially denoised 3D scene flow, and its hidden states condition the Action Expert without full multi‑step reconstruction. We evaluate LaMP on the LIBERO, LIBERO‑Plus, and SimplerEnv‑WidowX simulation benchmarks as well as real‑world experiments. LaMP consistently outperforms evaluated VLA baselines across LIBERO, LIBERO‑Plus, and SimplerEnv‑WidowX benchmarks, achieving the highest reported average success rates under the same training budgets. On LIBERO‑Plus OOD perturbations, LaMP shows improved robustness with an average 9.7% gain over the strongest prior baseline. Our project page is available at https://summerwxk.github.io/lamp‑project‑page/.
Authors:Niccolò Cavagnero, Narges Norouzi, Gijs Dubbelman, Daan de Geus
Abstract:
Vision Foundation Models (VFMs) pre‑trained at scale enable a single frozen encoder to serve multiple downstream tasks simultaneously. Recent VFM‑based encoder‑only models for image and video segmentation, such as EoMT and VidEoMT, achieve competitive accuracy with remarkably low latency, yet they require finetuning the encoder, sacrificing the multi‑task encoder sharing that makes VFMs practically attractive for large‑scale deployment. To reconcile encoder‑only simplicity and speed with frozen VFM features, we propose the Plain Mask Decoder (PMD), a fast Transformer‑based segmentation decoder that operates on top of frozen VFM features. The resulting model, the Plain Mask Transformer (PMT), preserves the architectural simplicity and low latency of encoder‑only designs while keeping the encoder representation unchanged and shareable. The design seamlessly applies to both image and video segmentation, inheriting the generality of the encoder‑only framework. On standard image segmentation benchmarks, PMT matches the frozen‑encoder state of the art while running up to ~3x faster. For video segmentation, it even performs on par with fully finetuned methods, while being up to 8x faster than state‑of‑the‑art frozen‑encoder models. Code: https://github.com/tue‑mps/pmt.
Authors:Yingmei Zhang, Wangtao Bao, Yong Yang, Weiguo Wan, Qin Xiao, Xueting Zou
Abstract:
Infrared small target detection (IRSTD) aims to identify and distinguish small targets from complex backgrounds. Leveraging the powerful multi‑scale feature fusion capability of the U‑Net architecture, IRSTD has achieved significant progress. However, U‑Net suffers from semantic degradation when transferring high‑level features from deep to shallow layers, limiting the precise localization of small targets. To address this issue, this paper proposes FSGNet, a lightweight and effective detection framework incorporating frequency‑aware and semantic guidance mechanisms. Specifically, a multi‑directional interactive attention module is proposed throughout the encoder to capture fine‑grained and directional features, enhancing the network's sensitivity to small, low‑contrast targets. To suppress background interference propagated through skip connections, a multi‑scale frequency‑aware module leverages Fast Fourier transform to filter out target‑similar clutter while preserving salient target structures. At the deepest layer, a global pooling module captures high‑level semantic information, which is subsequently upsampled and propagated to each decoder stage through the global semantic guidance flows, ensuring semantic consistency and precise localization across scales. Extensive experiments on four public IRSTD datasets demonstrate that FSGNet achieves superior detection performance and maintains high efficiency, highlighting its practical applicability and robustness. The codes will be released on https://github.com/Wangtao‑Bao/FSGNet.
Authors:Shengbin Guo, Hang Zhao, Senqiao Yang, Chenyang Jiang, Yuhang Cheng, Xiangru Peng, Rui Shao, Zhuotao Tian
Abstract:
Multimodal dataset distillation aims to construct compact synthetic datasets that enable efficient compression and knowledge transfer from large‑scale image‑text data. However, existing approaches often fail to capture the complex, dynamically evolving knowledge embedded in the later training stages of teacher models. This limitation leads to degraded student performance and compromises the quality of the distilled data. To address critical challenges such as pronounced cross‑stage performance gaps and unstable teacher trajectories, we propose Phased Teacher Model with Shortcut Trajectory (PTM‑ST) ‑‑ a novel phased distillation framework. PTM‑ST leverages stage‑aware teacher modeling and a shortcut‑based trajectory construction strategy to accurately fit the teacher's learning dynamics across distinct training phases. This enhances both the stability and expressiveness of the distillation process. Through theoretical analysis and comprehensive experiments, we show that PTM‑ST significantly mitigates optimization oscillations and inter‑phase knowledge gaps, while also reducing storage overhead. Our method consistently surpasses state‑of‑the‑art baselines on Flickr30k and COCO, achieving up to 13.5% absolute improvement and an average gain of 9.53% on Flickr30k. Code: https://github.com/Previsior/PTM‑ST.
Authors:Yongsung Kim, Wooseok Song, Jaihyun Lew, Hun Hwangbo, Jaehoon Lee, Sungroh Yoon
Abstract:
Visual Geometry Grounded Transformer (VGGT) has advanced 3D vision, yet its global attention layers suffer from quadratic computational costs that hinder scalability. Several sparsification‑based acceleration techniques have been proposed to alleviate this issue, but they often suffer from substantial accuracy degradation. We hypothesize that the accuracy degradation stems from the heterogeneity in head‑wise sparsification sensitivity, as the existing methods apply a uniform sparsity pattern across all heads. Motivated by this hypothesis, we present a two‑stage sparsification pipeline that effectively quantifies and exploits headwise sparsification sensitivity. In the first stage, we measure head‑wise sparsification sensitivity using a novel metric, the Head Sensitivity Score (HeSS), which approximates the Hessian with respect to two distinct error terms on a small calibration set. In the inference stage, we perform HeSS‑Guided Sparsification, leveraging the pre‑computed HeSS to reallocate the total attention budget‑assigning denser attention to sensitive heads and sparser attention to more robust ones. We demonstrate that HeSS effectively captures head‑wise sparsification sensitivity and empirically confirm that attention heads in the global attention layers exhibit heterogeneous sensitivity characteristics. Extensive experiments further show that our method effectively mitigates performance degradation under high sparsity, demonstrating strong robustness across varying sparsification levels. Code is available at https://github.com/libary753/HeSS.
Authors:Yunuo Chen, Bing He, Zezheng Lyu, Hongwei Hu, Qunshan Gu, Yuan Tian, Guo Lu
Abstract:
Efficient image compression relies on modeling both local and global redundancy. Most state‑of‑the‑art (SOTA) learned image compression (LIC) methods are based on CNNs or Transformers, which are inherently rigid. Standard CNN kernels and window‑based attention mechanisms impose fixed receptive fields and static connectivity patterns, which potentially couple non‑redundant pixels simply due to their proximity in Euclidean space. This rigidity limits the model's ability to adaptively capture spatially varying redundancy across the image, particularly at the global level. To overcome these limitations, we propose a content‑adaptive image compression framework based on Graph Neural Networks (GNNs). Specifically, our approach constructs dual‑scale graphs that enable flexible, data‑driven receptive fields. Furthermore, we introduce adaptive connectivity by dynamically adjusting the number of neighbors for each node based on local content complexity. These innovations empower our Graph‑based Learned Image Compression (GLIC) model to effectively model diverse redundancy patterns across images, leading to more efficient and adaptive compression. Experiments demonstrate that GLIC achieves state‑of‑the‑art performance, achieving BD‑rate reductions of 19.29%, 21.69%, and 18.71% relative to VTM‑9.1 on Kodak, Tecnick, and CLIC, respectively. Code will be released at https://github.com/UnoC‑727/GLIC.
Authors:Weijia Li, Haoen Xiang, Tianxu Wang, Shuaibing Wu, Qiming Xia, Cheng Wang, Chenglu Wen
Abstract:
Modern autonomous vehicle perception systems are often constrained by occlusions, blind spots, and limited sensing range. While existing cooperative perception paradigms, such as Vehicle‑to‑Vehicle (V2V) and Vehicle‑to‑Infrastructure (V2I), have demonstrated their effectiveness in mitigating these challenges, they remain limited to ground‑level collaboration and cannot fully address large‑scale occlusions or long‑range perception in complex environments. To advance research in cross‑view cooperative perception, we present V2U4Real, the first large‑scale real‑world multi‑modal dataset for Vehicle‑to‑UAV (V2U) cooperative object perception. V2U4Real is collected by a ground vehicle and a UAV equipped with multi‑view LiDARs and RGB cameras. The dataset covers urban streets, university campuses, and rural roads under diverse traffic scenarios, comprising over 56K LiDAR frames, 56K multi‑view camera images, and 700K annotated 3D bounding boxes across four classes. To support a wide range of research tasks, we establish benchmarks for single‑agent 3D object detection, cooperative 3D object detection, and object tracking. Comprehensive evaluations of several state‑of‑the‑art models demonstrate the effectiveness of V2U cooperation in enhancing perception robustness and long‑range awareness. The V2U4Real dataset and codebase is available at https://github.com/VjiaLi/V2U4Real.
Authors:Yuhan Chen, Pengwen Dai, Chuan Wang, Dayan Wu, Xiaochun Cao
Abstract:
Text‑video retrieval tasks have seen significant improvements due to the recent development of large‑scale vision‑language pre‑trained models. Traditional methods primarily focus on video representations or cross‑modal alignment, while recent works shift toward enriching text expressiveness to better match the rich semantics in videos. However, these methods use only interactions between text and frames/video, and ignore rich interactions among the internal frames within a video, so the final expanded text cannot capture frame contextual information, leading to disparities between text and video. In response, we introduce Energy‑Aware Fine‑Grained Relationship Learning Network (EagleNet) to generate accurate and context‑aware enriched text embeddings. Specifically, the proposed Fine‑Grained Relationship Learning mechanism (FRL) first constructs a text‑frame graph by the generated text candidates and frames, then learns relationships among texts and frames, which are finally used to aggregate text candidates into an enriched text embedding that incorporates frame contextual information. To further improve fine‑grained relationship learning in FRL, we design Energy‑Aware Matching (EAM) to model the energy of text‑frame interactions and thus accurately capture the distribution of real text‑video pairs. Moreover, for more effective cross‑modal alignment and stable training, we replace the conventional softmax‑based contrastive loss with the sigmoid loss. Extensive experiments have demonstrated the superiority of EagleNet across MSRVTT, DiDeMo, MSVD, and VATEX. Codes are available at https://github.com/draym28/EagleNet.
Authors:Pengpeng Yu, Haoran Li, Runqing Jiang, Dingquan Li, Jing Wang, Liang Lin, Yulan Guo
Abstract:
LiDAR point clouds are fundamental to various applications, yet the extreme sparsity of high‑precision geometric details hinders efficient context modeling, thereby limiting the compression speed and performance of existing methods. To address this challenge, we propose a compact representation for efficient predictive lossless coding. Our framework comprises two lightweight modules. First, the Geometry Re‑Densification Module iteratively densifies encoded sparse geometry, extracts features at a dense scale, and then sparsifies the features for predictive coding. This module avoids costly computation on highly sparse details while maintaining a lightweight prediction head. Second, the Cross‑scale Feature Propagation Module leverages occupancy cues from multiple resolution levels to guide hierarchical feature propagation, enabling information sharing across scales and reducing redundant feature extraction. Additionally, we introduce an integer‑only inference pipeline to enable bit‑exact cross‑platform consistency, which avoids the entropy‑coding collapse observed in existing neural compression methods and further accelerates coding. Experiments demonstrate competitive compression performance at real‑time speed. Code will be released upon acceptance. Code is available at https://github.com/pengpeng‑yu/FastPCC.
Authors:Yabin Zhang, Maya Varma, Yunhe Gao, Jean-Benoit Delbrouck, Jiaming Liu, Chong Wang, Curtis Langlotz
Abstract:
Out‑of‑distribution (OOD) detection aims to identify samples that deviate from in‑distribution (ID). One popular pipeline addresses this by introducing negative labels distant from ID classes and detecting OOD based on their distance to these labels. However, such labels may present poor activation on OOD samples, failing to capture the OOD characteristics. To address this, we propose \underlineTest‑time \underlineActivated \underlineNegative \underlineLabels (TANL) by dynamically evaluating activation levels across the corpus dataset and mining candidate labels with high activation responses during the testing process. Specifically, TANL identifies high‑confidence test images online and accumulates their assignment probabilities over the corpus to construct a label activation metric. Such a metric leverages historical test samples to adaptively align with the test distribution, enabling the selection of distribution‑adaptive activated negative labels. By further exploring the activation information within the current testing batch, we introduce a more fine‑grained, batch‑adaptive variant. To fully utilize label activation knowledge, we propose an activation‑aware score function that emphasizes negative labels with stronger activations, boosting performance and enhancing its robustness to the label number. Our TANL is training‑free, test‑efficient, and grounded in theoretical justification. Experiments on diverse backbones and wide task settings validate its effectiveness. Notably, on the large‑scale ImageNet benchmark, TANL significantly reduces the FPR95 from 17.5% to 9.8%. Codes are available at \hrefhttps://github.com/YBZh/OpenOOD‑VLMYBZh/OpenOOD‑VLM.
Authors:Taejin Jeong, Joohyeok Kim, Jinyeong Kim, Chanyoung Kim, Seong Jae Hwang
Abstract:
Spatial Transcriptomics (ST) provides spatially‑resolved gene expression, offering crucial insights into tissue architecture and complex diseases. However, its prohibitive cost limits widespread adoption, leading to significant attention on inferring spatial gene expression from readily available whole slide images. While graph neural networks have been proposed to model interactions between tissue regions, their reliance on pre‑defined sparse graphs prevents them from considering potentially interacting spot pairs, resulting in a structural limitation in capturing complex biological relationships. To address this, we propose FEAST (Fully connected Expressive Attention for Spatial Transcriptomics), an attention‑based framework that models the tissue as a fully connected graph, enabling the consideration of all pairwise interactions. To better reflect biological interactions, we introduce negative‑aware attention, which models both excitatory and inhibitory interactions, capturing essential negative relationships that standard attention often overlooks. Furthermore, to mitigate the information loss from truncated or ignored context in standard spot image extraction, we introduce an off‑grid sampling strategy that gathers additional images from intermediate regions, allowing the model to capture a richer morphological context. Experiments on public ST datasets show that FEAST surpasses state‑of‑the‑art methods in gene expression prediction while providing biologically plausible attention maps that clarify positive and negative interactions. Our code is available at https://github.com/starforTJ/ FEAST.
Authors:Jiahao Tian, Chenxi Song, Wei Cheng, Chi Zhang
Abstract:
Generating long videos using pre‑trained video diffusion models, which are typically trained on short clips, presents a significant challenge. Directly applying these models for long‑video inference often leads to a notable degradation in visual quality. This paper identifies that this issue primarily stems from two out‑of‑distribution (O.O.D) problems: frame‑level relative position O.O.D and context‑length O.O.D. To address these challenges, we propose FreeLOC, a novel training‑free, layer‑adaptive framework that introduces two core techniques: Video‑based Relative Position Re‑encoding (VRPR) for frame‑level relative position O.O.D, a multi‑granularity strategy that hierarchically re‑encodes temporal relative positions to align with the model's pre‑trained distribution, and Tiered Sparse Attention (TSA) for context‑length O.O.D, which preserves both local detail and long‑range dependencies by structuring attention density across different temporal scales. Crucially, we introduce a layer‑adaptive probing mechanism that identifies the sensitivity of each transformer layer to these O.O.D issues, allowing for the selective and efficient application of our methods. Extensive experiments demonstrate that our approach significantly outperforms existing training‑free methods, achieving state‑of‑the‑art results in both temporal consistency and visual quality. Code is available at https://github.com/Westlake‑AGI‑Lab/FreeLOC.
Authors:Marvin Seyfarth, Sarah Kaye Müller, Arman Ghanaat, Isabelle Ayx, Fabian Fastenrath, Philipp Wild, Alexander Hertel, Theano Papavassiliu, Salman Ul Hassan Dar, Sandy Engelhardt
Abstract:
Latent diffusion models (LDMs) have recently achieved strong performance in 3D medical image synthesis. However, modalities like cine cardiac MRI (CMR), representing a temporally synchronized 3D volume across the cardiac cycle, add an additional dimension that most generative approaches do not model directly. Instead, they factorize space and time or enforce temporal consistency through auxiliary mechanisms such as anatomical masks. Such strategies introduce structural biases that may limit global context integration and lead to subtle spatiotemporal discontinuities or physiologically inconsistent cardiac dynamics. We investigate whether a unified 4D generative model can learn continuous cardiac dynamics without architectural factorization. We propose CardioDiT, a fully 4D latent diffusion framework for short‑axis cine CMR synthesis based on diffusion transformers. A spatiotemporal VQ‑VAE encodes 2D+t slices into compact latents, which a diffusion transformer then models jointly as complete 3D+t volumes, coupling space and time throughout the generative process. We evaluate CardioDiT on public CMR datasets and a larger private cohort, comparing it to baselines with progressively stronger spatiotemporal coupling. Results show improved inter‑slice consistency, temporally coherent motion, and realistic cardiac function distributions, suggesting that explicit 4D modeling with a diffusion transformer provides a principled foundation for spatiotemporal cardiac image synthesis. Code and models trained on public data are available at https://github.com/Cardio‑AI/cardiodit.
Authors:Marvin Seyfarth, Salman Ul Hassan Dar, Yannik Frisch, Philipp Wild, Norbert Frey, Florian André, Sandy Engelhardt
Abstract:
Diffusion models have become a leading approach for high‑fidelity medical image synthesis. However, most existing methods for 3D medical image generation rely on convolutional U‑Net backbones within latent diffusion frameworks. While effective, these architectures impose strong locality biases and limited receptive fields, which may constrain scalability, global context integration, and flexible conditioning. In this work, we introduce VolDiT, the first purely transformer‑based 3D Diffusion Transformer for volumetric medical image synthesis. Our approach extends diffusion transformers to native 3D data through volumetric patch embeddings and global self‑attention operating directly over 3D tokens. To enable structured control, we propose a timestep‑gated control adapter that maps segmentation masks into learnable control tokens that modulate transformer layers during denoising. This token‑level conditioning mechanism allows precise spatial guidance while preserving the modeling advantages of transformer architectures. We evaluate our model on high‑resolution 3D medical image synthesis tasks and compare it to state‑of‑the‑art 3D latent diffusion models based on U‑Nets. Results demonstrate improved global coherence, superior generative fidelity, and enhanced controllability. Our findings suggest that fully transformerbased diffusion models provide a flexible foundation for volumetric medical image synthesis. The code and models trained on public data are available at https://github.com/Cardio‑AI/voldit.
Authors:Wanjiang Weng, Xiaofeng Tan, Xiangbo Shu, Guo-Sen Xie, Pan Zhou, Hongsong Wang
Abstract:
Text‑to‑motion generation holds significant potential for cross‑linguistic applications, yet it is hindered by the lack of bilingual datasets and the poor cross‑lingual semantic understanding of existing language models. To address these gaps, we introduce BiHumanML3D, the first bilingual text‑to‑motion benchmark, constructed via LLM‑assisted annotation and rigorous manual correction. Furthermore, we propose a simple yet effective baseline, Bilingual Motion Diffusion (BiMD), featuring Cross‑Lingual Alignment (CLA). CLA explicitly aligns semantic representations across languages, creating a robust conditional space that enables high‑quality motion generation from bilingual inputs, including zero‑shot code‑switching scenarios. Extensive experiments demonstrate that BiMD with CLA achieves an FID of 0.045 vs. 0.169 and R@3 of 82.8% vs. 80.8%, significantly outperforms monolingual diffusion models and translation baselines on BiHumanML3D, underscoring the critical necessity and reliability of our dataset and the effectiveness of our alignment strategy for cross‑lingual motion synthesis. The dataset and code are released at \hrefhttps://wengwanjiang.github.io/BilingualT2M‑pagehttps://wengwanjiang.github.io/BilingualT2M‑page
Authors:Md Mushfiqur Azam, John Quarles, Kevin Desai
Abstract:
Egocentric 3D human pose estimation remains challenging due to severe perspective distortion, limited body visibility, and complex camera motion inherent in first‑person viewpoints. Existing methods typically rely on single‑frame analysis or limited temporal fusion, which fails to effectively leverage the rich motion context available in egocentric videos. We introduce AG‑EgoPose, a novel dual‑stream framework that integrates short‑ and long‑range motion context with fine‑grained spatial cues for robust pose estimation from fisheye camera input. Our framework features two parallel streams: A spatial stream uses a weight‑sharing ResNet‑18 encoder‑decoder to generate 2D joint heatmaps and corresponding joint‑specific spatial feature tokens. Simultaneously, a temporal stream uses a ResNet‑50 backbone to extract visual features, which are then processed by an action recognition backbone to capture the motion dynamics. These complementary representations are fused and refined in a transformer decoder with learnable joint tokens, which allows for the joint‑level integration of spatial and temporal evidence while maintaining anatomical constraints. Experiments on real‑world datasets demonstrate that AG‑EgoPose achieves state‑of‑the‑art performance in both quantitative and qualitative metrics. Code is available at: https://github.com/Mushfiq5647/AG‑EgoPose.
Authors:Taegyoon Yoon, Yegyu Han, Seojin Ji, Jaewoo Park, Sojeong Kim, Taein Kwon, Hyung-Sin Kim
Abstract:
Smart glass is emerging as an useful device since it provides plenty of insights under hands‑busy, eyes‑on‑task situations. To understand the context of the wearer, 6D object pose estimation in egocentric view is becoming essential. However, existing 6D object pose estimation benchmarks fail to capture the challenges of real‑world egocentric applications, which are often dominated by severe motion blur, dynamic illumination, and visual obstructions. This discrepancy creates a significant gap between controlled lab data and chaotic real‑world application. To bridge this gap, we introduce EgoXtreme, a new large‑scale 6D pose estimation dataset captured entirely from an egocentric perspective. EgoXtreme features three challenging scenarios ‑ industrial maintenance, sports, and emergency rescue ‑ designed to introduce severe perceptual ambiguities through extreme lighting, heavy motion blur, and smoke. Evaluations of state‑of‑the‑art generalizable pose estimators on EgoXtreme indicate that their generalization fails to hold in extreme conditions, especially under low light. We further demonstrate that simply applying image restoration (e.g., deblurring) offers no positive improvement for extreme conditions. While performance gain has appeared in tracking‑based approach, implying using temporal information in fast‑motion scenarios is meaningful. We conclude that EgoXtreme is an essential resource for developing and evaluating the next generation of pose estimation models robust enough for real‑world egocentric vision. The dataset and code are available at https://taegyoun88.github.io/EgoXtreme/
Authors:Yinjian Wang, Wei Li, Yuanyuan Gui, James E. Fowler, Gemine Vivone
Abstract:
Robust principal component analysis (RPCA) seeks a low‑rank component and a sparse component from their summation. Yet, in many applications of interest, the sparse foreground actually replaces, or occludes, elements from the low‑rank background. To address this mismatch, a new framework is proposed in which the sparse component is identified indirectly through determining its support. This approach, called robust principal component completion (RPCC), is solved via variational Bayesian inference applied to a fully probabilistic Bayesian sparse tensor factorization. Convergence to a hard classifier for the support is shown, thereby eliminating the post‑hoc thresholding required of most prior RPCA‑driven approaches. Experimental results reveal that the proposed approach delivers near‑optimal estimates on synthetic data as well as robust foreground‑extraction and anomaly‑detection performance on real color video and hyperspectral datasets, respectively. Source implementation and Appendices are available at https://github.com/WongYinJ/BCP‑RPCC.
Authors:Minh-Quan Viet Bui, Jaeho Moon, Munchurl Kim
Abstract:
While 3D Vision Foundation Models (3DVFMs) have demonstrated remarkable zero‑shot capabilities in visual geometry estimation, their direct application to generalizable novel view synthesis (NVS) remains challenging.
In this paper, we propose AirSplat, a novel training framework that effectively adapts the robust geometric priors of 3DVFMs into high‑fidelity, pose‑free NVS. Our approach introduces two key technical contributions:
(1) Self‑Consistent Pose Alignment (SCPA), a training‑time feedback loop that ensures pixel‑aligned supervision to resolve pose‑geometry discrepancy; and
(2) Rating‑based Opacity Matching (ROM), which leverages the local 3D geometry consistency knowledge from a sparse‑view NVS teacher model to filter out degraded primitives.
Experimental results on large‑scale benchmarks demonstrate that our method significantly outperforms state‑of‑the‑art pose‑free NVS approaches in reconstruction quality.
Our AirSplat highlights the potential of adapting 3DVFMs to enable simultaneous visual geometry estimation and high‑quality view synthesis.
Authors:Chenglong Wang, Yifu Huo, Yang Gan, Qiaozhi He, Qi Meng, Bei Li, Yan Wang, Junfu Liu, Tianhua Zhou, Jingbo Zhu, Tong Xiao
Abstract:
Recent advances in multimodal reward modeling have been largely driven by a paradigm shift from discriminative to generative approaches. Building on this progress, recent studies have further employed reinforcement learning from verifiable rewards (RLVR) to enhance multimodal reward models (MRMs). Despite their success, RLVR‑based training typically relies on labeled multimodal preference data, which are costly and labor‑intensive to obtain, making it difficult to scale MRM training. To overcome this limitation, we propose a Multi‑Stage Reinforcement Learning (MSRL) approach, which can achieve scalable RL for MRMs with limited multimodal data. MSRL replaces the conventional RLVR‑based training paradigm by first learning a generalizable reward reasoning capability from large‑scale textual preference data, and then progressively transferring this capability to multimodal tasks through caption‑based and fully multimodal reinforcement‑learning stages. Furthermore, we introduce a cross‑modal knowledge distillation approach to improve preference generalization within MSRL. Extensive experiments demonstrate that MSRL effectively scales the RLVR‑based training of generative MRMs and substantially improves their performance across both visual understanding and visual generation tasks (e.g., from 66.6% to 75.9% on VL‑RewardBench and from 70.2% to 75.7% on GenAI‑Bench), without requiring additional multimodal preference annotations. Our code is available at: https://github.com/wangclnlp/MSRL.
Authors:Xuankai Zhang, Junjin Xiao, Shangwei Huang, Wei-shi Zheng, Qing Zhang
Abstract:
We present an approach for high‑quality dynamic Gaussian Splatting from monocular videos. To this end, we in this work go one step further beyond previous methods to explicitly model continuous position and orientation deformation of dynamic Gaussians, using an SE(3) B‑spline motion bases with a compact set of control points. To improve computational efficiency while enhancing the ability to model complex motions, an adaptive control mechanism is devised to dynamically adjust the number of motion bases and control points. Besides, we develop a soft segment reconstruction strategy to mitigate long‑interval motion interference, and employ a multi‑view diffusion model to provide multi‑view cues for avoiding overfitting to training views. Extensive experiments demonstrate that our method outperforms state‑of‑the‑art methods in novel view synthesis. Our code is available at https://github.com/hhhddddddd/se3bsplinegs.
Authors:Yinyi Luo, Hrishikesh Gokhale, Marios Savvides, Jindong Wang, Shengfeng He
Abstract:
Despite significant progress in text‑to‑image generation, aligning outputs with complex prompts remains challenging, particularly for fine‑grained semantics and spatial relations. This difficulty stems from the feed‑forward nature of generation, which requires anticipating alignment without fully understanding the output. In contrast, evaluating generated images is more tractable. Motivated by this asymmetry, we propose xLARD, a self‑correcting framework that uses multimodal large language models to guide generation through Explainable LAtent RewarDs. xLARD introduces a lightweight corrector that refines latent representations based on structured feedback from model‑generated references. A key component is a differentiable mapping from latent edits to interpretable reward signals, enabling continuous latent‑level guidance from non‑differentiable image‑level evaluations. This mechanism allows the model to understand, assess, and correct itself during generation. Experiments across diverse generation and editing tasks show that xLARD improves semantic alignment and visual fidelity while maintaining generative priors. Code is available at https://yinyiluo.github.io/xLARD/.
Authors:Jing Yang, Krithika Dharanikota, Emily Jia, Haiwei Chen, Yajie Zhao
Abstract:
Accurately modeling how real‑world materials reflect light remains a core challenge in inverse rendering, largely due to the scarcity of real measured reflectance data. Existing approaches rely heavily on synthetic datasets with simplified illumination and limited material realism, preventing models from generalizing to real‑world images. We introduce a large‑scale polarized reflection and material dataset of real‑world objects, captured with an 8‑camera, 346‑light Light Stage equipped with cross/parallel polarization. Our dataset spans 218 everyday objects across five acquisition dimensions‑multiview, multi‑illumination, polarization, reflectance separation, and material attributes‑yielding over 1.2M high‑resolution images with diffuse‑specular separation and analytically derived diffuse albedo, specular albedo, and surface normals. Using this dataset, we train and evaluate state‑of‑the‑art inverse and forward rendering models on intrinsic decomposition, relighting, and sparse‑view 3D reconstruction, demonstrating significant improvements in material separation, illumination fidelity, and geometric consistency. We hope that our work can establish a new foundation for physically grounded material understanding and enable real‑world generalization beyond synthetic training regimes. Project page: https://jingyangcarl.github.io/ICTPolarReal/
Authors:Luyu Yang, Yutong Dai, An Yan, Viraj Prabhu, Ran Xu, Zeyuan Chen
Abstract:
The physical world is not merely visual; it is governed by rigorous structural and procedural constraints. Yet, the evaluation of vision‑language models (VLMs) remains heavily skewed toward perceptual realism, prioritizing the generation of visually plausible 3D layouts, shapes, and appearances. Current benchmarks rarely test whether models grasp the step‑by‑step processes and physical dependencies required to actually build these artifacts, a capability essential for automating design‑to‑construction pipelines. To address this, we introduce DreamHouse, a novel benchmark for physical generative reasoning: the capacity to synthesize artifacts that concurrently satisfy geometric, structural, constructability, and code‑compliance constraints. We ground this benchmark in residential timber‑frame construction, a domain with fully codified engineering standards and objectively verifiable correctness. We curate over 26,000 structures spanning 13 architectural styles, ach verified to construction‑document standards (LOD 350) and develop a deterministic 10‑test structural validation framework. Unlike static benchmarks that assess only final outputs, DreamHouse supports iterative agentic interaction. Models observe intermediate build states, generate construction actions, and receive structured environmental feedback, enabling a fine‑grained evaluation of planning, structural reasoning, and self‑correction. Extensive experiments with state‑of‑the‑art VLMs reveal substantial capability gaps that are largely invisible on existing leaderboards. These findings establish physical validity as a critical evaluation axis orthogonal to visual realism, highlighting physical generative reasoning as a distinct and underdeveloped frontier in multimodal intelligence. Available at https://luluyuyuyang.github.io/dreamhouse
Authors:Yihan Wang, Jia Deng
Abstract:
We introduce WAFT‑Stereo, a simple and effective warping‑based method for stereo matching. WAFT‑Stereo demonstrates that cost volumes, a common design used in many leading methods, are not necessary for strong performance and can be replaced by warping with improved efficiency. WAFT‑Stereo ranks first on ETH3D (BP‑0.5), Middlebury (RMSE), and KITTI (all metrics), reducing the zero‑shot error by 81% on ETH3D, while being 1.8‑6.7x faster than competitive methods. Code and model weights are available at https://github.com/princeton‑vl/WAFT‑Stereo.
Authors:Junyi Ouyang, Wenbin Teng, Gonglin Chen, Yajie Zhao, Haiwei Chen
Abstract:
Long‑trajectory video generation is a crucial yet challenging task for world modeling primarily due to the limited scalability of existing video diffusion models (VDMs). Autoregressive models, while offering infinite rollout, suffer from visual drift and poor controllability. To address these issues, we propose DCARL, a novel divide‑and‑conquer, autoregressive framework that effectively combines the structural stability of the divide‑and‑conquer scheme with the high‑fidelity generation of VDMs. Our approach first employs a dedicated Keyframe Generator trained without temporal compression to establish long‑range, globally consistent structural anchors. Subsequently, an Interpolation Generator synthesizes the dense frames in an autoregressive manner with overlapping segments, utilizing the keyframes for global context and a single clean preceding frame for local coherence. Trained on a large‑scale internet long trajectory video dataset, our method achieves superior performance in both visual quality (lower FID and FVD) and camera adherence (lower ATE and ARE) compared to state‑of‑the‑art autoregressive and divide‑and‑conquer baselines, demonstrating stable and high‑fidelity generation for long trajectory videos up to 32 seconds in length.
Authors:Alabi Mehzabin Anisha, Guangjing Wang, Sriram Chellappan
Abstract:
State‑of‑the‑art crowd counting and localization are primarily modeled using two paradigms: density maps and point regression. Given the field's security ramifications, there is active interest in model robustness against adversarial attacks. Recent studies have demonstrated transferability across density‑map‑based approaches via adversarial patches, but cross‑paradigm attacks (i.e., across both density map‑based models and point regression‑based models) remain unexplored. We introduce a novel adversarial framework that compromises both density map and point regression architectural paradigms through a comprehensive multi‑task loss optimization. For point‑regression models, we employ scene‑density‑specific high‑confidence logit suppression; for density‑map approaches, we use peak‑targeted density map suppression. Both are combined with model‑agnostic perceptual constraints to ensure that perturbations are effective and imperceptible to the human eye. Extensive experiments demonstrate the effectiveness of our attack, achieving on average a 7X increase in Mean Absolute Error compared to clean images while maintaining competitive visual quality, and successfully transferring across seven state‑of‑the‑art crowd models with transfer ratios ranging from 0.55 to 1.69. Our approach strikes a balance between attack effectiveness and imperceptibility compared to state‑of‑the‑art transferable attack strategies. The source code is available at https://github.com/simurgh7/CrowdGen
Authors:Danil Tokhchukov, Aysel Mirzoeva, Andrey Kuznetsov, Konstantin Sobolev
Abstract:
In this paper, we uncover the hidden potential of Diffusion Transformers (DiTs) to significantly enhance generative tasks. Through an in‑depth analysis of the denoising process, we demonstrate that introducing a single learned scaling parameter can significantly improve the performance of DiT blocks. Building on this insight, we propose Calibri, a parameter‑efficient approach that optimally calibrates DiT components to elevate generative quality. Calibri frames DiT calibration as a black‑box reward optimization problem, which is efficiently solved using an evolutionary algorithm and modifies just ~100 parameters. Experimental results reveal that despite its lightweight design, Calibri consistently improves performance across various text‑to‑image models. Notably, Calibri also reduces the inference steps required for image generation, all while maintaining high‑quality outputs.
Authors:Matan Ben-Yosef, Tavi Halperin, Naomi Ken Korem, Mohammad Salama, Harel Cain, Asaf Joseph, Anthony Chen, Urska Jelercic, Ofir Bibi
Abstract:
Controlling video and audio generation requires diverse modalities, from depth and pose to camera trajectories and audio transformations, yet existing approaches either train a single monolithic model for a fixed set of controls or introduce costly architectural changes for each new modality. We introduce AVControl, a lightweight, extendable framework built on LTX‑2, a joint audio‑visual foundation model, where each control modality is trained as a separate LoRA on a parallel canvas that provides the reference signal as additional tokens in the attention layers, requiring no architectural changes beyond the LoRA adapters themselves. We show that simply extending image‑based in‑context methods to video fails for structural control, and that our parallel canvas approach resolves this. On the VACE Benchmark, we outperform all evaluated baselines on depth‑ and pose‑guided generation, inpainting, and outpainting, and show competitive results on camera control and audio‑visual benchmarks. Our framework supports a diverse set of independently trained modalities: spatially‑aligned controls such as depth, pose, and edges, camera trajectory with intrinsics, sparse motion control, video editing, and, to our knowledge, the first modular audio‑visual controls for a joint generation model. Our method is both compute‑ and data‑efficient: each modality requires only a small dataset and converges within a few hundred to a few thousand training steps, a fraction of the budget of monolithic alternatives. We publicly release our code and trained LoRA checkpoints.
Authors:Manglam Kartik, Neel Tushar Shah
Abstract:
Standard vision models treat objects as independent points in Euclidean space, unable to capture hierarchical structure like parts within wholes. We introduce Worldline Slot Attention, which models objects as persistent trajectories through spacetime worldlines, where each object has multiple slots at different hierarchy levels sharing the same spatial position but differing in temporal coordinates. This architecture consistently fails without geometric structure: Euclidean worldlines achieve 0.078 level accuracy, below random chance (0.33), while Lorentzian worldlines achieve 0.479‑0.661 across three datasets: a 6x improvement replicated over 20+ independent runs. Lorentzian geometry also outperforms hyperbolic embeddings showing visual hierarchies require causal structure (temporal dependency) rather than tree structure (radial branching). Our results demonstrate that hierarchical object discovery requires geometric structure encoding asymmetric causality, an inductive bias absent from Euclidean space but natural to Lorentzian light cones, achieved with only 11K parameters. The code is available at: https://github.com/iclrsubmissiongram/loco.
Authors:Lukas Radl, Felix Windisch, Andreas Kurz, Thomas Köhler, Michael Steiner, Markus Steinberger
Abstract:
Recently, 3D Gaussian Splatting (3DGS) greatly accelerated mesh extraction from posed images due to its explicit representation and fast software rasterization. While the addition of geometric losses and other priors has improved the accuracy of extracted surfaces, mesh extraction remains difficult in scenes with abundant view‑dependent effects. To resolve the resulting ambiguities, prior works rely on multi‑view techniques, iterative mesh extraction, or large pre‑trained models, sacrificing the inherent efficiency of 3DGS. In this work, we present a simple and efficient alternative by introducing a self‑supervised confidence framework to 3DGS: within this framework, learnable confidence values dynamically balance photometric and geometric supervision. Extending our confidence‑driven formulation, we introduce losses which penalize per‑primitive color and normal variance and demonstrate their benefits to surface extraction. Finally, we complement the above with an improved appearance model, by decoupling the individual terms of the D‑SSIM loss. Our final approach delivers state‑of‑the‑art results for unbounded meshes while remaining highly efficient.
Authors:Daniele Agostinelli, Thomas Agostinelli, Andrea Generosi, Maura Mengoni
Abstract:
Appearance‑based gaze estimation frequently relies on deep Convolutional Neural Networks (CNNs). These models are accurate, but computationally expensive and act as "black boxes", offering little interpretability. Geometric methods based on facial landmarks are a lightweight alternative, but their performance limits and generalization capabilities remain underexplored in modern benchmarks. In this study, we conduct a comprehensive evaluation of landmark‑based gaze estimation. We introduce a standardized pipeline to extract and normalize landmarks from three large‑scale datasets (Gaze360, ETH‑XGaze, and GazeGene) and train lightweight regression models, specifically Extreme Gradient Boosted trees and two neural architectures: a holistic Multi‑Layer Perceptron (MLP) and a siamese MLP designed to capture binocular geometry. We find that landmark‑based models exhibit lower performance in within‑domain evaluation, likely due to noise introduced into the datasets by the landmark detector. Nevertheless, in cross‑domain evaluation, the proposed MLP architectures show generalization capabilities comparable to those of ResNet18 baselines. These findings suggest that sparse geometric features encode sufficient information for robust gaze estimation, paving the way for efficient, interpretable, and privacy‑friendly edge applications. The source code and generated landmark‑based datasets are available at https://github.com/daniele‑agostinelli/LandmarkGaze.git.
Authors:Shengli Zhou, Minghang Zheng, Feng Zheng, Yang Liu
Abstract:
Spatial reasoning focuses on locating target objects based on spatial relations in 3D scenes, which plays a crucial role in developing intelligent embodied agents. Due to the limited availability of 3D scene‑language paired data, it is challenging to train models with strong reasoning ability from scratch. Previous approaches have attempted to inject 3D scene representations into the input space of Large Language Models (LLMs) and leverage the pretrained comprehension and reasoning abilities for spatial reasoning. However, models encoding absolute positions struggle to extract spatial relations from prematurely fused features, while methods explicitly encoding all spatial relations (which is quadratic in the number of objects) as input tokens suffer from poor scalability. To address these limitations, we propose QuatRoPE, a novel positional embedding method with an input length that is linear to the number of objects, and explicitly calculates pairwise spatial relations through the dot product in attention layers. QuatRoPE's holistic vector encoding of 3D coordinates guarantees a high degree of spatial consistency, maintaining fidelity to the scene's geometric integrity. Additionally, we introduce the Isolated Gated RoPE Extension (IGRE), which effectively limits QuatRoPE's influence to object‑related tokens, thereby minimizing interference with the LLM's existing positional embeddings and maintaining the LLM's original capabilities. Extensive experiments demonstrate the effectiveness of our approaches. The code and data are available at https://github.com/oceanflowlab/QuatRoPE.
Authors:Deyan Deng, Rongjun Qin
Abstract:
3D Gaussian Splatting (3DGS) has revolutionized real‑time rendering with its state‑of‑the‑art novel view synthesis, but its utility for accurate geometric measurement remains underutilized. Compared to multi‑view stereo (MVS) point clouds or meshes, 3DGS rendered views present superior visual quality and completeness. However, current point measurement methods still rely on demanding stereoscopic workstations or direct picking on often‑incomplete and inaccurate 3D meshes. As a novel view synthesizer, 3DGS renders exact source views and smoothly interpolates in‑between views. This allows users to intuitively pick congruent points across different views while operating 3DGS models. By triangulating these congruent points, one can precisely generate 3D point measurements. This approach mimics traditional stereoscopic measurement but is significantly less demanding: it requires neither a stereo workstation nor specialized operator stereoscopic capability. Furthermore, it enables multi‑view intersection (more than two views) for higher measurement accuracy. We implemented a web‑based application to demonstrate this proof‑of‑concept (PoC). Using several UAV aerial datasets, we show this PoC allows users to successfully perform highly accurate point measurements, achieving accuracy matching or exceeding traditional stereoscopic methods on standard hardware. Specifically, our approach significantly outperforms direct mesh‑based measurements. Quantitatively, our method achieves RMSEs in the 1‑2 cm range on well‑defined points. More critically, on challenging thin structures where mesh‑based RMSE was 0.062 m, our method achieved 0.037 m. On sharp corners poorly reconstructed in the mesh, our method successfully measured all points with a 0.013 m RMSE, whereas the mesh method failed entirely. Code is available at: https://github.com/GDAOSU/3dgs_measurement_tool.
Authors:Chandan Yeshwanth, Angela Dai
Abstract:
3D object understanding and generation methods produce impressive results, yet they often overlook a pervasive source of information in real‑world scenes: repeated objects. We introduce the task of lookalike object detection in indoor scenes, which leverages repeated and complementary cues from identical and near‑identical object pairs. Given an input scene, the task is to classify pairs of objects as identical, similar or different using multiview images as input. To address this, we present Lookalike3D, a multiview image transformer that effectively distinguishes such object pairs by harnessing strong semantic priors from large image foundation models. To support this task, we collected the 3DTwins dataset, containing 76k manually annotated identical, similar and different pairs of objects based on ScanNet++, and show an improvement of 104% IoU over baselines. We demonstrate how our method improves downstream tasks such as enabling joint 3D object reconstruction and part co‑segmentation, turning repeated and lookalike objects into a powerful cue for consistent, high‑quality 3D perception. Our code, dataset and models will be made publicly available.
Authors:Gokce Inal, Pouyan Navard, Alper Yilmaz
Abstract:
Recent advances in multimodal vision‑language models (VLMs) have enabled joint reasoning over visual and textual information, yet their application to planetary science remains largely unexplored. A key hindrance is the absence of large‑scale datasets that pair real planetary imagery with detailed scientific descriptions. In this work, we introduce LLaVA‑LE (Large Language‑and‑Vision Assistant for Lunar Exploration), a vision‑language model specialized for lunar surface and subsurface characterization. To enable this capability, we curate a new large‑scale multimodal lunar dataset, LUCID (LUnar Caption Image Dataset) consisting of 96k high‑resolution panchromatic images paired with detailed captions describing lunar terrain characteristics, and 81k question‑answer (QA) pairs derived from approximately 20k images in the LUCID dataset. Leveraging this dataset, we fine‑tune LLaVA using a two‑stage training curriculum: (1) concept alignment for domain‑specific terrain description, and (2) instruction‑tuned visual question answering. We further design evaluation benchmarks spanning multiple levels of reasoning complexity relevant to lunar terrain analysis. Evaluated against GPT and Gemini judges, LLaVA‑LE achieves a 3.3x overall performance gain over Base LLaVA and 2.1x over our Stage 1 model, with a reasoning score of 1.070, exceeding the judge's own reference score, highlighting the effectiveness of domain‑specific multimodal data and instruction tuning to advance VLMs in planetary exploration. Code is available at https://github.com/OSUPCVLab/LLaVA‑LE.
Authors:Bentao Song, Jun Huang, Qingfeng Wang
Abstract:
In mixed domain semi‑supervised medical image segmentation (MiDSS), achieving superior performance under domain shift and limited annotations is challenging. This scenario presents two primary issues: (1) distributional differences between labeled and unlabeled data hinder effective knowledge transfer, and (2) inefficient learning from unlabeled data causes severe confirmation bias. In this paper, we propose the bidirectional correlation maps domain adaptation (BCMDA) framework to overcome these issues. On the one hand, we employ knowledge transfer via virtual domain bridging (KTVDB) to facilitate cross‑domain learning. First, to construct a distribution‑aligned virtual domain, we leverage bidirectional correlation maps between labeled and unlabeled data to synthesize both labeled and unlabeled images, which are then mixed with the original images to generate virtual images using two strategies, a fixed ratio and a progressive dynamic MixUp. Next, dual bidirectional CutMix is used to enable initial knowledge transfer within the fixed virtual domain and gradual knowledge transfer from the dynamically transitioning labeled domain to the real unlabeled domains. On the other hand, to alleviate confirmation bias, we adopt prototypical alignment and pseudo label correction (PAPLC), which utilizes learnable prototype cosine similarity classifiers for bidirectional prototype alignment between the virtual and real domains, yielding smoother and more compact feature representations. Finally, we use prototypical pseudo label correction to generate more reliable pseudo labels. Empirical evaluations on three public multi‑domain datasets demonstrate the superiority of our method, particularly showing excellent performance even with very limited labeled samples. Code available at https://github.com/pascalcpp/BCMDA.
Authors:Yicheng Xu, Jiangning Zhang, Zhucun Xue, Teng Hu, Ran Yi, Xiaobin Hu, Yong Liu, Dacheng Tao
Abstract:
In‑context Learning enables training‑free adaptation via demonstrations but remains highly sensitive to example selection and formatting. In unified multimodal models spanning understanding and generation, this sensitivity is exacerbated by cross‑modal interference and varying cognitive demands. Consequently, In‑context Learning efficacy is often non‑monotonic and highly task‑dependent. To diagnose these behaviors, we introduce a six‑level capability‑oriented taxonomy that categorizes the functional role of demonstrations from basic perception to high‑order discernment. Guided by this cognitive framework, we construct UniICL‑760K, a large‑scale corpus featuring curated 8‑shot In‑context Learning episodes across 15 subtasks, alongside UniICL‑Bench for rigorous, controlled evaluation. As an architectural intervention to stabilize few‑shot adaptation, we propose the Context‑Adaptive Prototype Modulator, a lightweight, plug‑and‑play module. Evaluations on UniICL‑Bench show that our approach yields highly competitive unified results, outperforming larger‑parameter multimodal large language model baselines on most understanding In‑context Learning tasks. Data and code will be available soon at https://github.com/xuyicheng‑zju/UniICL.
Authors:An Yu, Ting Yu Tsai, Zhenfei Zhang, Weiheng Lu, Felix X. -F. Ye, Ming-Ching Chang
Abstract:
Recent multimodal large language models are computationally expensive because Transformers must process a large number of visual tokens. We present ReDiPrune, a training‑free token pruning method applied before the vision‑language projector, where visual features remain rich and discriminative. Unlike post‑projection pruning methods that operate on compressed representations, ReDiPrune selects informative tokens directly from vision encoder outputs, preserving fine‑grained spatial and semantic cues. Each token is scored by a lightweight rule that jointly consider text‑conditioned relevance and max‑min diversity, ensuring the selected tokens are both query‑relevant and non‑redundant. ReDiPrune is fully plug‑and‑play, requiring no retraining or architectural modifications, and can be seamlessly inserted between the encoder and projector. Across four video and five image benchmarks, it consistently improves the accuracy‑efficiency trade‑off. For example, on EgoSchema with LLaVA‑NeXT‑Video‑7B, retaining only 15% of visual tokens yields a +2.0% absolute accuracy gain while reducing computation by more than 6× in TFLOPs. Code is available at https://github.com/UA‑CVML/ReDiPrune.
Authors:Francesco Gentile, Nicola Dall'Asen, Francesco Tonini, Massimiliano Mancini, Lorenzo Vaquero, Elisa Ricci
Abstract:
As vision‑language models are deployed at scale, understanding their internal mechanisms becomes increasingly critical. Existing interpretability methods predominantly rely on activations, making them dataset‑dependent, vulnerable to data bias, and often restricted to coarse head‑level explanations. We introduce SITH (Semantic Inspection of Transformer Heads), a fully data‑free, training‑free framework that directly analyzes CLIP's vision transformer in weight space. For each attention head, we decompose its value‑output matrix into singular vectors and interpret each one via COMP (Coherent Orthogonal Matching Pursuit), a new algorithm that explains them as sparse, semantically coherent combinations of human‑interpretable concepts. We show that SITH yields coherent, faithful intra‑head explanations, validated through reconstruction fidelity and interpretability experiments. This allows us to use SITH for precise, interpretable weight‑space model edits that amplify or suppress specific concepts, improving downstream performance without retraining. Furthermore, we use SITH to study model adaptation, showing how fine‑tuning primarily reweights a stable semantic basis rather than learning entirely new features.
Authors:Xinying Guo, Chenxi Jiang, Hyun Bin Kim, Ying Sun, Yang Xiao, Yuhang Han, Jianfei Yang
Abstract:
Robotic manipulation often requires memory: occlusion and state changes can make decision‑time observations perceptually aliased, making action selection non‑Markovian at the observation level because the same observation may arise from different interaction histories. Most embodied agents implement memory via semantically compressed traces and similarity‑based retrieval, which discards disambiguating fine‑grained perceptual cues and can return perceptually similar but decision‑irrelevant episodes. Inspired by human episodic memory, we propose Chameleon, which writes geometry‑grounded multimodal tokens to preserve disambiguating context and produces goal‑directed recall through a differentiable memory stack. We also introduce Camo‑Dataset, a real‑robot UR5e dataset spanning episodic recall, spatial tracking, and sequential manipulation under perceptual aliasing. Across tasks, Chameleon consistently improves decision reliability and long‑horizon control over strong baselines in perceptually confusable settings.
Authors:Yubo Li, Xugong Qin, Peng Zhang, Hailun Lin, Gangyan Zeng, Kexin Zhang
Abstract:
Scene text editing seeks to modify textual content in natural images while maintaining visual realism and semantic consistency. Existing methods often require task‑specific training or paired data, limiting their scalability and adaptability. In this paper, we propose TextFlow, a training‑free scene text editing framework that integrates the strengths of Attention Boost (AttnBoost) and Flow Manifold Steering (FMS) to enable flexible, high‑fidelity text manipulation without additional training. Specifically, FMS preserves the structural and style consistency by modeling the visual flow of characters and background regions, while AttnBoost enhances the rendering of textual content through attention‑based guidance. By jointly leveraging these complementary modules, our approach performs end‑to‑end text editing through semantic alignment and spatial refinement in a plug‑and‑play manner. Extensive experiments demonstrate that our framework achieves visual quality and text accuracy comparable to or superior to those of training‑based counterparts, generalizing well across diverse scenes and languages. This study advances scene text editing toward a more efficient, generalizable, and training‑free paradigm. Code is available at https://github.com/lyb18758/TextFlow
Authors:Florian Stilz, Vinkle Srivastav, Nassir Navab, Nicolas Padoy
Abstract:
Video‑language foundation models have proven to be highly effective in zero‑shot applications across a wide range of tasks. A particularly challenging area is the intraoperative surgical procedure domain, where labeled data is scarce, and precise temporal understanding is often required for complex downstream tasks. To address this challenge, we introduce CliPPER (Contextual Video‑Language Pretraining on Long‑form Intraoperative Surgical Procedures for Event Recognition), a novel video‑language pretraining framework trained on surgical lecture videos. Our method is designed for fine‑grained temporal video‑text recognition and introduces several novel pretraining strategies to improve multimodal alignment in long‑form surgical videos. Specifically, we propose Contextual Video‑Text Contrastive Learning (VTC_CTX) and Clip Order Prediction (COP) pretraining objectives, both of which leverage temporal and contextual dependencies to enhance local video understanding. In addition, we incorporate a Cycle‑Consistency Alignment over video‑text matches within the same surgical video to enforce bidirectional consistency and improve overall representation coherence. Moreover, we introduce a more refined alignment loss, Frame‑Text Matching (FTM), to improve the alignment between video frames and text. As a result, our model establishes a new state‑of‑the‑art across multiple public surgical benchmarks, including zero‑shot recognition of phases, steps, instruments, and triplets. The source code and pretraining captions can be found at https://github.com/CAMMA‑public/CliPPER.
Authors:Zichuan Lin, Feiyu Liu, Yijun Yang, Jiafei Lyu, Yiming Gao, Yicheng Liu, Zhicong Lu, Yangbin Yu, Mingyu Yang, Junyou Li, Deheng Ye, Jie Jiang
Abstract:
Autonomous mobile GUI agents have attracted increasing attention along with the advancement of Multimodal Large Language Models (MLLMs). However, existing methods still suffer from inefficient learning from failed trajectories and ambiguous credit assignment under sparse rewards for long‑horizon GUI tasks. To that end, we propose UI‑Voyager, a novel two‑stage self‑evolving mobile GUI agent. In the first stage, we employ Rejection Fine‑Tuning (RFT), which enables the continuous co‑evolution of data and models in a fully autonomous loop. The second stage introduces Group Relative Self‑Distillation (GRSD), which identifies critical fork points in group rollouts and constructs dense step‑level supervision from successful trajectories to correct failed ones. Extensive experiments on AndroidWorld show that our 4B model achieves an 81.0% Pass@1 success rate, outperforming numerous recent baselines and exceeding human‑level performance. Ablation and case studies further verify the effectiveness of GRSD. Our method represents a significant leap toward efficient, self‑evolving, and high‑performance mobile GUI automation without expensive manual data annotation.
Authors:Jiawei Zhou, Zhenxin Zhu, Lingyi Du, Linye Lyu, Lijun Zhou, Zhanqian Wu, Hongcheng Luo, Zhuotao Tian, Bing Wang, Guang Chen, Hangjun Ye, Haiyang Sun, Yu Li
Abstract:
Video generation models have shown strong potential as world models for autonomous driving simulation. However, existing approaches are primarily trained on real‑world driving datasets, which mostly contain natural and safe driving scenarios. As a result, current models often fail when conditioned on challenging or counterfactual trajectories‑such as imperfect trajectories generated by simulators or planning systems‑producing videos with severe physical inconsistencies and artifacts. To address this limitation, we propose PhyGenesis, a world model designed to generate driving videos with high visual fidelity and strong physical consistency. Our framework consists of two key components: (1) a physical condition generator that transforms potentially invalid trajectory inputs into physically plausible conditions, and (2) a physics‑enhanced video generator that produces high‑fidelity multi‑view driving videos under these conditions. To effectively train these components, we construct a large‑scale, physics‑rich heterogeneous dataset. Specifically, in addition to real‑world driving videos, we generate diverse challenging driving scenarios using the CARLA simulator, from which we derive supervision signals that guide the model to learn physically grounded dynamics under extreme conditions. This challenging‑trajectory learning strategy enables trajectory correction and promotes physically consistent video generation. Extensive experiments demonstrate that PhyGenesis consistently outperforms state‑of‑the‑art methods, especially on challenging trajectories. Our project page is available at: https://wm‑research.github.io/PhyGenesis/.
Authors:Siqi Liu, Xinyang Li, Bochao Zou, Junbao Zhuo, Huimin Ma, Jiansheng Chen
Abstract:
As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human‑like Theory of Mind (ToM). Most existing ToM evaluations, however, are centered on text‑based inputs, while scenarios relying solely on visual information receive far less attention. This leaves a gap, since real‑world human‑AI interaction typically requires multimodal understanding. In addition, many current methods regard the model as a black box and rarely probe how its internal attention behaves in multiple‑choice question answering (QA). The impact of LLM hallucinations on such tasks is also underexplored from an interpretability perspective. To address these issues, we introduce VisionToM, a vision‑oriented intervention framework designed to strengthen task‑aware reasoning. The core idea is to compute intervention vectors that align visual representations with the correct semantic targets, thereby steering the model's attention through different layers of visual features. This guidance reduces the model's reliance on spurious linguistic priors, leading to more reliable multimodal language model (MLLM) outputs and better QA performance. Experiments on the EgoToM benchmark‑an egocentric, real‑world video dataset for ToM with three multiple‑choice QA settings‑demonstrate that our method substantially improves the ToM abilities of MLLMs. Furthermore, results on an additional open‑ended generation task show that VisionToM enables MLLMs to produce free‑form explanations that more accurately capture agents' mental states, pushing machine‑human collaboration toward greater alignment.
Authors:Kaihang Pan, Qi Tian, Jianwei Zhang, Weijie Kong, Jiangfeng Xiong, Yanxin Long, Shixue Zhang, Haiyi Qiu, Tan Wang, Zheqi Lv, Yue Wu, Liefeng Bo, Siliang Tang, Zhao Zhong
Abstract:
While proprietary systems such as Seedance‑2.0 have achieved remarkable success in omni‑capable video generation, open‑source alternatives significantly lag behind. Most academic models remain heavily fragmented, and the few existing efforts toward unified video generation still struggle to seamlessly integrate diverse tasks within a single framework. To bridge this gap, we propose OmniWeaving, an omni‑level video generation model featuring powerful multimodal composition and reasoning‑informed capabilities. By leveraging a massive‑scale pretraining dataset that encompasses diverse compositional and reasoning‑augmented scenarios, OmniWeaving learns to temporally bind interleaved text, multi‑image, and video inputs while acting as an intelligent agent to infer complex user intentions for sophisticated video creation. Furthermore, we introduce IntelligentVBench, the first comprehensive benchmark designed to rigorously assess next‑level intelligent unified video generation. Extensive experiments demonstrate that OmniWeaving achieves SoTA performance among open‑source unified models. The codes and model have already been publicly available. Project Page: https://omniweaving.github.io.
Authors:Jiawen Zhu, Yunqi Miao, Xueyi Zhang, Jiankang Deng, Guansong Pang
Abstract:
Recent Deepfake Video Detection (DFD) studies have demonstrated that pre‑trained Vision‑Language Models (VLMs) such as CLIP exhibit strong generalization capabilities in detecting artifacts across different identities. However, existing approaches focus on leveraging visual features only, overlooking their most distinctive strength ‑‑ the rich vision‑language semantics embedded in the latent space. We propose VLAForge, a novel DFD framework that unleashes the potential of such cross‑modal semantics to enhance model's discriminability in deepfake detection. This work i) enhances the visual perception of VLM through a ForgePerceiver, which acts as an independent learner to capture diverse, subtle forgery cues both granularly and holistically, while preserving the pretrained Vision‑Language Alignment (VLA) knowledge, and ii) provides a complementary discriminative cue ‑‑ Identity‑Aware VLA score, derived by coupling cross‑modal semantics with the forgery cues learned by ForgePerceiver. Notably, the VLA score is augmented by an identity prior‑informed text prompting to capture authenticity cues tailored to each identity, thereby enabling more discriminative cross‑modal semantics. Comprehensive experiments on video DFD benchmarks, including classical face‑swapping forgeries and recent full‑face generation forgeries, demonstrate that our VLAForge substantially outperforms state‑of‑the‑art methods at both frame and video levels. Code is available at https://github.com/mala‑lab/VLAForge.
Authors:Cheng Cui, Yubo Zhang, Ting Sun, Xueqing Wang, Hongen Liu, Manhui Lin, Yue Zhang, Tingquan Gao, Changda Zhou, Jiaxuan Liu, Zelun Zhang, Jing Zhang, Jun Zhang, Yi Liu
Abstract:
The advent of "OCR 2.0" and large‑scale vision‑language models (VLMs) has set new benchmarks in text recognition. However, these unified architectures often come with significant computational demands, challenges in precise text localization within complex layouts, and a propensity for textual hallucinations. Revisiting the prevailing notion that model scale is the sole path to high accuracy, this paper introduces PP‑OCRv5, a meticulously optimized, lightweight OCR system with merely 5 million parameters. We demonstrate that PP‑OCRv5 achieves performance competitive with many billion‑parameter VLMs on standard OCR benchmarks, while offering superior localization precision and reduced hallucinations. The cornerstone of our success lies not in architectural expansion but in a data‑centric investigation. We systematically dissect the role of training data by quantifying three critical dimensions: data difficulty, data accuracy, and data diversity. Our extensive experiments reveal that with a sufficient volume of high‑quality, accurately labeled, and diverse data, the performance ceiling for traditional, efficient two‑stage OCR pipelines is far higher than commonly assumed. This work provides compelling evidence for the viability of lightweight, specialized models in the large‑model era and offers practical insights into data curation for OCR. The source code and models are publicly available at https://github.com/PaddlePaddle/PaddleOCR.
Authors:Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, Yue Zhang, Yubo Zhang, Jing Zhang, Jun Zhang, Xing Wei, Yi Liu, Dianhai Yu, Yanjun Ma
Abstract:
Document parsing is a fine‑grained task where image resolution significantly impacts performance. While advanced research leveraging vision‑language models benefits from high‑resolution input to boost model performance, this often leads to a quadratic increase in the number of vision tokens and significantly raises computational costs. We attribute this inefficiency to substantial visual regions redundancy in document images, like background. To tackle this, we propose PaddleOCR‑VL, a novel coarse‑to‑fine architecture that focuses on semantically relevant regions while suppressing redundant ones, thereby improving both efficiency and performance. Specifically, we introduce a lightweight Valid Region Focus Module (VRFM) which leverages localization and contextual relationship prediction capabilities to identify valid vision tokens. Subsequently, we design and train a compact yet powerful 0.9B vision‑language model (PaddleOCR‑VL‑0.9B) to perform detailed recognition, guided by VRFM outputs to avoid direct processing of the entire large image. Extensive experiments demonstrate that PaddleOCR‑VL achieves state‑of‑the‑art performance in both page‑level parsing and element‑level recognition. It significantly outperforms existing solutions, exhibits strong competitiveness against top‑tier VLMs, and delivers fast inference while utilizing substantially fewer vision tokens and parameters, highlighting the effectiveness of targeted coarse‑to‑fine parsing for accurate and efficient document understanding. The source code and models are publicly available at https://github.com/PaddlePaddle/PaddleOCR.
Authors:Kai Zhu, Zhenyu Cui, Zehua Zang, Jiahuan Zhou
Abstract:
Recently, state space models have demonstrated efficient video segmentation through linear‑complexity state space compression. However, Video Semantic Segmentation (VSS) requires pixel‑level spatiotemporal modeling capabilities to maintain temporal consistency in segmentation of semantic objects. While state space models can preserve common semantic information during state space compression, the fixed‑size state space inevitably forgets specific information, which limits the models' capability for pixel‑level segmentation. To tackle the above issue, we proposed a Refining Specifics State Space Model approach (RS‑SSM) for video semantic segmentation, which performs complementary refining of forgotten spatiotemporal specifics. Specifically, a Channel‑wise Amplitude Perceptron (CwAP) is designed to extract and align the distribution characteristics of specific information in the state space. Besides, a Forgetting Gate Information Refiner (FGIR) is proposed to adaptively invert and refine the forgetting gate matrix in the state space model based on the specific information distribution. Consequently, our RS‑SSM leverages the inverted forgetting gate to complementarily refine the specific information forgotten during state space compression, thereby enhancing the model's capability for spatiotemporal pixel‑level segmentation. Extensive experiments on four VSS benchmarks demonstrate that our RS‑SSM achieves state‑of‑the‑art performance while maintaining high computational efficiency. The code is available at https://github.com/zhoujiahuan1991/CVPR2026‑RS‑SSM.
Authors:Guan Luo, Xiu Li, Rui Chen, Xuanyu Yi, Jing Lin, Chia-Hao Chen, Jiahang Liu, Song-Hai Zhang, Jianfeng Zhang
Abstract:
The dominant paradigm for high‑fidelity 3D generation relies on a VAE‑Diffusion pipeline, where the VAE's reconstruction capability sets a firm upper bound on generation quality. A fundamental challenge limiting existing VAEs is the representation mismatch between ground‑truth meshes and network predictions: GT meshes have arbitrary, variable topology, while VAEs typically predict fixed‑structure implicit fields (\eg, SDF on regular grids). This inherent misalignment prevents establishing explicit mesh‑level correspondences, forcing prior work to rely on indirect supervision signals such as SDF or rendering losses. Consequently, fine geometric details, particularly sharp features, are poorly preserved during reconstruction. To address this, we introduce TopoMesh, a sparse voxel‑based VAE that unifies both GT and predicted meshes under a shared Dual Marching Cubes (DMC) topological framework. Specifically, we convert arbitrary input meshes into DMC‑compliant representations via a remeshing algorithm that preserves sharp edges using an L\infty distance metric. Our decoder outputs meshes in the same DMC format, ensuring that both predicted and target meshes share identical topological structures. This establishes explicit correspondences at the vertex and face level, allowing us to derive explicit mesh‑level supervision signals for topology, vertex positions, and face orientations with clear gradients. Our sparse VAE architecture employs this unified framework and is trained with Teacher Forcing and progressive resolution training for stable and efficient convergence. Extensive experiments demonstrate that TopoMesh significantly outperforms existing VAEs in reconstruction fidelity, achieving superior preservation of sharp features and geometric details.
Authors:Tommaso Galliena, Stefano Rosa, Tommaso Apicella, Pietro Morerio, Alessio Del Bue, Lorenzo Natale
Abstract:
Vision‑Language Models (VLMs) often yield inconsistent descriptions of the same object across viewpoints, hindering the ability of embodied agents to construct consistent semantic representations over time. Previous methods resolved inconsistencies using offline multi‑view aggregation or multi‑stage pipelines that decouple exploration, data association, and caption learning, with limited capacity to reason over previously observed objects. In this paper, we introduce a unified, memory‑augmented Vision‑Language agent that simultaneously handles data association, object captioning, and exploration policy within a single autoregressive framework. The model processes the current RGB observation, a top‑down explored map, and an object‑level episodic memory serialized into object‑level tokens, ensuring persistent object identity and semantic consistency across extended sequences. To train the model in a self‑supervised manner, we collect a dataset in photorealistic 3D environments using a disagreement‑based policy and a pseudo‑captioning model that enforces consistency across multi‑view caption histories. Extensive evaluation on a manually annotated object‑level test set, demonstrate improvements of up to +11.86% in standard captioning scores and +7.39% in caption self‑similarity over baseline models, while enabling scalable performance through a compact scene representation. Code, model weights, and data are available at https://hsp‑iit.github.io/epos‑vlm/.
Authors:Nicanor Mayumu, Zeenath Khan, Melodena Stephens, Patrick Mukala, Farhad Oroumchian
Abstract:
Medical AI systems face two fundamental limitations. First, conventional vision‑language models (VLMs) perform single‑pass inference, yielding black‑box predictions that cannot be audited or explained in clinical terms. Second, iterative reasoning systems that expose intermediate steps rely on fixed iteration budgets wasting compute on simple cases while providing insufficient depth for complex ones. We address both limitations with a unified framework. RVLM replaces single‑pass inference with an iterative generate‑execute loop: at each step, the model writes Python code, invokes vision sub‑agents, manipulates images, and accumulates evidence. Every diagnostic claim is grounded in executable code, satisfying auditability requirements of clinical AI governance frameworks. RRouter makes iteration depth adaptive: a lightweight controller predicts the optimal budget from task‑complexity features, then monitors progress and terminates early when reasoning stalls. We evaluate on BraTS 2023 Meningioma (brain MRI) and MIMIC‑CXR (chest X‑ray) using Gemini 2.5 Flash without fine‑tuning. Across repeated runs, RVLM shows high consistency on salient findings (e.g., mass presence and enhancement) and can detect cross‑modal discrepancies between Fluid‑Attenuated Inversion Recovery (FLAIR) signal characteristics and segmentation boundaries. On MIMIC‑CXR, it generates structured reports and correctly recognises view‑specific artefacts. Code: https://github.com/nican2018/rvlm.
Authors:Minjun Kim, Minje Kim
Abstract:
Personalized Federated Learning (PFL) aims to deliver effective client‑specific models under heterogeneous distributions, yet existing methods suffer from shallow prototype alignment and brittle server‑side distillation. We propose HEART‑PFL, a dual‑sided framework that (i) performs depth‑aware Hierarchical Directional Alignment (HDA) using cosine similarity in the early stage and MSE matching in the deep stage to preserve client specificity, and (ii) stabilizes global updates through Adversarial Knowledge Transfer (AKT) with symmetric KL distillation on clean and adversarial proxy data. Using lightweight adapters with only 1.46M trainable parameters, HEART‑PFL achieves state‑of‑the‑art personalized accuracy on CIFAR‑100, Flowers‑102, and Caltech‑101 (63.42%, 84.23%, and 95.67%, respectively) under Dirichlet non‑IID partitions, and remains robust to out‑of‑domain proxy data. Ablation studies further confirm that HDA and AKT provide complementary gains in alignment, robustness, and optimization stability, offering insights into how the two components mutually reinforce effective personalization. Overall, these results demonstrate that HEART‑PFL simultaneously enhances personalization and global stability, highlighting its potential as a strong and scalable solution for PFL(code available at https://github.com/danny0628/HEART‑PFL).
Authors:Jaehun Bang, Jinhyeok Kim, Minji Kim, Seungheon Jeong, Kyungdon Joo
Abstract:
Open‑vocabulary 3D scene understanding enables users to segment novel objects in complex 3D environments through natural language. However, existing approaches remain slow, memory‑intensive, and overly complex due to iterative optimization and dense per‑Gaussian feature assignments. To address this, we propose LightSplat, a fast and memory‑efficient training‑free framework that injects compact 2‑byte semantic indices into 3D representations from multi‑view images. By assigning semantic indices only to salient regions and managing them with a lightweight index‑feature mapping, LightSplat eliminates costly feature optimization and storage overhead. We further ensure semantic consistency and efficient inference via single‑step clustering that links geometrically and semantically related masks in 3D. We evaluate our method on LERF‑OVS, ScanNet, and DL3DV‑OVS across complex indoor‑outdoor scenes. As a result, LightSplat achieves state‑of‑the‑art performance with up to 50‑400x speedup and 64x lower memory, enabling scalable language‑driven 3D understanding. For more details, visit our project page https://vision3d‑lab.github.io/lightsplat/.
Authors:Zhanhe Lei, Zhongyuan Wang, Jikang Cheng, Baojin Huang, Yuhong Yang, Zhen Han, Chao Liang, Dengpan Ye
Abstract:
Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor‑Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curriculum. Our method models the training process as a Markov Decision Process where a ``Tutor'' agent learns to guide a ``Student'' (the deepfake detector). The Tutor, implemented as a Proximal Policy Optimization (PPO) agent, observes a rich state representation for each training sample, encapsulating not only its visual features but also its historical learning dynamics, such as EMA loss and forgetting counts. Based on this state, the Tutor takes an action by assigning a continuous weight (0‑1) to the sample's loss, thereby dynamically re‑weighting the training batch. The Tutor is rewarded based on the Student's immediate performance change, specifically rewarding transitions from incorrect to correct predictions. This strategy encourages the Tutor to learn a curriculum that prioritizes high‑value samples, such as hard‑but‑learnable examples, leading to a more efficient and effective training process. We demonstrate that this adaptive curriculum improves the Student's generalization capabilities against unseen manipulation techniques compared to traditional training methods. Code is available at https://github.com/wannac1/TSRL.
Authors:Haoyu Ji, Bowen Chen, Zhihao Yang, Wenze Huang, Yu Gao, Xueting Liu, Weihong Ren, Zhiyong Wang, Honghai Liu
Abstract:
Skeleton‑based Temporal Action Segmentation (STAS) seeks to densely segment and classify diverse actions within long, untrimmed skeletal motion sequences. However, existing STAS methodologies face challenges of limited inter‑class discriminability and blurred segmentation boundaries, primarily due to insufficient distinction of spatio‑temporal patterns between adjacent actions. To address these limitations, we propose Spectral Scalpel, a frequency‑selective filtering framework aimed at suppressing shared frequency components between adjacent distinct actions while amplifying their action‑specific frequencies, thereby enhancing inter‑action discrepancies and sharpening transition boundaries. Specifically, Spectral Scalpel employs adaptive multi‑scale spectral filters as scalpels to edit frequency spectra, coupled with a discrepancy loss between adjacent actions serving as the surgical objective. This design amplifies representational disparities between neighboring actions, effectively mitigating boundary localization ambiguities and inter‑class confusion. Furthermore, complementing long‑term temporal modeling, we introduce a frequency‑aware channel mixer to strengthen channel evolution by aggregating spectra across channels. This work presents a novel paradigm for STAS that extends conventional spatio‑temporal modeling by incorporating frequency‑domain analysis. Extensive experiments on five public datasets demonstrate that Spectral Scalpel achieves state‑of‑the‑art performance. Code is available at https://github.com/HaoyuJi/SpecScalpel.
Authors:Mayssa Soussia, Gita Ayu Salsabila, Mohamed Ali Mahjoub, Islem Rekik
Abstract:
Message passing is a core mechanism in Graph Neural Networks (GNNs), enabling the iterative update of node embeddings by aggregating information from neighboring nodes. Graph Convolutional Networks (GCNs) exemplify this approach by adapting convolutional operations for graph structures, allowing features from adjacent nodes to be combined effectively. However, GCNs encounter challenges with complex or dynamic data. Capturing long‑range dependencies often requires deeper layers, which not only increase computational costs but also lead to over‑smoothing, where node embeddings become indistinguishable. To overcome these challenges, reservoir computing has been integrated into GNNs, leveraging iterative message‑passing dynamics for stable information propagation without extensive parameter tuning. Despite its promise, existing reservoir‑based models lack structured convolutional mechanisms, limiting their ability to accurately aggregate multi‑hop neighborhood information. To address these limitations, we propose RGC‑Net (Reservoir‑based Graph Convolutional Network), which integrates reservoir dynamics with structured graph convolution. Key contributions include: (i) a reimagined convolutional framework with fixed random reservoir weights and a leaky integrator to enhance feature retention; (ii) a robust, adaptable model for graph classification; and (iii) an RGC‑Net‑powered transformer for graph generation with application to dynamic brain connectivity. Extensive experiments show that RGC‑Net achieves state‑of‑the‑art performance in classification and generative tasks, including brain graph evolution, with faster convergence and reduced over‑smoothing. Source code is available at https://github.com/basiralab/RGC‑Net .
Authors:Haoyu Ji, Xueting Liu, Yu Gao, Wenze Huang, Zhihao Yang, Weihong Ren, Zhiyong Wang, Honghai Liu
Abstract:
Skeleton‑based Temporal Action Segmentation (STAS) aims to densely parse untrimmed skeletal sequences into frame‑level action categories. However, existing methods, while proficient at capturing spatio‑temporal kinematics, neglect the underlying physical dynamics that govern human motion. This oversight limits inter‑class discriminability between actions with similar kinematics but distinct dynamic intents, and hinders precise boundary localization where dynamic force profiles shift. To address these, we propose the Lagrangian‑Dynamic Informed Network (LaDy), a framework integrating principles of Lagrangian dynamics into the segmentation process. Specifically, LaDy first computes generalized coordinates from joint positions and then estimates Lagrangian terms under physical constraints to explicitly synthesize the generalized forces. To further ensure physical coherence, our Energy Consistency Loss enforces the work‑energy theorem, aligning kinetic energy change with the work done by the net force. The learned dynamics then drive a Spatio‑Temporal Modulation module: Spatially, generalized forces are fused with spatial representations to provide more discriminative semantics. Temporally, salient dynamic signals are constructed for temporal gating, thereby significantly enhancing boundary awareness. Experiments on challenging datasets show that LaDy achieves state‑of‑the‑art performance, validating the integration of physical dynamics for action segmentation. Code is available at https://github.com/HaoyuJi/LaDy.
Authors:Yuheng Feng, Wen Zhang, Haodong Duan, Xingxing Zou
Abstract:
We present PosterIQ, a design‑driven benchmark for poster understanding and generation, annotated across composition structure, typographic hierarchy, and semantic intent. It includes 7,765 image‑annotation instances and 822 generation prompts spanning real, professional, and synthetic cases. To bridge visual design cognition and generative modeling, we define tasks for layout parsing, text‑image correspondence, typography/readability and font perception, design quality assessment, and controllable, composition‑aware generation with metaphor. We evaluate state‑of‑the‑art MLLMs and diffusion‑based generators, finding persistent gaps in visual hierarchy, typographic semantics, saliency control, and intention communication; commercial models lead on high‑level reasoning but act as insensitive automatic raters, while generators render text well yet struggle with composition‑aware synthesis. Extensive analyses show PosterIQ is both a quantitative benchmark and a diagnostic tool for design reasoning, offering reproducible, task‑specific metrics. We aim to catalyze models' creativity and integrate human‑centred design principles into generative vision‑language systems.
Authors:Haiyang Xu, Ronghuan Wu, Li-Yi Wei, Nanxuan Zhao, Chenxi Liu, Cuong Nguyen, Zhuowen Tu, Zhaowen Wang
Abstract:
Graphic icons are a cornerstone of modern design workflows, yet they are often distributed as flattened single‑path or compound‑path graphics, where the original semantic layering is lost. This absence of semantic decomposition hinders downstream tasks such as editing, restyling, and animation. We formalize this problem as semantic layer construction for flattened vector art and introduce SemLayer, a visual generation empowered pipeline that restores editable layered structures. Given an abstract icon, SemLayer first generates a chromatically differentiated representation in which distinct semantic components become visually separable. To recover the complete geometry of each part, including occluded regions, we then perform a semantic completion step that reconstructs coherent object‑level shapes. Finally, the recovered parts are assembled into a layered vector representation with inferred occlusion relationships. Extensive qualitative comparisons and quantitative evaluations demonstrate the effectiveness of SemLayer, enabling editing workflows previously inapplicable to flattened vector graphics and establishing semantic layer reconstruction as a practical and valuable task. Project page: https://xxuhaiyang.github.io/SemLayer/
Authors:Kaiyuan Ji, Yixuan Gao, Lu Sun, Yushuo Zheng, Zijian Chen, Jianbo Zhang, Xiangyang Zhu, Yuan Tian, Zicheng Zhang, Guangtao Zhai
Abstract:
Advertising images significantly impact commercial conversion rates and brand equity, yet current evaluation methods rely on subjective judgments, lacking scalability, standardized criteria, and interpretability. To address these challenges, we present A^3 (Advertising Aesthetic Assessment), a comprehensive framework encompassing four components: a paradigm (A^3‑Law), a dataset (A^3‑Dataset), a multimodal large language model (A^3‑Align), and a benchmark (A^3‑Bench). Central to A^3 is a theory‑driven paradigm, A^3‑Law, comprising three hierarchical stages: (1) Perceptual Attention, evaluating perceptual image signals for their ability to attract attention; (2) Formal Interest, assessing formal composition of image color and spatial layout in evoking interest; and (3) Desire Impact, measuring desire evocation from images and their persuasive impact. Building on A^3‑Law, we construct A^3‑Dataset with 120K instruction‑response pairs from 30K advertising images, each richly annotated with multi‑dimensional labels and Chain‑of‑Thought (CoT) rationales. We further develop A^3‑Align, trained under A^3‑Law with CoT‑guided learning on A^3‑Dataset. Extensive experiments on A^3‑Bench demonstrate that A^3‑Align achieves superior alignment with A^3‑Law compared to existing models, and this alignment generalizes well to quality advertisement selection and prescriptive advertisement critique, indicating its potential for broader deployment. Dataset, code, and models can be found at: https://github.com/euleryuan/A3‑Align.
Authors:Avigail Cohen Rimon, Amir Mann, Mirela Ben Chen, Or Litany
Abstract:
3D Gaussian Splatting (3DGS) enables real‑time, photorealistic novel view synthesis, making it a highly attractive representation for model‑based video tracking. However, leveraging the differentiability of the 3DGS renderer "in the wild" remains notoriously fragile. A fundamental bottleneck lies in the compact, local support of the Gaussian primitives. Standard photometric objectives implicitly rely on spatial overlap; if severe camera misalignment places the rendered object outside the target's local footprint, gradients strictly vanish, leaving the optimizer stranded. We introduce SpectralSplats, a robust tracking framework that resolves this "vanishing gradient" problem by shifting the optimization objective from the spatial to the frequency domain. By supervising the rendered image via a set of global complex sinusoidal features (Spectral Moments), we construct a global basin of attraction, ensuring that a valid, directional gradient toward the target exists across the entire image domain, even when pixel overlap is completely nonexistent. To harness this global basin without introducing periodic local minima associated with high frequencies, we derive a principled Frequency Annealing schedule from first principles, gracefully transitioning the optimizer from global convexity to precise spatial alignment. We demonstrate that SpectralSplats acts as a seamless, drop‑in replacement for spatial losses across diverse deformation parameterizations (from MLPs to sparse control points), successfully recovering complex deformations even from severely misaligned initializations where standard appearance‑based tracking catastrophically fails.
Authors:Yumeng Liu, Xiao-Xiao Long, Marc Habermann, Xuanze Yang, Cheng Lin, Yuan Liu, Yuexin Ma, Wenping Wang, Ligang Liu
Abstract:
Recovering high‑fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially, scalable applications demand both accuracy and deployment flexibility, requiring the ability to leverage massive amounts of unstructured image data from the internet or enable deployment on consumer‑grade RGB cameras without complex calibration. However, current methods face a dilemma. While single‑view approaches are easy to deploy, they suffer from depth ambiguity and occlusion. Conversely, multi‑view systems resolve these uncertainties but typically demand fixed, calibrated setups, limiting their real‑world utility. To bridge this gap, we draw inspiration from 3D foundation models that learn explicit geometry directly from visual data. By reformulating hand reconstruction from arbitrary views as a visual‑geometry grounded task, we propose a feed‑forward architecture that, for the first time in literature, jointly infers 3D hand meshes and camera poses from uncalibrated views. Extensive evaluations show that our approach outperforms state‑of‑the‑art benchmarks and demonstrates strong generalization to uncalibrated, in‑the‑wild scenarios. Here is the link of our project page: https://lym29.github.io/HGGT/.
Authors:Jielun Peng, Yabin Wang, Yaqi Li, Long Kong, Xiaopeng Hong
Abstract:
The rapid progress of generative AI has enabled hyper‑realistic audio‑visual deepfakes, intensifying threats to personal security and social trust. Most existing deepfake detectors rely either on uni‑modal artifacts or audio‑visual discrepancies, failing to jointly leverage both sources of information. Moreover, detectors that rely on generator‑specific artifacts tend to exhibit degraded generalization when confronted with unseen forgeries. We argue that robust and generalizable detection should be grounded in intrinsic audio‑visual coherence within and across modalities. Accordingly, we propose HAVIC, a Holistic Audio‑Visual Intrinsic Coherence‑based deepfake detector. HAVIC first learns priors of modality‑specific structural coherence, inter‑modal micro‑ and macro‑coherence by pre‑training on authentic videos. Based on the learned priors, HAVIC further performs holistic adaptive aggregation to dynamically fuse audio‑visual features for deepfake detection. Additionally, we introduce HiFi‑AVDF, a high‑fidelity audio‑visual deepfake dataset featuring both text‑to‑video and image‑to‑video forgeries from state‑of‑the‑art commercial generators. Extensive experiments across several benchmarks demonstrate that HAVIC significantly outperforms existing state‑of‑the‑art methods, achieving improvements of 9.39% AP and 9.37% AUC on the most challenging cross‑dataset scenario. Our code and dataset are available at https://github.com/tuffy‑studio/HAVIC.
Authors:Qi Zhang, Daijie Chen, Yunfei Gong, Hui Huang
Abstract:
Existing multi‑view crowd counting and localization methods are evaluated under relatively small scenes with limited crowd numbers, camera views, and frames. This makes the evaluation and comparison of existing methods impractical, as small datasets are easily overfit by these methods. To avoid these issues, 3DROM proposes a data augmentation method. Instead, in this paper, we propose a large synthetic benchmark, SynMVCrowd, for more practical evaluation and comparison of multi‑view crowd counting and localization tasks. The SynMVCrowd benchmark consists of 50 synthetic scenes with a large number of multi‑view frames and camera views and a much larger crowd number (up to 1000), which is more suitable for large‑scene multi‑view crowd vision tasks. Besides, we propose strong multi‑view crowd localization and counting baselines that outperform all comparison methods on the new SynMVCrowd benchmark. Moreover, we prove that better domain transferring multi‑view and single‑image counting performance could be achieved with the aid of the benchmark on novel new real scenes. As a result, the proposed benchmark could advance the research for multi‑view and single‑image crowd counting and localization to more practical applications. The codes and datasets are here: https://github.com/zqyq/SynMVCrowd.
Authors:Kai-Yu Fu, Yi-Ting Chen
Abstract:
We study object importance‑based vision risk object identification (Vision‑ROI), a key capability for hazard detection in intelligent driving systems. Existing approaches make deterministic decisions and ignore uncertainty, which could lead to safety‑critical failures. Specifically, in ambiguous scenarios, fixed decision thresholds may cause premature or delayed risk detection and temporally unstable predictions, especially in complex scenes with multiple interacting risks. Despite these challenges, current methods lack a principled framework to model risk uncertainty jointly across space and time. We propose Conformal Risk Tube Prediction, a unified formulation that captures spatiotemporal risk uncertainty, provides coverage guarantees for true risks, and produces calibrated risk scores with uncertainty estimates. To conduct a systematic evaluation, we present a new dataset and metrics probing diverse scenario configurations with multi‑risk coupling effects, which are not supported by existing datasets. We systematically analyze factors affecting uncertainty estimation, including scenario variations, per‑risk category behavior, and perception error propagation. Our method delivers substantial improvements over prior approaches, enhancing vision‑ROI robustness and downstream performance, such as reducing nuisance braking alerts. For more qualitative results, please visit our project webpage: https://hcis‑lab.github.io/CRTP/
Authors:Yixian Wang, Haolin Yu, Jiadong Tang, Yu Gao, Xihan Wang, Yufeng Yue, Yi Yang
Abstract:
3D Gaussian Splatting has revolutionized neural rendering with real‑time performance. However, scaling this approach to large scenes using Level‑of‑Detail methods faces critical challenges: inefficient serial traversal consuming over 60% of rendering time, and redundant Gaussian‑tile pairs that incur unnecessary processing overhead. To address these limitations, we introduce FilterGS, featuring a parallel filtering mechanism with two complementary filters that select Gaussian elements efficiently without tree traversal. Additionally, we propose a novel GTC metric that quantifies the redundancy of Gaussian‑tile key‑value pairs. Based on this metric, we introduce a scene‑adaptive Gaussian shrinking strategy that effectively reduces redundant pairs. Extensive experiments demonstrate that FilterGS achieves state‑of‑the‑art rendering speeds while maintaining competitive visual quality across multiple large‑scale datasets. Project page: https://github.com/xenon‑w/FilterGS
Authors:Risa Shinoda, Kaede Shiohara, Nakamasa Inoue, Kuniaki Saito, Hiroaki Santo, Fumio Okura
Abstract:
Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual taxonomic information for species identification, the integration of the audio modality remains an open problem. We propose BioVITA, a novel visual‑textual‑acoustic alignment framework for biological applications. BioVITA involves (i) a training dataset, (ii) a representation model, and (iii) a retrieval benchmark. First, we construct a large‑scale training dataset comprising 1.3 million audio clips and 2.3 million images, covering 14,133 species annotated with 34 ecological trait labels. Second, building upon BioCLIP2, we introduce a two‑stage training framework to effectively align audio representations with visual and textual representations. Third, we develop a cross‑modal retrieval benchmark that covers all possible directional retrieval across the three modalities (i.e., image‑to‑audio, audio‑to‑text, text‑to‑image, and their reverse directions), with three taxonomic levels: Family, Genus, and Species. Extensive experiments demonstrate that our model learns a unified representation space that captures species‑level semantics beyond taxonomy, advancing multimodal biodiversity understanding. The project page is available at: https://dahlian00.github.io/BioVITA_Page/
Authors:Bingxue Zhao, Qi Zhang, Hui Huang
Abstract:
Modeling realistic pedestrian trajectories requires accounting for both social interactions and environmental context, yet most existing approaches largely emphasize social dynamics. We propose EnvSocial‑Diff: a diffusion‑based crowd simulation model informed by social physics and augmented with environmental conditioning and individual‑‑group interaction. Our structured environmental conditioning module explicitly encodes obstacles, objects of interest, and lighting levels, providing interpretable signals that capture scene constraints and attractors. In parallel, the individual‑‑group interaction module goes beyond individual‑level modeling by capturing both fine‑grained interpersonal relations and group‑level conformity through a graph‑based design. Experiments on multiple benchmark datasets demonstrate that EnvSocial‑Diff outperforms the latest state‑of‑the‑art methods, underscoring the importance of explicit environmental conditioning and multi‑level social interaction for realistic crowd simulation. Code is here: https://github.com/zqyq/EnvSocial‑Diff.
Authors:Philipp Wesp, Robbie Holland, Vasiliki Sideri-Lampretsa, Sergios Gatidis
Abstract:
Vision foundation models (FMs) achieve state‑of‑the‑art performance in medical imaging. However, they encode information in abstract latent representations that clinicians cannot interrogate or verify. The goal of this study is to investigate Sparse Autoencoders (SAEs) for replacing opaque FM image representations with human‑interpretable, sparse features. We train SAEs on embeddings from BiomedParse (biomedical) and DINOv3 (general‑purpose) using 909,873 CT and MRI 2D image slices from the TotalSegmentator dataset. We find that learned sparse features: (a) reconstruct original embeddings with high fidelity (R2 up to 0.941) and recover up to 87.8% of downstream performance using only 10 features (99.4% dimensionality reduction), (b) preserve semantic fidelity in image retrieval tasks, (c) correspond to specific concepts that can be expressed in language using large language model (LLM)‑based auto‑interpretation. (d) bridge clinical language and abstract latent representations in zero‑shot language‑driven image retrieval. Our work indicates SAEs are a promising pathway towards interpretable, concept‑driven medical vision systems. Code repository: https://github.com/pwesp/sail.
Authors:Shreen Gul, Mohamed Elmahallawy, Ardhendu Tripathy, Sanjay Madria
Abstract:
Deep learning models are increasingly deployed in safety‑critical applications, where reliable out‑of‑distribution (OOD) detection is essential to ensure robustness. Existing methods predominantly rely on the penultimate‑layer activations of neural networks, assuming they encapsulate the most informative in‑distribution (ID) representations. In this work, we revisit this assumption to show that intermediate layers encode equally rich and discriminative information for OOD detection. Based on this observation, we propose a simple yet effective model‑agnostic approach that leverages internal representations across multiple layers. Our scheme aggregates features from successive convolutional blocks, computes class‑wise mean embeddings, and applies L_2 normalization to form compact ID prototypes capturing class semantics. During inference, cosine similarity between test features and these prototypes serves as an OOD score‑‑ID samples exhibit strong affinity to at least one prototype, whereas OOD samples remain uniformly distant. Extensive experiments on state‑of‑the‑art OOD benchmarks across diverse architectures demonstrate that our approach delivers robust, architecture‑agnostic performance and strong generalization for image classification. Notably, it improves AUROC by up to 4.41% and reduces FPR by 13.58%, highlighting multi‑layer feature aggregation as a powerful yet underexplored signal for OOD detection, challenging the dominance of penultimate‑layer‑based methods. Our code is available at: https://github.com/sgchr273/cosine‑layers.git.
Authors:Jannik Endres, Etienne Laliberté, David Rolnick, Arthur Ouaknine
Abstract:
Accurate estimation of forest biomass, a major carbon sink, relies heavily on tree‑level traits such as height and species. Unoccupied Aerial Vehicles (UAVs) capturing high‑resolution imagery from a single RGB camera offer a cost‑effective and scalable approach for mapping and measuring individual trees. We introduce BIRCH‑Trees, the first benchmark for individual tree height and species estimation from tree‑centered UAV images, spanning three datasets: temperate forests, tropical forests, and boreal plantations. We also present DINOvTree, a unified approach using a Vision Foundation Model (VFM) backbone with task‑specific heads for simultaneous height and species prediction. Through extensive evaluations on BIRCH‑Trees, we compare DINOvTree against commonly used vision methods, including VFMs, as well as biological allometric equations. We find that DINOvTree achieves top overall results with accurate height predictions and competitive classification accuracy while using only 54% to 58% of the parameters of the second‑best approach.
Authors:Peiyu Xu, Xin Sun, Krishna Mullia, Raymond Fei, Iliyan Georgiev, Shuang Zhao
Abstract:
Ray‑tracing‑based 3D Gaussian splatting (3DGS) methods overcome the limitations of rasterization ‑‑ rigid pinhole camera assumptions, inaccurate shadows, and lack of native reflection or refraction ‑‑ but remain slower due to the cost of sorting all intersecting Gaussians along every ray. Moreover, existing ray‑tracing methods still rely on rasterization‑style approximations such as shadow mapping for relightable scenes, undermining the generality that ray tracing promises.
We present a differentiable, sorting‑free stochastic formulation for ray‑traced 3DGS ‑‑ the first framework that uses stochastic ray tracing to both reconstruct and render standard and relightable 3DGS scenes. At its core is an unbiased Monte Carlo estimator for pixel‑color gradients that evaluates only a small sampled subset of Gaussians per ray, bypassing the need for sorting. For standard 3DGS, our method matches the reconstruction quality and speed of rasterization‑based 3DGS while substantially outperforming sorting‑based ray tracing. For relightable 3DGS, the same stochastic estimator drives per‑Gaussian shading with fully ray‑traced shadow rays, delivering notably higher reconstruction fidelity than prior work.
Authors:Anh-Quan Cao, Tuan-Hung Vu
Abstract:
Relying on in‑domain annotations and precise sensor‑rig priors, existing 3D occupancy prediction methods are limited in both scalability and out‑of‑domain generalization. While recent visual geometry foundation models exhibit strong generalization capabilities, they were mainly designed for general purposes and lack one or more key ingredients required for urban occupancy prediction, namely metric prediction, geometry completion in cluttered scenes and adaptation to urban scenarios. We address this gap and present OccAny, the first unconstrained urban 3D occupancy model capable of operating on out‑of‑domain uncalibrated scenes to predict and complete metric occupancy coupled with segmentation features. OccAny is versatile and can predict occupancy from sequential, monocular, or surround‑view images. Our contributions are three‑fold: (i) we propose the first generalized 3D occupancy framework with (ii) Segmentation Forcing that improves occupancy quality while enabling mask‑level prediction, and (iii) a Novel View Rendering pipeline that infers novel‑view geometry to enable test‑time view augmentation for geometry completion. Extensive experiments demonstrate that OccAny outperforms all visual geometry baselines on 3D occupancy prediction task, while remaining competitive with in‑domain self‑supervised methods across three input settings on two established urban occupancy prediction datasets. Our code is available at https://github.com/valeoai/OccAny .
Authors:Jaewon Min, Jaeeun Lee, Yeji Choi, Paul Hyunbin Cho, Jin Hyeon Kim, Tae-Young Lee, Jongsik Ahn, Hwayeong Lee, Seonghyun Park, Seungryong Kim
Abstract:
Optical flow models trained on high‑quality data often degrade severely when confronted with real‑world corruptions such as blur, noise, and compression artifacts. To overcome this limitation, we formulate Degradation‑Aware Optical Flow, a new task targeting accurate dense correspondence estimation from real‑world corrupted videos. Our key insight is that the intermediate representations of image restoration diffusion models are inherently corruption‑aware but lack temporal awareness. To address this limitation, we lift the model to attend across adjacent frames via full spatio‑temporal attention, and empirically demonstrate that the resulting features exhibit zero‑shot correspondence capabilities. Based on this finding, we present DA‑Flow, a hybrid architecture that fuses these diffusion features with convolutional features within an iterative refinement framework. DA‑Flow substantially outperforms existing optical flow methods under severe degradation across multiple benchmarks.
Authors:Zhen Li, Zian Meng, Shuwei Shi, Wenshuo Peng, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang
Abstract:
Dynamical systems theory and reinforcement learning view world evolution as latent‑state dynamics driven by actions, with visual observations providing partial information about the state. Recent video world models attempt to learn this action‑conditioned dynamics from data. However, existing datasets rarely match the requirement: they typically lack diverse and semantically meaningful action spaces, and actions are directly tied to visual observations rather than mediated by underlying states. As a result, actions are often entangled with pixel‑level changes, making it difficult for models to learn structured world dynamics and maintain consistent evolution over long horizons. In this paper, we propose WildWorld, a large‑scale action‑conditioned world modeling dataset with explicit state annotations, automatically collected from a photorealistic AAA action role‑playing game (Monster Hunter: Wilds). WildWorld contains over 108 million frames and features more than 450 actions, including movement, attacks, and skill casting, together with synchronized per‑frame annotations of character skeletons, world states, camera poses, and depth maps. We further derive WildBench to evaluate models through Action Following and State Alignment. Extensive experiments reveal persistent challenges in modeling semantically rich actions and maintaining long‑horizon state consistency, highlighting the need for state‑aware video generation. The project page is https://shandaai.github.io/wildworld‑project/.
Authors:Brian Chao, Lior Yariv, Howard Xiao, Gordon Wetzstein
Abstract:
Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming video generation. The growing demand for higher resolutions, frame rates, and context lengths, however, makes efficient generation increasingly challenging, as computational complexity grows quadratically with the number of generated tokens. Our work seeks to optimize the efficiency of the generation process in settings where the user's gaze location is known or can be estimated, for example, by using eye tracking. In these settings, we leverage the eccentricity‑dependent acuity of human vision: while a user perceives very high‑resolution visual information in a small region around their gaze location (the foveal region), the ability to resolve detail quickly degrades in the periphery of the visual field. Our approach starts with a mask modeling the foveated resolution to allocate tokens non‑uniformly, assigning higher token density to foveal regions and lower density to peripheral regions. An image or video is generated in a mixed‑resolution token setting, yielding results perceptually indistinguishable from full‑resolution generation, while drastically reducing the token count and generation time. To this end, we develop a principled mechanism for constructing mixed‑resolution tokens directly from high‑resolution data, allowing a foveated diffusion model to be post‑trained from an existing base model while maintaining content consistency across resolutions. We validate our approach through extensive analysis and a carefully designed user study, demonstrating the efficacy of foveation as a practical and scalable axis for efficient generation.
Authors:Woojeong Jin, Jaeho Lee, Heeseong Shin, Seungho Jang, Junhwan Heo, Seungryong Kim
Abstract:
Referring Video Object Segmentation (RVOS) aims to segment a target object throughout a video given a natural language query. Training‑free methods for this task follow a common pipeline: a MLLM selects keyframes, grounds the referred object within those frames, and a video segmentation model propagates the results. While intuitive, this design asks the MLLM to make temporal decisions before any object‑level evidence is available, limiting both reasoning quality and spatio‑temporal coverage. To overcome this, we propose AgentRVOS, a training‑free agentic pipeline built on the complementary strengths of SAM3 and a MLLM. Given a concept derived from the query, SAM3 provides reliable perception over the full spatio‑temporal extent through generated mask tracks. The MLLM then identifies the target through query‑grounded reasoning over this object‑level evidence, iteratively pruning guided by SAM3's temporal existence information. Extensive experiments show that AgentRVOS achieves state‑of‑the‑art performance among training‑free methods across multiple benchmarks, with consistent results across diverse MLLM backbones. Our project page is available at: https://cvlab‑kaist.github.io/AgentRVOS/.
Authors:Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez, David Picard
Abstract:
Monocular novel‑view synthesis has long required multi‑view image pairs for supervision, limiting training data scale and diversity. We argue it is not necessary: one view is enough. We present OVIE, trained entirely on unpaired internet images. We leverage a monocular depth estimator as a geometric scaffold at training time: we lift a source image into 3D, apply a sampled camera transformation, and project to obtain a pseudo‑target view. To handle disocclusions, we introduce a masked training formulation that restricts geometric, perceptual, and textural losses to valid regions, enabling training on 30 million uncurated images. At inference, OVIE is geometry‑free, requiring no depth estimator or 3D representation. Trained exclusively on in‑the‑wild images, OVIE outperforms prior methods in a zero‑shot setting, while being 600x faster than the second‑best baseline. Code and models are publicly available at https://github.com/AdrienRR/ovie.
Authors:Haoyu Huang, Jinfa Huang, Zhongwei Wan, Xiawu Zheng, Rongrong Ji, Jiebo Luo
Abstract:
Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through iterative visual tool invocation. However, the cascaded perception, reasoning, and tool‑calling loops introduce significant sequential overhead. This overhead, termed agentic depth, incurs prohibitive latency and seriously limits system‑level concurrency. To this end, we propose SpecEyes, an agentic‑level speculative acceleration framework that breaks this sequential bottleneck. Our key insight is that a lightweight, tool‑free MLLM can serve as a speculative planner to predict the execution trajectory, enabling early termination of expensive tool chains without sacrificing accuracy. To regulate this speculative planning, we introduce a cognitive gating mechanism based on answer separability, which quantifies the model's confidence for self‑verification without requiring oracle labels. Furthermore, we design a heterogeneous parallel funnel that exploits the stateless concurrency of the small model to mask the stateful serial execution of the large model, maximizing system throughput. Extensive experiments on V Bench, HR‑Bench, and POPE demonstrate that SpecEyes achieves 1.1‑3.35x speedup over the agentic baseline while preserving or even improving accuracy (up to +6.7%), thereby boosting serving throughput under concurrent workloads.
Authors:Haoran Yuan, Weigang Yi, Zhenyu Zhang, Wendi Chen, Yuchen Mo, Jiashi Yin, Xinzhuo Li, Xiangyu Zeng, Chuan Wen, Cewu Lu, Katherine Driggs-Campbell, Ismini Lourentzou
Abstract:
Video‑Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent action predictions. Although such models demonstrate strong performance on long‑horizon tasks through visual reasoning, they remain limited in contact‑rich scenarios where critical interaction states are only partially observable from vision alone. In particular, fine‑grained force modulation and contact transitions are not reliably encoded in visual tokens, leading to unstable or imprecise behaviors. To bridge this gap, we introduce the Video‑Tactile Action Model (VTAM), a multimodal world modeling framework that incorporates tactile perception as a complementary grounding signal. VTAM augments a pretrained video transformer with tactile streams via a lightweight modality transfer finetuning, enabling efficient cross‑modal representation learning without tactile‑language paired data or independent tactile pretraining. To stabilize multimodal fusion, we introduce a tactile regularization loss that enforces balanced cross‑modal attention, preventing visual latent dominance in the action model. VTAM demonstrates superior performance in contact‑rich manipulation, maintaining a robust success rate of 90 percent on average. In challenging scenarios such as potato chip pick‑and‑place requiring high‑fidelity force awareness, VTAM outperforms the pi 0.5 baseline by 80 percent. Our findings demonstrate that integrating tactile feedback is essential for correcting visual estimation errors in world action models, providing a scalable approach to physically grounded embodied foundation models.
Authors:Dana Cohen-Bar, Ido Sobol, Raphael Bensadoun, Shelly Sheynin, Oran Gafni, Or Patashnik, Daniel Cohen-Or, Amit Zohar
Abstract:
State‑of‑the‑art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these models cannot guarantee 3D consistency. Conversely, 3D engines offer granular control over every scene element and provide native 3D consistency by design, yet their output often remains trapped in the "uncanny valley". Bridging this sim‑to‑real gap requires both structural precision, where the output must exactly preserve the geometry and dynamics of the input, and global semantic transformation, where materials, lighting, and textures must be holistically transformed to achieve photorealism. We present RealMaster, a method that leverages video diffusion models to lift rendered video into photorealistic video while maintaining full alignment with the output of the 3D engine. To train this model, we generate a paired dataset via an anchor‑based propagation strategy, where the first and last frames are enhanced for realism and propagated across the intermediate frames using geometric conditioning cues. We then train an IC‑LoRA on these paired videos to distill the high‑quality outputs of the pipeline into a model that generalizes beyond the pipeline's constraints, handling objects and characters that appear mid‑sequence and enabling inference without requiring anchor frames. Evaluated on complex GTA‑V sequences, RealMaster significantly outperforms existing video editing baselines, improving photorealism while preserving the geometry, dynamics, and identity specified by the original 3D control.
Authors:Gautam Rajendrakumar Gare, Neehar Peri, Matvei Popov, Shruti Jain, John Galeotti, Deva Ramanan
Abstract:
Multi‑Modal LLMs (MLLMs) demonstrate strong visual grounding capabilities on popular object detection benchmarks like OdinW‑13 and RefCOCO. However, state‑of‑the‑art models still struggle to generalize to out‑of‑distribution classes, tasks and imaging modalities not typically found in their pre‑training. While in‑context prompting is a common strategy to improve performance across diverse tasks, we find that it often yields lower detection accuracy than prompting with class names alone. This suggests that current MLLMs cannot yet effectively leverage few‑shot visual examples and rich textual descriptions for object detection. Since frontier MLLMs are typically only accessible via APIs, and state‑of‑the‑art open‑weights models are prohibitively expensive to fine‑tune on consumer‑grade hardware, we instead explore black‑box prompt optimization for few‑shot object detection. To this end, we propose Detection Prompt Optimization (DetPO), a gradient‑free test‑time optimization approach that refines text‑only prompts by maximizing detection accuracy on few‑shot visual training examples while calibrating prediction confidence. Our proposed approach yields consistent improvements across generalist MLLMs on Roboflow20‑VL and LVIS, outperforming prior black‑box approaches by up to 9.7%. Our code is available at https://github.com/ggare‑cmu/DetPO
Authors:Yiping Chen, Jinpeng Li, Wenyu Ke, Yang Luo, Jie Ouyang, Zhongjie He, Li Liu, Hongchao Fan, Hao Wu
Abstract:
While multi‑modality large language models excel in object‑centric or indoor scenarios, scaling them to 3D city‑scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity‑LLM, a unified framework designed for 3D city‑scale vision‑language perception and understanding. 3DCity‑LLM employs a coarse‑to‑fine feature encoding strategy comprising three parallel branches for target object, inter‑object relationship, and global scene. To facilitate large‑scale training, we introduce 3DCity‑LLM‑1.2M dataset that comprises approximately 1.2 million high‑quality samples across seven representative task categories, ranging from fine‑grained object analysis to multi‑faceted scene planning. This strictly quality‑controlled dataset integrates explicit 3D numerical information and diverse user‑oriented simulations, enriching the question‑answering diversity and realism of urban scenarios. Furthermore, we apply a multi‑dimensional protocol based on text‑similarity metrics and LLM‑based semantic assessment to ensure faithful and comprehensive evaluations for all methods. Extensive experiments on two benchmarks demonstrate that 3DCity‑LLM significantly outperforms existing state‑of‑the‑art methods, offering a promising and meaningful direction for advancing spatial reasoning and urban intelligence. The source code and dataset are available at https://github.com/SYSU‑3DSTAILab/3D‑City‑LLM.
Authors:Jia Li, Han Yan, Yihang Chen, Siqi Li, Xibin Song, Yifu Wang, Jianfei Cai, Tien-Tsin Wong, Pan Ji
Abstract:
Despite remarkable progress in video generation, maintaining long‑term scene consistency upon revisiting previously explored areas remains challenging. Existing solutions rely either on explicitly constructing 3D geometry, which suffers from error accumulation and scale ambiguity, or on naive camera Field‑of‑View (FoV) retrieval, which typically fails under complex occlusions. To overcome these limitations, we propose I3DM, a novel implicit 3D‑aware memory mechanism for consistent video scene generation that bypasses explicit 3D reconstruction. At the core of our approach is a 3D‑aware memory retrieval strategy, which leverages the intermediate features of a pre‑trained Feed‑Forward Novel View Synthesis (FF‑NVS) model to score view relevance, enabling robust retrieval even in highly occluded scenarios. Furthermore, to fully utilize the retrieved historical frames, we introduce a 3D‑aligned memory injection module. This module implicitly warps historical content to the target view and adaptively conditions the generation on reliable warping regions, leading to improved revisit consistency and accurate camera control. Extensive experiments demonstrate that our method outperforms state‑of‑the‑art approaches, achieving superior revisit consistency, generation fidelity, and camera control precision.
Authors:Joelle Hanna, Damian Falk, Stella X. Yu, Damian Borth
Abstract:
Recent advances in remote sensing have led to an increase in the number of available foundation models; each trained on different modalities, datasets, and objectives, yet capturing only part of the vast geospatial knowledge landscape. While these models show strong results within their respective domains, their capabilities remain complementary rather than unified. Therefore, instead of choosing one model over another, we aim to combine their strengths into a single shared representation. We introduce GeoSANE, a geospatial model foundry that learns a unified neural representation from the weights of existing foundation models and task‑specific models, able to generate novel neural networks weights on‑demand. Given a target architecture, GeoSANE generates weights ready for finetuning for classification, segmentation, and detection tasks across multiple modalities. Models generated by GeoSANE consistently outperform their counterparts trained from scratch, match or surpass state‑of‑the‑art remote sensing foundation models, and outperform models obtained through pruning or knowledge distillation when generating lightweight networks. Evaluations across ten diverse datasets and on GEO‑Bench confirm its strong generalization capabilities. By shifting from pre‑training to weight generation, GeoSANE introduces a new framework for unifying and transferring geospatial knowledge across models and tasks. Code is available at \hrefhttps://hsg‑aiml.github.io/GeoSANE/hsg‑aiml.github.io/GeoSANE/.
Authors:Xinyu Liu, Zhen Chen, Wuyang Li, Chenxin Li, Yixuan Yuan
Abstract:
Transformers have shown remarkable performance in 3D medical image segmentation, but their high computational requirements and need for large amounts of labeled data limit their applicability. To address these challenges, we consider two crucial aspects: model efficiency and data efficiency. Specifically, we propose Light‑UNETR, a lightweight transformer designed to achieve model efficiency. Light‑UNETR features a Lightweight Dimension Reductive Attention (LIDR) module, which reduces spatial and channel dimensions while capturing both global and local features via multi‑branch attention. Additionally, we introduce a Compact Gated Linear Unit (CGLU) to selectively control channel interaction with minimal parameters. Furthermore, we introduce a Contextual Synergic Enhancement (CSE) learning strategy, which aims to boost the data efficiency of Transformers. It first leverages the extrinsic contextual information to support the learning of unlabeled data with Attention‑Guided Replacement, then applies Spatial Masking Consistency that utilizes intrinsic contextual information to enhance the spatial context reasoning for unlabeled data. Extensive experiments on various benchmarks demonstrate the superiority of our approach in both performance and efficiency. For example, with only 10% labeled data on the Left Atrial Segmentation dataset, our method surpasses BCP by 1.43% Jaccard while drastically reducing the FLOPs by 90.8% and parameters by 85.8%. Code is released at https://github.com/CUHK‑AIM‑Group/Light‑UNETR.
Authors:Feifan Luo, Hongyang Chen
Abstract:
Shape matching is a fundamental task in computer graphics and vision, with deep functional maps becoming a prominent paradigm. However, existing methods primarily focus on learning informative feature representations by constraining pointwise and functional maps, while neglecting the optimization of the spectral basis‑a critical component of the functional map pipeline. This oversight often leads to suboptimal matching results. Furthermore, many current approaches rely on conventional, time‑consuming functional map solvers, incurring significant computational overhead. To bridge these gaps, we introduce Advanced Functional Maps, a framework that generalizes standard functional maps by replacing fixed basis functions with learnable ones, supported by rigorous theoretical guarantees. Specifically, the spectral basis is optimized through a set of learned inhibition functions. Building on this, we propose the first unsupervised spectral basis learning method for robust non‑rigid 3D shape matching, enabling the joint, end‑to‑end optimization of feature extraction and basis functions. Our approach incorporates a novel heat diffusion module and an unsupervised loss function, alongside a streamlined architecture that bypasses expensive solvers and auxiliary losses. Extensive experiments demonstrate that our method significantly outperforms state‑of‑the‑art feature‑learning approaches, particularly in challenging non‑isometric and topological noise scenarios, while maintaining high efficiency. Finally, we reveal that optimizing basis functions is equivalent to spectral convolution, where inhibition functions act as filters. This insight enables enhanced representations inspired by spectral graph networks, opening new avenues for future research. Our code is available at https://github.com/LuoFeifan77/Unsupervised‑Spectral‑Basis‑Learning.
Authors:Yuzhi Chen, Ronghan Chen, Dongjie Huo, Yandan Yang, Dekang Qi, Haoyun Liu, Tong Lin, Shuang Zeng, Junjin Xiao, Xinyuan Chang, Feng Xiong, Xing Wei, Zhiheng Ma, Mu Xu
Abstract:
Video‑based world models offer a powerful paradigm for embodied simulation and planning, yet state‑of‑the‑art models often generate physically implausible manipulations ‑ such as object penetration and anti‑gravity motion ‑ due to training on generic visual data and likelihood‑based objectives that ignore physical laws. We present ABot‑PhysWorld, a 14B Diffusion Transformer model that generates visually realistic, physically plausible, and action‑controllable videos. Built on a curated dataset of three million manipulation clips with physics‑aware annotation, it uses a novel DPO‑based post‑training framework with decoupled discriminators to suppress unphysical behaviors while preserving visual quality. A parallel context block enables precise spatial action injection for cross‑embodiment control. To better evaluate generalization, we introduce EZSbench, the first training‑independent embodied zero‑shot benchmark combining real and synthetic unseen robot‑task‑scene combinations. It employs a decoupled protocol to separately assess physical realism and action alignment. ABot‑PhysWorld achieves new state‑of‑the‑art performance on PBench and EZSbench, surpassing Veo 3.1 and Sora v2 Pro in physical plausibility and trajectory consistency. We will release EZSbench to promote standardized evaluation in embodied video generation.
Authors:Weihang Li, Lorenzo Garattoni, Fabien Despinoy, Nassir Navab, Benjamin Busam
Abstract:
Learning model‑free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category‑level approaches predict absolute poses in a canonical space but rely on predefined taxonomies, while relative pose methods estimate cross‑view transformations but cannot recover single‑view absolute pose. In this work, we propose Object Pose Transformer (\ours), a unified feed‑forward framework that bridges these paradigms through task factorization within a single model. \ours jointly predicts depth, point maps, camera parameters, and normalized object coordinates (NOCS) from RGB inputs, enabling both category‑level absolute SA(3) pose and unseen‑object relative SE(3) pose. Our approach leverages contrastive object‑centric latent embeddings for canonicalization without requiring semantic labels at inference time, and uses point maps as a camera‑space representation to enable multi‑view relative geometric reasoning. Through cross‑frame feature interaction and shared object embeddings, our model leverages relative geometric consistency across views to improve absolute pose estimation, reducing ambiguity in single‑view predictions. Furthermore, \ours is camera‑agnostic, learning camera intrinsics on‑the‑fly and supporting optional depth input for metric‑scale recovery, while remaining fully functional in RGB‑only settings. Extensive experiments on diverse benchmarks (NOCS, HouseCat6D, Omni6DPose, Toyota‑Light) demonstrate state‑of‑the‑art performance in both absolute and relative pose estimation tasks within a single unified architecture.
Authors:Yunfeng Wu, Hongying Cheng, Zihao He, Songhua Liu
Abstract:
Transformer‑based video diffusion models rely on 3D attention over spatial and temporal tokens, which incurs quadratic time and memory complexity and makes end‑to‑end training for ultra‑high‑resolution videos prohibitively expensive. To overcome this bottleneck, we propose a pure image adaptation framework that upgrades a video Diffusion Transformer pre‑trained at its native scale to synthesize higher‑resolution videos. Unfortunately, naively fine‑tuning with high‑resolution images alone often introduces noticeable noise due to the image‑video modality gap. To address this, we decouple the learning objective to separately handle modality alignment and spatial extrapolation. At the core of our approach is Relay LoRA, a two‑stage adaptation strategy. In the first stage, the video diffusion model is adapted to the image domain using low‑resolution images to bridge the modality gap. In the second stage, the model is further adapted with high‑resolution images to acquire spatial extrapolation capability. During inference, only the high‑resolution adaptation is retained to preserve the video generation modality while enabling high‑resolution video synthesis. To enhance fine‑grained detail synthesis, we further propose a High‑Frequency‑Awareness‑Training‑Objective, which explicitly encourages the model to recover high‑frequency components from degraded latent representations via a dedicated reconstruction loss. Extensive experiments demonstrate that our method produces ultra‑high‑resolution videos with rich visual details without requiring any video training data, even outperforming previous state‑of‑the‑art models trained on high‑resolution videos by 0.8 on the VBench benchmark. Code will be available at https://github.com/WillWu111/ViBe.
Authors:Chuanqing Zhuang, Xin Lu, Zehui Deng, Zhengda Lu, Yiqun Wang, Junqi Diao, Jun Xiao
Abstract:
Omnidirectional 3D Gaussian Splatting with panoramas is a key technique for 3D scene representation, and existing methods typically rely on slow SfM to provide camera poses and sparse points priors. In this work, we propose a pose‑free omnidirectional 3DGS method, named PFGS360, that reconstructs 3D Gaussians from unposed omnidirectional videos. To achieve accurate camera pose estimation, we first construct a spherical consistency‑aware pose estimation module, which recovers poses by establishing consistent 2D‑3D correspondences between the reconstructed Gaussians and the unposed images using Gaussians' internal depth priors. Besides, to enhance the fidelity of novel view synthesis, we introduce a depth‑inlier‑aware densification module to extract depth inliers and Gaussian outliers with consistent monocular depth priors, enabling efficient Gaussian densification and achieving photorealistic novel view synthesis. The experiments show significant outperformance over existing pose‑free and pose‑aware 3DGS methods on both real‑world and synthetic 360‑degree videos. Code is available at https://github.com/zcq15/PFGS360.
Authors:Ezgi Ozyilkan, Zhiqi Chen, Oren Rippel, Jona Ballé, Kedar Tatwawadi
Abstract:
Despite their output being ultimately consumed by human viewers, 3D Gaussian Splatting (3DGS) methods often rely on ad‑hoc combinations of pixel‑level losses, resulting in blurry renderings. To address this, we systematically explore perceptual optimization strategies for 3DGS by searching over a diverse set of distortion losses. We conduct the first‑of‑its‑kind large‑scale human subjective study on 3DGS, involving 39,320 pairwise ratings across several datasets and 3DGS frameworks. A regularized version of Wasserstein Distortion, which we call WD‑R, emerges as the clear winner, excelling at recovering fine textures without incurring a higher splat count. WD‑R is preferred by raters more than 2.3× over the original 3DGS loss, and 1.5× over current best method Perceptual‑GS. WD‑R also consistently achieves state‑of‑the‑art LPIPS, DISTS, and FID scores across various datasets, and generalizes across recent frameworks, such as Mip‑Splatting and Scaffold‑GS, where replacing the original loss with WD‑R consistently enhances perceptual quality within a similar resource budget (number of splats for Mip‑Splatting, model size for Scaffold‑GS), and leads to reconstructions being preferred by human raters 1.8× and 3.6×, respectively. We also find that this carries over to the task of 3DGS scene compression, with \approx 50% bitrate savings for comparable perceptual metric performance.
Authors:Xinyong Cai, Runming Xie, Hu Chen, Yuankai Wu
Abstract:
Spatiotemporal predictive learning aims to forecast future frames from historical observations in an unsupervised manner, and is critical to a wide range of applications. The key challenge is to model long‑range dynamics while preserving high‑frequency details for sharp multi‑step predictions. Existing efficient recurrent‑free frameworks typically rely on strided convolutions or pooling for sampling, which tends to discard textures and boundaries, while purely spatial operators often struggle to balance local interactions with global propagation. To address these issues, we propose WaveSFNet, an efficient framework that unifies a wavelet‑based codec with a spatial‑‑frequency dual‑domain gated spatiotemporal translator. The wavelet‑based codec preserves high‑frequency subband cues during downsampling and reconstruction. Meanwhile, the translator first injects adjacent‑frame differences to explicitly enhance dynamic information, and then performs dual‑domain gated fusion between large‑kernel spatial local modeling and frequency‑domain global modulation, together with gated channel interaction for cross‑channel feature exchange. Extensive experiments demonstrate that WaveSFNet achieves competitive prediction accuracy on Moving MNIST, TaxiBJ, and WeatherBench, while maintaining low computational complexity. Our code is available at https://github.com/fhjdqaq/WaveSFNet.
Authors:Yuchen Wu, Kun Wang, Yining Pan, Na Zhao
Abstract:
Multi‑modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual‑branch proposal‑level detectors, we identify two factors that limit robust cross‑domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised. To address these challenges, we propose three components. First, Query‑Decoupled Loss provides independent supervision for 2D‑only, 3D‑only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR‑Guided Depth Prior augments 2D queries with instance‑aware geometric priors through probabilistic fusion of image‑predicted and LiDAR‑derived depth distributions, improving their spatial initialization. Third, Complementary Cross‑Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion. Extensive experiments demonstrate substantial gains over state‑of‑the‑art baselines while preserving source‑domain performance. Code and models are publicly available at https://github.com/IMPL‑Lab/CCF.
Authors:Zekai Gu, Shuoxuan Feng, Yansong Wang, Hanzhuo Huang, Zhongshuo Du, Chengfeng Zhao, Chengwei Ren, Peng Wang, Yuan Liu
Abstract:
Reconstructing a renderable 3D model from images is a useful but challenging task. Recent feedforward 3D reconstruction methods have demonstrated remarkable success in efficiently recovering geometry, but still cannot accurately model the complex appearances of these 3D reconstructed models. Recent diffusion‑based generative models can synthesize realistic images or videos of an object using reference images without explicitly modeling its appearance, which provides a promising direction for object rendering, but lacks accurate control over the viewpoints. In this paper, we propose GO‑Renderer, a unified framework integrating the reconstructed 3D proxies to guide the video generative models to achieve high‑quality object rendering on arbitrary viewpoints under arbitrary lighting conditions. Our method not only enjoys the accurate viewpoint control using the reconstructed 3D proxy but also enables high‑quality rendering in different lighting environments using diffusion generative models without explicitly modeling complex materials and lighting. Extensive experiments demonstrate that GO‑Renderer achieves state‑of‑the‑art performance across the object rendering tasks, including synthesizing images on new viewpoints, rendering the objects in a novel lighting environment, and inserting an object into an existing video.
Authors:Yukinori Yamamoto, Kazuya Nishimura, Tsukasa Fukusato, Hirokazu Nosato, Tetsuya Ogata, Hirokatsu Kataoka
Abstract:
Deep learning‑based 3D medical image segmentation methods relies on large‑scale labeled datasets, yet acquiring such data is difficult due to privacy constraints and the high cost of expert annotation. Formula‑Driven Supervised Learning (FDSL) offers an appealing alternative by generating training data and labels directly from mathematical formulas. However, existing voxel‑based approaches are limited in geometric expressiveness and cannot synthesize realistic textures. We introduce Formula‑Driven supervised learning with Implicit Functions (FDIF), a framework that enables scalable pre‑training without using any real data and medical expert annotations. FDIF introduces an implicit‑function representation based on signed distance functions (SDFs), enabling compact modeling of complex geometries while exploiting the surface representation of SDFs to support controllable synthesis of both geometric and intensity textures. Across three medical image segmentation benchmarks (AMOS, ACDC, and KiTS) and three architectures (SwinUNETR, nnUNet ResEnc‑L, and nnUNet Primus‑M), FDIF consistently improves over a formula‑driven method, and achieves performance comparable to self‑supervised approaches pre‑trained on large‑scale real datasets. We further show that FDIF pre‑training also benefits 3D classification tasks, highlighting implicit‑function‑based formula supervision as a promising paradigm for data‑free representation learning. Code is available at https://github.com/yamanoko/FDIF.
Authors:Yuanhang Lei, Tao Cheng, Xingxuan Li, Boming Zhao, Siyuan Huang, Ruizhen Hu, Peter Yichen Chen, Hujun Bao, Zhaopeng Cui
Abstract:
Achieving real‑time physics‑based animation that generalizes across diverse 3D shapes and discretizations remains a fundamental challenge. We introduce PhysSkin, a physics‑informed framework that addresses this challenge. In the spirit of Linear Blend Skinning, we learn continuous skinning fields as basis functions lifting motion subspace coordinates to full‑space deformation, with subspace defined by handle transformations. To generate mesh‑free, discretization‑agnostic, and physically consistent skinning fields that generalize well across diverse 3D shapes, PhysSkin employs a new neural skinning fields autoencoder which consists of a transformer‑based encoder and a cross‑attention decoder. Furthermore, we also develop a novel physics‑informed self‑supervised learning strategy that incorporates on‑the‑fly skinning‑field normalization and conflict‑aware gradient correction, enabling effective balancing of energy minimization, spatial smoothness, and orthogonality constraints. PhysSkin shows outstanding performance on generalizable neural skinning and enables real‑time physics‑based animation.
Authors:Yuqin Lu, Haofeng Liu, Yang Zhou, Jun Liang, Shengfeng He, Jing Li
Abstract:
Diffusion models excel at 2D outpainting, but extending them to 360^\circ panoramic completion from unposed perspective images is challenging due to the geometric and topological mismatch between perspective projections and spherical panoramas. We present Gimbal360, a principled framework that explicitly bridges perspective observations and spherical panoramas. We introduce a Canonical Viewing Space that regularizes projective geometry and provides a consistent intermediate representation between the two domains. To anchor in‑the‑wild inputs to this space, we propose a Differentiable Auto‑Leveling module that stabilizes feature orientation without requiring camera parameters at inference. Panoramic generation also introduces a topological challenge. Standard generative architectures assume a bounded Euclidean image plane, while Equirectangular Projection (ERP) panoramas exhibit intrinsic S^1 periodicity. Euclidean operations therefore break boundary continuity. We address this mismatch by enforcing topological equivariance in the latent space to preserve seamless periodic structure. To support this formulation, we introduce Horizon360, a curated large‑scale dataset of gravity‑aligned panoramic environments. Extensive experiments show that explicitly standardizing geometric and topological priors enables Gimbal360 to achieve state‑of‑the‑art performance in structurally consistent 360^\circ scene completion.
Authors:Jingtao Zhou, Xuan Gao, Dongyu Liu, Junhui Hou, Yudong Guo, Juyong Zhang
Abstract:
We present GSwap, a novel consistent and realistic video head‑swapping system empowered by dynamic neural Gaussian portrait priors, which significantly advances the state of the art in face and head replacement. Unlike previous methods that rely primarily on 2D generative models or 3D Morphable Face Models (3DMM), our approach overcomes their inherent limitations, including poor 3D consistency, unnatural facial expressions, and restricted synthesis quality. Moreover, existing techniques struggle with full head‑swapping tasks due to insufficient holistic head modeling and ineffective background blending, often resulting in visible artifacts and misalignments. To address these challenges, GSwap introduces an intrinsic 3D Gaussian feature field embedded within a full‑body SMPL‑X surface, effectively elevating 2D portrait videos into a dynamic neural Gaussian field. This innovation ensures high‑fidelity, 3D‑consistent portrait rendering while preserving natural head‑torso relationships and seamless motion dynamics. To facilitate training, we adapt a pretrained 2D portrait generative model to the source head domain using only a few reference images, enabling efficient domain adaptation. Furthermore, we propose a neural re‑rendering strategy that harmoniously integrates the synthesized foreground with the original background, eliminating blending artifacts and enhancing realism. Extensive experiments demonstrate that GSwap surpasses existing methods in multiple aspects, including visual quality, temporal coherence, identity preservation, and 3D consistency.
Authors:August Leander Høeg, Sophia Wiinberg Bardenfleth, Hans Martin Kjer, Tim Bjørn Dyrby, Vedrana Andersen Dahl, Anders Bjorholm Dahl
Abstract:
Recent advances in volumetric super‑resolution (SR) have demonstrated strong performance in medical and scientific imaging, with transformer‑ and CNN‑based approaches achieving impressive results even at extreme scaling factors. In this work, we show that much of this performance stems from training on downsampled data rather than real low‑resolution scans. This reliance on downsampling is partly driven by the scarcity of paired high‑ and low‑resolution 3D datasets. To address this, we introduce VoDaSuRe, a large‑scale volumetric dataset containing paired high‑ and low‑resolution scans. When training models on VoDaSuRe, we reveal a significant discrepancy: SR models trained on downsampled data produce substantially sharper predictions than those trained on real low‑resolution scans, which smooth fine structures. Conversely, applying models trained on downsampled data to real scans preserves more structure but is inaccurate. Our findings suggest that current SR methods are overstated ‑ when applied to real data, they do not recover structures lost in low‑resolution scans and instead predict a smoothed average. We argue that progress in deep learning‑based volumetric SR requires datasets with paired real scans of high complexity, such as VoDaSuRe. Our dataset and code are publicly available through: https://augusthoeg.github.io/VoDaSuRe/
Authors:Dongwei Pan, Longwei Guo, Jiazhi Guan, Luying Huang, Yiding Li, Haojie Liu, Haocheng Feng, Wei He, Kaisiyuan Wang, Hang Zhou
Abstract:
Despite progress in speech‑to‑video synthesis, existing methods often struggle to capture cross‑individual dependencies and provide fine‑grained control over reactive behaviors in dyadic settings. To address these challenges, we propose InterDyad, a framework that enables naturalistic interactive dynamics synthesis via querying structural motion guidance. Specifically, we first design an Interactivity Injector that achieves video reenactment based on identity‑agnostic motion priors extracted from reference videos. Building upon this, we introduce a MetaQuery‑based modality alignment mechanism to bridge the gap between conversational audio and these motion priors. By leveraging a Multimodal Large Language Model (MLLM), our framework is able to distill linguistic intent from audio to dictate the precise timing and appropriateness of reactions. To further improve lip‑sync quality under extreme head poses, we propose Role‑aware Dyadic Gaussian Guidance (RoDG) for enhanced lip‑synchronization and spatial consistency. Finally, we introduce a dedicated evaluation suite with novelly designed metrics to quantify dyadic interaction. Comprehensive experiments demonstrate that InterDyad significantly outperforms state‑of‑the‑art methods in producing natural and contextually grounded two‑person interactions. Please refer to our project page for demo videos: https://interdyad.github.io/.
Authors:Jinzhe Tu, Ruilei Guo, Zihan Guo, Junxiao Yang, Shiyao Cui, Minlie Huang
Abstract:
Recent works have shown that Multimodal Large Language Models (MLLMs) are highly vulnerable to hidden‑pattern visual illusions, where the hidden content is imperceptible to models but obvious to humans. This deficiency highlights a perceptual misalignment between current MLLMs and humans, and also introduces potential safety concerns. To systematically investigate this failure, we introduce IlluChar, a comprehensive and challenging illusion dataset, and uncover a key underlying mechanism for the models' failure: high‑frequency attention bias, where the models are easily distracted by high‑frequency background textures in illusion images, causing them to overlook hidden patterns. To address the issue, we propose the Strategy of Multi‑Scale Perception (SMSP), a plug‑and‑play framework that aligns with human visual perceptual strategies. By suppressing distracting high‑frequency backgrounds, SMSP generates images closer to human perception. Our experiments demonstrate that SMSP significantly improves the performance of all evaluated MLLMs on illusion images, for instance, increasing the accuracy of Qwen3‑VL‑8B‑Instruct from 13.0% to 84.0%. Our work provides novel insights into MLLMs' visual perception, and offers a practical and robust solution to enhance it. Our code is publicly available at https://github.com/Tujz2023/SMSP.
Authors:Yik San Cheng, Runkai Zhao, Weidong Cai
Abstract:
2D visual foundation models, such as DINOv3, a self‑supervised model trained on large‑scale natural images, have demonstrated strong zero‑shot generalization, capturing both rich global context and fine‑grained structural cues. However, an analogous 3D foundation model for downstream volumetric neuroimaging remains lacking, largely due to the challenges of 3D image acquisition and the scarcity of high‑quality annotations. To address this gap, we propose to adapt the 2D visual representations learned by DINOv3 to a 3D biomedical segmentation model, enabling more data‑efficient and morphologically faithful neuronal reconstruction. Specifically, we design an inflation‑based adaptation strategy that inflates 2D filters into 3D operators, preserving semantic priors from DINOv3 while adapting to 3D neuronal volume patches. In addition, we introduce a topology‑aware skeleton loss to explicitly enforce structural fidelity of graph‑based neuronal arbor reconstruction. Extensive experiments on four neuronal imaging datasets, including two from BigNeuron and two public datasets, NeuroFly and CWMBS, demonstrate consistent improvements in reconstruction accuracy over SoTA methods, with average gains of 2.9% in Entire Structure Average, 2.8% in Different Structure Average, and 3.8% in Percentage of Different Structure. Code: https://github.com/yy0007/NeurINO.
Authors:Basit Alawode, Arif Mahmood, Muaz Khalifa Al-Radi, Shahad Albastaki, Asim Khan, Muhammad Bilal, Moshira Ali Abdalla, Mohammed Bennamoun, Sajid Javed
Abstract:
Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic information emerges from cellular morphology, regional tissue organization, and global context. Existing Computational Pathology (CPath) Multimodal Large Language Models (MLLMs) typically compress an entire WSI into a single embedding, which hinders fine‑grained grounding and ignores how pathologists synthesize evidence across different scales. We introduce MLLM‑HWSI, a Hierarchical WSI‑level MLLM that aligns visual features with pathology language at four distinct scales, cell as word, patch as phrase, region as sentence, and WSI as paragraph to support interpretable evidence‑grounded reasoning. MLLM‑HWSI decomposes each WSI into multi‑scale embeddings with scale‑specific projectors and jointly enforces (i) a hierarchical contrastive objective and (ii) a cross‑scale consistency loss, preserving semantic coherence from cells to the WSI. We compute diagnostically relevant patches and aggregate segmented cell embeddings into a compact cellular token per‑patch using a lightweight Cell‑Cell Attention Fusion (CCAF) transformer. The projected multi‑scale tokens are fused with text tokens and fed to an instruction‑tuned LLM for open‑ended reasoning, VQA, report, and caption generation tasks. Trained in three stages, MLLM‑HWSI achieves new SOTA results on 13 WSI‑level benchmarks across six CPath tasks. By aligning language with multi‑scale visual evidence, MLLM‑HWSI provides accurate, interpretable outputs that mirror diagnostic workflows and advance holistic WSI understanding. Code is available at: \hrefhttps://github.com/BasitAlawode/HWSI‑MLLMGitHub.
Authors:Guoyang Zhao, Weiqing Qi, Kai Zhang, Chenguang Zhang, Zeying Gong, Zhihai Bi, Kai Chen, Benshan Ma, Ming Liu, Jun Ma
Abstract:
Traffic Sign Recognition (TSR) is a core perception capability for autonomous driving, where robustness to cross‑region variation, long‑tailed categories, and semantic ambiguity is essential for reliable real‑world deployment. Despite steady progress in recognition accuracy, existing traffic sign datasets and benchmarks offer limited diagnostic insight into how different modeling paradigms behave under these practical challenges. We present TS‑1M, a large‑scale and globally diverse traffic sign dataset comprising over one million real‑world images across 454 standardized categories, together with a diagnostic benchmark designed to analyze model capability boundaries. Beyond standard train‑test evaluation, we provide a suite of challenge‑oriented settings, including cross‑region recognition, rare‑class identification, low‑clarity robustness, and semantic text understanding, enabling systematic and fine‑grained assessment of modern TSR models. Using TS‑1M, we conduct a unified benchmark across three representative learning paradigms: classical supervised models, self‑supervised pretrained models, and multimodal vision‑language models (VLMs). Our analysis reveals consistent paradigm‑dependent behaviors, showing that semantic alignment is a key factor for cross‑region generalization and rare‑category recognition, while purely visual models remain sensitive to appearance shift and data imbalance. Finally, we validate the practical relevance of TS‑1M through real‑scene autonomous driving experiments, where traffic sign recognition is integrated with semantic reasoning and spatial localization to support map‑level decision constraints. Overall, TS‑1M establishes a reference‑level diagnostic benchmark for TSR and provides principled insights into robust and semantic‑aware traffic sign perception. Project page: https://guoyangzhao.github.io/projects/ts1m.
Authors:ByeongCheol Lee, Hyun Seok Seong, Sangeek Hyun, Gilhan Park, WonJun Moon, Jae-Pil Heo
Abstract:
A sliding‑window inference strategy is commonly adopted in recent training‑free open‑vocabulary semantic segmentation methods to overcome limitation of the CLIP in processing high‑resolution images. However, this approach introduces a new challenge: each window is processed independently, leading to semantic discrepancy across windows. To address this issue, we propose Global‑Local Aligned CLIP~(GLA‑CLIP), a framework that facilitates comprehensive information exchange across windows. Rather than limiting attention to tokens within individual windows, GLA‑CLIP extends key‑value tokens to incorporate contextual cues from all windows. Nevertheless, we observe a window bias: outer‑window tokens are less likely to be attended, since query features are produced through interactions within the inner window patches, thereby lacking semantic grounding beyond their local context. To mitigate this, we introduce a proxy anchor, constructed by aggregating tokens highly similar to the given query from all windows, which provides a unified semantic reference for measuring similarity across both inner‑ and outer‑window patches. Furthermore, we propose a dynamic normalization scheme that adjusts attention strength according to object scale by dynamically scaling and thresholding the attention map to cope with small‑object scenarios. Moreover, GLA‑CLIP can be equipped on existing methods and broad their receptive field. Extensive experiments validate the effectiveness of GLA‑CLIP in enhancing training‑free open‑vocabulary semantic segmentation performance. Code is available at https://github.com/2btlFe/GLA‑CLIP.
Authors:Jintao Cheng, Haozhe Wang, Weibin Li, Gang Wang, Yipu Zhang, Xiaoyu Tang, Jin Wu, Xieyuanli Chen, Yunhui Liu, Wei Zhang
Abstract:
Vision‑Language‑Action (VLA) models have rapidly advanced embodied intelligence, enabling robots to execute complex, instruction‑driven tasks. However, as model capacity and visual context length grow, the inference cost of VLA systems becomes a major bottleneck for real‑world deployment on resource‑constrained platforms. Existing visual token pruning methods mainly rely on semantic saliency or simple temporal cues, overlooking the continuous physical interaction, a fundamental property of VLA tasks. Consequently, current approaches often prune visually sparse yet structurally critical regions that support manipulation, leading to unstable behavior during early task phases. To overcome this, we propose a shift toward an explicit Interaction‑First paradigm. Our proposed training‑free method, VLA‑IAP (Interaction‑Aligned Pruning), introduces a geometric prior mechanism to preserve structural anchors and a dynamic scheduling strategy that adapts pruning intensity based on semantic‑motion alignment. This enables a conservative‑to‑aggressive transition, ensuring robustness during early uncertainty and efficiency once interaction is locked. Extensive experiments show that VLA‑IAP achieves a 97.8% success rate with a 1.25× speedup on the LIBERO benchmark, and up to 1.54× speedup while maintaining performance comparable to the unpruned backbone. Moreover, the method demonstrates superior and consistent performance across multiple model architectures and three different simulation environments, as well as a real robot platform, validating its strong generalization capability and practical applicability. Our project website is: \hrefhttps://chengjt1999.github.io/VLA‑IAP.github.io/VLA‑IAP.com.
Authors:Manuel-Andreas Schneider, Angela Dai
Abstract:
Recent progress in image and video synthesis has inspired their use in advancing 3D scene generation. However, we observe that text‑to‑image and video approaches struggle to maintain scene‑ and object‑level consistency beyond a limited environment scale due to the absence of explicit geometry. We thus present a geometry‑first approach that decouples this complex problem of large‑scale 3D scene synthesis into its structural composition, represented as a mesh scaffold, and realistic appearance synthesis, which leverages powerful image synthesis models conditioned on the mesh scaffold. From an input text description, we first construct a mesh capturing the environment's geometry (walls, floors, etc.), and then use image synthesis, segmentation and object reconstruction to populate the mesh structure with objects in realistic layouts. This mesh scaffold is then rendered to condition image synthesis, providing a structural backbone for consistent appearance generation. This enables scalable, arbitrarily‑sized 3D scenes of high object richness and diversity, combining robust 3D consistency with photorealistic detail. We believe this marks a significant step toward generating truly environment‑scale, immersive 3D worlds.
Authors:Yue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang, Mingzhe Zheng, Xiangpeng Yang, Hao Li, Chongbo Zhao, Jixuan Ying, Harry Yang, Hongyu Liu, Qifeng Chen
Abstract:
In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images may vary significantly in pose, viewpoint, and spatial layout. Achieving coherent edits requires establishing reliable correspondences across the images, so that modifications can be applied accurately to semantically aligned regions. To address this, we propose GroupEditing, a novel framework that builds both explicit and implicit relationships among images within a group. On the explicit side, we extract geometric correspondences using VGGT, which provides spatial alignment based on visual features. On the implicit side, we reformulate the image group as a pseudo‑video and leverage the temporal coherence priors learned by pre‑trained video models to capture latent relationships. To effectively fuse these two types of correspondences, we inject the explicit geometric cues from VGGT into the video model through a novel fusion mechanism. To support large‑scale training, we construct GroupEditData, a new dataset containing high‑quality masks and detailed captions for numerous image groups. Furthermore, to ensure identity preservation during editing, we introduce an alignment‑enhanced RoPE module, which improves the model's ability to maintain consistent appearance across multiple images. Finally, we present GroupEditBench, a dedicated benchmark designed to evaluate the effectiveness of group‑level image editing. Extensive experiments demonstrate that GroupEditing significantly outperforms existing methods in terms of visual quality, cross‑view consistency, and semantic alignment.
Authors:Wei Luo, Haiming Yao, Wenyong Yu
Abstract:
Industrial anomaly detection plays a crucial role in ensuring product quality control. Therefore, proposing an effective anomaly detection model is of great significance. While existing feature‑reconstruction methods have demonstrated excellent performance, they face challenges with shortcut learning, which can lead to undesirable reconstruction of anomalous features. To address this concern, we present a novel feature‑reconstruction model called the Template‑based Feature Aggregation Network (TFA‑Net) for anomaly detection via template‑based feature aggregation. Specifically, TFA‑Net first extracts multiple hierarchical features from a pre‑trained convolutional neural network for a fixed template image and an input image. Instead of directly reconstructing input features, TFA‑Net aggregates them onto the template features, effectively filtering out anomalous features that exhibit low similarity to normal template features. Next, TFA‑Net utilizes the template features that have already fused normal features in the input features to refine feature details and obtain the reconstructed feature map. Finally, the defective regions can be located by comparing the differences between the input and reconstructed features. Additionally, a random masking strategy for input features is employed to enhance the overall inspection performance of the model. Our template‑based feature aggregation schema yields a nontrivial and meaningful feature reconstruction task. The simple, yet efficient, TFA‑Net exhibits state‑of‑the‑art detection performance on various real‑world industrial datasets. Additionally, it fulfills the real‑time demands of industrial scenarios, rendering it highly suitable for practical applications in the industry. Code is available at https://github.com/luow23/TFA‑Net.
Authors:Amber Yijia Zheng, Yu-Shan Tai, Raymond A. Yeh
Abstract:
Recent advances in machine unlearning have focused on developing algorithms to remove specific training samples from a trained model. In contrast, we observe that not all models are equally easy to unlearn. Hence, we introduce a family of deep semi‑parametric models (SPMs) that exhibit non‑parametric behavior during unlearning. SPMs use a fusion module that aggregates information from each training sample, enabling explicit test‑time deletion of selected samples without altering model parameters. Empirically, we demonstrate that SPMs achieve competitive task performance to parametric models in image classification and generation, while being significantly more efficient for unlearning. Notably, on ImageNet classification, SPMs reduce the prediction gap relative to a retrained (oracle) baseline by 11% and achieve over 10× faster unlearning compared to existing approaches on parametric models. The code is available at https://github.com/amberyzheng/spm_unlearning.
Authors:Wei Luo, Haiming Yao, Zhenfeng Qiang, Xiaotian Zhang, Weihang Zhang
Abstract:
Unsupervised anomaly detection is vital in industrial fields, with reconstruction‑based methods favored for their simplicity and effectiveness. However, reconstruction methods often encounter an identical shortcut issue, where both normal and anomalous regions can be well reconstructed and fail to identify outliers. The severity of this problem increases with the complexity of the normal data distribution. Consequently, existing methods may exhibit excellent detection performance in a specific scenario, but their performance sharply declines when transferred to another scenario. This paper focuses on establishing a universal model applicable to anomaly detection tasks across different settings, termed as universal anomaly detection. In this work, we introduce a novel, straightforward yet efficient framework for universal anomaly detection: \ulineFeature \ulineShuffling and \ulineRestoration (FSR), which can alleviate the identical shortcut issue across different settings. First and foremost, FSR employs multi‑scale features with rich semantic information as reconstruction targets, rather than raw image pixels. Subsequently, these multi‑scale features are partitioned into non‑overlapping feature blocks, which are randomly shuffled and then restored to their original state using a restoration network. This simple paradigm encourages the model to focus more on global contextual information. Additionally, we introduce a novel concept, the shuffling rate, to regulate the complexity of the FSR task, thereby alleviating the identical shortcut across different settings. Furthermore, we provide theoretical explanations for the effectiveness of FSR framework from two perspectives: network structure and mutual information. Extensive experimental results validate the superiority and efficiency of the FSR framework across different settings.Code is available at https://github.com/luow23/FSR.
Authors:Yunheng Li, Hangyi Kuang, Hengrui Zhang, Jiangxia Cao, Zhaojie Liu, Qibin Hou, Ming-Ming Cheng
Abstract:
Multimodal Chain‑of‑Thought (CoT) reasoning requires large vision‑language models to construct reasoning trajectories that interleave perceptual grounding with multi‑step inference. However, existing Reinforcement Learning with Verifiable Rewards (RLVR) methods typically optimize reasoning at a coarse granularity, treating CoT uniformly without distinguishing their varying degrees of visual grounding. In this work, we conduct a token‑level analysis of multimodal reasoning trajectories and show that successful reasoning is characterized by structured token dynamics reflecting both perceptual grounding and exploratory inference. Building upon this analysis, we propose Perception‑Exploration Policy Optimization (PEPO), which derives a perception prior from hidden state similarity and integrates it with token entropy through a smooth gating mechanism to produce token‑level advantages. PEPO integrates seamlessly with existing RLVR frameworks such as GRPO and DAPO, requiring neither additional supervision nor auxiliary branches. Extensive experiments across diverse multimodal benchmarks demonstrate consistent and robust improvements over strong RL baselines, spanning geometry reasoning, visual grounding, visual puzzle solving, and few‑shot classification, while maintaining stable training dynamics. Code: https://github.com/xzxxntxdy/PEPO
Authors:Jun Yang, Dong Wang, Hongxu Yin, Hongpeng Li, Jianxiong Yu
Abstract:
Drone detection is pivotal in numerous security and counter‑UAV applications. However, existing deep learning‑based methods typically struggle to balance robust feature representation with computational efficiency. This challenge is particularly acute when detecting miniature drones against complex backgrounds under severe environmental interference. To address these issues, we introduce UAV‑DETR, a novel framework that integrates a small‑target‑friendly architecture with real‑time detection capabilities. Specifically, UAV‑DETR features a WTConv‑enhanced backbone and a Sliding Window Self‑Attention (SWSA‑IFI) encoder, capturing the high‑frequency structural details of tiny targets while drastically reducing parameter overhead. Furthermore, we propose an Efficient Cross‑Scale Feature Recalibration and Fusion Network (ECFRFN) to suppress background noise and aggregate multi‑scale semantics. To further enhance accuracy, UAV‑DETR incorporates a hybrid Inner‑CIoU and NWD loss strategy, mitigating the extreme sensitivity of standard IoU metrics to minor positional deviations in small objects. Extensive experiments demonstrate that UAV‑DETR significantly outperforms the baseline RT‑DETR on our custom UAV dataset (+6.61% in mAP50:95, with a 39.8% reduction in parameters) and the public DUT‑ANTI‑UAV benchmark (+1.4% in Precision, +1.0% in F1‑Score). These results establish UAV‑DETR as a superior trade‑off between efficiency and precision in counter‑UAV object detection. The code is available at https://github.com/wd‑sir/UAVDETR.
Authors:Shiyu Li, Hannah Schieber, Kristoffer Waldow, Benjamin Busam, Julian Kreimeier, Daniel Roth
Abstract:
Multi‑camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information, such as markers or objects, across multiple views. While commonly cameras are calibrated in an initial step or updated through the constant use of markers, another option is to leverage information already present in the scene, like known objects. Another downside of marker‑based tracking is that markers have to be tracked inside the field‑of‑view (FoV) of the cameras.
To overcome these limitations, we propose a constant dynamic camera pose estimation leveraging spatiotemporal FoV overlaps of known objects on the fly. To achieve that, we enhance the state‑of‑the‑art object pose estimator to update our spatiotemporal scene graph, enabling a relation even among non‑overlapping FoV cameras. To evaluate our approach, we introduce a multi‑camera, multi‑object pose estimation dataset with temporal FoV overlap, including static and dynamic cameras. Furthermore, in FoV overlapping scenarios, we outperform the state‑of‑the‑art on the widely used YCB‑V and T‑LESS dataset in camera pose accuracy. Our performance on both previous and our proposed datasets validates the effectiveness of our marker‑less approach for AR applications.
The code and dataset are available on https://github.com/roth‑hex‑lab/IEEE‑VR‑2026‑MultiCam.
Authors:Chunxia Qin, Chenyu Liu, Pengcheng Xia, Jun Du, Baocai Yin, Bing Yin, Cong Liu
Abstract:
Tables are pervasive in diverse documents, making table recognition (TR) a fundamental task in document analysis. Existing modular TR pipelines separately model table structure and content, leading to suboptimal integration and complex workflows. End‑to‑end approaches rely heavily on large‑scale TR data and struggle in data‑constrained scenarios. To address these issues, we propose TDATR (Table Detail‑Aware Table Recognition) improves end‑to‑end TR through table detail‑aware learning and cell‑level visual alignment. TDATR adopts a ``perceive‑then‑fuse'' strategy. The model first performs table detail‑aware learning to jointly perceive table structure and content through multiple structure understanding and content recognition tasks designed under a language modeling paradigm. These tasks can naturally leverage document data from diverse scenarios to enhance model robustness. The model then integrates implicit table details to generate structured HTML outputs, enabling more efficient TR modeling when trained with limited data. Furthermore, we design a structure‑guided cell localization module integrated into the end‑to‑end TR framework, which efficiently locates cell and strengthens vision‑language alignment. It enhances the interpretability and accuracy of TR. We achieve state‑of‑the‑art or highly competitive performance on seven benchmarks without dataset‑specific fine‑tuning.
Authors:Lishen Qu, Shihao Zhou, Jie Liang, Hui Zeng, Lei Zhang, Jufeng Yang
Abstract:
Flicker artifacts, arising from unstable illumination and row‑wise exposure inconsistencies, pose a significant challenge in short‑exposure photography, severely degrading image quality. Unlike typical artifacts, e.g., noise and low‑light, flicker is a structured degradation with specific spatial‑temporal patterns, which are not accounted for in current generic restoration frameworks, leading to suboptimal flicker suppression and ghosting artifacts. In this work, we reveal that flicker artifacts exhibit two intrinsic characteristics, periodicity and directionality, and propose Flickerformer, a transformer‑based architecture that effectively removes flicker without introducing ghosting. Specifically, Flickerformer comprises three key components: a phase‑based fusion module (PFM), an autocorrelation feed‑forward network (AFFN), and a wavelet‑based directional attention module (WDAM). Based on the periodicity, PFM performs inter‑frame phase correlation to adaptively aggregate burst features, while AFFN exploits intra‑frame structural regularities through autocorrelation, jointly enhancing the network's ability to perceive spatially recurring patterns. Moreover, motivated by the directionality of flicker artifacts, WDAM leverages high‑frequency variations in the wavelet domain to guide the restoration of low‑frequency dark regions, yielding precise localization of flicker artifacts. Extensive experiments demonstrate that Flickerformer outperforms state‑of‑the‑art approaches in both quantitative metrics and visual quality. The source code is available at https://github.com/qulishen/Flickerformer.
Authors:Chamuditha Jayanga Galappaththige, Thomas Gottwald, Peter Stehr, Edgar Heinert, Niko Suenderhauf, Dimity Miller, Matthias Rottmann
Abstract:
Recent advances in 3D Gaussian Splatting have enabled impressive photorealistic novel view synthesis. However, to transition from a pure rendering engine to a reliable spatial map for autonomous agents and safety‑critical applications, knowing where the representation is uncertain is as important as the rendering fidelity itself. We bridge this critical gap by introducing a lightweight, plug‑and‑play framework for pixel‑wise, view‑dependent predictive uncertainty estimation. Our post‑hoc method formulates uncertainty as a Bayesian‑regularized linear least‑squares optimization over reconstruction residuals. This architecture‑agnostic approach extracts a per‑primitive uncertainty channel without modifying the underlying scene representation or degrading baseline visual fidelity. Crucially, we demonstrate that providing this actionable reliability signal successfully translates 3D Gaussian splatting into a trustworthy spatial map, further improving state‑of‑the‑art performance across three critical downstream perception tasks: active view selection, pose‑agnostic scene change detection, and pose‑agnostic anomaly detection.
Authors:Wenyue Chen, Wenjue Chen, Peng Li, Qinghe Wang, Xu Jia, Heliang Zheng, Rongfei Jia, Yuan Liu, Ronggang Wang
Abstract:
Recent advances in 3D generation have improved the fidelity and geometric details of synthesized 3D assets. However, due to the inherent ambiguity of single‑view observations and the lack of robust global structural priors caused by limited 3D training data, the unseen regions generated by existing models are often stochastic and difficult to control, which may sometimes fail to align with user intentions or produce implausible geometries. In this paper, we propose Know3D, a novel framework that incorporates rich knowledge from multimodal large language models into 3D generative processes via latent hidden‑state injection, enabling language‑controllable generation of the back‑view for 3D assets. We utilize a VLM‑diffusion‑based model, where the VLM is responsible for semantic understanding and guidance. The diffusion model acts as a bridge that transfers semantic knowledge from the VLM to the 3D generation model. In this way, we successfully bridge the gap between abstract textual instructions and the geometric reconstruction of unobserved regions, transforming the traditionally stochastic back‑view hallucination into a semantically controllable process, demonstrating a promising direction for future 3D generation models.
Authors:Jingwei Liao, Bo Chen, Klara Nahrstedt, Zhisheng Yan
Abstract:
Given the popularity of 360° images on social media platforms, 360° image compression becomes a critical technology for media storage and transmission. Conventional 360° image compression pipeline projects the spherical image into a single 2D plane, leading to issues of oversampling and distortion. In this paper, we propose a novel viewport‑based neural compression pipeline for 360° images. By replacing the image projection in conventional 360° image compression pipelines with viewport extraction and efficiently compressing multiple viewports, the proposed pipeline minimizes the inherent oversampling and distortion issues. However, viewport extraction impedes information sharing between multiple viewports during compression, causing the loss of global information about the spherical image. To tackle this global information loss, we design a neural viewport codec to capture global prior information across multiple viewports and maximally compress the viewport data. The viewport codec is empowered by a transformer‑based ViewPort ConText (VPCT) module that can be integrated with canonical learning‑based 2D image compression structures. We compare the proposed pipeline with existing 360° image compression models and conventional 360° image compression pipelines building on learning‑based 2D image codecs and standard hand‑crafted codecs. Results show that our pipeline saves an average of 14.01% bit consumption compared to the best‑performing 360° image compression methods without compromising quality. The proposed VPCT‑based codec also outperforms existing 2D image codecs in the viewport‑based neural compression pipeline. Our code can be found at: https://github.com/Jingwei‑Liao/VPCT.
Authors:Ao Cheng, Xingming Li, Xuanyu Ji, Xixiang He, Qiyao Sun, Chunping Qiu, Runke Huang, Qingyong Hu
Abstract:
Electronic Navigational Charts (ENCs) are the safety‑critical backbone of modern maritime navigation, yet it remains unclear whether multimodal large language models (MLLMs) can reliably interpret them. Unlike natural images or conventional charts, ENCs encode regulations, bathymetry, and route constraints via standardized vector symbols, scale‑dependent rendering, and precise geometric structure ‑‑ requiring specialized maritime expertise for interpretation. We introduce ENC‑Bench, the first benchmark dedicated to professional ENC understanding. ENC‑Bench contains 20,490 expert‑validated samples from 840 authentic National Oceanic and Atmospheric Administration (NOAA) ENCs, organized into a three‑level hierarchy: Perception (symbol and feature recognition), Spatial Reasoning (coordinate localization, bearing, distance), and Maritime Decision‑Making (route legality, safety assessment, emergency planning under multiple constraints). All samples are generated from raw S‑57 data through a calibrated vector‑to‑image pipeline with automated consistency checks and expert review. We evaluate 10 state‑of‑the‑art MLLMs such as GPT‑4o, Gemini 2.5, Qwen3‑VL, InternVL‑3, and GLM‑4.5V, under a unified zero‑shot protocol. The best model achieves only 47.88% accuracy, with systematic challenges in symbolic grounding, spatial computation, multi‑constraint reasoning, and robustness to lighting and scale variations. By establishing the first rigorous ENC benchmark, we open a new research frontier at the intersection of specialized symbolic reasoning and safety‑critical AI, providing essential infrastructure for advancing MLLMs toward professional maritime applications.
Authors:Purui Bai, Tao Wu, Jiayang Sun, Xinyue Liu, Huaibo Huang, Ran He
Abstract:
The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi‑modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however, are limited to static images or single videos, overlooking the complex interactions across multiple videos. To address this gap, we introduce the Multi‑Video Perception Evaluation Benchmark (MVPBench), a new benchmark featuring 14 subtasks across diverse visual domains designed to evaluate models on extracting relevant information from video sequences to make informed decisions. MVPBench includes 5K question‑answering tests involving 2.7K video clips sourced from existing datasets and manually annotated clips. Extensive evaluations reveal that current models struggle to process multi‑video inputs effectively, underscoring substantial limitations in their multi‑video comprehension. We anticipate MVPBench will drive advancements in multi‑video perception.
Authors:Jiayin Sun, Caixia Sun, Boyu Yang, Hailin Li, Xiao Chen, Yi Zhang, Errui Ding, Liang Li, Chao Deng, Junlan Feng
Abstract:
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine‑grained geometric structures, constraining their ability of geometric understanding and visual reasoning. To address this, we propose GeoTikzBridge, a framework that enhances local geometric perception and visual reasoning through tikz‑based code generation. Within this framework, we build two models supported by two complementary datasets. The GeoTikzBridge‑Base model is trained on GeoTikz‑Base dataset, the largest image‑to‑tikz dataset to date with 2.5M pairs (16 × larger than existing open‑sourced datasets). This process is achieved via iterative data expansion and a localized geometric transformation strategy. Subsequently, GeoTikzBridge‑Instruct is fine‑tuned on GeoTikz‑Instruct dataset which is the first instruction‑augmented tikz dataset supporting visual reasoning. Extensive experimental results demonstrate that our models achieve state‑of‑the‑art performance among open‑sourced MLLMs. Furthermore, GeoTikzBridge models can serve as plug‑and‑play reasoning modules for any MLLM(LLM), enhancing reasoning performance in geometric problem‑solving. Datasets and codes are publicly available at: https://github.com/sjy‑1995/GeoTikzBridge.
Authors:Shiyao Li, Antoine Guédon, Shizhe Chen, Vincent Lepetit
Abstract:
Active mapping aims to determine how an agent should move to efficiently reconstruct unknown environments. Most existing approaches rely on greedy next‑best‑view prediction, resulting in inefficient exploration and incomplete reconstruction. To address this, we introduce MAGICIAN, a novel long‑term planning framework that maximizes accumulated surface coverage gain through Imagined Gaussians, a scene representation based on 3D Gaussian Splatting, derived from a pre‑trained occupancy network with strong structural priors. This representation enables efficient coverage gain computation for any novel viewpoint via fast volumetric rendering, allowing its integration into a tree‑search algorithm for long‑horizon planning. We update Imagined Gaussians and refine the trajectory in a closed loop. Our method achieves state‑of‑the‑art performance across indoor and outdoor benchmarks with varying action spaces, highlighting the advantage of long‑term planning in active mapping.
Authors:Heejong Kim, Abhishek Thanki, Roel van Herten, Daniel Margolis, Mert R Sabuncu
Abstract:
Clinical MRI frequently acquires anisotropic volumes with high in‑plane resolution and low through‑plane resolution to reduce acquisition time. Multiple orientations are therefore acquired to provide complementary anatomical information. Conventional integration of these views relies on registration followed by interpolation, which can degrade fine structural details. Recent deep learning‑based super‑resolution (SR) approaches have demonstrated strong performance in enhancing single‑view images. However, their clinical reliability is often limited by the need for large‑scale training datasets, resulting in increased dependence on cohort‑level priors. Self‑supervised strategies offer an alternative by learning directly from the target scans. Prior work either neglects the existence of multi‑view information or assumes that in‑plane information can supervise through‑plane reconstruction under the assumption of pre‑alignment between images. However, this assumption is rarely satisfied in clinical settings. In this work, we introduce Single‑Subject Implicit Multi‑View Super‑Resolution for MRI (SIMS‑MRI), a framework that operates solely on anisotropic multi‑view scans from a single patient without requiring pre‑ or post‑processing. Our method combines a multi‑resolution hash‑encoded implicit representation with learned inter‑view alignment to generate a spatially consistent isotropic reconstruction. We validate the SIMS‑MRI pipeline on both simulated brain and clinical prostate MRI datasets. Code will be made publicly available for reproducibility: https://github.com/abhshkt/SIMS‑MRI
Authors:Dinglun He, Baoming Zhang, Xu Wang, Yao Hao, Deshan Yang, Ye Duan
Abstract:
Abdominal CT data are limited by high annotation costs and privacy constraints, which hinder the development of robust segmentation and diagnostic models. We present a Prior‑Integrated Variation Modeling (PIVM) framework, a diffusion‑based method for anatomically accurate CT image synthesis. Instead of generating full images from noise, PIVM predicts voxel‑wise intensity variations relative to organ‑specific intensity priors derived from segmentation labels. These priors and labels jointly guide the diffusion process, ensuring spatial alignment and realistic organ boundaries. Unlike latent‑space diffusion models, our approach operates directly in image space while preserving the full Hounsfield Unit (HU) range, capturing fine anatomical textures without smoothing. Source code is available at https://github.com/BZNR3/PIVM.
Authors:Abu Noman Md Sakib, OFM Riaz Rahman Aranya, Kevin Desai, Zijie Zhang
Abstract:
Attribution maps for semantic segmentation are almost always judged by visual plausibility. Yet looking convincing does not guarantee that the highlighted pixels actually drive the model's prediction, nor that attribution credit stays within the target region. These questions require a dedicated evaluation protocol. We introduce a reproducible benchmark that tests intervention‑based faithfulness, off‑target leakage, perturbation robustness, and runtime on Pascal VOC and SBD across three pretrained backbones. To further demonstrate the benchmark, we propose Dual‑Evidence Attribution (DEA), a lightweight correction that fuses gradient evidence with region‑level intervention signals through agreement‑weighted fusion. DEA increases emphasis where both sources agree and retains causal support when gradient responses are unstable. Across all completed runs, DEA consistently improves deletion‑based faithfulness over gradient‑only baselines and preserves strong robustness, at the cost of additional compute from intervention passes. The benchmark exposes a faithfulness‑stability tradeoff among attribution families that is entirely hidden under visual evaluation, providing a foundation for principled method selection in segmentation explainability. Code is available at https://github.com/anmspro/DEA.
Authors:OFM Riaz Rahman Aranya, Kevin Desai
Abstract:
Vision‑language models (VLMs) adapted to the medical domain have shown strong performance on visual question answering benchmarks, yet their robustness against two critical failure modes, hallucination and sycophancy, remains poorly understood, particularly in combination. We evaluate six VLMs (three general‑purpose, three medical‑specialist) on three medical VQA datasets and uncover a grounding‑sycophancy tradeoff: models with the lowest hallucination propensity are the most sycophantic, while the most pressure‑resistant model hallucinates more than all medical‑specialist models. To characterize this tradeoff, we propose three metrics: L‑VASE, a logit‑space reformulation of VASE that avoids its double‑normalization; CCS, a confidence‑calibrated sycophancy score that penalizes high‑confidence capitulation; and Clinical Safety Index (CSI), a unified safety index that combines grounding, autonomy, and calibration via a geometric mean. Across 1,151 test cases, no model achieves a CSI above 0.35, indicating that none of the evaluated 7‑8B parameter VLMs is simultaneously well‑grounded and robust to social pressure. Our findings suggest that joint evaluation of both properties is necessary before these models can be considered for clinical use. Code is available at https://github.com/UTSA‑VIRLab/AgreeOrRight
Authors:Fulvio Sanguigni, Davide Lobba, Bin Ren, Marcella Cornia, Nicu Sebe, Rita Cucchiara
Abstract:
Recent advances in Virtual Try‑On (VTON) and Virtual Try‑Off (VTOFF) have greatly improved photo‑realistic fashion synthesis and garment reconstruction. However, existing datasets remain static, lacking instruction‑driven editing for controllable and interactive fashion generation. In this work, we introduce the Dress Editing Dataset (Dress‑ED), the first large‑scale benchmark that unifies VTON, VTOFF, and text‑guided garment editing within a single framework. Each sample in Dress‑ED includes an in‑shop garment image, the corresponding person image wearing the garment, their edited counterparts, and a natural‑language instruction of the desired modification. Built through a fully automated multimodal pipeline that integrates MLLM‑based garment understanding, diffusion‑based editing, and LLM‑guided verification, Dress‑ED comprises over 146k verified quadruplets spanning three garment categories and seven edit types, including both appearance (e.g., color, pattern, material) and structural (e.g., sleeve length, neckline) modifications. Based on this benchmark, we further propose a unified multimodal diffusion framework that jointly reasons over linguistic instructions and visual garment cues, serving as a strong baseline for instruction‑driven VTON and VTOFF. Dataset and code will be made publicly available. Project page: https://furio1999.github.io/Dress‑ED/
Authors:Zewei Zhang, Jia Jun Cheng Xian, Kaiwen Liu, Ming Liang, Hang Chu, Jun Chen, Renjie Liao
Abstract:
Predicting future motion is crucial in video understanding and controllable video generation. Dense point trajectories are a compact, expressive motion representation, but modeling their future evolution from observed video remains challenging. We propose a framework that predicts future trajectories and visibility from past trajectories and video context. Our method has three components: (1) Grid‑Anchor Offset Encoding, which reduces location‑dependent bias by representing each point as an offset from its pixel‑center anchor; (2) TrajLoom‑VAE, which learns a compact spatiotemporal latent space for dense trajectories with masked reconstruction and a spatiotemporal consistency regularizer; and (3) TrajLoom‑Flow, which generates future trajectories in latent space via flow matching, with boundary cues and on‑policy K‑step fine‑tuning for stable sampling. We also introduce TrajLoomBench, a unified benchmark spanning real and synthetic videos with a standardized setup aligned with video‑generation benchmarks. Compared with state‑of‑the‑art methods, our approach extends the prediction horizon from 24 to 81 frames while improving motion realism and stability across datasets. The predicted trajectories directly support downstream video generation and editing. Code, model checkpoints, and datasets are available at https://trajloom.github.io/.
Authors:Yalda Foroutan, Ipek Oztas, Daniel Rebain, Aysegul Dundar, Kwang Moo Yi, Lily Goli, Andrea Tagliasacchi
Abstract:
Radiance fields have emerged as powerful tools for 3D scene reconstruction. However, casual capture remains challenging due to the narrow field of view of perspective cameras, which limits viewpoint coverage and feature correspondences necessary for reliable camera calibration and reconstruction. While commercially available 360^\circ cameras offer significantly broader coverage than perspective cameras for the same capture effort, existing 360^\circ reconstruction methods require special capture protocols and pre‑processing steps that undermine the promise of radiance fields: effortless workflows to capture and reconstruct 3D scenes. We propose a practical pipeline for reconstructing 3D scenes directly from raw 360^\circ camera captures. We require no special capture protocols or pre‑processing, and exhibit robustness to a prevalent source of reconstruction errors: the human operator that is visible in all 360^\circ imagery. To facilitate evaluation, we introduce a multi‑tiered dataset of scenes captured as raw dual‑fisheye images, establishing a benchmark for robust casual 360^\circ reconstruction. Our method significantly outperforms not only vanilla 3DGS for 360^\circ cameras but also robust perspective baselines when perspective cameras are simulated from the same capture, demonstrating the advantages of 360^\circ capture for casual reconstruction. Additional results are available at: https://theialab.github.io/fullcircle
Authors:Yohaï-Eliel Berreby, Sabrina Du, Audrey Durand, B. Suresh Krishna
Abstract:
Active computer vision promises efficient, biologically plausible perception through sequential, localized glimpses, but lacks scalable general‑purpose architectures and pretraining pipelines, leaving Active‑Vision Foundation Models (AVFMs) underexplored. We introduce CanViT, the first task‑ and policy‑agnostic AVFM. CanViT uses scene‑relative RoPE to bind a retinotopic Vision Transformer backbone and a spatiotopic scene‑wide latent workspace, the canvas. Efficient interaction with this high‑capacity working memory is supported by Canvas Attention, a novel asymmetric cross‑attention mechanism. We decouple thinking (backbone‑level) and memory (canvas‑level), eliminating canvas‑side self‑attention and fully‑connected layers to achieve fast sequential inference and scalability to high output resolutions. We propose a label‑free active vision pretraining scheme, policy‑agnostic passive‑to‑active dense latent distillation: reconstructing scene‑wide DINOv3 embeddings from sequences of low‑resolution glimpses with randomized locations, zoom levels, and lengths. We pretrain CanViT‑B from a random initialization on 13.2 million ImageNet‑21k scenes‑‑an order of magnitude more than previous active models‑‑and 1 billion random glimpses, in 166 hours on a single H100. On ADE20K segmentation, a frozen CanViT‑B achieves 38.5% mIoU in a single low‑resolution glimpse, outperforming the best active model's 27.6% with 20x fewer inference FLOPs as well as its FLOP‑ or input‑matched DINOv3 teacher. Given additional glimpses, CanViT‑B reaches 45.9% ADE20K mIoU. On ImageNet‑1k classification, CanViT‑B also sets a new active‑vision state of the art, with 84.5% top‑1 accuracy after fine‑tuning. CanViT generalizes to longer rollouts, larger scenes, and new policies. Our work narrows the wide gap between passive and active computer vision, demonstrating the potential of task‑ and policy‑agnostic AVFM pretraining.
Authors:Delin An, Chaoli Wang
Abstract:
Diffusion probabilistic models have demonstrated significant potential in generating high‑quality, realistic medical images, providing a promising solution to the persistent challenge of data scarcity in the medical field. Nevertheless, producing 3D medical volumes with anatomically consistent structures under multimodal conditions remains a complex and unresolved problem. We introduce Sketch2CT, a multimodal diffusion framework for structure‑aware 3D medical volume generation, jointly guided by a user‑provided 2D sketch and a textual description that captures 3D geometric semantics. The framework initially generates 3D segmentation masks of the target organ from random noise, conditioned on both modalities. To effectively align and fuse these inputs, we propose two key modules that refine sketch features with localized textual cues and integrate global sketch‑text representations. Built upon a capsule‑attention backbone, these modules leverage the complementary strengths of sketches and text to produce anatomically accurate organ shapes. The synthesized segmentation masks subsequently guide a latent diffusion model for 3D CT volume synthesis, enabling realistic reconstruction of organ appearances that are consistent with user‑defined sketches and descriptions. Extensive experiments on public CT datasets demonstrate that Sketch2CT achieves superior performance in generating multimodal medical volumes. Its controllable, low‑cost generation pipeline enables principled, efficient augmentation of medical datasets. Code is available at https://github.com/adlsn/Sketch2CT.
Authors:Davide Bucciarelli, Evelyn Turri, Lorenzo Baraldi, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Abstract:
Inference‑time scaling has emerged as an effective way to improve generative models at test time by using a verifier to score and select candidate outputs. A common choice is to employ Multimodal Large Language Models (MLLMs) as verifiers, which can improve performance but introduce substantial inference‑time cost. Indeed, diffusion pipelines operate in an autoencoder latent space to reduce computation, yet MLLM verifiers still require decoding candidates to pixel space and re‑encoding them into the visual embedding space, leading to redundant and costly operations. In this work, we propose Verifier on Hidden States (VHS), a verifier that operates directly on intermediate hidden representations of Diffusion Transformer (DiT) single‑step generators. VHS analyzes generator features without decoding to pixel space, thereby reducing the per‑candidate verification cost while improving or matching the performance of MLLM‑based competitors. We show that, under tiny inference budgets with only a small number of candidates per prompt, VHS enables more efficient inference‑time scaling reducing joint generation‑and‑verification time by 63.3%, compute FLOPs by 51% and VRAM usage by 14.5% with respect to a standard MLLM verifier, achieving a +2.7% improvement on GenEval at the same inference‑time budget.
Authors:Chenchen Zhu, Saksham Suri, Cijo Jose, Maxime Oquab, Marc Szafraniec, Wei Wen, Yunyang Xiong, Patrick Labatut, Piotr Bojanowski, Raghuraman Krishnamoorthi, Vikas Chandra
Abstract:
Running AI models on smart edge devices can unlock versatile user experiences, but presents challenges due to limited compute and the need to handle multiple tasks simultaneously. This requires a vision encoder with small size but powerful and versatile representations. We present our method, Efficient Universal Perception Encoder (EUPE), which offers both inference efficiency and universally good representations for diverse downstream tasks. We achieve this by distilling from multiple domain‑expert foundation vision encoders. Unlike previous agglomerative methods that directly scale down from multiple teachers to an efficient encoder, we demonstrate the importance of first scaling up to a large proxy teacher and then scaling down from this single teacher. Experiments show that EUPE achieves on‑par or better performance than individual domain experts of the same size on diverse task domains and also outperforms previous agglomerative encoders. We release the full family of EUPE models and the code to foster future research.
Authors:Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker, Salman Khan, Fahad Shahbaz Khan
Abstract:
Diffusion Transformers (DiTs) power high‑fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio‑temporal attention. Training‑free feature caching accelerates inference by reusing intermediate activations across denoising steps; however, existing methods largely rely on a Zero‑Order Hold assumption i.e., reusing cached features as static snapshots when global drift is small. This often leads to ghosting artifacts, blur, and motion inconsistencies in dynamic scenes. We propose WorldCache, a Perception‑Constrained Dynamical Caching framework that improves both when and how to reuse features. WorldCache introduces motion‑adaptive thresholds, saliency‑weighted drift estimation, optimal approximation via blending and warping, and phase‑aware threshold scheduling across diffusion steps. Our cohesive approach enables adaptive, motion‑consistent feature reuse without retraining. On Cosmos‑Predict2.5‑2B evaluated on PAI‑Bench, WorldCache achieves 2.3× inference speedup while preserving 99.4% of baseline quality, substantially outperforming prior training‑free caching approaches. Our code can be accessed on \hrefhttps://umair1221.github.io/World‑Cache/World‑Cache.
Authors:Shivam Duggal, Xingjian Bai, Zongze Wu, Richard Zhang, Eli Shechtman, Antonio Torralba, Phillip Isola, William T. Freeman
Abstract:
Latent diffusion models (LDMs) enable high‑fidelity synthesis by operating in learned latent spaces. However, training state‑of‑the‑art LDMs requires complex staging: a tokenizer must be trained first, before the diffusion model can be trained in the frozen latent space. We propose UNITE ‑ an autoencoder architecture for unified tokenization and latent diffusion. UNITE consists of a Generative Encoder that serves as both image tokenizer and latent generator via weight sharing. Our key insight is that tokenization and generation can be viewed as the same latent inference problem under different conditioning regimes: tokenization infers latents from fully observed images, whereas generation infers them from noise together with text or class conditioning. Motivated by this, we introduce a single‑stage training procedure that jointly optimizes both tasks via two forward passes through the same Generative Encoder. The shared parameters enable gradients to jointly shape the latent space, encouraging a "common latent language". Across image and molecule modalities, UNITE achieves near state of the art performance without adversarial losses or pretrained encoders (e.g., DINO), reaching FID 2.12 and 1.73 for Base and Large models on ImageNet 256 x 256. We further analyze the Generative Encoder through the lenses of representation alignment and compression. These results show that single stage joint training of tokenization & generation from scratch is feasible.
Authors:Wooseok Jang, Seonghu Jeon, Jisang Han, Jinhyeok Choi, Minkyung Kwon, Seungryong Kim, Saining Xie, Sainan Liu
Abstract:
While recent advances in generative latent spaces have driven substantial progress in single‑image generation, the optimal latent space for novel view synthesis (NVS) remains largely unexplored. In particular, NVS requires geometrically consistent generation across viewpoints, but existing approaches typically operate in a view‑independent VAE latent space. In this paper, we propose Geometric Latent Diffusion (GLD), a framework that repurposes the geometrically consistent feature space of geometric foundation models as the latent space for multi‑view diffusion. We show that these features not only support high‑fidelity RGB reconstruction but also encode strong cross‑view geometric correspondences, providing a well‑suited latent space for NVS. Our experiments demonstrate that GLD outperforms both VAE and RAE on 2D image quality and 3D consistency metrics, while accelerating training by more than 4.4x compared to the VAE latent space. Notably, GLD remains competitive with state‑of‑the‑art methods that leverage large‑scale text‑to‑image pretraining, despite training its diffusion model from scratch without such generative pretraining.
Authors:Jeffri Murrugarra-Llerena, Pranav Chitale, Zicheng Liu, Kai Ao, Yujin Ham, Guha Balakrishnan, Paola Cascante-Bonilla
Abstract:
Social group detection, or the identification of humans involved in reciprocal interpersonal interactions (e.g., family members, friends, and customers and merchants), is a crucial component of social intelligence needed for agents transacting in the world. The few existing benchmarks for social group detection are limited by low scene diversity and reliance on third‑person camera sources (e.g., surveillance footage). Consequently, these benchmarks generally lack real‑world evaluation on how groups form and evolve in diverse cultural contexts and unconstrained settings. To address this gap, we introduce EgoGroups, a first‑person view dataset that captures social dynamics in cities around the world. EgoGroups spans 65 countries covering low, medium, and high‑crowd settings under four weather/time‑of‑day conditions. We include dense human annotations for person and social groups, along with rich geographic and scene metadata. Using this dataset, we performed an extensive evaluation of state‑of‑the‑art VLM/LLMs and supervised models on their group detection capabilities. We found several interesting findings, including VLMs and LLMs can outperform supervised baselines in a zero‑shot setting, while crowd density and cultural regions clearly influence model performance.
Authors:Daniel Shao, Joel Runevic, Richard J. Chen, Drew F. K. Williamson, Ahrong Kim, Andrew H. Song, Faisal Mahmood
Abstract:
Multiple Instance Learning (MIL) is the predominant framework for classifying gigapixel whole‑slide images in computational pathology. MIL follows a sequence of 1) extracting patch features, 2) applying a linear layer to obtain task‑specific patch features, and 3) aggregating the patches into a slide feature for classification. While substantial efforts have been devoted to optimizing patch feature extraction and aggregation, none have yet addressed the second point, the critical layer which transforms general‑purpose features into task‑specific features. We hypothesize that this layer constitutes an overlooked performance bottleneck and that stronger representations can be achieved with a low‑rank transformation tailored to each patch's phenotype, yielding synergistic effects with any of the existing MIL approaches. To this end, we introduce MAMMOTH, a parameter‑efficient, multi‑head mixture of experts module designed to improve the performance of any MIL model with minimal alterations to the total number of parameters. Across eight MIL methods and 19 different classification tasks, we find that such task‑specific transformation has a larger effect on performance than the choice of aggregation method. For instance, when equipped with MAMMOTH, even simple methods such as max or mean pooling attain higher average performance than any method with the standard linear layer. Overall, MAMMOTH improves performance in 130 of the 152 examined configurations, with an average +3.8% change in performance. Code is available at https://github.com/mahmoodlab/mammoth.
Authors:Mingju Gao, Kaisen Yang, Huan-ang Gao, Bohan Li, Ao Ding, Wenyi Li, Yangcheng Yu, Jinkun Liu, Shaocong Xu, Yike Niu, Haohan Chi, Hao Chen, Hao Tang, Yu Zhang, Li Yi, Hao Zhao
Abstract:
Hand‑object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose‑only synthesis that predicts MANO trajectories without producing pixels; (2) single‑image HOI generation that hallucinates appearance from masks or 2D cues but lacks dynamics; and (3) video generation methods that require both the entire pose sequence and the ground‑truth first frame as inputs, preventing true sim‑to‑real deployment. Inspired by the philosophy of Joo et al. (2018), we think that HOI generation requires a unified engine that brings together pose, appearance, and motion within one coherent framework. Thus we introduce PAM: a Pose‑Appearance‑Motion Engine for controllable HOI video generation. The performance of our engine is validated by: (1) On DexYCB, we obtain an FVD of 29.13 (vs. 38.83 for InterDyn), and MPJPE of 19.37 mm (vs. 30.05 mm for CosHand), while generating higher‑resolution 480x720 videos compared to 256x256 and 256x384 baselines. (2) On OAKINK2, our full multi‑condition model improves FVD from 68.76 to 46.31. (3) An ablation over input conditions on DexYCB shows that combining depth, segmentation, and keypoints consistently yields the best results. (4) For a downstream hand pose estimation task using SimpleHand, augmenting training with 3,400 synthetic videos (207k frames) allows a model trained on only 50% of the real data plus our synthetic data to match the 100% real baseline.
Authors:Junrong Guo, Shancheng Fang, Yadong Qu, Hongtao Xie
Abstract:
Recent advances in Multimodal Large Language Models (MLLMs) have enabled automated generation of structured layouts from natural language descriptions. Existing methods typically follow a code‑only paradigm that generates code to represent layouts, which are then rendered by graphic engines to produce final images. However, they are blind to the rendered visual outcome, making it difficult to guarantee readability and aesthetics. In this paper, we identify visual feedback as a critical factor in layout generation and propose Visual Feedback Layout Model (VFLM), a self‑improving framework that leverages visual feedback iterative refinement. VFLM is capable of performing adaptive reflective generation, which leverages visual information to reflect on previous issues and iteratively generates outputs until satisfactory quality is achieved. It is achieved through reinforcement learning with a visually grounded reward model that incorporates OCR accuracy. By rewarding only the final generated outcome, we can effectively stimulate the model's iterative and reflective generative capabilities. Experiments across multiple benchmarks show that VFLM consistently outperforms advanced MLLMs, existing layout models, and code‑only baselines, establishing visual feedback as critical for design‑oriented MLLMs. Our code and data are available at https://github.com/FolSpark/VFLM.
Authors:Rui Zhao, Mike Zheng Shou
Abstract:
Recent advancements in video generation models have significantly improved their ability to follow text prompts. However, the customization of dynamic visual effects, defined as temporally evolving and appearance‑driven visual phenomena like object crushing or explosion, remains underexplored. Prior works on motion customization or control mainly focus on low‑level motions of the subject or camera, which can be guided using explicit control signals such as motion trajectories. In contrast, dynamic visual effects involve higher‑level semantics that are more naturally suited for control via text prompts. However, it is hard and time‑consuming for humans to craft a single prompt that accurately specifies these effects, as they require complex temporal reasoning and iterative refinement over time. To address this challenge, we propose P‑Flow, a novel training‑free framework for customizing dynamic visual effects in video generation without modifying the underlying model. By leveraging the semantic and temporal reasoning capabilities of vision‑language models, P‑Flow performs test‑time prompt optimization, refining prompts based on the discrepancy between the visual effects of the reference video and the generated output. Through iterative refinement, the prompts evolve to better induce the desired dynamic effect in novel scenes. Experiments demonstrate that P‑Flow achieves high‑fidelity and diverse visual effect customization and outperforms other models on both text‑to‑video and image‑to‑video generation tasks. Code is available at https://github.com/showlab/P‑Flow.
Authors:Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun
Abstract:
While Vision‑Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part‑to‑whole or parent‑child structures, and often face challenges in multi‑object compositional scenarios. Hyperbolic VLMs mitigate this issue by better preserving hierarchical structures and modeling part‑whole relations (i.e., whole scene and its part images) through entailment. However, existing approaches do not model that each part has a different level of semantic representativeness to the whole. We propose UNcertainty‑guided Compositional Hyperbolic Alignment (UNCHA) for enhancing hyperbolic VLMs. UNCHA models part‑to‑whole semantic representativeness with hyperbolic uncertainty, by assigning lower uncertainty to more representative parts and higher uncertainty to less representative ones for the whole scene. This representativeness is then incorporated into the contrastive objective with uncertainty‑guided weights. Finally, the uncertainty is further calibrated with an entailment loss regularized by entropy‑based term. With the proposed losses, UNCHA learns hyperbolic embeddings with more accurate part‑whole ordering, capturing the underlying compositional structure in an image and improving its understanding of complex multi‑object scenes. UNCHA achieves state‑of‑the‑art performance on zero‑shot classification, retrieval, and multi‑label classification benchmarks. Our code and models are available at: https://github.com/jeeit17/UNCHA.git.
Authors:Jianlin Chen, Gongyang Li, Zhijiang Zhang, Liang Chang, Dan Zeng
Abstract:
Transformer‑based methods for RGB‑D Salient Object Detection (SOD) have gained significant interest, owing to the transformer's exceptional capacity to capture long‑range pixel dependencies. Nevertheless, current RGB‑D SOD methods face challenges, such as the quadratic complexity of the attention mechanism and the limited local detail extraction. To overcome these limitations, we propose a novel Superpixel Token Enhancing Network (STENet), which introduces superpixels into cross‑modal interaction. STENet follows the two‑stream encoder‑decoder structure. Its cores are two tailored superpixel‑driven cross‑modal interaction modules, responsible for global and local feature enhancement. Specifically, we update the superpixel generation method by expanding the neighborhood range of each superpixel, allowing for flexible transformation between pixels and superpixels. With the updated superpixel generation method, we first propose the Superpixel Attention Global Enhancing Module to model the global pixel‑to‑superpixel relationship rather than the traditional global pixel‑to‑pixel relationship, which can capture region‑level information and reduce computational complexity. We also propose the Superpixel Attention Local Refining Module, which leverages pixel similarity within superpixels to filter out a subset of pixels (i.e., local pixels) and then performs feature enhancement on these local pixels, thereby capturing concerned local details. Furthermore, we fuse the globally and locally enhanced features along with the cross‑scale features to achieve comprehensive feature representation. Experiments on seven RGB‑D SOD datasets reveal that our STENet achieves competitive performance compared to state‑of‑the‑art methods. The code and results of our method are available at https://github.com/Mark9010/STENet.
Authors:Nour Alhuda Albashir, Lars Pernickel, Danial Hamoud, Idriss Gouigah, Eren Erdal Aksoy
Abstract:
Autonomous vehicles face major perception and navigation challenges in adverse weather such as rain, fog, and snow, which degrade the performance of LiDAR, RADAR, and RGB camera sensors. While each sensor type offers unique strengths, such as RADAR robustness in poor visibility and LiDAR precision in clear conditions, they also suffer distinct limitations when exposed to environmental obstructions. This study proposes LRC‑WeatherNet, a novel multi‑sensor fusion framework that integrates LiDAR, RADAR, and camera data for real‑time classification of weather conditions. By employing both early fusion using a unified Bird's Eye View representation and mid‑level gated fusion of modality‑specific feature maps, our approach adapts to the varying reliability of each sensor under changing weather. Evaluated on the extensive MSU‑4S dataset covering nine weather types, LRC‑WeatherNet achieves superior classification performance and computational efficiency, significantly outperforming unimodal baselines in adverse conditions. This work is the first to combine all three modalities for robust, real‑time weather classification in autonomous driving. We release our trained models and source code in https://github.com/nouralhudaalbashir/LRC‑WeatherNet.
Authors:Swan Htet Aung, Hein Htet, Htoo Say Wah Khaing, Thuya Myo Nyunt
Abstract:
We introduce the Burmese Handwritten Digit Dataset (BHDD), a collection of 87,561 grayscale images of handwritten Burmese digits in ten classes. Each image is 28x28 pixels, following the MNIST format. The training set has 60,000 samples split evenly across classes; the test set has 27,561 samples with class frequencies as they arose during collection. Over 150 people of different ages and backgrounds contributed samples. We analyze the dataset's class distribution, pixel statistics, and morphological variation, and identify digit pairs that are easily confused due to the round shapes of the Myanmar script. Simple baselines (an MLP, a two‑layer CNN, and an improved CNN with batch normalization and augmentation) reach 99.40%, 99.75%, and 99.83% test accuracy respectively. BHDD is available under CC BY‑SA 4.0 at https://github.com/baseresearch/BHDD
Authors:Youbin Kim, Jinho Park, Hogun Park, Eunbyung Park
Abstract:
Open‑vocabulary 3D object detection aims to localize and recognize objects beyond a fixed training taxonomy. In multi‑view RGB settings, recent approaches often decouple geometry‑based instance construction from semantic labeling, generating class‑agnostic fragments and assigning open‑vocabulary categories post hoc. While flexible, such decoupling leaves instance construction governed primarily by geometric consistency, without semantic constraints during merging. When geometric evidence is view‑dependent and incomplete, this geometry‑only merging can lead to irreversible association errors, including over‑merging of distinct objects or fragmentation of a single instance. We propose Group3D, a multi‑view open‑vocabulary 3D detection framework that integrates semantic constraints directly into the instance construction process. Group3D maintains a scene‑adaptive vocabulary derived from a multimodal large language model (MLLM) and organizes it into semantic compatibility groups that encode plausible cross‑view category equivalence. These groups act as merge‑time constraints: 3D fragments are associated only when they satisfy both semantic compatibility and geometric consistency. This semantically gated merging mitigates geometry‑driven over‑merging while absorbing multi‑view category variability. Group3D supports both pose‑known and pose‑free settings, relying only on RGB observations. Experiments on ScanNet and ARKitScenes demonstrate that Group3D achieves state‑of‑the‑art performance in multi‑view open‑vocabulary 3D detection, while exhibiting strong generalization in zero‑shot scenarios. The project page is available at https://ubin108.github.io/Group3D/.
Authors:Roy Amoyal, Oren Freifeld, Chaim Baskin
Abstract:
We present Gaussian Splatting Alignment (GSA), a novel method for aligning two independent 3D Gaussian Splatting (3DGS) models via a similarity transformation (rotation, translation, and scale), even when they are of different objects in the same category (e.g., different cars). In contrast, existing methods can only align 3DGS models of the same object (e.g., the same car) and often must be given true scale as input, while we estimate it successfully. GSA leverages viewpoint‑guided spherical map features to obtain robust correspondences and introduces a two‑step optimization framework that aligns 3DGS models while keeping them fixed. First, we apply an iterative feature‑guided absolute orientation solver as our coarse registration, which is robust to poor initialization (e.g., 180 degrees misalignment or a 10x scale gap). Next, we use a fine registration step that enforces multi‑view feature consistency, inspired by inverse radiance‑field formulations. The first step already achieves state‑of‑the‑art performance, and the second further improves results. In the same‑object case, GSA outperforms prior works, often by a large margin, even when the other methods are given the true scale. In the harder case of different objects in the same category, GSA vastly surpasses them, providing the first effective solution for category‑level 3DGS registration and unlocking new applications. Project webpage: https://bgu‑cs‑vil.github.io/GSA‑project/
Authors:Clemens Watzenböck, Daniel Aletaha, Michaël Deman, Thomas Deimel, Jana Eder, Ivana Janickova, Robert Janiczek, Peter Mandl, Philipp Seeböck, Gabriela Supp, Paul Weiser, Georg Langs
Abstract:
Quantitative disease severity scoring in medical imaging is costly, time‑consuming, and subject to inter‑reader variability. At the same time, clinical archives contain far more longitudinal imaging data than expert‑annotated severity scores. Existing self‑supervised methods typically ignore this chronological structure. We introduce ChronoCon, a contrastive learning approach that replaces label‑based ranking losses with rankings derived solely from the visitation order of a patient's longitudinal scans. Under the clinically plausible assumption of monotonic progression in irreversible diseases, the method learns disease‑relevant representations without using any expert labels. This generalizes the idea of Rank‑N‑Contrast from label distances to temporal ordering. Evaluated on rheumatoid arthritis radiographs for severity assessment, the learned representations substantially improve label efficiency. In low‑label settings, ChronoCon significantly outperforms a fully supervised baseline initialized from ImageNet weights. In a few‑shot learning experiment, fine‑tuning ChronoCon on expert scores from only five patients yields an intraclass correlation coefficient of 86% for severity score prediction. These results demonstrate the potential of chronological contrastive learning to exploit routinely available imaging metadata to reduce annotation requirements in the irreversible disease domain. Code is available at https://github.com/cirmuw/ChronoCon.
Authors:Guannan Lai, Da-Wei Zhou, Zhenguo Li, Han-Jia Ye
Abstract:
Continual Test‑Time Adaptation (CTTA) aims to enable models to adapt online to unlabeled data streams under distribution shift without accessing source data. Existing CTTA methods face an efficiency‑generalization trade‑off: updating more parameters improves adaptation but severely reduces online inference efficiency. An ideal solution is to achieve comparable adaptation with minimal feature updates; we call this minimal subspace the golden subspace. We prove its existence in a single‑step adaptation setting and show that it coincides with the row space of the pretrained classifier. To enable online maintenance of this subspace, we introduce the sample‑wise Average Gradient Outer Product (AGOP) as an efficient proxy for estimating the classifier weights without retraining. Building on these insights, we propose Guided Online Low‑rank Directional adaptation (GOLD), which uses a lightweight adapter to project features onto the golden subspace and learns a compact scaling vector while the subspace is dynamically updated via AGOP. Extensive experiments on classification and segmentation benchmarks, including autonomous‑driving scenarios, demonstrate that GOLD attains superior efficiency, stability, and overall performance. Our code is available at https://github.com/AIGNLAI/GOLD.
Authors:Linkuan Zhou, Yinghao Xia, Yufei Shen, Xiangyu Li, Wenjie Du, Cong Cong, Leyi Wei, Ran Su, Qiangguo Jin
Abstract:
Unsupervised Domain Adaptation (UDA) is essential for deploying medical segmentation models across diverse clinical environments. Existing methods are fundamentally limited, suffering from semantically unaware feature alignment that results in poor distributional fidelity and from pseudo‑label validation that disregards global anatomical constraints, thus failing to prevent the formation of globally implausible structures. To address these issues, we propose SHAPE (Structure‑aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation), a framework that reframes adaptation towards global anatomical plausibility. Built on a DINOv3 foundation, its Hierarchical Feature Modulation (HFM) module first generates features with both high fidelity and class‑awareness. This shifts the core challenge to robustly validating pseudo‑labels. To augment conventional pixel‑level validation, we introduce Hypergraph Plausibility Estimation (HPE), which leverages hypergraphs to assess the global anatomical plausibility that standard graphs cannot capture. This is complemented by Structural Anomaly Pruning (SAP) to purge remaining artifacts via cross‑view stability. SHAPE significantly outperforms prior methods on cardiac and abdominal cross‑modality benchmarks, achieving state‑of‑the‑art average Dice scores of 90.08% (MRI‑>CT) and 78.51% (CT‑>MRI) on cardiac data, and 87.48% (MRI‑>CT) and 86.89% (CT‑>MRI) on abdominal data. The code is available at https://github.com/BioMedIA‑repo/SHAPE.
Authors:Donald Shenaj, Federico Errica, Antonio Carta
Abstract:
Low Rank Adaptation (LoRA) is the de facto fine‑tuning strategy to generate personalized images from pre‑trained diffusion models. Choosing a good rank is extremely critical, since it trades off performance and memory consumption, but today the decision is often left to the community's consensus, regardless of the personalized subject's complexity. The reason is evident: the cost of selecting a good rank for each LoRA component is combinatorial, so we opt for practical shortcuts such as fixing the same rank for all components. In this paper, we take a first step to overcome this challenge. Inspired by variational methods that learn an adaptive width of neural networks, we let the ranks of each layer freely adapt during fine‑tuning on a subject. We achieve it by imposing an ordering of importance on the rank's positions, effectively encouraging the creation of higher ranks when strictly needed. Qualitatively and quantitatively, our approach, LoRA^2, achieves a competitive trade‑off between DINO, CLIP‑I, and CLIP‑T across 29 subjects while requiring much less memory and lower rank than high rank LoRA versions. Code: https://github.com/donaldssh/NotAllLayersAreCreatedEqual.
Authors:Mingzhe Zheng, Weijie Kong, Yue Wu, Dengyang Jiang, Yue Ma, Xuanhua He, Bin Lin, Kaixiong Gong, Zhao Zhong, Liefeng Bo, Qifeng Chen, Harry Yang
Abstract:
Group Relative Policy Optimization (GRPO) methods for video generation like FlowGRPO remain far less reliable than their counterparts for language models and images. This gap arises because video generation has a complex solution space, and the ODE‑to‑SDE conversion used for exploration can inject excess noise, lowering rollout quality and making reward estimates less reliable, which destabilizes post‑training alignment. To address this problem, we view the pre‑trained model as defining a valid video data manifold and formulate the core problem as constraining exploration within the vicinity of this manifold, ensuring that rollout quality is preserved and reward estimates remain reliable. We propose SAGE‑GRPO (Stable Alignment via Exploration), which applies constraints at both micro and macro levels. At the micro level, we derive a precise manifold‑aware SDE with a logarithmic curvature correction and introduce a gradient norm equalizer to stabilize sampling and updates across timesteps. At the macro level, we use a dual trust region with a periodic moving anchor and stepwise constraints so that the trust region tracks checkpoints that are closer to the manifold and limits long‑horizon drift. We evaluate SAGE‑GRPO on HunyuanVideo1.5 using the original VideoAlign as the reward model and observe consistent gains over previous methods in VQ, MQ, TA, and visual metrics (CLIPScore, PickScore), demonstrating superior performance in both reward maximization and overall video quality. The code and visual gallery are available at https://dungeonmassster.github.io/SAGE‑GRPO‑Page/.
Authors:Shuxian Zhao, Jie Gui, Baosheng Yu, Dacheng Tao
Abstract:
Steel surface defect analysis is critical for industrial quality control, yet existing benchmarks rely primarily on label‑only annotations, limiting fine‑grained semantic understanding and systematic evaluation of vision‑language models. To address this gap, we introduce SteelDefectX, a vision‑language dataset with multi‑form textual annotations for steel surface defect analysis, comprising 7,778 images across 25 defect categories. At the class level, the dataset provides defect names, representative visual attributes, and industrial causes. At the sample level, each image is annotated with three forms of textual representations: (1) free‑form natural language descriptions, (2) structured attribute annotations, and (3) template‑based sentences. These annotations provide flexible textual supervision with varying levels of expressiveness and controllability. We further establish a comprehensive benchmark covering vision‑language classification, segmentation, and cross‑dataset transfer, along with additional evaluations such as retrieval and text‑guided localization. Experimental results reveal a trade‑off between structure and flexibility in textual representations. Structured attributes provide more stable semantic alignment, while natural language descriptions improve transferability and fine‑grained spatial grounding. These findings highlight the critical role of textual design in industrial vision‑language learning. SteelDefectX provides a new benchmark for studying semantic alignment and generalization in industrial vision‑language learning. The code and dataset are available at https://github.com/Zhaosxian/SteelDefectX.
Authors:Yanglin Deng, Tianyang Xu, Chunyang Cheng, Hui Li, Xiao-jun Wu, Josef Kittler
Abstract:
Infrared and visible image fusion(IVIF) combines complementary modalities while preserving natural textures and salient thermal signatures. Existing solutions predominantly rely on extensive sets of rigidly aligned image pairs for training. However, acquiring such data is often impractical due to the costly and labour‑intensive alignment process. Besides, maintaining a rigid pairing setting during training restricts the volume of cross‑modal relationships, thereby limiting generalisation performance. To this end, this work challenges the necessity of Strictly Paired Training Paradigm (SPTP) by systematically investigating UnPaired and Arbitrarily Paired Training Paradigms (UPTP and APTP) for high‑performance IVIF. We establish a theoretical objective of APTP, reflecting the complementary nature between UPTP and SPTP. More importantly, we develop a practical framework capable of significantly enriching cross‑modal relationships even with severely limited and unaligned training data. To validate our propositions, three end‑to‑end lightweight baselines, alongside a set of innovative loss functions, are designed to cover three classic frameworks (CNN, Transformer, GAN). Comprehensive experiments demonstrate that the proposed APTP and UPTP are feasible and capable of training models on a severely limited and content‑inconsistent infrared and visible dataset, achieving performance comparable to that of a dataset 100× larger in SPTP. This finding fundamentally alleviates the cost and difficulty of data collection while enhancing model robustness from the data perspective, delivering a feasible solution for IVIF studies. The code is available at \hrefhttps://github.com/yanglinDeng/IVIF_unpair\textcolorbluehttps://github.com/yanglinDeng/IVIF\_unpair.
Authors:Dillan Imans, Phuoc-Nguyen Bui, Duc-Tai Le, Hyunseung Choo
Abstract:
Retinal fundus imaging enables low‑cost and scalable hypertension (HTN) screening, but HTN‑related retinal cues are subtle, yielding high‑variance predictions. Brain MRI provides stronger vascular and small‑vessel‑disease markers of HTN, yet it is expensive and rarely acquired alongside fundus images, resulting in modality‑siloed datasets with disjoint MRI and fundus cohorts. We study this unpaired MRI‑fundus regime and introduce Clinical Graph‑Mediated Distillation (CGMD), a framework that transfers MRI‑derived HTN knowledge to a fundus model without paired multimodal data. CGMD leverages shared structured biomarkers as a bridge by constructing a clinical similarity kNN graph spanning both cohorts. We train an MRI teacher, propagate its representations over the graph, and impute brain‑informed representation targets for fundus patients. A fundus student is then trained with a joint objective combining HTN supervision, target distillation, and relational distillation. Experiments on our newly collected unpaired MRI‑fundus‑biomarker dataset show that CGMD consistently improves fundus‑based HTN prediction over standard distillation and non‑graph imputation baselines, with ablations confirming the importance of clinically grounded graph connectivity. Code is available at https://github.com/DillanImans/CGMD‑unpaired‑distillation.
Authors:Lev Ayzenberg, Shady Abu-Hussein, Raja Giryes, Hayit Greenspan
Abstract:
Full data acquisition in MRI is inherently slow, which limits clinical throughput and increases patient discomfort. Compressed Sensing MRI (CS‑MRI) seeks to accelerate acquisition by reconstructing images from under‑sampled k‑space data, requiring both an optimal sampling trajectory and a high‑fidelity reconstruction model. In this work, we propose a novel active sampling framework that leverages the inherent discrete structure of a pretrained medical image tokenizer and a latent transformer. By representing anatomy through a dictionary of quantized visual tokens, the model provides a well‑defined probability distribution over the latent space. We utilize this distribution to derive a principled uncertainty measure via token entropy, which guides the active sampling process. We introduce two strategies to exploit this latent uncertainty: (1) Latent Entropy Selection (LES), projecting patch‑wise token entropy into the k‑space domain to identify informative sampling lines, and (2) Gradient‑based Entropy Optimization (GEO), which identifies regions of maximum uncertainty reduction via the k‑space gradient of a total latent entropy loss. We evaluate our framework on the fastMRI singlecoil Knee and Brain datasets at × 8 and × 16 acceleration. Our results demonstrate that our active policies outperform state‑of‑the‑art baselines in perceptual metrics, and feature‑based distances. Our code is available at https://github.com/levayz/TRUST‑MRI.
Authors:Chen Tasker, Roy Betser, Eyal Gofer, Meir Yossef Levi, Guy Gilboa
Abstract:
Generative models and vision encoders have largely advanced on separate tracks, optimized for different goals and grounded in different mathematical principles. Yet, they share a fundamental property: latent space Gaussianity. Generative models map Gaussian noise to images, while encoders map images to semantic embeddings whose coordinates empirically behave as Gaussian. We hypothesize that both are views of a shared latent source, the Universal Normal Embedding (UNE): an approximately Gaussian latent space from which encoder embeddings and DDIM‑inverted noise arise as noisy linear projections. To test our hypothesis, we introduce NoiseZoo, a dataset of per‑image latents comprising DDIM‑inverted diffusion noise and matching encoder representations (CLIP, DINO). On CelebA, linear probes in both spaces yield strong, aligned attribute predictions, indicating that generative noise encodes meaningful semantics along linear directions. These directions further enable faithful, controllable edits (e.g., smile, gender, age) without architectural changes, where simple orthogonalization mitigates spurious entanglements. Taken together, our results provide empirical support for the UNE hypothesis and reveal a shared Gaussian‑like latent geometry that concretely links encoding and generation. Code and data are available https://rbetser.github.io/UNE/
Authors:Bingxuan Zhao, Qing Zhou, Chuang Yang, Qi Wang
Abstract:
Text‑to‑image generation powered by Diffusion Transformers (DiTs) has made remarkable strides, yet remote sensing (RS) synthesis lags behind due to two barriers: the absence of a domain‑specialized DiT prior and the prohibitive cost of training at the large resolutions that RS applications demand. Training‑free resolution promotion via Rotary Position Embedding (RoPE) rescaling offers a practical remedy, but every existing method applies a static positional scaling rule throughout the denoising process. This uniform compression is particularly harmful for RS imagery, whose substantially denser medium‑ and high‑frequency energy encodes the fine structures critical for aerial‑scene realism, such as vehicles, building contours, and road markings. Addressing both challenges requires a domain‑specialized generative prior coupled with a denoising‑aware positional adaptation strategy. To this end, we fine‑tune FLUX on over 100,000 curated RS images to build a strong domain prior (RS‑FLUX), and propose Spectrum‑aware Highly‑dynamic Adaptation for Resolution Promotion (SHARP), a training‑free method that introduces a rational fractional time schedule k_rs(t) into RoPE. SHARP applies strong positional promotion during the early layout‑formation stage and progressively relaxes it during detail recovery, aligning extrapolation strength with the frequency‑progressive nature of diffusion denoising. Its resolution‑agnostic formulation further enables robust multi‑scale generation from a single set of hyperparameters. Extensive experiments across six square and rectangular resolutions show that SHARP consistently outperforms all training‑free baselines on CLIP Score, Aesthetic Score, and HPSv2, with widening margins at more aggressive extrapolation factors and negligible computational overhead. Code and weights are available at https://github.com/bxuanz/SHARP.
Authors:Yi Wang, Haofei Zhang, Qihan Huang, Anda Cao, Gongfan Fang, Wei Wang, Xuan Jin, Jie Song, Mingli Song, Xinchao Wang
Abstract:
Large Vision‑Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Although recent token reduction methods mitigate this issue, they mainly target single‑turn Visual Question Answering (VQA), leaving the more practical multi‑turn VQA (MT‑VQA) scenario largely unexplored. MT‑VQA introduces additional challenges, as subsequent questions are unknown beforehand and may refer to arbitrary image regions, making existing reduction strategies ineffective. Specifically, current approaches fall into two categories: prompt‑dependent methods, which bias toward the initial text prompt and discard information useful for subsequent turns; prompt‑agnostic ones, which, though technically applicable to multi‑turn settings, rely on heuristic reduction metrics such as attention scores, leading to suboptimal performance. In this paper, we propose a learning‑based prompt‑agnostic method, termed MetaCompress, overcoming the limitations of heuristic designs. We begin by formulating token reduction as a learnable compression mapping, unifying existing formats such as pruning and merging into a single learning objective. Upon this formulation, we introduce a data‑efficient training paradigm capable of learning optimal compression mappings with limited computational costs. Extensive experiments on MT‑VQA benchmarks and across multiple LVLM architectures demonstrate that MetaCompress achieves superior efficiency‑accuracy trade‑offs while maintaining strong generalization across dialogue turns. Our code is available at https://github.com/MArSha1147/MetaCompress.
Authors:Yiming Shao, Qiyu Dai, Chong Gao, Guanbin Li, Yeqiang Wang, He Sun, Qiong Zeng, Baoquan Chen, Wenzheng Chen
Abstract:
Novel view synthesis (NVS) through non‑planar refractive surfaces presents fundamental challenges due to severe, spatially varying optical distortions. While recent representations like NeRF and 3D Gaussian Splatting (3DGS) excel at NVS, their assumption of straight‑line ray propagation fails under these conditions, leading to significant artifacts. To overcome this limitation, we introduce RefracGS, a framework that jointly reconstructs the refractive water surface and the scene beneath the interface. Our key insight is to explicitly decouple the refractive boundary from the target objects: the refractive surface is modeled via a neural height field, capturing wave geometry, while the underlying scene is represented as a 3D Gaussian field. We formulate a refraction‑aware Gaussian ray tracing approach that accurately computes non‑linear ray trajectories using Snell's law and efficiently renders the underlying Gaussian field while backpropagating the loss gradients to the parameterized refractive surface. Through end‑to‑end joint optimization of both representations, our method ensures high‑fidelity NVS and view‑consistent surface recovery. Experiments on both synthetic and real‑world scenes with complex waves demonstrate that RefracGS outperforms prior refractive methods in visual quality, while achieving 15x faster training and real‑time rendering at 200 FPS. The project page for RefracGS is available at https://yimgshao.github.io/refracgs/.
Authors:Wen Guo, Pengfei Zhao, Zongmeng Wang, Yufan Hu, Junyu Gao
Abstract:
Multiple Object Tracking (MOT) has long been a fundamental task in computer vision, with broad applications in various real‑world scenarios. However, due to distribution shifts in appearance, motion pattern, and catagory between the training and testing data, model performance degrades considerably during online inference in MOT. Test‑Time Adaptation (TTA) has emerged as a promising paradigm to alleviate such distribution shifts. However, existing TTA methods often fail to deliver satisfactory results in MOT, as they primarily focus solely on frame‑level adaptation while neglecting temporal consistency and identity association across frames and videos. Inspired by human decision‑making process, this paper propose a Test‑time Calibration from Experience and Intuition (TCEI) framework. In this framework, the Intuitive system utilizes transient memory to recall recently observed objects for rapid predictions, while the Experiential system leverages the accumulated experience from prior test videos to reassess and calibrate these intuitive predictions. Furthermore, both confident and uncertain objects during online testing are exploited as historical priors and reflective cases, respectively, enabling the model to adapt to the testing environment and alleviate performance degradation. Extensive experiments demonstrate that the proposed TCEI framework consistently achieves superior performance across multiple benchmark datasets and significantly enhances the model's adaptability under distribution shifts. The code will be released at https://github.com/1941Zpf/TCEI.
Authors:Jiacheng Lu, Hui Ding, Shiyu Zhang, Guoping Huo
Abstract:
Brain tumor MRI segmentation is essential for clinical diagnosis and treatment planning, enabling accurate lesion detection and radiotherapy target delineation. However, tumor lesions occupy only a small fraction of the volumetric space, resulting in severe spatial sparsity, while existing segmentation networks often overlook clinically observed spatial priors of tumor occurrence, leading to redundant feature computation over extensive background regions. To address this issue, we propose PGR‑Net (Prior‑Guided ROI Reasoning Network) ‑ an explicit ROI‑aware framework that incorporates a data‑driven spatial prior set to capture the distribution and scale characteristics of tumor lesions, providing global guidance for more stable segmentation. Leveraging these priors, PGR‑Net introduces a hierarchical Top‑K ROI decision mechanism that progressively selects the most confident lesion candidate regions across encoder layers to improve localization precision. We further develop the WinGS‑ROI (Windowed Gaussian‑Spatial Decay ROI) module, which uses multi‑window Gaussian templates with a spatial decay function to produce center‑enhanced guidance maps, thus directing feature learning throughout the network. With these ROI features, a windowed RetNet backbone is adopted to enhance localization reliability. Experiments on BraTS‑2019/2023 and MSD Task01 show that PGR‑Net consistently outperforms existing approaches while using only 8.64M Params, achieving Dice scores of 89.02%, 91.82%, and 89.67% on the Whole Tumor region. Code is available at https://github.com/CNU‑MedAI‑Lab/PGR‑Net.
Authors:Guandong Li, Zhaobin Chu
Abstract:
Inversion‑based image editing in flow matching models has emerged as a powerful paradigm for training‑free, text‑guided image manipulation. A central challenge in this paradigm is the injection dilemma: injecting source features during denoising preserves the background of the original image but simultaneously suppresses the model's ability to synthesize edited content. Existing methods address this with fixed injection strategies ‑‑ binary on/off temporal schedules, uniform spatial mixing ratios, and channel‑agnostic latent perturbation ‑‑ that ignore the inherently heterogeneous nature of injection demand across both the temporal and channel dimensions. In this paper, we present AdaEdit, a training‑free adaptive editing framework that resolves this dilemma through two complementary innovations. First, we propose a Progressive Injection Schedule that replaces hard binary cutoffs with continuous decay functions (sigmoid, cosine, or linear), enabling a smooth transition from source‑feature preservation to target‑feature generation and eliminating feature discontinuity artifacts. Second, we introduce Channel‑Selective Latent Perturbation, which estimates per‑channel importance based on the distributional gap between the inverted and random latents and applies differentiated perturbation strengths accordingly ‑‑ strongly perturbing edit‑relevant channels while preserving structure‑encoding channels. Extensive experiments on the PIE‑Bench benchmark (700 images, 10 editing types) demonstrate that AdaEdit achieves an 8.7% reduction in LPIPS, a 2.6% improvement in SSIM, and a 2.3% improvement in PSNR over strong baselines, while maintaining competitive CLIP similarity. AdaEdit is fully plug‑and‑play and compatible with multiple ODE solvers including Euler, RF‑Solver, and FireFlow. Code is available at https://github.com/leeguandong/AdaEdit
Authors:Yiwei Xie, Zheng Zhang, Ping Liu
Abstract:
Concept erasure techniques for text‑to‑video (T2V) diffusion models report substantial suppression of sensitive content, yet current evaluation is limited to checking whether the target concept is absent from generated frames, treating output‑level suppression as evidence of representational removal. We introduce PROBE, a diagnostic protocol that quantifies the reactivation potential of erased concepts in T2V models. With all model parameters frozen, PROBE optimizes a lightweight pseudo‑token embedding through a denoising reconstruction objective combined with a novel latent alignment constraint that anchors recovery to the spatiotemporal structure of the original concept. We make three contributions: (1) a multi‑level evaluation framework spanning classifier‑based detection, semantic similarity, temporal reactivation analysis, and human validation; (2) systematic experiments across three T2V architectures, three concept categories, and three erasure strategies revealing that all tested methods leave measurable residual capacity whose robustness correlates with intervention depth; and (3) the identification of temporal re‑emergence, a video‑specific failure mode where suppressed concepts progressively resurface across frames, invisible to frame‑level metrics. These findings suggest that current erasure methods achieve output‑level suppression rather than representational removal. We release our protocol to support reproducible safety auditing. Our code is available at https://github.com/YiweiXie/PRObingBasedEvaluation.
Authors:Kaiqiang Li, Gang Li, Mingle Zhou, Min Li, Delong Han, Jin Wan
Abstract:
Zero‑shot (ZS) 3D anomaly detection is crucial for reliable industrial inspection, as it enables detecting and localizing defects without requiring any target‑category training data. Existing approaches render 3D point clouds into 2D images and leverage pre‑trained Vision‑Language Models (VLMs) for anomaly detection. However, such strategies inevitably discard geometric details and exhibit limited sensitivity to local anomalies. In this paper, we revisit intrinsic 3D representations and explore the potential of pre‑trained Point‑Language Models (PLMs) for ZS 3D anomaly detection. We propose BTP (Back To Point), a novel framework that effectively aligns 3D point cloud and textual embeddings. Specifically, BTP aligns multi‑granularity patch features with textual representations for localized anomaly detection, while incorporating geometric descriptors to enhance sensitivity to structural anomalies. Furthermore, we introduce a joint representation learning strategy that leverages auxiliary point cloud data to improve robustness and enrich anomaly semantics. Extensive experiments on Real3D‑AD and Anomaly‑ShapeNet demonstrate that BTP achieves superior performance in ZS 3D anomaly detection. Code will be available at \hrefhttps://github.com/wistful‑8029/BTP‑3DADhttps://github.com/wistful‑8029/BTP‑3DAD.
Authors:Jayanie Bogahawatte, Sachith Seneviratne, Saman Halgamuge
Abstract:
Whole Slide Images (WSIs) are giga‑pixel in scale and are typically partitioned into small instances in WSI classification pipelines for computational feasibility. However, obtaining extensive instance level annotations is costly, making few‑shot weakly supervised WSI classification (FSWC) crucial for learning from limited slide‑level labels. Recently, pre‑trained vision‑language models (VLMs) have been adopted in FSWC, yet they exhibit several limitations. Existing prompt tuning methods in FSWC substantially increase both the number of trainable parameters and inference overhead. Moreover, current methods discard instances with low alignment to text embeddings from VLMs, potentially leading to information loss. To address these challenges, we propose two key contributions. First, we introduce a new parameter efficient prompt tuning method by scaling and shifting features in text encoder, which significantly reduces the computational cost. Second, to leverage not only the pre‑trained knowledge of VLMs, but also the inherent hierarchical structure of WSIs, we introduce a WSI representation learning approach with a soft hierarchical textual guidance strategy without utilizing hard instance filtering. Comprehensive evaluations on pathology datasets covering breast, lung, and ovarian cancer types demonstrate consistent improvements up‑to 10.9%, 7.8%, and 13.8% respectively, over the state‑of‑the‑art methods in FSWC. Our method reduces the number of trainable parameters by 18.1% on both breast and lung cancer datasets, and 5.8% on the ovarian cancer dataset, while also excelling at weakly‑supervised tumor localization. Code at https://github.com/Jayanie/HIPSS.
Authors:Guowei Tang, Tianwen Qian, Huanran Zheng, Yifei Wang, Xiaoling Wang
Abstract:
Real‑time, continuous understanding of visual signals is essential for real‑world interactive AI applications, and poses a fundamental system‑level challenge. Existing research on streaming video understanding, however, typically focuses on isolated aspects such as question‑answering accuracy under limited visual context or improvements in encoding efficiency, while largely overlooking practical deployability under realistic resource constraints. To bridge this gap, we introduce StreamingEval, a unified evaluation framework for assessing the streaming video understanding capabilities of Video‑LLMs under realistic constraints. StreamingEval benchmarks both mainstream offline models and recent online video models under a standardized protocol, explicitly characterizing the trade‑off between efficiency, storage and accuracy. Specifically, we adopt a fixed‑capacity memory bank to normalize accessible historical visual context, and jointly evaluate visual encoding efficiency, text decoding latency, and task performance to quantify overall system deployability. Extensive experiments across multiple datasets reveal substantial gaps between current Video‑LLMs and the requirements of realistic streaming applications, providing a systematic basis for future research in this direction. Codes will be released at https://github.com/wwgTang‑111/StreamingEval1.
Authors:Jingnan Luo, Mingqi Gao, Jun Liu, Bin-Bin Gao, Feng Zheng
Abstract:
The prosperity of Multimodal Large Language Models (MLLMs) has stimulated the demand for video reasoning segmentation, which aims to segment video objects based on human instructions. Previous studies rely on unidirectional and implicit text‑trajectory alignment, which struggles with trajectory perception when faced with severe video dynamics. In this work, we propose TrajSeg, a simple and unified framework built upon MLLMs. Concretely, we introduce bidirectional text‑trajectory alignment, where MLLMs accept grounding‑intended (text‑to‑trajectory) and captioning‑intended (trajectory‑to‑text) instructions. This way, MLLMs can benefit from enhanced correspondence and better perceive object trajectories in videos. The mask generation from trajectories is achieved via a frame‑level content integration (FCI) module and a unified mask decoder. The former adapts the MLLM‑parsed trajectory‑level token to frame‑specific information. The latter unifies segmentation for all frames into a single structure, enabling the proposed framework to be simplified and end‑to‑end trainable. Extensive experiments on referring and reasoning video segmentation datasets demonstrate the effectiveness of TrajSeg, which outperforms all video reasoning segmentation methods on all metrics. The code will be publicly available at https://github.com/haodi19/TrajSeg.
Authors:Jingchen Sun, Shaobo Han, Deep Patel, Wataru Kohno, Can Jin, Changyou Chen
Abstract:
Knowledge distillation establishes a learning paradigm that leverages both data supervision and teacher guidance. However, determining the optimal balance between learning from data and learning from the teacher is challenging, as some samples may be noisy while others are subject to teacher uncertainty. This motivates the need for adaptively balancing data and teacher supervision. We propose Beta‑weighted Knowledge Distillation (Beta‑KD), an uncertainty‑aware distillation framework that adaptively modulates how much the student relies on teacher guidance. Specifically, we formulate teacher‑‑student learning from a unified Bayesian perspective and interpret teacher supervision as a Gibbs prior over student activations. This yields a closed‑form, uncertainty‑aware weighting mechanism and supports arbitrary distillation objectives and their combinations. Extensive experiments on multimodal VQA benchmarks demonstrate that distilling student Vision‑Language Models from a large teacher VLM consistently improves performance. The results show that Beta‑KD outperforms existing knowledge distillation methods. The code is available at https://github.com/Jingchensun/beta‑kd.
Authors:Nikolay Kormushev, Josip Šarić, Matej Kristan
Abstract:
Open‑vocabulary panoptic segmentation remains hindered by two coupled issues: (i) mask selection bias, where objectness heads trained on closed vocabularies suppress masks of categories not observed in training, and (ii) limited regional understanding in vision‑language models such as CLIP, which were optimized for global image classification rather than localized segmentation. We introduce OVRCOAT, a simple, modular framework that tackles both. First, a CLIP‑conditioned objectness adjustment (COAT) updates background/foreground probabilities, preserving high‑quality masks for out‑of‑vocabulary objects. Second, an open‑vocabulary mask‑to‑text refinement (OVR) strengthens CLIP's region‑level alignment to improve classification of both seen and unseen classes with markedly lower memory cost than prior fine‑tuning schemes. The two components combine to jointly improve objectness estimation and mask recognition, yielding consistent panoptic gains. Despite its simplicity, OVRCOAT sets a new state of the art on ADE20K (+5.5% PQ) and delivers clear gains on Mapillary Vistas and Cityscapes (+7.1% and +3% PQ, respectively). The code is available at: https://github.com/nickormushev/OVRCOAT
Authors:Mohamed A Mabrok
Abstract:
We present HamVision, a framework for medical image analysis that uses the damped harmonic oscillator, a fundamental building block of signal processing, as a structured inductive bias for both segmentation and classification tasks. The oscillator's phase‑space decomposition yields three functionally distinct representations: position~q (feature content), momentum~p (spatial gradients that encode boundary and texture information), and energy H = \tfrac12|z|^2 (a parameter‑free saliency map). These representations emerge from the dynamics, not from supervision, and can be exploited by different task‑specific heads without any modification to the oscillator itself. For segmentation, energy gates the skip connections while momentum injects boundary information at every decoder level (HamSeg). For classification, the three representations are globally pooled and concatenated into a phase‑space feature vector (HamCls). We evaluate HamVision across ten medical imaging benchmarks spanning five imaging modalities. On segmentation, HamSeg achieves state‑of‑the‑art Dice scores on ISIC\,2018 (89.38%), ISIC\,2017 (88.40%), TN3K (87.05%), and ACDC (92.40%), outperforming most baselines with only 8.57M parameters. On classification, HamCls achieves state‑of‑the‑art accuracy on BloodMNIST (98.85%) and PathMNIST (96.65%), and competitive results on the remaining MedMNIST datasets against MedMamba and MedViT. Diagnostic analysis confirms that the oscillator's momentum consistently encodes an interior\,>\,boundary\,>\,exterior gradient for segmentation and that the energy map correlates with discriminative regions for classification, properties that emerge entirely from the Hamiltonian dynamics. Code is available at https://github.com/Minds‑R‑Lab/hamvision.
Authors:Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras
Abstract:
Autoregressive (AR) video diffusion has recently emerged as a promising paradigm for long video generation, enabling causal synthesis beyond the limits of bidirectional models. To address training‑inference mismatch, a series of self‑forcing strategies have been proposed to improve rollout stability by conditioning the model on its own predictions during training. While these approaches substantially mitigate exposure bias, extending generation to minute‑scale horizons remains challenging due to progressive temporal degradation. In this work, we show that this limitation is not primarily caused by insufficient memory, but by how temporal memory is utilised during inference. Through empirical analysis, we find that increasing memory does not consistently improve long‑horizon generation, and that the temporal placement of historical context significantly influences motion dynamics while leaving visual quality largely unchanged. These findings suggest that temporal memory should not be treated as a homogeneous buffer. Motivated by this insight, we introduce Relax Forcing, a structured temporal memory mechanism for AR diffusion. Instead of attending to the dense generated history, Relax Forcing decomposes temporal context into three functional roles: Sink for global stability, Tail for short‑term continuity, and dynamically selected History for structural motion guidance, and selectively incorporates only the most relevant past information. This design mitigates error accumulation during extrapolation while preserving motion evolution. Experiments on VBench‑Long demonstrate that Relax Forcing improves motion dynamics and overall temporal consistency while reducing attention overhead. Our results suggest that structured temporal memory is essential for scalable long video generation, complementing existing forcing‑based training strategies.
Authors:Yuqiu Liu, Jialin Song, Marissa Ramirez de Chanlatte, Rochishnu Chowdhury, Rushil Paresh Desai, Wuyang Chen, Daniel Martin, Michael W. Mahoney
Abstract:
Real objects that inhabit the physical world follow physical laws and thus behave plausibly during interaction with other physical objects. However, current methods that perform 3D reconstructions of real‑world scenes from multi‑view 2D images optimize primarily for visual fidelity, i.e., they train with photometric losses and reason about uncertainty in the image or representation space. This appearance‑centric view overlooks body contacts and couplings, conflates function‑critical regions (e.g., aerodynamic or hydrodynamic surfaces) with ornamentation, and reconstructs structures suboptimally, even when physical regularizers are added. All these can lead to unphysical and implausible interactions. To address this, we consider the question: How can 3D reconstruction become aware of real‑world interactions and underlying object functionality, beyond visual cues? To answer this question, we propose FluidGaussian, a plug‑and‑play method that tightly couples geometry reconstruction with ubiquitous fluid‑structure interactions to assess surface quality at high granularity. We define a simulation‑based uncertainty metric induced by fluid simulations and integrate it with active learning to prioritize views that improve both visual and physical fidelity. In an empirical evaluation on NeRF Synthetic (Blender), Mip‑NeRF 360, and DrivAerNet++, our FluidGaussian method yields up to +8.6% visual PSNR (Peak Signal‑to‑Noise Ratio) and ‑62.3% velocity divergence during fluid simulations. Our code is available at https://github.com/delta‑lab‑ai/FluidGaussian.
Authors:Idris Zakariyya, Pai Chet Ng, Kaushik Bhargav Sivangi, S. Mohammad Sheikholeslami, Konstantinos N. Plataniotis, Fani Deligianni
Abstract:
Federated video action recognition enables collaborative model training without sharing raw video data, yet remains vulnerable to two key challenges: model exposure and communication overhead. Gradients exchanged between clients and the server can leak private motion patterns, while full‑model synchronization of high‑dimensional video networks causes significant bandwidth and communication costs. To address these issues, we propose Federated Differential Privacy with Selective Tuning and Efficient Communication for Action Recognition, namely FedDP‑STECAR. Our FedDP‑STECAR framework selectively fine‑tunes and perturbs only a small subset of task‑relevant layers under Differential Privacy (DP), reducing the surface of information leakage while preserving temporal coherence in video features. By transmitting only the tuned layers during aggregation, communication traffic is reduced by over 99% compared to full‑model updates. Experiments on the UCF‑101 dataset using the MViT‑B‑16x4 transformer show that FedDP‑STECAR achieves up to 70.2% higher accuracy under strict privacy (ε=0.65) in centralized settings and 48% faster training with 73.1% accuracy in federated setups, enabling scalable and privacy‑preserving video action recognition. Code available at https://github.com/izakariyya/mvit‑federated‑videodp
Authors:Injae Kim, Chaehyeon Kim, Minseong Bae, Minseok Joo, Hyunwoo J. Kim
Abstract:
Feed‑forward 3D Gaussian Splatting methods enable single‑pass reconstruction and real‑time rendering. However, they typically adopt rigid pixel‑to‑Gaussian or voxel‑to‑Gaussian pipelines that uniformly allocate Gaussians, leading to redundant Gaussians across views. Moreover, they lack an effective mechanism to control the total number of Gaussians while maintaining reconstruction fidelity. To address these limitations, we present F4Splat, which performs Feed‑Forward predictive densification for Feed‑Forward 3D Gaussian Splatting, introducing a densification‑score‑guided allocation strategy that adaptively distributes Gaussians according to spatial complexity and multi‑view overlap. Our model predicts per‑region densification scores to estimate the required Gaussian density and allows explicit control over the final Gaussian budget without retraining. This spatially adaptive allocation reduces redundancy in simple regions and minimizes duplicate Gaussians across overlapping views, producing compact yet high‑quality 3D representations. Extensive experiments demonstrate that our model achieves superior novel‑view synthesis performance compared to prior uncalibrated feed‑forward methods, while using significantly fewer Gaussians.
Authors:Jiazhong Cen, Jiemin Fang, Sikuang Li, Guanjun Wu, Chen Yang, Taoran Yi, Zanwei Zhou, Zhikuan Bao, Lingxi Xie, Wei Shen, Qi Tian
Abstract:
High‑quality 3D assets are essential for VR/AR, industrial design, and entertainment, motivating growing interest in generative models that create 3D content from user prompts. Most existing 3D generators, however, rely on a single conditioning modality: image‑conditioned models achieve high visual fidelity by exploiting pixel‑aligned cues but suffer from viewpoint bias when the input view is limited or ambiguous, while text‑conditioned models provide broad semantic guidance yet lack low‑level visual detail. This limits how users can express intent and raises a natural question: can these two modalities be combined for more flexible and faithful 3D generation? Our diagnostic study shows that even simple late fusion of text‑ and image‑conditioned predictions outperforms single‑modality models, revealing strong cross‑modal complementarity. We therefore formalize Text‑Image Conditioned 3D Generation, which requires joint reasoning over a visual exemplar and a textual specification. To address this task, we introduce TIGON, a minimalist dual‑branch baseline with separate image‑ and text‑conditioned backbones and lightweight cross‑modal fusion. Extensive experiments show that text‑image conditioning consistently improves over single‑modality methods, highlighting complementary vision‑language guidance as a promising direction for future 3D generation research. Project page: https://jumpat.github.io/tigon‑page
Authors:Zhengxian Wu, Kai Shi, Chuanrui Zhang, Zirui Liao, Jun Yang, Ni Yang, Qiuying Peng, Luyuan Zhang, Hangrui Xu, Tianhuang Su, Zhenyu Yang, Haonan Lu, Haoqian Wang
Abstract:
Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high‑quality annotated data or teacher‑model distillation, both of which are costly and difficult to scale. To address this, we propose an unsupervised self‑evolution training framework for multimodal reasoning that achieves stable performance improvements without using human‑annotated answers or external reward models. For each input, we sample multiple reasoning trajectories and jointly model their within group structure. We use the Actor's self‑consistency signal as a training prior, and introduce a bounded Judge based modulation to continuously reweight trajectories of different quality. We further model the modulated scores as a group level distribution and convert absolute scores into relative advantages within each group, enabling more robust policy updates. Trained with Group Relative Policy Optimization (GRPO) on unlabeled data, our method consistently improves reasoning performance and generalization on five mathematical reasoning benchmarks, offering a scalable path toward self‑evolving multimodal models. The code are available at https://github.com/OPPO‑Mente‑Lab/LLM‑Self‑Judge.
Authors:Yuntian Bo, Yazhou Zhu, Piotr Koniusz, Haofeng Zhang
Abstract:
Conventional few‑shot medical image segmentation (FSMIS) approaches face performance bottlenecks that hinder broader clinical applicability. Although the Segment Anything Model (SAM) exhibits strong category‑agnostic segmentation capabilities, its direct application to medical images often leads to over‑segmentation due to ambiguous anatomical boundaries. In this paper, we reformulate SAM‑based FSMIS as a prompt localization task and propose FoB (Focus on Background), a background‑centric prompt generator that provides accurate background prompts to constrain SAM's over‑segmentation. Specifically, FoB bridges the gap between segmentation and prompt localization by category‑agnostic generation of support background prompts and localizing them directly in the query image. To address the challenge of prompt localization for novel categories, FoB models rich contextual information to capture foreground‑background spatial dependencies. Moreover, inspired by the inherent structural patterns of background prompts in medical images, FoB models this structure as a constraint to progressively refine background prompt predictions. Experiments on three diverse medical image datasets demonstrate that FoB outperforms other baselines by large margins, achieving state‑of‑the‑art performance on FSMIS, and exhibiting strong cross‑domain generalization. Our code is available at https://github.com/primebo1/FoB_SAM.
Authors:Osamu Hirose, Emanuele Rodola
Abstract:
Nonrigid registration is conventionally divided into point set registration, which aligns sparse geometries, and image registration, which aligns continuous intensity fields on regular grids. However, this dichotomy creates a critical bottleneck for emerging scientific data, such as spatial transcriptomics, where high‑dimensional vector‑valued functions, e.g., gene expression, are defined on irregular, sparse manifolds. Consequently, researchers currently face a forced choice: either sacrifice single‑cell resolution via voxelization to utilize image‑based tools, or ignore the critical functional signal to utilize geometric tools. To resolve this dilemma, we propose Domain Elastic Transform (DET), a grid‑free probabilistic framework that unifies geometric and functional alignment. By treating data as functions on irregular domains, DET registers high‑dimensional signals directly without binning. We formulate the problem within a rigorous Bayesian framework, modeling domain deformation as an elastic motion guided by a joint spatial‑functional likelihood. The method is fully unsupervised and scalable, utilizing feature‑sensitive downsampling to handle massive atlases. We demonstrate that DET achieves 92% topological preservation on MERFISH data where state‑of‑the‑art optimal transport methods struggle (<5%), and successfully registers whole‑embryo Stereo‑seq atlases across developmental stages ‑‑ a task involving massive scale and complex nonrigid growth. The implementation of DET is available on https://github.com/ohirose/bcpd (since Mar, 2025).
Authors:Jinyu Xu, Tianqi Hu, Xiaonan Hu, Letian Zhou, Songliang Cao, Meng Zhang, Hao Lu
Abstract:
Visually cataloging and quantifying the natural world requires pushing the boundaries of both detailed visual classification and counting at scale. Despite significant progress, particularly in crowd and traffic analysis, the fine‑grained, taxonomy‑aware plant counting remains underexplored in vision. In contrast to crowds, plants exhibit nonrigid morphologies and physical appearance variations across growth stages and environments. To fill this gap, we present TPC‑268, the first plant counting benchmark incorporating plant taxonomy. Our dataset couples instance‑level point annotations with Linnaean labels (kingdom ‑> species) and organ categories, enabling hierarchical reasoning and species‑aware evaluation. The dataset features 10,000 images with 678,050 point annotations, includes 268 countable plant categories over 242 plant species in Plantae and Fungi, and spans observation scales from canopy‑level remote sensing imagery to tissue‑level microscopy. We follow the problem setting of class‑agnostic counting (CAC), provide taxonomy‑consistent, scale‑aware data splits, and benchmark state‑of‑the‑art regression‑ and detection‑based CAC approaches. By capturing the biodiversity, hierarchical structure, and multi‑scale nature of botanical and mycological taxa, TPC‑268 provides a biologically grounded testbed to advance fine‑grained class‑agnostic counting. Dataset and code are available at https://github.com/tiny‑smart/TPC‑268.
Authors:Shenghan Chen, Yiming Liu, Yanzhen Wang, Yujia Wang, Xiankai Lu
Abstract:
Balancing performance trade‑off on long‑tail (LT) data distributions remains a long‑standing challenge. In this paper, we posit that this dilemma stems from a phenomenon called "tail performance degradation" (the model tends to severely overfit on head classes while quickly forgetting tail classes) and pose a solution from a loss landscape perspective. We observe that different classes possess divergent convergence points in the loss landscape. Besides, this divergence is aggravated when the model settles into sharp and non‑robust minima, rather than a shared and flat solution that is beneficial for all classes. In light of this, we propose a continual learning inspired framework to prevent "tail performance degradation". To avoid inefficient per‑class parameter preservation, a Grouped Knowledge Preservation module is proposed to memorize group‑specific convergence parameters, promoting convergence towards a shared solution. Concurrently, our framework integrates a Grouped Sharpness Aware module to seek flatter minima by explicitly addressing the geometry of the loss landscape. Notably, our framework requires neither external training samples nor pre‑trained models, facilitating the broad applicability. Extensive experiments on four benchmarks demonstrate significant performance gains over state‑of‑the‑art methods. The code is available at:https://gkp‑gsa.github.io/.
Authors:Thomas Mendelson, Joshua Francois, Galit Lahav, Tammy Riklin-Raviv
Abstract:
Accurate delineation of individual cells in microscopy videos is essential for studying cellular dynamics, yet separating touching or overlapping instances remains a persistent challenge. Although foundation‑model for segmentation such as SAM have broadened the accessibility of image segmentation, they still struggle to separate nearby cell instances in dense microscopy scenes without extensive prompting.
We propose a prompt‑free, boundary‑aware instance segmentation framework that predicts signed distance functions (SDFs) instead of binary masks, enabling smooth and geometry‑consistent modeling of cell contours. A learned sigmoid mapping converts SDFs into probability maps, yielding sharp boundary localization and robust separation of adjacent instances. Training is guided by a unified Modified Hausdorff Distance (MHD) loss that integrates region‑ and boundary‑based terms.
Evaluations on both public and private high‑throughput microscopy datasets demonstrate improved boundary accuracy and instance‑level performance compared to recent SAM‑based and foundation‑model approaches. Source code is available at: https://github.com/ThomasMendelson/BAISeg.git
Authors:Jiatong Xia, Lingqiao Liu
Abstract:
We introduce a novel, training‑free system for reconstructing, understanding, and rendering 3D indoor scenes from a sparse set of unposed RGB images. Unlike traditional radiance field approaches that require dense views and per‑scene optimization, our pipeline achieves high‑fidelity results without any training or pose preprocessing. The system integrates three key innovations: (1) A robust point cloud reconstruction module that filters unreliable geometry using a warping‑based anomaly removal strategy; (2) A warping‑guided 2D‑to‑3D instance lifting mechanism that propagates 2D segmentation masks into a consistent, instance‑aware 3D representation; and (3) A novel rendering approach that projects the point cloud into new views and refines the renderings with a 3D‑aware diffusion model. Our method leverages the generative power of diffusion to compensate for missing geometry and enhances realism, especially under sparse input conditions. We further demonstrate that object‑level scene editing such as instance removal can be naturally supported in our pipeline by modifying only the point cloud, enabling the synthesis of consistent, edited views without retraining. Our results establish a new direction for efficient, editable 3D content generation without relying on scene‑specific optimization. Project page: https://jiatongxia.github.io/TID3R/
Authors:Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Rubaya Tabassum, Ariful Ekraj Hridoy, Mehraj Mahmood, Mahbub E Sobhani, Md. Tarek Hasan, Swakkhar Shatabda
Abstract:
Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded benchmark for evaluating multilingual vision‑language models (VLMs) on Bengali culture across historically linked languages and regional dialects. Built from 1,152 manually curated images across nine domains, the benchmark supports visual question answering and captioning, and is expanded into four languages and five Bangla dialects, yielding ~32.2K artifacts. Our experiments show that evaluating only standard Bangla overestimates true model capability: performance drops under dialectal variation, especially for caption generation, while historically linked languages such as Hindi and Urdu retain some cultural meaning but remain weaker for structured reasoning. Across domains, the main bottleneck is missing cultural knowledge rather than visual grounding alone, with knowledge‑intensive categories. These findings position BanglaVerse as a more realistic test bed for measuring culturally grounded multimodal understanding under linguistic variation.
Authors:Bo Li, Tingting Bao, Lingling Zhang, Weiping Fu, Yaxian Wang, Jun Liu
Abstract:
Diffusion models have achieved impressive performance on multi‑focus image fusion (MFIF). However, a key challenge in applying diffusion models to the ill‑posed MFIF problem is that defocus blur can make common symmetric geometric structures (e.g., textures and edges) appear warped and deformed, often leading to unexpected artifacts in the fused images. Therefore, embedding rotation equivariance into diffusion networks is essential, as it enables the fusion results to faithfully preserve the original orientation and structural consistency of geometric patterns underlying the input images. Motivated by this, we propose ReDiffuse, a rotation‑equivariant diffusion model for MFIF. Specifically, we carefully construct the basic diffusion architectures to achieve end‑to‑end rotation equivariance. We also provide a rigorous theoretical analysis to evaluate its intrinsic equivariance error, demonstrating the validity of embedding equivariance structures. ReDiffuse is comprehensively evaluated against various MFIF methods across four datasets (Lytro, MFFW, MFI‑WHU, and Road‑MF). Results demonstrate that ReDiffuse achieves competitive performance, with improvements of 0.28‑6.64% across six evaluation metrics. The code is available at https://github.com/MorvanLi/ReDiffuse.
Authors:Xiaoshan Wu, Xiaoyang Lyu, Yifei Yu, Bo Wang, Zhongrui Wang, Xiaojuan Qi
Abstract:
Dense semantic segmentation in dynamic environments is fundamentally limited by the low‑frame‑rate (LFR) nature of standard cameras, which creates critical perceptual gaps between frames. To solve this, we introduce Anytime Interframe Semantic Segmentation: a new task for predicting segmentation at any arbitrary time using only a single past RGB frame and a stream of asynchronous event data. This task presents a core challenge: how to robustly propagate dense semantic features using a motion field derived from sparse and often noisy event data, all while mitigating feature degradation in highly dynamic scenes. We propose LiFR‑Seg, a novel framework that directly addresses these challenges by propagating deep semantic features through time. The core of our method is an uncertainty‑aware warping process, guided by an event‑driven motion field and its learned, explicit confidence. A temporal memory attention module further ensures coherence in dynamic scenarios. We validate our method on the DSEC dataset and a new high‑frequency synthetic benchmark (SHF‑DSEC) we contribute. Remarkably, our LFR system achieves performance (73.82% mIoU on DSEC) that is statistically indistinguishable from an HFR upper‑bound (within 0.09%) that has full access to the target frame. This work presents a new, efficient paradigm for achieving robust, high‑frame‑rate perception with low‑frame‑rate hardware. Project Page: https://candy‑crusher.github.io/LiFR_Seg_Proj/#; Code: https://github.com/Candy‑Crusher/LiFR‑Seg.git.
Authors:Shanmukha Vellamcheti, Uday Kiran Kothapalli, Disharee Bhowmick, Sathyanarayanan N. Aakur
Abstract:
Multimodal large language models (MLLMs) achieve strong performance on single‑view spatial reasoning tasks, yet it remains unclear whether they maintain stable spatial state representations under counterfactual viewpoint changes. We introduce a controlled diagnostic benchmark that evaluates relational consistency under hypothetical camera orbit transformations without re‑rendering images. Across 100 synthetic scenes and 6,000 relational queries, we measure viewpoint consistency, 360° cycle agreement, and relational stability over sequential transformations. Despite high single‑view accuracy, state‑of‑the‑art MLLMs exhibit systematic degradation under counterfactual viewpoint changes, with frequent violations of cycle consistency and rapid decay in relational stability. We further evaluate multiple input representations, visual input, textual bounding boxes, and structured scene graphs, and show that increasing representational structure improves stability. Our results suggest that single‑view spatial accuracy overestimates the robustness of induced spatial representations and that representation structure plays a critical role in counterfactual spatial reasoning.
Authors:Shih-Wen Liu, Yen-Chang Chen, Wei-Ta Chu, Fu-En Yang, Yu-Chiang Frank Wang
Abstract:
Multi‑task learning (MTL) aims to enable a single model to solve multiple tasks efficiently; however, current parameter‑efficient fine‑tuning (PEFT) methods remain largely limited to single‑task adaptation. We introduce Free Sinewich, a parameter‑efficient multi‑task learning framework that enables near‑zero‑cost weight modulation via frequency switching (Free). Specifically, a Sine‑AWB (Sinewich) layer combines low‑rank factors and convolutional priors into a single kernel, which is then modulated elementwise by a sinusoidal transformation to produce task‑specialized weights. A lightweight Clock Net is introduced to produce bounded frequencies that stabilize this modulation during training. Theoretically, sine modulation enhances the rank of low‑rank adapters, while frequency separation decorrelates the weights of different tasks. On dense prediction benchmarks, Free Sinewich achieves state‑of‑the‑art performance‑efficiency trade‑offs (e.g., up to +5.39% improvement over single‑task fine‑tuning with only 6.53M trainable parameters), offering a compact and scalable paradigm based on frequency‑based parameter sharing. Project page: \hrefhttps://casperliuliuliu.github.io/projects/Free‑Sinewich/https://casperliuliuliu.github.io/projects/Free‑Sinewich.
Authors:He Wang, Tianyang Xu, Zhangyong Tang, Xiao-Jun Wu, Josef Kittler
Abstract:
Due to the limited availability of paired multi‑modal data, multi‑modal trackers are typically built by adopting pre‑trained RGB models with parameter‑efficient fine‑tuning modules. However, these fine‑tuning methods overlook advanced adaptations for applying RGB pre‑trained models and fail to modulate a single specific modality, cross‑modal interactions, and the prediction head. To address the issues, we propose to perform Progressive Adaptation for Multi‑Modal Tracking (PATrack). This innovative approach incorporates modality‑dependent, modality‑entangled, and task‑level adapters, effectively bridging the gap in adapting RGB pre‑trained networks to multi‑modal data through a progressive strategy. Specifically, modality‑specific information is enhanced through the modality‑dependent adapter, decomposing the high‑ and low‑frequency components, which ensures a more robust feature representation within each modality. The inter‑modal interactions are introduced in the modality‑entangled adapter, which implements a cross‑attention operation guided by inter‑modal shared information, ensuring the reliability of features conveyed between modalities. Additionally, recognising that the strong inductive bias of the prediction head does not adapt to the fused information, a task‑level adapter specific to the prediction head is introduced. In summary, our design integrates intra‑modal, inter‑modal, and task‑level adapters into a unified framework. Extensive experiments on RGB+Thermal, RGB+Depth, and RGB+Event tracking tasks demonstrate that our method shows impressive performance against state‑of‑the‑art methods. Code is available at https://github.com/ouha1998/Learning‑Progressive‑Adaptation‑for‑Multi‑Modal‑Tracking.
Authors:Ping Guo, Chengzhou Li, Guanchen Meng, Qi Jia, Jinyuan Liu, Zhu Liu, Yu Liu, Zhongxuan Luo, Xin Fan
Abstract:
As one of the most important underwater sensing technologies, forward‑looking sonar exhibits unique imaging characteristics. Sonar images are often affected by severe speckle noise, low texture contrast, acoustic shadows, and geometric distortions. These factors make it difficult for traditional teacher‑student frameworks to achieve satisfactory performance in sonar semantic segmentation tasks under extremely limited labeled data conditions. To address this issue, we propose a Collaborative Teacher Semantic Segmentation Framework for forward‑looking sonar images. This framework introduces a multi‑teacher collaborative mechanism composed of one general teacher and multiple sonar‑specific teachers. By adopting a multi‑teacher alternating guidance strategy, the student model can learn general semantic representations while simultaneously capturing the unique characteristics of sonar images, thereby achieving more comprehensive and robust feature modeling. Considering the challenges of sonar images, which can lead teachers to generate a large number of noisy pseudo‑labels, we further design a cross‑teacher reliability assessment mechanism. This mechanism dynamically quantifies the reliability of pseudo‑labels by evaluating the consistency and stability of predictions across multiple views and multiple teachers, thereby mitigating the negative impact caused by noisy pseudo‑labels. Notably, on the FLSMD dataset, when only 2% of the data is labeled, our method achieves a 5.08% improvement in mIoU compared to other state‑of‑the‑art approaches.
Authors:Hwasik Jeong, Seungryong Lee, Gyeongjin Kang, Seungkwon Yang, Xiangyu Sun, Seungtae Nam, Eunbyung Park
Abstract:
Pose‑free feed‑forward 3D Gaussian Splatting (3DGS) has opened a new frontier for rapid 3D modeling, enabling high‑quality Gaussian representations to be generated from uncalibrated multi‑view images in a single forward pass. The dominant approach in this space adopts unified monolithic architectures, often built on geometry‑centric 3D foundation models, to jointly estimate camera poses and synthesize 3DGS representations within a single network. While architecturally streamlined, such "all‑in‑one" designs may be suboptimal for high‑fidelity 3DGS generation, as they entangle geometric reasoning and appearance modeling within a shared representation. In this work, we introduce 2Xplat, a pose‑free feed‑forward 3DGS framework based on a two‑expert design that explicitly separates geometry estimation from Gaussian generation. A dedicated geometry expert first predicts camera poses, which are then explicitly passed to a powerful appearance expert that synthesizes 3D Gaussians. Despite its conceptual simplicity, being largely underexplored in prior works, the proposed approach proves highly effective. In fewer than 5K training iterations, the proposed two‑experts pipeline substantially outperforms prior pose‑free feed‑forward 3DGS approaches and achieves performance on par with state‑of‑the‑art posed methods. These results challenge the prevailing unified paradigm and suggest the potential advantages of modular design principles for complex 3D geometric estimation and appearance synthesis tasks.
Authors:Pengchong Hu, Zhizhong Han
Abstract:
3D Gaussian Splatting (3DGS) has made remarkable progress in RGBD SLAM. Current methods usually use 3D Gaussians or view‑tied 3D Gaussians to represent radiance fields in tracking and mapping. However, these Gaussians are either too flexible or too limited in movements, resulting in slow convergence or limited rendering quality. To resolve this issue, we adopt pixel‑aligned Gaussians but allow each Gaussian to adjust its position along its ray to maximize the rendering quality, even if Gaussians are simplified to improve system scalability. To speed up the tracking, we model the depth distribution around each pixel as a Gaussian distribution, and then use these distributions to align each frame to the 3D scene quickly. We report our evaluations on widely used benchmarks, justify our designs, and show advantages over the latest methods in view rendering, camera tracking, runtime, and storage complexity. Please see our project page for code and videos at https://machineperceptionlab.github.io/SGAD‑SLAM‑Project .
Authors:Shuwei Huang, Shizhuo Liu, Zijun Wei
Abstract:
Diffusion‑based image super‑resolution (SR) aims to reconstruct high‑resolution (HR) images from low‑resolution (LR) observations. However, the inherent randomness injected during the reverse diffusion process causes the performance of diffusion‑based SR models to vary significantly across different sampling runs, particularly when the sampling trajectory is compressed into a limited number of steps. A critical yet underexplored question is: what is the optimal noise to inject at each intermediate diffusion step? In this paper, we establish a theoretical framework that derives the closed‑form analytical solution for optimal intermediate noise in diffusion models from a maximum likelihood estimation perspective, revealing a consistent conditional dependence structure that generalizes across diffusion paradigms. We instantiate this framework under the residual‑shifting diffusion paradigm and accordingly design an LR‑guided multi‑input‑aware noise predictor to replace random Gaussian noise. We further mitigate initialization bias with a high‑quality pre‑upsampling network. The compact 4‑step trajectory uniquely enables end‑to‑end optimization of the entire reverse chain, which is computationally prohibitive for conventional long‑trajectory diffusion models. Extensive experiments demonstrate that LPNSR achieves state‑of‑the‑art perceptual performance on both synthetic and real‑world datasets, without relying on any large‑scale text‑to‑image priors. The source code of our method can be found at https://github.com/Faze‑Hsw/LPNSR.
Authors:Uzair Shah, Marco Agus, Mahmoud Gamal, Mahmood Alzubaidi, Corrado Cali, Pierre J. Magistretti, Abdesselam Bouzerdoum, Mowafa Househ
Abstract:
Neuronal morphology encodes critical information about circuit function, development, and disease, yet current methods analyze topology or graph structure in isolation. We introduce GraPHFormer, a multimodal architecture that unifies these complementary views through CLIP‑style contrastive learning.
Our vision branch processes a novel three‑channel persistence image encoding unweighted, persistence‑weighted, and radius‑weighted topological densities via DINOv2‑ViT‑S. In parallel, a TreeLSTM encoder captures geometric and radial attributes from skeleton graphs. Both project to a shared embedding space trained with symmetric InfoNCE loss, augmented by persistence‑space transformations that preserve topological semantics.
Evaluated on six benchmarks (BIL‑6, ACT‑4, JML‑4, N7, M1‑Cell, M1‑REG) spanning self‑supervised and supervised settings, GraPHFormer achieves state‑of‑the‑art performance on five benchmarks, significantly outperforming topology‑only, graph‑only, and morphometrics baselines. We demonstrate practical utility by discriminating glial morphologies across cortical regions and species, and detecting signatures of developmental and degenerative processes.
Code: https://github.com/Uzshah/GraPHFormer
Authors:Xu Zhang, Jin Yuan, BinHong Yang, Xuan Liu, Qianjun Zhang, Yuyi Wang, Zhiyong Li, Hanwang Zhang
Abstract:
Recent advancements in multimodal large models have significantly bridged the representation gap between diverse modalities, catalyzing the evolution of video multimodal interpretation, which enhances users' understanding of video content by generating correlated modalities. However, most existing video multimodal interpretation methods primarily concentrate on global comprehension with limited user interaction. To address this, we propose a novel task, Controllable Video Segmentation and Captioning (SegCaptioning), which empowers users to provide specific prompts, such as a bounding box around an object of interest, to simultaneously generate correlated masks and captions that precisely embody user intent. An innovative framework Scene Graph‑guided Fine‑grained SegCaptioning Transformer (SG‑FSCFormer) is designed that integrates a Prompt‑guided Temporal Graph Former to effectively captures and represents user intent through an adaptive prompt adaptor, ensuring that the generated content well aligns with the user's requirements. Furthermore, our model introduces a Fine‑grained Mask‑linguistic Decoder to collaboratively predict high‑quality caption‑mask pairs using a Multi‑entity Contrastive loss, as well as provide fine‑grained alignment between each mask and its corresponding caption tokens, thereby enhancing users' comprehension of videos. Comprehensive experiments conducted on two benchmark datasets demonstrate that SG‑FSCFormer achieves remarkable performance, effectively capturing user intent and generating precise multimodal outputs tailored to user specifications. Our code is available at https://github.com/XuZhang1211/SG‑FSCFormer.
Authors:Xiefan Guo, Xinzhu Ma, Haoxiang Ma, Zihao Zhou, Di Huang
Abstract:
Text‑to‑image diffusion models have achieved remarkable fidelity in synthesizing images from explicit text prompts, yet exhibit a critical deficiency in processing implicit prompts that require deep‑level world knowledge, ranging from natural sciences to cultural commonsense, resulting in counter‑factual synthesis. This paper traces the root of this limitation to a fundamental dislocation of the underlying knowledge structures, manifesting as a chaotic organization of implicit prompts compared to their explicit counterparts. In this paper, we propose EruDiff, which aims to refactor the knowledge within diffusion models. Specifically, we develop the Diffusion Knowledge Distribution Matching (DK‑DM) to register the knowledge distribution of intractable implicit prompts with that of well‑defined explicit anchors. Furthermore, to rectify the inherent biases in explicit prompt rendering, we employ the Negative‑Only Reinforcement Learning (NO‑RL) strategy for fine‑grained correction. Rigorous empirical evaluations demonstrate that our method significantly enhances the performance of leading diffusion models, including FLUX and Qwen‑Image, across both the scientific knowledge benchmark (i.e., Science‑T2I) and the world knowledge benchmark (i.e., WISE), underscoring the effectiveness and generalizability. Our code is available at https://github.com/xiefan‑guo/erudiff.
Authors:Hanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan Shen
Abstract:
While structure‑based relocalizers have long strived for point correspondences when establishing or regressing query‑map associations, in this paper, we pioneer the use of planar primitives and 3D planar maps for lightweight 6‑DoF camera relocalization in structured environments. Planar primitives, beyond being fundamental entities in projective geometry, also serve as region‑based representations that encapsulate both structural and semantic richness. This motivates us to introduce PlanaReLoc, a streamlined plane‑centric paradigm where a deep matcher associates planar primitives across the query image and the map within a learned unified embedding space, after which the 6‑DoF pose is solved and refined under a robust framework. Through comprehensive experiments on the ScanNet and 12Scenes datasets across hundreds of scenes, our method demonstrates the superiority of planar primitives in facilitating reliable cross‑modal structural correspondences and achieving effective camera relocalization without requiring realistically textured/colored maps, pose priors, or per‑scene training. The code and data are available at https://github.com/3dv‑casia/PlanaReLoc .
Authors:Chenxing Meng, Wuzhou Quan, Yingjie Cai, Liqun Cao, Liyan Zhang, Mingqiang Wei
Abstract:
Cloud occlusion severely degrades the semantic integrity of optical remote sensing imagery. While incorporating Synthetic Aperture Radar (SAR) provides complementary observations, achieving efficient global modeling and reliable cross‑modal fusion under cloud interference remains challenging. Existing methods rely on dense global attention to capture long‑range dependencies, yet such aggregation indiscriminately propagates cloud‑induced noise. Improving robustness typically entails enlarging model capacity, which further increases computational overhead. Given the large‑scale and high‑resolution nature of remote sensing applications, such computational demands hinder practical deployment, leading to an efficiency‑reliability trade‑off. To address this dilemma, we propose EDC, an efficiency‑oriented and discrepancy‑conditioned optical‑SAR semantic segmentation framework. A tri‑stream encoder with Carrier Tokens enables compact global context modeling with reduced complexity. To prevent noise contamination, we introduce a Discrepancy‑Conditioned Hybrid Fusion (DCHF) mechanism that selectively suppresses unreliable regions during global aggregation. In addition, an auxiliary cloud removal branch with teacher‑guided distillation enhances semantic consistency under occlusion. Extensive experiments demonstrate that EDC achieves superior accuracy and efficiency, improving mIoU by 0.56% and 0.88% on M3M‑CR and WHU‑OPT‑SAR, respectively, while reducing the number of parameters by 46.7% and accelerating inference by 1.98×. Our implementation is available at https://github.com/mengcx0209/EDC.
Authors:Xiaoya Cheng, Long Wang, Yan Liu, Xinyi Liu, Hanlin Tan, Yu Liu, Maojun Zhang, Shen Yan
Abstract:
We present PiLoT, a unified framework that tackles UAV‑based ego and target geo‑localization. Conventional approaches rely on decoupled pipelines that fuse GNSS and Visual‑Inertial Odometry (VIO) for ego‑pose estimation, and active sensors like laser rangefinders for target localization. However, these methods are susceptible to failure in GNSS‑denied environments and incur substantial hardware costs and complexity. PiLoT breaks this paradigm by directly registering live video stream against a geo‑referenced 3D map. To achieve robust, accurate, and real‑time performance, we introduce three key contributions: 1) a Dual‑Thread Engine that decouples map rendering from core localization thread, ensuring both low latency while maintaining drift‑free accuracy; 2) a large‑scale synthetic dataset with precise geometric annotations (camera pose, depth maps). This dataset enables the training of a lightweight network that generalizes in a zero‑shot manner from simulation to real data; and 3) a Joint Neural‑Guided Stochastic‑Gradient Optimizer (JNGO) that achieves robust convergence even under aggressive motion. Evaluations on a comprehensive set of public and newly collected benchmarks show that PiLoT outperforms state‑of‑the‑art methods while running over 25 FPS on NVIDIA Jetson Orin platform. Our code and dataset is available at: https://github.com/Choyaa/PiLoT.
Authors:Xiefan Guo, Xinzhu Ma, Haiyu Zhang, Di Huang
Abstract:
Recent advancements in text‑to‑image synthesis have been largely propelled by diffusion‑based models, yet achieving precise alignment between text prompts and generated images remains a persistent challenge. We find that this difficulty arises primarily from the limitations of conventional diffusion loss, which provides only implicit supervision for modeling fine‑grained text‑image correspondence. In this paper, we introduce Cross‑Timestep Self‑Calibration (CTCal), founded on the supporting observation that establishing accurate text‑image alignment within diffusion models becomes progressively more difficult as the timestep increases. CTCal leverages the reliable text‑image alignment (i.e., cross‑attention maps) formed at smaller timesteps with less noise to calibrate the representation learning at larger timesteps with more noise, thereby providing explicit supervision during training. We further propose a timestep‑aware adaptive weighting to achieve a harmonious integration of CTCal and diffusion loss. CTCal is model‑agnostic and can be seamlessly integrated into existing text‑to‑image diffusion models, encompassing both diffusion‑based (e.g., SD 2.1) and flow‑based approaches (e.g., SD 3). Extensive experiments on T2I‑Compbench++ and GenEval benchmarks demonstrate the effectiveness and generalizability of the proposed CTCal. Our code is available at https://github.com/xiefan‑guo/ctcal.
Authors:Qunjie Huang, Weina Zhu
Abstract:
Cross‑subject EEG‑to‑image retrieval for visual decoding is challenged by subject shift and hubness in the embedding space, which distort similarity geometry and destabilize top‑k rankings, making small‑k shortlists unreliable. We introduce SATTC (Structure‑Aware Test‑Time Calibration), a label‑free calibration head that operates directly on the similarity matrix of frozen EEG and image encoders. SATTC combines a geometric expert, subject‑adaptive whitening of EEG embeddings with an adaptive variant of Cross‑domain Similarity Local Scaling (CSLS), and a structural expert built from mutual nearest neighbors, bidirectional top‑k ranks, and class popularity, fused via a simple Product‑of‑Experts rule. On THINGS‑EEG2 under a strict leave‑one‑subject‑out protocol, standardized inference with cosine similarities, L2‑normalized embeddings, and candidate whitening already yields a strong cross‑subject baseline over the original ATM retrieval setup. Building on this baseline, SATTC further improves Top‑1 and Top‑5 accuracy, reduces hubness and per‑class imbalance, and produces more reliable small‑k shortlists. These gains transfer across multiple EEG encoders, supporting SATTC as an encoder‑agnostic, label‑free test‑time calibration layer for cross‑subject neural decoding.
Authors:Ivan Desiatov, Torsten Sattler
Abstract:
3D Gaussian Splatting (3DGS) has become the method of choice for photo‑realistic 3D reconstruction of scenes, due to being able to efficiently and accurately recover the scene appearance and geometry from images. 3DGS represents the scene through a set of 3D Gaussians, parameterized by their position, spatial extent, and view‑dependent color. Starting from an initial point cloud, 3DGS refines the Gaussians' parameters as to reconstruct a set of training images as accurately as possible. Typically, a sparse Structure‑from‑Motion point cloud is used as initialization. In order to obtain dense Gaussian clouds, 3DGS methods thus rely on a densification stage. In this paper, we systematically study the relation between densification and initialization. Proposing a new benchmark, we study combinations of different types of initializations (dense laser scans, dense (multi‑view) stereo point clouds, dense monocular depth estimates, sparse SfM point clouds) and different densification schemes. We show that current densification approaches are not able to take full advantage of dense initialization as they are often unable to (significantly) improve over sparse SfM‑based initialization. We will make our benchmark publicly available.
Authors:Xiaoran Zhang, Jian Ding, Yuxing Duan, Haoyue Liu, Gang Chen, Yi Chang, Luxin Yan
Abstract:
Turbulence mitigation (TM) is highly ill‑posed due to the stochastic nature of atmospheric turbulence. Most methods rely on multiple frames recorded by conventional cameras to capture stable patterns in natural scenarios. However, they inevitably suffer from a trade‑off between accuracy and efficiency: more frames enhance restoration at the cost of higher system latency and larger data overhead. Event cameras, equipped with microsecond temporal resolution and efficient sensing of dynamic changes, offer an opportunity to break the bottleneck. In this work, we present EHETM, a high‑quality and efficient TM method inspired by the superiority of events to model motions in continuous sequences. We discover two key phenomena: (1) turbulence‑induced events exhibit distinct polarity alternation correlated with sharp image gradients, providing structural cues for restoring scenes; and (2) dynamic objects form spatiotemporally coherent ``event tubes'' in contrast to irregular patterns within turbulent events, providing motion priors for disentangling objects from turbulence. Based on these insights, we design two complementary modules that respectively leverage polarity‑weighted gradients for scene refinement and event‑tube constraints for motion decoupling, achieving high‑quality restoration with few frames. Furthermore, we construct two real‑world event‑frame turbulence datasets covering atmospheric and thermal cases. Experiments show that EHETM outperforms SOTA methods, especially under scenes with dynamic objects, while reducing data overhead and system latency by approximately 77.3% and 89.5%, respectively. Our code is available at: https://github.com/Xavier667/EHETM.
Authors:Canqun Xiang, Chen Yang, Jiaoyan Zhao
Abstract:
Capsule networks (CapsNets) are superior at modeling hierarchical spatial relationships but suffer from two critical limitations: high computational cost due to iterative dynamic routing and poor robustness under input corruptions. To address these issues, we propose IBCapsNet, a novel capsule architecture grounded in the Information Bottleneck (IB) principle. Instead of iterative routing, IBCapsNet employs a one‑pass variational aggregation mechanism, where primary capsules are first compressed into a global context representation and then processed by class‑specific variational autoencoders (VAEs) to infer latent capsules regularized by the KL divergence. This design enables efficient inference while inherently filtering out noise. Experiments on MNIST, Fashion‑MNIST, SVHN and CIFAR‑10 show that IBCapsNet matches CapsNet in clean‑data accuracy (achieving 99.41% on MNIST and 92.01% on SVHN), yet significantly outperforms it under four types of synthetic noise ‑ demonstrating average improvements of +17.10% and +14.54% for clamped additive and multiplicative noise, respectively. Moreover, IBCapsNet achieves 2.54x faster training and 3.64x higher inference throughput compared to CapsNet, while reducing model parameters by 4.66%. Our work bridges information‑theoretic representation learning with capsule networks, offering a principled path toward robust, efficient, and interpretable deep models. Code is available at https://github.com/cxiang26/IBCapsnet
Authors:Ling Xiao, Toshihiko Yamasaki
Abstract:
Most fine‑grained fashion image retrieval (FIR) methods assume a static setting, requiring full retraining when new attributes appear, which is costly and impractical for dynamic scenarios. Although pretrained models support zero‑shot inference, their accuracy drops without supervision, and no prior work explores class‑incremental learning (CIL) for fine‑grained FIR. We propose a multihead continual learning framework for fine‑grained fashion image retrieval with contrastive learning and exponential moving average (EMA) distillation (MCL‑FIR). MCL‑FIR adopts a multi‑head design to accommodate evolving classes across increments, reformulates triplet inputs into doublets with InfoNCE for simpler and more effective training, and employs EMA distillation for efficient knowledge transfer. Experiments across four datasets demonstrate that, beyond its scalability, MCL‑FIR achieves a strong balance between efficiency and accuracy. It significantly outperforms CIL baselines under similar training cost, and compared with static methods, it delivers comparable performance while using only about 30% of the training cost. The source code is publicly available in https://github.com/Dr‑LingXiao/MCL‑FIR.
Authors:Liangyu Yuan, Yufei Huang, Mingkun Lei, Tong Zhao, Ruoyu Wang, Changxi Chi, Yiwei Wang, Chi Zhang
Abstract:
Diffusion models generate synthetic images through an iterative refinement process. However, the misalignment between the simulation‑free objective and the iterative process often causes accumulated gradient error along the sampling trajectory, which leads to unsatisfactory results and a failure to generalize. Guidance techniques like Classifier Free Guidance (CFG) and AutoGuidance (AG) alleviate this by extrapolating between the main and inferior signal for stronger generalization. Despite empirical success, the effective operational regimes of prevalent guidance methods are still under‑explored, leading to ambiguity when selecting the appropriate guidance method given a precondition. In this work, we first conduct synthetic comparisons to isolate and demonstrate the effective regime of guidance methods represented by CFG and AG from the perspective of weak‑to‑strong principle. Based on this, we propose a hybrid instantiation called SGG under the principle, taking the benefits of both. Furthermore, we demonstrate that the W2S principle along with SGG can be migrated into the training objective, improving the generalization ability of unguided diffusion models. We validate our approach with comprehensive experiments. At inference time, evaluations on SD3 and SD3.5 confirm that SGG outperforms existing training‑free guidance variants. Training‑time experiments on transformer architectures demonstrate the effective migration and performance gains in both conditional and unconditional settings. Code is available at https://github.com/851695e35/SGG.
Authors:Fawaz Sammani, Tzoulio Chamiti, Paul Gavrikov, Nikos Deligiannis
Abstract:
Joint Vision‑Language Embedding models such as CLIP typically fail at understanding negation in text queries, for example, failing to distinguish "no" in the query: "a plain blue shirt with no logos". Prior work has largely addressed this limitation through data‑centric approaches, fine‑tuning CLIP on large‑scale synthetic negation datasets. However, these efforts are commonly evaluated using retrieval‑based metrics that cannot reliably reflect whether negation is actually understood. In this paper, we identify two key limitations of such evaluation metrics and investigate an alternative evaluation framework based on Multimodal LLMs‑as‑a‑judge, which typically excel at understanding simple yes/no questions about image content, providing a fair evaluation of negation understanding in CLIP models. We then ask whether there already exists a direction in the CLIP embedding space associated with negation. We find evidence that such a direction exists, and show that it can be manipulated through test‑time intervention via representation engineering to steer CLIP toward negation‑aware behavior without any fine‑tuning. Finally, we test negation understanding on non‑common image‑text samples to evaluate generalization under distribution shifts. Code is at https://github.com/fawazsammani/negation‑steering
Authors:Rui Zhou, Xander Yap, Jianwen Cao, Allison Lau, Boyang Sun, Marc Pollefeys
Abstract:
Target localization is a prerequisite for embodied tasks such as navigation and manipulation. Conventional approaches rely on constructing explicit 3D scene representations to enable target localization, such as point clouds, voxel grids, or scene graphs. While effective, these pipelines incur substantial mapping time, storage overhead, and scalability limitations. Recent advances in vision‑language models suggest that rich semantic reasoning can be performed directly on 2D observations, raising a fundamental question: is a complete 3D scene reconstruction necessary for object localization? In this work, we revisit object localization and propose a map‑free pipeline that stores only posed RGB‑D keyframes as a lightweight visual memory‑‑without constructing any global 3D representation of the scene. At query time, our method retrieves candidate views, re‑ranks them with a vision‑language model, and constructs a sparse, on‑demand 3D estimate of the queried target through depth backprojection and multi‑view fusion. Compared to reconstruction‑based pipelines, this design drastically reduces preprocessing cost, enabling scene indexing that is over two orders of magnitude faster to build while using substantially less storage. We further validate the localized targets on downstream object‑goal navigation tasks. Despite requiring no task‑specific training, our approach achieves strong performance across multiple benchmarks, demonstrating that direct reasoning over image‑based scene memory can effectively replace dense 3D reconstruction for object‑centric robot navigation. Project page: https://ruizhou‑cn.github.io/memory‑over‑maps/
Authors:Yuanhong Zheng, Ruichuan An, Xiaopeng Lin, Yuxing Liu, Sihan Yang, Huanyu Zhang, Haodong Li, Qintong Zhang, Renrui Zhang, Guopeng Li, Yifan Zhang, Yuheng Li, Wentao Zhang
Abstract:
Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimodal personalization methods are largely limited to static images or offline videos. This disconnects continuous visual input from instant real‑world feedback, limiting their ability to provide the real‑time, interactive personalized responses essential for future AI assistants. To bridge this gap, we first propose and formally define the novel task of Personalized Streaming Video Understanding (PSVU). To facilitate research in this new direction, we introduce PEARL‑Bench, the first comprehensive benchmark designed specifically to evaluate this challenging setting. It evaluates a model's ability to respond to personalized concepts at exact timestamps under two modes: (1) Frame‑level, focusing on a specific person or object in discrete frames, and (2) a novel Video‑level, focusing on personalized actions unfolding across continuous frames. PEARL‑Bench comprises 132 unique videos and 2,173 fine‑grained annotations with precise timestamps. Concept diversity and annotation quality are strictly ensured through a combined pipeline of automated generation and human verification. To tackle this challenging new setting, we further propose PEARL, a plug‑and‑play, training‑free strategy that serves as a strong baseline. Extensive evaluations across 8 offline and online models demonstrate that PEARL achieves state‑of‑the‑art performance. Notably, it brings consistent PSVU improvements when applied to 3 distinct architectures, proving to be a highly effective and robust strategy. We hope this work advances vision‑language model (VLM) personalization and inspires further research into streaming personalized AI assistants. Code is available at https://github.com/Yuanhong‑Zheng/PEARL.
Authors:Liu hung ming
Abstract:
Video world models trained with Joint Embedding Predictive Architectures (JEPA) acquire rich spatiotemporal representations by predicting masked regions in latent space rather than reconstructing pixels. This removes the visual verification pathway of generative models, creating a structural interpretability gap: the encoder has learned physical structure inaccessible in any inspectable form. Existing probing methods either operate in continuous space without a structured intermediate layer, or attach generative components whose parameters confound attribution of behavior to the encoder.
We propose the AI Mother Tongue (AIM) framework as a passive quantization probe: a lightweight, vocabulary‑free probe that converts V‑JEPA 2 continuous latent vectors into discrete symbol sequences without task‑specific supervision or modifying the encoder. Because the encoder is kept completely frozen, any symbolic structure in the AIM codebook is attributable entirely to V‑JEPA 2 pre‑trained representations ‑‑ not to the probe.
We evaluate through category‑contrast experiments on Kinetics‑mini along three physical dimensions: grasp angle, object geometry, and motion temporal structure. AIM symbol distributions differ significantly across all three experiments (chi^2 p < 10^‑4; MI 0.036‑‑0.117 bits, NMI 1.2‑‑3.9% of the 3‑bit maximum; JSD up to 0.342; codebook active ratio 62.5%). The experiments reveal that V‑JEPA 2 latent space is markedly compact: diverse action categories share a common representational core, with semantic differences encoded as graded distributional variations rather than categorical boundaries. These results establish Stage 1 of a four‑stage roadmap toward an action‑conditioned symbolic world model, demonstrating that structured symbolic manifolds are discoverable properties of frozen JEPA latent spaces.
Authors:Xiaojian Lin, Yaomin Shen, Junyuan Ma, Yujie Sun, Chengqing Bu, Wenxin Zhang, Zongzheng Zhang, Hao Fei, Lei Jin, Hao Zhao
Abstract:
Monocular vertex‑level human‑scene contact prediction is a fundamental capability for interactive systems such as assistive monitoring, embodied AI, and rehabilitation analysis. In this work, we study this task jointly with single‑image 3D human mesh reconstruction, using reconstructed body geometry as a scaffold for contact reasoning. Existing approaches either focus on contact prediction without sufficiently exploiting explicit 3D human priors, or emphasize pose/mesh reconstruction without directly optimizing robust vertex‑level contact inference under occlusion and perceptual noise. To address this gap, we propose GraphiContact, a pose‑aware framework that transfers complementary human priors from two pretrained Transformer encoders and predicts per‑vertex human‑scene contact on the reconstructed mesh. To improve robustness in real‑world scenarios, we further introduce a Single‑Image Multi‑Infer Uncertainty (SIMU) training strategy with token‑level adaptive routing, which simulates occlusion and noisy observations during training while preserving efficient single‑branch inference at test time. Experiments on five benchmark datasets show that GraphiContact achieves consistent gains on both contact prediction and 3D human reconstruction. Our code, based on the GraphiContact method, provides comprehensive 3D human reconstruction and interaction analysis, and will be publicly available at https://github.com/Aveiro‑Lin/GraphiContact.
Authors:Qihao Lin, Borui Chen, Yuping Zhou, Jianing Wu, Yulan Guo, Weishi Zheng, Chongkun Xia
Abstract:
The contour estimation of transparent fragments is very important for autonomous reassembly, especially in the fields of precision optical instrument repair, cultural relic restoration, and identification of other precious device broken accidents. Different from general intact transparent objects, the contour estimation of transparent fragments face greater challenges due to strict optical properties, irregular shapes and edges. To address this issue, a general transparent fragments contour estimation framework based on visual‑tactile fusion is proposed in this paper. First, we construct the transparent fragment dataset named TransFrag27K, which includes a multiscene synthetic data of broken fragments from multiple types of transparent objects, and a scalable synthetic data generation pipeline. Secondly, we propose a visual grasping position detection network named TransFragNet to identify, locate and segment the sampling grasping position. And, we use a two‑finger gripper with Gelsight Mini sensors to obtain reconstructed tactile information of the lateral edge of the fragments. By fusing this tactile information with visual cues, a visual‑tactile fusion material classifier is proposed. Inspired by the way humans estimate a fragment's contour combining vision and touch, we introduce a general transparent fragment contour estimation framework based on visual‑tactile fusion, demonstrates strong performance in real‑world validation. Finally, a multi‑dimensional similarity metrics based contour matching and reassembly algorithm is proposed, providing a reproducible benchmark for evaluating visual‑tactile contour estimation and fragment reassembly. The experimental results demonstrate the validity of the proposed framework. The dataset and codes are available at https://github.com/Keithllin/Transparent‑Fragments‑Contour‑Estimation.
Authors:Heng Zhou, Xiaoxiong Liu, Zhenxi Zhang, Jieheng Yun, Chengyang Li, Yunchu Yang, Dongyi Xia, Chunna Tian, Xiao-Jun Wu
Abstract:
Remote sensing images (RSIs) are frequently degraded by haze, fog, and thin clouds, which obscure surface reflectance and hinder downstream applications. This study presents the first systematic and unified survey of RSIs dehazing, integrating methodological evolution, benchmark assessment, and physical consistency analysis. We categorize existing approaches into a three‑stage progression: from handcrafted physical priors, to data‑driven deep restoration, and finally to hybrid physical‑intelligent generation, and summarize more than 30 representative methods across CNNs, GANs, Transformers, and diffusion models. To provide a reliable empirical reference, we conduct large‑scale quantitative experiments on five public datasets using 12 metrics, including PSNR, SSIM, CIEDE, LPIPS, FID, SAM, ERGAS, UIQI, QNR, NIQE, and HIST. Cross‑domain comparison reveals that recent Transformer‑ and diffusion‑based models improve SSIM by 12%~18% and reduce perceptual errors by 20%~35% on average, while hybrid physics‑guided designs achieve higher radiometric stability. A dedicated physical radiometric consistency experiment further demonstrates that models with explicit transmission or airlight constraints reduce color bias by up to 27%. Based on these findings, we summarize open challenges: dynamic atmospheric modeling, multimodal fusion, lightweight deployment, data scarcity, and joint degradations, and outline promising research directions for future development of trustworthy, controllable, and efficient (TCE) dehazing systems. All reviewed resources, including source code, benchmark datasets, evaluation metrics, and reproduction configurations are publicly available at https://github.com/VisionVerse/RemoteSensing‑Restoration‑Survey.
Authors:Behnood Rasti, Bikram Koirala, Paul Scheunders
Abstract:
This paper proposes a semisupervised geometric unmixing approach called minimum simplex semisupervised unmixing (MiSiSUn). The geometry of the data was incorporated for the first time into library‑based unmixing using a simplex‑volume‑flavored penalty based on an archetypal analysis‑type linear model. The experimental results were performed on two simulated datasets considering different levels of mixing ratios and spatial instruction at varying input noise. MiSiSUn considerably outperforms state‑of‑the‑art semisupervised unmixing methods. The improvements vary from 1 dB to over 3 dB in different scenarios. The proposed method was also applied to a real dataset where visual interpretation is close to the geological map. MiSiSUn was implemented using PyTorch, which is open‑source and available at https://github.com/BehnoodRasti/MiSiSUn. Moreover, we provide a dedicated Python package for Semisupervised Unmixing, which is open‑source and includes all the methods used in the experiments for the sake of reproducibility.
Authors:Xinyi Shang, Yi Tang, Jiacheng Cui, Ahmed Elhagry, Salwa K. Al Khatib, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Jing-Hao Xue, Hao Li, Salman Khan, Zhiqiang Shen
Abstract:
Existing tampering detection benchmarks largely rely on object masks, which severely misalign with the true edit signal: many pixels inside a mask are untouched or only trivially modified, while subtle yet consequential edits outside the mask are treated as natural. We reformulate VLM image tampering from coarse region labels to a pixel‑grounded, meaning and language‑aware task. First, we introduce a taxonomy spanning edit primitives (replace/remove/splice/inpaint/attribute/colorization, etc.) and their semantic class of tampered object, linking low‑level changes to high‑level understanding. Second, we release a new benchmark with per‑pixel tamper maps and paired category supervision to evaluate detection and classification within a unified protocol. Third, we propose a training framework and evaluation metrics that quantify pixel‑level correctness with localization to assess confidence or prediction on true edit intensity, and further measure tamper meaning understanding via semantics‑aware classification and natural language descriptions for the predicted regions. We also re‑evaluate the existing strong segmentation/localization baselines on recent strong tamper detectors and reveal substantial over‑ and under‑scoring using mask‑only metrics, and expose failure modes on micro‑edits and off‑mask changes. Our framework advances the field from masks to pixels, meanings and language descriptions, establishing a rigorous standard for tamper localization, semantic classification and description. Code and benchmark data are available at https://github.com/VILA‑Lab/PIXAR.
Authors:Jiazheng Xing, Fei Du, Hangjie Yuan, Pengwei Liu, Hongbin Xu, Hai Ci, Ruigang Niu, Weihua Chen, Fan Wang, Yong Liu
Abstract:
Recent advances in diffusion models have significantly improved text‑to‑video generation, enabling personalized content creation with fine‑grained control over both foreground and background elements. However, precise face‑attribute alignment across subjects remains challenging, as existing methods lack explicit mechanisms to ensure intra‑group consistency. Addressing this gap requires both explicit modeling strategies and face‑attribute‑aware data resources. We therefore propose LumosX, a framework that advances both data and model design. On the data side, a tailored collection pipeline orchestrates captions and visual cues from independent videos, while multimodal large language models (MLLMs) infer and assign subject‑specific dependencies. These extracted relational priors impose a finer‑grained structure that amplifies the expressive control of personalized video generation and enables the construction of a comprehensive benchmark. On the modeling side, Relational Self‑Attention and Relational Cross‑Attention intertwine position‑aware embeddings with refined attention dynamics to inscribe explicit subject‑attribute dependencies, enforcing disciplined intra‑group cohesion and amplifying the separation between distinct subject clusters. Comprehensive evaluations on our benchmark demonstrate that LumosX achieves state‑of‑the‑art performance in fine‑grained, identity‑consistent, and semantically aligned personalized multi‑subject video generation. Code and models are available at https://jiazheng‑xing.github.io/lumosx‑home/.
Authors:Omkar Thawakar, Dmitry Demidov, Vaishnav Potlapalli, Sai Prasanna Teja Reddy Bogireddy, Viswanatha Reddy Gajjala, Alaa Mostafa Lasheen, Rao Muhammad Anwer, Fahad Khan
Abstract:
Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification text fully specifies the visual changes, overlooking after‑effects and implicit consequences (e.g., motion, state transitions, viewpoint or duration cues) that emerge from the edit. We argue that successful CoVR requires reasoning about these after‑effects. We introduce a reasoning‑first, zero‑shot approach that leverages large multimodal models to (i) infer causal and temporal consequences implied by the edit, and (ii) align the resulting reasoned queries to candidate videos without task‑specific finetuning. To evaluate reasoning in CoVR, we also propose CoVR‑Reason, a benchmark that pairs each (reference, edit, target) triplet with structured internal reasoning traces and challenging distractors that require predicting after‑effects rather than keyword matching. Experiments show that our zero‑shot method outperforms strong retrieval baselines on recall at K and particularly excels on implicit‑effect subsets. Our automatic and human analysis confirm higher step consistency and effect factuality in our retrieved results. Our findings show that incorporating reasoning into general‑purpose multimodal models enables effective CoVR by explicitly accounting for causal and temporal after‑effects. This reduces dependence on task‑specific supervision, improves generalization to challenging implicit‑effect cases, and enhances interpretability of retrieval outcomes. These results point toward a scalable and principled framework for explainable video search. The model, code, and benchmark are available at https://github.com/mbzuai‑oryx/CoVR‑R.
Authors:Sebastian Gerard, Josephine Sullivan
Abstract:
Predicting future states in uncertain environments, such as wildfire spread, medical diagnosis, or autonomous driving, requires models that can consider multiple plausible outcomes. While diffusion models can effectively learn such multi‑modal distributions, naively sampling from these models is computationally inefficient, potentially requiring hundreds of samples to find low‑probability modes that may still be operationally relevant. In this work, we address the challenge of sample‑efficient ambiguous segmentation by evaluating several training‑free sampling methods that encourage diverse predictions. We adapt two techniques, particle guidance and SPELL, originally designed for the generation of diverse natural images, to discrete segmentation tasks, and additionally propose a simple clustering‑based technique. We validate these approaches on the LIDC medical dataset, a modified version of the Cityscapes dataset, and MMFire, a new simulation‑based wildfire spread dataset introduced in this paper. Compared to naive sampling, these approaches increase the HM IoU metric by up to 7.5% on MMFire and 16.4% on Cityscapes, demonstrating that training‑free methods can be used to efficiently increase the sample diversity of segmentation diffusion models with little cost to image quality and runtime.
Code and dataset: https://github.com/SebastianGer/wildfire‑spread‑scenarios
Authors:Yuan Zhou, Yongzhi Li, Yanqi Dai, Xingyu Zhu, Yi Tan, Qingshan Xu, Beier Zhu, Richang Hong, Hanwang Zhang
Abstract:
Video‑driven human reaction generation aims to synthesize 3D human motions that directly react to observed video sequences, which is crucial for building human‑like interactive AI systems. However, existing methods often fail to effectively leverage video inputs to steer human reaction synthesis, resulting in reaction motions that are mismatched with the content of video sequences. We reveal that this limitation arises from a severe relational distortion between visual observations and reaction types. In light of this, we propose MuSteerNet, a simple yet effective framework that generates 3D human reactions from videos via observation‑reaction mutual steering. Specifically, we first propose a Prototype Feedback Steering mechanism to mitigate relational distortion by refining visual observations with a gated delta‑rectification modulator and a relational margin constraint, guided by prototypical vectors learned from human reactions. We then introduce Dual‑Coupled Reaction Refinement that fully leverages rectified visual cues to further steer the refinement of generated reaction motions, thereby effectively improving reaction quality and enabling MuSteerNet to achieve competitive performance. Extensive experiments and ablation studies validate the effectiveness of our method. Code coming soon: https://github.com/zhouyuan888888/MuSteerNet.
Authors:Puskal Khadka, KC Santosh
Abstract:
State Space Models (SSMs), especially recent Mamba architecture, have achieved remarkable success in sequence modeling tasks. However, extending SSMs to computer vision remains challenging due to the non‑sequential structure of visual data and its complex 2D spatial dependencies. Although several early studies have explored adapting selective SSMs for vision applications, most approaches primarily depend on employing various traversal strategies over the same input. This introduces redundancy and distorts the intricate spatial relationships within images. To address these challenges, we propose MFil‑Mamba, a novel visual state space architecture built on a multi‑filter scanning backbone. Unlike fixed multi‑directional traversal methods, our design enables each scan to capture unique and contextually relevant spatial information while minimizing redundancy. Furthermore, we incorporate an adaptive weighting mechanism to effectively fuse outputs from multiple scans in addition to architectural enhancements. MFil‑Mamba achieves superior performance over existing state‑of‑the‑art models across various benchmarks that include image classification, object detection, instance segmentation, and semantic segmentation. For example, our tiny variant attains 83.2% top‑1 accuracy on ImageNet‑1K, 47.3% box AP and 42.7% mask AP on MS COCO, and 48.5% mIoU on the ADE20K dataset. Code and models are available at https://github.com/puskal‑khadka/MFil‑Mamba.
Authors:Tianling Liu, Hongying Liu, Fanhua Shang, Lequan Yu, Tong Han, Liang Wan
Abstract:
In clinical practice, crossmodal information including medical images and tabular data is essential for disease diagnosis. There exists a significant modality gap between these data types, which obstructs advancements in crossmodal diagnostic accuracy. Most existing crossmodal learning (CML) methods primarily focus on exploring relationships among high‑level encoder outputs, leading to the neglect of local information in images. Additionally, these methods often overlook the extraction of task‑relevant information. In this paper, we propose a novel coarse‑to‑fine crossmodal learning (CFCML) framework to progressively reduce the modality gap between multimodal images and tabular data, by thoroughly exploring inter‑modal relationships. At the coarse stage, we explore the relationships between multi‑granularity features from various image encoder stages and tabular information, facilitating a preliminary reduction of the modality gap. At the fine stage, we generate unimodal and crossmodal prototypes that incorporate class‑aware information, and establish hierarchical anchor‑based relationship mining (HRM) strategy to further diminish the modality gap and extract discriminative crossmodal information. This strategy utilize modality samples, unimodal prototypes, and crossmodal prototypes as anchors to develop contrastive learning approaches, effectively enhancing inter‑class disparity while reducing intra‑class disparity from multiple perspectives. Experimental results indicate that our method outperforms the state‑of‑the‑art (SOTA) methods, achieving improvements of 1.53% and 0.91% in AUC metrics on the MEN and Derm7pt datasets, respectively. The code is available at https://github.com/IsDling/CFCML.
Authors:Zheng Gao, Debin Meng, Yunqi Miao, Zhensong Zhang, Songcen Xu, Ioannis Patras, Jifei Song
Abstract:
Current diffusion‑based makeup transfer methods commonly use the makeup information encoded by off‑the‑shelf foundation models (e.g., CLIP) as condition to preserve the makeup style of reference image in the generation. Although effective, these works mainly have two limitations: (1) foundation models pre‑trained for generic tasks struggle to capture makeup styles; (2) the makeup features of reference image are injected to the diffusion denoising model as a whole for global makeup transfer, overlooking the facial region‑aware makeup features (i.e., eyes, mouth, etc) and limiting the regional controllability for region‑specific makeup transfer. To address these, in this work, we propose Facial Region‑Aware Makeup features (FRAM), which has two stages: (1) makeup CLIP fine‑tuning; (2) identity and facial region‑aware makeup injection. For makeup CLIP fine‑tuning, unlike prior works using off‑the‑shelf CLIP, we synthesize annotated makeup style data using GPT‑o3 and text‑driven image editing model, and then use the data to train a makeup CLIP encoder through self‑supervised and image‑text contrastive learning. For identity and facial region‑aware makeup injection, we construct before‑and‑after makeup image pairs from the edited images in stage 1 and then use them to learn to inject identity of source image and makeup of reference image to the diffusion denoising model for makeup transfer. Specifically, we use learnable tokens to query the makeup CLIP encoder to extract facial region‑aware makeup features for makeup injection, which is learned via an attention loss to enable regional control. As for identity injection, we use a ControlNet Union to encode source image and its 3D mesh simultaneously. The experimental results verify the superiority of our regional controllability and our makeup transfer performance. Code is available at https://github.com/zaczgao/Facial_Region‑Aware_Makeup.
Authors:Haoyue Liu, Jinghan Xu, Luxin Feng, Hanyu Zhou, Haozhi Zhao, Yi Chang, Luxin Yan
Abstract:
High‑quality imaging of dynamic scenes in extremely low‑light conditions is highly challenging. Photon scarcity induces severe noise and texture loss, causing significant image degradation. Event cameras, featuring a high dynamic range (120 dB) and high sensitivity to motion, serve as powerful complements to conventional cameras by offering crucial cues for preserving subtle textures. However, most existing approaches emphasize texture recovery from events, while paying little attention to image noise or the intrinsic noise of events themselves, which ultimately hinders accurate pixel reconstruction under photon‑starved conditions. In this work, we propose NEC‑Diff, a novel diffusion‑based event‑RAW hybrid imaging framework that extracts reliable information from heavily noisy signals to reconstruct fine scene structures. The framework is driven by two key insights: (1) combining the linear light‑response property of RAW images with the brightness‑change nature of events to establish a physics‑driven constraint for robust dual‑modal denoising; and (2) dynamically estimating the SNR of both modalities based on denoising results to guide adaptive feature fusion, thereby injecting reliable cues into the diffusion process for high‑fidelity visual reconstruction. Furthermore, we construct the REAL (Raw and Event Acquired in Low‑light) dataset which provides 47,800 pixel‑aligned low‑light RAW images, events, and high‑quality references under 0.001‑0.8 lux illumination. Extensive experiments demonstrate the superiority of NEC‑Diff under extreme darkness. The project are available at: https://github.com/jinghan‑xu/NEC‑Diff.
Authors:Rozain Shakeel, Abdul Rahman Mohammad Ali, Muneeb Mushtaq, Tausifa Jan Saleem, Tajamul Ashraf
Abstract:
Despite the rapid progress of Multimodal Large Language Models (MLLMs), their ability to perform reliable visual grounding in high‑stakes clinical software environments remains underexplored. Existing GUI benchmarks largely focus on isolated, single‑step grounding queries, overlooking the sequential, workflow‑driven reasoning required in real‑world medical interfaces, where tasks evolve across independent steps and dynamic interface states. We introduce MedSPOT, a workflow‑aware sequential grounding benchmark for clinical GUI environments. Unlike prior benchmarks that treat grounding as a standalone prediction task, MedSPOT models procedural interaction as a sequence of structured spatial decisions. The benchmark comprises 216 task‑driven videos with 597 annotated keyframes, in which each task consists of 2 to 3 interdependent grounding steps within realistic medical workflows. This design captures interface hierarchies, contextual dependencies, and fine‑grained spatial precision under evolving conditions. To evaluate procedural robustness, we propose a strict sequential evaluation protocol that terminates task assessment upon the first incorrect grounding prediction, explicitly measuring error propagation in multi‑step workflows. We further introduce a comprehensive failure taxonomy, including edge bias, small‑target errors, no prediction, near miss, far miss, and toolbar confusion, to enable systematic diagnosis of model behavior in clinical GUI settings. By shifting evaluation from isolated grounding to workflow‑aware sequential reasoning, MedSPOT establishes a realistic and safety‑critical benchmark for assessing multimodal models in medical software environments. Code and data are available at: https://github.com/Tajamul21/MedSPOT.
Authors:Jizhou Han, Chenhao Ding, Yuhang He, Qiang Wang, Shaokun Wang, SongLin Dong, Yihong Gong
Abstract:
Generalized Category Discovery (GCD) seeks to uncover novel categories in unlabeled data while preserving recognition of known categories, yet prevailing visual‑only pipelines and the loose coupling between supervised learning and discovery often yield brittle boundaries on fine‑grained, look‑alike categories. We introduce the Analogical Textual Concept Generator (ATCG), a plug‑and‑play module that analogizes from labeled knowledge to new observations, forming textual concepts for unlabeled samples. Fusing these analogical textual concepts with visual features turns discovery into a visual‑textual reasoning process, transferring prior knowledge to novel data and sharpening category separation. ATCG attaches to both parametric and clustering style GCD pipelines and requires no changes to their overall design. Across six benchmarks, ATCG consistently improves overall, known‑class, and novel‑class performance, with the largest gains on fine‑grained data. Our code is available at: https://github.com/zhou‑9527/AnaLogical‑GCD.
Authors:Minyue Dai, Ke Fan, Anyi Rao, Jingbo Wang, Bo Dai
Abstract:
Text‑to‑motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on cumbersome, high‑dimensional joint constraints (e.g., trajectories), which hinder user‑friendly, iterative refinement. To address this, we propose Modular Body‑Part Phase Control, a plug‑and‑play framework enabling structured, localized editing via a compact, scalar‑based phase interface. By modeling body‑part latent motion channels as sinusoidal phase signals characterized by amplitude, frequency, phase shift, and offset, we extract interpretable codes that capture part‑specific dynamics. A modular Phase ControlNet branch then injects this signal via residual feature modulation, seamlessly decoupling control from the generative backbone. Experiments on both diffusion‑ and flow‑based models demonstrate that our approach provides predictable and fine‑grained control over motion magnitude, speed, and timing. It preserves global motion coherence and offers a practical paradigm for controllable T2M generation. Project page: https://jixiii.github.io/bp‑phase‑project‑page/
Authors:Yifei Zhao, Fanyu Zhao, Zhongyuan Zhang, Shengtang Wu, Yixuan Lin, Yinsheng Li
Abstract:
Generalized few‑shot 3D point cloud segmentation aims to adapt to novel classes from only a few annotations while maintaining strong performance on base classes, but this remains challenging due to the inherent stability‑plasticity trade‑off: adapting to novel classes can interfere with shared representations and cause base‑class forgetting. We present HOP3D, a unified framework that learns hierarchical orthogonal prototypes with an entropy‑based few‑shot regularizer to enable robust novel‑class adaptation without degrading base‑class performance. HOP3D introduces hierarchical orthogonalization that decouples base and novel learning at both the gradient and representation levels, effectively mitigating base‑novel interference. To further enhance adaptation under sparse supervision, we incorporate an entropy‑based regularizer that leverages predictive uncertainty to refine prototype learning and promote balanced predictions. Extensive experiments on ScanNet200 and ScanNet++ demonstrate that HOP3D consistently outperforms state‑of‑the‑art baselines under both 1‑shot and 5‑shot settings. The code is available at https://fdueblab‑hop3d.github.io/.
Authors:Hantao Zheng, Ning Han, Yawen Zeng, Hao Chen
Abstract:
Recent weakly supervised video anomaly detection methods have achieved significant advances by employing unified frameworks for joint optimization. However, this paradigm is limited by a fundamental sensitivity‑stability trade‑off, as the conflicting objectives for detecting transient and sustained anomalies lead to either fragmented predictions or over‑smoothed responses. To address this limitation, we propose DeSC, a novel Decoupled Sensitivity‑Consistency framework that trains two specialized streams using distinct optimization strategies. The temporal sensitivity stream adopts an aggressive optimization strategy to capture high‑frequency abrupt changes, whereas the semantic consistency stream applies robust constraints to maintain long‑term coherence and reduce noise. Their complementary strengths are fused through a collaborative inference mechanism that reduces individual biases and produces balanced predictions. Extensive experiments demonstrate that DeSC establishes new state‑of‑the‑art performance by achieving 89.37% AUC on UCF‑Crime (+1.29%) and 87.18% AP on XD‑Violence (+2.22%). Code is available at https://github.com/imzht/DeSC.
Authors:Wen Yin, Cencen Liu, Dingrui Liu, Bing Su, Yuan-Fang Li, Tao He
Abstract:
Unifying Image Quality Assessment (IQA) and Image Aesthetic Assessment (IAA) in a single multimodal large language model is appealing, yet existing methods adopt a task‑agnostic recipe that applies the same reasoning strategy and reward to both tasks. We show this is fundamentally misaligned: IQA relies on low‑level, objective perceptual cues and benefits from concise distortion‑focused reasoning, whereas IAA requires deliberative semantic judgment and is poorly served by point‑wise score regression. We identify these as a reasoning mismatch and an optimization mismatch, and provide empirical evidence for both through controlled probes. Motivated by these findings, we propose TATAR (Task‑Aware Thinking with Asymmetric Rewards), a unified framework that shares the visual‑language backbone while conditioning post‑training on each task's nature. TATAR combines three components: fast‑‑slow task‑specific reasoning construction that pairs IQA with concise perceptual rationales and IAA with deliberative aesthetic narratives; two‑stage SFT+GRPO learning that establishes task‑aware behavioral priors before reward‑driven refinement; and asymmetric rewards that apply Gaussian score shaping for IQA and Thurstone‑style completion ranking for IAA. Extensive experiments across eight benchmarks demonstrate that TATAR consistently outperforms prior unified baselines on both tasks under in‑domain and cross‑domain settings, remains competitive with task‑specific specialized models, and yields more stable training dynamics for aesthetic assessment. Our results establish task‑conditioned post‑training as a principled paradigm for unified perceptual scoring. Our code is publicly available at https://github.com/yinwen2019/TATAR.
Authors:Chengzhi Hong, Bijun Li
Abstract:
Monocular 3D lane detection remains challenging due to depth ambiguity and weak geometric constraints. Mainstream methods rely on depth guidance, BEV projection, and anchor‑ or curve‑based heads with simplified physical assumptions, remapping high‑dimensional image features while only weakly encoding road geometry. Lacking an invariant geometric‑topological coupling between lanes and the underlying road surface, 2D‑to‑3D lifting is ill‑posed and brittle, often degenerating into concavities, bulges, and twists. To address this, we propose the Road‑Manifold Assumption: the road is a smooth 2D manifold in \mathbbR^3, lanes are embedded 1D submanifolds, and sampled lane points are dense observations, thereby coupling metric and topology across surfaces, curves, and point sets. Building on this, we propose ReManNet, which first produces initial lane predictions with an image backbone and detection heads, then encodes geometry as Riemannian Gaussian descriptors on the symmetric positive‑definite (SPD) manifold, and fuses these descriptors with visual features through a lightweight gate to maintain coherent 3D reasoning. We also propose the 3D Tunnel Lane IoU (3D‑TLIoU) loss, a joint point‑curve objective that computes slice‑wise overlap of tubular neighborhoods along each lane to improve shape‑level alignment. Extensive experiments on standard benchmarks demonstrate that ReManNet achieves state‑of‑the‑art (SOTA) or competitive results. On OpenLane, it improves F1 by +8.2% over the baseline and by +1.8% over the previous best, with scenario‑level gains of up to +6.6%. The code will be publicly available at https://github.com/changehome717/ReManNet.
Authors:Yifei Zhao, Fanyu Zhao, Yinsheng Li
Abstract:
Few‑shot 3D semantic segmentation aims to generate accurate semantic masks for query point clouds with only a few annotated support examples. Existing prototype‑based methods typically construct compact and deterministic prototypes from the support set to guide query segmentation. However, such rigid representations are unable to capture the intrinsic uncertainty introduced by scarce supervision, which often results in degraded robustness and limited generalization. In this work, we propose UPL (Uncertainty‑aware Prototype Learning), a probabilistic approach designed to incorporate uncertainty modeling into prototype learning for few‑shot 3D segmentation. Our framework introduces two key components. First, UPL introduces a dual‑stream prototype refinement module that enriches prototype representations by jointly leveraging limited information from both support and query samples. Second, we formulate prototype learning as a variational inference problem, regarding class prototypes as latent variables. This probabilistic formulation enables explicit uncertainty modeling, providing robust and interpretable mask predictions. Extensive experiments on the widely used ScanNet and S3DIS benchmarks show that our UPL achieves consistent state‑of‑the‑art performance under different settings while providing reliable uncertainty estimation. The code is available at https://fdueblab‑upl.github.io/.
Authors:Jiadong Liang, Bojun Xiong, Jie Tian, Hua Li, Xiao Long, Yong Zheng, Huan Fu
Abstract:
This paper primarily investigates the task of expression‑only portrait video performance editing based on a driving video, which plays a crucial role in animation and film industries. Most existing research mainly focuses on portrait animation, which aims to animate a static portrait image according to the facial motion from the driving video. As a consequence, it remains challenging for them to disentangle the facial expression from head pose rotation and thus lack the ability to edit facial expression independently. In this paper, we propose PerformRecast, a versatile expression‑only video editing method which is dedicated to recast the performance in existing film and animation. The key insight of our method comes from the characteristics of 3D Morphable Face Model (3DMM), which models the face identity, facial expression and head pose of 3D face mesh with separate parameters. Therefore, we improve the keypoints transformation formula in previous methods to make it more consistent with 3DMM model, which achieves a better disentanglement and provides users with much more fine‑grained control. Furthermore, to avoid the misalignment around the boundary of face in generated results, we decouple the facial and non‑facial regions of input portrait images and pre‑train a teacher model to provide separate supervision for them. Extensive experiments show that our method produces high‑quality results which are more faithful to the driving video, outperforming existing methods in both controllability and efficiency. Our code, data and trained models are available at https://youku‑aigc.github.io/PerformRecast.
Authors:Phuong-Anh Nguyen, Tien Anh Pham, Duc-Trong Le, Cam-Van Thi Nguyen
Abstract:
Learning from multiple modalities often suffers from imbalance, where information‑rich modalities dominate optimization while weaker or partially missing modalities contribute less. This imbalance becomes severe in realistic settings with imbalanced missing rates (IMR), where each modality is absent with different probabilities, distorting representation learning and gradient dynamics. We revisit this issue from a training‑process perspective and propose BALM, a model‑agnostic plug‑in framework to achieve balanced multimodal learning under IMR. The framework comprises two complementary modules: the Feature Calibration Module (FCM), which recalibrates unimodal features using global context to establish a shared representation basis across heterogeneous missing patterns; the Gradient Rebalancing Module (GRM), which balances learning dynamics across modalities by modulating gradient magnitudes and directions from both distributional and spatial perspectives. BALM can be seamlessly integrated into diverse backbones, including multimodal emotion recognition (MER) models, without altering their architectures. Experimental results across multiple MER benchmarks confirm that BALM consistently enhances robustness and improves performance under diverse missing and imbalance settings. Code available at: https://github.com/np4s/BALM_CVPR2026.git
Authors:Chaoqin Huang, Zi Zeng, Aofan Jiang, Yuchen Xu, Qing Cao, Kang Chen, Chenfei Chi, Yanfeng Wang, Ya Zhang
Abstract:
Rare cardiac anomalies are difficult to detect from electrocardiograms (ECGs) due to their long‑tailed distribution with extremely limited case counts and demographic disparities in diagnostic performance. These limitations contribute to delayed recognition and uneven quality of care, creating an urgent need for a generalizable framework that enhances sensitivity while ensuring equity across diverse populations. In this study, we developed an AI‑assisted two‑stage ECG framework integrating self‑supervised anomaly detection with demographic‑aware representation learning. The first stage performs self‑supervised anomaly detection pretraining by reconstructing masked global and local ECG signals, modeling signal trends, and predicting patient attributes to learn robust ECG representations without diagnostic labels. The pretrained model is then fine‑tuned for multi‑label ECG classification using asymmetric loss to better handle long‑tail cardiac abnormalities, and additionally produces anomaly score maps for localization, with CPU‑based optimization enabling practical deployment. Evaluated on a longitudinal cohort of over one million clinical ECGs, our method achieves an AUROC of 94.7% for rare anomalies and reduces the common‑rare performance gap by 73%, while maintaining consistent diagnostic accuracy across age and sex groups. In conclusion, the proposed equity‑aware AI framework demonstrates strong clinical utility, interpretable anomaly localization, and scalable performance across multiple cohorts, highlighting its potential to mitigate diagnostic disparities and advance equitable anomaly detection in biomedical signals and digital health. Source code is available at https://github.com/MediaBrain‑SJTU/Rare‑ECG.
Authors:Takeshi Noda, Yu-Shen Liu, Zhizhong Han
Abstract:
Rendering 3D surfaces has been revolutionized within the modeling of radiance fields through either 3DGS or NeRF. Although 3DGS has shown advantages over NeRF in terms of rendering quality or speed, there is still room for improvement in recovering high fidelity surfaces through 3DGS. To resolve this issue, we propose a self‑constrained prior to constrain the learning of 3D Gaussians, aiming for more accurate depth rendering. Our self‑constrained prior is derived from a TSDF grid that is obtained by fusing the depth maps rendered with current 3D Gaussians. The prior measures a distance field around the estimated surface, offering a band centered at the surface for imposing more specific constraints on 3D Gaussians, such as removing Gaussians outside the band, moving Gaussians closer to the surface, and encouraging larger or smaller opacity in a geometry‑aware manner. More importantly, our prior can be regularly updated by the most recent depth images which are usually more accurate and complete. In addition, the prior can also progressively narrow the band to tighten the imposed constraints. We justify our idea and report our superiority over the state‑of‑the‑art methods in evaluations on widely used benchmarks.
Authors:Shicai Wei, Kaijie Zhang, Luyi Chen, Tao He, Guiduo Duan
Abstract:
Traditional multimodal methods often assume static modality quality, which limits their adaptability in dynamic real‑world scenarios. Thus, dynamical multimodal methods are proposed to assess modality quality and adjust their contribution accordingly. However, they typically rely on empirical metrics, failing to measure the modality quality when noise levels are extremely low or high. Moreover, existing methods usually assume that the initial contribution of each modality is the same, neglecting the intrinsic modality dependency bias. As a result, the modality hard to learn would be doubly penalized, and the performance of dynamical fusion could be inferior to that of static fusion. To address these challenges, we propose the Unbiased Dynamic Multimodal Learning (UDML) framework. Specifically, we introduce a noise‑aware uncertainty estimator that adds controlled noise to the modality data and predicts its intensity from the modality feature. This forces the model to learn a clear correspondence between feature corruption and noise level, allowing accurate uncertainty measure across both low‑ and high‑noise conditions. Furthermore, we quantify the inherent modality reliance bias within multimodal networks via modality dropout and incorporate it into the weighting mechanism. This eliminates the dual suppression effect on the hard‑to‑learn modality. Extensive experiments across diverse multimodal benchmark tasks validate the effectiveness, versatility, and generalizability of the proposed UDML. The code is available at https://github.com/shicaiwei123/UDML.
Authors:Kunlun Xu, Haotong Cheng, Jiangmeng Li, Xu Zou, Jiahuan Zhou
Abstract:
Lifelong person re‑identification (LReID) aims to learn from varying domains to obtain a unified person retrieval model. Existing LReID approaches typically focus on learning from scratch or a visual classification‑pretrained model, while the Vision‑Language Model (VLM) has shown generalizable knowledge in a variety of tasks. Although existing methods can be directly adapted to the VLM, since they only consider global‑aware learning, the fine‑grained attribute knowledge is underleveraged, leading to limited acquisition and anti‑forgetting capacity. To address this problem, we introduce a novel VLM‑driven LReID approach named Vision‑Language Attribute Disentanglement and Reinforcement (VLADR). Our key idea is to explicitly model the universally shared human attributes to improve inter‑domain knowledge transfer, thereby effectively utilizing historical knowledge to reinforce new knowledge learning and alleviate forgetting. Specifically, VLADR includes a Multi‑grain Text Attribute Disentanglement mechanism that mines the global and diverse local text attributes of an image. Then, an Inter‑domain Cross‑modal Attribute Reinforcement scheme is developed, which introduces cross‑modal attribute alignment to guide visual attribute extraction and adopts inter‑domain attribute alignment to achieve fine‑grained knowledge transfer. Experimental results demonstrate that our VLADR outperforms the state‑of‑the‑art methods by 1.9%‑2.2% and 2.1%‑2.5% on anti‑forgetting and generalization capacity. Our source code is available at https://github.com/zhoujiahuan1991/CVPR2026‑VLADR
Authors:Xiaolu Liu, Yicong Li, Song Wang, Junbo Chen, Angela Yao, Jianke Zhu
Abstract:
Recently, world models have been incorporated into the autonomous driving systems to improve the planning reliability. Existing approaches typically predict future states through appearance generation or deterministic regression, which limits their ability to capture trajectory‑conditioned scene evolution and leads to unreliable action planning. To address this, we propose DynFlowDrive, a latent world model that leverages flow‑based dynamics to model the transition of world states under different driving actions. By adopting the rectifiedflow formulation, the model learns a velocity field that describes how the scene state changes under different driving actions, enabling progressive prediction of future latent states. Building upon this, we further introduce a stability‑aware multi‑mode trajectory selection strategy that evaluates candidate trajectories according to the stability of the induced scene transitions. Extensive experiments on the nuScenes and NavSim benchmarks demonstrate consistent improvements across diverse driving frameworks without introducing additional inference overhead. Source code will be abaliable at https://github.com/xiaolul2/DynFlowDrive.
Authors:Daniel Ajisafe, Eric Hedlin, Helge Rhodin, Kwang Moo Yi
Abstract:
With the recent drastic advancements in text‑to‑video diffusion models, controlling their generations has drawn interest. A popular way for control is through bounding boxes or layouts. However, enforcing adherence to these control inputs is still an open problem. In this work, we show that by slightly adjusting user‑provided bounding boxes we can improve both the quality of generations and the adherence to the control inputs. This is achieved by simply optimizing the bounding boxes to better align with the internal attention maps of the video diffusion model while carefully balancing the focus on foreground and background. In a sense, we are modifying the bounding boxes to be at places where the model is familiar with. Surprisingly, we find that even with small modifications, the quality of generations can vary significantly. To do so, we propose a smooth mask to make the bounding box position differentiable and an attention‑maximization objective that we use to alter the bounding boxes. We conduct thorough experiments, including a user study to validate the effectiveness of our method. Our code is made available on the project webpage to foster future research from the community.
Authors:Yichen Zeng, Hebaixu Wang, Meng Liu, Yu Zhou, Chen Gao, Kehan Chen, Gongping Huang
Abstract:
Audio‑visual navigation enables embodied agents to navigate toward sound‑emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and leading to spatially discontinuous observations. To establish a more realistic setting, we introduce Semantic Audio‑Visual Navigation in Continuous Environments (SAVN‑CE), where agents can move freely in 3D spaces and perceive temporally and spatially coherent audio‑visual streams. In this setting, targets may intermittently become silent or stop emitting sound entirely, causing agents to lose goal information. To tackle this challenge, we propose MAGNet, a multimodal transformer‑based model that jointly encodes spatial and semantic goal representations and integrates historical context with self‑motion cues to enable memory‑augmented goal reasoning. Comprehensive experiments demonstrate that MAGNet significantly outperforms state‑of‑the‑art methods, achieving up to a 12.1% absolute improvement in success rate. These results also highlight its robustness to short‑duration sounds and long‑distance navigation scenarios. The code is available at https://github.com/yichenzeng24/SAVN‑CE.
Authors:Caiyi Sun, Yujing Sun, Xiangyu Li, Yuhang Zheng, Yiming Ren, Jiamin Wang, Yuexin Ma, Siu-Ming Yiu
Abstract:
Deepface generation has traditionally followed a task‑driven paradigm, where distinct tasks (e.g., face transfer and hair transfer) are addressed by task‑specific models. Nevertheless, this single‑task setting severely limits model generalization and scalability. A unified model capable of solving multiple deepface generation tasks in a single pass represents a promising and practical direction, yet remains challenging due to data scarcity and cross‑task conflicts arising from heterogeneous attribute transformations. To this end, we propose UniBioTransfer, the first unified framework capable of handling both conventional deepface tasks (e.g., face transfer and face reenactment) and shape‑varying transformations (e.g., hair transfer and head transfer). Besides, UniBioTransfer naturally generalizes to unseen tasks, like lip, eye, and glasses transfer, with minimal fine‑tuning. Generally, UniBioTransfer addresses data insufficiency in multi‑task generation through a unified data construction strategy, including a swapping‑based corruption mechanism designed for spatially dynamic attributes like hair. It further mitigates cross‑task interference via an innovative BioMoE, a mixture‑of‑experts based model coupled with a novel two‑stage training strategy that effectively disentangles task‑specific knowledge. Extensive experiments demonstrate the effectiveness, generalization, and scalability of UniBioTransfer, outperforming both existing unified models and task‑specific methods across a wide range of deepface generation tasks. Project page is at https://scy639.github.io/UniBioTransfer.github.io/
Authors:Yiheng Wang, Changhong Fu, Liangliang Yao, Haobo Zuo, Zijie Zhang
Abstract:
Robust feature encoding constitutes the foundation of UAV tracking by enabling the nuanced perception of target appearance and motion, thereby playing a pivotal role in ensuring reliable tracking. However, existing feature encoding methods often overlook critical illumination and viewpoint cues, which are essential for robust perception under challenging nighttime conditions, leading to degraded tracking performance. To overcome the above limitation, this work proposes a dual prompt‑driven feature encoding method that integrates prompt‑conditioned feature adaptation and context‑aware prompt evolution to promote domain‑invariant feature encoding. Specifically, the pyramid illumination prompter is proposed to extract multi‑scale frequency‑aware illumination prompts. %The dynamic viewpoint prompter adapts the sampling to different viewpoints, enabling the tracker to learn view‑invariant features. The dynamic viewpoint prompter modulates deformable convolution offsets to accommodate viewpoint variations, enabling the tracker to learn view‑invariant features. Extensive experiments validate the effectiveness of the proposed dual prompt‑driven tracker (DPTracker) in tackling nighttime UAV tracking. Ablation studies highlight the contribution of each component in DPTracker. Real‑world tests under diverse nighttime UAV tracking scenarios further demonstrate the robustness and practical utility. The code and demo videos are available at https://github.com/yiheng‑wang‑duke/DPTracker.
Authors:Shuaibang Peng, Juelin Zhu, Xia Li, Kun Yang, Maojun Zhang, Yu Liu, Shen Yan
Abstract:
We present LoD‑Loc v3, a novel method for generalized aerial visual localization in dense urban environments. While prior work LoD‑Loc v2 achieves localization through semantic building silhouette alignment with low‑detail city models, it suffers from two key limitations: poor cross‑scene generalization and frequent failure in dense building scenes. Our method addresses these challenges through two key innovations. First, we develop a new synthetic data generation pipeline that produces InsLoD‑Loc ‑ the largest instance segmentation dataset for aerial imagery to date, comprising 100k images with precise instance building annotations. This enables trained models to exhibit remarkable zero‑shot generalization capability. Second, we reformulate the localization paradigm by shifting from semantic to instance silhouette alignment, which significantly reduces pose estimation ambiguity in dense scenes. Extensive experiments demonstrate that LoD‑Loc v3 outperforms existing state‑of‑the‑art (SOTA) baselines, achieving superior performance in both cross‑scene and dense urban scenarios with a large margin. The project is available at https://nudt‑sawlab.github.io/LoD‑Locv3/.
Authors:Kaixin Cai, Pengzhen Ren, Jianhua Han, Yi Zhu, Hang Xu, Jianzhuang Liu, Xiaodan Liang
Abstract:
Open‑world semantic segmentation presently relies significantly on extensive image‑text pair datasets, which often suffer from a lack of fine‑grained pixel annotations on sufficient categories. The acquisition of such data is rendered economically prohibitive due to the substantial investments of both human labor and time. In light of the formidable image generation capabilities of diffusion models, we introduce a novel diffusion model‑driven pipeline for automatically generating datasets tailored to the needs of open‑world semantic segmentation, named "MagicSeg". Our MagicSeg initiates from class labels and proceeds to generate high‑fidelity textual descriptions, which in turn serve as guidance for the diffusion model to generate images. Rather than only generating positive samples for each label, our process encompasses the simultaneous generation of corresponding negative images, designed to serve as paired counterfactual samples for contrastive training. Then, to provide a self‑supervised signal for open‑world segmentation pretraining, our MagicSeg integrates an open‑vocabulary detection model and an interactive segmentation model to extract object masks as precise segmentation labels from images based on the provided category labels. By applying our dataset to the contrastive language‑image pretraining model with the pseudo mask supervision and the auxiliary counterfactual contrastive training, the downstream model obtains strong performance on open‑world semantic segmentation. We evaluate our model on PASCAL VOC, PASCAL Context, and COCO, achieving SOTA with performance of 62.9%, 26.7%, and 40.2%, respectively, demonstrating our dataset's effectiveness in enhancing open‑world semantic segmentation capabilities. Project website: https://github.com/ckxhp/magicseg.
Authors:Chao Wang, Xudong Tan, Jianjian Cao, Kangcong Li, Tao Chen
Abstract:
Multimodal Large Language Models have achieved significant success in offline video understanding, yet their application to streaming videos is severely limited by the linear explosion of visual tokens, which often leads to Out‑of‑Memory (OOM) errors or catastrophic forgetting. Existing visual retention and memory management methods typically rely on uniform sampling, low‑level physical metrics, or passive cache eviction. However, these strategies often lack intrinsic semantic awareness, potentially disrupting contextual coherence and blurring transient yet critical semantic transitions. To address these limitations, we propose CurveStream, a training‑free, curvature‑aware hierarchical visual memory management framework. Our approach is motivated by the key observation that high‑curvature regions along continuous feature trajectories closely align with critical global semantic transitions. Based on this geometric insight, CurveStream evaluates real‑time semantic intensity via a Curvature Score and integrates an online K‑Sigma dynamic threshold to adaptively route frames into clear and fuzzy memory states under a strict token budget. Evaluations across diverse temporal scales confirm that this lightweight framework, CurveStream, consistently yields absolute performance gains of over 10% (e.g., 10.69% on StreamingBench and 13.58% on OVOBench) over respective baselines, establishing new state‑of‑the‑art results for streaming video perception.The code will be released at https://github.com/streamingvideos/CurveStream.
Authors:Minghe Xu, Rouying Wu, ChiaWei Chu, Xiao Wang, Yu Li
Abstract:
Event‑based pedestrian attribute recognition (PAR) leverages motion cues to enhance RGB cameras in low‑light and motion‑blur scenarios, enabling more accurate inference of attributes like age and emotion. However, existing two‑stream multimodal fusion methods introduce significant computational overhead and neglect the valuable guidance from contextual samples. To address these limitations, this paper proposes an Event Prompter. Discarding the computationally expensive auxiliary backbone, this module directly applies extremely lightweight and efficient Discrete Cosine Transform (DCT) and Inverse DCT (IDCT) operations to the event data. This design extracts frequency‑domain event features at a minimal computational cost, thereby effectively augmenting the RGB branch. Furthermore, an external memory bank designed to provide rich prior knowledge, combined with modern Hopfield networks, enables associative memory‑augmented representation learning. This mechanism effectively mines and leverages global relational knowledge across different samples. Finally, a cross‑attention mechanism fuses the RGB and event modalities, followed by feed‑forward networks for attribute prediction. Extensive experiments on multiple benchmark datasets fully validate the effectiveness of the proposed RGB‑Event PAR framework. The source code of this paper will be released on https://github.com/Event‑AHU/OpenPAR
Authors:Haoyu Zhang, Zhihao Yu, Rui Wang, Yaochu Jin, Qiqi Liu, Ran Cheng
Abstract:
Modern computer vision requires balancing predictive accuracy with real‑time efficiency, yet the high inference cost of large vision models (LVMs) limits deployment on resource‑constrained edge devices. Although Evolutionary Neural Architecture Search (ENAS) is well suited for multi‑objective optimization, its practical use is hindered by two issues: expensive candidate evaluation and ranking inconsistency among subnetworks. To address them, we propose EvoNAS, an efficient distributed framework for multi‑objective evolutionary architecture search. We build a hybrid supernet that integrates Vision State Space and Vision Transformer (VSS‑ViT) modules, and optimize it with a Cross‑Architecture Dual‑Domain Knowledge Distillation (CA‑DDKD) strategy. By coupling the computational efficiency of VSS blocks with the semantic expressiveness of ViT modules, CA‑DDKD improves the representational capacity of the shared supernet and enhances ranking consistency, enabling reliable fitness estimation during evolution without extra fine‑tuning. To reduce the cost of large‑scale validation, we further introduce a Distributed Multi‑Model Parallel Evaluation (DMMPE) framework based on GPU resource pooling and asynchronous scheduling. Compared with conventional data‑parallel evaluation, DMMPE improves efficiency by over 70% through concurrent multi‑GPU, multi‑model execution. Experiments on COCO, ADE20K, KITTI, and NYU‑Depth v2 show that the searched architectures, termed EvoNets, consistently achieve Pareto‑optimal trade‑offs between accuracy and efficiency. Compared with representative CNN‑, ViT‑, and Mamba‑based models, EvoNets deliver lower inference latency and higher throughput under strict computational budgets while maintaining strong generalization on downstream tasks such as novel view synthesis. Code is available at https://github.com/EMI‑Group/evonas
Authors:Xiao Fang, Yiming Gong, Stanislav Panev, Celso de Melo, Shuowen Hu, Shayok Chakraborty, Fernando De la Torre
Abstract:
Deep neural networks (DNNs) have achieved remarkable success in computer vision but remain highly vulnerable to adversarial attacks. Among them, camouflage attacks manipulate an object's visible appearance to deceive detectors while remaining stealthy to humans. In this paper, we propose a new framework that formulates vehicle camouflage attacks as a conditional image‑editing problem. Specifically, we explore both image‑level and scene‑level camouflage generation strategies, and fine‑tune a ControlNet to synthesize camouflaged vehicles directly on real images. We design a unified objective that jointly enforces vehicle structural fidelity, style consistency, and adversarial effectiveness. Extensive experiments on the COCO and LINZ datasets show that our method achieves significantly stronger attack effectiveness, leading to more than 38% AP50 decrease, while better preserving vehicle structure and improving human‑perceived stealthiness compared to existing approaches. Furthermore, our framework generalizes effectively to unseen black‑box detectors and exhibits promising transferability to the physical world. Project page is available at https://humansensinglab.github.io/CtrlCamo
Authors:Ufaq Khan, L. D. M. S. Sai Teja, Ayuba Shakiru, Mai A. Shaaban, Yutong Xie, Muhammad Bilal, Muhammad Haris Khan
Abstract:
Ultrasound images vary widely across scanners, operators, and anatomical targets, which often causes models trained in one setting to generalize poorly to new hospitals and clinical conditions. The Foundation Model Challenge for Ultrasound Image Analysis (FMC‑UIA) reflects this difficulty by requiring a single model to handle multiple tasks, including segmentation, detection, classification, and landmark regression across diverse organs and datasets. We propose a unified multi‑task framework based on a transformer visual encoder from the Qwen3‑VL family. Intermediate token features are projected into spatial feature maps and fused using a lightweight multi‑scale feature pyramid, enabling both pixel‑level predictions and global reasoning within a shared representation. Each task is handled by a small task‑specific prediction head, while training uses task‑aware sampling and selective loss balancing to manage heterogeneous supervision and reduce task imbalance. Our method is designed to be simple to optimize and adaptable across a wide range of ultrasound analysis tasks. The performance improved from 67% to 85% on the validation set and achieved an average score of 81.84% on the official test set across all tasks. The code is publicly available at: https://github.com/saitejalekkala33/FMCUIA‑ISBI.git
Authors:Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia, Yumeng Zhang, Xiaofan Li, Xiao Tan, Xiang Bai
Abstract:
While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine‑grained geometric reasoning and physical dynamics. Existing solutions typically rely on explicit 3D modalities or complex geometric scaffolding, which are limited by data scarcity and generalization challenges. In this work, we propose a paradigm shift by leveraging the implicit spatial prior within large‑scale video generation models. We posit that to synthesize temporally coherent videos, these models inherently learn robust 3D structural priors and physical laws. We introduce VEGA‑3D (Video Extracted Generative Awareness), a plug‑and‑play framework that repurposes a pre‑trained video diffusion model as a Latent World Simulator. By extracting spatiotemporal features from intermediate noise levels and integrating them with semantic representations via a token‑level adaptive gated fusion mechanism, we enrich MLLMs with dense geometric cues without explicit 3D supervision. Extensive experiments across 3D scene understanding, spatial reasoning, and embodied manipulation benchmarks demonstrate that our method outperforms state‑of‑the‑art baselines, validating that generative priors provide a scalable foundation for physical‑world understanding. Code is publicly available at https://github.com/H‑EmbodVis/VEGA‑3D.
Authors:Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Jeffrey Hu, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Josef Bengtson, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos, Cengiz Oztireli
Abstract:
The ability to render scenes at adjustable fidelity from a single model, known as level of detail (LoD), is crucial for practical deployment of 3D Gaussian Splatting (3DGS). Existing discrete LoD methods expose only a limited set of operating points, while concurrent continuous LoD approaches enable smoother scaling but often suffer noticeable quality degradation at full capacity, making LoD a costly design decision. We introduce Matryoshka Gaussian Splatting (MGS), a training framework that enables continuous LoD for standard 3DGS pipelines without sacrificing full‑capacity rendering quality. MGS learns a single ordered set of Gaussians such that rendering any prefix, the first k splats, produces a coherent reconstruction whose fidelity improves smoothly with increasing budget. Our key idea is stochastic budget training: each iteration samples a random splat budget and optimises both the corresponding prefix and the full set. This strategy requires only two forward passes and introduces no architectural modifications. Experiments across four benchmarks and six baselines show that MGS matches the full‑capacity performance of its backbone while enabling a continuous speed‑quality trade‑off from a single model. Extensive ablations on ordering strategies, training objectives, and model capacity further validate the designs.
Authors:Yuqing Wang, Chuofan Ma, Zhijie Lin, Yao Teng, Lijun Yu, Shuai Wang, Jiaming Han, Jiashi Feng, Yi Jiang, Xihui Liu
Abstract:
Visual generation with discrete tokens has gained significant attention as it enables a unified token prediction paradigm shared with language models, promising seamless multimodal architectures. However, current discrete generation methods remain limited to low‑dimensional latent tokens (typically 8‑32 dims), sacrificing the semantic richness essential for understanding. While high‑dimensional pretrained representations (768‑1024 dims) could bridge this gap, their discrete generation poses fundamental challenges. In this paper, we present Cubic Discrete Diffusion (CubiD), the first discrete generation model for high‑dimensional representations. CubiD performs fine‑grained masking throughout the high‑dimensional discrete representation ‑‑ any dimension at any position can be masked and predicted from partial observations. This enables the model to learn rich correlations both within and across spatial positions, with the number of generation steps fixed at T regardless of feature dimensionality, where T \ll hwd. On ImageNet‑256, CubiD achieves state‑of‑the‑art discrete generation with strong scaling behavior from 900M to 3.7B parameters. Crucially, we validate that these discretized tokens preserve original representation capabilities, demonstrating that the same discrete tokens can effectively serve both understanding and generation tasks. We hope this work will inspire future research toward unified multimodal architectures. Code is available at: https://github.com/YuqingWang1029/CubiD.
Authors:Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, Ziwei Liu
Abstract:
Prior motion generation largely follows two paradigms: continuous diffusion models that excel at kinematic control, and discrete token‑based generators that are effective for semantic conditioning. To combine their strengths, we propose a three‑stage framework comprising condition feature extraction (Perception), discrete token generation (Planning), and diffusion‑based motion synthesis (Control). Central to this framework is MoTok, a diffusion‑based discrete motion tokenizer that decouples semantic abstraction from fine‑grained reconstruction by delegating motion recovery to a diffusion decoder, enabling compact single‑layer tokens while preserving motion fidelity. For kinematic conditions, coarse constraints guide token generation during planning, while fine‑grained constraints are enforced during control through diffusion‑based optimization. This design prevents kinematic details from disrupting semantic token planning. On HumanML3D, our method significantly improves controllability and fidelity over MaskControl while using only one‑sixth of the tokens, reducing trajectory error from 0.72 cm to 0.08 cm and FID from 0.083 to 0.029. Unlike prior methods that degrade under stronger kinematic constraints, ours improves fidelity, reducing FID from 0.033 to 0.014.
Authors:Dong Zhuo, Wenzhao Zheng, Sicheng Zuo, Siming Yan, Lu Hou, Jie Zhou, Jiwen Lu
Abstract:
With the growing adoption of vision‑language‑action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visual modality. However, most existing tokenizers are designed for monocular and 2D scenes, leading to inefficiency and inter‑view inconsistency when applied to high‑resolution multi‑view driving scenes. To address this, we propose DriveTok, an efficient 3D driving scene tokenizer for unified multi‑view reconstruction and understanding. DriveTok first obtains semantically rich visual features from vision foundation models and then transforms them into the scene tokens with 3D deformable cross‑attention. For decoding, we employ a multi‑view transformer to reconstruct multi‑view features from the scene tokens and use multiple heads to obtain RGB, depth, and semantic reconstructions. We also add a 3D head directly on the scene tokens for 3D semantic occupancy prediction for better spatial awareness. With the multiple training objectives, DriveTok learns unified scene tokens that integrate semantic, geometric, and textural information for efficient multi‑view tokenization. Extensive experiments on the widely used nuScenes dataset demonstrate that the scene tokens from DriveTok perform well on image reconstruction, semantic segmentation, depth prediction, and 3D occupancy prediction tasks.
Authors:Keda Tao, Yuhua Zheng, Jia Xu, Wenjie Du, Kele Shao, Hesong Wang, Xueyi Chen, Xin Jin, Junhan Zhu, Bohan Yu, Weiqiang Wang, Jian Liu, Can Qin, Yulun Zhang, Ming-Hsuan Yang, Huan Wang
Abstract:
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds to 5 minutes, failing to reflect the demands of real‑world applications, where videos typically run for tens of minutes. To address this critical gap, we introduce LVOmniBench, a new benchmark designed specifically for the cross‑modal comprehension of long‑form audio and video. This dataset comprises high‑quality videos sourced from open platforms that feature rich audio‑visual dynamics. Through rigorous manual selection and annotation, LVOmniBench comprises 275 videos, ranging in duration from 10 to 90 minutes, and 1,014 question‑answer (QA) pairs. LVOmniBench aims to rigorously evaluate the capabilities of OmniLLMs across domains, including long‑term memory, temporal localization, fine‑grained understanding, and multimodal perception. Our extensive evaluation reveals that current OmniLLMs encounter significant challenges when processing extended audio‑visual inputs. Open‑source models generally achieve accuracies below 35%, whereas the Gemini 3 Pro reaches a peak accuracy of approximately 65%. We anticipate that this dataset, along with our empirical findings, will stimulate further research and the development of advanced models capable of resolving complex cross‑modal understanding problems within long‑form audio‑visual contexts.
Authors:Shang-Jui Ray Kuo, Paola Cascante-Bonilla
Abstract:
Large vision‑‑language models (VLMs) often use a frozen vision backbone, whose image features are mapped into a large language model through a lightweight connector. While transformer‑based encoders are the standard visual backbone, we ask whether state space model (SSM) vision backbones can be a strong alternative. We systematically evaluate SSM vision backbones for VLMs in a controlled setting. Under matched ImageNet‑1K initialization, the SSM backbone achieves the strongest overall performance across both VQA and grounding/localization. We further adapt both SSM and ViT‑family backbones with detection or segmentation training and find that dense‑task tuning generally improves performance across families; after this adaptation, the SSM backbone remains competitive while operating at a substantially smaller model scale. We further observe that (i) higher ImageNet accuracy or larger backbones do not reliably translate into better VLM performance, and (ii) some visual backbones are unstable in localization. Based on these findings, we propose stabilization strategies that improve robustness for both backbone families and highlight SSM backbones as a strong alternative to transformer‑based vision encoders in VLMs.
Authors:Wan-Cyuan Fan, Jiayun Luo, Declan Kutscher, Leonid Sigal, Ritwik Gupta
Abstract:
Vision‑Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In this work, we demonstrate that VLMs are selectively blind. They modulate the amount of attention applied to visual inputs based on linguistic framing even when alternative framings demand identical visual reasoning. Using visual attention as a probe, we quantify how framing alters both the amount and distribution of attention over the image. Constrained framings, such as multiple choice and yes/no, induce substantially lower attention to image context compared to open‑ended, reduce focus on task‑relevant regions, and shift attention towards uninformative tokens. We further demonstrate that this attention misallocation is the principal cause of degraded accuracy and cross‑framing inconsistency. Building on this mechanistic insight, we introduce a lightweight prompt‑tuning method using learnable tokens that encourages the robust, visually grounded attention patterns observed in open‑ended settings, improving visual grounding and improving performance across framings.
Authors:Yuxiang Lu, Zhe Liu, Xianzhe Fan, Zhenya Yang, Jinghua Hou, Junyi Li, Kaixin Ding, Hengshuang Zhao
Abstract:
Real‑time execution is crucial for deploying Vision‑Language‑Action (VLA) models in the physical world. Existing asynchronous inference methods primarily optimize trajectory smoothness, but neglect the critical latency in reacting to environmental changes. By rethinking the notion of reaction in action chunking policies, this paper presents a systematic analysis of the factors governing reaction time. We show that reaction time follows a uniform distribution determined jointly by the Time to First Action (TTFA) and the execution horizon. Moreover, we reveal that the standard practice of applying a constant schedule in flow‑based VLAs can be inefficient and forces the system to complete all sampling steps before any movement can start, forming the bottleneck in reaction latency. To overcome this issue, we propose Fast Action Sampling for ImmediaTE Reaction (FASTER). By introducing a Horizon‑Aware Schedule, FASTER adaptively prioritizes near‑term actions during flow sampling, compressing the denoising of the immediate reaction by tenfold (e.g., in π_0.5 and X‑VLA) into a single step, while preserving the quality of long‑horizon trajectory. Coupled with a streaming client‑server pipeline, FASTER substantially reduces the effective reaction latency on real robots, especially when deployed on consumer‑grade GPUs. Real‑world experiments, including a highly dynamic table tennis task, prove that FASTER unlocks substantially improved real‑time responsiveness for generalist policies, enabling rapid generation of accurate and smooth trajectories.
Authors:Yiren Lu, Xin Ye, Burhaneddin Yaman, Jingru Luo, Zhexiao Xiong, Liu Ren, Yu Yin
Abstract:
Bird's‑Eye‑View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding‑view images to enable reasoning for various downstream tasks, such as semantic segmentation, 3D object detection, and motion prediction. However, most existing BEV perception frameworks adopt an end‑to‑end training paradigm, where image features are directly transformed into the BEV space and optimized solely through downstream task supervision. This formulation treats the entire perception process as a black box, often lacking explicit 3D geometric understanding and interpretability, leading to suboptimal performance. In this paper, we claim that an explicit 3D representation matters for accurate BEV perception, and we propose Splat2BEV, a Gaussian Splatting‑assisted framework for BEV tasks. Splat2BEV aims to learn BEV feature representations that are both semantically rich and geometrically precise. We first pre‑train a Gaussian generator that explicitly reconstructs 3D scenes from multi‑view inputs, enabling the generation of geometry‑aligned feature representations. These representations are then projected into the BEV space to serve as inputs for downstream tasks. Extensive experiments on nuScenes and argoverse dataset demonstrate that Splat2BEV achieves state‑of‑the‑art performance and validate the effectiveness of incorporating explicit 3D reconstruction into BEV perception.
Authors:Amandine Brunetto
Abstract:
Generating audio that is acoustically consistent with a scene is essential for immersive virtual environments. Recent neural acoustic field methods enable spatially continuous sound rendering but remain scene‑specific, requiring dense audio measurements and costly training for each environment. Few‑shot approaches improve scalability across rooms but still rely on multiple recordings and, being deterministic, fail to capture the inherent uncertainty of scene acoustics under sparse context. We introduce flow‑matching acoustic generation (FLAC), a probabilistic method for few‑shot acoustic synthesis that models the distribution of plausible room impulse responses (RIRs) given minimal scene context. FLAC leverages a diffusion transformer trained with a flow‑matching objective to generate RIRs at arbitrary positions in novel scenes, conditioned on spatial, geometric, and acoustic cues. FLAC outperforms state‑of‑the‑art eight‑shot baselines with one‑shot on both the AcousticRooms and Hearing Anything Anywhere datasets. To complement standard perceptual metrics, we further introduce AGREE, a joint acoustic‑geometry embedding, enabling geometry‑consistent evaluation of generated RIRs through retrieval and distributional metrics. This work is the first to apply generative flow matching to explicit RIR synthesis, establishing a new direction for robust and data‑efficient acoustic synthesis.
Authors:Swagat Padhan, Lakshya Jain, Bhavya Minesh Shah, Omkar Patil, Thao Nguyen, Nakul Gopalan
Abstract:
Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the right of the fridge" requires grounding semantic references, spatial relations, and metric constraints within a 3D scene. While recent vision language models (VLMs) demonstrate strong semantic grounding capabilities, they are not explicitly designed to reason about metric constraints in physically defined spaces. In this work, we empirically demonstrate that state‑of‑the‑art VLM‑based grounding approaches struggle with complex metric‑semantic language queries. To address this limitation, we propose MAPG (Multi‑Agent Probabilistic Grounding), an agentic framework that decomposes language queries into structured subcomponents and queries a VLM to ground each component. MAPG then probabilistically composes these grounded outputs to produce metrically consistent, actionable decisions in 3D space. We evaluate MAPG on the HM‑EQA benchmark and show consistent performance improvements over strong baselines. Furthermore, we introduce a new benchmark, MAPG‑Bench, specifically designed to evaluate metric‑semantic goal grounding, addressing a gap in existing language grounding evaluations. We also present a real‑world robot demonstration showing that MAPG transfers beyond simulation when a structured scene representation is available.
Authors:Yiren Lu, Yi Du, Disheng Liu, Yunlai Zhou, Chen Wang, Yu Yin
Abstract:
Effective embodied exploration requires agents to accumulate and retain spatial knowledge over time. However, existing scene representations, such as discrete scene graphs or static view‑based snapshots, lack post‑hoc re‑observability. If an initial observation misses a target, the resulting memory omission is often irrecoverable. To bridge this gap, we propose GSMem, a zero‑shot embodied exploration and reasoning framework built upon 3D Gaussian Splatting (3DGS). By explicitly parameterizing continuous geometry and dense appearance, 3DGS serves as a persistent spatial memory that endows the agent with Spatial Recollection: the ability to render photorealistic novel views from optimal, previously unoccupied viewpoints. To operationalize this, GSMem employs a retrieval mechanism that simultaneously leverages parallel object‑level scene graphs and semantic‑level language fields. This complementary design robustly localizes target regions, enabling the agent to ``hallucinate'' optimal views for high‑fidelity Vision‑Language Model (VLM) reasoning. Furthermore, we introduce a hybrid exploration strategy that combines VLM‑driven semantic scoring with a 3DGS‑based coverage objective, balancing task‑aware exploration with geometric coverage. Extensive experiments on embodied question answering and lifelong navigation demonstrate the robustness and effectiveness of our framework
Authors:Yuqiang Lin, Kehua Chen, Sam Lockyer, Arjun Yadav, Mingxuan Sui, Shucheng Zhang, Yan Shi, Bingzhang Wang, Yuang Zhang, Markus Zarbock, Florain Stanek, Adrian Evans, Wenbin Li, Yinhai Wang, Nic Zhang
Abstract:
Traffic Anomaly Understanding (TAU) is important for traffic safety in Intelligent Transportation Systems. Recent vision‑language models (VLMs) have shown strong capabilities in video understanding. However, progress on TAU remains limited due to the lack of benchmarks and task‑specific methodologies. To address this limitation, we introduce Roundabout‑TAU, a dataset constructed from real‑world roundabout videos collected in collaboration with the City of Carmel, Indiana. The dataset contains 342 clips and is annotated with more than 2,000 question‑answer pairs covering multiple aspects of traffic anomaly understanding. Building on this benchmark, we propose TAU‑R1, a two‑layer vision‑language framework for TAU. The first layer is a lightweight anomaly classifier that performs coarse anomaly categorisation, while the second layer is a larger anomaly reasoner that generates detailed event summaries. To improve task‑specific reasoning, we introduce a two‑stage training strategy consisting of decomposed‑QA‑enhanced supervised fine‑tuning followed by TAU‑GRPO, a GRPO‑based post‑training method with TAU‑specific reward functions. Experimental results show that TAU‑R1 achieves strong performance on both anomaly classification and reasoning tasks while maintaining deployment efficiency. The dataset and code are available at: https://github.com/siri‑rouser/TAU‑R1
Authors:Ye Wang, Wei Lu, Zhihui You, Keyan Chen, Tongfei Liu, Kaiyu Li, Hongruixuan Chen, Qingling Shu, Sibao Chen
Abstract:
Change detection in optical remote sensing imagery is susceptible to illumination fluctuations, seasonal changes, and variations in surface land‑cover materials. Relying solely on RGB imagery often produces pseudo‑changes and leads to semantic ambiguity in features. Incorporating near‑infrared (NIR) information provides heterogeneous physical cues that are complementary to visible light, thereby enhancing the discriminability of building materials and tiny structures while improving detection accuracy. However, existing multi‑modal datasets generally lack high‑resolution and accurately registered bi‑temporal imagery, and current methods often fail to fully exploit the inherent heterogeneity between these modalities. To address these issues, we introduce the Large‑scale Small‑change Multi‑modal Dataset (LSMD), a bi‑temporal RGB‑NIR building change detection benchmark dataset targeting small changes in realistic scenarios, providing a rigorous testing platform for evaluating multi‑modal change detection methods in complex environments. Based on LSMD, we further propose the Multi‑modal Spectral Complementarity Network (MSCNet) to achieve effective cross‑modal feature fusion. MSCNet comprises three key components: the Neighborhood Context Enhancement Module (NCEM) to strengthen local spatial details, the Cross‑modal Alignment and Interaction Module (CAIM) to enable deep interaction between RGB and NIR features, and the Saliency‑aware Multisource Refinement Module (SMRM) to progressively refine fused features. Extensive experiments demonstrate that MSCNet effectively leverages multi‑modal information and consistently outperforms existing methods under multiple input configurations, validating its efficacy for fine‑grained building change detection. The source code will be made publicly available at: https://github.com/AeroVILab‑AHU/LSMD
Authors:Moyang Li, Zihan Zhu, Marc Pollefeys, Daniel Barath
Abstract:
We present a robust, real‑time RGB SLAM system that handles dynamic environments by leveraging differentiable Uncertainty‑aware Bundle Adjustment. Traditional SLAM methods typically assume static scenes, leading to tracking failures in the presence of motion. Recent dynamic SLAM approaches attempt to address this challenge using predefined dynamic priors or uncertainty‑aware mapping, but they remain limited when confronted with unknown dynamic objects or highly cluttered scenes where geometric mapping becomes unreliable. In contrast, our method estimates per‑pixel uncertainty by exploiting multi‑view visual feature inconsistency, enabling robust tracking and reconstruction even in real‑world environments. The proposed system achieves state‑of‑the‑art camera poses and scene geometry in cluttered dynamic scenarios while running in real time at around 10 FPS. Code and datasets are available at https://github.com/MoyangLi00/DROID‑W.git.
Authors:Weijia Dou, Wenzhao Zheng, Weiliang Chen, Yu Zheng, Jie Zhou, Jiwen Lu
Abstract:
Recent generative models can produce high‑fidelity videos, yet they often exhibit 3D spatial geometric inconsistencies. Existing evaluation methods fail to accurately characterize these inconsistencies: fidelity‑centric metrics like FVD are insensitive to geometric distortions, while consistency‑focused benchmarks often penalize valid foreground dynamics. To address this gap, we introduce SGC, a metric for evaluating 3D Spatial Geometric Consistency in dynamically generated videos. We quantify geometric consistency by measuring the divergence among multiple camera poses estimated from distinct local regions. Our approach first separates static from dynamic regions, then partitions the static background into spatially coherent sub‑regions. We predict depth for each pixel, estimate a local camera pose for each subregion, and compute the divergence among these poses to quantify geometric consistency. Experiments on real and generative videos demonstrate that SGC robustly quantifies geometric inconsistencies, effectively identifying critical failures missed by existing metrics.
Authors:Telang Xu, Chaoyang Zhang, Guangtao Zhai, Xiaohong Liu
Abstract:
Single image reflection removal (SIRR) is challenging in real scenes, where reflection strength varies spatially and reflection patterns are tightly entangled with transmission structures. This paper presents a diffusion model with prior modulation framework (FUMO) that introduces explicit guidance signals to improve spatial controllability and structural faithfulness. Two priors are extracted directly from the mixed image, an intensity prior that estimates spatial reflection severity and a high‑frequency prior that captures detail‑sensitive responses via multi‑scale residual aggregation. We propose a coarse‑to‑fine training paradigm. In the first stage, these cues are combined to gate the conditional residual injections, focusing the conditioning on regions that are both reflection‑dominant and structure‑sensitive. In the second stage, a fine‑grained refinement network corrects local misalignment and sharpens fine details in the image space. Experiments conducted on both standard benchmarks and challenging images in the wild demonstrate competitive quantitative results and consistently improved perceptual quality. The code is released at https://github.com/Lucious‑Desmon/FUMO.
Authors:Anqi Zhang, Xiaokang Ji, Guangyu Gao, Jianbo Jiao, Chi Harold Liu, Yunchao Wei
Abstract:
Recent segmentation methods leveraging Multi‑modal Large Language Models (MLLMs) have shown reliable object‑level segmentation and enhanced spatial perception. However, almost all previous methods predominantly rely on specialist mask decoders to interpret masks from generated segmentation‑related embeddings and visual features, or incorporate multiple additional tokens to assist. This paper aims to investigate whether and how we can unlock segmentation from MLLM itSELF with 1 segmentation Embedding (SELF1E) while achieving competitive results, which eliminates the need for external decoders. To this end, our approach targets the fundamental limitation of resolution reduction in pixel‑shuffled image features from MLLMs. First, we retain image features at their original uncompressed resolution, and refill them with residual features extracted from MLLM‑processed compressed features, thereby improving feature precision. Subsequently, we integrate pixel‑unshuffle operations on image features with and without LLM processing, respectively, to unleash the details of compressed features and amplify the residual features under uncompressed resolution, which further enhances the resolution of refilled features. Moreover, we redesign the attention mask with dual perception pathways, i.e., image‑to‑image and image‑to‑segmentation, enabling rich feature interaction between pixels and the segmentation token. Comprehensive experiments across multiple segmentation tasks validate that SELF1E achieves performance competitive with specialist mask decoder‑based methods, demonstrating the feasibility of decoder‑free segmentation in MLLMs. Project page: https://github.com/ANDYZAQ/SELF1E.
Authors:Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini, Abdalla Arafa, Jameel Malik, Ahmed Elhayek, Didier Stricker
Abstract:
Understanding realistic hand‑object interactions from monocular RGB videos is essential for AR/VR, robotics, and embodied AI. Existing methods rely on category‑specific templates or heavy computation, yet still produce physically inconsistent hand‑object alignment in 3D. We introduce GHOST (Gaussian Hand‑Object Splatting), a fast, category‑agnostic framework for reconstructing dynamic hand‑object interactions using 2D Gaussian Splatting. GHOST represents both hands and objects as dense, view‑consistent Gaussian discs and introduces three key innovations: (1) a geometric‑prior retrieval and consistency loss that completes occluded object regions, (2) a grasp‑aware alignment that refines hand translations and object scale to ensure realistic contact, and (3) a hand‑aware background loss that prevents penalizing hand‑occluded object regions. GHOST achieves complete, physically consistent, and animatable reconstructions from a single RGB video while running an order of magnitude faster than prior category‑agnostic methods. Extensive experiments on ARCTIC, HO3D, and in‑the‑wild datasets demonstrate state‑of‑the‑art accuracy in 3D reconstruction and 2D rendering quality, establishing GHOST as an efficient and robust solution for realistic hand‑object interaction modeling. Code is available at https://github.com/ATAboukhadra/GHOST.
Authors:Yitong Li, Igor Yakushev, Dennis M. Hedderich, Christian Wachinger
Abstract:
Positron emission tomography (PET) is a widely recognized technique for diagnosing neurodegenerative diseases, offering critical functional insights. However, its high costs and radiation exposure hinder its widespread use. In contrast, magnetic resonance imaging (MRI) does not involve such limitations. While MRI also detects neurodegenerative changes, it is less sensitive for diagnosis compared to PET. To overcome such limitations, one approach is to generate synthetic PET from MRI. Recent advances in generative models have paved the way for cross‑modality medical image translation; however, existing methods largely emphasize structural preservation while neglecting the critical need for pathology awareness. To address this gap, we propose PASTA, a novel image translation framework built on conditional diffusion models with enhanced pathology awareness. PASTA surpasses state‑of‑the‑art methods by preserving both structural and pathological details through its highly interactive dual‑arm architecture and multi‑modal condition integration. Additionally, we introduce a novel cycle exchange consistency and volumetric generation strategy that significantly enhances PASTA's ability to produce high‑quality 3D PET images. Our qualitative and quantitative results demonstrate the high quality and pathology awareness of the synthesized PET scans. For Alzheimer's diagnosis, the performance of these synthesized scans improves over MRI by 4%, almost reaching the performance of actual PET. Our code is available at https://github.com/ai‑med/PASTA.
Authors:Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang
Abstract:
Spatial reasoning is foundational for Vision‑Language Models (VLMs), particularly when deployed as Vision‑Language‑Action (VLA) agents in physical environments. However, existing benchmarks predominantly focus on elementary, single‑hop relations, neglecting the multi‑hop compositional reasoning and precise visual grounding essential for real‑world scenarios. To address this, we introduce MultihopSpatial, offering three key contributions: (1) A comprehensive benchmark designed for multi‑hop and compositional spatial reasoning, featuring 1‑ to 3‑hop complex queries across diverse spatial perspectives. (2) Acc@50IoU, a complementary metric that simultaneously evaluates reasoning and visual grounding by requiring both answer selection and precise bounding box prediction ‑ capabilities vital for robust VLA deployment. (3) MultihopSpatial‑Train, a dedicated large‑scale training corpus to foster spatial intelligence. Extensive evaluation of 37 state‑of‑the‑art VLMs yields eight key insights, revealing that compositional spatial reasoning remains a formidable challenge. Finally, we demonstrate that reinforcement learning post‑training on our corpus enhances both intrinsic VLM spatial reasoning and downstream embodied manipulation performance.
Authors:Tianci Luo, Jinpeng Wang, Shiyu Qin, Niu Lian, Yan Feng, Bin Chen, Chun Yuan, Shu-Tao Xia
Abstract:
Visual In‑Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way to extend VICL. Unfortunately, the patch‑wise fusion framework and model‑agnostic supervision hinder the exploitation of informative cues, thereby limiting performance gains. To overcome this deficiency, we introduce PromptHub, a framework that holistically strengthens multi‑prompting through locality‑aware fusion, concentration and alignment. PromptHub exploits spatial priors to capture richer contextual information, employs complementary concentration, alignment, and prediction objectives to mutually guide training, and incorporates data augmentation to further reinforce supervision. Extensive experiments on three fundamental vision tasks demonstrate the superiority of PromptHub. Moreover, we validate its universality, transferability, and robustness across out‑of‑distribution settings, and various retrieval scenarios. This work establishes a reliable locality‑aware paradigm for prompt fusion, moving beyond prior patch‑wise approaches. Code is available at https://github.com/luotc‑why/ICLR26‑PromptHub.
Authors:Bishoy Galoaa, Shayda Moezzi, Xiangyu Bai, Sarah Ostadabbas
Abstract:
Recent video reasoning models increasingly produce spatio‑temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grounding \emphwhere and \emphwhen evidence appears, they often leave the motion connecting observations, the how, implicit. This makes dynamic and trajectory‑dependent claims difficult to supervise, verify, or penalize when unsupported by the video. We formalize this missing component as Spatial‑Temporal‑Trajectory (STT) reasoning and introduce Motion‑o, a motion‑centric extension to vision‑language models (VLMs) that makes trajectories explicit and verifiable. Motion‑o augments evidence chains with Motion Chain of Thought (MCoT), a structured pathway that represents object motion through a discrete \texttt<motion/> tag summarizing direction, speed, and scale change. To supervise MCoT, we densify sparse spatio‑temporal annotations into object tracks and derive motion descriptors from centroid displacement and box‑area change. We then train with complementary rewards for trajectory consistency and visual grounding, including a perturbation‑based signal that penalizes motion descriptions that remain unchanged when temporal evidence is removed. Across multiple video understanding benchmarks, Motion‑o consistently improves trajectory‑faithful reasoning without architectural modifications. These results suggest that an explicit motion interface can complement existing VLM pipelines by converting implicit dynamics into verifiable evidence. Code is available at~\hrefhttps://github.com/ostadabbas/Motion‑o\faGithub\ \textttostadabbas/Motion‑o.
Authors:Xiangyu Bai, Bishoy Galoaa, Sarah Ostadabbas
Abstract:
Video question answering (VQA) with vision‑language models (VLMs) depends critically on which frames are selected from the input video, yet most systems rely on uniform or heuristic sampling that cannot be optimized for downstream answering quality. We introduce HORNet, a lightweight frame selection policy trained with Group Relative Policy Optimization (GRPO) to learn which frames a frozen VLM needs to answer questions correctly. With fewer than 1M trainable parameters, HORNet reduces input frames by up to 99% and VLM processing time by up to 93%, while improving answer quality on short‑form benchmarks (+1.7% F1 on MSVD‑QA) and achieving strong performance on temporal reasoning tasks (+7.3 points over uniform sampling on NExT‑QA). We formalize this as Select Any Frames (SAF), a task that decouples visual input curation from VLM reasoning, and show that GRPO‑trained selection generalizes better out‑of‑distribution than supervised and PPO alternatives. HORNet's policy further transfers across VLM answerers without retraining, yielding an additional 8.5% relative gain when paired with a stronger model. Evaluated across six benchmarks spanning 341,877 QA pairs and 114.2 hours of video, our results demonstrate that optimizing \emphwhat a VLM sees is a practical and complementary alternative to optimizing what it generates while improving efficiency. Code is available at https://github.com/ostadabbas/HORNet.
Authors:Hesong Li, Ziqi Wu, Ruiwen Shao, Ying Fu
Abstract:
High‑Resolution Transmission Electron Microscopy (HRTEM) enables atomic‑scale observation of nucleation dynamics, which boosts the studies of advanced solid materials. Nonetheless, due to the millisecond‑scale rapid change of nucleation, it requires short‑exposure rapid imaging, leading to severe noise that obscures atomic positions. In this work, we propose a statistical characteristic‑guided denoising network, which utilizes statistical characteristics to guide the denoising process in both spatial and frequency domains. In the spatial domain, we present spatial deviation‑guided weighting to select appropriate convolution operations for each spatial position based on deviation characteristic. In the frequency domain, we present frequency band‑guided weighting to enhance signals and suppress noise based on band characteristics. We also develop an HRTEM‑specific noise calibration method and generate a dataset with disordered structures and realistic HRTEM image noises. It can ensure the denoising performance of models on real images for nucleation observation. Experiments on synthetic and real data show our method outperforms the state‑of‑the‑art methods in HRTEM image denoising, with effectiveness in the localization downstream task. Code will be available at https://github.com/HeasonLee/SCGN.
Authors:Jiatong Xia, Zicheng Duan, Anton van den Hengel, Lingqiao Liu
Abstract:
Recent progress in 3D generation has been driven largely by models conditioned on images or text, while readily available 3D priors are still underused. In many real‑world scenarios, the visible‑region point cloud are easy to obtain from active sensors such as LiDAR or from feed‑forward predictors like VGGT, offering explicit geometric constraints that current methods fail to exploit. In this work, we introduce Points‑to‑3D, a diffusion‑based framework that leverages point cloud priors for geometry‑controllable 3D asset and scene generation. Built on a latent 3D diffusion model TRELLIS, Points‑to‑3D first replaces pure‑noise sparse structure latent initialization with a point cloud priors tailored input formulation.A structure inpainting network, trained within the TRELLIS framework on task‑specific data designed to learn global structural inpainting, is then used for inference with a staged sampling strategy (structural inpainting followed by boundary refinement), completing the global geometry while preserving the visible regions of the input priors. In practice, Points‑to‑3D can take either accurate point‑cloud priors or VGGT‑estimated point clouds from single images as input. Experiments on both objects and scene scenarios consistently demonstrate superior performance over state‑of‑the‑art baselines in terms of rendering quality and geometric fidelity, highlighting the effectiveness of explicitly embedding point‑cloud priors for achieving more accurate and structurally controllable 3D generation. Project page: https://jiatongxia.github.io/points2‑3D/
Authors:Ying Zheng, Yiyi Zhang, Yi Wang, Lap-Pui Chau
Abstract:
Source‑Free Domain Adaptation (SFDA) adapts pre‑trained models to unlabeled target domains without requiring access to source data. Although state‑of‑the‑art methods leveraging local neighborhood structures show promise for SFDA, they tend to over‑rely on prediction similarity among neighbors. This over‑reliance accelerates the forgetting of source knowledge and increases susceptibility to local noise overfitting. To address these issues, we introduce ProCal, a probability calibration method that dynamically calibrates neighborhood‑based predictions through a dual‑model collaborative prediction mechanism. ProCal integrates the source model's initial predictions with the current model's online outputs to effectively calibrate neighbor probabilities. This strategy not only mitigates the interference of local noise but also preserves the discriminative information from the source model, thereby achieving a balance between knowledge retention and domain adaptation. Furthermore, we design a joint optimization objective that combines a soft supervision loss with a diversity loss to guide the target model. Our theoretical analysis shows that ProCal converges to an equilibrium where source knowledge and target information are effectively fused, reducing both knowledge forgetting and overfitting. We validate the effectiveness of our approach through extensive experiments on 31 cross‑domain tasks across four public datasets. Our code is available at: https://github.com/zhengyinghit/ProCal.
Authors:Longfei Liu, Yongjie Hou, Yang Li, Qirui Wang, Youyang Sha, Yongjun Yu, Yinzhi Wang, Peizhe Ru, Xuanlong Yu, Xi Shen
Abstract:
Deploying high‑performance dense prediction models on resource‑constrained edge devices remains challenging due to strict limits on computation and memory. In practice, lightweight systems for object detection, instance segmentation, and pose estimation are still dominated by CNN‑based architectures such as YOLO, while compact Vision Transformers (ViTs) often struggle to achieve similarly strong accuracy efficiency tradeoff, even with large scale pretraining. We argue that this gap is largely due to insufficient task specific representation learning in small scale ViTs, rather than an inherent mismatch between ViTs and edge dense prediction. To address this issue, we introduce EdgeCrafter, a unified compact ViT framework for edge dense prediction centered on ECDet, a detection model built from a distilled compact backbone and an edge‑friendly encoder decoder design. On the COCO dataset, ECDet‑S achieves 51.7 AP with fewer than 10M parameters using only COCO annotations. For instance segmentation, ECInsSeg achieves performance comparable to RF‑DETR while using substantially fewer parameters. For pose estimation, ECPose‑X reaches 74.8 AP, significantly outperforming YOLO26Pose‑X (71.6 AP). These results show that compact ViTs, when paired with task‑specialized distillation and edge‑aware design, can be a practical and competitive option for edge dense prediction. Code is available at: https://intellindust‑ai‑lab.github.io/projects/EdgeCrafter/
Authors:Juan Miguel Valverde, Dim P. Papadopoulos, Rasmus Larsen, Anders Bjorholm Dahl
Abstract:
Standard deep learning models for image segmentation cannot guarantee topology accuracy, failing to preserve the correct number of connected components or structures. This, in turn, affects the quality of the segmentations and compromises the reliability of the subsequent quantification analyses. Previous works have proposed to enhance topology accuracy with specialized frameworks, architectures, and loss functions. However, these methods are often cumbersome to integrate into existing training pipelines, they are computationally very expensive, or they are restricted to structures with tubular morphology. We present SCNP, an efficient method that improves topology accuracy by penalizing the logits with their poorest‑classified neighbor, forcing the model to improve the prediction at the pixels' neighbors before allowing it to improve the pixels themselves. We show the effectiveness of SCNP across 13 datasets, covering different structure morphologies and image modalities, and integrate it into three frameworks for semantic and instance segmentation. Additionally, we show that SCNP can be integrated into several loss functions, making them improve topology accuracy. Our code can be found at https://jmlipman.github.io/SCNP‑SameClassNeighborPenalization.
Authors:Jingguo Qu, Xinyang Han, Yao Pu, Man-Lik Chui, Simon Takadiyi Gunda, Ziman Chen, Jing Qin, Ann Dorothy King, Winnie Chiu-Wing Chu, Jing Cai, Michael Tin-Cheung Ying
Abstract:
Medical ultrasound image segmentation faces significant challenges due to limited labeled data and characteristic imaging artifacts including speckle noise and low‑contrast boundaries. While semi‑supervised learning (SSL) approaches have emerged to address data scarcity, existing methods suffer from suboptimal unlabeled data utilization and lack robust feature representation mechanisms. In this paper, we propose Switch, a novel SSL framework with two key innovations: (1) Multiscale Switch (MSS) strategy that employs hierarchical patch mixing to achieve uniform spatial coverage; (2) Frequency Domain Switch (FDS) with contrastive learning that performs amplitude switching in Fourier space for robust feature representations. Our framework integrates these components within a teacher‑student architecture to effectively leverage both labeled and unlabeled data. Comprehensive evaluation across six diverse ultrasound datasets (lymph nodes, breast lesions, thyroid nodules, and prostate) demonstrates consistent superiority over state‑of‑the‑art methods. At 5% labeling ratio, Switch achieves remarkable improvements: 80.04% Dice on LN‑INT, 85.52% Dice on DDTI, and 83.48% Dice on Prostate datasets, with our semi‑supervised approach even exceeding fully supervised baselines. The method maintains parameter efficiency (1.8M parameters) while delivering superior performance, validating its effectiveness for resource‑constrained medical imaging applications. The source code is publicly available at https://github.com/jinggqu/Switch
Authors:Teer Song, Yue Zhang, Yu Tian, Ziyang Wang, Xianlin Zhang, Guixuan Zhang, Xuan Liu, Xueming Li, Yasen Zhang
Abstract:
To better preserve an individual's identity, face restoration has evolved from reference‑free to reference‑based approaches, which leverage high‑quality reference images of the same identity to enhance identity fidelity in the restored outputs. However, most existing methods implicitly assume that the reference and degraded input are age‑aligned, limiting their effectiveness in real‑world scenarios where only cross‑age references are available, such as historical photo restoration. This paper proposes MeInTime, a diffusion‑based face restoration method that extends reference‑based restoration from same‑age to cross‑age settings. Given one or few reference images along with an age prompt corresponding to the degraded input, MeInTime achieves faithful restoration with both identity fidelity and age consistency. Specifically, we decouple the modeling of identity and age conditions. During training, we focus solely on effectively injecting identity features through a newly introduced attention mechanism and introduce Gated Residual Fusion modules to facilitate the integration between degraded features and identity representations. At inference, we propose Age‑Aware Gradient Guidance, a training‑free sampling strategy, using an age‑driven direction to iteratively nudge the identity‑aware denoising latent toward the desired age semantic manifold. Extensive experiments demonstrate that MeInTime outperforms existing face restoration methods in both identity preservation and age consistency. Our code is available at: https://github.com/teer4/MeInTime
Authors:Jiayi Luo, Jiayu Chen, Jiankun Wang, Cong Wang, Hanxin Zhu, Qingyun Sun, Chen Gao, Zhibo Chen, Jianxin Li
Abstract:
Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, motivating sparse attention techniques for improving efficiency. However, existing training‑free sparse attention methods for video generation still face two unresolved limitations: ignoring layer heterogeneity in attention pruning and ignoring query‑key coupling in block partitioning, which hinder a better quality‑speedup trade‑off. In this work, we uncover a critical insight: attention sparsity is an intrinsic layer‑wise property, with only minor variation across different inputs. Motivated by this observation, we propose SVOO, a training‑free sparse attention framework for fast video generation via offline layer‑wise sparsity profiling and online bidirectional co‑clustering. Specifically, SVOO adopts a two‑stage paradigm: (i) offline layer‑wise sensitivity profiling to derive intrinsic per‑layer pruning levels, and (ii) online block‑wise sparse attention via a bidirectional co‑clustering algorithm. Extensive experiments on seven widely used video generation models demonstrate that SVOO achieves a superior quality‑speedup trade‑off over state‑of‑the‑art methods, delivering up to 1.93x speedup while maintaining a PSNR of up to 29 dB on Wan2.1. Code is available at: https://github.com/Mutual‑Luo/SVOO.
Authors:Xuan Liu, Xiaobin Chang
Abstract:
Weight regularization methods in continual learning (CL) alleviate catastrophic forgetting by assessing and penalizing changes to important model weights. Elastic Weight Consolidation (EWC) is a foundational and widely used approach within this framework that estimates weight importance based on gradients. However, it has consistently shown suboptimal performance. In this paper, we conduct a systematic analysis of importance estimation in EWC from a gradient‑based perspective. For the first time, we find that EWC's reliance on the Fisher Information Matrix (FIM) results in gradient vanishing and inaccurate importance estimation in certain scenarios. Our analysis also reveals that Memory Aware Synapses (MAS), a variant of EWC, imposes unnecessary constraints on parameters irrelevant to prior tasks, termed the redundant protection. Consequently, both EWC and its variants exhibit fundamental misalignments in estimating weight importance, leading to inferior performance. To tackle these issues, we propose the Logits Reversal (LR) operation, a simple yet effective modification that rectifies EWC's importance estimation. Specifically, reversing the logit values during the calculation of FIM can effectively prevent both gradient vanishing and redundant protection. Extensive experiments across various CL tasks and datasets show that the proposed method significantly outperforms existing EWC and its variants. Therefore, we refer to it as EWC Done Right (EWC‑DR). Code is available at https://github.com/scarlet0703/EWC‑DR.
Authors:Swarnendu Banik, Manish Das, Shiv Ram Dubey, Satish Kumar Singh
Abstract:
Vision Transformers have excelled in computer vision but their attention mechanisms operate independently across layers, limiting information flow and feature learning. We propose an effective cross‑layer attention propagation method that preserves and integrates historical attention matrices across encoder layers, offering a principled refinement of inter‑layer information flow in Vision Transformers. This approach enables progressive refinement of attention patterns throughout the transformer hierarchy, enhancing feature acquisition and optimization dynamics. The method requires minimal architectural changes, adding only attention matrix storage and blending operations. Comprehensive experiments on CIFAR‑100 and TinyImageNet demonstrate consistent accuracy improvements, with ViT performance increasing from 75.74% to 77.07% on CIFAR‑100 (+1.33%) and from 57.82% to 59.07% on TinyImageNet (+1.25%). Cross‑architecture validation shows similar gains across transformer variants, with CaiT showing 1.01% enhancement. Systematic analysis identifies the blending hyperparameter of historical attention (alpha = 0.45) as optimal across all configurations, providing the ideal balance between current and historical attention information. Random initialization consistently outperforms zero initialization, indicating that diverse initial attention patterns accelerate convergence and improve final performance. Our code is publicly available at https://github.com/banik‑s/HAViT.
Authors:Xiang Zhou, Hong Shang, Zijian Zhan, Tianyu He, Jintao Meng, Dong Liang
Abstract:
Deep unrolled models (DUMs) have become the state of the art for accelerated MRI reconstruction, yet their robustness under domain shift remains a critical barrier to clinical adoption. In this work, we identify coil sensitivity map (CSM) estimation as the primary bottleneck limiting generalization. To address this, we propose UEPS, a novel DUM architecture featuring three key innovations: (i) an Unrolled Expanded (UE) design that eliminates CSM dependency by reconstructing each coil independently; (ii) progressive resolution, which leverages k‑space‑to‑image mapping for efficient coarse‑to‑fine refinement; and (iii) sparse attention tailored to MRI's 1D undersampling nature. These physics‑grounded designs enable simultaneous gains in robustness and computational efficiency. We construct a large‑scale zero‑shot transfer benchmark comprising 10 out‑of‑distribution test sets spanning diverse clinical shifts ‑‑ anatomy, view, contrast, vendor, field strength, and coil configurations. Extensive experiments demonstrate that UEPS consistently and substantially outperforms existing DUM, end‑to‑end, diffusion, and untrained methods across all OOD tests, achieving state‑of‑the‑art robustness with low‑latency inference suitable for real‑time deployment.
Authors:Hyun-kyu Ko, Jihyeon Park, Younghyun Kim, Dongheok Park, Eunbyung Park
Abstract:
Creating dynamic, view‑consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual production, and next‑generation e‑commerce. However, despite rapid progress in subject‑driven video generation, existing methods predominantly treat subjects as 2D entities, focusing on transferring identity through single‑view visual features or textual prompts. Because real‑world subjects are inherently 3D, applying these 2D‑centric approaches to 3D object customization reveals a fundamental limitation: they lack the comprehensive spatial priors necessary to reconstruct the 3D geometry. Consequently, when synthesizing novel views, they must rely on generating plausible but arbitrary details for unseen regions, rather than preserving the true 3D identity. Achieving genuine 3D‑aware customization remains challenging due to the scarcity of multi‑view video datasets. While one might attempt to fine‑tune models on limited video sequences, this often leads to temporal overfitting. To resolve these issues, we introduce a novel framework for 3D‑aware video customization, comprising 3DreamBooth and 3Dapter. 3DreamBooth decouples spatial geometry from temporal motion through a 1‑frame optimization paradigm. By restricting updates to spatial representations, it effectively bakes a robust 3D prior into the model without the need for exhaustive video‑based training. To enhance fine‑grained textures and accelerate convergence, we incorporate 3Dapter, a visual conditioning module. Following single‑view pre‑training, 3Dapter undergoes multi‑view joint optimization with the main generation branch via an asymmetrical conditioning strategy. This design allows the module to act as a dynamic selective router, querying view‑specific geometric hints from a minimal reference set. Project page: https://ko‑lani.github.io/3DreamBooth/
Authors:Mingde Zhou, Zheng Chen, Yulun Zhang
Abstract:
Video compression aims to maximize reconstruction quality with minimal bitrates. Beyond standard distortion metrics, perceptual quality and temporal consistency are also critical. However, at ultra‑low bitrates, traditional end‑to‑end compression models tend to produce blurry images of poor perceptual quality. Besides, existing generative compression methods often treat video frames independently and show limitations in time coherence and efficiency. To address these challenges, we propose the Efficient Video Diffusion with Sparse Information Transmission (Diff‑SIT), which comprises the Sparse Temporal Encoding Module (STEM) and the One‑Step Video Diffusion with Frame Type Embedder (ODFTE). The STEM sparsely encodes the original frame sequence into an information‑rich intermediate sequence, achieving significant bitrate savings. Subsequently, the ODFTE processes this intermediate sequence as a whole, which exploits the temporal correlation. During this process, our proposed Frame Type Embedder (FTE) guides the diffusion model to perform adaptive reconstruction according to different frame types to optimize the overall quality. Extensive experiments on multiple datasets demonstrate that Diff‑SIT establishes a new state‑of‑the‑art in perceptual quality and temporal consistency, particularly in the challenging ultra‑low‑bitrate regime. Code is released at https://github.com/MingdeZhou/Diff‑SIT.
Authors:Seonghyun Jin, Jong Chul Ye
Abstract:
Streaming 3D reconstruction maintains a persistent latent state that is updated online from incoming frames, enabling constant‑memory inference. A key failure mode is the state update rule: aggressive overwrites forget useful history, while conservative updates fail to track new evidence, and both behaviors become unstable beyond the training horizon. To address this challenge, we propose FILT3R, a training‑free latent filtering layer that casts recurrent state updates as stochastic state estimation in token space. FILT3R maintains a per‑token variance and computes a Kalman‑style gain that adaptively balances memory retention against new observations. Process noise ‑‑ governing how much the latent state is expected to change between frames ‑‑ is estimated online from EMA‑normalized temporal drift of candidate tokens. Using extensive experiments, we demonstrate that FILT3R yields an interpretable, plug‑in update rule that generalizes common overwrite and gating policies as special cases. Specifically, we show that gains shrink in stable regimes as uncertainty contracts with accumulated evidence, and rise when genuine scene change increases process uncertainty, improving long‑horizon stability for depth, pose, and 3D reconstruction, compared to the existing methods. Code will be released at https://github.com/jinotter3/FILT3R.
Authors:Bo Zhao, Yihang Liu, Chenfeng Zhang, Huan Yang, Kun Gai, Wei Ji
Abstract:
Text‑guided texture editing aims to modify object appearance while preserving the underlying geometric structure. However, our empirical analysis reveals that even SOTA editing models frequently struggle to maintain structural consistency during texture editing, despite the intended changes being purely appearance‑related. Motivated by this observation, we jointly enhance structure preservation from both data and training perspectives, and build TexEditor, a dedicated texture editing model based on Qwen‑Image‑Edit‑2509. Firstly, we construct TexBlender, a high‑quality SFT dataset generated with Blender, which provides strong structural priors for a cold start. Sec‑ ondly, we introduce StructureNFT, a RL‑based approach that integrates structure‑preserving losses to transfer the structural priors learned during SFT to real‑world scenes. Moreover, due to the limited realism and evaluation coverage of existing benchmarks, we introduce TexBench, a general‑purpose real‑world benchmark for text‑guided texture editing. Extensive experiments on existing Blender‑based texture benchmarks and our TexBench show that TexEditor consistently outperforms strong baselines such as Nano Banana Pro. In addition, we assess TexEditor on the general purpose benchmark ImgEdit to validate its generalization. Our code and data are available at https://github.com/KlingAIResearch/TexEditor.
Authors:Yuqi Yang, Dongliang Chang, Yijia Ling, Ruoyi Du, Zhanyu Ma
Abstract:
Colour is one of the most perceptually salient yet least controllable attributes in image generation. Although recent diffusion models can modify object colours from user instructions, their results often deviate from the intended hue, especially for fine‑grained and local edits. Early text‑driven methods rely on discrete language descriptions that cannot accurately represent continuous chromatic variations. To overcome this limitation, we propose ColourCrafter, a unified diffusion framework that transforms colour editing from global tone transfer into a structured, region‑aware generation process. Unlike traditional colour driven methods, ColourCrafter performs token‑level fusion of RGB colour tokens and image tokens in latent space, selectively propagating colour information to semantically relevant regions while preserving structural fidelity. A perceptual Lab‑space Loss further enhances pixel‑level precision by decoupling luminance and chrominance and constraining edits within masked areas. Additionally, we build ColourfulSet, a largescale dataset of high‑quality image pairs with continuous and diverse colour variations. Extensive experiments demonstrate that ColourCrafter achieves state‑of‑the‑art colour accuracy, controllability and perceptual fidelity in fine‑grained colour editing. Our project is available at https://yangyuqi317.github.io/ColourCrafter.github.io/.
Authors:Kazuya Nishimura, Ryoma Bise, Shinnosuke Matsuo, Haruka Hirose, Yasuhiro Kojima
Abstract:
Estimating slide‑ and patch‑level gene expression profiles from pathology images enables rapid and low‑cost molecular analysis with broad clinical impact. Despite strong results, existing approaches treat gene expression as a mere slide‑ or spot‑level signal and do not incorporate the fact that the measured expression arises from the aggregation of underlying cell‑level expression. To explicitly introduce this missing cell‑resolved guidance, we propose a Cell‑type Prototype‑informed Neural Network (CPNN) that leverages publicly available single‑cell RNA‑sequencing datasets. Since single‑cell measurements are noisy and not paired with histology images, we first estimate cell‑type prototypes‑mean expression profiles that reflect stable gene‑gene co‑variation patterns.CPNN then learns cell‑type compositional weights directly from images and models the relationship between prototypes and observed bulk or spatial expression, providing a biologically grounded and structurally regularized prediction framework. We evaluate CPNN on three slide‑level datasets and three patch‑level spatial transcriptomics datasets. Across all settings, CPNN achieves the highest performance in terms of Spearman correlation. Moreover, by visualizing the inferred compositional weights, our framework provides interpretable insights into which cell types drive the predicted expression. Code is publicly available at https://github.com/naivete5656/CPNN.
Authors:Leyuan Fang, Zan Mao, Zijing Wang, Yinlong Yan
Abstract:
Zero‑shot object‑goal navigation aims to find target objects in unseen environments using only egocentric observation. Recent methods leverage foundation models' comprehension and reasoning capabilities to enhance navigation performance. However, when faced with poor viewpoints or weak semantic cues, foundation models often fail to support reliable reasoning in both perception and planning, resulting in inefficient or failed navigation. We observe that inherent relationships among objects and regions encode structured scene priors, which help agents infer plausible target locations even under partial observations. Motivated by this insight, we propose Spatial Relation‑aware Navigation (SR‑Nav), a framework that models both observed and experience‑based spatial relationships to enhance both perception and planning. Specifically, SR‑Nav first constructs a Dynamic Spatial Relationship Graph (DSRG) that encodes the target‑centered spatial relationships through the foundation models and updates dynamically with real‑time observations. We then introduce a Relation‑aware Matching Module. It utilizes relationship matching instead of naive detection, leveraging diverse relationships in the DSRG to verify and correct errors, enhancing visual perception robustness. Finally, we design a Dynamic Relationship Planning Module to reduce the planning search space by dynamically computing the optimal paths based on the DSRG from the current position, thereby guiding planning and reducing exploration redundancy. Experiments on HM3D show that our method achieves state‑of‑the‑art performance in both success rate and navigation efficiency. The code will be publicly available at https://github.com/Mzyw‑1314/SR‑Nav
Authors:Yibo Shi, Jungang Li, Linghao Zhang, Zihao Dongfang, Biao Wu, Sicheng Tao, Yibo Yan, Chenxi Qin, Weiting Liu, Zhixin Lin, Hanqian Li, Yu Huang, Song Dai, Yonghua Hei, Yue Ding, Xiang Li, Shikang Wang, Chengdong Xu, Jingqi Liu, Xueying Ma, Zhiwen Zheng, Xiaofei Zhang, Bincheng Wang, Nichen Yang, Jie Wu, Lihua Tian, Chen Li, Xuming Hu
Abstract:
Long‑horizon GUI agents are a key step toward real‑world deployment, yet effective interaction memory under prevailing paradigms remains under‑explored. Replaying full interaction sequences is redundant and amplifies noise, while summaries often erase dependency‑critical information and traceability. We present AndroTMem, a diagnostic framework for anchored memory in long‑horizon Android GUI agents. Its core benchmark, AndroTMem‑Bench, comprises 1,069 tasks with 34,473 interaction steps (avg. 32.1 per task, max. 65). We evaluate agents with TCR (Task Complete Rate), focusing on tasks whose completion requires carrying forward critical intermediate state; AndroTMem‑Bench is designed to enforce strong step‑to‑step causal dependencies, making sparse yet essential intermediate states decisive for downstream actions and centering interaction memory in evaluation. Across open‑ and closed‑source GUI agents, we observe a consistent pattern: as interaction sequences grow longer, performance drops are driven mainly by within‑task memory failures, not isolated perception errors or local action mistakes. Guided by this diagnosis, we propose Anchored State Memory (ASM), which represents interaction sequences as a compact set of causally linked intermediate‑state anchors to enable subgoal‑targeted retrieval and attribution‑aware decision making. Across multiple settings and 12 evaluated GUI agents, ASM consistently outperforms full‑sequence replay and summary‑based baselines, improving TCR by 5%‑30.16% and AMS by 4.93%‑24.66%, indicating that anchored, structured memory effectively mitigates the interaction‑memory bottleneck in long‑horizon GUI tasks. The code, benchmark, and related resources are publicly available at [https://github.com/CVC2233/AndroTMem](https://github.com/CVC2233/AndroTMem).
Authors:Huy Che, Dinh-Duy Phan, Duc-Khai Lam
Abstract:
Collecting and annotating datasets for pixel‑level semantic segmentation tasks are highly labor‑intensive. Data augmentation provides a viable solution by enhancing model generalization without additional real‑world data collection. Traditional augmentation techniques, such as translation, scaling, and color transformations, create geometric variations but fail to generate new structures. While generative models have been employed to extend semantic information of datasets, they often struggle to maintain consistency between the original and generated images, particularly for pixel‑level tasks. In this work, we propose a novel synthetic data augmentation pipeline that integrates controllable diffusion models. Our approach balances diversity and reliability data, effectively bridging the gap between synthetic and real data. We utilize class‑aware prompting and visual prior blending to improve image quality further, ensuring precise alignment with segmentation labels. By evaluating benchmark datasets such as PASCAL VOC and BDD100K, we demonstrate that our method significantly enhances semantic segmentation performance, especially in data‑scarce scenarios, while improving model robustness in real‑world applications. Our code is available at \hrefhttps://github.com/chequanghuy/Enhanced‑Generative‑Data‑Augmentation‑for‑Semantic‑Segmentation‑via‑Stronger‑Guidancehttps://github.com/chequanghuy/Enhanced‑Generative‑Data‑Augmentation‑for‑Semantic‑Segmentation‑via‑Stronger‑Guidance.
Authors:Zilin Huang, Zihao Sheng, Zhengyang Wan, Yansong Qu, Junwei You, Sicong Jiang, Sikai Chen
Abstract:
Ensuring safe decision‑making in autonomous vehicles remains a fundamental challenge despite rapid advances in end‑to‑end learning approaches. Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the rich contextual understanding required for safe driving and make unsafe exploration unavoidable in real‑world settings. Recent vision‑language models (VLMs) offer promising semantic understanding capabilities; however, their high inference latency and susceptibility to hallucination hinder direct application to real‑time vehicle control. To address these limitations, this paper proposes DriveVLM‑RL, a neuroscience‑inspired framework that integrates VLMs into RL through a dual‑pathway architecture for safe and deployable autonomous driving. The framework decomposes semantic reward learning into a Static Pathway for continuous spatial safety assessment using CLIP‑based contrasting language goals, and a Dynamic Pathway for attention‑gated multi‑frame semantic risk reasoning using a lightweight detector and a large VLM. A hierarchical reward synthesis mechanism fuses semantic signals with vehicle states, while an asynchronous training pipeline decouples expensive VLM inference from environment interaction. All VLM components are used only during offline training and are removed at deployment, ensuring real‑time feasibility. Experiments in the CARLA simulator show significant improvements in collision avoidance, task success, and generalization across diverse traffic scenarios, including strong robustness under settings without explicit collision penalties. These results demonstrate that DriveVLM‑RL provides a practical paradigm for integrating foundation models into autonomous driving without compromising real‑time feasibility. Demo video and code are available at: https://zilin‑huang.github.io/DriveVLM‑RL‑website/
Authors:Mohammed Rahman Sherif Khan Mohammad, Ardhendu Behera, Sandip Pradhan, Swagat Kumar, Amr Ahmed
Abstract:
Recent adapter‑based CLIP tuning (e.g., Tip‑Adapter) is a strong few‑shot learner, achieving efficiency by caching support features for fast prototype matching. However, these methods rely on global uni‑modal feature vectors, overlooking fine‑grained patch relations and their structural alignment with class text. To bridge this gap without incurring inference costs, we introduce a novel asymmetric training‑only framework. Instead of altering the lightweight adapter, we construct a high‑capacity auxiliary Heterogeneous Graph Teacher that operates solely during training. This teacher (i) integrates multi‑scale visual patches and text prompts into a unified graph, (ii) performs deep cross‑modal reasoning via a Modality‑aware Graph Transformer (MGT), and (iii) applies discriminative node filtering to extract high‑fidelity class features. Crucially, we employ a cache‑aware dual‑objective strategy to supervise this relational knowledge directly into the Tip‑Adapter's key‑value cache, effectively upgrading the prototypes while the graph teacher is discarded at test time. Thus, inference remains identical to Tip‑Adapter with zero extra latency or memory. Across standard 1‑16‑shot benchmarks, our method consistently establishes a new state‑of‑the‑art. Ablations confirm that the auxiliary graph supervision, text‑guided reasoning, and node filtering are the essential ingredients for robust few‑shot adaptation. Code is available at https://github.com/MR‑Sherif/TOGA.git.
Authors:Haoxiang Rao, Zhao Wang, Chenyang Si, Yan Lyu, Yuanyi Duan, Fang Zhao, Caifeng Shan
Abstract:
Industrial anomaly detection (AD) is characterized by an abundance of normal images but a scarcity of anomalous ones. Although numerous few‑shot anomaly synthesis methods have been proposed to augment anomalous data for downstream AD tasks, most existing approaches require time‑consuming training and struggle to learn distributions that are faithful to real anomalies, thereby restricting the efficacy of AD models trained on such data. To address these limitations, we propose a training‑free few‑shot anomaly generation method, namely O2MAG, which leverages the self‑attention in One reference anomalous image to synthesize More realistic anomalies, supporting effective downstream anomaly detection. Specifically, O2MAG manipulates three parallel diffusion processes via self‑attention grafting and incorporates the anomaly mask to mitigate foreground‑background query confusion, synthesizing text‑guided anomalies that closely adhere to real anomalous distributions. To bridge the semantic gap between the encoded anomaly text prompts and the true anomaly semantics, Anomaly‑Guided Optimization is further introduced to align the synthesis process with the target anomalous distribution, steering the generation toward realistic and text‑consistent anomalies. Moreover, to mitigate faint anomaly synthesis inside anomaly masks, Dual‑Attention Enhancement is adopted during generation to reinforce both self‑ and cross‑attention on masked regions. Extensive experiments validate the effectiveness of O2MAG, demonstrating its superior performance over prior state‑of‑the‑art methods on downstream AD tasks.
Authors:Wei Tang, Xuejing Liu, Yanpeng Sun, Zechao Li
Abstract:
The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Expression Segmentation (RES). Toward this end, we propose SSP‑SAM, a framework that fully utilizes SAM's segmentation capabilities by integrating a Semantic‑Spatial Prompt (SSP) encoder. Specifically, we incorporate both visual and linguistic attention adapters into the SSP encoder, which highlight salient objects within the visual features and discriminative phrases within the linguistic features. This design enhances the referent representation for the prompt generator, resulting in high‑quality SSPs that enable SAM to generate precise masks guided by language. Although not specifically designed for Generalized RES (GRES), where the referent may correspond to zero, one, or multiple objects, SSP‑SAM naturally supports this more flexible setting without additional modifications. Extensive experiments on widely used RES and GRES benchmarks confirm the superiority of our method. Notably, our approach generates segmentation masks of high quality, achieving strong precision even at strict thresholds such as Pr@0.9. Further evaluation on the PhraseCut dataset demonstrates improved performance in open‑vocabulary scenarios compared to existing state‑of‑the‑art RES methods. The code and checkpoints are available at: https://github.com/WayneTomas/SSP‑SAM.
Authors:Wuqi Wang, Haochen Yang, Baolu Li, Jiaqi Sun, Xiangmo Zhao, Zhigang Xu, Qing Guo, Haigen Min, Tianyun Zhang, Hongkai Yu
Abstract:
The low‑light conditions are challenging to the vision‑centric perception systems for autonomous driving in the dark environment. In this paper, we propose a new benchmark dataset (named DarkDriving) to investigate the low‑light enhancement for autonomous driving. The existing real‑world low‑light enhancement benchmark datasets can be collected by controlling various exposures only in small‑ranges and static scenes. The dark images of the current nighttime driving datasets do not have the precisely aligned daytime counterparts. The extreme difficulty to collect a real‑world day and night aligned dataset in the dynamic driving scenes significantly limited the research in this area. With a proposed automatic day‑night Trajectory Tracking based Pose Matching (TTPM) method in a large real‑world closed driving test field (area: 69 acres), we collected the first real‑world day and night aligned dataset for autonomous driving in the dark environment. The DarkDriving dataset has 9,538 day and night image pairs precisely aligned in location and spatial contents, whose alignment error is in just several centimeters. For each pair, we also manually label the object 2D bounding boxes. DarkDriving introduces four perception related tasks, including low‑light enhancement, generalized low‑light enhancement, and low‑light enhancement for 2D detection and 3D detection of autonomous driving in the dark environment. The experimental results show that our DarkDriving dataset provides a comprehensive benchmark for evaluating low‑light enhancement for autonomous driving and it can also be generalized to enhance dark images and promote detection in some other low‑light driving environment, such as nuScenes.The code and dataset will be publicly available at https://github.com/DriveMindLab/DarkDriving‑ICRA‑2026.
Authors:Ziyi Wang, Peiming Li, Xinshun Wang, Yang Tang, Kai-Kuang Ma, Mengyuan Liu
Abstract:
Multimodal large language models (MLLMs) exhibit strong visual‑language reasoning, yet cannot process structured, non‑visual data such as human skeletons. Existing methods either compress skeleton dynamics into lossy feature vectors for text alignment, or quantize motion into discrete tokens that generalize poorly across heterogeneous skeleton formats. We present SkeletonLLM, which achieves universal skeleton understanding by translating arbitrary skeleton sequences into the MLLM's native visual modality. At its core is DrAction, a differentiable, format‑agnostic renderer that converts skeletal kinematics into compact image sequences. Because the pipeline is end‑to‑end differentiable, MLLM gradients can directly guide the rendering to produce task‑informative visual tokens. To further enhance reasoning capabilities, we introduce a cooperative training strategy: Causal Reasoning Distillation transfers structured, step‑by‑step reasoning from a teacher model, while Discriminative Finetuning sharpens decision boundaries between confusable actions. SkeletonLLM demonstrates strong generalization \revisein open‑vocabulary action recognition, while its learned reasoning capabilities naturally extend to motion captioning and question answering across heterogeneous skeleton formats ‑‑ suggesting a viable path for applying MLLMs to non‑native modalities. Code: https://github.com/wangzy01/SkeletonLLM.
Authors:Kevin Qu, Haozhe Qi, Mihai Dusmanu, Mahdi Rad, Rui Wang, Marc Pollefeys
Abstract:
Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint‑aware reasoning. Recent efforts aim to augment the input representations with geometric cues rather than explicitly teaching models to reason in 3D space. We introduce Loc3R‑VLM, a framework that equips 2D Vision‑Language Models with advanced 3D understanding capabilities from monocular video input. Inspired by human spatial cognition, Loc3R‑VLM relies on two joint objectives: global layout reconstruction to build a holistic representation of the scene structure, and explicit situation modeling to anchor egocentric perspective. These objectives provide direct spatial supervision that grounds both perception and language in a 3D context. To ensure geometric consistency and metric‑scale alignment, we leverage lightweight camera pose priors extracted from a pre‑trained 3D foundation model. Loc3R‑VLM achieves state‑of‑the‑art performance in language‑based localization and outperforms existing 2D‑ and video‑based approaches on situated and general 3D question‑answering benchmarks, demonstrating that our spatial supervision framework enables strong 3D understanding. Project page: https://kevinqu7.github.io/loc3r‑vlm
Authors:Yigit Ekin, Yossi Gandelsman
Abstract:
We present a training‑free framework for continuous and controllable image editing at test time for text‑conditioned generative models. In contrast to prior approaches that rely on additional training or manual user intervention, we find that a simple steering in the text‑embedding space is sufficient to produce smooth edit control. Given a target concept (e.g., enhancing photorealism or changing facial expression), we use a large language model to automatically construct a small set of debiased contrastive prompt pairs, from which we compute a steering vector in the generator's text‑encoder space. We then add this vector directly to the input prompt representation to control generation along the desired semantic axis. To obtain a continuous control, we propose an elastic range search procedure that automatically identifies an effective interval of steering magnitudes, avoiding both under‑steering (no‑edit) and over‑steering (changing other attributes). Adding the scaled versions of the same vector within this interval yields smooth and continuous edits. Since our method modifies only textual representations, it naturally generalizes across text‑conditioned modalities, including image and video generation. To quantify the steering continuity, we introduce a new evaluation metric that measures the uniformity of semantic change across edit strengths. We compare the continuous editing behavior across methods and find that, despite its simplicity and lightweight design, our approach is comparable to training‑based alternatives, outperforming other training‑free methods.
Authors:Huajian Zeng, Abhishek Saroha, Daniel Cremers, Xi Wang
Abstract:
Synthesizing controllable 6‑DOF object manipulation trajectories in 3D environments is essential for enabling robots to interact with complex scenes, yet remains challenging due to the need for accurate spatial reasoning, physical feasibility, and multimodal scene understanding. Existing approaches often rely on 2D or partial 3D representations, limiting their ability to capture full scene geometry and constraining trajectory precision. We present GMT, a multimodal transformer framework that generates realistic and goal‑directed object trajectories by jointly leveraging 3D bounding box geometry, point cloud context, semantic object categories, and target end poses. The model represents trajectories as continuous 6‑DOF pose sequences and employs a tailored conditioning strategy that fuses geometric, semantic, contextual, and goaloriented information. Extensive experiments on synthetic and real‑world benchmarks demonstrate that GMT outperforms state‑of‑the‑art human motion and human‑object interaction baselines, such as CHOIS and GIMO, achieving substantial gains in spatial accuracy and orientation control. Our method establishes a new benchmark for learningbased manipulation planning and shows strong generalization to diverse objects and cluttered 3D environments. Project page: https://huajian‑ zeng.github. io/projects/gmt/.
Authors:Jinho Park, Se Young Chun, Mingoo Seok
Abstract:
Radar is a critical perception modality in autonomous driving systems due to its all‑weather characteristics and ability to measure range and Doppler velocity. However, the sheer volume of high‑dimensional raw radar data saturates the communication link to the computing engine (e.g., an NPU), which is often a low‑bandwidth interface with data rate provisioned only for a few low‑resolution range‑Doppler frames. A generalized codec for utilizing high‑dimensional radar data is notably absent, while existing image‑domain approaches are unsuitable, as they typically operate at fixed compression ratios and fail to adapt to varying or adversarial conditions. In light of this, we propose radar data compression with adaptive feedback. It dynamically adjusts the compression ratio by performing gradient descent from the proxy gradient of detection confidence with respect to the compression rate. We employ a zeroth‑order gradient approximation as it enables gradient computation even with non‑differentiable core operations‑‑pruning and quantization. This also avoids transmitting the gradient tensors over the band‑limited link, which, if estimated, would be as large as the original radar data. In addition, we have found that radar feature maps are heavily concentrated on a few frequency components. Thus, we apply the discrete cosine transform to the radar data cubes and selectively prune out the coefficients effectively. We preserve the dynamic range of each radar patch through scaled quantization. Combining those techniques, our proposed online adaptive compression scheme achieves over 100x feature size reduction at minimal performance drop (~1%p). We validate our results on the RADIal, CARRADA, and Radatron datasets.
Authors:Aymen Mir, Riza Alp Guler, Xiangjun Tang, Peter Wonka, Gerard Pons-Moll
Abstract:
We present AHOY, a method for reconstructing complete, animatable 3D Gaussian avatars from in‑the‑wild monocular video despite heavy occlusion. Existing methods assume unoccluded input‑a fully visible subject, often in a canonical pose‑excluding the vast majority of real‑world footage where people are routinely occluded by furniture, objects, or other people. Reconstructing from such footage poses fundamental challenges: large body regions may never be observed, and multi‑view supervision per pose is unavailable. We address these challenges with four contributions: (i) a hallucination‑as‑supervision pipeline that uses identity‑finetuned diffusion models to generate dense supervision for previously unobserved body regions; (ii) a two‑stage canonical‑to‑pose‑dependent architecture that bootstraps from sparse observations to full pose‑dependent Gaussian maps; (iii) a map‑pose/LBS‑pose decoupling that absorbs multi‑view inconsistencies from the generated data; (iv) a head/body split supervision strategy that preserves facial identity. We evaluate on YouTube videos and on multi‑view capture data with significant occlusion and demonstrate state‑of‑the‑art reconstruction quality. We also demonstrate that the resulting avatars are robust enough to be animated with novel poses and composited into 3DGS scenes captured using cell‑phone video. Our project page is available at https://miraymen.github.io/ahoy/
Authors:Markus Gross, Sai Bharadhwaj Matha, Rui Song, Viswanathan Muthuveerappan, Conrad Christoph, Julius Huber, Daniel Cremers
Abstract:
Semantic segmentation for uncrewed aerial vehicles (UAVs) is fundamental for aerial scene understanding, yet existing RGB and RGB‑T datasets remain limited in scale, diversity, and annotation efficiency due to the high cost of manual labeling and the difficulties of accurate RGB‑T alignment on off‑the‑shelf UAVs. To address these challenges, we propose a scalable geometry‑driven 2D‑3D‑2D paradigm that leverages multi‑view redundancy in high‑overlap aerial imagery to automatically propagate labels from a small subset of manually annotated RGB images to both RGB and thermal modalities within a unified framework. By lifting less than 3% of RGB images into a semantic 3D point cloud and reprojecting it into all views, our approach enables dense pseudo ground‑truth generation across large image collections, automatically producing 97% of RGB labels and 100% of thermal labels while achieving 91% and 88% annotation accuracy without any 2D manual refinement. We further extend this 2D‑3D‑2D paradigm to cross‑modal image registration, using 3D geometry as an intermediate alignment space to obtain fully automatic, strong pixel‑level RGB‑T alignment with 87% registration accuracy and no hardware‑level synchronization. Applying our framework to existing geo‑referenced aerial imagery, we construct SegFly, a large‑scale benchmark with over 20,000 high‑resolution RGB images and more than 15,000 geometrically aligned RGB‑T pairs spanning diverse urban, industrial, and rural environments across multiple altitudes and seasons. On SegFly, we establish the Firefly baseline for RGB and thermal semantic segmentation and show that both conventional architectures and vision foundation models benefit substantially from SegFly supervision, highlighting the potential of geometry‑driven 2D‑3D‑2D pipelines for scalable multi‑modal scene understanding. Data and Code available at https://github.com/markus‑42/SegFly.
Authors:Yingjie Chen, Shilun Lin, Cai Xing, Binxin Yang, Long Zhou, Qixin Yan, Wenjing Wang, Dingming Liu, Hao Liu, Chen Li, Jing Lyu
Abstract:
Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity‑aware content creation. Nevertheless, an openly accessible framework enabling fine‑grained control over facial appearance and voice timbre across multiple identities remains unavailable. In this work, we present a unified and scalable framework for identity‑aware joint audio‑video generation, enabling high‑fidelity and consistent personalization. Specifically, we introduce a data curation pipeline that automatically extracts identity‑bearing information with paired annotations across audio and visual modalities, covering diverse scenarios from single‑subject to multi‑subject interactions. We further propose a flexible and scalable identity injection mechanism for single‑ and multi‑subject scenarios, in which both facial appearance and vocal timbre act as identity‑bearing control signals. Moreover, in light of modality disparity, we design a multi‑stage training strategy to accelerate convergence and enforce cross‑modal coherence. Experiments demonstrate the superiority of the proposed framework. For more details and qualitative results, please refer to our webpage: \hrefhttps://chen‑yingjie.github.io/projects/Identity‑as‑PresenceIdentity‑as‑Presence.
Authors:Anwai Archit, Constantin Pape
Abstract:
Cell segmentation is a fundamental task in microscopy image analysis. Several foundation models for cell segmentation have been introduced, virtually all of them are extensions of Segment Anything Model (SAM), improving it for microscopy data. Recently, SAM2 and SAM3 have been published, further improving and extending the capabilities of general‑purpose segmentation foundation models. Here, we comprehensively evaluate foundation models for cell segmentation (CellPoseSAM, CellSAM, μSAM) and for general‑purpose segmentation (SAM, SAM2, SAM3) on a diverse set of (light) microscopy datasets, for tasks including cell, nucleus and organoid segmentation. Furthermore, we introduce a new instance segmentation strategy called automatic prompt generation (APG) that can be used to further improve SAM‑based microscopy foundation models. APG consistently improves segmentation results for μSAM, which is used as the base model, and is competitive with the state‑of‑the‑art model CellPoseSAM. Moreover, our work provides important lessons for adaptation strategies of SAM‑style models to microscopy and provides a strategy for creating even more powerful microscopy foundation models. Our code is publicly available at https://github.com/computational‑cell‑analytics/micro‑sam.
Authors:Ziwei Xiang, Fanhu Zeng, Hongjian Fang, Rui-Qi Wang, Renxing Chen, Yanan Zhu, Yi Chen, Peipei Yang, Xu-Yao Zhang
Abstract:
Large Vision Language Models (LVLMs) have achieved remarkable success in a range of downstream tasks that require multimodal interaction, but their capabilities come with substantial computational and memory overhead, which hinders practical deployment. Among numerous acceleration techniques, post‑training quantization is a popular and effective strategy for reducing memory cost and accelerating inference. However, existing LVLM quantization methods typically measure token sensitivity at the modality level, which fails to capture the complex cross‑token interactions and falls short in quantitatively measuring the quantization error at the token level. As tokens interact within the model, the distinction between modalities gradually diminishes, suggesting the need for fine‑grained calibration. Inspired by axiomatic attribution in mechanistic interpretability, we introduce a fine‑grained quantization strategy on Quantization‑aware Integrated Gradients (QIG), which leverages integrated gradients to quantitatively evaluate token sensitivity and push the granularity from modality level to token level, reflecting both inter‑modality and intra‑modality dynamics. Extensive experiments on multiple LVLMs under both W4A8 and W3A16 settings show that our method improves accuracy across models and benchmarks with negligible latency overhead. For example, under 3‑bit weight‑only quantization, our method improves the average accuracy of LLaVA‑onevision‑7B by 1.60%, reducing the gap to its full‑precision counterpart to only 1.33%. The code is available at https://github.com/ucas‑xiang/QIG.
Authors:Haoyun Chen, Fenghe Tang, Wenxin Ma, Shaohua Kevin Zhou
Abstract:
Universal medical image segmentation seeks to use a single foundational model to handle diverse tasks across multiple imaging modalities. However, existing approaches often rely heavily on manual visual prompts or retrieved reference images, which limits their automation and robustness. In addition, naive joint training across modalities often fails to address large domain shifts. To address these limitations, we propose Concept‑to‑Pixel (C2P), a novel prompt‑free universal segmentation framework. C2P explicitly separates anatomical knowledge into two components: Geometric and Semantic representations. It leverages Multimodal Large Language Models (MLLMs) to distill abstract, high‑level medical concepts into learnable Semantic Tokens and introduces explicitly supervised Geometric Tokens to enforce universal physical and structural constraints. These disentangled tokens interact deeply with image features to generate input‑specific dynamic kernels for precise mask prediction. Furthermore, we introduce a Geometry‑Aware Inference Consensus mechanism, which utilizes the model's predicted geometric constraints to assess prediction reliability and suppress outliers. Extensive experiments and analysis on a unified benchmark comprising eight diverse datasets across seven modalities demonstrate the significant superiority of our jointly trained approach, compared to universe‑ or single‑model approaches. Remarkably, our unified model demonstrates strong generalization, achieving impressive results not only on zero‑shot tasks involving unseen cases but also in cross‑modal transfers across similar tasks. Code is available at: https://github.com/Yundi218/Concept‑to‑Pixel
Authors:Yuhe Tian, Kun Zhang, Haoran Ma, Rui Yan, Yingtai Li, Rongsheng Wang, Shaohua Kevin Zhou
Abstract:
While large language models (LLMs) have advanced CT report generation, existing methods typically encode 3D volumes holistically, failing to distinguish informative cues from redundant anatomical background. Inspired by radiological cognitive subtraction, we propose Differential Visual Prompting (DiffVP), which conditions report generation on explicit, high‑level semantic scan‑to‑reference differences rather than solely on absolute visual features. DiffVP employs a hierarchical difference extractor to capture complementary global and local semantic discrepancies into a shared latent space, along with a difference‑to‑prompt generator that transforms these signals into learnable visual prefix tokens for LLM conditioning. These difference prompts serve as structured conditioning signals that implicitly suppress invariant anatomy while amplifying diagnostically relevant visual evidence, thereby facilitating accurate report generation without explicit lesion localization. On two large‑scale benchmarks, DiffVP consistently outperforms prior methods, improving the average BLEU‑1‑4 by +10.98 and +4.36, respectively, and further boosts clinical efficacy on RadGenome‑ChestCT (F1 score 0.421). All codes will be released at https://github.com/ArielTYH/DiffVP/.
Authors:Haocheng Li, Juepeng Zheng, Shuangxi Miao, Ruibo Lu, Guosheng Cai, Haohuan Fu, Jianxi Huang
Abstract:
Multimodal remote sensing semantic segmentation enhances scene interpretation by exploiting complementary physical cues from heterogeneous data. Although pretrained Vision Foundation Models (VFMs) provide strong general‑purpose representations, adapting them to multimodal tasks often incurs substantial computational overhead and is prone to modality imbalance, where the contribution of auxiliary modalities is suppressed during optimization. To address these challenges, we propose MoBaNet, a parameter‑efficient and modality‑balanced symmetric fusion framework. Built upon a largely frozen VFM backbone, MoBaNet adopts a symmetric dual‑stream architecture to preserve generalizable representations while minimizing the number of trainable parameters. Specifically, we design a Cross‑modal Prompt‑Injected Adapter (CPIA) to enable deep semantic interaction by generating shared prompts and injecting them into bottleneck adapters under the frozen backbone. To obtain compact and discriminative multimodal representations for decoding, we further introduce a Difference‑Guided Gated Fusion Module (DGFM), which adaptively fuses paired stage features by explicitly leveraging cross‑modal discrepancy to guide feature selection. Furthermore, we propose a Modality‑Conditional Random Masking (MCRM) strategy to mitigate modality imbalance by masking one modality only during training and imposing hard‑pixel auxiliary supervision on modality‑specific branches. Extensive experiments on the ISPRS Vaihingen and Potsdam benchmarks demonstrate that MoBaNet achieves state‑of‑the‑art performance with significantly fewer trainable parameters than full fine‑tuning, validating its effectiveness for robust and balanced multimodal fusion. The source code in this work is available at https://github.com/sauryeo/MoBaNet.
Authors:Liangyu Yuan, Ruoyu Wang, Tong Zhao, Dingwen Fu, Mingkun Lei, Beier Zhu, Chi Zhang
Abstract:
Diffusion and flow matching models generate high‑fidelity data by simulating paths defined by Ordinary or Stochastic Differential Equations (ODEs/SDEs), starting from a tractable prior distribution. The probability flow ODE formulation enables the use of advanced numerical solvers to accelerate sampling. Orthogonal yet vital to solver design is the discretization strategy. While early approaches employed handcrafted heuristics and recent methods adopt optimization‑based techniques, most existing strategies enforce a globally shared timestep schedule across all samples. This uniform treatment fails to account for instance‑specific complexity in the generative process, potentially limiting performance. Motivated by controlled experiments on synthetic data, which reveals the suboptimality of global schedules under instance‑specific dynamics, we propose an instance‑aware discretization framework. Our method learns to adapt timestep allocations based on input‑dependent priors, extending gradient‑based discretization search to the conditional generative setting. Empirical results across diverse settings, including synthetic data, pixel‑space diffusion, latent‑space images and video flow matching models, demonstrate that our method consistently improves generation quality with marginal tuning cost compared to training and negligible inference overhead.
Authors:Rui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata, Stephan Alaniz
Abstract:
Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine‑grained queries, a challenge underrepresented by existing benchmarks that focus on coarse image‑related questions. We introduce FIne‑grained NEgative queRies (FINER), alongside two benchmarks: FINER‑CompreCap and FINER‑DOCCI. Using FINER, we analyze hallucinations across four settings: multi‑object, multi‑attribute, multi‑relation, and ``what'' questions. Our benchmarks reveal that MLLMs hallucinate when fine‑grained mismatches co‑occur with genuinely present elements in the image. To address this, we propose FINER‑Tuning, leveraging Direct Preference Optimization (DPO) on FINER‑inspired data. Finetuning four frontier MLLMs with FINER‑Tuning yields up to 24.2% gains (InternVL3.5‑14B) on hallucinations from our benchmarks, while simultaneously improving performance on eight existing hallucination suites, and enhancing general multimodal capabilities across six benchmarks. Code, benchmark, and models are available at \hrefhttps://explainableml.github.io/finer‑project/https://explainableml.github.io/finer‑project/.
Authors:Chaokang Jiang, Desen Zhou, Jiuming Liu, Kevin Li Sun
Abstract:
Closed‑loop evaluation of autonomous‑driving policies requires interactive simulation beyond log replay. However, existing generative world models often degrade in closed loop due to (i) history‑free initialization that mismatches policy inputs, (ii) multi‑step sampling latency that violates real‑time budgets, and (iii) compounding kinematic infeasibility over long horizons. We propose VectorWorld, a streaming world model that incrementally generates ego‑centric 64 \mathrmm× 64\mathrmm lane‑‑agent vector‑graph tiles during rollout. VectorWorld aligns initialization with history‑conditioned policies by producing a policy‑compatible interaction state via a motion‑aware gated VAE. It enables real‑time outpainting via solver‑free one‑step masked completion with an edge‑gated relational DiT trained with interval‑conditioned MeanFlow and JVP‑based large‑step supervision. To stabilize long‑horizon rollouts, we introduce ΔSim, a physics‑aligned non‑ego (NPC) policy with hybrid discrete‑‑continuous actions and differentiable kinematic logit shaping. On Waymo open motion and nuPlan, VectorWorld improves map‑structure fidelity and initialization validity, and supports stable, real‑time 1\mathrmkm+ closed‑loop rollouts (\hrefhttps://github.com/jiangchaokang/VectorWorldcode).
Authors:Tae Eun Choi, Sumin Shim, Junhyeok Kim, Seong Jae Hwang
Abstract:
Generative inbetweening (GI) seeks to synthesize realistic intermediate frames between the first and last keyframes beyond mere interpolation. As sequences become sparser and motions larger, previous GI models struggle with inconsistent frames with unstable pacing and semantic misalignment. Since GI involves fixed endpoints and numerous plausible paths, this task requires additional guidance gained from the keyframes and text to specify the intended path. Thus, we give semantic and temporal guidance from the keyframes and text onto each intermediate frame through Keyframe‑anchored Attention Bias. We also better enforce frame consistency with Rescaled Temporal RoPE, which allows self‑attention to attend to keyframes more faithfully. TGI‑Bench, the first benchmark specifically designed for text‑conditioned GI evaluation, enables challenge‑targeted evaluation to analyze GI models. Without additional training, our method achieves state‑of‑the‑art frame consistency, semantic fidelity, and pace stability for both short and long sequences across diverse challenges.
Authors:Xinze Li, Pengxu Chen, Yiyuan Wang, Weifeng Su, Wentao Cheng
Abstract:
Feed‑forward 3D foundation models face a key challenge: the quadratic computational cost introduced by global attention, which severely limits scalability as input length increases. Concurrent acceleration methods, such as token merging, operate at the token level. While they offer local savings, the required nearest‑neighbor searches introduce undesirable overhead. Consequently, these techniques fail to tackle the fundamental issue of structural redundancy dominant in dense capture data. In this work, we introduce S‑VGGT, a novel approach that addresses redundancy at the structural frame level, drastically shifting the optimization focus. We first leverage the initial features to build a dense scene graph, which characterizes structural scene redundancy and guides the subsequent scene partitioning. Using this graph, we softly assign frames to a small number of subscenes, guaranteeing balanced groups and smooth geometric transitions. The core innovation lies in designing the subscenes to share a common reference frame, establishing a parallel geometric bridge that enables independent and highly efficient processing without explicit geometric alignment. This structural reorganization provides strong intrinsic acceleration by cutting the global attention cost at its source. Crucially, S‑VGGT is entirely orthogonal to token‑level acceleration methods, allowing the two to be seamlessly combined for compounded speedups without compromising reconstruction fidelity. Code is available at https://github.com/Powertony102/S‑VGGT.
Authors:Yaxu Xie, Abdalla Arafa, Alireza Javanmardi, Christen Millerdurai, Jia Cheng Hu, Shaoxiang Wang, Alain Pagani, Didier Stricker
Abstract:
Achieving unified 3D perception and reasoning across tasks such as segmentation, retrieval, and relation understanding remains challenging, as existing methods are either object‑centric or rely on costly training for inter‑object reasoning. We present a novel framework that constructs a hierarchical language‑distilled Gaussian scene and its 3D semantic scene graph without scene‑specific training. A Gaussian pruning mechanism refines scene geometry, while a robust multi‑view language alignment strategy aggregates noisy 2D features into accurate 3D object embeddings. On top of this hierarchy, we build an open‑vocabulary 3D scene graph with Vision Language derived annotations and Graph Neural Network‑based relational reasoning. Our approach enables efficient and scalable open‑vocabulary 3D reasoning by jointly modeling hierarchical semantics and inter/intra‑object relationships, validated across tasks including open‑vocabulary segmentation, scene graph generation, and relation‑guided retrieval. Project page: https://dfki‑av.github.io/ReLaGS/
Authors:Qihong Tang, Changhan Liu, Shaofeng Zhang, Wenbin Li, Qi Fan, Yang Gao
Abstract:
Identifying potential objects is critical for object recognition and analysis across various computer vision applications. Existing methods typically localize potential objects by relying on exemplar images, predefined categories, or textual descriptions. However, their reliance on image and text prompts often limits flexibility, restricting adaptability in real‑world scenarios. In this paper, we introduce a novel Prompt‑Free Universal Region Proposal Network (PF‑RPN), which identifies potential objects without relying on external prompts. First, the Sparse Image‑Aware Adapter (SIA) module performs initial localization of potential objects using a learnable query embedding dynamically updated with visual features. Next, the Cascade Self‑Prompt (CSP) module identifies the remaining potential objects by leveraging the self‑prompted learnable embedding, autonomously aggregating informative visual features in a cascading manner. Finally, the Centerness‑Guided Query Selection (CG‑QS) module facilitates the selection of high‑quality query embeddings using a centerness scoring network. Our method can be optimized with limited data (e.g., 5% of MS COCO data) and applied directly to various object detection application domains for identifying potential objects without fine‑tuning, such as underwater object detection, industrial defect detection, and remote sensing image object detection. Experimental results across 19 datasets validate the effectiveness of our method. Code is available at https://github.com/tangqh03/PF‑RPN.
Authors:Yimin Wei, Aoran Xiao, Hongruixuan Chen, Junshi Xia, Naoto Yokoya
Abstract:
Open‑vocabulary segmentation enables pixel‑level recognition from an open set of textual categories, allowing generalization beyond fixed classes. Despite great potential in remote sensing, progress in this area remains largely limited to clear‑sky optical data and struggles under cloudy or haze‑contaminated conditions. We present MM‑OVSeg, a multimodal Optical‑SAR fusion framework for resilient open‑vocabulary segmentation under adverse weather conditions. MM‑OVSeg leverages the complementary strengths of the two modalities‑‑optical imagery provides rich spectral semantics, while synthetic aperture radar (SAR) offers cloud‑penetrating structural cues. To address the cross‑modal domain gap and the limited dense prediction capability of current vision‑language models, we propose two key designs: a cross‑modal unification process for multi‑sensor representation alignment, and a dual‑encoder fusion module that integrates hierarchical features from multiple vision foundation models for text‑aligned multimodal segmentation. Extensive experiments demonstrate that MM‑OVSeg achieves superior robustness and generalization across diverse cloud conditions. The source dataset and code are available at https://github.com/Jimmyxichen/MM‑OVSeg.
Authors:Jiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang, Haibo Qiu, Lihuo He, Dengpan Ye, Xinbo Gao, Jing Zhang
Abstract:
We present Omni‑I2C, a comprehensive benchmark designed to evaluate the capability of Large Multimodal Models (LMMs) in converting complex, structured digital graphics into executable code. We argue that this task represents a non‑trivial challenge for the current generation of LMMs: it demands an unprecedented synergy between high‑fidelity visual perception ‑‑ to parse intricate spatial hierarchies and symbolic details ‑‑ and precise generative expression ‑‑ to synthesize syntactically sound and logically consistent code. Unlike traditional descriptive tasks, Omni‑I2C requires a holistic understanding where any minor perceptual hallucination or coding error leads to a complete failure in visual reconstruction. Omni‑I2C features 1080 meticulously curated samples, defined by its breadth across subjects, image modalities, and programming languages. By incorporating authentic user‑sourced cases, the benchmark spans a vast spectrum of digital content ‑‑ from scientific visualizations to complex symbolic notations ‑‑ each paired with executable reference code. To complement this diversity, our evaluation framework provides necessary depth; by decoupling performance into perceptual fidelity and symbolic precision, it transcends surface‑level accuracy to expose the granular structural failures and reasoning bottlenecks of current LMMs. Our evaluation reveals a substantial performance gap among leading LMMs; even state‑of‑the‑art models struggle to preserve structural integrity in complex scenarios, underscoring that multimodal code generation remains a formidable challenge. Data and code are available at https://github.com/MiliLab/Omni‑I2C.
Authors:Segyu Lee, Boryeong Cho, Hojung Jung, Seokhyun An, Juhyeong Kim, Jaehyun Kwak, Yongjin Yang, Sangwon Jang, Youngrok Park, Wonjun Chang, Se-Young Yun
Abstract:
Unified Multimodal Models (UMMs) offer powerful cross‑modality capabilities but introduce new safety risks not observed in single‑task models. Despite their emergence, existing safety benchmarks remain fragmented across tasks and modalities, limiting the comprehensive evaluation of complex system‑level vulnerabilities. To address this gap, we introduce UniSAFE, the first comprehensive benchmark for system‑level safety evaluation of UMMs across 7 I/O modality combinations, spanning conventional tasks and novel multimodal‑context image generation settings. UniSAFE is built with a shared‑target design that projects common risk scenarios across task‑specific I/O configurations, enabling controlled cross‑task comparisons of safety failures. Comprising 6,802 curated instances, we use UniSAFE to evaluate 15 state‑of‑the‑art UMMs, both proprietary and open‑source. Our results reveal critical vulnerabilities across current UMMs, including elevated safety violations in multi‑image composition and multi‑turn settings, with image‑output tasks consistently more vulnerable than text‑output tasks. These findings highlight the need for stronger system‑level safety alignment for UMMs. Our code and data are publicly available at https://github.com/segyulee/UniSAFE
Authors:Chupeng Liu, Jiyong Rao, Shangquan Sun, Runkai Zhao, Weidong Cai
Abstract:
Monocular 3D object detection typically relies on pseudo‑labeling techniques to reduce dependency on real‑world annotations. Recent advances demonstrate that deterministic linguistic cues can serve as effective auxiliary weak supervision signals, providing complementary semantic context. However, hand‑crafted textual descriptions struggle to capture the inherent visual diversity of individuals across scenes, limiting the model's ability to learn scene‑aware representations. To address this challenge, we propose Visual‑referred Probabilistic Prompt Learning (VirPro), an adaptive multi‑modal pretraining paradigm that can be seamlessly integrated into diverse weakly supervised monocular 3D detection frameworks. Specifically, we generate a diverse set of learnable, instance‑conditioned prompts across scenes and store them in an Adaptive Prompt Bank (APB). Subsequently, we introduce Multi‑Gaussian Prompt Modeling (MGPM), which incorporates scene‑based visual features into the corresponding textual embeddings, allowing the text prompts to express visual uncertainties. Then, from the fused vision‑language embeddings, we decode a prompt‑targeted Gaussian, from which we derive a unified object‑level prompt embedding for each instance. RoI‑level contrastive matching is employed to enforce modality alignment, bringing embeddings of co‑occurring objects within the same scene closer in the latent space, thus enhancing semantic coherence. Extensive experiments on the KITTI benchmark demonstrate that integrating our pretraining paradigm consistently yields substantial performance gains, achieving up to a 4.8% average precision improvement than the baseline. Code is available at https://github.com/AustinLCP/VirPro.
Authors:Chaeyun Kim, Seunghoon Yi, Yejin Kim, Yohan Jo, Joonseok Lee
Abstract:
Referring Image Segmentation (RIS) requires identifying objects from images based on textual descriptions. We observe that existing methods significantly underperform on motion‑related queries compared to appearance‑based ones. To address this, we first introduce an efficient data augmentation scheme that extracts motion‑centric phrases from original captions, exposing models to more motion expressions without additional annotations. Second, since the same object can be described differently depending on the context, we propose Multimodal Radial Contrastive Learning (MRaCL), performed on fused image‑text embeddings rather than unimodal representations. For comprehensive evaluation, we introduce a new test split focusing on motion‑centric queries, and introduce a new benchmark called M‑Bench, where objects are distinguished primarily by actions. Extensive experiments show our method substantially improves performance on motion‑centric queries across multiple RIS models, maintaining competitive results on appearance‑based descriptions. Codes are available at https://github.com/snuviplab/MRaCL
Authors:Yang-Tian Sun, Zehuan Huang, Yifan Niu, Lin Ma, Yan-Pei Cao, Yuewen Ma, Xiaojuan Qi
Abstract:
We present StereoWorld, a camera‑conditioned stereo world model that jointly learns appearance and binocular geometry for end‑to‑end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry directly from disparity. To efficiently achieve consistent stereo generation, our approach introduces two key designs: (1) a unified camera‑frame RoPE that augments latent tokens with camera‑aware rotary positional encoding, enabling relative, view‑ and time‑consistent conditioning while preserving pretrained video priors via a stable attention initialization; and (2) a stereo‑aware attention decomposition that factors full 4D attention into 3D intra‑view attention plus horizontal row attention, leveraging the epipolar prior to capture disparity‑aligned correspondences with substantially lower compute. Across benchmarks, StereoWorld improves stereo consistency, disparity accuracy, and camera‑motion fidelity over strong monocular‑then‑convert pipelines, achieving more than 3x faster generation with an additional 5% gain in viewpoint consistency. Beyond benchmarks, StereoWorld enables end‑to‑end binocular VR rendering without depth estimation or inpainting, enhances embodied policy learning through metric‑scale depth grounding, and is compatible with long‑video distillation for extended interactive stereo synthesis.
Authors:Umangi Jain, Vladimir Kim, Matheus Gadelha, Igor Gilitschenski, Zhiqin Chen
Abstract:
We introduce the problem of material‑aware part grouping in untextured meshes. Many real‑world shapes, such as scales of pinecones or windows of buildings, contain repeated structures that share the same material but exhibit geometric variations. When assigning materials to such meshes, these repeated parts often require piece‑by‑piece manual identification and selection, which is tedious and time‑consuming. To address this, we propose Material Magic Wand, a tool that allows artists to select part groups based on their estimated material properties ‑‑ when one part is selected, our algorithm automatically retrieves all other parts likely to share the same material. The key component of our approach is a part encoder that generates a material‑aware embedding for each 3D part, accounting for both local geometry and global context. We train our model with a supervised contrastive loss that brings embeddings of material‑consistent parts closer while separating those of different materials; therefore, part grouping can be achieved by retrieving embeddings that are close to the embedding of the selected part. To benchmark this task, we introduce a curated dataset of 100 shapes with 241 part‑level queries. We verify the effectiveness of our method through extensive experiments and demonstrate its practical value in an interactive material assignment application.
Authors:Yiwen Zhao, Ce Zheng, Yufu Wang, Hsueh-Han Daniel Yang, Liting Wen, Laszlo A. Jeni
Abstract:
Human mesh recovery (HMR) models 3D human body from monocular videos, with recent works extending it to world‑coordinate human trajectory and motion reconstruction. However, most existing methods remain offline, relying on future frames or global optimization, which limits their applicability in interactive feedback and perception‑action loop scenarios such as AR/VR and telepresence. To address this, we propose OnlineHMR, a fully online framework that jointly satisfies four essential criteria of online processing, including system‑level causality, faithfulness, temporal consistency, and efficiency. Built upon a two‑branch architecture, OnlineHMR enables streaming inference via a causal key‑value cache design and a curated sliding‑window learning strategy. Meanwhile, a human‑centric incremental SLAM provides online world‑grounded alignment under physically plausible trajectory correction. Experimental results show that our method achieves performance comparable to existing chunk‑based approaches on the standard EMDB benchmark and highly dynamic custom videos, while uniquely supporting online processing. Page and code are available at https://tsukasane.github.io/Video‑OnlineHMR/.
Authors:Thuy Truong Tran, Minh Kha Do, Phuc Nguyen Duy, Min Hun Lee
Abstract:
Medical anomaly detection (MAD) and segmentation play a critical role in assisting clinical diagnosis by identifying abnormal regions in medical images and localizing pathological regions. Recent CLIP‑based studies are promising for anomaly detection in zero‑/few‑shot settings, and typically rely on global representations and weak supervision, often producing coarse localization and limited segmentation quality. In this work, we study supervised adaptation of CLIP for MAD under a realistic clinical setting where a limited yet meaningful amount of labeled abnormal data is available. Our model MedSAD‑CLIP leverages fine‑grained text‑visual cues via the Token‑Patch Cross‑Attention(TPCA) to improve lesion localization while preserving the generalization capability of CLIP representations. Lightweight image adapters and learnable prompt tokens efficiently adapt the pretrained CLIP encoder to the medical domain while preserving its rich semantic alignment. Furthermore, a Margin‑based image‑text Contrastive Loss is designed to enhance global feature discrimination between normal and abnormal representations. Extensive experiments on four diverse benchmarks‑Brain, Retina, Lung, and Breast datasets‑demonstrate the effectiveness of our approach, achieving superior performance in both pixel‑level segmentation and image‑level classification over state‑of‑the‑art methods. Our results highlight the potential of supervised CLIP adaptation as a unified and scalable paradigm for medical anomaly understanding. Code will be made available at https://github.com/thuy4tbn99/MedSAD‑CLIP
Authors:Yuelin Zhang, Sijie Cheng, Chen Li, Zongzhao Li, Yuxin Huang, Yang Liu, Wenbing Huang
Abstract:
Accurately estimating task progress is critical for embodied agents to plan and execute long‑horizon, multi‑step tasks. Despite promising advances, existing Vision‑Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning potential. Furthermore, processing long video trajectories with VLMs is computationally prohibitive for real‑world deployment. To address these challenges, we propose the Recurrent Reasoning Vision‑Language Model (\textR^2VLM). Our model features a recurrent reasoning framework that processes local video snippets iteratively, maintaining a global context through an evolving Chain of Thought (CoT). This CoT explicitly records task decomposition, key steps, and their completion status, enabling the model to reason about complex temporal dependencies. This design avoids the high cost of processing long videos while preserving essential reasoning capabilities. We train \textR^2VLM on large‑scale, automatically generated datasets from ALFRED and Ego4D. Extensive experiments on progress estimation and downstream applications, including progress‑enhanced policy learning, reward modeling for reinforcement learning, and proactive assistance, demonstrate that \textR^2VLM achieves strong performance and generalization, achieving a new state‑of‑the‑art in long‑horizon task progress estimation. The models and benchmarks are publicly available at \hrefhttps://huggingface.co/collections/zhangyuelin/r2vlmhuggingface.
Authors:Haiyang Yan, Hongyun Zhou, Peng Xu, Xiaoxue Feng, Mengyi Liu
Abstract:
Despite rapid developments and widespread applications of MLLM agents, they still struggle with long‑form video understanding (LVU) tasks, which are characterized by high information density and extended temporal spans. Recent research on LVU agents demonstrates that simple task decomposition and collaboration mechanisms are insufficient for long‑chain reasoning tasks. Moreover, directly reducing the time context through embedding‑based retrieval may lose key information of complex problems. In this paper, we propose Symphony, a multi‑agent system, to alleviate these limitations. By emulating human cognition patterns, Symphony decomposes LVU into fine‑grained subtasks and incorporates a deep reasoning collaboration mechanism enhanced by reflection, effectively improving the reasoning capability. Additionally, Symphony provides a VLM‑based grounding approach to analyze LVU tasks and assess the relevance of video segments, which significantly enhances the ability to locate complex problems with implicit intentions and large temporal spans. Experimental results show that Symphony achieves state‑of‑the‑art performance on LVBench, LongVideoBench, VideoMME, and MLVU, with a 5.0% improvement over the prior state‑of‑the‑art method on LVBench. Code is available at https://github.com/Haiyang0226/Symphony.
Authors:Zhuojiang Cai, Zhenghui Sun, Feng Lu
Abstract:
We present GazeOnce360, a novel end‑to‑end model for multi‑person gaze estimation from a single tabletop‑mounted upward‑facing fisheye camera. Unlike conventional approaches that rely on forward‑facing cameras in constrained viewpoints, we address the underexplored setting of estimating the 3D gaze direction of multiple people distributed across a 360° scene from an upward fisheye perspective. To support research in this setting, we introduce MPSGaze360, a large‑scale synthetic dataset rendered using Unreal Engine, featuring diverse multi‑person configurations with accurate 3D gaze and eye landmark annotations. Our model tackles the severe distortion and perspective variation inherent in fisheye imagery by incorporating rotational convolutions and eye landmark supervision. To better capture fine‑grained eye features crucial for gaze estimation, we propose a dual‑resolution architecture that fuses global low‑resolution context with high‑resolution local eye regions. Experimental results demonstrate the effectiveness of each component in our model. This work highlights the feasibility and potential of fisheye‑based 360° gaze estimation in practical multi‑person scenarios. Project page: https://caizhuojiang.github.io/GazeOnce360/.
Authors:Wei Yu, Runjia Qian, Yumeng Li, Liquan Wang, Songheng Yin, Sri Siddarth Chakaravarthy P, Dennis Anthony, Yang Ye, Yidi Li, Weiwei Wan, Animesh Garg
Abstract:
Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial memory remains a key bottleneck: explicit 3D structures can improve reprojection‑based consistency but struggle to depict moving objects, while implicit memory often produces inaccurate camera motion even with correct poses. We propose Mosaic Memory (MosaicMem), a hybrid spatial memory that lifts patches into 3D for reliable localization and targeted retrieval, while exploiting the model's native conditioning to preserve prompt‑following generation. MosaicMem composes spatially aligned patches in the queried view via a patch‑and‑compose interface, preserving what should persist while allowing the model to inpaint what should evolve. With PRoPE camera conditioning and two new memory alignment methods, experiments show improved pose adherence compared to implicit memory and stronger dynamic modeling than explicit baselines. MosaicMem further enables minute‑level navigation, memory‑based scene editing, and autoregressive rollout.
Authors:M. Arda Aydın, Melih B. Yilmaz, Aykut Koç, Tolga Çukur
Abstract:
The success of CLIP‑like vision‑language models (VLMs) on natural images has inspired medical counterparts, yet existing approaches largely fall into two extremes: specialist models trained on single‑domain data, which capture domain‑specific details but generalize poorly, and generalist medical VLMs trained on multi‑domain data, which retain broad semantics but dilute fine‑grained diagnostic cues. Bridging this specialization‑generalization trade‑off remains challenging. To address this problem, we propose ACE‑LoRA, a parameter‑efficient adaptation framework for generalist medical VLMs that maintains robust zero‑shot generalization. ACE‑LoRA integrates Low‑Rank Adaptation (LoRA) modules into frozen image‑text encoders and introduces an Attention‑based Context Enhancement Hypergraph Neural Network (ACE‑HGNN) module that captures higher‑order contextual interactions beyond pairwise similarity to enrich global representations with localized diagnostic cues, addressing a key limitation of prior Parameter‑Efficient Fine‑Tuning (PEFT) methods that overlook fine‑grained details. To further enhance cross‑modal alignment, we formulate a label‑guided InfoNCE loss to effectively suppress false negatives between semantically related image‑text pairs. Despite adding only 0.95M trainable parameters, ACE‑LoRA consistently outperforms state‑of‑the‑art medical VLMs and PEFT baselines across zero‑shot classification, segmentation, and detection benchmarks spanning multiple domains. Our code is available at https://github.com/icon‑lab/ACE‑LoRA.
Authors:Yasaswini Chebolu
Abstract:
Reliable terrain perception is a fundamental requirement for autonomous navigation in unstructured, off‑road environments. Desert landscapes present unique challenges due to low chromatic contrast between terrain categories, extreme lighting variability, and sparse vegetation that defy the assumptions of standard road‑scene segmentation models. We present DesertFormer, a semantic segmentation pipeline for off‑road desert terrain analysis based on SegFormer B2 with a hierarchical Mix Transformer (MiT‑B2) backbone. The system classifies terrain into ten ecologically meaningful categories ‑‑ Trees, Lush Bushes, Dry Grass, Dry Bushes, Ground Clutter, Flowers, Logs, Rocks, Landscape, and Sky ‑‑ enabling safety‑aware path planning for ground robots and autonomous vehicles. Trained on a purpose‑built dataset of 4,176 annotated off‑road images at 512x512 resolution, DesertFormer achieves a mean Intersection‑over‑Union (mIoU) of 64.4% and pixel accuracy of 86.1%, representing a +24.2% absolute improvement over a DeepLabV3 MobileNetV2 baseline (41.0% mIoU). We further contribute a systematic failure analysis identifying the primary confusion patterns ‑‑ Ground Clutter to Landscape and Dry Grass to Landscape ‑‑ and propose class‑weighted training and copy‑paste augmentation for rare terrain categories. Code, checkpoints, and an interactive inference dashboard are released at https://github.com/Yasaswini‑ch/Vision‑based‑Desert‑Terrain‑Segmentation‑using‑SegFormer.
Authors:Yijian Wang, Qingsen Yan, Jiantao Zhou, Duwei Dai, Wei Dong
Abstract:
Image Restoration (IR) agents, leveraging multimodal large language models to perceive degradation and invoke restoration tools, have shown promise in automating IR tasks. However, existing IR agents typically lack an insight summarization mechanism for past interactions, which results in an exhaustive search for the optimal IR tool. To address this limitation, we propose a portrait‑aware IR agent, dubbed PaAgent, which incorporates a self‑evolving portrait bank for IR tools and Retrieval‑Augmented Generation (RAG) to select a suitable IR tool for input. Specifically, to construct and evolve the portrait bank, the PaAgent continuously enriches it by summarizing the characteristics of various IR tools with restored images, selected IR tools, and degraded images. In addition, the RAG is employed to select the optimal IR tool for the input image by retrieving relevant insights from the portrait bank. Furthermore, to enhance PaAgent's ability to perceive degradation in complex scenes, we propose a subjective‑objective reinforcement learning strategy that considers both image quality scores and semantic insights in reward generation, which accurately provides the degradation information even under partial and non‑uniform degradation. Extensive experiments across 8 IR benchmarks, covering six single‑degradation and eight mixed‑degradation scenarios, validate PaAgent's superiority in addressing complex IR tasks. Our project page is \hrefhttps://wyjgr.github.io/PaAgent.htmlPaAgent.
Authors:Pengyu Zhang, Klim Zaporojets, Jie Liu, Jia-Hong Huang, Paul Groth
Abstract:
Multi‑Modal Knowledge Graphs (MMKGs) benefit from visual information, yet large‑scale image collection is hard to curate and often excludes ambiguous but relevant visuals (e.g., logos, symbols, abstract scenes). We present Beyond Images, an automatic data‑centric enrichment pipeline with optional human auditing. This pipeline operates in three stages: (1) large‑scale retrieval of additional entity‑related images, (2) conversion of all visual inputs into textual descriptions to ensure that ambiguous images contribute usable semantics rather than noise, and (3) fusion of multi‑source descriptions using a large language model (LLM) to generate concise, entity‑aligned summaries. These summaries replace or augment the text modality in standard MMKG models without changing their architectures or loss functions. Across three public MMKG datasets and multiple baseline models, we observe consistent gains (up to 7% Hits@1 overall). Furthermore, on a challenging subset of entities with visually ambiguous logos and symbols, converting images into text yields large improvements (201.35% MRR and 333.33% Hits@1). Additionally, we release a lightweight Text‑Image Consistency Check Interface for optional targeted audits, improving description quality and dataset reliability. Our results show that scaling image coverage and converting ambiguous visuals into text is a practical path to stronger MMKG completion. Code, datasets, and supplementary materials are available at https://github.com/pengyu‑zhang/Beyond‑Images.
Authors:Hisayuki Yokomizo, Taiki Miyanishi, Yan Gang, Shuhei Kurita, Nakamasa Inoue, Yusuke Iwasawa
Abstract:
Vision‑Language Models (VLMs) are increasingly applied to robotic perception and manipulation, yet their ability to infer physical properties required for manipulation remains limited. In particular, estimating the mass of real‑world objects is essential for determining appropriate grasp force and ensuring safe interaction. However, current VLMs lack reliable mass reasoning capabilities, and most existing benchmarks do not explicitly evaluate physical quantity estimation under realistic sensing conditions. In this work, we propose PhysQuantAgent, a framework for real‑world object mass estimation using VLMs, together with VisPhysQuant, a new benchmark dataset for evaluation. VisPhysQuant consists of RGB‑D videos of real objects captured from multiple viewpoints, annotated with precise mass measurements. To improve estimation accuracy, we introduce three visual prompting methods that enhance the input image with object detection, scale estimation, and cross‑sectional image generation to help the model comprehend the size and internal structure of the target object. Experiments show that visual prompting significantly improves mass estimation accuracy on real‑world data, suggesting the efficacy of integrating spatial reasoning with VLM knowledge for physical inference.
Authors:Zongshun Zhang, Yao Liu, Qiao Liu, Xuefeng Peng, Peiyuan Jiang, Jiaye Yang, Daibing Yao, Wei Lin
Abstract:
Video‑based lie detection aims to identify deceptive behaviors from visual cues. Despite recent progress, its core challenge lies in learning sparse yet discriminative representations. Deceptive signals are typically subtle and short‑lived, easily overwhelmed by redundant information, while individual and contextual variations introduce strong identity‑related noise. To address this issue, we propose GenLie, a Global‑Enhanced Lie Detection Network that performs local feature modeling under global supervision. Specifically, sparse and subtle deceptive cues are captured at the local level, while global supervision and optimization ensure robust and discriminative representations by suppressing identity‑related noise. Experiments on three public datasets, covering both high‑ and low‑stakes scenarios, show that GenLie consistently outperforms state‑of‑the‑art methods. Source code is available at https://github.com/AliasDictusZ1/GenLie.
Authors:Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Abstract:
The deployment of Multimodal Large Language Models (MLLMs) in agriculture is currently stalled by a critical trade‑off: the existing literature lacks the large‑scale agricultural datasets required for robust model development and evaluation, while current state‑of‑the‑art models lack the verified domain expertise necessary to reason across diverse taxonomies. To address these challenges, we propose the Vision‑to‑Verified‑Knowledge (V2VK) pipeline, a novel generative AI‑driven annotation framework that integrates visual captioning with web‑augmented scientific retrieval to autonomously generate the AgriMM benchmark, effectively eliminating biological hallucinations by grounding training data in verified phytopathological literature. The AgriMM benchmark contains over 3,000 agricultural classes and more than 607k VQAs spanning multiple tasks, including fine‑grained plant species identification, plant disease symptom recognition, crop counting, and ripeness assessment. Leveraging this verifiable data, we present AgriChat, a specialized MLLM that presents broad knowledge across thousands of agricultural classes and provides detailed agricultural assessments with extensive explanations. Extensive evaluation across diverse tasks, datasets, and evaluation conditions reveals both the capabilities and limitations of current agricultural MLLMs, while demonstrating AgriChat's superior performance over other open‑source models, including internal and external benchmarks. The results validate that preserving visual detail combined with web‑verified knowledge constitutes a reliable pathway toward robust and trustworthy agricultural AI. The code and dataset are publicly available at https://github.com/boudiafA/AgriChat .
Authors:Nimrod Shabtay, Moshe Kimhi, Artem Spector, Sivan Haray, Ehud Rivlin, Chaim Baskin, Raja Giryes, Eli Schwartz
Abstract:
Vision‑language models (VLMs) typically process images at a native high‑resolution, forcing a trade‑off between accuracy and computational efficiency: high‑resolution inputs capture fine details but incur significant computational costs, while low‑resolution inputs advocate for efficiency, they potentially miss critical visual information, like small text. We present AwaRes, a spatial‑on‑demand framework that resolves this accuracy‑efficiency trade‑off by operating on a low‑resolution global view and using tool‑calling to retrieve only high‑resolution segments needed for a given query. We construct supervised data automatically: a judge compares low‑ vs.\ high‑resolution answers to label whether cropping is needed, and an oracle grounding model localizes the evidence for the correct answer, which we map to a discrete crop set to form multi‑turn tool‑use trajectories. We train our framework with cold‑start SFT followed by multi‑turn GRPO with a composite reward that combines semantic answer correctness with explicit crop‑cost penalties. Project page: https://nimrodshabtay.github.io/AwaRes
Authors:Jisu Nam, Yicong Hong, Chun-Hao Paul Huang, Feng Liu, JoungBin Lee, Jiyoung Kim, Siyoon Jin, Yunsung Lee, Jaeyoon Jung, Suhwan Choi, Seungryong Kim, Yang Zhou
Abstract:
Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and long‑horizon 3D consistency. Most prior works treat user actions as abstract conditioning signals, overlooking the fundamental geometric coupling between actions and the 3D world, whereby actions induce relative camera motions that accumulate into a global camera pose within a 3D world. In this paper, we establish camera pose as a unifying geometric representation to jointly ground immediate action control and long‑term 3D consistency. First, we define a physics‑based continuous action space and represent user inputs in the Lie algebra to derive precise 6‑DoF camera poses, which are injected into the generative model via a camera embedder to ensure accurate action alignment. Second, we use global camera poses as spatial indices to retrieve relevant past observations, enabling geometrically consistent revisiting of locations during long‑horizon navigation. To support this research, we introduce a large‑scale dataset comprising 3,000 minutes of authentic human gameplay annotated with camera trajectories and textual descriptions. Extensive experiments show that our approach substantially outperforms state‑of‑the‑art interactive gaming world models in action controllability, long‑horizon visual quality, and 3D spatial consistency.
Authors:Lin Li, Haoran Feng, Zehuan Huang, Haohua Chen, Wenbo Nie, Shaohua Hou, Keqing Fan, Pan Hu, Sheng Wang, Buyu Li, Lu Sheng
Abstract:
We introduce SegviGen, a framework that repurposes native 3D generative models for 3D part segmentation. Existing pipelines either lift strong 2D priors into 3D via distillation or multi‑view mask aggregation, often suffering from cross‑view inconsistency and blurred boundaries, or explore native 3D discriminative segmentation, which typically requires large‑scale annotated 3D data and substantial training resources. In contrast, SegviGen leverages the structured priors encoded in pretrained 3D generative model to induce segmentation through distinctive part colorization, establishing a novel and efficient framework for part segmentation. Specifically, SegviGen encodes a 3D asset and predicts part‑indicative colors on active voxels of a geometry‑aligned reconstruction. It supports interactive part segmentation, full segmentation, and full segmentation with 2D guidance in a unified framework. Extensive experiments show that SegviGen improves over the prior state of the art by 40% on interactive part segmentation and by 15% on full segmentation, while using only 0.32% of the labeled training data. It demonstrates that pretrained 3D generative priors transfer effectively to 3D part segmentation, enabling strong performance with limited supervision. See our project page at https://fenghora.github.io/SegviGen‑Page/.
Authors:Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
Abstract:
Monocular 3D scene reconstruction has recently seen significant progress. Powered by the modern neural architectures and large‑scale data, recent methods achieve high performance in depth estimation from a single image. Meanwhile, reconstructing and decomposing common scenes into individual 3D objects remains a hard challenge due to the large variety of objects, frequent occlusions and complex object relations. Notably, beyond shape and pose estimation of individual objects, applications in robotics and animation require physically‑plausible scene reconstruction where objects obey physical principles of non‑penetration and realistic contacts. In this work we advance object‑level scene reconstruction along two directions. First, we introduceMessyKitchens, a new dataset with real‑world scenes featuring cluttered environments and providing high‑fidelity object‑level ground truth in terms of 3D object shapes, poses and accurate object contacts. Second, we build on the recent SAM 3D approach for single‑object reconstruction and extend it with Multi‑Object Decoder (MOD) for joint object‑level scene reconstruction. To validate our contributions, we demonstrate MessyKitchens to significantly improve previous datasets in registration accuracy and inter‑object penetration. We also compare our multi‑object reconstruction approach on three datasets and demonstrate consistent and significant improvements of MOD over the state of the art. Our new benchmark, code and pre‑trained models will become publicly available on our project website: https://messykitchens.github.io/.
Authors:Kerui Ren, Guanghao Li, Changjian Jiang, Yingxiang Xu, Tao Lu, Linning Xu, Junting Dong, Jiangmiao Pang, Mulin Yu, Bo Dai
Abstract:
Streaming reconstruction from uncalibrated monocular video remains challenging, as it requires both high‑precision pose estimation and computationally efficient online refinement in dynamic environments. While coupling 3D foundation models with SLAM frameworks is a promising paradigm, a critical bottleneck persists: most multi‑view foundation models estimate poses in a feed‑forward manner, yielding pixel‑level correspondences that lack the requisite precision for rigorous geometric optimization. To address this, we present M^3, which augments the Multi‑view foundation model with a dedicated Matching head to facilitate fine‑grained dense correspondences and integrates it into a robust Monocular Gaussian Splatting SLAM. M^3 further enhances tracking stability by incorporating dynamic area suppression and cross‑inference intrinsic alignment. Extensive experiments on diverse indoor and outdoor benchmarks demonstrate state‑of‑the‑art accuracy in both pose estimation and scene reconstruction. Notably, M^3 reduces ATE RMSE by 64.3% compared to VGGT‑SLAM 2.0 and outperforms ARTDECO by 2.11 dB in PSNR on the ScanNet++ dataset.
Authors:Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
Abstract:
Pixel‑space diffusion has recently re‑emerged as a strong alternative to latent diffusion, enabling high‑quality generation without pretrained autoencoders. However, standard pixel‑space diffusion models receive relatively weak semantic supervision and are not explicitly designed to capture high‑level visual structure. Recent representation‑alignment methods (e.g., REPA) suggest that pretrained visual features can substantially improve diffusion training, and visual co‑denoising has emerged as a promising direction for incorporating such features into the generative process. However, existing co‑denoising approaches often entangle multiple design choices, making it unclear which design choices are truly essential. Therefore, we present V‑Co, a systematic study of visual co‑denoising in a unified JiT‑based framework. This controlled setting allows us to isolate the ingredients that make visual co‑denoising effective. Our study reveals four key ingredients for effective visual co‑denoising. First, preserving feature‑specific computation while enabling flexible cross‑stream interaction motivates a fully dual‑stream architecture. Second, effective classifier‑free guidance (CFG) requires a structurally defined unconditional prediction. Third, stronger semantic supervision is best provided by a perceptual‑drifting hybrid loss. Fourth, stable co‑denoising further requires proper cross‑stream calibration, which we realize through RMS‑based feature rescaling. Together, these findings yield a simple recipe for visual co‑denoising. Experiments on ImageNet‑256 show that, at comparable model sizes, V‑Co outperforms the underlying pixel‑space diffusion baseline and strong prior pixel‑diffusion methods while using fewer training epochs, offering practical guidance for future representation‑aligned generative models.
Authors:Qiaosi Yi, Shuai Li, Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, Lei Zhang
Abstract:
Recently, reinforcement learning (RL) has been employed for improving generative image super‑resolution (ISR) performance. However, the current efforts are focused on multi‑step generative ISR, while one‑step generative ISR remains underexplored due to its limited stochasticity. In addition, RL methods such as Direct Preference Optimization (DPO) require the generation of positive and negative sample pairs offline, leading to a limited number of samples, while Group Relative Policy Optimization (GRPO) only calculates the likelihood of the entire image, ignoring local details that are crucial for ISR. In this paper, we propose Group Direct Preference Optimization (GDPO), a novel approach to integrate RL into one‑step generative ISR model training. First, we introduce a noise‑aware one‑step diffusion model that can generate diverse ISR outputs. To prevent performance degradation caused by noise injection, we introduce an unequal‑timestep strategy to decouple the timestep of noise addition from that of diffusion. We then present the GDPO strategy, which integrates the principle of GRPO into DPO, to calculate the group‑relative advantage of each online generated sample for model optimization. Meanwhile, an attribute‑aware reward function is designed to dynamically evaluate the score of each sample based on its statistics of smooth and texture areas. Experiments demonstrate the effectiveness of GDPO in enhancing the performance of one‑step generative ISR models. Code: https://github.com/Joyies/GDPO.
Authors:Chenggong Hu, Yi Wang, Mengqi Xue, Haofei Zhang, Jie Song, Li Sun
Abstract:
Textile pattern generation (TPG) aims to synthesize fine‑grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated TPG, existing image‑to‑image models appear to be natural candidates for this task. However, when applied directly, these methods often produce unfaithful results, failing to preserve fine‑grained details due to feature confusion between complex textile patterns and the inherent non‑rigid texture distortions in clothing images. In this paper, we propose a novel method, SLDDM‑TPG, for faithful and high‑fidelity TPG. Our method consists of two stages: (1) a latent disentangled network (LDN) that resolves feature confusion in clothing representations and constructs a multi‑dimensional, independent clothing feature space; and (2) a semi‑supervised latent diffusion model (S‑LDM), which receives guidance signals from LDN and generates faithful results through semi‑supervised diffusion training, combined with our designed fine‑grained alignment strategy. Extensive evaluations show that SLDDM‑TPG reduces FID by 4.1 and improves SSIM by up to 0.116 on our CTP‑HD dataset, and also demonstrate good generalization on the VITON‑HD dataset.
Authors:Guangzhi Xiong, Sanchit Sinha, Zhenghao He, Aidong Zhang
Abstract:
Vision‑language models (VLMs) have achieved impressive performance across a wide range of multimodal reasoning tasks, but they often struggle to disentangle fine‑grained visual attributes and reason about underlying causal relationships. In‑context learning (ICL) offers a promising avenue for VLMs to adapt to new tasks, but its effectiveness critically depends on the selection of demonstration examples. Existing retrieval‑augmented approaches typically rely on passive similarity‑based retrieval, which tends to select correlated but non‑causal examples, amplifying spurious associations and limiting model robustness. We introduce CIRCLES (Composed Image Retrieval for Causal Learning Example Selection), a novel framework that actively constructs demonstration sets by retrieving counterfactual‑style examples through targeted, attribute‑guided composed image retrieval. By incorporating counterfactual‑style examples, CIRCLES enables VLMs to implicitly reason about the causal relations between attributes and outcomes, moving beyond superficial correlations and fostering more robust and grounded reasoning. Comprehensive experiments on four diverse datasets demonstrate that CIRCLES consistently outperforms existing methods across multiple architectures, especially on small‑scale models, with pronounced gains under information scarcity. Furthermore, CIRCLES retrieves more diverse and causally informative examples, providing qualitative insights into how models leverage in‑context demonstrations for improved reasoning. Our code is available at https://github.com/gzxiong/CIRCLES.
Authors:Lukas Höllein, Matthias Nießner
Abstract:
Video diffusion models generate high‑quality and diverse worlds; however, individual frames often lack 3D consistency across the output sequence, which makes the reconstruction of 3D worlds difficult. To this end, we propose a new method that handles these inconsistencies by non‑rigidly aligning the video frames into a globally‑consistent coordinate frame that produces sharp and detailed pointcloud reconstructions. First, a geometric foundation model lifts each frame into a pixel‑wise 3D pointcloud, which contains unaligned surfaces due to these inconsistencies. We then propose a tailored non‑rigid iterative frame‑to‑model ICP to obtain an initial alignment across all frames, followed by a global optimization that further sharpens the pointcloud. Finally, we leverage this pointcloud as initialization for 3D reconstruction and propose a novel inverse deformation rendering loss to create high quality and explorable 3D environments from inconsistent views. We demonstrate that our 3D scenes achieve higher quality than baselines, effectively turning video models into 3D‑consistent world generators.
Authors:Mutian Xu, Tianbao Zhang, Tianqi Liu, Zhaoxi Chen, Xiaoguang Han, Ziwei Liu
Abstract:
Simulating robot‑world interactions is a cornerstone of Embodied AI. Recently, a few works have shown promise in leveraging video generations to transcend the rigid visual/physical constraints of traditional simulators. However, they primarily operate in 2D space or are guided by static environmental cues, ignoring the fundamental reality that robot‑world interactions are inherently 4D spatiotemporal events that require precise interactive modeling. To restore this 4D essence while ensuring the precise robot control, we introduce Kinema4D, a new action‑conditioned 4D generative robotic simulator that disentangles the robot‑world interaction into: i) Precise 4D representation of robot controls: we drive a URDF‑based 3D robot via kinematics, producing a precise 4D robot control trajectory. ii) Generative 4D modeling of environmental reactions: we project the 4D robot trajectory into a pointmap as a spatiotemporal visual signal, controlling the generative model to synthesize complex environments' reactive dynamics into synchronized RGB/pointmap sequences. To facilitate training, we curated a large‑scale dataset called Robo4D‑200k, comprising 201,426 robot interaction episodes with high‑quality 4D annotations. Extensive experiments demonstrate that our method effectively simulates physically‑plausible, geometry‑consistent, and embodiment‑agnostic interactions that faithfully mirror diverse real‑world dynamics. For the first time, it shows potential zero‑shot transfer capability, providing a high‑fidelity foundation for advancing next‑generation embodied simulation.
Authors:Tianyuan Yuan, Zibin Dong, Yicheng Liu, Hang Zhao
Abstract:
World Action Models (WAMs) have emerged as a promising alternative to Vision‑Language‑Action (VLA) models for embodied control because they explicitly model how visual observations may evolve under action. Most existing WAMs follow an imagine‑then‑execute paradigm, incurring substantial test‑time latency from iterative video denoising, yet it remains unclear whether explicit future imagination is actually necessary for strong action performance. In this paper, we ask whether WAMs need explicit future imagination at test time, or whether their benefit comes primarily from video modeling during training. We disentangle the role of video modeling during training from explicit future generation during inference by proposing Fast‑WAM, a WAM architecture that retains video co‑training during training but skips future prediction at test time. We further instantiate several Fast‑WAM variants to enable a controlled comparison of these two factors. Across these variants, we find that Fast‑WAM remains competitive with imagine‑then‑execute variants, while removing video co‑training causes a much larger performance drop. Empirically, Fast‑WAM achieves competitive results with state‑of‑the‑art methods both on simulation benchmarks (LIBERO and RoboTwin) and real‑world tasks, without embodied pretraining. It runs in real time with 190ms latency, over 4× faster than existing imagine‑then‑execute WAMs. These results suggest that the main value of video prediction in WAMs may lie in improving world representations during training rather than generating future observations at test time. Project page: https://yuantianyuan01.github.io/FastWAM/
Authors:Md Jahidul Islam
Abstract:
Adapting large‑scale Vision‑Language Models (VLMs) like CLIP to downstream tasks often suffers from a "one‑size‑fits‑all" architectural approach, where visual and textual tokens are processed uniformly by wide, generic adapters. We argue that this homogeneity ignores the distinct structural nature of the modalities ‑‑ spatial locality in images versus semantic density in text. To address this, we propose HeBA (Heterogeneous Bottleneck Adapter), a unified architectural framework that introduces modality‑specific structural inductive biases. HeBA departs from conventional designs through three key architectural innovations: (1) Heterogeneity: It processes visual tokens via 2D depthwise‑separable convolutions to preserve spatial correlations, while distinctively processing text tokens via dense linear projections to capture semantic relationships; (2) Bottleneck Regularization: Unlike standard expanding adapters, HeBA employs a compression bottleneck (D ‑> D/4) that explicitly forces the model to learn compact, robust features and acts as a structural regularizer; and (3) Active Gradient Initialization: We challenge the restrictive zero‑initialization paradigm, utilizing a Kaiming initialization strategy that ensures sufficient initial gradient flow to accelerate convergence without compromising the frozen backbone's pre‑trained knowledge. Extensive experiments demonstrate that HeBA's architecturally specialized design achieves superior stability and accuracy, establishing a new state‑of‑the‑art on 11 few‑shot benchmarks. Code is available at https://github.com/Jahid12012021/VLM‑HeBA.
Authors:Shihao Zhu, Ziheng Ouyang, Yijia Kang, Qilong Wang, Mi Zhou, Bo Li, Ming-Ming Cheng, Qibin Hou
Abstract:
Diffusion‑based stylization has advanced significantly, yet existing methods are limited to color‑driven transformations, neglecting complex semantics and material details. We introduce StyleExpert, a semantic‑aware framework based on the Mixture of Experts (MoE). Our framework employs a unified style encoder, trained on our large‑scale dataset of content‑style‑stylized triplets, to embed diverse styles into a consistent latent space. This embedding is then used to condition a similarity‑aware gating mechanism, which dynamically routes styles to specialized experts within the MoE architecture. Leveraging this MoE architecture, our method adeptly handles diverse styles spanning multiple semantic levels, from shallow textures to deep semantics. Extensive experiments show that StyleExpert outperforms existing approaches in preserving semantics and material details, while generalizing to unseen styles. Our code and collected images are available at the project page: https://hh‑lg.github.io/StyleExpert‑Page/.
Authors:Melissa Schween, Mathis Kruse, Bodo Rosenhahn
Abstract:
We propose Bijective Universal Scene‑Specific Anomalous Relationship Detection (BUSSARD), a normalizing flow‑based model for detecting anomalous relations in scene graphs, generated from images. Our work follows a multimodal approach, embedding object and relationship tokens from scene graphs with a language model to leverage semantic knowledge from the real world. A normalizing flow model is used to learn bijective transformations that map object‑relation‑object triplets from scene graphs to a simple base distribution (typically Gaussian), allowing anomaly detection through likelihood estimation. We evaluate our approach on the SARD dataset containing office and dining room scenes. Our method achieves around 10% better AUROC results compared to the current state‑of‑the‑art model, while simultaneously being five times faster. Through ablation studies, we demonstrate superior robustness and universality, particularly regarding the use of synonyms, with our model maintaining stable performance while the baseline shows 17.5% deviation. This work demonstrates the strong potential of learning‑based methods for relationship anomaly detection in scene graphs. Our code is available at https://github.com/mschween/BUSSARD .
Authors:Redwan Sony, Anil K Jain, Arun Ross
Abstract:
Multimodal Large Language Models (MLLMs) have recently been proposed as a means to generate natural‑language explanations for face recognition decisions. While such explanations facilitate human interpretability, their reliability on unconstrained face images remains underexplored. In this work, we systematically analyze MLLM‑generated explanations for the unconstrained face verification task on the challenging IJB‑S dataset, with a particular focus on extreme pose variation and surveillance imagery. Our results show that even when MLLMs produce correct verification decisions, the accompanying explanations frequently rely on non‑verifiable or hallucinated facial attributes that are not supported by visual evidence. We further study the effect of incorporating information from traditional face recognition systems, viz., scores and decisions, alongside the input images. Although such information improves categorical verification performance, it does not consistently lead to faithful explanations. To evaluate the explanations beyond decision accuracy, we introduce a likelihood‑ratio‑based framework that measures the evidential strength of textual explanations. Our findings highlight fundamental limitations of current MLLMs for explainable face recognition and underscore the need for a principled evaluation of reliable and trustworthy explanations in biometric applications. Code is available at https://github.com/redwankarimsony/LR‑MLLMFR‑Explainability.
Authors:Weiqin Jiao, Hao Cheng, George Vosselman, Claudio Persello
Abstract:
We tackle the problem of generating a complete vector map representation from aerial imagery in a single run: producing polygons for all land‑cover classes with shared boundaries and without gaps or overlaps. Existing polygonization methods are typically class‑specific; extending them to multiple classes via per‑class runs commonly leads to topological inconsistencies, such as duplicated edges, gaps, and overlaps. We formalize this new task as All‑Class Polygonal Vectorization (ACPV) and release the first public benchmark, Deventer‑512, with standardized metrics jointly evaluating semantic fidelity, geometric accuracy, vertex efficiency, per‑class topological fidelity and global topological consistency. To realize ACPV, we propose ACPV‑Net, a unified framework introducing a novel Semantically Supervised Conditioning (SSC) mechanism coupling semantic perception with geometric primitive generation, along with a topological reconstruction that enforces shared‑edge consistency by design. While enforcing such strict topological constraints, ACPV‑Net surpasses all class‑specific baselines in polygon quality across classes on Deventer‑512. It also applies to single‑class polygonal vectorization without any architectural modification, achieving the best‑reported results on WHU‑Building. Data, code, and models will be released at: https://github.com/HeinzJiao/ACPV‑Net.
Authors:Weijie Qiu, Dai Guan, Junxin Wang, Zhihang Li, Yongbo Gai, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang
Abstract:
Generative reward models (GRMs) for vision‑language models (VLMs) often evaluate outputs via a three‑stage pipeline: rubric generation, criterion‑based scoring, and a final verdict. However, the intermediate rubric is rarely optimized directly. Prior work typically either treats rubrics as incidental or relies on expensive LLM‑as‑judge checks that provide no differentiable signal and limited training‑time guidance. We propose Proxy‑GRM, which introduces proxy‑guided rubric verification into Reinforcement Learning (RL) to explicitly enhance rubric quality. Concretely, we train lightweight proxy agents (Proxy‑SFT and Proxy‑RL) that take a candidate rubric together with the original query and preference pair, and then predict the preference ordering using only the rubric as evidence. The proxy's prediction accuracy serves as a rubric‑quality reward, incentivizing the model to produce rubrics that are internally consistent and transferable. With ~50k data samples, Proxy‑GRM reaches state‑of‑the‑art results on the VL‑Reward Bench, Multimodal Reward Bench, and MM‑RLHF‑Reward Bench, outperforming the methods trained on four times the data. Ablations show Proxy‑SFT is a stronger verifier than Proxy‑RL, and implicit reward aggregation performs best. Crucially, the learned rubrics transfer to unseen evaluators, improving reward accuracy at test time without additional training. Our code is available at https://github.com/Qwen‑Applications/Proxy‑GRM.
Authors:Fangjing Li, Zhihai Wang, Xinxin Ding, Haiyang Liu, Ronghua Gao, Rong Wang, Yao Zhu, Ming Jin
Abstract:
Mounting posture is an important visual indicator of estrus in dairy cattle. However, achieving reliable mounting pose estimation in real‑world environments remains challenging due to cluttered backgrounds and frequent inter‑animal occlusion. We present FSMC‑Pose, a top‑down framework that integrates a lightweight frequency‑spatial fusion backbone, CattleMountNet, and a multiscale self‑calibration head, SC2Head. Specifically, we design two algorithmic components for CattleMountNet: the Spatial Frequency Enhancement Block (SFEBlock) and the Receptive Aggregation Block (RABlock). SFEBlock separates cattle from cluttered backgrounds, while RABlock captures multiscale contextual information. The Spatial‑Channel Self‑Calibration Head (SC2Head) attends to spatial and channel dependencies and introduces a self‑calibration branch to mitigate structural misalignment under inter‑animal overlap. We construct a mounting dataset, MOUNT‑Cattle, covering 1176 mounting instances, which follows the COCO format and supports drop‑in training across pose estimation models. Using a comprehensive dataset that combines MOUNT‑Cattle with the public NWAFU‑Cattle dataset, FSMC‑Pose achieves higher accuracy than strong baselines, with markedly lower computational and parameter costs, while maintaining real‑time inference on commodity GPUs. Extensive experiments and qualitative analyses show that FSMC‑Pose effectively captures and estimates cattle mounting pose in complex and cluttered environments. Dataset and code are available at https://github.com/elianafang/FSMC‑Pose.
Authors:Yong Zou, Haoran Li, Fanxiao Li, Shenyang Wei, Yunyun Dong, Li Tang, Wei Zhou, Renyang Liu
Abstract:
Recent progress in image generation models (IGMs) enables high‑fidelity content creation but also amplifies risks, including the reproduction of copyrighted content and the generation of offensive content. Image Generation Model Unlearning (IGMU) mitigates these risks by removing harmful concepts without full retraining. Despite growing attention, the robustness under adversarial inputs, particularly image‑side threats in black‑box settings, remains underexplored. To bridge this gap, we present REFORGE, a black‑box red‑teaming framework that evaluates IGMU robustness via adversarial image prompts. REFORGE initializes stroke‑based images and optimizes perturbations with a cross‑attention‑guided masking strategy that allocates noise to concept‑relevant regions, balancing attack efficacy and visual fidelity. Extensive experiments across representative unlearning tasks and defenses demonstrate that REFORGE significantly improves attack success rate while achieving stronger semantic alignment and higher efficiency than involved baselines. These results expose persistent vulnerabilities in current IGMU methods and highlight the need for robustness‑aware unlearning against multi‑modal adversarial attacks. Our code is at: https://github.com/Imfatnoily/REFORGE.
Authors:Florian Bürger, Martim Dias Gomes, Adrián E. Granada, Noémie Moreau, Katarzyna Bozek
Abstract:
Understanding non‑genetic determinants of cell fate is critical for developing and improving cancer therapies, as genetically identical cells can exhibit divergent outcomes under the same treatment conditions. In this work, we present a deep learning approach for cell fate prediction from raw long‑term live‑cell recordings of cancer cell populations under chemotherapeutic treatment. Our Transformer model is trained to predict cell fate directly from raw image sequences, without relying on predefined morphological or molecular features. Beyond classification, we introduce a comprehensive explainability framework for interpreting the temporal and morphological cues guiding the model's predictions. We demonstrate that prediction of cell outcomes is possible based on the video only, our model achieves balanced accuracy of 0.94 and an F1‑score of 0.93. Attention and masking experiments further indicate that the signal predictive of the cell fate is not uniquely located in the final frames of a cell trajectory, as reliable predictions are possible up to 10 h before the event. Our analysis reveals distinct temporal distribution of predictive information in the mitotic and apoptotic sequences, as well as the role of cell morphology and p53 signaling in determining cell outcomes. Together, these findings demonstrate that attention‑based temporal models enable accurate cell fate prediction while providing biologically interpretable insights into non‑genetic determinants of cellular decision‑making. The code is available at https://github.com/bozeklab/Cell‑Fate‑Prediction.
Authors:Kaiwen Song, Jinkai Cui, Juyong Zhang
Abstract:
In practical real‑time XR and telepresence applications, network and computing resources fluctuate frequently. Therefore, a progressive 3D representation is needed. To this end, we propose ProgressiveAvatars, a progressive avatar representation built on a hierarchy of 3D Gaussians grown by adaptive implicit subdivision on a template mesh. 3D Gaussians are defined in face‑local coordinates to remain animatable under varying expressions and head motion across multiple detail levels. The hierarchy expands when screen‑space signals indicate a lack of detail, allocating resources to important areas. Leveraging importance ranking, ProgressiveAvatars supports incremental loading and rendering, adding new Gaussians as they arrive while preserving previous content, thus achieving smooth quality improvements across varying bandwidths. ProgressiveAvatars enables progressive delivery and progressive rendering under fluctuating network bandwidth and varying compute and memory resources.
Authors:Hunain Ahmed Jillani, Ahmed Tawfik Aboukhadra, Ahmed Elhayek, Jameel Malik, Nadia Robertini, Didier Stricker
Abstract:
Fast and accurate 3D hand reconstruction is essential for real‑time applications in VR/AR, human‑computer interaction, robotics, and healthcare. Most state‑of‑the‑art methods rely on heavy models, limiting their use on resource‑constrained devices like headsets, smartphones, and embedded systems. In this paper, we investigate how the use of lightweight neural networks, combined with Knowledge Distillation, can accelerate complex 3D hand reconstruction models by making them faster and lighter, while maintaining comparable reconstruction accuracy. While our approach is suited for various hand reconstruction frameworks, we focus primarily on boosting the HaMeR model, currently the leading method in terms of reconstruction accuracy. We replace its original ViT‑H backbone with lighter alternatives, including MobileNet, MobileViT, ConvNeXt, and ResNet, and evaluate three knowledge distillation strategies: output‑level, feature‑level, and a hybrid of both. Our experiments show that using lightweight backbones that are only 35% the size of the original achieves 1.5x faster inference speed while preserving similar performance quality with only a minimal accuracy difference of 0.4mm. More specifically, we show how output‑level distillation notably improves student performance, while feature‑level distillation proves more effective for higher‑capacity students. Overall, the findings pave the way for efficient real‑world applications on low‑power devices. The code and models are publicly available under https://github.com/hunainahmedj/Fast‑HaMeR.
Authors:Joona Kareinen, Veikka Immonen, Tuomas Eerola, Lumi Haraguchi, Lasse Lensu, Kaisa Kraft, Sanna Suikkanen, Heikki Kälviäinen
Abstract:
This paper considers self‑supervised cross‑modal coordination as a strategy enabling utilization of multiple modalities and large volumes of unlabeled plankton data to build models for plankton recognition. Automated imaging instruments facilitate the continuous collection of plankton image data on a large scale. Current methods for automatic plankton image recognition rely primarily on supervised approaches, which require labeled training sets that are labor‑intensive to collect. On the other hand, some modern plankton imaging instruments complement image information with optical measurement data, such as scatter and fluorescence profiles, which currently are not widely utilized in plankton recognition. In this work, we explore the possibility of using such measurement data to guide the learning process without requiring manual labeling. Inspired by the concepts behind Contrastive Language‑Image Pre‑training, we train encoders for both modalities using only binary supervisory information indicating whether a given image and profile originate from the same particle or from different particles. For plankton recognition, we employ a small labeled gallery of known plankton species combined with a k‑NN classifier. This approach yields a recognition model that is inherently multimodal, i.e., capable of utilizing information extracted from both image and profile data. We demonstrate that the proposed method achieves high recognition accuracy while requiring only a minimal number of labeled images. Furthermore, we show that the approach outperforms an image‑only self‑supervised baseline. Code available at https://github.com/Jookare/cross‑modal‑plankton.
Authors:Jing Dai, Chen Wu, Ming Wu, Qibin Zhang, Zexi Wu, Jingdong Zhang, Hongming Xu
Abstract:
Recent advances in multimodal learning have significantly improved cancer survival risk prediction. However, the joint prognostic potential of protein markers and histopathology images remains underexplored, largely due to the high cost and limited availability of protein expression profiling. To address this challenge, we propose HGP‑Mamba, a Mamba‑based multimodal framework that efficiently integrates histological with generated protein features for survival risk prediction. Specifically, we introduce a protein feature extractor (PFE) that leverages pretrained foundation models to derive high‑throughput protein embeddings directly from Whole Slide Images (WSIs), enabling data‑efficient incorporation of molecular information. Together with histology embeddings that capture morphological patterns, we further introduce the Local Interaction‑aware Mamba (LiAM) for fine‑grained feature interaction and the Global Interaction‑enhanced Mamba (GiEM) to promote holistic modality fusion at the slide level, thus capture complex cross‑modal dependencies. Experiments on four public cancer datasets demonstrate that HGP‑Mamba achieves state‑of‑the‑art performance while maintaining superior computational efficiency compared with existing methods. Our source code is publicly available at https://github.com/Daijing‑ai/HGP‑Mamba.git.
Authors:Zhan Tong, ChenXu Zhou, Fei Tang, Yiming Tu, Tianyu Qin, Kaihao Fang
Abstract:
Defense Meteorological Satellite Program (DMSP‑OLS) and Suomi National Polar‑orbiting Partnership (SNPP‑VIIRS) nighttime light (NTL) data are vital for monitoring urbanization, yet sensor incompatibilities hinder long‑term analysis. This study proposes a cross‑sensor calibration method using Contrastive Unpaired Translation (CUT) network to transform DMSP data into VIIRS‑like format, correcting DMSP defects. The method employs multilayer patch‑wise contrastive learning to maximize mutual information between corresponding patches, preserving content consistency while learning cross‑domain similarity. Utilizing 2012‑2013 overlapping data for training, the network processes 1992‑2013 DMSP imagery to generate enhanced VIIRS‑style raster data. Validation results demonstrate that generated VIIRS‑like data exhibits high consistency with actual VIIRS observations (R‑squared greater than 0.87) and socioeconomic indicators. This approach effectively resolves cross‑sensor data fusion issues and calibrates DMSP defects, providing reliable attempt for extended NTL time‑series.
Authors:Zhengbo Zhang, Jinbo Su, Zhaowen Zhou, Changtao Miao, Yuhan Hong, Qimeng Wu, Yumeng Liu, Feier Wu, Yihe Tian, Yuhao Liang, Zitong Shan, Wanke Xia, Yi-Fan Zhang, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
Abstract:
The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the real world. But existing benchmarks suffer from two limitations: insufficient evaluation of visual reasoning ability and the neglect of native visual information of web pages in the reasoning chains. To address these challenges, we introduce a new benchmark for visual‑native search, VisBrowse‑Bench. It contains 169 VQA instances covering multiple domains and evaluates the models' visual reasoning capabilities during the search process through multimodal evidence cross‑validation via text‑image retrieval and joint reasoning. These data were constructed by human experts using a multi‑stage pipeline and underwent rigorous manual verification. We additionally propose an agent workflow that can effectively drive the browsing agent to actively collect and reason over visual information during the search process. We comprehensively evaluated both open‑source and closed‑source models in this workflow. Experimental results show that even the best‑performing model, Claude‑4.6‑Opus only achieves an accuracy of 47.6%, while the proprietary Deep Research model, o3‑deep‑research only achieves an accuracy of 41.1%. The code and data can be accessed at: https://github.com/ZhengboZhang/VisBrowse‑Bench
Authors:Tiantian Dang, Chao Bi, Shufan Shen, Jinzhe Liu, Qingming Huang, Shuhui Wang
Abstract:
Despite the significant advancements in Large Vision‑Language Models (LVLMs), their tendency to generate hallucinations undermines reliability and restricts broader practical deployment. Among the hallucination mitigation methods, feature steering emerges as a promising approach that reduces erroneous outputs in LVLMs without increasing inference costs. However, current methods apply uniform feature steering across all layers. This heuristic strategy ignores inter‑layer differences, potentially disrupting layers unrelated to hallucinations and ultimately leading to performance degradation on general tasks. In this paper, we propose Locate‑Then‑Sparsify for Feature Steering (LTS‑FS), a plug‑and‑play framework which controls the steering intensity according to the hallucination relevance of each layer. We first construct a dataset comprising token‑level and sentence‑level hallucination cases. Based on this dataset, we introduce an attribution method based on causal interventions to quantify the hallucination relevance of each layer. With the attribution scores across layers, we propose a layerwise strategy that converts these scores into feature steering intensities for individual layers, enabling more precise adjustments specifically on hallucination‑relevant layers. Extensive experiments across multiple LVLMs and benchmarks demonstrate that LTS‑FS effectively mitigates hallucination while preserving strong performance. Codes are available at https://github.com/huttersadan/LTS‑FS.
Authors:Hongwei Lin, Xun Huang, Chenglu Wen, Cheng Wang
Abstract:
Robust 3D object detection under adverse weather conditions is crucial for autonomous driving. However, most existing methods simply combine all weather samples for training while overlooking data distribution discrepancies across different weather scenarios, leading to performance conflicts. To address this issue, we introduce AW‑MoE, the framework that innovatively integrates Mixture of Experts (MoE) into weather‑robust multi‑modal 3D object detection approaches. AW‑MoE incorporates Image‑guided Weather‑aware Routing (IWR), which leverages the superior discriminability of image features across weather conditions and their invariance to scene variations for precise weather classification. Based on this accurate classification, IWR selects the top‑K most relevant Weather‑Specific Experts (WSE) that handle data discrepancies, ensuring optimal detection under all weather conditions. Additionally, we propose a Unified Dual‑Modal Augmentation (UDMA) for synchronous LiDAR and 4D Radar dual‑modal data augmentation while preserving the realism of scenes. Extensive experiments on the real‑world dataset demonstrate that AW‑MoE achieves ~ 15% improvement in adverse‑weather performance over state‑of‑the‑art methods, while incurring negligible inference overhead. Moreover, integrating AW‑MoE into established baseline detectors yields performance improvements surpassing current state‑of‑the‑art methods. These results show the effectiveness and strong scalability of our AW‑MoE. We will release the code publicly at https://github.com/windlinsherlock/AW‑MoE.
Authors:Weihua Gao, Wenlong Niu, Jie Tang, Man Yang, Jiafeng Zhang, Xiaodong Peng
Abstract:
Infrared small target detection (IRSTD) methods predominantly formulate the task as pixel‑level segmentation, which requires costly dense annotations and is not well suited to tiny targets with weak texture and ambiguous boundaries. To address this issue, we propose Point‑to‑Mask, a framework that bridges low‑cost point supervision and mask‑level detection through two components: a Physics‑driven Adaptive Mask Generation (PAMG) module that converts point annotations into compact target masks and geometric cues, and a lightweight Radius‑aware Point Regression Network (RPR‑Net) that reformulates IRSTD as target center localization and effective radius regression using spatiotemporal motion cues. The two modules form a closed loop: PAMG generates pseudo masks and geometric supervision during training, while the geometric predictions of RPR‑Net are fed back to PAMG for pixel‑level mask recovery during inference. To facilitate systematic evaluation, we further construct SIRSTD‑Pixel, a sequential dataset with refined pixel‑level annotations. Experiments show that the proposed framework achieves strong pseudo‑label quality, high detection accuracy, and efficient inference, approaching full‑supervision performance under point‑supervised settings with substantially lower annotation cost. Code and datasets will be available at: https://github.com/GaoScience/point‑to‑mask.
Authors:Junxin Wang, Dai Guan, Weijie Qiu, Zhihang Li, Yongbo Gai, Zhengyi Yang, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang
Abstract:
Vision‑language process reward models (VL‑PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test‑time scaling. However, they often function as black‑box judges: a low step score may reflect a genuine reasoning mistake or simply the verifier's misperception of the image. This entanglement between perception and reasoning leads to systematic false positives (rewarding hallucinated visual premises) and false negatives (penalizing correct grounded statements), undermining both reranking and error localization. We introduce Explicit Visual Premise Verification (EVPV), a lightweight verification interface that conditions step scoring on the reliability of the visual premises a step depends on. The policy is prompted to produce a step‑wise visual checklist that makes required visual facts explicit, while a constraint extractor independently derives structured visual constraints from the input image. EVPV matches checklist claims against these constraints to compute a scalar visual reliability signal, and calibrates PRM step rewards via reliability gating: rewards for visually dependent steps are attenuated when reliability is low and preserved when reliability is high. This decouples perceptual uncertainty from logical evaluation without per‑step tool calls. Experiments on VisualProcessBench and six multimodal reasoning benchmarks show that EVPV improves step‑level verification and consistently boosts Best‑of‑N reranking accuracy over strong baselines. Furthermore, injecting controlled corruption into the extracted constraints produces monotonic performance degradation, providing causal evidence that the gains arise from constraint fidelity and explicit premise verification rather than incidental prompt effects. Code is available at: https://github.com/Qwen‑Applications/EVPV‑PRM
Authors:Duc T. Nguyen, Hoang-Long Nguyen, Huy-Hieu Pham
Abstract:
Automated white blood cell (WBC) classification is essential for leukemia screening but remains challenged by extreme class imbalance, long‑tail distributions, and domain shift, leading deep models to overfit dominant classes and fail on rare subtypes. We propose a hybrid framework for rare‑class generalization that integrates a generative Pix2Pix‑based restoration module for artifact removal, a Swin Transformer ensemble with MedSigLIP contrastive embeddings for robust representation learning, and a biologically‑inspired refinement step using geometric spikiness and Mahalanobis‑based morphological constraints to recover out‑of‑distribution predictions. Evaluated on the WBCBench 2026 challenge, our method achieves a Macro‑F1 of 0.77139 on the private leaderboard, demonstrating strong performance under severe imbalance and highlighting the value of incorporating biological priors into deep learning for hematological image analysis. The code is available at https://github.com/trongduc‑nguyen/WBCBench2026
Authors:Ryutaro Miya, Kazuyoshi Fushinobu, Tatsuya Kawaguchi
Abstract:
We propose PureCLIP‑Depth, a completely prompt‑free, decoder‑free Monocular Depth Estimation (MDE) model that operates entirely within the Contrastive Language‑Image Pre‑training (CLIP) embedding space. Unlike recent models that rely heavily on geometric features, we explore a novel approach to MDE driven by conceptual information, performing computations directly within the conceptual CLIP space. The core of our method lies in learning a direct mapping from the RGB domain to the depth domain strictly inside this embedding space. Our approach achieves state‑of‑the‑art performance among CLIP embedding‑based models on both indoor and outdoor datasets. The code used in this research is available at: https://github.com/ryutaroLF/PureCLIP‑Depth
Authors:Haodong Yan, Zhide Zhong, Jiaguan Zhu, Junjie He, Weilin Yuan, Wenxuan Song, Xin Gong, Yingjie Cai, Guanyi Zhao, Xu Yan, Bingbing Liu, Ying-Cong Chen, Haoang Li
Abstract:
Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi‑step video generation or noisy one‑step feature extraction, cannot simultaneously guarantee real‑time inference and high‑fidelity foresight. To address this limitation, we propose S‑VAM, a shortcut video‑action model that foresees coherent geometric and semantic representations via a single forward pass. Serving as a stable blueprint, these foreseen representations significantly simplify the action prediction. To enable this efficient shortcut, we introduce a novel self‑distillation strategy that condenses structured generative priors of multi‑step denoising into one‑step inference. Specifically, vision foundation model (VFM) representations extracted from the diffusion model's own multi‑step generated videos provide teacher targets. Lightweight decouplers, as students, learn to directly map noisy one‑step features to these targets. Extensive experiments in simulation and the real world demonstrate that our S‑VAM outperforms state‑of‑the‑art methods, enabling efficient and precise manipulation in complex environments. Our project page is https://haodong‑yan.github.io/S‑VAM/
Authors:Peng Sun, Jun Xie, Tao Lin
Abstract:
Unified Multimodal Models (UMMs) are often constrained by the pre‑training of their visual generation components, which typically relies on inefficient paradigms and scarce, high‑quality text‑image paired data. In this paper, we systematically analyze pre‑training recipes for UMM visual generation and identify these two issues as the major bottlenecks.
To address them, we propose Image‑Only Training for UMMs (IOMM), a data‑efficient two‑stage training framework.
The first stage pre‑trains the visual generative component exclusively using abundant unlabeled image‑only data, thereby removing the dependency on paired data for this costly phase. The second stage fine‑tunes the model using a mixture of unlabeled images and a small curated set of text‑image pairs, leading to improved instruction alignment and generative quality.
Extensive experiments show that IOMM not only improves training efficiency but also achieves state‑of‑the‑art (SOTA) performance.
For example, our IOMM‑B (3.6B) model was trained from scratch using only ~ 1050 H800 GPU hours (with the vast majority, 1000 hours, dedicated to the efficient image‑only pre‑training stage). It achieves 0.89 on GenEval and 0.55 on WISE‑‑surpassing strong baselines such as BAGEL‑7B (0.82 & 0.55) and BLIP3‑o‑4B (0.84 & 0.50).
Code is available \hrefhttps://github.com/LINs‑lab/IOMMhttps://github.com/LINs‑lab/IOMM.
Authors:Zhiwei Wang, Yayu Zheng, Defeng He, Li Zhao, Xiaoqin Zhang, Yuxing Li, Edmund Y. Lam
Abstract:
Overexposure frequently occurs in practical scenarios, causing the loss of critical visual information. However, existing infrared and visible fusion methods still exhibit unsatisfactory performance in highly bright regions. To address this, we propose EPOFusion, an exposure‑aware fusion model. Specifically, a guidance module is introduced to facilitate the encoder in extracting fine‑grained infrared features from overexposed regions. Meanwhile, an iterative decoder incorporating a multiscale context fusion module is designed to progressively enhance the fused image, ensuring consistent details and superior visual quality. Finally, an adaptive loss function dynamically constrains the fusion process, enabling an effective balance between the modalities under varying exposure conditions. To achieve better exposure awareness, we construct the first infrared and visible overexposure dataset (IVOE) with high quality infrared guided annotations for overexposed regions. Extensive experiments show that EPOFusion outperforms existing methods. It maintains infrared cues in overexposed regions while achieving visually faithful fusion in non‑overexposed areas, thereby enhancing both visual fidelity and downstream task performance. Code, fusion results and IVOE dataset will be made available at https://github.com/warren‑wzw/EPOFusion.git.
Authors:Minbing Chen, Zhu Meng, Fei Su
Abstract:
Vision‑Language Models (VLMs) offer significant potential in computational pathology by enabling interpretable image analysis, automated reporting, and scalable decision support. However, their widespread clinical adoption remains limited due to the absence of reliable, automated evaluation metrics capable of identifying subtle failures such as hallucinations. To address this gap, we propose PathGLS, a novel reference‑free evaluation framework that assesses pathology VLMs across three dimensions: Grounding (fine‑grained visual‑text alignment), Logic (entailment graph consistency using Natural Language Inference), and Stability (output variance under adversarial visual‑semantic perturbations). PathGLS supports both patch‑level and whole‑slide image (WSI)‑level analysis, yielding a comprehensive trust score. Experiments on Quilt‑1M, TCGA, REG2025, PathMMU and TCGA‑Sarcoma datasets demonstrate the superiority of PathGLS. Specifically, on the Quilt‑1M dataset, PathGLS reveals a steep sensitivity drop of 40.2% for hallucinated reports compared to only 2.1% for BERTScore. Moreover, validation against expert‑defined clinical error hierarchies reveals that PathGLS achieves a strong Spearman's rank correlation of ρ=0.71 (p < 0.0001), significantly outperforming Large Language Model (LLM)‑based approaches (Gemini 3.0 Pro: ρ=0.39, p < 0.0001). These results establish PathGLS as a robust reference‑free metric. By directly quantifying hallucination rates and domain shift robustness, it serves as a reliable criterion for benchmarking VLMs on private clinical datasets and informing safe deployment. Code can be found at: https://github.com/My13ad/PathGLS
Authors:Butian Xiong, Rong Liu, Tiantian Zhou, Meida Chen, Zhiwen Fan, Andrew Feng
Abstract:
3D Gaussian Splat (3DGS) enables high‑fidelity, real‑time novel view synthesis by representing scenes with large sets of anisotropic primitives, but often requires millions of Splats, incurring significant storage and transmission costs. Most existing compression methods rely on GPU‑intensive post‑training optimization with calibrated images, limiting practical deployment. We introduce NanoGS, a training‑free and lightweight framework for Gaussian Splat simplification. Instead of relying on image‑based rendering supervision, NanoGS formulates simplification as local pairwise merging over a sparse spatial graph. The method approximates a pair of Gaussians with a single primitive using mass preserved moment matching and evaluates merge quality through a principled merge cost between the original mixture and its approximation. By restricting merge candidates to local neighborhoods and selecting compatible pairs efficiently, NanoGS produces compact Gaussian representations while preserving scene structure and appearance. NanoGS operates directly on existing Gaussian Splat models, runs efficiently on CPU, and preserves the standard 3DGS parameterization, enabling seamless integration with existing rendering pipelines. Experiments demonstrate that NanoGS substantially reduces primitive count while maintaining high rendering fidelity, providing an efficient and practical solution for Gaussian Splat simplification. Our project website is available at https://saliteta.github.io/NanoGS/.
Authors:Jonas Herzog, Yue Wang
Abstract:
Recent research suggested that the embeddings produced by CLIP‑like contrastive language‑image training are suboptimal for image‑only tasks. The main theory is that the inter‑modal (language‑image) alignment loss ignores intra‑modal (image‑image) alignment, leading to poorly calibrated distances between images. In this study, we question this intra‑modal misalignment hypothesis. We reexamine its foundational theoretical argument, the indicators used to support it, and the performance metrics affected. For the theoretical argument, we demonstrate that there are no such supposed degrees of freedom for image embedding distances. For the empirical measures, our findings reveal they yield similar results for language‑image trained models (CLIP, SigLIP) and image‑image trained models (DINO, SigLIP2). This indicates the observed phenomena do not stem from a misalignment specific to the former. Experiments on the commonly studied intra‑modal tasks retrieval and few‑shot classification confirm that addressing task ambiguity, not supposed misalignment, is key for best results.
Authors:Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Tongliang Liu, Mingming Gong, Jiawang Bian
Abstract:
Existing diffusion‑based 3D scene generation methods primarily operate in 2D image/video latent spaces, which makes maintaining cross‑view appearance and geometric consistency inherently challenging. To bridge this gap, we present OneWorld, a framework that performs diffusion directly within a coherent 3D representation space. Central to our approach is the 3D Unified Representation Autoencoder (3D‑URAE); it leverages pretrained 3D foundation models and augments their geometry‑centric nature by injecting appearance and distilling semantics into a unified 3D latent space. Furthermore, we introduce token‑level Cross‑View‑Correspondence (CVC) consistency loss to explicitly enforce structural alignment across views, and propose Manifold‑Drift Forcing (MDF) to mitigate train‑inference exposure bias and shape a robust 3D manifold by mixing drifted and original representations. Comprehensive experiments demonstrate that OneWorld generates high‑quality 3D scenes with superior cross‑view consistency compared to state‑of‑the‑art 2D‑based methods. Our code will be available at https://github.com/SensenGao/OneWorld.
Authors:Shin'ya Yamaguchi, Daiki Chijiwa, Tamao Sakao, Taku Hasegawa
Abstract:
Large vision‑language models (LVLMs) employ multi‑modal in‑context learning (MM‑ICL) to adapt to new tasks by leveraging demonstration examples. While increasing the number of demonstrations boosts performance, they incur significant inference latency due to the quadratic computational cost of Transformer attention with respect to the context length. To address this trade‑off, we propose Parallel In‑Context Learning (Parallel‑ICL), a plug‑and‑play inference algorithm. Parallel‑ICL partitions the long demonstration context into multiple shorter, manageable chunks. It processes these chunks in parallel and integrates their predictions at the logit level, using a weighted Product‑of‑Experts (PoE) ensemble to approximate the full‑context output. Guided by ensemble learning theory, we introduce principled strategies for Parallel‑ICL: (i) clustering‑based context chunking to maximize inter‑chunk diversity and (ii) similarity‑based context compilation to weight predictions by query relevance. Extensive experiments on VQA, image captioning, and classification benchmarks demonstrate that Parallel‑ICL achieves performance comparable to full‑context MM‑ICL, while significantly improving inference speed. Our work offers an effective solution to the accuracy‑efficiency trade‑off in MM‑ICL, enabling dynamic task adaptation with substantially reduced inference overhead.
Authors:Sijie Li, Biao Qian, Jungong Han
Abstract:
Network pruning is an effective technique for enabling lightweight Large Vision‑Language Models (LVLMs), which primarily incorporates both weights and activations into the importance metric. However, existing efforts typically process calibration data from different modalities in a unified manner, overlooking modality‑specific behaviors. This raises a critical challenge: how to address the divergent behaviors of textual and visual tokens for accurate pruning of LVLMs. To this end, we systematically investigate the sensitivity of visual and textual tokens to the pruning operation by decoupling their corresponding weights, revealing that: (i) the textual pathway should be calibrated via text tokens, since it exhibits higher sensitivity than the visual pathway; (ii) the visual pathway exhibits high redundancy, permitting even 50% sparsity. Motivated by these insights, we propose a simple yet effective Asymmetric Text‑Visual Weight Pruning method for LVLMs, dubbed ATV‑Pruning, which establishes the importance metric for accurate weight pruning by selecting the informative tokens from both textual and visual pathways. Specifically, ATV‑Pruning integrates two primary innovations: first, a calibration pool is adaptively constructed by drawing on all textual tokens and a subset of visual tokens; second, we devise a layer‑adaptive selection strategy to yield important visual tokens. Finally, extensive experiments across standard multimodal benchmarks verify the superiority of our ATV‑Pruning over state‑of‑the‑art methods.
Authors:Xiaoyan Cong, Zekun Li, Zhiyang Dou, Hongyu Li, Omid Taheri, Chuan Guo, Abhay Mittal, Sizhe An, Taku Komura, Wojciech Matusik, Michael J. Black, Srinath Sridhar
Abstract:
Large‑scale foundation models (LFMs) have recently made impressive progress in text‑to‑motion generation by learning strong generative priors from massive 3D human motion datasets and paired text descriptions. However, how to effectively and efficiently leverage such single‑purpose motion LFMs, i.e., text‑to‑motion synthesis, in more diverse cross‑modal and in‑context motion generation downstream tasks remains largely unclear. Prior work typically adapts pretrained generative priors to individual downstream tasks in a task‑specific manner. In contrast, our goal is to unlock such priors to support a broad spectrum of downstream motion generation tasks within a single unified framework. To bridge this gap, we present UMO, a simple yet general unified formulation that casts diverse downstream tasks into compositions of atomic per‑frame operations, enabling in‑context adaptation to unlock the generative priors of pretrained DiT‑based motion LFMs. Specifically, UMO introduces three learnable frame‑level meta‑operation embeddings to specify per‑frame intent and employs lightweight temporal fusion to inject in‑context cues into the pretrained backbone, with negligible runtime overhead compared to the base model. With this design, UMO finetunes the pretrained model, originally limited to text‑to‑motion generation, to support diverse previously unsupported tasks, including temporal inpainting, text‑guided motion editing, text‑serialized geometric constraints, and multi‑identity reaction generation. Experiments demonstrate that UMO consistently outperforms task‑specific and training‑free baselines across a wide range of benchmarks, despite using a single unified model. Code and model will be publicly available. Project Page: https://oliver‑cong02.github.io/UMO.github.io/
Authors:Jakaria Rabbi, Nilanjan Ray, Dana Cobzas
Abstract:
Disentangling pathological changes from physiological aging in 3D medical shapes is crucial for developing interpretable biomarkers and patient stratification. However, this separation is challenging when diagnosis labels are limited or unavailable, since disease and aging often produce overlapping effects on shape changes, obscuring clinically relevant shape patterns. To address this challenge, we propose a two‑stage framework combining unsupervised disease discovery with self‑supervised disentanglement of implicit shape representations. In the first stage, we train an implicit neural model with signed distance functions to learn stable shape embeddings. We then apply clustering on the shape latent space, which yields pseudo disease labels without using ground‑truth diagnosis during discovery. In the second stage, we disentangle factors in a compact variational space using pseudo disease labels discovered in the first stage and the ground truth age labels available for all subjects. We enforce separation and controllability with a multi‑objective disentanglement loss combining covariance and a supervised contrastive loss. On ADNI hippocampus and OAI distal femur shapes, we achieve near‑supervised performance, improving disentanglement and reconstruction over state‑of‑the‑art unsupervised baselines, while enabling high‑fidelity reconstruction, controllable synthesis, and factor‑based explainability. Code and checkpoints are available at https://github.com/anonymous‑submission01/medical‑shape‑disentanglement
Authors:Renjie Liang, Yiling Ma, Yang Xing, Zhengkang Fan, Jinqian Pan, Chengkun Sun, Li Li, Kuang Gong, Jie Xu
Abstract:
Automated radiology report generation from 3D CT volumes often suffers from incomplete pathology coverage. We provide empirical evidence that this limitation stems from a representational bottleneck: contrastive 3D CT embeddings encode discriminative pathology signals, yet exhibit severe dimensional concentration, with as few as 2 effective dimensions out of 512. Corroborating this, scaling the language model yields no measurable improvement, suggesting that the bottleneck lies in the visual representation rather than the generator. This bottleneck limits both generation and retrieval; naive static retrieval fails to improve clinical efficacy and can even degrade performance. We propose AdaRAG‑CT, an adaptive augmentation framework that compensates for this visual bottleneck by introducing supplementary textual information through controlled retrieval and selectively integrating it during generation. On the CT‑RATE benchmark, AdaRAG‑CT achieves state‑of‑the‑art clinical efficacy, improving Clinical F1 from 0.420 (CT‑Agent) to 0.480 (+6 points); ablation studies confirm that both the retrieval and generation components contribute to the improvement. Code is available at https://github.com/renjie‑liang/Adaptive‑RAG‑for‑3DCT‑Report‑Generation.
Authors:Malte Prinzler, Paulo Gotardo, Siyu Tang, Timo Bolkart
Abstract:
We present MATCH (Multi‑view Avatars from Topologically Corresponding Heads), a multi‑view Gaussian registration method for high‑quality head avatar creation and editing. State‑of‑the‑art multi‑view head avatar methods require time‑consuming head tracking followed by expensive avatar optimization, often resulting in a total creation time of more than one day. MATCH, in contrast, directly predicts Gaussian splat textures in correspondence from calibrated multi‑view images in just 0.5 seconds per frame, without requiring data preprocessing. The learned intra‑subject correspondence across frames enables fast creation of personalized head avatars, while correspondence across subjects supports applications such as expression transfer, optimization‑free tracking, semantic editing, and identity interpolation. We establish these correspondences end‑to‑end using a transformer‑based model that predicts Gaussian splat textures in the fixed UV layout of a template mesh. To achieve this, we introduce a novel registration‑guided attention block, where each UV‑map token attends exclusively to image tokens depicting its corresponding mesh region. This design improves efficiency and performance compared to dense cross‑view attention. MATCH outperforms existing methods in novel‑view synthesis, geometry registration, and head avatar generation, while making avatar creation 10 times faster than the closest competing baseline. The code and model weights are available on the project website.
Authors:Marcell Kegl, Andras Palffy, Csaba Benedek, Dariu M. Gavrila
Abstract:
In this paper, we address extrinsic calibration for camera, lidar, and 4D radar sensors. Accurate extrinsic calibration of radar remains a challenge due to the sparsity of its data. We propose CLRNet, a novel, multi‑modal end‑to‑end deep learning (DL) calibration network capable of addressing joint camera‑lidar‑radar calibration, or pairwise calibration between any two of these sensors. We incorporate equirectangular projection, camera‑based depth image prediction, additional radar channels, and leverage lidar with a shared feature space and loop closure loss. In extensive experiments using the View‑of‑Delft and Dual‑Radar datasets, we demonstrate superior calibration accuracy compared to existing state‑of‑the‑art methods, reducing both median translational and rotational calibration errors by at least 50%. Finally, we examine the domain transfer capabilities of the proposed network and baselines, when evaluating across datasets. The code will be made publicly available upon acceptance at: https://github.com/tudelft‑iv.
Authors:Bingzhou Li, Tao Huang
Abstract:
Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Existing compression methods typically rely on fixed window partitioning and attention‑based pruning, which overlook the piecewise semantic structure of audio‑visual signals and become fragile under aggressive token reduction. We propose Dynamic Audio‑driven Semantic cHunking (DASH), a training‑free framework that aligns token compression with semantic structure. DASH treats audio embeddings as a semantic anchor and detects boundary candidates via cosine‑similarity discontinuities, inducing dynamic, variable‑length segments that approximate the underlying piecewise‑coherent organization of the sequence. These boundaries are projected onto video tokens to establish explicit cross‑modal segmentation. Within each segment, token retention is determined by a tri‑signal importance estimator that fuses structural boundary cues, representational distinctiveness, and attention‑based salience, mitigating the sparsity bias of attention‑only selection. This structure‑aware allocation preserves transition‑critical tokens while reducing redundant regions. Extensive experiments on AVUT, VideoMME, and WorldSense demonstrate that DASH maintains superior accuracy while achieving higher compression ratios compared to prior methods. Code is available at: https://github.com/laychou666/DASH.
Authors:Heng Fang, Shangru Li, Shuhan Wang, Xuanyang Xi, Dingkang Liang, Xiang Bai
Abstract:
Vision‑Language‑Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap primarily stems from a scarcity of dynamic manipulation datasets and the reliance of mainstream VLAs on single‑frame observations, restricting their spatiotemporal reasoning capabilities. To address this, we introduce DOMINO, a large‑scale dataset and benchmark for generalizable dynamic manipulation, featuring 35 tasks with hierarchical complexities, over 110K expert trajectories, and a multi‑dimensional evaluation suite. Through comprehensive experiments, we systematically evaluate existing VLAs on dynamic tasks, explore effective training strategies for dynamic awareness, and validate the generalizability of dynamic data. Furthermore, we propose PUMA, a dynamics‑aware VLA architecture. By integrating scene‑centric historical optical flow and specialized world queries to implicitly forecast object‑centric future states, PUMA couples history‑aware perception with short‑horizon prediction. Results demonstrate that PUMA achieves state‑of‑the‑art performance, yielding a 6.3% absolute improvement in success rate over baselines. Moreover, we show that training on dynamic data fosters robust spatiotemporal representations that transfer to static tasks. All code and data are available at https://github.com/H‑EmbodVis/DOMINO.
Authors:Zhenghong Zhou, Xiaohang Zhan, Zhiqin Chen, Soo Ye Kim, Nanxuan Zhao, Haitian Zheng, Qing Liu, He Zhang, Zhe Lin, Yuqian Zhou, Jiebo Luo
Abstract:
Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine‑grained control remains a key bottleneck that limits practical customizability for content creation. For AI video creators, three forms of control are crucial: (i) scene composition, (ii) multi‑view consistent subject customization, and (iii) camera‑pose or object‑motion adjustment. Existing methods typically handle these dimensions in isolation, with limited support for multi‑view subject synthesis and identity preservation under arbitrary pose changes. This lack of a unified architecture makes it difficult to support versatile, jointly controllable video. We introduce Tri‑Prompting, a unified framework and two‑stage training paradigm that integrates scene composition, multi‑view subject consistency, and motion control. Our approach leverages a dual‑condition motion module driven by 3D tracking points for background scenes and downsampled RGB cues for foreground subjects. To ensure a balance between controllability and visual realism, we further propose an inference ControlNet scale schedule. Tri‑Prompting supports novel workflows, including 3D‑aware subject insertion into any scenes and manipulation of existing subjects in an image. Experimental results demonstrate that Tri‑Prompting significantly outperforms specialized baselines such as Phantom and DaS in multi‑view subject identity, 3D consistency, and motion accuracy.
Authors:Yukang Cao, Haozhe Xie, Fangzhou Hong, Long Zhuo, Zhaoxi Chen, Liang Pan, Ziwei Liu
Abstract:
We present HSImul3R, a unified framework for simulation‑ready 3D reconstruction of human‑scene interactions (HSI) from casual captures, including sparse‑view images and monocular videos. Existing methods suffer from a perception‑simulation gap: visually plausible reconstructions often violate physical constraints, leading to instability in physics engines and failure in embodied AI applications. To bridge this gap, we introduce a physically‑grounded bi‑directional optimization pipeline that treats the physics simulator as an active supervisor to jointly refine human dynamics and scene geometry. In the forward direction, we employ Scene‑targeted Reinforcement Learning to optimize human motion under dual supervision of motion fidelity and contact stability. In the reverse direction, we propose Direct Simulation Reward Optimization, which leverages simulation feedback on gravitational stability and interaction success to refine scene geometry. We further present HSIBench, a new benchmark with diverse objects and interaction scenarios. Extensive experiments demonstrate that HSImul3R produces the first stable, simulation‑ready HSI reconstructions and can be directly deployed to real‑world humanoid robots.
Authors:Pengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen, Ser-Nam Lim, Harry Yang
Abstract:
Existing video‑to‑audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating acoustically distinct sounds under coarse labels, and textual ambiguity in describing micro‑acoustic features. These bottlenecks make it difficult to perform fine‑grained sound synthesis using text‑controlled modes. To address these limitations, we propose AC‑Foley, an audio‑conditioned V2A model that directly leverages reference audio to achieve precise and fine‑grained control over generated sounds. This approach enables fine‑grained sound synthesis, timbre transfer, zero‑shot sound generation, and improved audio quality. By directly conditioning on audio signals, our approach bypasses the semantic ambiguities of text descriptions while enabling precise manipulation of acoustic attributes. Empirically, AC‑Foley achieves state‑of‑the‑art performance for Foley generation when conditioned on reference audio, while remaining competitive with state‑of‑the‑art video‑to‑audio methods even without audio conditioning. Code and demo are available at: https://ff2416.github.io/AC‑Foley‑Page
Authors:Zeyu Ding, Yong Zhou, Jiaqi Zhao, Wen-Liang Du, Xixi Li, Rui Yao, Abdulmotaleb El Saddik
Abstract:
Recent real‑time detection transformers have gained popularity due to their simplicity and efficiency. However, these detectors do not explicitly model object rotation, especially in remote sensing imagery where objects appear at arbitrary angles, leading to challenges in angle representation, matching cost, and training stability. In this paper, we propose a real‑time oriented object detection transformer, the first real‑time end‑to‑end oriented object detector to the best of our knowledge, that addresses the above issues. Specifically, angle distribution refinement is proposed to reformulate angle regression as an iterative refinement of probability distributions, thereby capturing the uncertainty of object rotation and providing a more fine‑grained angle representation. Then, we incorporate a Chamfer distance cost into bipartite matching, measuring box distance via vertex sets, enabling more accurate geometric alignment and eliminating ambiguous matches. Moreover, we propose oriented contrastive denoising to stabilize training and analyze four noise modes. We observe that a ground truth can be assigned to different index queries across different decoder layers, and analyze this issue using the proposed instability metric. We design a series of model variants and experiments to validate the proposed method. Notably, our O2‑DFINE‑L, O2‑RTDETR‑R50 and O2‑DEIM‑R50 achieve 77.73%/78.45%/80.15% AP50 on DOTA1.0 and 132/119/119 FPS on the 2080ti GPU. Code is available at https://github.com/wokaikaixinxin/ai4rs.
Authors:Xianbao Hou, Yonghao He, Zeyd Boukhers, John See, Hu Su, Wei Sui, Cong Yang
Abstract:
Diffusion models have significantly mitigated the impact of annotated data scarcity in remote sensing (RS). Although recent approaches have successfully harnessed these models to enable diverse and controllable Layout‑to‑Image (L2I) synthesis, they still suffer from limited fine‑grained control and fail to strictly adhere to bounding box constraints. To address these limitations, we propose RSGen, a plug‑and‑play framework that leverages diverse edge guidance to enhance layout‑driven RS image generation. Specifically, RSGen employs a progressive enhancement strategy: 1) it first enriches the diversity of edge maps composited from retrieved training instances via Image‑to‑Image generation; and 2) subsequently utilizes these diverse edge maps as conditioning for existing L2I models to enforce pixel‑level control within bounding boxes, ensuring the generated instances strictly adhere to the layout. Extensive experiments across three baseline models demonstrate that RSGen significantly boosts the capabilities of existing L2I models. For instance, with CC‑Diff on the DOTA dataset for oriented object detection, we achieve remarkable gains of +9.8/+12.0 in YOLOScore mAP50/mAP50‑95 and +1.6 in mAP on the downstream detection task. Our code will be publicly available: https://github.com/D‑Robotics‑AI‑Lab/RSGen
Authors:Ruonan Yu, Zhenxiong Tan, Zigeng Chen, Songhua Liu, Xinchao Wang
Abstract:
Diffusion Transformers (DiTs) have demonstrated remarkable scalability and quality in image and video generation, prompting growing interest in extending them to controllable generation and editing tasks. However, compared to the image counterparts, progress in video control and editing remains limited, mainly due to the scarcity of paired video data and the high computational cost of training video diffusion models. To address this issue, in this paper, we propose a video‑free tuning framework termed ViFeEdit for video diffusion transformers. Without requiring any forms of video training data, ViFeEdit achieves versatile video generation and editing, adapted solely with 2D images. At the core of our approach is an architectural reparameterization that decouples spatial independence from the full 3D attention in modern video diffusion transformers, which enables visually faithful editing while maintaining temporal consistency with only minimal additional parameters. Moreover, this design operates in a dual‑path pipeline with separate timestep embeddings for noise scheduling, exhibiting strong adaptability to diverse conditioning signals. Extensive experiments demonstrate that our method delivers promising results of controllable video generation and editing with only minimal training on 2D image data. Codes are available https://github.com/Lexie‑YU/ViFeEdit.
Authors:Yuanfan Zheng, Kunyu Peng, Xu Zheng, Kailun Yang
Abstract:
Cross‑domain panoramic semantic segmentation has attracted growing interest as it enables comprehensive 360° scene understanding for real‑world applications. However, it remains particularly challenging due to severe geometric Field of View (FoV) distortions and inconsistent open‑set semantics across domains. In this work, we formulate an open‑set domain adaptation setting, and propose Extrapolative Domain Adaptive Panoramic Segmentation (EDA‑PSeg) framework that trains on local perspective views and tests on full 360° panoramic images, explicitly tackling both geometric FoV shifts across domains and semantic uncertainty arising from previously unseen classes. To this end, we propose the Euler‑Margin Attention (EMA), which introduces an angular margin to enhance viewpoint‑invariant semantic representation, while performing amplitude and phase modulation to improve generalization toward unseen classes. Additionally, we design the Graph Matching Adapter (GMA), which builds high‑order graph relations to align shared semantics across FoV shifts while effectively separating novel categories through structural adaptation. Extensive experiments on four benchmark datasets under camera‑shift, weather‑condition, and open‑set scenarios demonstrate that EDA‑PSeg achieves state‑of‑the‑art performance, robust generalization to diverse viewing geometries, and resilience under varying environmental conditions. The code is available at https://github.com/zyfone/EDA‑PSeg.
Authors:Corentin Dumery, Noa Etté, Aoxiang Fan, Ren Li, Jingyi Xu, Hieu Le, Pascal Fua
Abstract:
Visual object counting is a fundamental computer vision task in industrial inspection, where accurate, high‑throughput inventory tracking and quality assurance are critical. Moreover, manufactured parts are often too light to reliably deduce their count from their weight, or too heavy to move the stack on a scale safely and practically, making automated visual counting the more robust solution in many scenarios. However, existing methods struggle with stacked 3D items in containers, pallets, or bins, where most objects are heavily occluded and only a few are directly visible. To address this important yet underexplored challenge, we propose a novel 3D counting approach that decomposes the task into two complementary subproblems: estimating the 3D geometry of the stack and its occupancy ratio from multi‑view images. By combining geometric reconstruction with deep learning‑based depth analysis, our method can accurately count identical manufactured parts inside containers, even when they are irregularly stacked and partially hidden. We validate our 3D counting pipeline on large‑scale synthetic and diverse real‑world data with manually verified total counts, demonstrating robust performance under realistic inspection conditions.
Authors:Grzegorz Wilczyński, Mikołaj Zieliński, Krzysztof Byrski, Joanna Waczyńska, Dominik Belter, Przemysław Spurek
Abstract:
Neural Radiance Fields achieve high‑fidelity scene representation but suffer from costly training and rendering, while 3D Gaussian splatting offers real‑time performance with strong empirical results. Recently, solutions that harness the best of both worlds by using Gaussians as proxies to guide neural field evaluations, still suffer from significant computational inefficiencies. They typically rely on stochastic volumetric sampling to aggregate features, which severely limits rendering performance. To address this issue, a novel framework named IRIS (Intersection‑aware Ray‑based Implicit Editable Scenes) is introduced as a method designed for efficient and interactive scene editing. To overcome the limitations of standard ray marching, an analytical sampling strategy is employed that precisely identifies interaction points between rays and scene primitives, effectively eliminating empty space processing. Furthermore, to address the computational bottleneck of spatial neighbor lookups, a continuous feature aggregation mechanism is introduced that operates directly along the ray. By interpolating latent attributes from sorted intersections, costly 3D searches are bypassed, ensuring geometric consistency, enabling high‑fidelity, real‑time rendering, and flexible shape editing. Code can be found at https://github.com/gwilczynski95/iris.
Authors:Jiacheng Dong, Huan Li, Sicheng Zhou, Wenhao Hu, Weili Xu, Yan Wang
Abstract:
Reconstruction is a fundamental task in 3D vision and a fundamental capability for spatial intelligence. Particularly, streaming 3D reconstruction is central to real‑time spatial perception, yet existing recurrent online models often suffer from progressive degradation on long sequences due to state drift and forgetting, motivating inference‑time remedies. We present MeMix, a training‑free, plug‑and‑play module that improves streaming reconstruction by recasting the recurrent state into a Memory Mixture. MeMix partitions the state into multiple independent memory patches and updates only the least‑aligned memory patches while exactly preserving others. This selective update mitigates catastrophic forgetting while retaining O(1) inference memory, and requires no fine‑tuning or additional learnable parameters, making it directly applicable to existing recurrent reconstruction models. Across standard benchmarks (ScanNet, 7‑Scenes, KITTI, etc.), under identical backbones and inference settings, MeMix reduces reconstruction completeness error by 15.3% on average (up to 40.0%) across 300‑‑500 frame streams on 7‑Scenes. The code is available at https://dongjiacheng06.github.io/MeMix/
Authors:Aggelos Psiris, Yannis Panagakis, Maria Vakalopoulou, Georgios Th. Papadopoulos
Abstract:
Few‑Shot Industrial Visual Anomaly Detection (FS‑IVAD) comprises a critical task in modern manufacturing settings, where automated product inspection systems need to identify rare defects using only a handful of normal/defect‑free training samples. In this context, the current study introduces a novel reconstruction‑based approach termed GATE‑AD. In particular, the proposed framework relies on the employment of a masked, representation‑aligned Graph Attention Network (GAT) encoding scheme to learn robust appearance patterns of normal samples. By leveraging dense, patch‑level, visual feature tokens as graph nodes, the model employs stacked self‑attentional layers to adaptively encode complex, irregular, non‑Euclidean, local relations. The graph is enhanced with a representation alignment component grounded on a learnable, latent space, where high reconstruction residual areas (i.e., defects) are assessed using a Scaled Cosine Error (SCE) objective function. Extensive comparative evaluation on the MVTec AD, VisA, and MPDD industrial defect detection benchmarks demonstrates that GATE‑AD achieves state‑of‑the‑art performance across the 1‑ to 8‑shot settings, combining the highest detection accuracy (increase up to 1.8% in image AUROC in the 8‑shot case in MPDD) with the lowest per‑image inference latency (at least 25.05% faster), compared to the best‑performing literature methods. In order to facilitate reproducibility and further research, the source code of GATE‑AD is available at https://github.com/gthpapadopoulos/GATE‑AD.
Authors:Aram Davtyan, Leello Tadesse Dadi, Volkan Cevher, Paolo Favaro
Abstract:
Conditional Flow Matching (CFM), a simulation‑free method for training continuous normalizing flows, provides an efficient alternative to diffusion models for key tasks like image and video generation. The performance of CFM in solving these tasks depends on the way data is coupled with noise. A recent approach uses minibatch optimal transport (OT) to reassign noise‑data pairs in each training step to streamline sampling trajectories and thus accelerate inference. However, its optimization is restricted to individual minibatches, limiting its effectiveness on large datasets. To address this shortcoming, we introduce LOOM‑CFM (Looking Out Of Minibatch‑CFM), a novel method to extend the scope of minibatch OT by preserving and optimizing these assignments across minibatches over training time. Our approach demonstrates consistent improvements in the sampling speed‑quality trade‑off across multiple datasets. LOOM‑CFM also enhances distillation initialization and supports high‑resolution synthesis in latent space training.
Authors:Théo Sourget, Niclas Claßen, Jack Junchi Xu, Rob van der Goot, Veronika Cheplygina
Abstract:
The diversity of training datasets is usually perceived as an important aspect to obtain a robust model. However, the definition of diversity is often not defined or differs across papers, and while some metrics exist, the quantification of this diversity is often overlooked when developing new algorithms. In this work, we study the behaviour of multiple dataset diversity metrics for image, text and metadata using MorphoMNIST, a toy dataset with controlled perturbations, and PadChest, a publicly available chest X‑ray dataset. We evaluate whether these metrics correlate with each other but also with the intuition of a clinical expert. We also assess whether they correlate with downstream‑task performance and how they impact the training dynamic of the models. We find limited correlations between the AUC and image or metadata reference‑free diversity metrics, but higher correlations with the FID and the semantic diversity metrics. Finally, the clinical expert indicates that scanners are the main source of diversity in practice. However, we find that the addition of another scanner to the training set leads to shortcut learning. The code used in this study is available at https://github.com/TheoSourget/dataset_diversity_evaluation
Authors:Junlong Ke, Zichen Wen, Boxue Yang, Yantai Yang, Xuyang Liu, Chenfei Liao, Zhaorun Chen, Shaobo Wang, Linfeng Zhang
Abstract:
Native unified multimodal models, which integrate both generative and understanding capabilities, face substantial computational overhead that hinders their real‑world deployment. Existing acceleration techniques typically employ a static, monolithic strategy, ignoring the fundamental divergence in computational profiles between iterative generation tasks (e.g., image generation) and single‑pass understanding tasks (e.g., VQA). In this work, we present the first systematic analysis of unified models, revealing pronounced parameter specialization, where distinct neuron sets are critical for each task. This implies that, at the parameter level, unified models have implicitly internalized separate inference pathways for generation and understanding within a single architecture. Based on these insights, we introduce a training‑free and task‑aware acceleration framework, FlashU, that tailors optimization to each task's demands. Across both tasks, we introduce Task‑Specific Network Pruning and Dynamic Layer Skipping, aiming to eliminate inter‑layer and task‑specific redundancy. For visual generation, we implement a time‑varying control signal for the guidance scale and a temporal approximation for the diffusion head via Diffusion Head Cache. For multimodal understanding, building upon the pruned model, we introduce Dynamic Token Pruning via a V‑Norm Proxy to exploit the spatial redundancy of visual inputs. Extensive experiments on Show‑o2 demonstrate that FlashU achieves 1.78× to 2.01× inference acceleration across both understanding and generation tasks while maintaining SOTA performance, outperforming competing unified models and validating our task‑aware acceleration paradigm. Our code is publicly available at https://github.com/Rirayh/FlashU.
Authors:Victor Wåhlstrand, Jennifer Alvén, Ida Häggström
Abstract:
We present a framework to take advantage of existing labels at inference, called exemplars, in order to improve the performance of object detection in medical images. The method, exemplar diffusion, leverages existing diffusion methods for object detection to enable a training‑free approach to adding information of known bounding boxes at test time. We demonstrate that for medical image datasets with clear spatial structure, the method yields an across‑the‑board increase in average precision and recall, and a robustness to exemplar quality, enabling non‑expert annotation. Moreover, we demonstrate how our method may also be used to quantify predictive uncertainty in diffusion detection methods. Source code and data splits openly available online: https://github.com/waahlstrand/ExemplarDiffusion
Authors:Kuniaki Saito, Risa Shinoda, Shohei Tanaka, Tosho Hirasawa, Fumio Okura, Yoshitaka Ushiku
Abstract:
Hallucination detection in captions (HalDec) assesses a vision‑language model's ability to correctly align image content with text by identifying errors in captions that misrepresent the image. Beyond evaluation, effective hallucination detection is also essential for curating high‑quality image‑caption pairs used to train VLMs. However, the generalizability of VLMs as hallucination detectors across different captioning models and hallucination types remains unclear due to the lack of a comprehensive benchmark. In this work, we introduce HalDec‑Bench, a benchmark designed to evaluate hallucination detectors in a principled and interpretable manner. HalDec‑Bench contains captions generated by diverse VLMs together with human annotations indicating the presence of hallucinations, detailed hallucination‑type categories, and segment‑level labels. The benchmark provides tasks with a wide range of difficulty levels and reveals performance differences across models that are not visible in existing multimodal reasoning or alignment benchmarks. Our analysis further uncovers two key findings. First, detectors tend to recognize sentences appearing at the beginning of a response as correct, regardless of their actual correctness. Second, our experiments suggest that dataset noise can be substantially reduced by using strong VLMs as filters while employing recent VLMs as caption generators. Our project page is available at https://dahlian00.github.io/HalDec‑Bench‑Page/.
Authors:David Holtz, Niklas Hanselmann, Simon Doll, Marius Cordts, Bernt Schiele
Abstract:
End‑to‑end autonomous driving has gained significant attention for its potential to learn robust behavior in interactive scenarios and scale with data. Popular architectures often build on separate modules for perception and planning connected through latent representations, such as bird's eye view feature grids, to maintain end‑to‑end differentiability. This paradigm emerged mostly on open‑loop datasets, with evaluation focusing not only on driving performance, but also intermediate perception tasks. Unfortunately, architectural advances that excel in open‑loop often fail to translate to scalable learning of robust closed‑loop driving. In this paper, we systematically re‑examine the impact of common architectural patterns on closed‑loop performance: (1) high‑resolution perceptual representations, (2) disentangled trajectory representations, and (3) generative planning. Crucially, our analysis evaluates the combined impact of these patterns, revealing both unexpected limitations as well as underexplored synergies. Building on these insights, we introduce BevAD, a novel lightweight and highly scalable end‑to‑end driving architecture. BevAD achieves 72.7% success rate on the Bench2Drive benchmark and demonstrates strong data‑scaling behavior using pure imitation learning. Our code and models are publicly available here: https://dmholtz.github.io/bevad/
Authors:Hua Chang, Xin Xu, Wei Liu, Jiayi Wu, Kui Jiang, Fei Ma, Qi Tian
Abstract:
Many classic opera videos exhibit poor visual quality due to the limitations of early filming equipment and long‑term degradation during storage. Although real‑world video super‑resolution (RWVSR) has achieved significant advances in recent years, directly applying existing methods to degraded opera videos remains challenging. The difficulties are twofold. First, accurately modeling real‑world degradations is complex: simplistic combinations of classical degradation kernels fail to capture the authentic noise distribution, while methods that extract real noise patches from external datasets are prone to style mismatches that introduce visual artifacts. Second, current RWVSR methods, which rely solely on degraded image features, struggle to reconstruct realistic and detailed textures due to a lack of high‑level semantic guidance. To address these issues, we propose a Text‑guided Dual‑Branch Opera Video Super‑Resolution (TextOVSR) network, which introduces two types of textual prompts to guide the super‑resolution process. Specifically, degradation‑descriptive text, derived from the degradation process, is incorporated into the negative branch to constrain the solution space. Simultaneously, content‑descriptive text is incorporated into a positive branch and our proposed Text‑Enhanced Discriminator (TED) to provide semantic guidance for enhanced texture reconstruction. Furthermore, we design a Degradation‑Robust Feature Fusion (DRF) module to facilitate cross‑modal feature fusion while suppressing degradation interference. Experiments on our OperaLQ benchmark show that TextOVSR outperforms state‑of‑the‑art methods both qualitatively and quantitatively. The code is available at https://github.com/ChangHua0/TextOVSR.
Authors:Hainuo Wang, Mingjia Li, Xiaojie Guo
Abstract:
While recent Flow Matching models avoid the reconstruction bottlenecks of latent autoencoders by operating directly in pixel space, the lack of semantic continuity in the pixel manifold severely intertwines optimal transport paths. This induces severe trajectory conflicts near intersections, yielding sub‑optimal solutions. Rather than bypassing this issue via information‑lossy latent representations, we directly untangle the pixel‑space trajectories by proposing Waypoint Diffusion Transformers (WiT). WiT factorizes the continuous vector field via intermediate semantic waypoints projected from pre‑trained vision models. It effectively disentangles the generation trajectories by breaking the optimal transport into prior‑to‑waypoint and waypoint‑to‑pixel segments. Specifically, during the iterative denoising process, a lightweight generator dynamically infers these intermediate waypoints from the current noisy state. They then continuously condition the primary diffusion transformer via the Just‑Pixel AdaLN mechanism, steering the evolution towards the next state, ultimately yielding the final RGB pixels. Evaluated on ImageNet 256x256, WiT beats strong pixel‑space baselines, accelerating JiT training convergence by 2.2x. Code will be publicly released at https://github.com/hainuo‑wang/WiT.git.
Authors:Udi Barzelay, Ophir Azulai, Inbar Shapira, Idan Friedman, Foad Abo Dahood, Madison Lee, Abraham Daniels
Abstract:
We introduce VAREX (VARied‑schema EXtraction), a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. VAREX employs a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic values, producing deterministic ground truth validated through three‑phase quality assurance. The benchmark comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities: plain text, layout‑preserving text (whitespace‑aligned to approximate column positions), document image, or both text and image combined. Unlike existing benchmarks that evaluate from a single input representation, VAREX provides four controlled modalities per document, enabling systematic ablation of how input format affects extraction accuracy ‑‑ a capability absent from prior benchmarks. We evaluate 20 models from frontier proprietary models to small open models, with particular attention to models <=4B parameters suitable for cost‑sensitive and latency‑constrained deployment. Results reveal that (1) below 4B parameters, structured output compliance ‑‑ not extraction capability ‑‑ is a dominant bottleneck; in particular, schema echo (models producing schema‑conforming structure instead of extracted values) depresses scores by 45‑65 pp (percentage points) in affected models; (2) extraction‑specific fine‑tuning at 2B yields +81 pp gains, demonstrating that the instruction‑following deficit is addressable without scale; (3) layout‑preserving text provides the largest accuracy gain (+3‑18 pp), exceeding pixel‑level visual cues; and (4) the benchmark most effectively discriminates models in the 60‑95% accuracy band. Dataset and evaluation code are publicly available.
Authors:Omer Ben Hayun, Roy Betser, Meir Yossef Levi, Levi Kassel, Guy Gilboa
Abstract:
Following major advances in text and image generation, the video domain has surged, producing highly realistic and controllable sequences. Along with this progress, these models also raise serious concerns about misinformation, making reliable detection of synthetic videos increasingly crucial. Image‑based detectors are fundamentally limited because they operate per frame and ignore temporal dynamics, while supervised video detectors generalize poorly to unseen generators, a critical drawback given the rapid emergence of new models. These challenges motivate zero‑shot approaches, which avoid synthetic data and instead score content against real‑data statistics, enabling training‑free, model‑agnostic detection. We introduce STALL, a simple, training‑free, theoretically justified detector that provides likelihood‑based scoring for videos, jointly modeling spatial and temporal evidence within a probabilistic framework. We evaluate STALL on two public benchmarks and introduce ComGenVid, a new benchmark with state‑of‑the‑art generative models. STALL consistently outperforms prior image‑ and video‑based baselines. Code and data are available at https://omerbenhayun.github.io/stall‑video.
Authors:Yiqi Nie, Fei Wang, Junjie Chen, Kun Li, Yudi Cai, Dan Guo, Chenglong Li, Meng Wang
Abstract:
Memes represent a tightly coupled, multimodal form of social expression, in which visual context and overlaid text jointly convey nuanced affect and commentary. Inspired by cognitive reappraisal in psychology, we introduce Meme Reappraisal, a novel multimodal generation task that aims to transform negatively framed memes into constructive ones while preserving their underlying scenario, entities, and structural layout. Unlike prior works on meme understanding or generation, Meme Reappraisal requires emotion‑controllable, structure‑preserving multimodal transformation under multiple semantic and stylistic constraints. To support this task, we construct MER‑Bench, a benchmark of real‑world memes with fine‑grained multimodal annotations, including source and target emotions, positively rewritten meme text, visual editing specifications, and taxonomy labels covering visual type, sentiment polarity, and layout structure. We further propose a structured evaluation framework based on a multimodal large language model (MLLM)‑as‑a‑Judge paradigm, decomposing performance into modality‑level generation quality, affect controllability, structural fidelity, and global affective alignment. Extensive experiments across representative image‑editing and multimodal‑generation systems reveal substantial gaps in satisfying the constraints of structural preservation, semantic consistency, and affective transformation. We believe MER‑Bench establishes a foundation for research on controllable meme editing and emotion‑aware multimodal generation. Our code is available at: https://github.com/one‑seven17/MER‑Bench.
Authors:Xingtai Gui, Meijie Zhang, Tianyi Yan, Wencheng Han, Jiahao Gong, Feiyang Tan, Cheng-zhong Xu, Jianbing Shen
Abstract:
End‑to‑end autonomous driving aims to generate safe and plausible planning policies from raw sensor input. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving scene. However, existing driving world models primarily focus on visual scene representation, and motion representation is not explicitly designed to be planner‑shared and inheritable, leaving a schism between the optimization of visual scene generation and the requirements of precise motion planning. We present WorldDrive, a holistic framework that couples scene generation and real‑time planning via unifying vision and motion representation. We first introduce a Trajectory‑aware Driving World Model, which conditions on a trajectory vocabulary to enforce consistency between visual dynamics and motion intentions, enabling the generation of diverse and plausible future scenes conditioned on a specific trajectory. We transfer the vision and motion encoders to a downstream Multi‑modal Planner, ensuring the driving policy operates on mature representations pre‑optimized by scene generation. A simple interaction between motion representation, visual representation, and ego status can generate high‑quality, multi‑modal trajectories. Furthermore, to exploit the world model's foresight, we propose a Future‑aware Rewarder, which distills future latent representation from the frozen world model to evaluate and select optimal trajectories in real‑time. Extensive experiments on the NAVSIM, NAVSIM‑v2, and nuScenes benchmarks demonstrate that WorldDrive achieves leading planning performance among vision‑only methods while maintaining high‑fidelity action‑controlled video generation capabilities, providing strong evidence for the effectiveness of unifying vision and motion representation for robust autonomous driving.
Authors:Tianyu Zhang, Dongchi Li, Keiichi Sawada, Haoran Xie
Abstract:
Recent generative image editing methods adopt layered representations to mitigate the entangled nature of raster images and improve controllability, typically relying on object‑based segmentation. However, such strategies may fail to capture the structural and stylized properties of human‑created images, such as anime illustrations. To solve this issue, we propose a workflow‑aware structured layer decomposition framework tailored to the illustration production of anime artwork. Inspired by the creation pipeline of anime production, our method decomposes the illustration into semantically meaningful production layers, including line art, flat color, shadow, and highlight. To decouple all these layers, we introduce lightweight layer semantic embeddings to provide specific task guidance for each layer. Furthermore, a set of layer‑wise losses is incorporated to supervise the training process of individual layers. To overcome the lack of ground‑truth layered data, we construct a high‑quality illustration dataset that simulated the standard anime production workflow. Experiments demonstrate that the accurate and visually coherent layer decompositions were achieved by using our method. We believe that the resulting layered representation further enables downstream tasks such as recoloring and embedding texture, supporting content creation, and illustration editing. Code is available at: https://github.com/zty0304/Anime‑layer‑decomposition
Authors:Zitong Xu, Huiyu Duan, Zhongpeng Ji, Xinyun Zhang, Yutao Liu, Xiongkuo Min, Ke Gu, Jian Zhang, Shusong Xu, Jinwei Chen, Bo Li, Guangtao Zhai
Abstract:
Recent text‑guided image editing (TIE) models have achieved remarkable progress, while many edited images still suffer from issues such as artifacts, unexpected editings, unaesthetic contents. Although some benchmarks and methods have been proposed for evaluating edited images, scalable evaluation models are still lacking, which limits the development of human feedback reward models for image editing. To address the challenges, we first introduce EditHF‑1M, a million‑scale image editing dataset with over 29M human preference pairs and 148K human mean opinion ratings, both evaluated from three dimensions, i.e., visual quality, instruction alignment, and attribute preservation. Based on EditHF‑1M, we propose EditHF, a multimodal large language model (MLLM) based evaluation model, to provide human‑aligned feedback from image editing. Finally, we introduce EditHF‑Reward, which utilizes EditHF as the reward signal to optimize the text‑guided image editing models through reinforcement learning. Extensive experiments show that EditHF achieves superior alignment with human preferences and demonstrates strong generalization on other datasets. Furthermore, we fine‑tune the Qwen‑Image‑Edit using EditHF‑Reward, achieving significant performance improvements, which demonstrates the ability of EditHF to serve as a reward model to scale‑up the image editing. Both the dataset and code will be released in our GitHub repository: https://github.com/IntMeGroup/EditHF.
Authors:Seungryong Lee, Woojeong Baek, Joosang Lee, Eunbyung Park
Abstract:
A long‑term goal in CT imaging is to achieve fast and accurate 3D reconstruction from sparse‑view projections, thereby reducing radiation exposure, lowering system cost, and enabling timely imaging in clinical workflows. Recent feed‑forward approaches have shown strong potential toward this overarching goal, yet their results still suffer from artifacts and loss of fine details. In this work, we introduce Iterative Latent Volumes (ILV), a feed‑forward framework that integrates data‑driven priors with classical iterative reconstruction principles to overcome key limitations of prior feed‑forward models in sparse‑view CBCT reconstruction. At its core, ILV constructs an explicit 3D latent volume that is repeatedly updated by conditioning on multi‑view X‑ray features and the learned anatomical prior, enabling the recovery of fine structural details beyond the reach of prior feed‑forward models. In addition, we develop and incorporate several key architectural components, including an X‑ray feature volume, group cross‑attention, efficient self‑attention, and view‑wise feature aggregation, that efficiently realize its core latent volume refinement concept. Extensive experiments on a large‑scale dataset of approximately 14,000 CT volumes demonstrate that ILV significantly outperforms existing feed‑forward and optimization‑based methods in both reconstruction quality and speed. These results show that ILV enables fast and accurate sparse‑view CBCT reconstruction suitable for clinical use. The project page is available at: https://sngryonglee.github.io/ILV/.
Authors:Yaoyu Liu, Minghui Zhang, Junjun He, Yun Gu
Abstract:
Automatic extraction of vessel skeletons is crucial for many clinical applications. However, achieving topologically faithful delineation of thin vessel skeletons remains highly challenging, primarily due to frequent discontinuities and the presence of spurious skeleton segments. To address these difficulties, we propose TopoVST, a topology‑fidelitious vessel skeleton tracker. TopoVST constructs multi‑scale sphere graphs to sample the input image and employs graph neural networks to jointly estimate tracking directions and vessel radii. The utilization of multi‑scale representations is enhanced through a gating‑based feature fusion mechanism, while the issue of class imbalance during training is mitigated by embedding a geometry‑aware weighting scheme into the directional loss. In addition, we design a wave‑propagation‑based skeleton tracking algorithm that explicitly mitigates the generation of spurious skeletons through space‑occupancy filtering. We evaluate TopoVST on two vessel datasets with different geometries. Extensive comparisons with state‑of‑the‑art baselines demonstrate that TopoVST achieves competitive performance in both overlapping and topological metrics. Our source code is available at: https://github.com/EndoluminalSurgicalVision‑IMR/TopoVST.
Authors:Huanjing Yue, Shangbin Xie, Cong Cao, Qian Wu, Lei Zhang, Lei Zhao, Jingyu Yang
Abstract:
RAW images preserve superior fidelity and rich scene information compared to RGB, making them essential for tasks in challenging imaging conditions. To alleviate the high cost of data collection, recent RGB‑to‑RAW conversion methods aim to synthesize RAW images from RGB. However, they overlook two key challenges: (i) the reconstruction difficulty varies with pixel intensity, and (ii) multi‑camera conversion requires camera‑specific adaptation. To address these issues, we propose SpiralDiff, a diffusion‑based framework tailored for RGB‑to‑RAW conversion with a signal‑dependent noise weighting strategy that adapts reconstruction fidelity across intensity levels. In addition, we introduce CamLoRA, a camera‑aware lightweight adaptation module that enables a unified model to adapt to different camera‑specific ISP characteristics. Extensive experiments on four benchmark datasets demonstrate the superiority of SpiralDiff in RGB‑to‑RAW conversion quality and its downstream benefits in RAW‑based object detection. Our code and model are available at https://github.com/Chuancy‑TJU/SpiralDiff.
Authors:Linfei Li, Lin Zhang, Ying Shen
Abstract:
Visual‑language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches focus on coarse‑grained, object‑level localization, while traditional robotic grasping methods rely predominantly on geometric cues and lack language guidance, which limits their applicability in language‑driven manipulation scenarios. To address these limitations, we propose the RealVLG framework, which integrates the RealVLG‑11B dataset and the RealVLG‑R1 model to unify real‑world visual‑language grounding and grasping tasks. RealVLG‑11B dataset provides multi‑granularity annotations including bounding boxes, segmentation masks, grasp poses, contact points, and human‑verified fine‑grained language descriptions, covering approximately 165,000 images, over 800 object instances, 1.3 million segmentation, detection, and language annotations, and roughly 11 billion grasping examples. Building on this dataset, RealVLG‑R1 employs Reinforcement Fine‑tuning on pretrained large‑scale vision‑language models to predict bounding boxes, segmentation masks, grasp poses, and contact points in a unified manner given natural language instructions. Experimental results demonstrate that RealVLG supports zero‑shot perception and manipulation in real‑world unseen environments, establishing a unified semantic‑visual multimodal benchmark that provides a comprehensive data and evaluation platform for language‑driven robotic perception and grasping policy learning. All data and code are publicly available at https://github.com/lif314/RealVLG‑R1.
Authors:Tuan-Anh Yang, Bao V. Q. Bui, Chanh-Quang Vo-Van, Truong-Son Hy
Abstract:
We propose a deep learning framework for COVID‑19 detection and disease classification from chest CT scans that integrates both 2.5D and 3D representations to capture complementary slice‑level and volumetric information. The 2.5D branch processes multi‑view CT slices (axial, coronal, sagittal) using a DINOv3 vision transformer to extract robust visual features, while the 3D branch employs a ResNet‑18 architecture to model volumetric context and is pretrained with Variance Risk Extrapolation (VREx) followed by supervised contrastive learning to improve cross‑source robustness. Predictions from both branches are combined through logit‑level ensemble inference. Experiments on the PHAROS‑AIF‑MIH benchmark demonstrate the effectiveness of the proposed approach: for binary COVID‑19 detection, the ensemble achieves 94.48% accuracy and a 0.9426 Macro F1‑score, outperforming both individual models, while for multi‑class disease classification the 2.5D DINOv3 model achieves the best performance with 79.35% accuracy and a 0.7497 Macro F1‑score. These results highlight the benefit of combining pretrained slice‑based representations with volumetric modeling for robust multi‑source medical imaging analysis. Code is available at https://github.com/HySonLab/PHAROS‑AIF‑MIH
Authors:Shiwei Wang, Yongzhen Wang, Bingwen Hu, Liyan Zhang, Xiao-Ping Zhang, Mingqiang Wei
Abstract:
While Transformer‑based architectures have dominated recent advances in all‑in‑one image restoration, they remain fundamentally reactive: propagating degradations rather than proactively suppressing them. In the absence of explicit suppression mechanisms, degraded signals interfere with feature learning, compelling the decoder to balance artifact removal and detail preservation, thereby increasing model complexity and limiting adaptability. To address these challenges, we propose M2IR, a novel restoration framework that proactively regulates degradation propagation during the encoding stage and efficiently eliminates residual degradations during decoding. Specifically, the Mamba‑Style Transformer (MST) block performs pixel‑wise selective state modulation to mitigate degradations while preserving structural integrity. In parallel, the Adaptive Degradation Expert Collaboration (ADEC) module utilizes degradation‑specific experts guided by a DA‑CLIP‑driven router and complemented by a shared expert to eliminate residual degradations through targeted and cooperative restoration. By integrating the MST block and ADEC module, M2IR transitions from passive reaction to active degradation control, effectively harnessing learned representations to achieve superior generalization, enhanced adaptability, and refined recovery of fine‑grained details across diverse all‑in‑one image restoration benchmarks. Our source codes are available at https://github.com/Im34v/M2IR.
Authors:Kailin Lyu, Kangyi Wu, Pengna Li, Xiuyu Hu, Qingyi Si, Cui Miao, Ning Yang, Zihang Wang, Long Xiao, Lianyu Hu, Jingyuan Sun, Ce Hao
Abstract:
LLM‑based agents have demonstrated impressive zero‑shot performance in vision‑language navigation (VLN) tasks. However, most zero‑shot methods primarily rely on closed‑source LLMs as navigators, which face challenges related to high token costs and potential data leakage risks. Recent efforts have attempted to address this by using open‑source LLMs combined with a spatiotemporal CoT framework, but they still fall far short compared to closed‑source models. In this work, we identify a critical issue, Navigation Amnesia, through a detailed analysis of the navigation process. This issue leads to navigation failures and amplifies the gap between open‑source and closed‑source methods. To address this, we propose HiMemVLN, which incorporates a Hierarchical Memory System into a multimodal large model to enhance visual perception recall and long‑term localization, mitigating the amnesia issue and improving the agent's navigation performance. Extensive experiments in both simulated and real‑world environments demonstrate that HiMemVLN achieves nearly twice the performance of the open‑source state‑of‑the‑art method. The code is available at https://github.com/lvkailin0118/HiMemVLN.
Authors:Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
Abstract:
Computer‑using agents (CUAs) act directly on graphical user interfaces, yet their perception of the screen is often unreliable. Existing work largely treats these failures as performance limitations, asking whether an action succeeds, rather than whether the agent is acting on the correct object at all. We argue that this is fundamentally a security problem. We formalize the visual confused deputy: a failure mode in which an agent authorizes an action based on a misperceived screen state, due to grounding errors, adversarial screenshot manipulation, or time‑of‑check‑to‑time‑of‑use (TOCTOU) races. This gap is practically exploitable: even simple screen‑level manipulations can redirect routine clicks into privileged actions while remaining indistinguishable from ordinary agent mistakes. To mitigate this threat, we propose the first guardrail that operates outside the agent's perceptual loop. Our method, dual‑channel contrastive classification, independently evaluates (1) the visual click target and (2) the agent's reasoning about the action against deployment‑specific knowledge bases, and blocks execution if either channel indicates risk. The key insight is that these two channels capture complementary failure modes: visual evidence detects target‑level mismatches, while textual reasoning reveals dangerous intent behind visually innocuous controls. Across controlled attacks, real GUI screenshots, and agent traces, the combined guardrail consistently outperforms either channel alone. Our results suggest that CUA safety requires not only better action generation, but independent verification of what the agent believes it is clicking and why. Materials are provided\footnoteModel, benchmark, and code: https://github.com/vllm‑project/semantic‑router.
Authors:Salim Khazem
Abstract:
Frozen‑backbone transfer with Vision Transformers faces two under‑addressed issues: optimization instability when adapters are naively inserted into a fixed feature extractor, and the absence of principled guidance for setting adapter capacity. We introduce AdapterTune, which augments each transformer block with a residual low‑rank bottleneck whose up‑projection is zero‑initialized, guaranteeing that the adapted network starts exactly at the pretrained function and eliminates early‑epoch representation drift. On the analytical side, we formalize adapter rank as a capacity budget for approximating downstream task shifts in feature space. The resulting excess‑risk decomposition predicts monotonic but diminishing accuracy gains with increasing rank, an ``elbow'' behavior we confirm through controlled sweeps. We evaluate on 9 datasets and 3 backbone scales with multi‑seed reporting throughout. On a core 5 dataset transfer suite, AdapterTune improves top‑1 accuracy over head‑only transfer by +14.9 points on average while training only 0.92 of the parameters required by full fine‑tuning, and outperforms full fine‑tuning on 10 of 15 dataset‑backbone pairs. Across the full benchmark, AdapterTune improves over head‑only transfer on every dataset‑backbone pair tested. Ablations on rank, placement, and initialization isolate each design choice. The code is available at: https://github.com/salimkhazem/adaptertune
Authors:Ping Chen, Xiang Liu, Xingpeng Zhang, Fei Shen, Xun Gong, Zhaoxiang Liu, Zezhou Chen, Huan Hu, Kai Wang, Shiguo Lian
Abstract:
Diffusion models operate in a reflexive System 1 mode, constrained by a fixed, content‑agnostic sampling schedule. This rigidity arises from the curse of state dimensionality, where the combinatorial explosion of possible states in the high‑dimensional noise manifold renders explicit trajectory planning intractable and leads to systematic computational misallocation. To address this, we introduce Chain‑of‑Trajectories (CoTj), a train‑free framework enabling System 2 deliberative planning. Central to CoTj is Diffusion DNA, a low‑dimensional signature that quantifies per‑stage denoising difficulty and serves as a proxy for the high‑dimensional state space, allowing us to reformulate sampling as graph planning on a directed acyclic graph. Through a Predict‑Plan‑Execute paradigm, CoTj dynamically allocates computational effort to the most challenging generative phases. Experiments across multiple generative models demonstrate that CoTj discovers context‑aware trajectories, improving output quality and stability while reducing redundant computation. This work establishes a new foundation for resource‑aware, planning‑based diffusion modeling. The code is available at https://github.com/UnicomAI/CoTj.
Authors:Mang Ning, Mingxiao Li, Le Zhang, Lanmiao Liu, Matthew B. Blaschko, Albert Ali Salah, Itir Onal Ertugrul
Abstract:
In this paper, we study the diffusability (learnability) of variational autoencoders (VAE) in latent diffusion. First, we show that pixel‑space diffusion trained with an MSE objective is inherently biased toward learning low and mid spatial frequencies, and that the power‑law power spectral density (PSD) of natural images makes this bias perceptually beneficial. Motivated by this result, we propose the \emphSpectrum Matching Hypothesis: latents with superior diffusability should (i) follow a flattened power‑law PSD (\emphEncoding Spectrum Matching, ESM) and (ii) preserve frequency‑to‑frequency semantic correspondence through the decoder (\emphDecoding Spectrum Matching, DSM). In practice, we apply ESM by matching the PSD between images and latents, and DSM via shared spectral masking with frequency‑aligned reconstruction. Importantly, Spectrum Matching provides a unified view that clarifies prior observations of over‑noisy or over‑smoothed latents, and interprets several recent methods as special cases (e.g., VA‑VAE, EQ‑VAE). Experiments suggest that Spectrum Matching yields superior diffusion generation on CelebA and ImageNet datasets, and outperforms prior approaches. Finally, we extend the spectral view to representation alignment (REPA): we show that the directional spectral energy of the target representation is crucial for REPA, and propose a DoG‑based method to further improve the performance of REPA. Our code is available https://github.com/forever208/SpectrumMatching.
Authors:Zengqun Zhao, Ziquan Liu, Yu Cao, Shaogang Gong, Zhensong Zhang, Jifei Song, Jiankang Deng, Ioannis Patras
Abstract:
The recent success of inference‑time scaling in large language models has inspired similar explorations in video diffusion. In particular, motivated by the existence of "golden noise" that enhances video quality, prior work has attempted to improve inference by optimising or searching for better initial noise. However, these approaches have notable limitations: they either rely on priors imposed at the beginning of noise sampling or on rewards evaluated only on the denoised and decoded videos. This leads to error accumulation, delayed and sparse reward signals, and prohibitive computational cost, which prevents the use of stronger search algorithms. Crucially, stronger search algorithms are precisely what could unlock substantial gains in controllability, sample efficiency and generation quality for video diffusion, provided their computational cost can be reduced. To fill in this gap, we enable efficient inference‑time scaling for video diffusion through latent reward guidance, which provides intermediate, informative and efficient feedback along the denoising trajectory. We introduce a latent reward model that scores partially denoised latents at arbitrary timesteps with respect to visual quality, motion quality, and text alignment. Building on this model, we propose LatSearch, a novel inference‑time search mechanism that performs Reward‑Guided Resampling and Pruning (RGRP). In the resampling stage, candidates are sampled according to reward‑normalised probabilities to reduce over‑reliance on the reward model. In the pruning stage, applied at the final scheduled step, only the candidate with the highest cumulative reward is retained, improving both quality and efficiency. We evaluate LatSearch on the VBench‑2.0 benchmark and demonstrate that it consistently improves video generation across multiple evaluation dimensions compared to the baseline Wan2.1 model.
Authors:Chaoyang Wang, Wenrui Bao, Sicheng Gao, Bingxin Xu, Yu Tian, Yogesh S. Rawat, Yunhao Ge, Yuzhang Shang
Abstract:
Vision‑Language‑Action (VLA) models have shown promising capabilities for embodied intelligence, but most existing approaches rely on text‑based chain‑of‑thought reasoning where visual inputs are treated as static context. This limits the ability of the model to actively revisit the environment and resolve ambiguities during long‑horizon tasks. We propose VLA‑Thinker, a thinking‑with‑image reasoning framework that models perception as a dynamically invocable reasoning action. To train such a system, we introduce a two‑stage training pipeline consisting of (1) an SFT cold‑start phase with curated visual Chain‑of‑Thought data to activate structured reasoning and tool‑use behaviors, and (2) GRPO‑based reinforcement learning to align complete reasoning‑action trajectories with task‑level success. Extensive experiments on LIBERO and RoboTwin 2.0 benchmarks demonstrate that VLA‑Thinker significantly improves manipulation performance, achieving 97.5% success rate on LIBERO and strong gains across long‑horizon robotic tasks. Project and Codes: https://cywang735.github.io/VLA‑Thinker/ .
Authors:Zhuoxuan Peng, Boan Zhu, Xingjian Zhang, Wenying Li, S. -H. Gary Chan
Abstract:
Current millimeter‑wave (mmWave) datasets for human pose estimation (HPE) are scarce and lack diversity in both point cloud (PC) attributes and human poses, hindering the generalization ability of their trained models. On the other hand, unlabeled mmWave HPE data and diverse LiDAR HPE datasets are readily available. We propose EMDUL, a novel approach to expand the volume and diversity of an existing mmWave dataset using unlabeled mmWave data and LiDAR datasets. EMDUL consists of two independent modules, namely a pseudo‑label estimator to annotate unlabeled mmWave data, and a closed‑form converter that translates an annotated LiDAR PC to its mmWave counterpart. Expanding the original dataset with both LiDAR‑converted and pseudo‑labeled mmWave PCs significantly boosts the performance and generalization ability of all the examined HPE models, reducing 15.1% and 18.9% error for in‑domain and out‑of‑domain settings, respectively. Code is available at https://github.com/Shimmer93/EMDUL.
Authors:Yuhao Zhang, Wanxi Dong, Yue Shi, Yi Liang, Jingnan Gao, Qiaochu Yang, Yaxing Lyu, Zhixuan Liang, Yibin Liu, Congsheng Xu, Xianda Guo, Wei Sui, Yaohui Jin, Xiaokang Yang, Yanyan Xu, Yao Mu
Abstract:
Embodied manipulation requires accurate 3D understanding of objects and their spatial relations to plan and execute contact‑rich actions. While large‑scale 3D vision models provide strong priors, their computational cost incurs prohibitive latency for real‑time control. We propose Real‑time 3D‑aware Policy (R3DP), which integrates powerful 3D priors into manipulation policies without sacrificing real‑time performance. A core innovation of R3DP is the asynchronous fast‑slow collaboration module, which seamlessly integrates large‑scale 3D priors into the policy without compromising real‑time performance. The system maintains real‑time efficiency by querying the pre‑trained slow system (VGGT) only on sparse key frames, while simultaneously employing a lightweight Temporal Feature Prediction Network (TFPNet) to predict features for all intermediate frames. By leveraging historical data to exploit temporal correlations, TFPNet explicitly improves task success rates through consistent feature estimation. Additionally, to enable more effective multi‑view fusion, we introduce a Multi‑View Feature Fuser (MVFF) that aggregates features across views by explicitly incorporating camera intrinsics and extrinsics. R3DP offers a plug‑and‑play solution for integrating large models into real‑time inference systems. We evaluate R3DP against multiple baselines across different visual configurations. R3DP effectively harnesses large‑scale 3D priors to achieve superior results, outperforming single‑view and multi‑view DP by 32.9% and 51.4% in average success rate, respectively. Furthermore, by decoupling heavy 3D reasoning from policy execution, R3DP achieves a 44.8% reduction in inference time compared to a naive DP+VGGT integration.
Authors:Haoyu Zhang, Wei Zhai, Yuhang Yang, Yang Cao, Zheng-Jun Zha
Abstract:
Monocular 4D human‑object interaction (HOI) reconstruction ‑ recovering a moving human and a manipulated object from a single RGB video ‑ remains challenging due to depth ambiguity and frequent occlusions. Existing methods often rely on multi‑stage pipelines or iterative optimization, leading to high inference latency, failing to meet real‑time requirements, and susceptibility to error accumulation. To address these limitations, we propose THO, an end‑to‑end Spatial‑Temporal Transformer that predicts human motion and coordinated object motion in a forward fashion from the given video and 3D template. THO achieves this by leveraging spatial‑temporal HOI tuple priors. Spatial priors exploit contact‑region proximity to infer occluded object features from human cues, while temporal priors capture cross‑frame kinematic correlations to refine object representations and enforce physical coherence. Extensive experiments demonstrate that THO operates at an inference speed of 31.5 FPS on a single RTX 4090 GPU, achieving a >600x speedup over prior optimization‑based methods while simultaneously improving reconstruction accuracy and temporal consistency. The project page is available at: https://nianheng.github.io/THO‑project/
Authors:Kuanning Wang, Ke Fan, Yuqian Fu, Siyu Lin, Hu Luo, Daniel Seita, Yanwei Fu, Yu-Gang Jiang, Xiangyang Xue
Abstract:
We present OCRA, an Object‑Centric framework for video‑based human‑to‑Robot Action transfer that learns directly from human demonstration videos to enable robust manipulation. Object‑centric learning emphasizes task‑relevant objects and their interactions while filtering out irrelevant background, providing a natural and scalable way to teach robots. OCRA leverages multi‑view RGB videos, the state‑of‑the‑art 3D foundation model VGGT, and advanced detection and segmentation models to reconstruct object‑centric 3D point clouds, capturing rich interactions between objects. To handle properties not easily perceived by vision alone, we incorporate tactile priors via a large‑scale dataset of over one million tactile images. These 3D and tactile priors are fused through a multimodal module (ResFiLM) and fed into a Diffusion Policy to generate robust manipulation actions. Extensive experiments on both vision‑only and visuo‑tactile tasks show that OCRA significantly outperforms existing baselines and ablations, demonstrating its effectiveness for learning from human demonstration videos.
Authors:Seokju Yun, Dongheon Lee, Noori Bae, Jaesung Jun, Chanseul Cho, Youngmin Ro
Abstract:
As AI systems are being integrated more rapidly into diverse and complex real‑world environments, the ability to perform holistic reasoning over an implicit query and an image to localize a target is becoming increasingly important. However, recent reasoning segmentation methods fail to sufficiently elicit the visual reasoning capabilities of the base mode. In this work, we present Segment Anything Reasoner (StAR), a comprehensive framework that refines the design space from multiple perspectives‑including parameter‑tuning scheme, reward functions, learning strategies and answer format‑and achieves substantial improvements over recent baselines. In addition, for the first time, we successfully introduce parallel test‑time scaling to the segmentation task, pushing the performance boundary even further. To extend the scope and depth of reasoning covered by existing benchmark, we also construct the ReasonSeg‑X, which compactly defines reasoning types and includes samples that require deeper reasoning. Leveraging this dataset, we train StAR with a rollout‑expanded selective‑tuning approach to activate the base model's latent reasoning capabilities, and establish a rigorous benchmark for systematic, fine‑grained evaluation of advanced methods. With only 5k training samples, StAR achieves significant gains over its base counterparts across extensive benchmarks, demonstrating that our method effectively brings dormant reasoning competence to the surface.
Authors:Xiangbo Gao, Mingyang Wu, Siyuan Yang, Jiongze Yu, Pardis Taghavi, Fangzhou Lin, Zhengzhong Tu
Abstract:
While recent generative video models have achieved remarkable visual realism and are being explored as world models, true physical simulation requires mastering both space and time. Current models can produce visually smooth kinematics, yet they lack a reliable internal motion pulse to ground these motions in a consistent, real‑world time scale. This temporal ambiguity stems from the common practice of indiscriminately training on videos with vastly different real‑world speeds, forcing them into standardized frame rates. This leads to what we term chronometric hallucination: generated sequences exhibit ambiguous, unstable, and uncontrollable physical motion speeds. To address this, we propose Visual Chronometer, a predictor that recovers the Physical Frames Per Second (PhyFPS) directly from the visual dynamics of an input video. Trained via controlled temporal resampling, our method estimates the true temporal scale implied by the motion itself, bypassing unreliable metadata. To systematically quantify this issue, we establish two benchmarks, PhyFPS‑Bench‑Real and PhyFPS‑Bench‑Gen. Our evaluations reveal a harsh reality: state‑of‑the‑art video generators suffer from severe PhyFPS misalignment and temporal instability. Finally, we demonstrate that applying PhyFPS corrections significantly improves the human‑perceived naturalness of AI‑generated videos. Our project page is https://xiangbogaobarry.github.io/Visual_Chronometer/.
Authors:Xiaoya Lu, Yijin Zhou, Zeren Chen, Ruocheng Wang, Bingrui Sima, Enshen Zhou, Lu Sheng, Dongrui Liu, Jing Shao
Abstract:
Vision‑Language Models (VLMs) empower embodied agents to execute complex instructions, yet they remain vulnerable to contextual safety risks where benign commands become hazardous due to subtle environmental states. Existing safeguards often prove inadequate. Rule‑based methods lack scalability in object‑dense scenes, whereas model‑based approaches relying on prompt engineering suffer from unfocused perception, resulting in missed risks or hallucinations. To address this, we propose an architecture‑agnostic safeguard featuring Context‑Guided Chain‑of‑Thought (CG‑CoT). This mechanism decomposes risk assessment into active perception that sequentially anchors attention to interaction targets and relevant spatial neighborhoods, followed by semantic judgment based on this visual evidence. We support this approach with a curated grounding dataset and a two‑stage training strategy utilizing Reinforcement Fine‑Tuning (RFT) with process rewards to enforce precise intermediate grounding. Experiments demonstrate that our model HomeGuard significantly enhances safety, improving risk match rates by over 30% compared to base models while reducing oversafety. Beyond hazard detection, the generated visual anchors serve as actionable spatial constraints for downstream planners, facilitating explicit collision avoidance and safety trajectory generation. Code and data are released under https://github.com/AI45Lab/HomeGuard
Authors:Jaeyo Shin, Jiwook Kim, Hyunjung Shim
Abstract:
Representation Alignment (REPA) has emerged as a simple way to accelerate Diffusion Transformers training in latent space. At the same time, pixel‑space diffusion transformers such as Just image Transformers (JiT) have attracted growing attention because they remove a dependency on a pretrained tokenizer, and then avoid the reconstruction bottleneck of latent diffusion. This paper shows that the REPA can fail for JiT. REPA yields worse FID for JiT as training proceeds and collapses diversity on image subsets that are tightly clustered in the representation space of pretrained semantic encoder on ImageNet. We trace the failure to an information asymmetry: denoising occurs in the high dimensional image space, while the semantic target is strongly compressed, making direct regression a shortcut objective. We propose PixelREPA, which transforms the alignment target and constrains alignment with a Masked Transformer Adapter that combines a shallow transformer adapter with partial token masking. PixelREPA improves both training convergence and final quality. PixelREPA reduces FID from 3.66 to 3.17 for JiT‑B/16 and improves Inception Score (IS) from 275.1 to 284.6 on ImageNet 256 × 256, while achieving > 2× faster convergence. Finally, PixelREPA‑H/16 achieves FID=1.81 and IS=317.2. Our code is available at https://github.com/kaist‑cvml/PixelREPA.
Authors:Peng Xu, Zhengnan Deng, Jiayan Deng, Zonghua Gu, Shaohua Wan
Abstract:
Vision‑Language Navigation (VLN) for Unmanned Aerial Vehicles (UAVs) demands complex visual interpretation and continuous control in dynamic 3D environments. Existing hierarchical approaches rely on dense oracle guidance or auxiliary object detectors, creating semantic gaps and limiting genuine autonomy. We propose AerialVLA, a minimalist end‑to‑end Vision‑Language‑Action framework mapping raw visual observations and fuzzy linguistic instructions directly to continuous physical control signals. First, we introduce a streamlined dual‑view perception strategy that reduces visual redundancy while preserving essential cues for forward navigation and precise grounding, which additionally facilitates future simulation‑to‑reality transfer. To reclaim genuine autonomy, we deploy a fuzzy directional prompting mechanism derived solely from onboard sensors, completely eliminating the dependency on dense oracle guidance. Ultimately, we formulate a unified control space that integrates continuous 3‑Degree‑of‑Freedom (3‑DoF) kinematic commands with an intrinsic landing signal, freeing the agent from external object detectors for precision landing. Extensive experiments on the TravelUAV benchmark demonstrate that AerialVLA achieves state‑of‑the‑art performance in seen environments. Furthermore, it exhibits superior generalization in unseen scenarios by achieving nearly three times the success rate of leading baselines, validating that a minimalist, autonomy‑centric paradigm captures more robust visual‑motor representations than complex modular systems.
Authors:Liyuan Cui, Wentao Hu, Wenyuan Zhang, Zesong Yang, Fan Shi, Xiaoqiang Liu
Abstract:
Real‑time talking avatar generation requires low latency and minute‑level temporal stability. Autoregressive (AR) forcing enables streaming inference but suffers from exposure bias, which causes errors to accumulate and become irreversible over long rollouts. In contrast, full‑sequence diffusion transformers mitigate drift but remain computationally prohibitive for real‑time long‑form synthesis. We present AvatarForcing, a one‑step streaming diffusion framework that denoises a fixed local‑future window with heterogeneous noise levels and emits one clean block per step under constant per‑step cost. To stabilize unbounded streams, the method introduces dual‑anchor temporal forcing: a style anchor that re‑indexes RoPE to maintain a fixed relative position with respect to the active window and applies anchor‑audio zero‑padding, and a temporal anchor that reuses recently emitted clean blocks to ensure smooth transitions. Real‑time one‑step inference is enabled by two‑stage streaming distillation with offline ODE backfill and distribution matching. Experiments on standard benchmarks and a new 400‑video long‑form benchmark show strong visual quality and lip synchronization at 34 ms/frame using a 1.3B‑parameter student model for realtime streaming. Our page is available at: https://cuiliyuan121.github.io/AvatarForcing/
Authors:Tongshun Zhang, Pingping Liu, Yuqing Lei, Zixuan Zhong, Qiuzhan Zhou, Zhiyuan Zha
Abstract:
Limited illumination often causes severe physical noise and detail degradation in images. Existing Low‑Light Image Enhancement (LLIE) methods frequently treat the enhancement process as a blind black‑box mapping, overlooking the physical noise transformation during imaging, leading to suboptimal performance. To address this, we propose a novel LLIE approach, conceptually formulated as a physics‑based attack and display‑adaptive defense paradigm. Specifically, on the attack side, we establish a physics‑based Degradation Synthesis (PDS) pipeline. Unlike standard data augmentation, PDS explicitly models Image Signal Processor (ISP) inversion to the RAW domain, injects physically plausible photon and read noise, and re‑projects the data to the sRGB domain. This generates high‑fidelity training pairs with explicitly parameterized degradation vectors, effectively simulating realistic attacks on clean signals. On the defense side, we construct a dual‑layer fortified system. A noise predictor estimates degradation parameters from the input sRGB image. These estimates guide a degradation‑aware Mixture of Experts (DA‑MoE), which dynamically routes features to experts specialized in handling specific noise intensities. Furthermore, we introduce an Adaptive Metric Defense (AMD) mechanism, dynamically calibrating the feature embedding space based on noise severity, ensuring robust representation learning under severe degradation. Extensive experiments demonstrate that our approach offers significant plug‑and‑play performance enhancement for existing benchmark LLIE methods, effectively suppressing real‑world noise while preserving structural fidelity. The sourced code is available at https://github.com/bywlzts/Attack‑defense‑llie.
Authors:Yujia Wang, Yuyan Li, Jiuming Liu, Fang-Lue Zhang, Xinhu Zheng, Neil. A Dodgson
Abstract:
Blind 360°image quality assessment (IQA) aims to predict perceptual quality for panoramic images without a pristine reference. Unlike conventional planar images, 360°content in immersive environments restricts viewers to a limited viewport at any moment, making viewing behaviors critical to quality perception. Although existing scanpath‑based approaches have attempted to model viewing behaviors by approximating the human view‑then‑rate paradigm, they treat scanpath generation and quality assessment as separate steps, preventing end‑to‑end optimization and task‑aligned exploration. To address this limitation, we propose RL‑ScanIQA, a reinforcement‑learned framework for blind 360°IQA. RL‑ScanIQA optimize a PPO‑trained scanpath policy and a quality assessor, where the policy receives quality‑driven feedback to learn task‑relevant viewing strategies. To improve training stability and prevent mode collapse, we design multi‑level rewards, including scanpath diversity and equator‑biased priors. We further boost cross‑dataset robustness using distortion‑space augmentation together with rank‑consistent losses that preserve intra‑image and inter‑image quality orderings. Extensive experiments on three benchmarks show that RL‑ScanIQA achieves superior in‑dataset performance and cross‑dataset generalization. Codes are available at https://github.com/wangyuji1/RLScanIQA.git.
Authors:Xingyuan Li, Songcheng Du, Yang Zou, HaoYuan Xu, Zhiying Jiang, Jinyuan Liu
Abstract:
Image fusion aims to integrate complementary information from multiple source images to produce a more informative and visually consistent representation, benefiting both human perception and downstream vision tasks. Despite recent progress, most existing fusion methods are designed for specific tasks (i.e., multi‑modal, multi‑exposure, or multi‑focus fusion) and struggle to effectively preserve source information during the fusion process. This limitation primarily arises from task‑specific architectures and the degradation of source information caused by deep‑layer propagation. To overcome these issues, we propose UniFusion, a unified image fusion framework designed to achieve cross‑task generalization. First, leveraging DINOv3 for modality‑consistent feature extraction, UniFusion establishes a shared semantic space for diverse inputs. Second, to preserve the understanding of each source image, we introduce a reconstruction‑alignment loss to maintain consistency between fused outputs and inputs. Finally, we employ a bilevel optimization strategy to decouple and jointly optimize reconstruction and fusion objectives, effectively balancing their coupling relationship and ensuring smooth convergence. Extensive experiments across multiple fusion tasks demonstrate UniFusion's superior visual quality, generalization ability, and adaptability to real‑world scenarios. Code is available at https://github.com/dusongcheng/UniFusion.
Authors:Shishi Xiao, Tongyu Zhou, David Laidlaw, Gromit Yeuk-Yin Chan
Abstract:
A pictorial chart is an effective medium for visual storytelling, seamlessly integrating visual elements with data charts. However, creating such images is challenging because the flexibility of visual elements often conflicts with the rigidity of chart structures. This process thus requires a creative deformation that maintains both data faithfulness and visual aesthetics. Current methods that extract dense structural cues from natural images (e.g., edge or depth maps) are ill‑suited as conditioning signals for pictorial chart generation. We present ChArtist, a domain‑specific diffusion model for generating pictorial charts automatically, offering two distinct types of control: 1) spatial control that aligns well with the chart structure, and 2) subject‑driven control that respects the visual characteristics of a reference image. To achieve this, we introduce a skeleton‑based spatial control representation. This representation encodes only the data‑encoding information of the chart, allowing for the easy incorporation of reference visuals without a rigid outline constraint. We implement our method based on the Diffusion Transformer (DiT) and leverage an adaptive position encoding mechanism to manage these two controls. We further introduce Spatially Gated Attention to modulate the interaction between spatial control and subject control. To support the fine‑tuning of pre‑trained models for this task, we created a large‑scale dataset of 30,000 triplets (skeleton, reference image, pictorial chart). We also propose a unified data accuracy metric to evaluate the data faithfulness of the generated charts. We believe this work demonstrates that current generative models can achieve data‑driven visual storytelling by moving beyond general‑purpose conditions to task‑specific representations. Project page: https://chartist‑ai.github.io/.
Authors:Kai Peng, Yunzhe Shen, Miao Zhang, Leiye Liu, Yidong Han, Wei Ji, Jingjing Li, Yongri Piao, Huchuan Lu
Abstract:
The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio‑Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction between audio and visual modalities still requires further exploration. In this work, we aim to answer the following questions: How can a model effectively suppress audio noise while enhancing relevant audio information? How can we achieve discriminative interaction between the audio and visual modalities? To this end, we propose SDAVS, equipped with the Selective Noise‑Resilient Processor (SNRP) module and the Discriminative Audio‑Visual Mutual Fusion (DAMF) strategy. The proposed SNRP mitigates audio noise interference by selectively emphasizing relevant auditory cues, while DAMF ensures more consistent audio‑visual representations. Experimental results demonstrate that our proposed method achieves state‑of‑the‑art performance on benchmark AVS datasets, especially in multi‑source and complex scenes. The code and model are available at https://github.com/happylife‑pk/SDAVS.
Authors:Zhiwei Wang, Yuxing Li, Meilu Zhu, Defeng He, Edmund Y. Lam
Abstract:
Accurate diagnosis of glaucoma is challenging, as early‑stage changes are subtle and often lack clear structural or appearance cues. Most existing approaches rely on a single modality, such as fundus or optical coherence tomography (OCT), capturing only partial pathological information and often missing early disease progression. In this paper, we propose an iterative multimodal optimization model (IMO) for joint segmentation and grading. IMO integrates fundus and OCT features through a mid‑level fusion strategy, enhanced by a cross‑modal feature alignment (CMFA) module to reduce modality discrepancies. An iterative refinement decoder progressively optimizes the multimodal features through a denoising diffusion mechanism, enabling fine‑grained segmentation of the optic disc and cup while supporting accurate glaucoma grading.
Extensive experiments show that our method effectively integrates multimodal features, providing a comprehensive and clinically significant approach to glaucoma assessment. Source codes are available at https://github.com/warren‑wzw/IMO.git.
Authors:Bang-Dang Pham, Anh Tran, Cuong Pham, Minh Hoai
Abstract:
This paper introduces a novel unsupervised approach for image deblurring that utilizes a simple process for training data collection, thereby enhancing the applicability and effectiveness of deblurring methods. Our technique does not require meticulously paired data of blurred and corresponding sharp images; instead, it uses unpaired blurred and sharp images of similar scenes to generate pseudo‑ground truth data by leveraging a dense matching model to identify correspondences between a blurry image and reference sharp images. Thanks to the simplicity of the training data collection process, our approach does not rely on existing paired training data or pre‑trained networks, making it more adaptable to various scenarios and suitable for networks of different sizes, including those designed for low‑resource devices. We demonstrate that this novel approach achieves state‑of‑the‑art performance, marking a significant advancement in the field of image deblurring.
Authors:Junyao Hu, Zhongwei Cheng, Waikeung Wong, Xingxing Zou
Abstract:
Virtual try‑on (VTON) has advanced single‑garment visualization, yet real‑world fashion centers on full outfits with multiple garments, accessories, fine‑grained categories, layering, and diverse styling, remaining beyond current VTON systems. Existing datasets are category‑limited and lack outfit diversity. We introduce Garments2Look, the first large‑scale multimodal dataset for outfit‑level VTON, comprising 80K many‑garments‑to‑one‑look pairs across 40 major categories and 300+ fine‑grained subcategories. Each pair includes an outfit with 3‑12 reference garment images (Average 4.48), a model image wearing the outfit, and detailed item and try‑on textual annotations. To balance authenticity and diversity, we propose a synthesis pipeline. It involves heuristically constructing outfit lists before generating try‑on results, with the entire process subjected to strict automated filtering and human validation to ensure data quality. To probe task difficulty, we adapt SOTA VTON methods and general‑purpose image editing models to establish baselines. Results show current methods struggle to try on complete outfits seamlessly and to infer correct layering and styling, leading to misalignment and artifacts.
Authors:Shahriar Kabir, Abdullah Muhammed Amimul Ehsan, Istiak Ahmmed Rifti, Md Kaykobad Reza
Abstract:
Automated segmentation of Martian landslides, particularly in tectonically active regions such as Valles Marineris,is important for planetary geology, hazard assessment, and future robotic exploration. However, detecting landslides from planetary imagery is challenging due to the heterogeneous nature of available sensing modalities and the limited number of labeled samples. Each observation combines RGB imagery with geophysical measurements such as digital elevation models, slope maps, thermal inertia, and contextual grayscale imagery, which differ significantly in resolution and statistical properties. To address these challenges, we propose DualSwinFusionSeg, a multimodal segmentation architecture that separates modality‑specific feature extraction and performs multi‑scale cross‑modal fusion. The model employs two parallel Swin Transformer V2 encoders to independently process RGB and auxiliary geophysical inputs, producing hierarchical feature representations. Corresponding features from the two streams are fused at multiple scales and decoded using a UNet++ decoder with dense nested skip connections to preserve fine boundary details. Extensive ablation studies evaluate modality contributions, loss functions, decoder architectures, and fusion strategies. Experiments on the MMLSv2 dataset from the PBVS 2026 Mars‑LS Challenge show that modality‑specific encoders and simple concatenation‑based fusion improve segmentation accuracy under limited training data. The final model achieves 0.867 mIoU and 0.905 F1 on the development benchmark and 0.783 mIoU on the held‑out test set, demonstrating strong performance for multimodal planetary surface segmentation.
Authors:Yichang Xu, Gaowen Liu, Ramana Rao Kompella, Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Ling Liu
Abstract:
This paper presents a multi‑agent perception‑action exploration alliance, dubbed A4VL, for efficient long‑video reasoning. A4VL operates in a multi‑round perception‑action exploration loop with a selection of VLM agents. In each round, the team of agents performs video question‑answer (VideoQA) via perception exploration followed by action exploration. During perception exploration, each agent learns to extract query‑specific perception clue(s) from a few sampled frames and performs clue‑based alignment to find the video block(s) that are most relevant to the query‑specific event. During action exploration, A4VL performs video reasoning in three steps: (1) each agent produces its initial answer with rational, (2) all agents collaboratively scores one another through cross‑reviews and relevance ranking, and (3) based on whether a satisfactory consensus is reached, the decision is made either to start a new round of perception‑action deliberation by pruning (e.g., filtering out the lowest performing agent) and re‑staging (e.g., new‑clue and matching block based perception‑action exploration), or to conclude by producing its final answer. The integration of the multi‑agent alliance through multi‑round perception‑action exploration, coupled with event‑driven partitioning and cue‑guided block alignment, enables A4VL to effectively scale to real world long videos while preserving high quality video reasoning. Evaluation Results on five popular VideoQA benchmarks show that A4VL outperforms 18 existing representative VLMs and 11 recent methods optimized for long‑video reasoning, while achieving significantly lower inference latency. Our code is released at https://github.com/git‑disl/A4VL.
Authors:Neelu Madan, Àlex Pujol, Andreas Møgelmose, Sergio Escalera, Kamal Nasrollahi, Graham W. Taylor, Thomas B. Moeslund
Abstract:
Slot attention has emerged as a powerful framework for unsupervised object‑centric learning, decomposing visual scenes into a small set of compact vector representations called \emphslots, each capturing a distinct region or object. However, these slots are learned in Euclidean space, which provides no geometric inductive bias for the hierarchical relationships that naturally structure visual scenes. In this work, we propose a simple post‑hoc pipeline to project Euclidean slot embeddings onto the Lorentz hyperboloid of hyperbolic space, without modifying the underlying training pipeline. We construct five‑level visual hierarchies directly from slot attention masks and analyse whether hyperbolic geometry reveals latent hierarchical structure that remains invisible in Euclidean space. Integrating our pipeline with SPOT (images), VideoSAUR (video), and SlotContrast (video), We find that hyperbolic projection exposes a consistent scene‑level to object‑level organisation, where coarse slots occupy greater manifold depth than fine slots, which is absent in Euclidean space. We further identify a "curvature‑‑task tradeoff": low curvature (c=0.2) matches or outperforms Euclidean on parent slot retrieval, while moderate curvature (c=0.5) achieves better inter‑level separation. Together, these findings suggest that slot representations already encode latent hierarchy that hyperbolic geometry reveals, motivating end‑to‑end hyperbolic training as a natural next step. Code and models are available at \hrefhttps://github.com/NeeluMadan/HHSgithub.com/NeeluMadan/HHS.
Authors:Wanhu Sun, Zhongjin Luo, Heliang Zheng, Jiahao Chang, Chongjie Ye, Huiang He, Shengchu Zhao, Rongfei Jia, Xiaoguang Han
Abstract:
Part‑level 3D generation is crucial for various downstream applications, including gaming, film production, and industrial design. However, decomposing a 3D shape into geometrically plausible and meaningful components remains a significant challenge. Previous part‑based generation methods often struggle to produce well‑constructed parts, exhibiting poor structural coherence, geometric implausibility, inaccuracy, or inefficiency.
To address these challenges, we introduce EI‑Part, a novel framework specifically designed to generate high‑quality 3D shapes with components, characterized by strong structural coherence, geometric plausibility, geometric fidelity, and generation efficiency. We propose utilizing distinct representations at different stages: an Explode state for part completion and an Implode state for geometry refinement. This strategy fully leverages spatial resolution, enabling flexible part completion and fine geometric detail generation. To maintain structural coherence between parts, a self‑attention mechanism is incorporated in both exploded and imploded states, facilitating effective information perception and feature fusion among components during generation.
Extensive experiments on multiple benchmarks demonstrate that EI‑Part efficiently produces semantically meaningful and structurally coherent parts with fine‑grained geometric details, achieving state‑of‑the‑art performance in part‑level 3D generation.
Project page: https://cvhadessun.github.io/EI‑Part/
Authors:Jiachen Li, Xiaojin Gong, Dongping Zhang
Abstract:
Domain Generalized person Re‑identification (DG Re‑ID) is a challenging task, where models are trained on source domains but tested on unseen target domains. Although previous pure vision‑based models have achieved significant progress, the performance remains further improved. Recently, Vision‑Language Models (VLMs) present outstanding generalization capabilities in various visual applications. However, directly adapting a VLM to Re‑ID shows limited generalization improvement. This is because the VLM only produces with global features that are insensitive to ID nuances. To tacle this problem, we propose a CLIP‑based multi‑grained vision‑language alignment framework in this work. Specifically, several multi‑grained prompts are introduced in language modality to describe different body parts and align with their counterparts in vision modality. To obtain fine‑grained visual information, an adaptively masked multi‑head self‑attention module is employed to precisely extract specific part features. To train the proposed module, an MLLM‑based visual grounding expert is employed to automatically generate pseudo labels of body parts for supervision. Extensive experiments conducted on both single‑ and multi‑source generalization protocols demonstrate the superior performance of our approach. The implementation code will be released at https://github.com/RikoLi/MUVA.
Authors:Emmanuel Oladokun, Sarina Thomas, Jurica Šprem, Vicente Grau
Abstract:
Echocardiography is widely used for assessing cardiac function, where clinically meaningful parameters such as left‑ventricular ejection fraction (EF) play a central role in diagnosis and management. Generative models capable of synthesising realistic echocardiogram videos with explicit control over such parameters are valuable for data augmentation, counterfactual analysis, and specialist training. However, existing approaches typically rely on computationally expensive multi‑step sampling and aggressive temporal normalisation, limiting efficiency and applicability to heterogeneous real‑world data.
We introduce EchoLVFM, a one‑step latent video flow‑matching framework for controllable echocardiogram generation. Operating in the latent space, EchoLVFM synthesises temporally coherent videos in a single inference step, achieving a \mathbf~ 50× improvement in sampling efficiency compared to multi‑step flow baselines while maintaining visual fidelity. The model supports global conditioning on clinical variables, demonstrated through precise control of EF, and enables reconstruction and counterfactual generation from partially observed sequences. A masked conditioning strategy further removes fixed‑length constraints, allowing shorter sequences to be retained rather than discarded.
We evaluate EchoLVFM on the CAMUS dataset under challenging single‑frame conditioning. Quantitative and qualitative results demonstrate competitive video quality, strong EF adherence, and 57.9% discrimination accuracy by expert clinicians which is close to chance. These findings indicate that efficient, one‑step flow matching can enable practical, controllable echocardiogram video synthesis without sacrificing fidelity. Code available at: https://github.com/EngEmmanuel/EchoLVFM
Authors:Hiroto Nakata, Yawen Zou, Shunsuke Sakai, Shun Maeda, Chunzhi Gu, Yijin Wei, Shangce Gao, Chao Zhang
Abstract:
Logical anomaly detection in industrial inspection remains challenging due to variations in visual appearance (e.g., background clutter, illumination shift, and blur), which often distract vision‑centric detectors from identifying rule‑level violations. However, existing benchmarks rarely provide controlled settings where logical states are fixed while such nuisance factors vary. To address this gap, we introduce VID‑AD, a dataset for logical anomaly detection under vision‑induced distraction. It comprises 10 manufacturing scenarios and five capture conditions, totaling 50 one‑class tasks and 10,395 images. Each scenario is defined by two logical constraints selected from quantity, length, type, placement, and relation, with anomalies including both single‑constraint and combined violations. We further propose a language‑based anomaly detection framework that relies solely on text descriptions generated from normal images. Using contrastive learning with positive texts and contradiction‑based negative texts synthesized from these descriptions, our method learns embeddings that capture logical attributes rather than low‑level features. Extensive experiments demonstrate consistent improvements over baselines across the evaluated settings. The dataset is available at: https://github.com/nkthiroto/VID‑AD.
Authors:Kursat Komurcu, Linas Petkevicius
Abstract:
Predicting satellite imagery requires a balance between structural accuracy and textural detail. Standard deterministic methods like PredRNN or SimVP minimize pixel‑based errors but suffer from the "regression to the mean" problem, producing blurry outputs that obscure subtle geographic‑spatial features. Generative models provide realistic textures but often misleadingly reveal structural anomalies. To bridge this gap, we introduce Sat‑JEPA‑Diff, which combines Self‑Supervised Learning (SSL) with Hidden Diffusion Models (LDM). An IJEPA module predicts stable semantic representations, which then route a frozen Stable Diffusion backbone via a lightweight cross‑attention adapter. This ensures that the synthesized high‑accuracy textures are based on absolutely accurate structural predictions. Evaluated on a global Sentinel‑2 dataset, Sat‑JEPA‑Diff excels at resolving sharp boundaries. It achieves leading perceptual scores (GSSIM: 0.8984, FID: 0.1475) and significantly outperforms deterministic baselines, despite standard autoregressive stability limits. The code and dataset are publicly available on https://github.com/VU‑AIML/SAT‑JEPA‑DIFF.
Authors:Jonas V. Funk, Lukas Roming, Andreas Michel, Paul Bäcker, Georg Maier, Thomas Längle, Markus Klute
Abstract:
Growing waste streams and the transition to a circular economy require efficient automated waste sorting. In industrial settings, materials move on fast conveyor belts, where reliable identification and ejection demand pixel‑accurate segmentation. RGB imaging delivers high‑resolution spatial detail, which is essential for accurate segmentation, but it confuses materials that look similar in the visible spectrum. Hyperspectral imaging (HSI) provides spectral signatures that separate such materials, yet its lower spatial resolution limits detail. Effective waste sorting therefore needs methods that fuse both modalities to exploit their complementary strengths. We present Bidirectional Cross‑Attention Fusion (BCAF), which aligns high‑resolution RGB with low‑resolution HSI at their native grids via localized, bidirectional cross‑attention, avoiding pre‑upsampling or early spectral collapse. BCAF uses two independent backbones: a standard Swin Transformer for RGB and an HSI‑adapted Swin backbone that preserves spectral structure through 3D tokenization with spectral self‑attention. We also analyze trade‑offs between RGB input resolution and the number of HSI spectral slices. Although our evaluation targets RGB‑HSI fusion, BCAF is modality‑agnostic and applies to co‑registered RGB with lower‑resolution, high‑channel auxiliary sensors. On the benchmark SpectralWaste dataset, BCAF achieves state‑of‑the‑art performance of 76.4% mIoU at 31 images/s and 75.4% mIoU at 55 images/s. We further evaluate a novel industrial dataset: K3I‑Cycling (first RGB subset already released on Fordatis). On this dataset, BCAF reaches 62.3% mIoU for material segmentation (paper, metal, plastic, etc.) and 66.2% mIoU for plastic‑type segmentation (PET, PP, HDPE, LDPE, PS, etc.). Code and model checkpoints are publicly available at https://github.com/jonasvilhofunk/BCAF_2026 .
Authors:Seokmin Lee, Yunghee Lee, Byeonghyun Pak, Byeongju Woo
Abstract:
For robotic agents operating in dynamic environments, learning visual state representations from streaming video observations is essential for sequential decision making. Recent self‑supervised learning methods have shown strong transferability across vision tasks, but they do not explicitly address what a good visual state should encode. We argue that effective visual states must capture what‑is‑where by jointly encoding the semantic identities of scene elements and their spatial locations, enabling reliable detection of subtle dynamics across observations. To this end, we propose CroBo, a visual state representation learning framework based on a global‑to‑local reconstruction objective. Given a reference observation compressed into a compact bottleneck token, CroBo learns to reconstruct heavily masked patches in a local target crop from sparse visible cues, using the global bottleneck token as context. This learning objective encourages the bottleneck token to encode a fine‑grained representation of scene‑wide semantic entities, including their identities, spatial locations, and configurations. As a result, the learned visual states reveal how scene elements move and interact over time, supporting sequential decision making. We evaluate CroBo on diverse vision‑based robot policy learning benchmarks, where it achieves state‑of‑the‑art performance. Reconstruction analyses and perceptual straightness experiments further show that the learned representations preserve pixel‑level scene composition and encode what‑moves‑where across observations. Project page available at: https://seokminlee‑chris.github.io/CroBo‑ProjectPage.
Authors:Qilong Li, Chongsheng Zhang
Abstract:
Scene text recognition (STR) methods have demonstrated their excellent capability in English text images. However, due to the complex inner structures of Chinese and the extensive character categories, it poses challenges for recognizing Chinese text in images. Recently, studies have shown that the methods designed for English text recognition encounter an accuracy bottleneck when recognizing Chinese text images. This raises the question: Is it appropriate to apply the model developed for English to the Chinese STR task? To explore this issue, we propose a novel method named LER, which explicitly decouples each character and independently recognizes characters while taking into account the complex inner structures of Chinese. LER consists of three modules: Localization, Extraction, and Recognition. Firstly, the localization module utilizes multimodal information to determine the character's position precisely. Then, the extraction module dissociates all characters in parallel. Finally, the recognition module considers the unique inner structures of Chinese to provide the text prediction results. Extensive experiments conducted on large‑scale Chinese benchmarks indicate that our method significantly outperforms existing methods. Furthermore, extensive experiments conducted on six English benchmarks and the Union14M benchmark show impressive results in English text recognition by LER. Code is available at https://github.com/Pandarenlql/LER.
Authors:Bohan Zhang, Weidong Tang, Zhixiang Chi, Yi Jin, Zhenbo Li, Yang Wang, Yanan Wu
Abstract:
On‑the‑Fly Category Discovery (OCD) aims to recognize known classes while simultaneously discovering emerging novel categories during inference, using supervision only from known classes during offline training. Existing approaches rely either on fixed label supervision or on diffusion‑based augmentations to enhance the backbone, yet none of them explicitly train the model to perform the discovery task required at test time. It is fundamentally unreasonable to expect a model optimized on limited labeled data to carry out a qualitatively different discovery objective during inference. This mismatch creates a clear optimization misalignment between the offline learning stage and the online discovery stage. In addition, prior methods often depend on hash‑based encodings or severe feature compression, which further limits representational capacity. To address these issues, we propose Learning through Creation (LTC), a fully feature‑based and hash‑free framework that injects novel‑category awareness directly into offline learning. At its core is a lightweight, online pseudo‑unknown generator driven by kernel‑energy minimization and entropy maximization (MKEE). Unlike previous methods that generate synthetic samples once before training, our generator evolves jointly with the model dynamics and synthesizes pseudo‑novel instances on the fly at negligible cost. These samples are incorporated through a dual max‑margin objective with adaptive thresholding, strengthening the model's ability to delineate and detect unknown regions through explicit creation. Extensive experiments across seven benchmarks show that LTC consistently outperforms prior work, achieving improvements ranging from 1.5 percent to 13.1 percent in all‑class accuracy. The code is available at https://github.com/brandinzhang/LTC
Authors:Jun Lu, Zehao Sang, Haoqi Wei, Xiangyun Liu, Kun Zhu, Haitao Guo, Zhihui Gong, Lei Ding
Abstract:
Cross‑View Geo‑Localization (CVGL) in remote sensing aims to locate a drone‑view query by matching it to geo‑tagged satellite images. Although supervised methods have achieved strong results on closeset benchmarks, they often fail to generalize to unconstrained, real‑world scenarios due to severe viewpoint differences and dataset bias. To overcome these limitations, we present VFM‑Loc, a training‑free framework for zero‑shot CVGL that leverages the generalizable visual representations from vision foundational models (VFMs). VFM‑Loc identifies and matches discriminative visual clues across different viewpoints through a progressive alignment strategy. First, we design a hierarchical clue extraction mechanism using Generalized Mean pooling and Scale‑Weighted RMAC to preserve distinctive visual clues across scales while maintaining hierarchical confidence. Second, we introduce a statistical manifold alignment pipeline based on domain‑wise PCA and Orthogonal Procrustes analysis, linearly aligning heterogeneous feature distributions in a shared metric space. Experiments demonstrate that VFM‑Loc exhibits strong zero‑shot accuracy on standard benchmarks and surpasses supervised methods by over 20% in Recall@1 on the challenging LO‑UCV dataset with large oblique angles. This work highlights that principled alignment of pre‑trained features can effectively bridge the cross‑view gap, establishing a robust and training‑free paradigm for real‑world CVGL. The relevant code is made available at: https://github.com/DingLei14/VFM‑Loc.
Authors:Xi Jiang, Yue Guo, Jian Li, Yong Liu, Bin-Bin Gao, Hanqiu Deng, Jun Liu, Heng Zhao, Chengjie Wang, Feng Zheng
Abstract:
Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform in industrial anomaly detection (IAD). This is because MLLMs trained mostly on general web data differ significantly from industrial images. Moreover, they encode each image independently and can only compare images in the language space, making them insensitive to subtle visual differences that are key to IAD. To tackle these issues, we present AD‑Copilot, an interactive MLLM specialized for IAD via visual in‑context comparison. We first design a novel data curation pipeline to mine inspection knowledge from sparsely labeled industrial images and generate precise samples for captioning, VQA, and defect localization, yielding a large‑scale multimodal dataset Chat‑AD rich in semantic signals for IAD. On this foundation, AD‑Copilot incorporates a novel Comparison Encoder that employs cross‑attention between paired image features to enhance multi‑image fine‑grained perception, and is trained with a multi‑stage strategy that incorporates domain knowledge and gradually enhances IAD skills. In addition, we introduce MMAD‑BBox, an extended benchmark for anomaly localization with bounding‑box‑based evaluation. The experiments show that AD‑Copilot achieves 82.3% accuracy on the MMAD benchmark, outperforming all other models without any data leakage. In the MMAD‑BBox test, it achieves a maximum improvement of 3.35× over the baseline. AD‑Copilot also exhibits excellent generalization of its performance gains across other specialized and general‑purpose benchmarks. Remarkably, AD‑Copilot surpasses human expert‑level performance on several IAD tasks, demonstrating its potential as a reliable assistant for real‑world industrial inspection. All datasets and models will be released for the broader benefit of the community.
Authors:Zhexiao Xiong, Yizhi Song, Liu He, Wei Xiong, Yu Yuan, Feng Qiao, Nathan Jacobs
Abstract:
Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incoherent content that violates basic physical intuition, significantly limiting their practical applicability. We propose PhysAlign, an efficient framework for physics‑coherent image‑to‑video (I2V) generation that explicitly addresses this limitation. To overcome the critical scarcity of physics‑annotated videos, we first construct a fully controllable synthetic data generation pipeline based on rigid‑body simulation, yielding a highly‑curated dataset with accurate, fine‑grained physics and 3D annotations. Leveraging this data, PhysAlign constructs a unified physical latent space by coupling explicit 3D geometry constraints with a Gram‑based spatio‑temporal relational alignment that extracts kinematic priors from video foundation models. Extensive experiments demonstrate that PhysAlign significantly outperforms existing VDMs on tasks requiring complex physical reasoning and temporal stability, without compromising zero‑shot visual quality. PhysAlign shows the potential to bridge the gap between raw visual synthesis and rigid‑body kinematics, establishing a practical paradigm for genuinely physics‑grounded video generation. The project page is available at https://physalign.github.io/PhysAlign.
Authors:Tajamul Ashraf, Tavaheed Tariq, Sonia Yadav, Abrar Ul Riyaz, Wasif Tak, Moloud Abdar, Janibul Bashir
Abstract:
Multi‑object tracking (MOT) has traditionally focused on estimating trajectories of all objects in a video, without selectively reasoning about user‑specified targets under semantic instructions. In this work, we introduce a query‑driven tracking paradigm that formulates tracking as a spatiotemporal reasoning problem conditioned on natural language queries. Given a reference frame, a video sequence, and a textual query, the goal is to localize and track only the target(s) specified in the query while maintaining temporal coherence and identity consistency. To support this setting, we construct RMOT26, a large‑scale benchmark with grounded queries and sequence‑level splits to prevent identity leakage and enable robust evaluation of generalization. We further present QTrack, an end‑to‑end vision‑language model that integrates multimodal reasoning with tracking‑oriented localization. Additionally, we introduce a Temporal Perception‑Aware Policy Optimization strategy with structured rewards to encourage motion‑aware reasoning. Extensive experiments demonstrate the effectiveness of our approach for reasoning‑centric, language‑guided tracking. Code and data are available at https://github.com/gaash‑lab/QTrack
Authors:Zengyan Wang, Sirshapan Mitra, Rajat Modi, Grace Lim, Yogesh Rawat
Abstract:
We introduce Sky2Ground, a three‑view dataset designed for varying altitude camera localization, correspondence learning, and reconstruction. The dataset combines structured synthetic imagery with real, in‑the‑wild images, providing both controlled multi‑view geometry and realistic scene noise. Each of the 51 sites contains thousands of satellite, aerial, and ground images spanning wide altitude ranges and nearly orthogonal viewing angles, enabling rigorous evaluation across global‑to‑local contexts. We benchmark state of the art pose estimation models, including MASt3R, DUSt3R, Map Anything, and VGGT, and observe that the use of satellite imagery often degrades performance, highlighting the challenges under large altitude variations. We also examine reconstruction methods, highlighting the challenges introduced by sparse geometric overlap, varying perspectives, and the use of real imagery, which often introduces noise and reduces rendering quality. To address some of these challenges, we propose SkyNet, a model which enhances cross‑view consistency when incorporating satellite imagery with a curriculum‑based training strategy to progressively incorporate more satellite views. SkyNet significantly strengthens multi‑view alignment and outperforms existing methods by 9.6% on RRA@5 and 18.1% on RTA@5 in terms of absolute performance. Sky2Ground and SkyNet together establish a comprehensive testbed and baseline for advancing large‑scale, multi‑altitude 3D perception and generalizable camera localization. Code and models will be released publicly for future research.Project page: https://sky2ground2026.github.io/sky2ground/
Authors:Bo Ma, Wei Qi Yan, Jinsong Wu
Abstract:
Learning systems that preserve privacy often inject noise into hierarchical visual representations; a central challenge is to \emphmodel how such perturbations align with a declared privacy budget in a way that is interpretable and applicable across vision backbones and vision‑‑language models (VLMs). We propose \emphBodhi VLM, a \emphprivacy‑alignment modeling framework for \emphhierarchical neural representations: it (1) links sensitive concepts to layer‑wise grouping via NCP and MDAV‑based clustering; (2) locates sensitive feature regions using bottom‑up (BUA) and top‑down (TDA) strategies over multi‑scale representations (e.g., feature pyramids or vision‑encoder layers); and (3) uses an Expectation‑Maximization Privacy Assessment (EMPA) module to produce an interpretable \emphbudget‑alignment signal by comparing the fitted sensitive‑feature distribution to an evaluator‑specified reference (e.g., Laplace or Gaussian with scale c/ε). The output is reference‑relative and is \emphnot a formal differential‑privacy estimator. We formalize BUA/TDA over hierarchical feature structures and validate the framework on object detectors (YOLO, PPDPTS, DETR) and on the \emphvisual encoders of VLMs (CLIP, LLaVA, BLIP). BUA and TDA yield comparable deviation trends; EMPA provides a stable alignment signal under the reported setups. We compare with generic discrepancy baselines (Chi‑square, K‑L, MMD) and with task‑relevant baselines (MomentReg, NoiseMLE, Wass‑1). Results are reported as mean\pmstd over multiple seeds with confidence intervals in the supplementary materials. This work contributes a learnable, interpretable modeling perspective for privacy‑aligned hierarchical representations rather than a post hoc audit only. Source code: \hrefhttps://github.com/mabo1215/bodhi‑vlm.gitBodhi‑VLM GitHub repository
Authors:Bo Ma, Jinsong Wu, Wei Qi Yan
Abstract:
Sensitive data release is vulnerable to output‑side privacy threats such as membership inference, attribute inference, and record linkage. This creates a practical need for release mechanisms that provide formal privacy guarantees while preserving utility in measurable ways. We propose REAEDP, a differential privacy framework that combines entropy‑calibrated histogram release, a synthetic‑data release mechanism, and attack‑based evaluation. On the theory side, we derive an explicit sensitivity bound for Shannon entropy, together with an extension to Rényi entropy, for adjacent histogram datasets, enabling calibrated differentially private release of histogram statistics. We further study a synthetic‑data mechanism \mathcalF with a privacy‑test structure and show that it satisfies a formal differential privacy guarantee under the stated parameter conditions. On multiple public tabular datasets, the empirical entropy change remains below the theoretical bound in the tested regime, standard Laplace and Gaussian baselines exhibit comparable trends, and both membership‑inference and linkage‑style attack performance move toward random‑guess behavior as the privacy parameter decreases. These results support REAEDP as a practically usable privacy‑preserving release pipeline in the tested settings. Source code: https://github.com/mabo1215/REAEDP.git
Authors:Chen Zhenyuan, Zhang Zechuan, Zhang Feng
Abstract:
In this paper, we explore text‑guided image editing in the remote sensing domain using generative modeling. We propose \rsedit, a collection of models from U‑Net to DiT with various configurations. Specifically, we present the first comprehensive study of conditioning strategies for building image editing models from off‑the‑shelf text‑to‑image ones. Our experiments show that \rsedit achieves the best instruction‑faithful edits while preserving geospatial structure. We release the code at \urlhttps://github.com/Bili‑Sakura/RSEdit‑Preview and checkpoints at \urlhttps://huggingface.co/collections/BiliSakura/rsedit.
Authors:Bo Ma, Jinsong Wu, Weiqi Yan
Abstract:
Multi‑object tracking in video often requires appearance or location cues that can reveal sensitive identity information, while adding privacy‑preserving noise typically disrupts cross‑frame association and causes ID switches or target loss. We propose TSDCRF, a plug‑in refinement framework that balances privacy and tracking by combining three components: (i) (\varepsilon,δ)‑differential privacy via calibrated Gaussian noise on sensitive regions under a configurable privacy budget; (ii) a Normalized Control Penalty (NCP) that down‑weights unstable or conflicting class predictions before noise injection to stabilize association; and (iii) a time‑series dynamic conditional random field (DCRF) that enforces temporal consistency and corrects trajectory deviation after noise, mitigating ID switches and resilience to trajectory hijacking. The pipeline is agnostic to the choice of detector and tracker (e.g., YOLOv4 and DeepSORT). We evaluate on MOT16, MOT17, Cityscapes, and KITTI. Results show that TSDCRF achieves a better privacy‑‑utility trade‑off than white noise and prior methods (NTPD, PPDTSA): lower KL‑divergence shift, lower tracking RMSE, and improved robustness under trajectory hijacking while preserving privacy. Source code in https://github.com/mabo1215/TSDCRF.git
Authors:Yunhe Gao, Yabin Zhang, Chong Wang, Jiaming Liu, Maya Varma, Jean-Benoit Delbrouck, Akshay Chaudhari, Curtis Langlotz
Abstract:
Foundation models have transformed vision and language by learning general‑purpose representations from large‑scale unlabeled data, yet 3D medical imaging lacks analogous approaches. Existing self‑supervised methods rely on low‑level reconstruction or contrastive objectives that fail to capture the anatomical semantics critical for medical image analysis, limiting transfer to downstream tasks. We present MASS (MAsk‑guided Self‑Supervised learning), which treats in‑context segmentation as the pretext task for learning general‑purpose medical imaging representations. MASS's key insight is that automatically generated class‑agnostic masks provide sufficient structural supervision for learning semantically rich representations. By training on thousands of diverse mask proposals spanning anatomical structures and pathological findings, MASS learns what semantically defines medical structures: the holistic combination of appearance, shape, spatial context, and anatomical relationships. We demonstrate effectiveness across data regimes: from small‑scale pretraining on individual datasets (20‑200 scans) to large‑scale multi‑modal pretraining on 5K CT, MRI, and PET volumes, all without annotations. MASS demonstrates: (i) few‑shot segmentation on novel structures, (ii) matching full supervision with only 20‑40% labeled data while outperforming self‑supervised baselines by over 20 in Dice score in low‑data regimes, and (iii) frozen‑encoder classification on unseen pathologies that matches full supervised training with thousands of samples. Mask‑guided self‑supervised pretraining captures broadly generalizable knowledge, opening a path toward 3D medical imaging foundation models without expert annotations. Code is available: https://github.com/Stanford‑AIMI/MASS.
Authors:Xiaoqiong Liu, Heng Fan
Abstract:
Recently, feature upsampling has gained increasing attention owing to its effectiveness in enhancing vision foundation models (VFMs) for pixel‑level understanding tasks. Existing methods typically rely on high‑resolution features from the same foundation model to achieve upsampling via self‑reconstruction. However, relying solely on intra‑model features forces the upsampler to overfit to the source model's inherent location misalignment and high‑norm artifacts. To address this fundamental limitation, we propose DiveUp, a novel framework that breaks away from single‑model dependency by introducing multi‑VFM relational guidance. Instead of naive feature fusion, DiveUp leverages diverse VFMs as a panel of experts, utilizing their structural consensus to regularize the upsampler's learning process, effectively preventing the propagation of inaccurate spatial structures from the source model. To reconcile the unaligned feature spaces across different VFMs, we propose a universal relational feature representation, formulated as a local center‑of‑mass (COM) field, that extracts intrinsic geometric structures, enabling seamless cross‑model interaction. Furthermore, we introduce a spikiness‑aware selection strategy that evaluates the spatial reliability of each VFM, effectively filtering out high‑norm artifacts to aggregate guidance from only the most reliable expert at each local region. DiveUp is a unified, encoder‑agnostic framework; a jointly‑trained model can universally upsample features from diverse VFMs without requiring per‑model retraining. Extensive experiments demonstrate that DiveUp achieves state‑of‑the‑art performance across various downstream dense prediction tasks, validating the efficacy of multi‑expert relational guidance. Our code and models are available at: https://github.com/Xiaoqiong‑Liu/DiveUp
Authors:Alessandro Pesci, Valerio Guarrasi, Marco Alì, Isabella Castiglioni, Paolo Soda
Abstract:
The translation from Magnetic resonance imaging (MRI) to Computed tomography (CT) has been proposed as an effective solution to facilitate MRI‑only clinical workflows while limiting exposure to ionizing radiation. Although numerous Generative Adversarial Network (GAN) architectures have been proposed for MRI‑to‑CT translation, systematic and fair comparisons across heterogeneous models remain limited. We present a comprehensive benchmark of ten GAN architectures evaluated on the SynthRAD2025 dataset across three anatomical districts (abdomen, thorax, head‑and‑neck). All models were trained under a unified validation protocol with identical preprocessing and optimization settings. Performance was assessed using complementary metrics capturing voxel‑wise accuracy, structural fidelity, perceptual quality, and distribution‑level realism, alongside an analysis of computational complexity. Supervised Paired models consistently outperformed Unpaired approaches, confirming the importance of voxel‑wise supervision. Pix2Pix achieved the most balanced performance across districts while maintaining a favorable quality‑to‑complexity trade‑off. Multi‑district training improved structural robustness, whereas intra‑district training maximized voxel‑wise fidelity. This benchmark provides quantitative and computational guidance for model selection in MRI‑only radiotherapy workflows and establishes a reproducible framework for future comparative studies. To ensure the reproducibility of our experiments we make our code public, together with the overall results, at the following link:https://github.com/arco‑group/MRI_TO_CT.git
Authors:Eric Nazarenus, Chuqiao Li, Yannan He, Xianghui Xie, Jan Eric Lenssen, Gerard Pons-Moll
Abstract:
We present ActionPlan, a unified motion diffusion framework that bridges real‑time streaming with high‑quality offline generation within a single model. The core idea is to introduce a per‑frame action plan: the model predicts frame‑level text latents that act as dense semantic anchors throughout denoising, and uses them to denoise the full motion sequence with combined semantic and motion cues. To support this structured workflow, we design latent‑specific diffusion steps, allowing each motion latent to be denoised independently and sampled in flexible orders at inference. As a result, ActionPlan can run in a history‑conditioned, future‑aware mode for real‑time streaming, while also supporting high‑quality offline generation. The same mechanism further enables zero‑shot motion editing and in‑betweening without additional models. Experiments demonstrate that our real‑time streaming is 5.25x faster while also achieving 18% motion quality improvement over the best previous method in terms of FID.
Authors:Pratik Ramesh, George Stoica, Arun Iyer, Leshem Choshen, Judy Hoffman
Abstract:
Model merging has shown that multitask models can be created by directly combining the parameters of different models that are each specialized on tasks of interest. However, models trained independently on distinct tasks often exhibit interference that degrades the merged model's performance. To solve this problem, we formally define the notion of Cross‑Task Interference as the drift in the representation of the merged model relative to its constituent models. Reducing cross‑task interference is key to improving merging performance. To address this issue, we propose our method, Resolving Interference (RI), a light‑weight adaptation framework which disentangles expert models to be functionally orthogonal to the space of other tasks, thereby reducing cross‑task interference. RI does this whilst using only unlabeled auxiliary data as input (i.e., no task‑data is needed), allowing it to be applied in data‑scarce scenarios. RI consistently improves the performance of state‑of‑the‑art merging methods by up to 3.8% and generalization to unseen domains by up to 2.3%. We also find RI to be robust to the source of auxiliary input while being significantly less sensitive to tuning of merging hyperparameters. Our codebase is available at: https://github.com/pramesh39/resolving_interference
Authors:Liang Tang, Hongda Li, Jiayu Zhang, Long Chen, Shuxian Li, Siqi Pei, Tiaonan Duan, Yuhao Cheng
Abstract:
Emotion recognition in videos is a pivotal task in affective computing, where identifying subtle psychological states such as Ambivalence and Hesitancy holds significant value for behavioral intervention and digital health. Ambivalence and Hesitancy states often manifest through cross‑modal inconsistencies such as discrepancies between facial expressions, vocal tones, and textual semantics, posing a substantial challenge for automated recognition. This paper proposes a recognition framework that integrates temporal segment modeling with Multimodal Large Language Models. To address computational efficiency and token constraints in long video processing, we employ a segment‑based strategy, partitioning videos into short clips with a maximum duration of 5 seconds. We leverage the Qwen3‑Omni‑30B‑A3B model, fine‑tuned on the BAH dataset using LoRA and full‑parameter strategies via the MS‑Swift framework, enabling the model to synergistically analyze visual and auditory signals. Experimental results demonstrate that the proposed method achieves an accuracy of 85.1% on the test set, significantly outperforming existing benchmarks and validating the superior capability of Multimodal Large Language Models in capturing complex and nuanced emotional conflicts. The code is released at https://github.com/dlnn123/A‑H‑Detection‑with‑Qwen‑Omni.git.
Authors:Yang Yang, Tianyi Zhang, Wei Huang, Jinwei Chen, Boxi Wu, Xiaofei He, Deng Cai, Bo Li, Peng-Tao Jiang
Abstract:
Interactive long video generation requires prompt switching to introduce new subjects or events, while maintaining perceptual fidelity and coherent motion over extended horizons. Recent distilled streaming video diffusion models reuse a rolling KV cache for long‑range generation, enabling prompt‑switch interaction through re‑cache at each switch. However, existing streaming methods still exhibit progressive quality degradation and weakened motion dynamics. We identify two failure modes specific to interactive streaming generation: (i) at each prompt switch, current cache maintenance cannot simultaneously retain KV‑based semantic context and recent latent cues, resulting in weak boundary conditioning and reduced perceptual quality; and (ii) during distillation, unbounded time indexing induces a positional distribution shift from the pretrained backbone's bounded RoPE regime, weakening pretrained motion priors and long‑horizon motion retention. To address these issues, we propose Anchor Forcing, a cache‑centric framework with two designs. First, an anchor‑guided re‑cache mechanism stores KV states in anchor caches and warm‑starts re‑cache from these anchors at each prompt switch, reducing post‑switch evidence loss and stabilizing perceptual quality. Second, a tri‑region RoPE with region‑specific reference origins, together with RoPE re‑alignment distillation, reconciles unbounded streaming indices with the pretrained RoPE regime to better retain motion priors. Experiments on long videos show that our method improves perceptual quality and motion metrics over prior streaming baselines in interactive settings. Project page: https://github.com/vivoCameraResearch/Anchor‑Forcing
Authors:Zhaoyu Liu, Xi Weng, Lianyu Hu, Zhe Hou, Kan Jiang, Jin Song Dong, Yang Liu
Abstract:
Tennis is one of the most widely followed sports, generating extensive broadcast footage with strong potential for professional analysis, automated coaching, and real‑time commentary. However, automatic tennis understanding remains underexplored due to two key challenges: (1) the lack of large‑scale benchmarks with fine‑grained annotations and expert‑level commentary, and (2) the difficulty of building accurate yet efficient multimodal systems suitable for real‑time deployment. To address these challenges, we introduce TennisVL, a large‑scale tennis benchmark comprising over 200 professional matches (471.9 hours) and 40,000+ rally‑level clips. Unlike existing commentary datasets that focus on descriptive play‑by‑play narration, TennisVL emphasizes expert analytical commentary capturing tactical reasoning, player decisions, and match momentum. Furthermore, we propose TennisExpert, a multimodal tennis understanding framework that integrates a video semantic parser with a memory‑augmented model built on Qwen3‑VL‑8B. The parser extracts key match elements (e.g., scores, shot sequences, ball bounces, and player locations), while hierarchical memory modules capture both short‑ and long‑term temporal context. Experiments show that TennisExpert consistently outperforms strong proprietary baselines, including GPT‑5, Gemini, and Claude, and demonstrates improved ability to capture tactical context and match dynamics. Our dataset and code are publicly available at https://github.com/LZYAndy/TennisExpert.
Authors:Sihan Cao, Jianwei Zhang, Pengcheng Zheng, Jiaxin Yan, Caiyan Qin, Yalan Ye, Wei Dong, Peng Wang, Yang Yang, Chaoning Zhang
Abstract:
Large Vision‑Language Models (LVLMs) incur substantial inference costs due to the processing of a vast number of visual tokens. Existing methods typically struggle to model progressive visual token reduction as a multi‑step decision process with sequential dependencies and often rely on hand‑engineered scoring rules that lack adaptive optimization for complex reasoning trajectories. To overcome these limitations, we propose TPRL, a reinforcement learning framework that learns adaptive pruning trajectories through language‑guided sequential optimization tied directly to end‑task performance. We formulate visual token pruning as a sequential decision process with explicit state transitions and employ a self‑supervised autoencoder to compress visual tokens into a compact state representation for efficient policy learning. The pruning policy is initialized through learning from demonstrations and subsequently fine‑tuned using Proximal Policy Optimization (PPO) to jointly optimize task accuracy and computational efficiency. Our experimental results demonstrate that TPRL removes up to 66.7% of visual tokens and achieves up to a 54.2% reduction in FLOPs during inference while maintaining a near‑lossless average accuracy drop of only 0.7%. Code is released at \hrefhttps://github.com/MagicVicCoder/TPRL\textcolormypinkhttps://github.com/MagicVicCoder/TPRL.
Authors:Zongqing Li, Zhihui Liu, Yujie Xie, Shansiyuan Wu, Hongshen Lv, Songzhi Su
Abstract:
Instruction‑based image editing aims to modify source content according to textual instructions. However, existing methods built upon flow matching often struggle to maintain consistency in non‑edited regions due to denoising‑induced reconstruction errors that cause drift in preserved content. Moreover, they typically lack fine‑grained control over edit strength. To address these limitations, we propose VeloEdit, a training‑free method that enables highly consistent and continuously controllable editing. VeloEdit dynamically identifies editing regions by quantifying the discrepancy between the velocity fields responsible for preserving source content and those driving the desired edits. Based on this partition, we enforce consistency in preservation regions by substituting the editing velocity with the source‑restoring velocity, while enabling continuous modulation of edit intensity in target regions via velocity interpolation. Unlike prior works that rely on complex attention manipulation or auxiliary trainable modules, VeloEdit operates directly on the velocity fields. Extensive experiments on Flux.1 Kontext and Qwen‑Image‑Edit demonstrate that VeloEdit improves visual consistency and editing continuity with negligible additional computational cost. Code is available at https://github.com/xmulzq/VeloEdit.
Authors:Jiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha, Qiaoyu Tan
Abstract:
Vision‑Language Models (VLMs) have demonstrated remarkable capabilities in aligning and understanding multimodal signals, yet their potential to reason over structured data, where multimodal entities are connected through explicit relational graphs, remains largely underexplored. Unlocking this capability is crucial for real‑world applications such as social networks, recommendation systems, and scientific discovery, where multimodal information is inherently structured. To bridge this gap, we present GraphVLM, a systematic benchmark designed to evaluate and harness the capabilities of VLMs for multimodal graph learning (MMGL). GraphVLM investigates three complementary paradigms for integrating VLMs with graph reasoning: (1) VLM‑as‑Encoder, which enriches graph neural networks through multimodal feature fusion; (2) VLM‑as‑Aligner, which bridges modalities in latent or linguistic space to facilitate LLM‑based structured reasoning; and (3) VLM‑as‑Predictor, which directly employs VLMs as multimodal backbones for graph learning tasks. Extensive experiments across six datasets from diverse domains demonstrate that VLMs enhance multimodal graph learning via all three roles. Among these paradigms, VLM‑as‑Predictor achieves the most substantial and consistent performance gains, revealing the untapped potential of vision‑language models as a new foundation for multimodal graph learning. The benchmark code is publicly available at https://github.com/oamyjin/GraphVLM.
Authors:Zhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li, Guangyao Chen
Abstract:
Source‑Free Cross‑Domain Few‑Shot Learning (SF‑CDFSL) focuses on fine‑tuning with limited training data from target domains (e.g., medical or satellite images), where Vision‑Language Models (VLMs) such as CLIP and SigLIP have shown promising results. Current works in traditional visual models suggest that improving visual discriminability enhances performance. However, in VLM‑based SF‑CDFSL tasks, we find that strengthening visual‑modal discriminability actually suppresses VLMs' performance. In this paper, we aim to delve into this phenomenon for an interpretation and a solution. By both theoretical and experimental proofs, our study reveals that fine‑tuning with the typical cross‑entropy loss (\mathcalL_\mathrmvlm) inherently includes a visual learning part and a cross‑modal learning part, where the cross‑modal part is crucial for rectifying the heavily disrupted modality misalignment in SF‑CDFSL. However, we find that the visual learning essentially acts as a shortcut that encourages the model to reduce \mathcalL_\mathrmvlm without considering the cross‑modal part, therefore hindering the cross‑modal alignment and harming the performance. Based on this interpretation, we further propose an approach to address this problem: first, we perturb the visual learning to guide the model to focus on the cross‑modal alignment. Then, we use the visual‑text semantic relationships to gradually align the visual and textual modalities during the fine‑tuning. Extensive experiments on various settings, backbones (CLIP, SigLip, PE‑Core), and tasks (4 CDFSL datasets and 11 FSL datasets) show that we consistently set new state‑of‑the‑art results. Code is available at https://github.com/zhenyuZ‑HUST/CVPR26‑Mind‑the‑Discriminability‑Trap.
Authors:Mingyu Kim, Young-Heon Kim, Mijung Park
Abstract:
Safety mechanisms for diffusion and flow models have recently been developed along two distinct paths. In robot planning, control barrier functions are employed to guide generative trajectories away from obstacles at every denoising step by explicitly imposing geometric constraints. In parallel, recent data‑driven, negative guidance approaches have been shown to suppress harmful content and promote diversity in generated samples. However, they rely on heuristics without clearly stating when safety guidance is actually necessary. In this paper, we first introduce a unified probabilistic framework using a Maximum Mean Discrepancy (MMD) potential for image generation tasks that recasts both Shielded Diffusion and Safe Denoiser as instances of our energy‑based negative guidance against unsafe data samples. Furthermore, we leverage control‑barrier functions analysis to justify the existence of a critical time window in which negative guidance must be strong; outside of this window, the guidance should decay to zero to ensure safe and high‑quality generation. We evaluate our unified framework on several realistic safe generation scenarios, confirming that negative guidance should be applied in the early stages of the denoising process for successful safe generation.
Authors:Idan Sulami, Alon Itzkovitch, Michael R. Kearney, Moni Shahar, Ofir Levy
Abstract:
Microclimate models are essential for linking climate to ecological processes, yet most physically based frameworks estimate temperature independently for each spatial unit and rely on simplified representations of lateral heat exchange. As a result, the spatial scales over which surrounding environmental conditions influence local microclimates remain poorly quantified. Here, we show how remote sensing can help quantify the contribution of spatial context to microclimate temperature predictions. Building on convolutional neural network principles, we designed a task‑specific deep neural network and trained a series of models in which the spatial extent of input data was systematically varied. Drone‑derived spatial layers and meteorological data were used to predict ground temperature at a focal location, allowing direct assessment of how prediction accuracy changes with increasing spatial context. Our results show that incorporating spatially adjacent information substantially improves prediction accuracy, with diminishing returns beyond spatial extents of approximately 5‑7 m. This characteristic scale indicates that ground temperatures are influenced not only by local surface properties, but also by horizontal heat transfer and radiative interactions operating across neighboring microhabitats. The magnitude of spatial effects varied systematically with time of day, microhabitat type, and local environmental characteristics, highlighting context‑dependent spatial coupling in microclimate formation. By treating deep learning as a diagnostic tool rather than solely a predictive one, our approach provides a general and transferable method for quantifying spatial dependencies in microclimate models and informing the development of hybrid mechanistic‑data‑driven approaches that explicitly account for spatial interactions while retaining physical interpretability.
Authors:Ozge Mercanoglu Sincan, Jian He Low, Sobhan Asasi, Richard Bowden
Abstract:
Sign Language Translation (SLT) aims to automatically convert visual sign language videos into spoken language text and vice versa.
While recent years have seen rapid progress, the true sources of performance improvements often remain unclear. Do reported performance gains come from methodological novelty, or from the choice of a different backbone, training optimizations, hyperparameter tuning, or even differences in the calculation of evaluation metrics? This paper presents a comprehensive study of recent gloss‑free SLT models by re‑implementing key contributions in a unified codebase. We ensure fair comparison by standardizing preprocessing, video encoders, and training setups across all methods. Our analysis shows that many of the performance gains reported in the literature often diminish when models are evaluated under consistent conditions, suggesting that implementation details and evaluation setups play a significant role in determining results. We make the codebase publicly available here (https://github.com/ozgemercanoglu/sltbaselines) to support transparency and reproducibility in SLT research.
Authors:Yangsong Zhang, Anujith Muraleedharan, Rikhat Akizhanov, Abdul Ahad Butt, Gül Varol, Pascal Fua, Fabio Pizzati, Ivan Laptev
Abstract:
Recent progress in text‑conditioned human motion generation has been largely driven by diffusion models trained on large‑scale human motion data. Building on this progress, recent methods attempt to transfer such models for character animation and real robot control by applying a Whole‑Body Controller (WBC) that converts diffusion‑generated motions into executable trajectories. While WBC trajectories become compliant with physics, they may expose substantial deviations from original motion. To address this issue, we here propose PhysMoDPO, a Direct Preference Optimization framework. Unlike prior work that relies on hand‑crafted physics‑aware heuristics such as foot‑sliding penalties, we integrate WBC into our training pipeline and optimize diffusion model such that the output of WBC becomes compliant both with physics and original text instructions. To train PhysMoDPO we deploy physics‑based and task‑specific rewards and use them to assign preference to synthesized trajectories. Our extensive experiments on text‑to‑motion and spatial control tasks demonstrate consistent improvements of PhysMoDPO in both physical realism and task‑related metrics on simulated robots. Moreover, we demonstrate that PhysMoDPO results in significant improvements when applied to zero‑shot motion transfer in simulation and for real‑world deployment on a G1 humanoid robot.
Authors:Helen Qu, Rudy Morel, Michael McCabe, Alberto Bietti, François Lanusse, Shirley Ho, Yann LeCun
Abstract:
Machine learning approaches to spatiotemporal physical systems have primarily focused on next‑frame prediction, with the goal of learning an accurate emulator for the system's evolution in time. However, these emulators are computationally expensive to train and are subject to performance pitfalls, such as compounding errors during autoregressive rollout. In this work, we take a different perspective and look at scientific tasks further downstream of predicting the next frame, such as estimation of a system's governing physical parameters. Accuracy on these tasks offers a uniquely quantifiable glimpse into the physical relevance of the representations of these models. We evaluate the effectiveness of general‑purpose self‑supervised methods in learning physics‑grounded representations that are useful for downstream scientific tasks. Surprisingly, we find that not all methods designed for physical modeling outperform generic self‑supervised learning methods on these tasks, and methods that learn in the latent space (e.g., joint embedding predictive architectures, or JEPAs) outperform those optimizing pixel‑level prediction objectives. Code is available at https://github.com/helenqu/physical‑representation‑learning.
Authors:Ziyu Liu, Shengyuan Ding, Xinyu Fang, Xuanlang Dai, Penghui Yang, Jianze Liang, Jiaqi Wang, Kai Chen, Dahua Lin, Yuhang Zang
Abstract:
Vision‑to‑code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured representations with high visual fidelity. While recent Large Vision Language Models (LVLMs) achieve strong results via supervised fine‑tuning, reinforcement learning remains challenging due to misaligned reward signals. Existing rewards either rely on textual rules or coarse visual embedding similarity, both of which fail to capture fine‑grained visual discrepancies and are vulnerable to reward hacking. We propose Visual Equivalence Reward Model (Visual‑ERM), a multimodal generative reward model that provides fine‑grained, interpretable, and task‑agnostic feedback to evaluate vision‑to‑code quality directly in the rendered visual space. Integrated into RL, Visual‑ERM improves Qwen3‑VL‑8B‑Instruct by +8.4 on chart‑to‑code and yields consistent gains on table and SVG parsing (+2.7, +4.1 on average), and further strengthens test‑time scaling via reflection and revision. We also introduce VisualCritic‑RewardBench (VC‑RewardBench), a benchmark for judging fine‑grained image‑to‑image discrepancies on structured visual data, where Visual‑ERM at 8B decisively outperforms Qwen3‑VL‑235B‑Instruct and approaches leading closed‑source models. Our results suggest that fine‑grained visual reward supervision is both necessary and sufficient for vision‑to‑code RL, regardless of task specificity.
Authors:Ziqi Ma, Mengzhan Liufu, Georgia Gkioxari
Abstract:
Evolutions in the world, such as water pouring or ice melting, happen regardless of being observed. Video world models generate "worlds" via 2D frame observations. Can these generated "worlds" evolve regardless of observation? To probe this question, we design a benchmark to evaluate whether video world models can decouple state evolution from observation. Our benchmark, STEVO‑Bench, applies observation control to evolving processes via instructions of occluder insertion, turning off the light, or specifying camera "lookaway" trajectories. By evaluating video models with and without camera control for a diverse set of naturally‑occurring evolutions, we expose their limitations in decoupling state evolution from observation. STEVO‑Bench proposes an evaluation protocol to automatically detect and disentangle failure modes of video world models across key aspects of natural state evolution. Analysis of STEVO‑Bench results provide new insight into potential data and architecture bias of present‑day video world models. Project website: https://glab‑caltech.github.io/STEVOBench/. Blog: https://ziqi‑ma.github.io/blog/2026/outofsight/
Authors:Rohith Peddi, Saurabh, Shravan Shanmugam, Likhitha Pallapothula, Yu Xiang, Parag Singla, Vibhav Gogate
Abstract:
Spatio‑temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame‑centric: they reason only about currently visible objects, discard entities upon occlusion, and operate in 2D. To address this, we first introduce ActionGenome4D, a dataset that upgrades Action Genome videos into 4D scenes via feed‑forward 3D reconstruction, world‑frame oriented bounding boxes for every object involved in actions, and dense relationship annotations including for objects that are temporarily unobserved due to occlusion or camera motion. Building on this data, we formalize World Scene Graph Generation (WSGG), the task of constructing a world scene graph at each timestamp that encompasses all interacting objects in the scene, both observed and unobserved. We then propose three complementary methods, each exploring a different inductive bias for reasoning about unobserved objects: PWG (Persistent World Graph), which implements object permanence via a zero‑order feature buffer; MWAE (Masked World Auto‑Encoder), which reframes unobserved‑object reasoning as masked completion with cross‑view associative retrieval; and 4DST (4D Scene Transformer), which replaces the static buffer with differentiable per‑object temporal attention enriched by 3D motion and camera‑pose features. We further design and evaluate the performance of strong open‑source Vision‑Language Models on the WSGG task via a suite of Graph RAG‑based approaches, establishing baselines for unlocalized relationship prediction. WSGG thus advances video scene understanding toward world‑centric, temporally persistent, and interpretable scene reasoning.
Authors:Hui Wei, Hao Yu, Guoying Zhao
Abstract:
Face de‑identification (FDeID) aims to remove personally identifiable information from facial images while preserving task‑relevant utility attributes such as age, gender, and expression. It is critical for privacy‑preserving computer vision, yet the field suffers from fragmented implementations, inconsistent evaluation protocols, and incomparable results across studies. These challenges stem from the inherent complexity of the task: FDeID spans multiple downstream applications (e.g., age estimation, gender recognition, expression analysis) and requires evaluation across three dimensions (e.g., privacy protection, utility preservation, and visual quality), making existing codebases difficult to use and extend. To address these issues, we present FDeID‑Toolbox, a comprehensive toolbox designed for reproducible FDeID research. Our toolbox features a modular architecture comprising four core components: (1) standardized data loaders for mainstream benchmark datasets, (2) unified method implementations spanning classical approaches to SOTA generative models, (3) flexible inference pipelines, and (4) systematic evaluation protocols covering privacy, utility, and quality metrics. Through experiments, we demonstrate that FDeID‑Toolbox enables fair and reproducible comparison of diverse FDeID methods under consistent conditions.
Authors:Sidaty El Hadramy, Nazim Haouchine, Michael Wehrli, Philippe C. Cattin
Abstract:
This paper presents NOIR, a framework that reframes core medical imaging tasks as operator learning between continuous function spaces, challenging the prevailing paradigm of discrete grid‑based deep learning. Instead of operating on fixed pixel or voxel grids, NOIR embeds discrete medical signals into shared Implicit Neural Representations and learns a Neural Operator that maps between their latent modulations, enabling resolution‑independent function‑to‑function transformations. We evaluate NOIR across multiple 2D and 3D downstream tasks, including segmentation, shape completion, image‑to‑image translation, and image synthesis, on several public datasets such as Shenzhen, OASIS‑4, SkullBreak, fastMRI, as well as an in‑house clinical dataset. It achieves competitive performance at native resolution while demonstrating strong robustness to unseen discretizations, and empirically satisfies key theoretical properties of neural operators. The project page is available here: https://github.com/Sidaty1/NOIR‑io.
Authors:Guoqiang Zhao, Zhe Yang, Sheng Wu, Fei Teng, Mengfei Duan, Yuanfan Zheng, Kai Luo, Kailun Yang
Abstract:
Panoramic imagery provides holistic 360° visual coverage for perception in quadruped robots. However, existing occupancy prediction methods are mainly designed for wheeled autonomous driving and rely heavily on RGB cues, limiting their robustness in complex environments. To bridge this gap, (1) we present PanoMMOcc, the first real‑world panoramic multimodal occupancy dataset for quadruped robots, featuring four sensing modalities across diverse scenes. (2) We propose a panoramic multimodal occupancy perception framework, VoxelHound, tailored for legged mobility and spherical imaging. Specifically, we design (i) a Vertical Jitter Compensation (VJC) module to mitigate severe viewpoint perturbations caused by body pitch and roll during mobility, enabling more consistent spatial reasoning, and (ii) an effective Multimodal Information Prompt Fusion (MIPF) module that jointly leverages panoramic visual cues and auxiliary modalities to enhance volumetric occupancy prediction. (3) We establish a benchmark based on PanoMMOcc and provide detailed data analysis to enable systematic evaluation of perception methods under challenging embodied scenarios. Extensive experiments demonstrate that VoxelHound achieves state‑of‑the‑art performance on PanoMMOcc (+4.16% in mIoU). The dataset and code will be publicly released to facilitate future research on panoramic multimodal 3D perception for embodied robotic systems at https://github.com/SXDR/PanoMMOcc, along with the calibration tools released at https://github.com/losehu/CameraLiDAR‑Calib.
Authors:Seunghwan Bang, Hwanjun Song
Abstract:
The growing interest in embodied agents increases the demand for spatiotemporal video understanding, yet existing benchmarks largely emphasize extractive reasoning, where answers can be explicitly presented within spatiotemporal events. It remains unclear whether multimodal large language models can instead perform abstractive spatiotemporal reasoning, which requires integrating observations over time, combining dispersed cues, and inferring implicit spatial and contextual structure. To address this gap, we formalize abstractive spatiotemporal reasoning from videos by introducing a structured evaluation taxonomy that systematically targets its core dimensions and constructs a controllable, scenario‑driven synthetic egocentric video dataset tailored to evaluate abstractive spatiotemporal reasoning capabilities, spanning object‑, room‑, and floor‑plan‑level scenarios. Based on this framework, we present VAEX‑BENCH, a benchmark comprising five abstractive reasoning tasks together with their extractive counterparts. Our extensive experiments compare the performance of state‑of‑the‑art MLLMs under extractive and abstractive settings, exposing their limitations on abstractive tasks and providing a fine‑grained analysis of the underlying bottlenecks. The dataset will be released soon.
Authors:Yebin Yang, Di Wen, Lei Qi, Weitong Kong, Junwei Zheng, Ruiping Liu, Yufan Chen, Chengzhi Wu, Kailun Yang, Yuqian Fu, Danda Pani Paudel, Luc Van Gool, Kunyu Peng
Abstract:
Text‑guided 3D motion editing has seen success in single‑person scenarios, but its extension to multi‑person settings is less explored due to limited paired data and the complexity of inter‑person interactions. We introduce the task of multi‑person 3D motion editing, where a target motion is generated from a source and a text instruction. To support this, we propose InterEdit3D, a new dataset with manual two‑person motion change annotations, and a Text‑guided Multi‑human Motion Editing (TMME) benchmark. We present InterEdit, a synchronized classifier‑free conditional diffusion model for TMME. It introduces Semantic‑Aware Plan Token Alignment with learnable tokens to capture high‑level interaction cues and an Interaction‑Aware Frequency Token Alignment strategy using DCT and energy pooling to model periodic motion dynamics. Experiments show that InterEdit improves text‑to‑motion consistency and edit fidelity, achieving state‑of‑the‑art TMME performance. The dataset and code will be released at https://github.com/YNG916/InterEdit.
Authors:Handong Zheng, Yumeng Li, Kaile Zhang, Liang Xin, Guangwei Zhao, Hao Liu, Jiayu Chen, Jie Lou, Qi Fu, Rui Yang, Shuo Jiang, Weijian Luo, Weijie Su, Weijun Zhang, Xingyu Zhu, Yabin Li, Yiwei ma, Yu Chen, Yuqiu Ji, Zhaohui Yu, Guang Yang, Colin Zhang, Lei Zhang, Yuliang Liu, Xiang Bai
Abstract:
We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus on text recognition and leave graphical regions as cropped pixels, our method, termed dots.mocr, treats visual elements such as charts, diagrams, tables, and icons as first‑class parsing targets, enabling systems to parse documents while preserving semantic relationships across elements. It offers several advantages: (1) it reconstructs both text and graphics as structured outputs, enabling more faithful document reconstruction; (2) it supports end‑to‑end training over heterogeneous document elements, allowing models to exploit semantic relations between textual and visual components; and (3) it converts previously discarded graphics into reusable code‑level supervision, unlocking multimodal supervision embedded in existing documents. To make this paradigm practical at scale, we build a comprehensive data engine from PDFs, rendered webpages, and native SVG assets, and train a compact 3B‑parameter model through staged pretraining and supervised fine‑tuning. We evaluate dots.mocr from two perspectives: document parsing and structured graphics parsing. On document parsing benchmarks, it ranks second only to Gemini 3 Pro on our OCR Arena Elo leaderboard, surpasses existing open‑source document parsing systems, and sets a new state of the art of 83.9 on olmOCR Bench. On structured graphics parsing, our model achieves higher reconstruction quality than Gemini 3 Pro across image‑to‑SVG benchmarks, demonstrating strong performance on charts, UI layouts, scientific figures, and chemical diagrams. These results show a scalable path toward building large‑scale image‑to‑code corpora for multimodal pretraining. Code and models are publicly available at https://github.com/rednote‑hilab/dots.mocr.
Authors:Tianhao Fu, Bingxuan Yang, Juncheng Guo, Shrena Sribalan, Yucheng Chen
Abstract:
Automatic identification of screw types is important for industrial automation, robotics, and inventory management. However, publicly available datasets for screw classification are scarce, particularly for controlled single‑object scenarios commonly encountered in automated sorting systems. In this work, we introduce SortScrews, a dataset for casewise visual classification of screws. The dataset contains 560 RGB images at 512×512 resolution covering six screw types and a background class. Images are captured using a standardized acquisition setup and include mild variations in lighting and camera perspective across four capture settings.
To facilitate reproducible research and dataset expansion, we also provide a reusable data collection script that allows users to easily construct similar datasets for custom hardware components using inexpensive camera setups.
We establish baseline results using transfer learning with EfficientNet‑B0 and ResNet‑18 classifiers pretrained on ImageNet. In addition, we conduct a well‑explored failure analysis. Despite the limited dataset size, these lightweight models achieve strong classification accuracy, demonstrating that controlled acquisition conditions enable effective learning even with relatively small datasets. The dataset, collection pipeline, and baseline training code are publicly available at https://github.com/ATATC/SortScrews.
Authors:Aditya Parikh, Aasa Feragen
Abstract:
We present a fairness‑aware framework for multi‑class lung disease diagnosis from chest CT volumes, developed for the Fair Disease Diagnosis Challenge at the PHAROS‑AIF‑MIH Workshop (CVPR 2026). The challenge requires classifying CT scans into four categories ‑‑ Healthy, COVID‑19, Adenocarcinoma, and Squamous Cell Carcinoma ‑‑ with performance measured as the average of per‑gender macro F1 scores, explicitly penalizing gender‑inequitable predictions. Our approach addresses two core difficulties: the sparse pathological signal across hundreds of slices, and a severe demographic imbalance compounded across disease class and gender. We propose an attention‑based Multiple Instance Learning (MIL) model on a ConvNeXt backbone that learns to identify diagnostically relevant slices without slice‑level supervision, augmented with a Gradient Reversal Layer (GRL) that adversarially suppresses gender‑predictive structure in the learned scan representation. Training incorporates focal loss with label smoothing, stratified cross‑validation over joint (class, gender) strata, and targeted oversampling of the most underrepresented subgroup. At inference, all five‑fold checkpoints are ensembled with horizontal‑flip test‑time augmentation via soft logit voting and out‑of‑the‑fold threshold optimization for robustness. Our model achieves a mean validation competition score of 0.685 (std ‑ 0.030), with the best single fold reaching 0.759. All training and inference code is publicly available at https://github.com/ADE‑17/cvpr‑fair‑chest‑ct
Authors:Riccardo Raciti, Lemuel Puglisi, Francesco Guarnera, Daniele Ravì, Sebastiano Battiato
Abstract:
Percentage Brain Volume Change (PBVC) derived from Magnetic Resonance Imaging (MRI) is a widely used biomarker of brain atrophy, with SIENA among the most established methods for its estimation. However, SIENA relies on classical image processing steps, particularly skull stripping and tissue segmentation, whose failures can propagate through the pipeline and bias atrophy estimates. In this work, we examine whether targeted deep learning substitutions can improve SIENA while preserving its established and interpretable framework. To this end, we integrate SynthStrip and SynthSeg into SIENA and evaluate three pipeline variants on the ADNI and PPMI longitudinal cohorts. Performance is assessed using three complementary criteria: correlation with longitudinal clinical and structural decline, scan‑order consistency, and end‑to‑end runtime. Replacing the skull‑stripping module yields the most consistent gains: in ADNI, it substantially strengthens associations between PBVC and multiple measures of disease progression relative to the standard SIENA pipeline, while across both datasets it markedly improves robustness under scan reversal. The fully integrated pipeline achieves the strongest scan‑order consistency, reducing the error by up to 99.1%. In addition, GPU‑enabled variants reduce execution time by up to 46% while maintaining CPU runtimes comparable to standard SIENA. Overall, these findings show that deep learning can meaningfully strengthen established longitudinal atrophy pipelines when used to reinforce their weakest image processing steps. More broadly, this study highlights the value of modularly modernizing clinically trusted neuroimaging tools without sacrificing their interpretability. Code is publicly available at https://github.com/Raciti/Enhanced‑SIENA.git.
Authors:Zikang Liu, Longteng Guo, Handong Li, Ru Zhen, Xingjian He, Ruyi Ji, Xiaoming Ren, Yanhao Zhang, Haonan Lu, Jing Liu
Abstract:
Real‑time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers reasoning until the full video context is observed, resulting in high latency and growing computational cost that are incompatible with streaming scenarios. In this paper, we introduce ThinkStream, a framework for streaming video reasoning based on a Watch‑‑Think‑‑Speak paradigm that enables models to incrementally update their understanding as new video observations arrive. At each step, the model performs a short reasoning update and decides whether sufficient evidence has accumulated to produce a response. To support long‑horizon streaming, we propose Reasoning‑Compressed Streaming Memory (RCSM), which treats intermediate reasoning traces as compact semantic memory that replaces outdated visual tokens while preserving essential context. We further train the model using a Streaming Reinforcement Learning with Verifiable Rewards scheme that aligns incremental reasoning and response timing with the requirements of streaming interaction. Experiments on multiple streaming video benchmarks show that ThinkStream significantly outperforms existing online video models while maintaining low latency and memory usage. Code, models and data will be released at https://github.com/johncaged/ThinkStream
Authors:Shaofeng Guo, Jiequan Cui, Richang Hong
Abstract:
With the rapid rise of Artificial Intelligence Generated Content (AIGC), image manipulation has become increasingly accessible, posing significant challenges for image forgery detection and localization (IFDL). In this paper, we study how to fully leverage vision‑language models (VLMs) to assist the IFDL task. In particular, we observe that priors from VLMs hardly benefit the detection and localization performance and even have negative effects due to their inherent biases toward semantic plausibility rather than authenticity. Additionally, the location masks explicitly encode the forgery concepts, which can serve as extra priors for VLMs to ease their training optimization, thus enhancing the interpretability of detection and localization results. Building on these findings, we propose a new IFDL pipeline named IFDL‑VLM. To demonstrate the effectiveness of our method, we conduct experiments on 9 popular benchmarks and assess the model performance under both in‑domain and cross‑dataset generalization settings. The experimental results show that we consistently achieve new state‑of‑the‑art performance in detection, localization, and interpretability.Code is available at: https://github.com/sha0fengGuo/IFDL‑VLM.
Authors:Xin Xu, Weilong Li, Wei Liu, Wenke Huang, Zhixi Yu, Bin Yang, Xiaoying Liao, Kui Jiang
Abstract:
Federated Domain Generalization for Person Re‑Identification (FedDG‑ReID) learns domain‑invariant representations from decentralized data. While Vision Transformer (ViT) is widely adopted, its global attention often fails to distinguish pedestrians from high similarity backgrounds or diverse viewpoints ‑‑ a challenge amplified by cross‑client distribution shifts in FedDG‑ReID. To address this, we propose Federated Body Distribution Aware Visual Prompt (FedBPrompt), introducing learnable visual prompts to guide Transformer attention toward pedestrian‑centric regions. FedBPrompt employs a Body Distribution Aware Visual Prompts Mechanism (BAPM) comprising: Holistic Full Body Prompts to suppress cross‑client background noise, and Body Part Alignment Prompts to capture fine‑grained details robust to pose and viewpoint variations. To mitigate high communication costs, we design a Prompt‑based Fine‑Tuning Strategy (PFTS) that freezes the ViT backbone and updates only lightweight prompts, significantly reducing communication overhead while maintaining adaptability. Extensive experiments demonstrate that BAPM effectively enhances feature discrimination and cross‑domain generalization, while PFTS achieves notable performance gains within only a few aggregation rounds. Moreover, both BAPM and PFTS can be easily integrated into existing ViT‑based FedDG‑ReID frameworks, making FedBPrompt a flexible and effective solution for federated person re‑identification. The code is available at https://github.com/leavlong/FedBPrompt.
Authors:David McAllister, Miika Aittala, Tero Karras, Janne Hellsten, Angjoo Kanazawa, Timo Aila, Samuli Laine
Abstract:
Reinforcement learning (RL) has become a standard technique for post‑training diffusion‑based image synthesis models, as it enables learning from reward signals to explicitly improve desirable aspects such as image quality and prompt alignment. In this paper, we propose an online RL variant that reduces the variance in the model updates by sampling paired trajectories and pulling the flow velocity in the direction of the more favorable image. Unlike existing methods that treat each sampling step as a separate policy action, we consider the entire sampling process as a single action. We experiment with both high‑quality vision language models and off‑the‑shelf quality metrics for rewards, and evaluate the outputs using a broad set of metrics. Our method converges faster and yields higher output quality and prompt alignment than previous approaches.
Authors:Lydia A. Schönpflug, Nikki van den Berg, Sonali Andani, Nanda Horeweg, Jurriaan Barkey Wolf, Tjalling Bosse, Viktor H. Koelzer, Maxime W. Lafarge
Abstract:
Sensitivity to staining variation remains a major barrier to deploying computational pathology (CPath) models as hematoxylin and eosin (H&E) staining varies across laboratories, requiring systematic assessment of how this variability affects model prediction. In this work, we developed a three‑step protocol for evaluating robustness to H&E staining variation in CPath models. Step 1: Select reference staining conditions, Step 2: Characterize test set staining properties, Step 3: Apply CPath model(s) under simulated reference staining conditions. Here, we first created a new reference staining library based on the PLISM dataset. As an exemplary use case, we applied the protocol to assess the robustness properties of 306 microsatellite instability (MSI) classification models on the unseen SurGen colorectal cancer dataset (n=738), including 300 attention‑based multiple instance learning models trained on the TCGA‑COAD/READ datasets across three feature extractors (UNI2‑h, H‑Optimus‑1, Virchow2), alongside six public MSI classification models. Classification performance was measured as AUC, and robustness as the min‑max AUC range across four simulated staining conditions (low/high H&E intensity, low/high H&E color similarity). Across models and staining conditions, classification performance ranged from AUC 0.769‑0.911 (Δ = 0.142). Robustness ranged from 0.007‑0.079 (Δ = 0.072), and showed a weak inverse correlation with classification performance (Pearson r=‑0.22, 95% CI [‑0.34, ‑0.11]). Thus, we show that the proposed evaluation protocol enables robustness‑informed CPath model selection and provides insight into performance shifts across H&E staining conditions, supporting the identification of operational ranges for reliable model deployment. Code is available at https://github.com/CTPLab/staining‑robustness‑evaluation .
Authors:Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
Abstract:
Computer Use Agents (CUAs) translate natural‑language instructions into Graphical User Interface (GUI) actions such as clicks, keystrokes, and scrolls by relying on a Vision‑Language Model (VLM) to interpret screenshots and predict grounded tool calls. However, grounding accuracy varies dramatically across VLMs, while current CUA systems typically route every action to a single fixed model regardless of difficulty. We propose Adaptive VLM Routing (AVR), a framework that inserts a lightweight semantic routing layer between the CUA orchestrator and a pool of VLMs. For each tool call, AVR estimates action difficulty from multimodal embeddings, probes a small VLM to measure confidence, and routes the action to the cheapest model whose predicted accuracy satisfies a target reliability threshold. For warm agents with memory of prior UI interactions, retrieved context further narrows the capability gap between small and large models, allowing many actions to be handled without escalation. We formalize routing as a cost‑‑accuracy trade‑off, derive a threshold‑based policy for model selection, and evaluate AVR using ScreenSpot‑Pro grounding data together with the OpenClaw agent routing benchmark. Across these settings, AVR projects inference cost reductions of up to 78% while staying within 2 percentage points of an all‑large‑model baseline. When combined with the Visual Confused Deputy guardrail, AVR also escalates high‑risk actions directly to the strongest available model, unifying efficiency and safety within a single routing framework. Materials are also provided Model, benchmark, and code: https://github.com/vllm‑project/semantic‑router.
Authors:Sen Nie, Jie Zhang, Zhongqi Wang, Zhaoyang Wei, Shiguang Shan, Xilin Chen
Abstract:
Achieving adversarial robustness in Vision‑Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long‑standing and challenging trade‑off. In this work, we revisit this trade‑off by investigating a fundamental question: What makes VLMs robust? Through a detailed analysis of adversarially fine‑tuned models, we examine how robustness mechanisms function internally and how they interact with clean accuracy. Our analysis reveals that adversarial robustness is not uniformly distributed across network depth. Instead, unexpectedly, it is primarily localized within the shallow layers, driven by a low‑frequency spectral bias and input‑insensitive attention patterns. Meanwhile, updates to the deep layers tend to undermine both clean accuracy and robust generalization. Motivated by these insights, we propose Adversarial Robustness Adaptation (R‑Adapt), a simple yet effective framework that freezes all pre‑trained weights and introduces minimal, insight‑driven adaptations only in the initial layers. This design achieves an exceptional balance between adversarial robustness and clean accuracy. R‑Adapt further supports training‑free, model‑guided, and data‑driven paradigms, offering flexible pathways to seamlessly equip standard models with robustness. Extensive evaluations on 18 datasets and diverse tasks demonstrate our state‑of‑the‑art performance under various attacks. Notably, R‑Adapt generalizes efficiently to large vision‑language models (e.g., LLaVA and Qwen‑VL) to enhance their robustness. Our project page is available at https://summu77.github.io/R‑Adapt.
Authors:Sangmin Kim, Minhyuk Hwang, Geonho Cha, Dongyoon Wee, Jaesik Park
Abstract:
Recent advances in 3D foundation models have led to growing interest in reconstructing humans and their surrounding environments. However, most existing approaches focus on monocular inputs, and extending them to multi‑view settings requires additional overhead modules or preprocessed data. To this end, we present CHROMM, a unified framework that jointly estimates cameras, scene point clouds, and human meshes from multi‑person multi‑view videos without relying on external modules or preprocessing. We integrate strong geometric and human priors from Pi3X and Multi‑HMR into a single trainable neural network architecture, and introduce a scale adjustment module to solve the scale discrepancy between humans and the scene. We also introduce a multi‑view fusion strategy to aggregate per‑view estimates into a single representation at test‑time. Finally, we propose a geometry‑based multi‑person association method, which is more robust than appearance‑based approaches. Experiments on EMDB, RICH, EgoHumans, and EgoExo4D show that CHROMM achieves competitive performance in global human motion and multi‑view pose estimation while running over 8x faster than prior optimization‑based multi‑view approaches. Project page: https://nstar1125.github.io/chromm.
Authors:Shuchang Lyu, Haiquan Wen, Guangliang Cheng, Meng Li, Zheng Zhou, You Zhou, Dingding Yao, Zhenwei Shi
Abstract:
Recent advances in reasoning language models and reinforcement learning with verifiable rewards have significantly enhanced multi‑step reasoning capabilities. This progress motivates the extension of reasoning paradigms to remote sensing visual grounding task. However, existing remote sensing grounding methods remain largely confined to perception‑level matching and single‑entity formulations, limiting the role of explicit reasoning and inter‑entity modeling. To address this challenge, we introduce a new benchmark dataset for Multi‑Entity Reasoning Grounding in Remote Sensing (ME‑RSRG). Based on ME‑RSRG, we reformulate remote sensing grounding as a multi‑entity reasoning task and propose an Entity‑Aware Reasoning (EAR) framework built upon visual‑linguistic foundation models. EAR generates structured reasoning traces and subject‑object grounding outputs. It adopts supervised fine‑tuning for cold‑start initialization and is further optimized via entity‑aware reward‑driven Group Relative Policy Optimization (GRPO). Extensive experiments on ME‑RSRG demonstrate the challenges of multi‑entity reasoning and verify the effectiveness of our proposed EAR framework. Our dataset, code, and models will be available at https://github.com/CV‑ShuchangLyu/ME‑RSRG.
Authors:Shifeng Chen, Yihui Li, Jun Liao, Hongyu Yang, Di Huang
Abstract:
Recent advances in 3D scene editing using NeRF and 3DGS enable high‑quality static scene editing. In contrast, dynamic scene editing remains challenging, as methods that directly extend 2D diffusion models to 4D often produce motion artifacts, temporal flickering, and inconsistent style propagation. We introduce Catalyst4D, a framework that transfers high‑quality 3D edits to dynamic 4D Gaussian scenes while maintaining spatial and temporal coherence. At its core, Anchor‑based Motion Guidance (AMG) builds a set of structurally stable and spatially representative anchors from both original and edited Gaussians. These anchors serve as robust region‑level references, and their correspondences are established via optimal transport to enable consistent deformation propagation without cross‑region interference or motion drift. Complementarily, Color Uncertainty‑guided Appearance Refinement (CUAR) preserves temporal appearance consistency by estimating per‑Gaussian color uncertainty and selectively refining regions prone to occlusion‑induced artifacts. Extensive experiments demonstrate that Catalyst4D achieves temporally stable, high‑fidelity dynamic scene editing and outperforms existing methods in both visual quality and motion coherence.
Authors:Xiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu, Fanman Meng, Linfeng Xu, Hongliang Li
Abstract:
Error detection is crucial in industrial training, healthcare, and assembly quality control. Most existing work assumes a single‑view setting and cannot handle the practical case where a third‑person (exo) demonstration is used to assess a first‑person (ego) imitation. We formalize Ego\rightarrowExo Imitation Error Detection: given asynchronous, length‑mismatched ego and exo videos, the model must localize procedural steps on the ego timeline and decide whether each is erroneous. This setting introduces cross‑view domain shift, temporal misalignment, and heavy redundancy. Under a unified protocol, we adapt strong baselines from dense video captioning and temporal action detection and show that they struggle in this cross‑view regime. We then propose SAVA‑X, an Align‑Fuse‑Detect framework with (i) view‑conditioned adaptive sampling, (ii) scene‑adaptive view embeddings, and (iii) bidirectional cross‑attention fusion. On the EgoMe benchmark, SAVA‑X consistently improves AUPRC and mean tIoU over all baselines, and ablations confirm the complementary benefits of its components. Code is available at https://github.com/jack1ee/SAVAX.
Authors:Xiaoyu Li, Yuhang Liu, Xuanshuo Kang, Zheng Luo, Fangqi Lou, Xiaohua Wu, Zihan Xiong
Abstract:
In‑Context Learning (ICL) is a significant paradigm for Large Multimodal Models (LMMs), using a few in‑context demonstrations (ICDs) for new task adaptation. However, its performance is sensitive to demonstration configurations and computationally expensive. Mathematically, the influence of these demonstrations can be decomposed into a dynamic mixture of the standard attention output and the context values. Current approximation methods simplify this process by learning a "shift vector". Inspired by the exact decomposition, we introduce High‑Fidelity In‑Context Learning (HIFICL) to more faithfully model the ICL mechanism. HIFICL consists of three key components: 1) a set of "virtual key‑value pairs" to act as a learnable context, 2) a low‑rank factorization for stable and regularized training, and 3) a simple end‑to‑end training objective. From another perspective, this mechanism constitutes a form of context‑aware Parameter‑Efficient Fine‑Tuning (PEFT). Extensive experiments show that HiFICL consistently outperforms existing approximation methods on several multimodal benchmarks. The code is available at https://github.com/bbbandari/HiFICL.
Authors:Lutao Jiang, Zidong Cao, Weikai Chen, Xu Zheng, Yuanhuiyi Lyu, Zhenyang Li, Zeyu HU, Yingda Yin, Keyang Luo, Runze Zhang, Kai Yan, Shengju Qian, Haidi Fan, Yifan Peng, Xin Wang, Hui Xiong, Ying-Cong Chen
Abstract:
Promptable instance segmentation is widely adopted in embodied and AR systems, yet the performance of foundation models trained on perspective imagery often degrades on 360° panoramas. In this paper, we introduce Segment Any 4K Panorama (SAP), a foundation model for 4K high‑resolution panoramic instance‑level segmentation. We reformulate panoramic segmentation as fixed‑trajectory perspective video segmentation, decomposing a panorama into overlapping perspective patches sampled along a continuous spherical traversal. This memory‑aligned reformulation preserves native 4K resolution while restoring the smooth viewpoint transitions required for stable cross‑view propagation. To enable large‑scale supervision, we synthesize 183,440 4K‑resolution panoramic images with instance segmentation labels using the InfiniGen engine. Trained under this trajectory‑aligned paradigm, SAP generalizes effectively to real‑world 360° images, achieving +17.2 zero‑shot mIoU gain over vanilla SAM2 of different sizes on real‑world 4K panorama benchmark.
Authors:Yuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu, Yibin Lou, Jie Wu, Jing Xu, Jian Zhang, Zheng Yang, Yunlong Lin, Chenxin Li, Panwang Pan, Junbin Lu, Jingyan Jiang, Xinghao Ding, Yue Huang, Zhi Wang
Abstract:
Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While current Multimodal Large Language Models (MLLMs) excel in static visual understanding, can they also be adept at "thinking in dynamics", i.e., perceive, track and reason about spatio‑temporal dynamics in evolving scenes? To systematically assess their spatio‑temporal reasoning and localized dynamics perception capabilities, we introduce Dyn‑Bench, a large‑scale benchmark built from diverse real‑world and synthetic video datasets, enabling robust and scalable evaluation of spatio‑temporal understanding. Through multi‑stage filtering from massive 2D and 4D data sources, Dyn‑Bench provides a high‑quality collection of dynamic scenes, comprising 1k videos, 7k visual question answering (VQA) pairs, and 3k dynamic object grounding pairs. We probe general, spatial and region‑level MLLMs to express how they think in dynamics both linguistically and visually, and find that existing models cannot simultaneously maintain strong performance in both spatio‑temporal reasoning and dynamic object grounding, often producing inconsistent interpretations of motion and interaction. Notably, conventional prompting strategies (e.g., chain‑of‑thought or caption‑based hints) provide limited improvement, whereas structured integration approaches, including Mask‑Guided Fusion and Spatio‑Temporal Textual Cognitive Map (ST‑TCM), significantly enhance MLLMs' dynamics perception and spatio‑temporal reasoning in the physical 4D world. Code and benchmark are available at https://dyn‑bench.github.io/.
Authors:Chenyang Zhu, Hongxiang Li, Xiu Li, Long Chen
Abstract:
Concept customization typically binds rare tokens to a target concept. Unfortunately, these approaches often suffer from unstable performance as the pretraining data seldom contains these rare tokens. Meanwhile, these rare tokens fail to convey the inherent knowledge of the target concept. Consequently, we introduce Knowledge‑aware Concept Customization, a novel task aiming at binding diverse textual knowledge to target visual concepts. This task requires the model to identify the knowledge within the text prompt to perform high‑fidelity customized generation. Meanwhile, the model should efficiently bind all the textual knowledge to the target concept. Therefore, we propose MoKus, a novel framework for knowledge‑aware concept customization. Our framework relies on a key observation: cross‑modal knowledge transfer, where modifying knowledge within the text modality naturally transfers to the visual modality during generation. Inspired by this observation, MoKus contains two stages: (1) In visual concept learning, we first learn the anchor representation to store the visual information of the target concept. (2) In textual knowledge updating, we update the answer for the knowledge queries to the anchor representation, enabling high‑fidelity customized generation. To further comprehensively evaluate our proposed MoKus on the new task, we introduce the first benchmark for knowledge‑aware concept customization: KnowCusBench. Extensive evaluations have demonstrated that MoKus outperforms state‑of‑the‑art methods. Moreover, the cross‑model knowledge transfer allows MoKus to be easily extended to other knowledge‑aware applications like virtual concept creation and concept erasure. We also demonstrate the capability of our method to achieve improvements on world knowledge benchmarks.
Authors:Kaifan Zhang, Lihuo He, Junjie Ke, Yuqi Ji, Lukun Wu, Lizi Wang, Xinbo Gao
Abstract:
Visual stimuli reconstruction from EEG remains challenging due to fidelity loss and representation shift. We propose CognitionCapturerPro, an enhanced framework that integrates EEG with multi‑modal priors (images, text, depth, and edges) via collaborative training. Our core contributions include an uncertainty‑weighted similarity scoring mechanism to quantify modality‑specific fidelity and a fusion encoder for integrating shared representations. By employing a simplified alignment module and a pre‑trained diffusion model, our method significantly outperforms the original CognitionCapturer on the THINGS‑EEG dataset, improving Top‑1 and Top‑5 retrieval accuracy by 25.9% and 10.6%, respectively. Code is available at: https://github.com/XiaoZhangYES/CognitionCapturerPro.
Authors:Dongxu Zhang, Yingsen Wang, Yiding Sun, Haoran Xu, Peilin Fan, Jihua Zhu
Abstract:
Robust point cloud registration is a fundamental task in 3D computer vision and geometric deep learning, essential for applications such as large‑scale 3D reconstruction, augmented reality, and scene understanding. However, the performance of established learning‑based methods often degrades in complex, real world scenarios characterized by incomplete data, sensor noise, and low overlap regions. To address these limitations, we propose CMHANet, a novel Cross‑Modal Hybrid Attention Network. Our method integrates the fusion of rich contextual information from 2D images with the geometric detail of 3D point clouds, yielding a comprehensive and resilient feature representation. Furthermore, we introduce an innovative optimization function based on contrastive learning, which enforces geometric consistency and significantly improves the model's robustness to noise and partial observations. We evaluated CMHANet on the 3DMatch and the challenging 3DLoMatch datasets. \revAdditionally, zero‑shot evaluations on the TUM RGB‑D SLAM dataset verify the model's generalization capability to unseen domains. The experimental results demonstrate that our method achieves substantial improvements in both registration accuracy and overall robustness, outperforming current techniques. We also release our code in \hrefhttps://github.com/DongXu‑Zhang/CMHANethttps://github.com/DongXu‑Zhang/CMHANet.
Authors:Dongxu Zhang, Jihua Zhu, Shiqi Li, Wenbiao Yan, Haoran Xu, Peilin Fan, Huimin Lu
Abstract:
Point cloud registration (PCR) is a fundamental task in 3D vision and provides essential support for applications such as autonomous driving, robotics, and environmental modeling. Despite its widespread use, existing methods often fail when facing real‑world challenges like heavy noise, significant occlusions, and large‑scale transformations. These limitations frequently result in compromised registration accuracy and insufficient robustness in complex environments. In this paper, we propose IGASA as a novel registration framework constructed upon a Hierarchical Pyramid Architecture (HPA) designed for robust multi‑scale feature extraction and fusion. The framework integrates two pivotal components consisting of the Hierarchical Cross‑Layer Attention (HCLA) module and the Iterative Geometry‑Aware Refinement (IGAR) module. The HCLA module utilizes skip attention mechanisms to align multi‑resolution features and enhance local geometric consistency. Simultaneously, the IGAR module is designed for the fine matching phase by leveraging reliable correspondences established during coarse matching. This synergistic integration within the architecture allows IGASA to adapt effectively to diverse point cloud structures and intricate transformations. We evaluate the performance of IGASA on four widely recognized benchmark datasets including 3D(Lo)Match, KITTI, and nuScenes. Our extensive experiments consistently demonstrate that IGASA significantly surpasses state‑of‑the‑art methods and achieves notable improvements in registration accuracy. This work provides a robust foundation for advancing point cloud registration techniques while offering valuable insights for practical 3D vision applications. The code for IGASA is available in \hrefhttps://github.com/DongXu‑Zhang/IGASAhttps://github.com/DongXu‑Zhang/IGASA.
Authors:Jillur Rahman Saurav, Thuong Le Hoai Pham, Pritam Mukherjee, Paul Yi, Brent A. Orr, Jacob M. Luber
Abstract:
Virtual immunohistochemistry (IHC) staining from hematoxylin and eosin (H&E) images can accelerate diagnostics by providing preliminary molecular insight directly from routine sections, reducing the need for repeat sectioning when tissue is limited. Existing methods improve realism through contrastive objectives, prototype matching, or domain alignment, yet the generator itself receives no direct guidance from pathology foundation models. We present UNIStainNet, a SPADE‑UNet conditioned on dense spatial tokens from a frozen pathology foundation model (UNI), providing tissue‑level semantic guidance for stain translation. A misalignment‑aware loss suite preserves stain quantification accuracy, and learned stain embeddings enable a single model to serve multiple IHC markers simultaneously. On MIST, UNIStainNet achieves state‑of‑the‑art distributional metrics on all four stains (HER2, Ki67, ER, PR) from a single unified model, where prior methods typically train separate per‑stain models. On BCI, it also achieves the best distributional metrics. A tissue‑type stratified failure analysis reveals that remaining errors are systematic, concentrating in non‑tumor tissue. Code is available at https://github.com/facevoid/UNIStainNet.
Authors:Pingping Zhang, Tianyu Yan, Yuhao Wang, Yang Liu, Tongdan Tang, Yili Ma, Long Lv, Feng Tian, Weibing Sun, and Huchuan Lu
Abstract:
Marine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning‑based MAS methods struggle with the long‑distance modeling issue. Recently, Segment Anything Model (SAM) has gained popularity in general image segmentation. However, it lacks of perceiving fine‑grained details and frequency information. To this end, we propose a novel learning framework, named Hierarchical Frequency Prompted SAM (HFP‑SAM) for high‑performance MAS. First, we design a Frequency Guided Adapter (FGA) to efficiently inject marine scene information into the frozen SAM backbone through frequency domain prior masks. Additionally, we introduce a Frequency‑aware Point Selection (FPS) to generate highlighted regions through frequency analysis. These regions are combined with the coarse predictions of SAM to generate point prompts and integrate into SAM's decoder for fine predictions. Finally, to obtain comprehensive segmentation masks, we introduce a Full‑View Mamba (FVM) to efficiently extract spatial and channel contextual information with linear computational complexity. Extensive experiments on four public datasets demonstrate the superior performance of our approach. The source code is publicly available at https://github.com/Drchip61/TIP‑HFP‑SAM.
Authors:Pengyiang Liu, Zhongyue Shi, Hongye Hao, Qi Fu, Xueting Bi, Siwei Zhang, Xiaoyang Hu, Zitian Wang, Linjiang Huang, Si Liu
Abstract:
Video understanding requires models to continuously track and update world state during playback. While existing benchmarks have advanced video understanding evaluation across multiple dimensions, the observation of how models maintain world state remains insufficient. We propose VCBench, a streaming counting benchmark that repositions counting as a minimal probe for diagnosing world state maintenance capability. We decompose this capability into object counting and event counting, forming 8 fine‑grained subcategories. Object counting covers tracking currently visible objects and cumulative unique identities, while event counting covers detecting instantaneous actions and tracking complete activity cycles. VCBench contains 406 videos with frame‑by‑frame annotations of 10,071 event occurrence moments and object state change moments, generating 1,000 streaming QA pairs with 4,576 query points along timelines. By observing state maintenance trajectories through streaming multi‑point queries, we design three complementary metrics to diagnose numerical precision, trajectory consistency, and temporal awareness. Evaluation on mainstream video‑language models shows that current models still exhibit significant deficiencies in spatial‑temporal state maintenance, particularly struggling with tasks like periodic event counting. VCBench provides a diagnostic framework for measuring and improving state maintenance in video understanding systems. Our code and data are available at https://github.com/buaaplay/VCBench.
Authors:Liangzheng Sun, Mengfan He, Xingyu Shao, Binbin Li, Zhiqiang Yan, Chunyu Li, Ziyang Meng, Fei Xing
Abstract:
Infrared‑visible (IR‑VIS) feature matching plays an essential role in cross‑modality visual localization, navigation and perception. Along with the rapid development of deep learning techniques, a number of representative image matching methods have been proposed. However, crossmodal feature matching is still a challenging task due to the significant appearance difference. A significant gap for cross‑modal feature matching research lies in the absence of standardized benchmarks and metrics for evaluations. In this paper, we introduce a comprehensive cross‑modal feature matching benchmark, CM‑Bench, which encompasses 30 feature matching algorithms across diverse cross‑modal datasets. Specifically, state‑of‑the‑art traditional and deep learning‑based methods are first summarized and categorized into sparse, semidense, and dense methods. These methods are evaluated by different tasks including homography estimation, relative pose estimation, and feature‑matching‑based geo‑localization. In addition, we introduce a classification‑network‑based adaptive preprocessing front‑end that automatically selects suitable enhancement strategies before matching. We also present a novel infrared‑satellite cross‑modal dataset with manually annotated ground‑truth correspondences for practical geo‑localization evaluation. The dataset and resource will be available at: https://github.com/SLZ98/CM‑Bench.
Authors:Selim Furkan Tekin, Yichang Xu, Gaowen Liu, Ramana Rao Kompella, Margaret L. Loper, Ling Liu
Abstract:
With the growing number and diversity of Vision‑Language Models (VLMs), many works explore language‑based ensemble, collaboration, and routing techniques across multiple VLMs to improve multi‑model reasoning. In contrast, we address the diverse model selection using both vision and language modalities. We introduce focal error diversity to capture complementary reasoning across VLMs and a CKA‑based focal diversity metric (CKA‑focal) to measure disagreement in their visual embeddings. On the constructed ensemble surface from a pool of candidate VLMs, we applied a Genetic Algorithm to effectively prune out those component VLMs that do not add value to the fusion performance. We identify the best combination for each task as well as fuse the outputs of each VLMs in the model pool, and show that heterogeneous models can capture epistemic uncertainty dynamically and mitigate hallucinations. Our V3Fusion approach is capable of producing dual focal‑diversity fused predictions with high performance for vision‑language reasoning, even when there is no majority consensus or the majority of VLMs make incorrect predictions. Extensive experiments validate V3Fusion on four popular VLM benchmarks (A‑OKVQA, MMMU, MMMU‑Pro, and OCR‑VQA). The results show that V3Fusion outperforms the best‑performing VLM on MMMU by 8.09% and MMMU‑Pro by 4.87% gain in accuracy. For generative tasks, V3Fusion outperforms Intern‑VL2‑8b and Qwen2.5‑VL‑7b, the top‑2 VLM performers on both A‑OKVQA and OCR‑VQA. Our code and datasets are available at https://github.com/sftekin/v3fusion.
Authors:Ty Valencia, Burak Barlas, Varun Singhal, Ruchir Bhatia, Wei Yang
Abstract:
Multimodal recommendation is commonly framed as a feature fusion problem, where textual and visual signals are combined to better model user preference. However, the effectiveness of multimodal recommendation may depend not only on how modalities are fused, but also on whether item content is represented in a semantic space aligned with preference matching. This issue is particularly important because raw visual features often preserve appearance similarity, while user decisions are typically driven by higher‑level semantic factors such as style, material, and usage context. Motivated by this observation, we propose LVLM‑grounded Multimodal Semantic Representation for Recommendation (VLM4Rec), a lightweight framework that organizes multimodal item content through semantic alignment rather than direct feature fusion. VLM4Rec first uses a large vision‑language model to ground each item image into an explicit natural‑language description, and then encodes the grounded semantics into dense item representations for preference‑oriented retrieval. Recommendation is subsequently performed through a simple profile‑based semantic matching mechanism over historical item embeddings, yielding a practical offline‑online decomposition. Extensive experiments on multiple multimodal recommendation datasets show that VLM4Rec consistently improves performance over raw visual features and several fusion‑based alternatives, suggesting that representation quality may matter more than fusion complexity in this setting. The code is released at https://github.com/tyvalencia/enhancing‑mm‑rec‑sys.
Authors:Guodong Sun, Qihang Liang, Xingyu Pan, Moyun Liu, Yang Zhang
Abstract:
Accurate visual fault detection in freight trains remains a critical challenge for intelligent transportation system maintenance, due to complex operational environments, structurally repetitive components, and frequent occlusions or contaminations in safety‑critical regions. Conventional instance segmentation methods based on convolutional neural networks and Transformers often suffer from poor generalization and limited boundary accuracy under such conditions. To address these challenges, we propose a lightweight self‑prompted instance segmentation framework tailored for freight train fault detection. Our method leverages the Segment Anything Model by introducing a self‑prompt generation module that automatically produces task‑specific prompts, enabling effective knowledge transfer from foundation models to domain‑specific inspection tasks. In addition, we adopt a Tiny Vision Transformer backbone to reduce computational cost, making the framework suitable for real‑time deployment on edge devices in railway monitoring systems. We construct a domain‑specific dataset collected from real‑world freight inspection stations and conduct extensive evaluations. Experimental results show that our method achieves 74.6 AP^\textbox and 74.2 AP^\textmask on the dataset, outperforming existing state‑of‑the‑art methods in both accuracy and robustness while maintaining low computational overhead. This work offers a deployable and efficient vision solution for automated freight train inspection, demonstrating the potential of foundation model adaptation in industrial‑scale fault diagnosis scenarios. Project page: https://github.com/MVME‑HBUT/SAM_FTI‑FDet.git
Authors:Furui Chen, Han Wang, Yuhan Sun, Jianing You, Yixuan Lv, Zhuang Zhou, Hong Tan, Shengyang Li
Abstract:
Cross‑modal ship re‑identification (ReID) between optical and synthetic aperture radar (SAR) imagery is fundamentally challenged by the severe radiometric discrepancy between passive optical imaging and coherent active radar sensing. While existing approaches primarily rely on statistical distribution alignment or semantic matching, they often overlook a critical physical prior: ships are rigid objects whose geometric structures remain stable across sensing modalities, whereas texture appearance is highly modality‑dependent. In this work, we propose SDF‑Net, a Structure‑Aware Disentangled Feature Learning Network that systematically incorporates geometric consistency into optical‑‑SAR ship ReID. Built upon a ViT backbone, SDF‑Net introduces a structure consistency constraint that extracts scale‑invariant gradient energy statistics from intermediate layers to robustly anchor representations against radiometric variations. At the terminal stage, SDF‑Net disentangles the learned representations into modality‑invariant identity features and modality‑specific characteristics. These decoupled cues are then integrated through a parameter‑free additive residual fusion, effectively enhancing discriminative power. Extensive experiments on the HOSS‑ReID dataset demonstrate that SDF‑Net consistently outperforms existing state‑of‑the‑art methods. The code and trained models are publicly available at https://github.com/cfrfree/SDF‑Net.
Authors:Jianqiang Lin, Zhiqiang Shen, Peng Cao, Jinzhu Yang, Osmar R. Zaiane, Xiaoli Liu
Abstract:
Although diffusion models have achieved remarkable progress in multi‑modal magnetic resonance imaging (MRI) translation tasks, existing methods still tend to suffer from anatomical inconsistencies or degraded texture details when handling arbitrary missing‑modality scenarios. To address these issues, we propose a latent diffusion‑based multi‑modal MRI translation framework, termed MSG‑LDM. By leveraging the available modalities, the proposed method infers complete structural information, which preserves reliable boundary details. Specifically, we introduce a style‑‑structure disentanglement mechanism in the latent space, which explicitly separates modality‑specific style features from shared structural representations, and jointly models low‑frequency anatomical layouts and high‑frequency boundary details in a multi‑scale feature space. During the structure disentanglement stage, high‑frequency structural information is explicitly incorporated to enhance feature representations, guiding the model to focus on fine‑grained structural cues while learning modality‑invariant low‑frequency anatomical representations. Furthermore, to reduce interference from modality‑specific styles and improve the stability of structure representations, we design a style consistency loss and a structure‑aware loss. Extensive experiments on the BraTS2020 and WMH datasets demonstrate that the proposed method outperforms existing MRI synthesis approaches, particularly in reconstructing complete structures. The source code is publicly available at https://github.com/ziyi‑start/MSG‑LDM.
Authors:Xuanhua Yin, Chuanzhi Xu, Haoxian Zhou, Boyu Wei, Weidong Cai
Abstract:
Diffusion Transformers (DiTs) are a dominant backbone for high‑fidelity text‑to‑image generation due to strong scalability and alignment at high resolutions. However, quadratic self‑attention over dense spatial tokens leads to high inference latency and limits deployment. We observe that denoising is spatially non‑uniform with respect to aesthetic descriptors in the prompt. Regions associated with aesthetic tokens receive concentrated cross‑attention and show larger temporal variation, while low‑affinity regions evolve smoothly with redundant computation. Based on this insight, we propose AccelAes, a training‑free framework that accelerates DiTs through aesthetics‑aware spatio‑temporal reduction while improving perceptual aesthetics. AccelAes builds AesMask, a one‑shot aesthetic focus mask derived from prompt semantics and cross‑attention signals. When localized computation is feasible, SkipSparse reallocates computation and guidance to masked regions. We further reduce temporal redundancy using a lightweight step‑level prediction cache that periodically replaces full Transformer evaluations. Experiments on representative DiT families show consistent acceleration and improved aesthetics‑oriented quality. On Lumina‑Next, AccelAes achieves a 2.11× speedup and improves ImageReward by +11.9% over the dense baseline. Code is available at https://github.com/xuanhuayin/AccelAes.
Authors:Songsong Ouyang, Yingying Zhu
Abstract:
Cross‑view geo‑localization (CVGL) aims to estimate the geographic location of a street image by matching it with a corresponding aerial image. This is critical for autonomous navigation and mapping in complex real‑world scenarios. However, the task remains challenging due to significant viewpoint differences and the influence of confounding factors. To tackle these issues, we propose the Causal Learning and Geometric Topology (CLGT) framework, which integrates two key components: a Causal Feature Extractor (CFE) that mitigates the influence of confounding factors by leveraging causal intervention to encourage the model to focus on stable, task‑relevant semantics; and a Geometric Topology Fusion (GT Fusion) module that injects Bird's Eye View (BEV) road topology into street features to alleviate cross‑view inconsistencies caused by extreme perspective changes. Additionally, we introduce a Data‑Adaptive Pooling (DA Pooling) module to enhance the representation of semantically rich regions. Extensive experiments on CVUSA, CVACT, and their robustness‑enhanced variants (CVUSA‑C‑ALL and CVACT‑C‑ALL) demonstrate that CLGT achieves state‑of‑the‑art performance, particularly under challenging real‑world corruptions. Our codes are available at https://github.com/oyss‑szu/CLGT.
Authors:Yura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi, Jiankang Deng, Stefanos Zafeiriou
Abstract:
Understanding and answering questions based on a user's pointing gesture is essential for next‑generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of gesture‑rich data and their limited ability to infer fine‑grained pointing intent from egocentric video. To address this, we introduce EgoPointVQA, a dataset and benchmark for gesture‑grounded egocentric question answering, comprising 4000 synthetic and 400 real‑world videos across multiple deictic reasoning tasks. Built upon it, we further propose Hand Intent Tokens (HINT), which encodes tokens derived from 3D hand keypoints using an off‑the‑shelf reconstruction model and interleaves them with the model input to provide explicit spatial and temporal context for interpreting pointing intent. We show that our model outperforms others in different backbones and model sizes. In particular, HINT‑14B achieves 68.1% accuracy, on average over 6 tasks, surpassing the state‑of‑the‑art, InternVL3‑14B, by 6.6%. To further facilitate the open research, we will release the code, model, and dataset. Project page: https://yuuraa.github.io/papers/choi2026egovqa
Authors:Shivam Chaudhary, Sheethal Bhat, Andreas Maier
Abstract:
Accurate detection and localization of traumatic injuries in abdominal CT scans remains a critical challenge in emergency radiology, primarily due to severe scarcity of annotated medical data. This paper presents a label‑efficient approach combining self‑supervised pre‑training with semi‑supervised detection for 3D medical image analysis. We employ patch‑based Masked Image Modeling (MIM) to pre‑train a 3D U‑Net encoder on 1,206 CT volumes without annotations, learning robust anatomical representations. The pretrained encoder enables two downstream clinical tasks: 3D injury detection using VDETR with Vertex Relative Position Encoding, and multi‑label injury classification. For detection, semi‑supervised learning with 2,000 unlabeled volumes and consistency regularization achieves 56.57% validation mAP@0.50 and 45.30% test mAP@0.50 with only 144 labeled training samples, representing a 115% improvement over supervised‑only training. For classification, expanding to 2,244 labeled samples yields 94.07% test accuracy across seven injury categories using only a frozen encoder, demonstrating immediately transferable self‑supervised features. Our results validate that self‑supervised pre‑training combined with semi‑supervised learning effectively addresses label scarcity in medical imaging, enabling robust 3D object detection with limited annotations.
Authors:Joong Ho Kim, Nicholas Thai, Souhardya Saha Dip, Dong Lao, Keith G. Mills
Abstract:
Text‑to‑Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slots at a casino, a DM will produce different results given the same user‑defined inputs. This imposes a gambler's burden: To perform multiple generation cycles to obtain a satisfactory result. However, even though DMs use stochastic sampling to seed generation, the distribution of generated content quality highly depends on the prompt and the generative ability of a DM with respect to it.
To account for this, we propose Naïve PAINE for improving the generative quality of Diffusion Models by leveraging T2I preference benchmarks. We directly predict the numerical quality of an image from the initial noise and given prompt. Naïve PAINE then selects a handful of quality noises and forwards them to the DM for generation. Further, Naïve PAINE provides feedback on the DM generative quality given the prompt and is lightweight enough to seamlessly fit into existing DM pipelines. Experimental results demonstrate that Naïve PAINE outperforms existing approaches on several prompt corpus benchmarks.
Authors:Rujie Wu, Haozhe Zhao, Hai Ci, Yizhou Wang
Abstract:
Multimodal instruction tuning is often compute‑inefficient because training budgets are spread across large mixed image‑video pools whose utility is highly uneven. We present Goal‑Driven Data Optimization (GDO), a framework that computes six sample descriptors for each candidate and constructs optimized 1× training subsets for different goals. Under a fixed one‑epoch Qwen3‑VL‑8B‑Instruct training and evaluation recipe on 8 H20 GPUs, GDO uses far fewer training samples than the Uni‑10x baseline while converging faster and achieving higher accuracy. Relative to the fixed 512k‑sample Uni‑10x baseline, GDO reaches the Uni‑10x reference after 35.4k samples on MVBench, 26.6k on VideoMME, 27.3k on MLVU, and 34.7k on LVBench, while improving Accuracy by +1.38, +1.67, +3.08, and +0.84 percentage points, respectively. The gains are largest on MVBench and MLVU, while LVBench improves more modestly, consistent with its ultra‑long‑video setting and the mismatch between that benchmark and the short‑video/image‑dominant training pool. Across MinLoss, Diverse, Temp, and Temp+, stronger temporal emphasis yields steadily better long‑video understanding behavior. Overall, GDO provides a goal‑driven data optimization framework that enables faster convergence with fewer training samples under a fixed training protocol. Code is available at https://github.com/rujiewu/GDO.
Authors:Alexis Guichemerre, Banafsheh Karimian, Soufiane Belharbi, Natacha Gillet, Nicolas Thome, Pourya Shamsolmoali, Mohammadhadi Shateri, Luke McCaffrey, Eric Granger
Abstract:
Weakly Supervised Object Localization (WSOL) models enable joint classification and region‑of‑interest localization in histology images using only image‑class supervision. When deployed in a target domain, distributions shift remains a major cause of performance degradation, especially when applied on new organs or institutions with different staining protocols and scanner characteristics. Under stronger cross‑domain shifts, WSOL predictions can become biased toward dominant classes, producing highly skewed pseudo‑label distributions in the target domain. Source‑Free (Unsupervised) Domain Adaptation (SFDA) methods are commonly employed to address domain shift. However, because they rely on self‑training, the initial bias is reinforced over training iterations, degrading both classification and localization tasks. We identify this amplification of prediction bias as a primary obstacle to the SFDA of WSOL models in histopathology. This paper introduces \sfdadep, a method inspired by machine unlearning that formulates SFDA as an iterative process of identifying and correcting prediction bias. It periodically identifies target images from over‑predicted classes and selectively reduces the predictive confidence for uncertain (high entropy) images, while preserving confident predictions. This process reduces the drift of decision boundaries and bias toward dominant classes. A jointly optimized pixel‑level classifier further restores discriminative localization features under distribution shift. Extensive experiments on cross‑organ and center histopathology benchmarks (glas, CAMELYON‑16, CAMELYON‑17) with several WSOL models show that SFDA‑DeP consistently improves classification and localization over state‑of‑the‑art SFDA baselines. \small Code: \hrefhttps://anonymous.4open.science/r/SFDA‑DeP‑1797/anonymous.4open.science/r/SFDA‑DeP‑1797/
Authors:Mohamad Alansari, Naufal Suryanto, Divya Velayudhan, Sajid Javed, Naoufel Werghi, Muzammal Naseer
Abstract:
Multimodal large language models (MLLMs) have advanced from image‑level reasoning to pixel‑level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temporally consistent reference tracking. Existing video MLLMs often rely on a static segmentation token ([SEG]) for frame‑wise grounding, which provides semantics but lacks temporal context, causing spatial drift, identity switches, and unstable initialization when objects move or reappear. We introduce SPARROW, a pixel‑grounded video MLLM that unifies spatial accuracy and temporal stability through two key components: (i) Target‑Specific Tracked Features (TSF), which inject temporally aligned referent cues during training, and (ii) a dual‑prompt design that decodes box ([BOX]) and segmentation ([SEG]) tokens to fuse geometric priors with semantic grounding. SPARROW is supported by a curated referential video dataset of 30,646 videos and 45,231 Q&A pairs and operates end‑to‑end without external detectors via a class‑agnostic SAM2‑based proposer. Integrated into three recent open‑source video MLLMs (UniPixel, GLUS, and VideoGLaMM), SPARROW delivers consistent gains across six benchmarks, improving up to +8.9 J&F on RVOS, +5 mIoU on visual grounding, and +5.4 CLAIR on GCG. These results demonstrate that SPARROW substantially improves referential stability, spatial precision, and temporal coherence in pixel‑grounded video understanding. Project page: https://risys‑lab.github.io/SPARROW
Authors:Tianwei Xiong, Jun Hao Liew, Zilong Huang, Zhijie Lin, Jiashi Feng, Xihui Liu
Abstract:
Autoregressive (AR) video generative models rely on video tokenizers that compress pixels into discrete token sequences. The length of these token sequences is crucial for balancing reconstruction quality against downstream generation computational cost. Traditional video tokenizers apply a uniform token assignment across temporal blocks of different videos, often wasting tokens on simple, static, or repetitive segments while underserving dynamic or complex ones. To address this inefficiency, we introduce EVATok, a framework to produce Efficient Video Adaptive Tokenizers. Our framework estimates optimal token assignments for each video to achieve the best quality‑cost trade‑off, develops lightweight routers for fast prediction of these optimal assignments, and trains adaptive tokenizers that encode videos based on the assignments predicted by routers. We demonstrate that EVATok delivers substantial improvements in efficiency and overall quality for video reconstruction and downstream AR generation. Enhanced by our advanced training recipe that integrates video semantic encoders, EVATok achieves superior reconstruction and state‑of‑the‑art class‑to‑video generation on UCF‑101, with at least 24.4% savings in average token usage compared to the prior state‑of‑the‑art LARP and our fixed‑length baseline.
Authors:Haozhan Shen, Shilin Yan, Hongwei Xue, Shuaiqi Lu, Xiaojun Tang, Guannan Zhang, Tiancheng Zhao, Jianwei Yin
Abstract:
Multimodal Large Language Models (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step depends on verified visual compositional conditions (e.g., "if a permission dialog appears and the color of the interface is green, click Allow") and the process may branch or terminate early. Yet this capability remains under‑evaluated: existing benchmarks focus on shallow‑compositions or independent‑constraints rather than deeply chained compositional conditionals. In this paper, we introduce MM‑CondChain, a benchmark for visually grounded deep compositional reasoning. Each benchmark instance is organized as a multi‑layer reasoning chain, where every layer contains a non‑trivial compositional condition grounded in visual evidence and built from multiple objects, attributes, or relations. To answer correctly, an MLLM must perceive the image in detail, reason over multiple visual elements at each step, and follow the resulting execution path to the final outcome. To scalably construct such workflow‑style data, we propose an agentic synthesis pipeline: a Planner orchestrates layer‑by‑layer generation of compositional conditions, while a Verifiable Programmatic Intermediate Representation (VPIR) ensures each layer's condition is mechanically verifiable. A Composer then assembles these verified layers into complete instructions. Using this pipeline, we construct benchmarks across three visual domains: natural images, data charts, and GUI trajectories. Experiments on a range of MLLMs show that even the strongest model attains only 53.33 Path F1, with sharp drops on hard negatives and as depth or predicate complexity grows, confirming that deep compositional reasoning remains a fundamental challenge.
Authors:Yibin Yan, Jilan Xu, Shangzhe Di, Haoning Wu, Weidi Xie
Abstract:
Modern visual agents require representations that are general, causal, and physically structured to operate in real‑time streaming environments. However, current vision foundation models remain fragmented, specializing narrowly in image semantic perception, offline temporal modeling, or spatial geometry. This paper introduces OmniStream, a unified streaming visual backbone that effectively perceives, reconstructs, and acts from diverse visual inputs. By incorporating causal spatiotemporal attention and 3D rotary positional embeddings (3D‑RoPE), our model supports efficient, frame‑by‑frame online processing of video streams via a persistent KV‑cache. We pre‑train OmniStream using a synergistic multi‑task framework coupling static and temporal representation learning, streaming geometric reconstruction, and vision‑language alignment on 29 datasets. Extensive evaluations show that, even with a strictly frozen backbone, OmniStream achieves consistently competitive performance with specialized experts across image and video probing, streaming geometric reconstruction, complex video and spatial reasoning, as well as robotic manipulation (unseen at training). Rather than pursuing benchmark‑specific dominance, our work demonstrates the viability of training a single, versatile vision backbone that generalizes across semantic, spatial, and temporal reasoning, i.e., a more meaningful step toward general‑purpose visual understanding for interactive and embodied agents.
Authors:Mingxin Liu, Ziqian Fan, Zhaokai Wang, Leyao Gu, Zirun Zhu, Yiguo He, Yuchen Yang, Changyao Tian, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Qibing Ren, Zhihang Zhong, Xuanhe Zhou, Junchi Yan, Xue Yang
Abstract:
Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability under structured, domain‑specific constraints. In this work, we introduce GRADE, the first benchmark to assess discipline‑informed knowledge and reasoning in image editing. GRADE comprises 520 carefully curated samples across 10 academic domains, spanning from natural science to social science. To support rigorous evaluation, we propose a multi‑dimensional evaluation protocol that jointly assesses Discipline Reasoning, Visual Consistency, and Logical Readability. Extensive experiments on 20 state‑of‑the‑art open‑source and closed‑source models reveal substantial limitations in current models under implicit, knowledge‑intensive editing settings, leading to large performance gaps. Beyond quantitative scores, we conduct rigorous analyses and ablations to expose model shortcomings and identify the constraints within disciplinary editing. Together, GRADE pinpoints key directions for the future development of unified multimodal models, advancing the research on discipline‑informed image editing and reasoning. Our benchmark and evaluation code are publicly released.
Authors:Yiran Guan, Liang Yin, Dingkang Liang, Jianzhong Ju, Zhenbo Luo, Jian Luan, Yuliang Liu, Xiang Bai
Abstract:
Online Video Large Language Models (VideoLLMs) play a critical role in supporting responsive, real‑time interaction. Existing methods focus on streaming perception, lacking a synchronized logical reasoning stream. However, directly applying test‑time scaling methods incurs unacceptable response latency. To address this trade‑off, we propose Video Streaming Thinking (VST), a novel paradigm for streaming video understanding. It supports a thinking while watching mechanism, which activates reasoning over incoming video clips during streaming. This design improves timely comprehension and coherent cognition while preserving real‑time responsiveness by amortizing LLM reasoning latency over video playback. Furthermore, we introduce a comprehensive post‑training pipeline that integrates VST‑SFT, which structurally adapts the offline VideoLLM to causal streaming reasoning, and VST‑RL, which provides end‑to‑end improvement through self‑exploration in a multi‑turn video interaction environment. Additionally, we devise an automated training‑data synthesis pipeline that uses video knowledge graphs to generate high‑quality streaming QA pairs, with an entity‑relation grounded streaming Chain‑of‑Thought to enforce multi‑evidence reasoning and sustained attention to the video stream. Extensive evaluations show that VST‑7B performs strongly on online benchmarks, e.g. 79.5% on StreamingBench and 59.3% on OVO‑Bench. Meanwhile, VST remains competitive on offline long‑form or reasoning benchmarks. Compared with Video‑R1, VST responds 15.7 times faster and achieves +5.4% improvement on VideoHolmes, demonstrating higher efficiency and strong generalization across diverse video understanding tasks. Code, data, and models will be released at https://github.com/1ranGuan/VST.
Authors:Mateusz Pach, Jessica Bader, Quentin Bouniot, Serge Belongie, Zeynep Akata
Abstract:
Text‑to‑image generation models have advanced rapidly, yet achieving fine‑grained control over generated images remains difficult, largely due to limited understanding of how semantic information is encoded. We develop an interpretation of the color representation in the Variational Autoencoder latent space of FLUX.1 [Dev], revealing a structure reflecting Hue, Saturation, and Lightness. We verify our Latent Color Subspace (LCS) interpretation by demonstrating that it can both predict and explicitly control color, introducing a fully training‑free method in FLUX based solely on closed‑form latent‑space manipulation. Code is available at https://github.com/ExplainableML/LCS.
Authors:Fangfu Liu, Diankun Wu, Jiawei Chi, Yimo Cai, Yi-Hsin Hung, Xumin Yu, Hao Li, Han Hu, Yongming Rao, Yueqi Duan
Abstract:
Humans perceive and understand real‑world spaces through a stream of visual observations. Therefore, the ability to streamingly maintain and update spatial evidence from potentially unbounded video streams is essential for spatial intelligence. The core challenge is not simply longer context windows but how spatial information is selected, organized, and retained over time. In this paper, we propose Spatial‑TTT towards streaming visual‑based spatial intelligence with test‑time training (TTT), which adapts a subset of parameters (fast weights) to capture and organize spatial evidence over long‑horizon scene videos. Specifically, we design a hybrid architecture and adopt large‑chunk updates parallel with sliding‑window attention for efficient spatial video processing. To further promote spatial awareness, we introduce a spatial‑predictive mechanism applied to TTT layers with 3D spatiotemporal convolution, which encourages the model to capture geometric correspondence and temporal continuity across frames. Beyond architecture design, we construct a dataset with dense 3D spatial descriptions, which guides the model to update its fast weights to memorize and organize global 3D spatial signals in a structured manner. Extensive experiments demonstrate that Spatial‑TTT improves long‑horizon spatial understanding and achieves state‑of‑the‑art performance on video spatial benchmarks. Project page: https://liuff19.github.io/Spatial‑TTT.
Authors:Baifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye, David Eigen, Aaron Reite, Boyi Li, Jan Kautz, Song Han, David M. Chan, Pavlo Molchanov, Trevor Darrell, Hongxu Yin
Abstract:
Multi‑modal large language models (MLLMs) have advanced general‑purpose video understanding but struggle with long, high‑resolution videos ‑‑ they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweight module that removes redundant patches before processed by a ViT or an MLLM. Trained with next‑token prediction and reinforcement learning, AutoGaze autoregressively selects a minimal set of multi‑scale patches that can reconstruct the video within a user‑specified error threshold, eliminating redundancy while preserving information. Empirically, AutoGaze reduces visual tokens by 4x‑100x and accelerates ViTs and MLLMs by up to 19x, enabling scaling MLLMs to 1K‑frame 4K‑resolution videos and achieving superior results on video benchmarks (e.g., 67.0% on VideoMME). Furthermore, we introduce HLVid: the first high‑resolution, long‑form video QA benchmark with 5‑minute 4K‑resolution videos, where an MLLM scaled with AutoGaze improves over the baseline by 10.1% and outperforms the previous best MLLM by 4.5%. Project page: https://autogaze.github.io/.
Authors:Xuanlang Dai, Yujie Zhou, Long Xing, Jiazi Bu, Xilin Wei, Yuhong Liu, Beichen Zhang, Kai Chen, Yuhang Zang
Abstract:
Recently, Multimodal Large Language Models (MLLMs) have been widely integrated into diffusion frameworks primarily as text encoders to tackle complex tasks such as spatial reasoning. However, this paradigm suffers from two critical limitations: (i) MLLMs text encoder exhibits insufficient reasoning depth. Single‑step encoding fails to activate the Chain‑of‑Thought process, which is essential for MLLMs to provide accurate guidance for complex tasks. (ii) The guidance remains invariant during the decoding process. Invariant guidance during decoding prevents DiT from progressively decomposing complex instructions into actionable denoising steps, even with correct MLLM encodings. To this end, we propose Endogenous Chain‑of‑Thought (EndoCoT), a novel framework that first activates MLLMs' reasoning potential by iteratively refining latent thought states through an iterative thought guidance module, and then bridges these states to the DiT's denoising process. Second, a terminal thought grounding module is applied to ensure the reasoning trajectory remains grounded in textual supervision by aligning the final state with ground‑truth answers. With these two components, the MLLM text encoder delivers meticulously reasoned guidance, enabling the DiT to execute it progressively and ultimately solve complex tasks in a step‑by‑step manner. Extensive evaluations across diverse benchmarks (e.g., Maze, TSP, VSP, and Sudoku) achieve an average accuracy of 92.1%, outperforming the strongest baseline by 8.3 percentage points. The code and dataset are publicly available at https://lennoxdai.github.io/EndoCoT‑Webpage/.
Authors:Moayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Dogyun Park, Anil Kag, Michael Vasilkovsky, Sergey Tulyakov, Vicente Ordonez, Aliaksandr Siarohin
Abstract:
Diffusion transformers (DiTs) achieve high generative quality but lock FLOPs to image resolution, limiting principled latency‑quality trade‑offs, and allocate computation uniformly across input spatial tokens, wasting resource allocation to unimportant regions. We introduce Elastic Latent Interface Transformer (ELIT), a drop‑in, DiT‑compatible mechanism that decouples input image size from compute. Our approach inserts a latent interface, a learnable variable‑length token sequence on which standard transformer blocks can operate. Lightweight Read and Write cross‑attention layers move information between spatial tokens and latents and prioritize important input regions. By training with random dropping of tail latents, ELIT learns to produce importance‑ordered representations with earlier latents capturing global structure while later ones contain information to refine details. At inference, the number of latents can be dynamically adjusted to match compute constraints. ELIT is deliberately minimal, adding two cross‑attention layers while leaving the rectified flow objective and the DiT stack unchanged. Across datasets and architectures (DiT, U‑ViT, HDiT, MM‑DiT), ELIT delivers consistent gains. On ImageNet‑1K 512px, ELIT delivers an average gain of 35.3% and 39.6% in FID and FDD scores. Project page: https://snap‑research.github.io/elit/
Authors:Jiacheng Liu, Shengkun Tang, Jiacheng Cui, Dongkuan Xu, Zhiqiang Shen
Abstract:
Acceleration methods for diffusion models (e.g., token merging or downsampling) typically optimize synthesis quality under reduced compute, yet often ignore discriminative capacity. We revisit token compression with a joint objective and present BiGain, a training‑free, plug‑and‑play framework that preserves generation quality while improving classification in accelerated diffusion models. Our key insight is frequency separation: mapping feature‑space signals into a frequency‑aware representation disentangles fine detail from global semantics, enabling compression that respects both generative fidelity and discriminative utility. BiGain reflects this principle with two frequency‑aware operators: (1) Laplacian‑gated token merging, which encourages merges among spectrally smooth tokens while discouraging merges of high‑contrast tokens, thereby retaining edges and textures; and (2) Interpolate‑Extrapolate KV Downsampling, which downsamples keys/values via a controllable interextrapolation between nearest and average pooling while keeping queries intact, thereby conserving attention precision. Across DiT‑ and U‑Net‑based backbones and ImageNet‑1K, ImageNet‑100, Oxford‑IIIT Pets, and COCO‑2017, our operators consistently improve the speed‑accuracy trade‑off for diffusion‑based classification, while maintaining or enhancing generation quality under comparable acceleration. For instance, on ImageNet‑1K, with 70% token merging on Stable Diffusion 2.0, BiGain increases classification accuracy by 7.15% while improving FID by 0.34 (1.85%). Our analyses indicate that balanced spectral retention, preserving high‑frequency detail and low/mid‑frequency semantics, is a reliable design rule for token compression in diffusion models. To our knowledge, BiGain is the first framework to jointly study and advance both generation and classification under accelerated diffusion, supporting lower‑cost deployment.
Authors:Jun Luo, Jiaxiang Tang, Ruijie Lu, Gang Zeng
Abstract:
Text‑to‑3D scene generation from natural language is highly desirable for digital content creation. However, existing methods are largely domain‑restricted or reliant on predefined spatial relationships, limiting their capacity for unconstrained, open‑vocabulary 3D scene synthesis. In this paper, we introduce SceneAssistant, a visual‑feedback‑driven agent designed for open‑vocabulary 3D scene generation. Our framework leverages modern 3D object generation model along with the spatial reasoning and planning capabilities of Vision‑Language Models (VLMs). To enable open‑vocabulary scene composition, we provide the VLMs with a comprehensive set of atomic operations (e.g., Scale, Rotate, FocusOn). At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achieve more coherent spatial arrangements and better alignment with the input text. Experimental results demonstrate that our method can generate diverse, open‑vocabulary, and high‑quality 3D scenes. Both qualitative analysis and quantitative human evaluations demonstrate the superiority of our approach over existing methods. Furthermore, our method allows users to instruct the agent to edit existing scenes based on natural language commands. Our code is available at https://github.com/ROUJINN/SceneAssistant
Authors:Görkay Aydemir, Fatma Güney, Weidi Xie
Abstract:
Models for long‑term point tracking are typically trained on large synthetic datasets. The performance of these models degrades in real‑world videos due to different characteristics and the absence of dense ground‑truth annotations. Self‑training on unlabeled videos has been explored as a practical solution, but the quality of pseudo‑labels strongly depends on the reliability of teacher models, which vary across frames and scenes. In this paper, we address the problem of real‑world fine‑tuning and introduce verifier, a meta‑model that learns to assess the reliability of tracker predictions and guide pseudo‑label generation. Given candidate trajectories from multiple pretrained trackers, the verifier evaluates them per frame and selects the most trustworthy predictions, resulting in high‑quality pseudo‑label trajectories. When applied for fine‑tuning, verifier‑guided pseudo‑labeling substantially improves the quality of supervision and enables data‑efficient adaptation to unlabeled videos. Extensive experiments on four real‑world benchmarks demonstrate that our approach achieves state‑of‑the‑art results while requiring less data than prior self‑training methods. Project page: https://kuis‑ai.github.io/track_on_r
Authors:Mengzhen Liu, Enshen Zhou, Cheng Chi, Yi Han, Shanyu Rong, Liming Chen, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
Abstract:
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic‑driven active perception with robust, viewpoint‑invariant execution. We propose SaPaVe, an end‑to‑end framework that jointly learns these capabilities in a data‑efficient manner. Our approach decouples camera and manipulation actions rather than placing them in a shared action space, and follows a bottom‑up training strategy: we first train semantic camera control on a large‑scale dataset, then jointly optimize both action types using hybrid data. To support this framework, we introduce ActiveViewPose‑200K, a dataset of 200k image‑language‑camera movement pairs for semantic camera movement learning, and a 3D geometry‑aware module that improves execution robustness under dynamic viewpoints. We also present ActiveManip‑Bench, the first benchmark for evaluating active manipulation beyond fixed‑view settings. Extensive experiments in both simulation and real‑world environments show that SaPaVe outperforms recent vision‑language‑action models such as GR00T N1 and \(π_0\), achieving up to 31.25% higher success rates in real‑world tasks. These results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Project page: https://lmzpai.github.io/SaPaVe
Authors:Zexuan Yan, Jiarui Jin, Yue Ma, Shijian Wang, Jiahui Hu, Wenxiang Jiao, Yuan Lu, Linfeng Zhang
Abstract:
Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable challenge. This difficulty primarily stems from the limited instruction‑following capabilities of current models when encountering out‑of‑distribution prompts. To address this, we introduce GlyphBanana, alongside a corresponding benchmark specifically designed for rendering complex characters and formulas. GlyphBanana employs an agentic workflow that integrates auxiliary tools to inject glyph templates into both the latent space and attention maps, facilitating the iterative refinement of generated images. Notably, our training‑free approach can be seamlessly applied to various Text‑to‑Image (T2I) models, achieving superior precision compared to existing baselines. Extensive experiments demonstrate the effectiveness of our proposed workflow. Associated code is publicly available at https://github.com/yuriYanZeXuan/GlyphBanana.
Authors:Mengfei Duan, Hao Shi, Fei Teng, Guoqiang Zhao, Yuheng Zhang, Zhiyong Li, Kailun Yang
Abstract:
Understanding and reconstructing the 3D world through omnidirectional perception is an inevitable trend in the development of autonomous agents and embodied intelligence. However, existing 3D occupancy prediction methods are constrained by limited perspective inputs and predefined training distribution, making them difficult to apply to embodied agents that require comprehensive and safe perception of scenes in open world exploration. To address this, we present O3N, the first purely visual, end‑to‑end Omnidirectional Open‑vocabulary Occupancy predictioN framework. O3N embeds omnidirectional voxels in a polar‑spiral topology via the Polar‑spiral Mamba (PsM) module, enabling continuous spatial representation and long‑range context modeling across 360°. The Occupancy Cost Aggregation (OCA) module introduces a principled mechanism for unifying geometric and semantic supervision within the voxel space, ensuring consistency between the reconstructed geometry and the underlying semantic structure. Moreover, Natural Modality Alignment (NMA) establishes a gradient‑free alignment pathway that harmonizes visual features, voxel embeddings, and text semantics, forming a consistent "pixel‑voxel‑text" representation triad. Extensive experiments on multiple models demonstrate that our method not only achieves state‑of‑the‑art performance on QuadOcc and Human360Occ benchmarks but also exhibits remarkable cross‑scene generalization and semantic scalability, paving the way toward universal 3D world modeling. The source code will be made publicly available at https://github.com/MengfeiD/O3N.
Authors:Xiaolong Qian, Qi Jiang, Yao Gao, Lei Sun, Zhonghua Yi, Kailun Yang, Luc Van Gool, Kaiwei Wang
Abstract:
Prevalent Computational Aberration Correction (CAC) methods are typically tailored to specific optical systems, leading to poor generalization and labor‑intensive re‑training for new lenses. Developing CAC paradigms capable of generalizing across diverse photographic lenses offers a promising solution to these challenges. However, efforts to achieve such cross‑lens universality within consumer photography are still in their early stages due to the lack of a comprehensive benchmark that encompasses a sufficiently wide range of optical aberrations. Furthermore, it remains unclear which specific factors influence existing CAC methods and how these factors affect their performance. In this paper, we present comprehensive experiments and evaluations involving 24 image restoration and CAC algorithms, utilizing our newly proposed UniCAC, a large‑scale benchmark for photographic cameras constructed via automatic optical design. The Optical Degradation Evaluator (ODE) is introduced as a novel framework to objectively assess the difficulty of CAC tasks, offering credible quantification of optical aberrations and enabling reliable evaluation. Drawing on our comparative analysis, we identify three key factors ‑‑ prior utilization, network architecture, and training strategy ‑‑ that most significantly influence CAC performance, and further investigate their respective effects. We believe that our benchmark, dataset, and observations contribute foundational insights to related areas and lay the groundwork for future investigations. Benchmarks, codes, and Zemax files will be available at https://github.com/XiaolongQian/UniCAC.
Authors:Zhaoyang Jiang, Zhizhong Fu, David McAllister, Yunsoo Kim, Honghan Wu
Abstract:
Longitudinal brain MRI is essential for characterizing the progression of neurological diseases such as Alzheimer's disease assessment. However, current deep‑learning tools fragment this process: classifiers reduce a scan to a label, volumetric pipelines produce uninterpreted measurements, and vision‑language models (VLMs) may generate fluent but potentially hallucinated conclusions. We present LoV3D, a pipeline for training 3D vision‑language models, which reads longitudinal T1‑weighted brain MRI, produces a region‑level anatomical assessment, conducts longitudinal comparison with the prior scan, and finally outputs a three‑class diagnosis (Cognitively Normal, Mild Cognitive Impairment, or Dementia) along with a synthesized diagnostic summary. The stepped pipeline grounds the final diagnosis by enforcing label consistency, longitudinal coherence, and biological plausibility, thereby reducing the risks of hallucinations. The training process introduces a clinically‑weighted Verifier that scores candidate outputs automatically against normative references derived from standardized volume metrics, driving Direct Preference Optimization without a single human annotation. On a subject‑level held‑out ADNI test set (479 scans, 258 subjects), LoV3D achieves 93.7% three‑class diagnostic accuracy (+34.8% over the no‑grounding baseline), 97.2% on two‑class diagnosis accuracy (+4% over the SOTA) and 82.6% region‑level anatomical classification accuracy (+33.1% over VLM baselines). Zero‑shot transfer yields 95.4% on MIRIAD (100% Dementia recall) and 82.9% three‑class accuracy on AIBL, confirming high generalizability across sites, scanners, and populations. Code is available at https://github.com/Anonymous‑TEVC/LoV‑3D.
Authors:Umberto Cappellazzo, Stavros Petridis, Maja Pantic
Abstract:
Audio‑Visual Speech Recognition (AVSR) leverages both acoustic and visual information for robust recognition under noise. However, how models balance these modalities remains unclear. We present Dr. SHAP‑AV, a framework using Shapley values to analyze modality contributions in AVSR. Through experiments on six models across two benchmarks and varying SNR levels, we introduce three analyses: Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input‑output correspondence. Our findings reveal that models shift toward visual reliance under noise yet maintain high audio contributions even under severe degradation. Modality balance evolves during generation, temporal alignment holds under noise, and SNR is the dominant factor driving modality weighting. These findings expose a persistent audio bias, motivating ad‑hoc modality‑weighting mechanisms and Shapley‑based attribution as a standard AVSR diagnostic.
Authors:Dichang Zhang, Yixuan Shao, Simon Birrer, Dimitris Samaras
Abstract:
The upcoming decade of observational cosmology will be shaped by large sky surveys, such as the ground‑based LSST at the Vera C. Rubin Observatory and the space‑based Euclid mission. While they promise an unprecedented view of the Universe across depth, resolution, and wavelength, their differences in observational modality, sky coverage, point‑spread function, and scanning cadence make joint analysis beneficial, but also challenging. To facilitate joint analysis, we introduce A(stronomical)S(urvey)‑Bridge, a bidirectional generative model that translates between ground‑ and space‑based observations. AS‑Bridge learns a diffusion model that employs a stochastic Brownian Bridge process between the LSST and Euclid observations. The two surveys have overlapping sky regions, where we can explicitly model the conditional probabilistic distribution between them. We show that this formulation enables new scientific capabilities beyond single‑survey analysis, including faithful probabilistic predictions of missing survey observations and inter‑survey detection of rare events. These results establish the feasibility of inter‑survey generative modeling. AS‑Bridge is therefore well‑positioned to serve as a complementary component of future LSST‑Euclid joint data pipelines, enhancing the scientific return once data from both surveys become available. Data and code are available at \hrefhttps://github.com/ZHANG7DC/AS‑Bridgehttps://github.com/ZHANG7DC/AS‑Bridge.
Authors:InSpatio Team, Donghui Shen, Guofeng Zhang, Haomin Liu, Haoyu Ji, Jialin Liu, Jing Guo, Nan Wang, Siji Pan, Weihong Pan, Weijian Xie, Xiaojun Xiang, Xiaoyu Zhang, Xianbin Liu, Yifu Wang, Yipeng Chen, Zhewen Le, Zhichao Ye, Ziqiang Zhao
Abstract:
We present InSpatio‑WorldFM, an open‑source real‑time frame model for spatial intelligence. Unlike video‑based world models that rely on sequential frame generation and incur substantial latency due to window‑level processing, InSpatio‑WorldFM adopts a frame‑based paradigm that generates each frame independently, enabling low‑latency real‑time spatial inference. By enforcing multi‑view spatial consistency through explicit 3D anchors and implicit spatial memory, the model preserves global scene geometry while maintaining fine‑grained visual details across viewpoint changes. We further introduce a progressive three‑stage training pipeline that transforms a pretrained image diffusion model into a controllable frame model and finally into a real‑time generator through few‑step distillation. Experimental results show that InSpatio‑WorldFM achieves strong multi‑view consistency while supporting interactive exploration on consumer‑grade GPUs, providing an efficient alternative to traditional video‑based world models for real‑time world simulation.
Authors:Lu Wang, Zhuoran Jin, Yupu Hao, Yubo Chen, Kang Liu, Yulong Ao, Jun Zhao
Abstract:
Multimodal large language models (MLLMs) have shown strong performance on offline video understanding, but most are limited to offline inference or have weak online reasoning, making multi‑turn interaction over continuously arriving video streams difficult. Existing streaming methods typically use an interleaved perception‑generation paradigm, which prevents concurrent perception and generation and leads to early memory decay as streams grow, hurting long‑range dependency modeling. We propose Think While Watching, a memory‑anchored streaming video reasoning framework that preserves continuous segment‑level memory during multi‑turn interaction. We build a three‑stage, multi‑round chain‑of‑thought dataset and adopt a stage‑matched training strategy, while enforcing strict causality through a segment‑level streaming causal mask and streaming positional encoding. During inference, we introduce an efficient pipeline that overlaps watching and thinking and adaptively selects the best attention backend. Under both single‑round and multi‑round streaming input protocols, our method achieves strong results. Built on Qwen3‑VL, it improves single‑round accuracy by 2.6% on StreamingBench and by 3.79% on OVO‑Bench. In the multi‑round setting, it maintains performance while reducing output tokens by 56%. Code is available at: https://github.com/wl666hhh/Think_While_Watching/
Authors:Jiahao Li, Qingwang Zhang, Qiuyu Chen, Guozhan Qiu, Yunzhong Lou, Xiangdong Zhou
Abstract:
The field of Computer‑Aided Design (CAD) generation has made significant progress in recent years. Existing methods typically fall into two separate categories: parametric CAD modeling and direct boundary representation (B‑Rep) synthesis. In modern feature‑based CAD systems, parametric modeling and B‑Rep are inherently intertwined, as advanced parametric operations (e.g., fillet and chamfer) require explicit selection of B‑Rep geometric primitives, and the B‑Rep itself is derived from parametric operations. Consequently, this paradigm gap remains a critical factor limiting AI‑driven CAD modeling for complex industrial product design. This paper presents FutureCAD, a novel text‑to‑CAD framework that leverages large language models (LLMs) and a B‑Rep grounding transformer (BRepGround) for high‑fidelity CAD generation. Our method generates executable CadQuery scripts, and introduces a text‑based query mechanism that enables the LLM to specify geometric selections via natural language, which BRepGround then grounds to the target primitives. To train our framework, we construct a new dataset comprising real‑world CAD models. For the LLM, we apply supervised fine‑tuning (SFT) to establish fundamental CAD generation capabilities, followed by reinforcement learning (RL) to improve generalization. Experiments show that FutureCAD achieves state‑of‑the‑art CAD generation performance. Code and dataset are available at: https://github.com/JohanStackk/FutureCAD
Authors:Yue Shi, Rui Shi, Yuxuan Xiong, Bingbing Ni, Wenjun Zhang
Abstract:
Existing 3D editing methods often produce unrealistic and unrefined results due to the deeply integrated nature of their reconstruction networks. To address the challenge, this paper introduces CEI‑3D, an editing‑oriented reconstruction pipeline designed to facilitate realistic and fine‑grained editing. Specifically, we propose a collaborative explicit‑implicit reconstruction approach, which represents the target object using an implicit SDF network and a differentially sampled, locally controllable set of handler points. The implicit network provides a smooth and continuous geometry prior, while the explicit handler points offer localized control, enabling mutual guidance between the global 3D structure and user‑specified local editing regions. To independently control each attribute of the handler points, we design a physical properties disentangling module to decouple the color of the handler points into separate physical properties. We also propose a dual‑diffuse‑albedo network in this module to process the edited and non‑edited regions through separate branches, thereby preventing undesired interference from editing operations. Building on the reconstructed collaborative explicit‑implicit representation with disentangled properties, we introduce a spatial‑aware editing module that enables part‑wise adjustment of relevant handler points. This module employs a cross‑view propagation‑based 3D segmentation strategy, which helps users to edit the specified physical attributes of a target part efficiently. Extensive experiments on both real and synthetic datasets demonstrate that our approach achieves more realistic and fine‑grained editing results than the state‑of‑the‑art (SOTA) methods while requiring less editing time. Our code is available on https://github.com/shiyue001/CEI‑3D.
Authors:Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Abstract:
Text‑to‑video (T2V) generation models have made rapid progress in producing visually high‑quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text‑video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object's state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action‑object interactions into regular, novel, and compositional scenarios to probe both in‑distribution performance and generalization. We evaluate six representative open‑source and proprietary T2V models using both human user study and multimodal large language model (MLLM)‑based automatic evaluation. Our results show that, despite strong performance on semantic and scene alignment, current T2V models consistently struggle with accurate and temporally consistent object state changes, especially in novel and compositional settings. These findings position OSC as a key bottleneck in text‑to‑video generation and establish OSCBench as a diagnostic benchmark for advancing state‑aware video generation models.
Authors:Meilu Zhu, Zhiwei Wang, Axiu Mao, Yuxing Li, Xiaohan Xing, Yixuan Yuan, Edmund Y. Lam
Abstract:
Federated learning (FL) offers a privacy‑preserving paradigm for collaborative medical image analysis without sharing raw data. However, the absence of standardized benchmarks for medical image segmentation hinders fair and comprehensive evaluation of FL methods. To address this gap, we introduce FL‑MedSegBench, the first comprehensive benchmark for federated learning on medical image segmentation. Our benchmark encompasses nine segmentation tasks across ten imaging modalities, covering both 2D and 3D formats with realistic clinical heterogeneity. We systematically evaluate eight generic FL (gFL) and five personalized FL (pFL) methods across multiple dimensions: segmentation accuracy, fairness, communication efficiency, convergence behavior, and generalization to unseen domains. Extensive experiments reveal several key insights: (i) pFL methods, particularly those with client‑specific batch normalization (e.g., FedBN), consistently outperform generic approaches; (ii) No single method universally dominates, with performance being dataset‑dependent; (iii) Communication frequency analysis shows normalization‑based personalization methods exhibit remarkable robustness to reduced communication frequency; (iv) Fairness evaluation identifies methods like Ditto and FedRDN that protect underperforming clients; (v) A method's generalization to unseen domains is strongly tied to its ability to perform well across participating clients. We will release an open‑source toolkit to foster reproducible research and accelerate clinically applicable FL solutions, providing empirically grounded guidelines for real‑world clinical deployment. The source code is available at https://github.com/meiluzhu/FL‑MedSegBench.
Authors:Yaofeng Su, Yuming Li, Zeyue Xue, Jie Huang, Siming Fu, Haoran Li, Ying Li, Zezhong Qian, Haoyang Huang, Nan Duan
Abstract:
Recent joint audio‑visual diffusion models achieve remarkable generation quality but suffer from high latency due to their bidirectional attention dependencies, hindering real‑time applications. We propose OmniForcing, the first framework to distill an offline, dual‑stream bidirectional diffusion model into a high‑fidelity streaming autoregressive generator. However, naively applying causal distillation to such dual‑stream architectures triggers severe training instability, due to the extreme temporal asymmetry between modalities and the resulting token sparsity. We address the inherent information density gap by introducing an Asymmetric Block‑Causal Alignment with a zero‑truncation Global Prefix that prevents multi‑modal synchronization drift. The gradient explosion caused by extreme audio token sparsity during the causal shift is further resolved through an Audio Sink Token mechanism equipped with an Identity RoPE constraint. Finally, a Joint Self‑Forcing Distillation paradigm enables the model to dynamically self‑correct cumulative cross‑modal errors from exposure bias during long rollouts. Empowered by a modality‑independent rolling KV‑cache inference scheme, OmniForcing achieves state‑of‑the‑art streaming generation at ~25 FPS on a single GPU, maintaining multi‑modal synchronization and visual quality on par with the bidirectional teacher.Project Page: \hrefhttps://omniforcing.comhttps://omniforcing.com
Authors:Baicheng Li, Dong Wu, Jun Li, Shunkai Zhou, Zecui Zeng, Lusong Li, Hongbin Zha
Abstract:
Recent unified 3D generation models have made remarkable progress in producing high‑quality 3D assets from a single image. Notably, layout‑aware approaches such as SAM3D can reconstruct multiple objects while preserving their spatial arrangement, opening the door to practical scene‑level 3D generation. However, current methods are limited to single‑view input and cannot leverage complementary multi‑view observations, while independently estimated object poses often lead to physically implausible layouts such as interpenetration and floating artifacts.
We present MV‑SAM3D, a training‑free framework that extends layout‑aware 3D generation with multi‑view consistency and physical plausibility. We formulate multi‑view fusion as a Multi‑Diffusion process in 3D latent space and propose two adaptive weighting strategies ‑‑ attention‑entropy weighting and visibility weighting ‑‑ that enable confidence‑aware fusion, ensuring each viewpoint contributes according to its local observation reliability. For multi‑object composition, we introduce physics‑aware optimization that injects collision and contact constraints both during and after generation, yielding physically plausible object arrangements. Experiments on standard benchmarks and real‑world multi‑object scenes demonstrate significant improvements in reconstruction fidelity and layout plausibility, all without any additional training. Code is available at https://github.com/devinli123/MV‑SAM3D.
Authors:Tong Zhao, Mingkun Lei, Liangyu Yuan, Yanming Yang, Chenxi Song, Yang Wang, Beier Zhu, Chi Zhang
Abstract:
Diffusion Models (DMs) have achieved state‑of‑the‑art generative performance across multiple modalities, yet their sampling process remains prohibitively slow due to the need for hundreds of function evaluations. Recent progress in multi‑step ODE solvers has greatly improved efficiency by reusing historical gradients, but existing methods rely on handcrafted coefficients that fail to adapt to the non‑stationary dynamics of diffusion sampling. To address this limitation, we propose Dynamic Gradient Weighting (DyWeight), a lightweight, learning‑based multi‑step solver that introduces a streamlined implicit coupling paradigm. By relaxing classical numerical constraints, DyWeight learns unconstrained time‑varying parameters that adaptively aggregate historical gradients while intrinsically scaling the effective step size. This implicit time calibration accurately aligns the solver's numerical trajectory with the model's internal denoising dynamics under large integration steps, avoiding complex decoupled parameterizations and optimizations. Extensive experiments on CIFAR‑10, FFHQ, AFHQv2, ImageNet64, LSUN‑Bedroom, Stable Diffusion and FLUX.1‑dev demonstrate that DyWeight achieves superior visual fidelity and stability with significantly fewer function evaluations, establishing a new state‑of‑the‑art among efficient diffusion solvers. Code is available at https://github.com/Westlake‑AGI‑Lab/DyWeight
Authors:Lijun Guo, Haoyu Zhao, Xingyue Zhao, Rong Fu, Linghao Zhuang, Siteng Huang, Zhongyu Li, Hua Zou
Abstract:
Building high‑fidelity digital twins of articulated objects from visual data remains a central challenge. Existing approaches depend on multi‑view captures of the object in discrete, static states, which severely constrains their real‑world scalability. In this paper, we introduce Articulat3D, a novel framework that constructs such digital twins from casually captured monocular videos by jointly enforcing explicit 3D geometric and motion constraints. We first propose Motion Prior‑Driven Initialization, which leverages 3D point tracks to exploit the low‑dimensional structure of articulated motion. By modeling scene dynamics with a compact set of motion bases, we facilitate soft decomposition of the scene into multiple rigidly‑moving groups. Building on this initialization, we introduce Geometric and Motion Constraints Refinement, which enforces physically plausible articulation through learnable kinematic primitives parameterized by a joint axis, a pivot point, and per‑frame motion scalars, yielding reconstructions that are both geometrically accurate and temporally coherent. Extensive experiments demonstrate that Articulat3D achieves state‑of‑the‑art performance on synthetic benchmarks and real‑world casually captured monocular videos, significantly advancing the feasibility of digital twin creation under uncontrolled real‑world conditions. Our project page is at https://maxwell‑zhao.github.io/Articulat3D.
Authors:Junkun Jiang, Ho Yin Au, Jingyu Xiang, Jie Chen
Abstract:
Human motion is highly expressive and naturally aligned with language, yet prevailing methods relying heavily on joint text‑motion embeddings struggle to synthesize temporally accurate, detailed motions and often lack explainability. To address these limitations, we introduce LabanLite, a motion representation developed by adapting and extending the Labanotation system. Unlike black‑box text‑motion embeddings, LabanLite encodes each atomic body‑part action (e.g., a single left‑foot step) as a discrete Laban symbol paired with a textual template. This abstraction decomposes complex motions into interpretable symbol sequences and body‑part instructions, establishing a symbolic link between high‑level language and low‑level motion trajectories. Building on LabanLite, we present LaMoGen, a Text‑to‑LabanLite‑to‑Motion Generation framework that enables large language models (LLMs) to compose motion sequences through symbolic reasoning. The LLM interprets motion patterns, relates them to textual descriptions, and recombines symbols into executable plans, producing motions that are both interpretable and linguistically grounded. To support rigorous evaluation, we introduce a Labanotation‑based benchmark with structured description‑motion pairs and three metrics that jointly measure text‑motion alignment across symbolic, temporal, and harmony dimensions. Experiments demonstrate that LaMoGen establishes a new baseline for both interpretability and controllability, outperforming prior methods on our benchmark and two public datasets. These results highlight the advantages of symbolic reasoning and agent‑based design for language‑driven motion synthesis.
Authors:Zhongyu Xia, Yousen Tang, Yongtao Wang, Zhifeng Wang, Weijun Qin
Abstract:
4D radar‑camera sensing configuration has gained increasing importance in autonomous driving. However, existing 3D object detection methods that fuse 4D Radar and camera data confront several challenges. First, their absolute depth estimation module is not robust and accurate enough, leading to inaccurate 3D localization. Second, the performance of their temporal fusion module will degrade dramatically or even fail when the ego vehicle's pose is missing or inaccurate. Third, for some small objects, the sparse radar point clouds may completely fail to reflect from their surfaces. In such cases, detection must rely solely on visual unimodal priors. To address these limitations, we propose R4Det, which enhances depth estimation quality via the Panoramic Depth Fusion module, enabling mutual reinforcement between absolute and relative depth. For temporal fusion, we design a Deformable Gated Temporal Fusion module that does not rely on the ego vehicle's pose. In addition, we built an Instance‑Guided Dynamic Refinement module that extracts semantic prototypes from 2D instance guidance. Experiments show that R4Det achieves state‑of‑the‑art 3D object detection results on the TJ4DRadSet and VoD datasets. The source code and models will be released at https://github.com/VDIGPKU/R4Det.
Authors:Robinson Umeike, Cuong Pham, Ryan Hausen, Thang Dao, Shane Crawford, Tanya Brown-Giammanco, Gerard Lemson, John van de Lindt, Blythe Johnston, Arik Mitschang, Trung Do
Abstract:
We present TornadoNet, a comprehensive benchmark for automated street‑level building damage assessment evaluating how modern real‑time object detection architectures and ordinal‑aware supervision strategies perform under realistic post‑disaster conditions. TornadoNet provides the first controlled benchmark demonstrating how architectural design and loss formulation jointly influence multi‑level damage detection from street‑view imagery, delivering methodological insights and deployable tools for disaster response. Using 3,333 high‑resolution geotagged images and 8,890 annotated building instances from the 2021 Midwest tornado outbreak, we systematically compare CNN‑based detectors from the YOLO family against transformer‑based models (RT‑DETR) for multi‑level damage detection. Models are trained under standardized protocols using a five‑level damage classification framework based on IN‑CORE damage states, validated through expert cross‑annotation. Baseline experiments reveal complementary architectural strengths. CNN‑based YOLO models achieve highest detection accuracy and throughput, with larger variants reaching 46.05% mAP@0.5 at 66‑276 FPS on A100 GPUs. Transformer‑based RT‑DETR models exhibit stronger ordinal consistency, achieving 88.13% Ordinal Top‑1 Accuracy and MAOE of 0.65, indicating more reliable severity grading despite lower baseline mAP. To align supervision with the ordered nature of damage severity, we introduce soft ordinal classification targets and evaluate explicit ordinal‑distance penalties. RT‑DETR trained with calibrated ordinal supervision achieves 44.70% mAP@0.5, a 4.8 percentage‑point improvement, with gains in ordinal metrics (91.15% Ordinal Top‑1 Accuracy, MAOE = 0.56). These findings establish that ordinal‑aware supervision improves damage severity estimation when aligned with detector architecture. Model & Data: https://github.com/crumeike/TornadoNet
Authors:Md Jahidul Islam
Abstract:
The adaptation of large‑scale Vision‑Language Models (VLMs) like CLIP to downstream tasks with extremely limited data ‑‑ specifically in the one‑shot regime ‑‑ is often hindered by a significant "Stability‑Plasticity" dilemma. While efficient caching mechanisms have been introduced by training‑free methods such as Tip‑Adapter, these approaches often function as local Nadaraya‑Watson estimators. Such estimators are characterized by inherent boundary bias and a lack of global structural regularization. In this paper, ReHARK (Refined Hybrid Adaptive RBF Kernels) is proposed as a synergistic training‑free framework that reinterprets few‑shot adaptation through global proximal regularization in a Reproducing Kernel Hilbert Space (RKHS). A multistage refinement pipeline is introduced, consisting of: (1) Hybrid Prior Construction, where zero‑shot textual knowledge from CLIP and GPT‑3 is fused with visual class prototypes to form a robust semantic‑visual anchor; (2) Support Set Augmentation (Bridging), where intermediate samples are generated to smooth the transition between visual and textual modalities; (3) Adaptive Distribution Rectification, where test feature statistics are aligned with the augmented support set to mitigate domain shifts; and (4) Multi‑Scale RBF Kernels, where an ensemble of kernels is employed to capture complex feature geometries across diverse scales. Superior stability and accuracy are demonstrated through extensive experiments on 11 diverse benchmarks. A new state‑of‑the‑art for one‑shot adaptation is established by ReHARK, which achieves an average accuracy of 65.83%, significantly outperforming existing baselines. Code is available at https://github.com/Jahid12012021/ReHARK.
Authors:Xiaobiao Du, Yida Wang, Kun Zhan, Xin Yu
Abstract:
3D Gaussian Splatting (3DGS) has emerged as a powerful representation for high‑quality rendering across a wide range of applications.However, its high computational demands and large storage costs pose significant challenges for deployment on mobile devices. In this work, we propose a mobile‑tailored real‑time Gaussian Splatting method, dubbed Mobile‑GS, enabling efficient inference of Gaussian Splatting on edge devices. Specifically, we first identify alpha blending as the primary computational bottleneck, since it relies on the time‑consuming Gaussian depth sorting process. To solve this issue, we propose a depth‑aware order‑independent rendering scheme that eliminates the need for sorting, thereby substantially accelerating rendering. Although this order‑independent rendering improves rendering speed, it may introduce transparency artifacts in regions with overlapping geometry due to the scarcity of rendering order. To address this problem, we propose a neural view‑dependent enhancement strategy, enabling more accurate modeling of view‑dependent effects conditioned on viewing direction, 3D Gaussian geometry, and appearance attributes. In this way, Mobile‑GS can achieve both high‑quality and real‑time rendering. Furthermore, to facilitate deployment on memory‑constrained mobile platforms, we also introduce first‑order spherical harmonics distillation, a neural vector quantization technique, and a contribution‑based pruning strategy to reduce the number of Gaussian primitives and compress the 3D Gaussian representation with the assistance of neural networks. Extensive experiments demonstrate that our proposed Mobile‑GS achieves real‑time rendering and compact model size while preserving high visual quality, making it well‑suited for mobile applications.
Authors:Xiaogang Du, Jiawei Zhang, Tongfei Liu, Tao Lei, Yingbo Wang
Abstract:
In medical image segmentation tasks, the domain gap caused by the difference in data collection between training and testing data seriously hinders the deployment of pre‑trained models in clinical practice. Continual Test‑Time Adaptation (CTTA) aims to enable pre‑trained models to adapt to continuously changing unlabeled domains, providing an effective approach to solving this problem. However, existing CTTA methods often rely on unreliable supervisory signals, igniting a self‑reinforcing cycle of error accumulation that culminates in catastrophic performance degradation. To overcome these challenges, we propose a CTTA via Semantic‑Prompt‑Enhanced Graph Clustering (SPEGC) for medical image segmentation. First, we design a semantic prompt feature enhancement mechanism that utilizes decoupled commonality and heterogeneity prompt pools to inject global contextual information into local features, alleviating their susceptibility to noise interference under domain shift. Second, based on these enhanced features, we design a differentiable graph clustering solver. This solver reframes global edge sparsification as an optimal transport problem, allowing it to distill a raw similarity matrix into a refined and high‑order structural representation in an end‑to‑end manner. Finally, this robust structural representation is used to guide model adaptation, ensuring predictions are consistent at a cluster‑level and dynamically adjusting decision boundaries. Extensive experiments demonstrate that SPEGC outperforms other state‑of‑the‑art CTTA methods on two medical image segmentation benchmarks. The source code is available at https://github.com/Jwei‑Z/SPEGC‑for‑MIS.
Authors:Seung hee Choi, MinJu Jeon, Hyunwoo Oh, Jihwan Lee, Dong-Jin Kim
Abstract:
Existing retrieval‑augmented approaches for Dense Video Captioning (DVC) often fail to achieve accurate temporal segmentation aligned with true event boundaries, as they rely on heuristic strategies that overlook ground truth event boundaries. The proposed framework, STaRC, overcomes this limitation by supervising frame‑level saliency through a highlight detection module. Note that the highlight detection module is trained on binary labels derived directly from DVC ground truth annotations without the need for additional annotation. We also propose to utilize the saliency scores as a unified temporal signal that drives retrieval via saliency‑guided segmentation and informs caption generation through explicit Saliency Prompts injected into the decoder. By enforcing saliency‑constrained segmentation, our method produces temporally coherent segments that align closely with actual event transitions, leading to more accurate retrieval and contextually grounded caption generation. We conduct comprehensive evaluations on the YouCook2 and ViTT benchmarks, where STaRC achieves state‑of‑the‑art performance across most of the metrics. Our code is available at https://github.com/ermitaju1/STaRC
Authors:Mehmet Kerem Turkcan
Abstract:
Recent advances in vision‑language modeling have produced promptable detection and segmentation systems that accept arbitrary natural language queries at inference time. Among these, SAM3 achieves state‑of‑the‑art accuracy by combining a ViT‑H/14 backbone with cross‑modal transformer decoding and learned object queries. However, SAM3 processes a single text prompt per forward pass. Detecting N categories requires N independent executions, each dominated by the 439M‑parameter backbone. We present Detect Anything in Real Time (DART), a training‑free framework that converts SAM3 into a real‑time multi‑class detector by exploiting a structural invariant: the visual backbone is class‑agnostic, producing image features independent of the text prompt. This allows the backbone computation to be shared between all classes, reducing its cost from O(N) to O(1). Combined with batched multi‑class decoding, detection‑only inference, and TensorRT FP16 deployment, these optimizations yield 5.6x cumulative speedup at 3 classes, scaling to 25x at 80 classes, without modifying any model weight. On COCO val2017 (5,000 images, 80 classes), DART achieves 55.8 AP at 15.8 FPS (4 classes, 1008x1008) on a single RTX 4080, surpassing purpose‑built open‑vocabulary detectors trained on millions of box annotations. For extreme latency targets, adapter distillation with a frozen encoder‑decoder achieves 38.7 AP with a 13.9 ms backbone. Code and models are available at https://github.com/mkturkcan/DART.
Authors:Mingzhe Tao, Ruiping Liu, Junwei Zheng, Yufan Chen, Kedi Ying, M. Saquib Sarfraz, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen
Abstract:
Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leveraging multi‑sensor information to understand adverse driving scenarios in autonomous vehicles. To address this gap, we propose the DriveXQA, a multimodal dataset for autonomous driving VQA. In addition to four visual modalities, five sensor failure cases, and five weather conditions, it includes 102,505 QA pairs categorized into three types: global scene level, allocentric level, and ego‑vehicle centric level. Since no existing MLLM framework adopts multiple complementary visual modalities as input, we design MVX‑LLM, a token‑efficient architecture with a Dual Cross‑Attention (DCA) projector that fuses the modalities to alleviate information redundancy. Experiments demonstrate that our DCA achieves improved performance under challenging conditions such as foggy (GPTScore: 53.5 vs. 25.1 for the baseline).
Authors:Yuto Shibata, Kashu Yamazaki, Lalit Jayanti, Yoshimitsu Aoki, Mariko Isogawa, Katerina Fragkiadaki
Abstract:
Humanoid robotics has strong potential to transform daily service and caregiving applications. Although recent advances in general motion tracking within physics engines (GMT) have enabled virtual characters and humanoid robots to reproduce a broad range of human motions, these behaviors are primarily limited to contact‑less social interactions or isolated movements. Assistive scenarios, by contrast, require continuous awareness of a human partner and rapid adaptation to their evolving posture and dynamics. In this paper, we formulate the imitation of closely interacting, force‑exchanging human‑human motion sequences as a multi‑agent reinforcement learning problem. We jointly train partner‑aware policies for both the supporter (assistant) agent and the recipient agent in a physics simulator to track assistive motion references. To make this problem tractable, we introduce a partner policies initialization scheme that transfers priors from single‑human motion‑tracking controllers, greatly improving exploration. We further propose dynamic reference retargeting and contact‑promoting reward, which adapt the assistant's reference motion to the recipient's real‑time pose and encourage physically meaningful support. We show that AssistMimic is the first method capable of successfully tracking assistive interaction motions on established benchmarks, demonstrating the benefits of a multi‑agent RL formulation for physically grounded and socially aware humanoid control.
Authors:Jérémy Scanvic, Quentin Barthélemy, Julián Tachella
Abstract:
The simplicity and effectiveness of the UNet architecture makes it ubiquitous in image restoration, image segmentation, and diffusion models. They are often assumed to be equivariant to translations, yet they traditionally consist of layers that are known to be prone to aliasing, which hinders their equivariance in practice. To overcome this limitation, we propose a new alias‑free UNet designed from a careful selection of state‑of‑the‑art translation‑equivariant layers. We evaluate the proposed equivariant architecture against non‑equivariant baselines on image restoration tasks and observe competitive performance with a significant increase in measured equivariance. Through extensive ablation studies, we also demonstrate that each change is crucial for its empirical equivariance. Our implementation is available at https://github.com/jscanvic/UNet‑AF
Authors:Benedikt Schwab, Thomas H. Kolbe
Abstract:
Although semantic 3D city models are internationally available and becoming increasingly detailed, the incorporation of material information remains largely untapped. However, a structured representation of materials and their physical properties could substantially broaden the application spectrum and analytical capabilities for urban digital twins. At the same time, the growing number of repeated mobile laser scans of cities and their street spaces yields a wealth of observations influenced by the material characteristics of the corresponding surfaces. To leverage this information, we propose radiometric fingerprints of object surfaces by grouping LiDAR observations reflected from the same semantic object under varying distances, incident angles, environmental conditions, sensors, and scanning campaigns. Our study demonstrates how 312.4 million individual beams acquired across four campaigns using five LiDAR sensors on the Audi Autonomous Driving Dataset (A2D2) vehicle can be automatically associated with 6368 individual objects of the semantic 3D city model. The model comprises a comprehensive and semantic representation of four inner‑city streets at Level of Detail (LOD) 3 with centimeter‑level accuracy. It is based on the CityGML 3.0 standard and enables fine‑grained sub‑differentiation of objects. The extracted radiometric fingerprints for object surfaces reveal recurring intra‑class patterns that indicate class‑dominant materials. The semantic model, the method implementations, and the developed geodatabase solution 3DSensorDB are released under: https://github.com/tum‑gis/sensordb
Authors:Yuehao Song, Shaoyu Chen, Hao Gao, Yifan Zhu, Weixiang Yue, Jialv Zou, Bo Jiang, Zihao Lu, Yu Wang, Qian Zhang, Xinggang Wang
Abstract:
Vision‑language models (VLMs) enhance the planning capability of end‑to‑end (E2E) driving policy by leveraging high‑level semantic reasoning. However, existing approaches often overlook the dual‑system consistency between VLM's high‑level decision and E2E's low‑level planning. As a result, the generated trajectories may misalign with the intended driving decisions, leading to weakened top‑down guidance and decision‑following ability of the system. To address this issue, we propose Senna‑2, an advanced VLM‑E2E driving policy that explicitly aligns the two systems for consistent decision‑making and planning. Our method follows a consistency‑oriented three‑stage training paradigm. In the first stage, we conduct driving pre‑training to achieve preliminary decision‑making and planning, with a decision adapter transmitting VLM decisions to E2E policy in the form of implicit embeddings. In the second stage, we align the VLM and the E2E policy in an open‑loop setting. In the third stage, we perform closed‑loop alignment via bottom‑up Hierarchical Reinforcement Learning in 3DGS environments to reinforce the safety and efficiency. Extensive experiments demonstrate that Senna‑2 achieves superior dual‑system consistency (19.3% F1 score improvement) and significantly enhances driving safety in both open‑loop (5.7% FDE reduction) and closed‑loop settings (30.6% AF‑CR reduction).
Authors:Yutong Chen, Yiming Wang, Xucong Zhang, Sergey Prokudin, Siyu Tang
Abstract:
Recent feed‑forward networks have achieved remarkable progress in sparse‑view 3D reconstruction by predicting dense point maps directly from RGB images. However, they often suffer from geometric inconsistencies and limited fine‑grained accuracy due to the absence of explicit multi‑view constraints. We introduce the Geometry‑Grounded Point Transformer (GGPT), a framework that augments feed‑forward reconstruction with reliable sparse geometric guidance. We first propose an improved Structure‑from‑Motion pipeline based on dense feature matching and lightweight geometric optimisation to efficiently estimate accurate camera poses and partial 3D point clouds from sparse input views. Building on this foundation, we propose a geometry‑guided 3D point transformer that refines dense point maps under explicit partial‑geometry supervision using an optimised guidance encoding. Extensive experiments demonstrate that our method provides a principled mechanism for integrating geometric priors with dense feed‑forward predictions, producing reconstructions that are both geometrically consistent and spatially complete, recovering fine structures and filling gaps in textureless areas. Trained solely on ScanNet++ with VGGT predictions, GGPT generalises across architectures and datasets, substantially outperforming state‑of‑the‑art feed‑forward 3D reconstruction models in both in‑domain and out‑of‑domain settings.
Authors:Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steve Seitz
Abstract:
We propose a fully automated AI system that produces short comedic videos similar to sketch shows such as Saturday Night Live. Starting with character references, the system employs a population of agents loosely based on real production studio roles, structured to optimize the quality and diversity of ideas and outputs through iterative competition, evaluation, and improvement. A key contribution is the introduction of LLM critics aligned with real viewer preferences through the analysis of a corpus of comedy videos on YouTube to automatically evaluate humor. Our experiments show that our framework produces results approaching the quality of professionally produced sketches while demonstrating state‑of‑the‑art performance in video generation.
Authors:Jen-Hao Rick Chang, Xiaoming Zhao, Dorian Chan, Oncel Tuzel
Abstract:
We propose a 3D latent representation that jointly models object geometry and view‑dependent appearance. Most prior works focus on either reconstructing 3D geometry or predicting view‑independent diffuse appearance, and thus struggle to capture realistic view‑dependent effects. Our approach leverages that RGB‑depth images provide samples of a surface light field. By encoding random subsamples of this surface light field into a compact set of latent vectors, our model learns to represent both geometry and appearance within a unified 3D latent space. This representation reproduces view‑dependent effects such as specular highlights and Fresnel reflections under complex lighting. We further train a latent flow matching model on this representation to learn its distribution conditioned on a single input image, enabling the generation of 3D objects with appearances consistent with the lighting and materials in the input. Experiments show that our approach achieves higher visual quality and better input fidelity than existing methods.
Authors:Tao Zhong, Yixun Hu, Dongzhe Zheng, Aditya Sood, Christine Allen-Blanchette
Abstract:
Inverse problems for stiff parabolic partial differential equations (PDEs), such as the inverse heat conduction problem (IHCP), are severely ill‑posed: the forward map rapidly damps high‑frequency interior structure before it reaches the boundary. Soft‑constrained physics‑informed neural networks (PINNs), which embed the PDE as a residual penalty, suffer from gradient pathology in this regime and tend to fit boundary measurements while leaving the interior field essentially untouched. We propose Neural Field Thermal Tomography (NeFTY), a hard‑constrained neural field framework for label‑free three‑dimensional inverse heat conduction. NeFTY represents the unknown diffusivity as a continuous coordinate‑based neural network, and at every optimization step passes the candidate field through a differentiable implicit‑Euler heat solver with harmonic‑mean interface flux, so that the governing PDE holds exactly on the discretization rather than as a soft penalty. Adjoint gradients propagate the surface reconstruction error back to the network weights at solver‑level memory cost, making test‑time inversion tractable on a single GPU. Across synthetic 3D benchmarks, NeFTY substantially outperforms soft‑constrained PINN variants and a voxel‑grid baseline on label‑free volumetric recovery, and it transfers to real thermography data, surpassing classical signal‑processing baselines in both defect segmentation and depth estimation. Additional details at https://cab‑lab‑princeton.github.io/nefty/
Authors:Yan-Bo Lin, Jonah Casebeer, Long Mai, Aniruddha Mahapatra, Gedas Bertasius, Nicholas J. Bryan
Abstract:
Generating music that temporally aligns with video events is challenging for existing text‑to‑music models, which lack fine‑grained temporal control. We introduce V2M‑ZERO, a video‑to‑music generation approach that generates time‑aligned music with disentangled time synchronization and semantic control (e.g., genre, mood) from video while requiring zero video‑music pairs at training time. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra‑modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine‑tune a text‑to‑music model on music‑event curves, then substitute video‑event curves at inference without cross‑modal training or paired data. Across OES‑Pub, MovieGenBench‑Music, and AIST++, V2M‑ZERO achieves state‑of‑the‑art performance without any paired music‑video data, surpassing the strongest prior baselines per metric with 5‑9% higher audio quality, 13‑15% better semantic alignment, 21‑52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd‑source subjective listening test. Our results validate that temporal alignment through within‑modality features is not only effective for video‑to‑music generation but also leads to better performance than paired cross‑modal supervision. Furthermore, our approach enables independent controls for timing and music style (e.g., genre, mood) for more controllable generation.
Authors:Shuyao Shang, Bing Zhan, Yunfei Yan, Yuqi Wang, Yingyan Li, Yasong An, Xiaoman Wang, Jierui Liu, Lu Hou, Lue Fan, Zhaoxiang Zhang, Tieniu Tan
Abstract:
We propose DynVLA, a driving VLA model that introduces a new CoT paradigm termed Dynamics CoT. DynVLA forecasts compact world dynamics before action generation, enabling more informed and physically grounded decision‑making. To obtain compact dynamics representations, DynVLA introduces a Dynamics Tokenizer that compresses future evolution into a small set of dynamics tokens. Considering the rich environment dynamics in interaction‑intensive driving scenarios, DynVLA decouples ego‑centric and environment‑centric dynamics, yielding more accurate world dynamics modeling. We then train DynVLA to generate dynamics tokens before actions through SFT and RFT, improving decision quality while maintaining latency‑efficient inference. Compared to Textual CoT, which lacks fine‑grained spatiotemporal understanding, and Visual CoT, which introduces substantial redundancy due to dense image prediction, Dynamics CoT captures the evolution of the world in a compact, interpretable, and efficient form. Extensive experiments on NAVSIM, Bench2Drive, and a large‑scale in‑house dataset demonstrate that DynVLA consistently outperforms Textual CoT and Visual CoT methods, validating the effectiveness and practical value of Dynamics CoT. Project Page: https://yaoyao‑jpg.github.io/dynvla.
Authors:Zhengyao Fang, Zexi Jia, Yijia Zhong, Pengcheng Luo, Jinchao Zhang, Guangming Lu, Jun Yu, Wenjie Pei
Abstract:
Recent advances in text‑to‑image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real‑world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference‑trained metrics often favor visually vivid images with exaggerated saturation and contrast, which make generations often too vivid to be real even when prompted for realistic‑style images. To address this issue, we present Color Fidelity Dataset (CFD) and Color Fidelity Metric (CFM) for objective evaluation of color fidelity in realistic‑style generations. CFD contains over 1.3M real and synthetic images with ordered levels of color realism, while CFM employs a multimodal encoder to learn perceptual color fidelity. In addition, we propose a training‑free Color Fidelity Refinement (CFR) that adaptively modulates spatial‑temporal guidance scale in generation, thereby enhancing color authenticity. Together, CFD supports CFM for assessment, whose learned attention further guides CFR to refine T2I fidelity, forming a progressive framework for assessing and improving color fidelity in realistic‑style T2I generation. The dataset and code are available at https://github.com/ZhengyaoFang/CFM.
Authors:Konrad Szafer, Marek Kraft, Dominik Belter
Abstract:
Foundation models for point cloud data have recently grown in capability, often leveraging extensive representation learning from language or vision. In this work, we take a more controlled approach by introducing a lightweight transformer‑based point cloud architecture. In contrast to the heavy reliance on cross‑modal supervision, our model is trained only on 39k point clouds ‑ yet it outperforms several larger foundation models trained on over 200k training samples. Interestingly, our method approaches state‑of‑the‑art results from models that have seen over a million point clouds, images, and text samples, demonstrating the value of a carefully curated training setup and architecture. To ensure rigorous evaluation, we conduct a comprehensive replication study that standardizes the training regime and benchmarks across multiple point cloud architectures. This unified experimental framework isolates the impact of architectural choices, allowing for transparent comparisons and highlighting the benefits of our design and other tokenizer‑free architectures. Our results show that simple backbones can deliver competitive results to more complex or data‑rich strategies. The implementation, including code, pre‑trained models, and training protocols, is available at https://github.com/KonradSzafer/Pointy.
Authors:Zegu Zhang, Jian Zhang
Abstract:
Variational autoencoders (VAEs) frequently suffer from posterior collapse, where the latent variables become uninformative as the approximate posterior degenerates to the prior. While recent work has characterized collapse as a phase transition determined by data covariance properties, existing approaches primarily aim to avoid rather than eliminate collapse. We introduce a novel framework that theoretically guarantees non‑collapsed solutions by leveraging spherical shell geometry and cluster‑aware constraints. Our method transforms data to a spherical shell, computes optimal cluster assignments via K‑means, and defines a feasible region between the within‑cluster variance W and collapse loss δ_\textcollapse. We prove that when the reconstruction loss is constrained to this region, the collapsed solution is mathematically excluded from the feasible parameter space. Critically, we introduce norm constraint mechanisms that ensure decoder outputs remain compatible with the spherical shell geometry without restricting representational capacity. Unlike prior approaches, our method provides a strict theoretical guarantee with minimal computational overhead without imposing constraints on decoder outputs. Experiments on synthetic and real‑world datasets demonstrate 100% collapse prevention under conditions where conventional VAEs completely fail, with reconstruction quality matching or exceeding state‑of‑the‑art methods. Our approach requires no explicit stability conditions (e.g., σ^2 < λ_\max) and works with arbitrary neural architectures. The code is available at https://github.com/tsegoochang/spherical‑vae‑with‑Cluster.
Authors:Fanqi Yu, Matteo Tiezzi, Tommaso Apicella, Cigdem Beyan, Vittorio Murino
Abstract:
We introduce a lifelong imitation learning framework that enables continual policy refinement across sequential tasks under realistic memory and data constraints. Our approach departs from conventional experience replay by operating entirely in a multimodal latent space, where compact representations of visual, linguistic, and robot's state information are stored and reused to support future learning. To further stabilize adaptation, we introduce an incremental feature adjustment mechanism that regularizes the evolution of task embeddings through an angular margin constraint, preserving inter‑task distinctiveness. Our method establishes a new state of the art in the LIBERO benchmarks, achieving 10‑17 point gains in AUC and up to 65% less forgetting compared to previous leading methods. Ablation studies confirm the effectiveness of each component, showing consistent gains over alternative strategies. The code is available at: https://github.com/yfqi/lifelong_mlr_ifa.
Authors:Yan Zhang, Long Ma, Yuxin Feng, Zhe Huang, Fan Zhou, Zhuo Su
Abstract:
Learning‑based real image dehazing methods have achieved notable progress, yet they still face adaptation challenges in diverse real haze scenes. These challenges mainly stem from the lack of effective unsupervised mechanisms for unlabeled data and the heavy cost of full model fine‑tuning. To address these challenges, we propose the haze‑to‑clear text‑directed loss that leverages CLIP's cross‑modal capabilities to reformulate real image dehazing as a semantic alignment problem in latent space, thereby providing explicit unsupervised cross‑modal guidance in the absence of reference images. Furthermore, we introduce the Bilevel Layer‑positioning LoRA (BiLaLoRA) strategy, which learns both the LoRA parameters and automatically search the injection layers, enabling targeted adaptation of critical network layers. Extensive experiments demonstrate our superiority against state‑of‑the‑art methods on multiple real‑world dehazing benchmarks. The code is publicly available at https://github.com/YanZhang‑zy/BiLaLoRA.
Authors:Lin Chen, Bolin Ni, Qi Yang, Zili Wang, Kun Ding, Ying Wang, Houwen Peng, Shiming Xiang
Abstract:
Despite the remarkable capabilities of Multimodal Large Language Models (MLLMs), they still suffer from visual fading in long‑context scenarios. Specifically, the attention to visual tokens diminishes as the text sequence lengthens, leading to text generation detached from visual constraints. We attribute this degradation to the inherent inductive bias of Multimodal RoPE, which penalizes inter‑modal attention as the distance between visual and text tokens increases. To address this, we propose inter‑modal Distance Invariant Position Encoding (DIPE), a simple but effective mechanism that disentangles position encoding based on modality interactions. DIPE retains the natural relative positioning for intra‑modal interactions to preserve local structure, while enforcing an anchored perceptual proximity for inter‑modal interactions. This strategy effectively mitigates the inter‑modal distance‑based penalty, ensuring that visual signals remain perceptually consistent regardless of the context length. Experimental results demonstrate that by integrating DIPE with Multimodal RoPE, the model maintains stable visual grounding in long‑context scenarios, significantly alleviating visual fading while preserving performance on standard short‑context benchmarks. Code is available at https://github.com/lchen1019/DIPE.
Authors:Shilong Han, Yuming Zhang, Hongxia Wang
Abstract:
Classifier‑Free Guidance (CFG) is a cornerstone of modern text‑to‑image models, yet its reliance on a semantically vacuous null prompt (\varnothing) generates a guidance signal prone to geometric entanglement. This is a key factor limiting its precision, leading to well‑documented failures in complex compositional tasks. We propose Condition‑Degradation Guidance (CDG), a novel paradigm that replaces the null prompt with a strategically degraded condition, \boldsymbolc_\textdeg. This reframes guidance from a coarse "good vs. null" contrast to a more refined "good vs. almost good" discrimination, thereby compelling the model to capture fine‑grained semantic distinctions. We find that tokens in transformer text encoders split into two functional roles: content tokens encoding object semantics, and context‑aggregating tokens capturing global context. By selectively degrading only the former, CDG constructs \boldsymbolc_\textdeg without external models or training. Validated across diverse architectures including Stable Diffusion 3, FLUX, and Qwen‑Image, CDG markedly improves compositional accuracy and text‑image alignment. As a lightweight, plug‑and‑play module, it achieves this with negligible computational overhead. Our work challenges the reliance on static, information‑sparse negative samples and establishes a new principle for diffusion guidance: the construction of adaptive, semantically‑aware negative samples is critical to achieving precise semantic control. Code is available at https://github.com/Ming‑321/Classifier‑Degradation‑Guidance.
Authors:Tongkun Guan, Zhibo Yang, Jianqiang Wan, Mingkun Yang, Zhengtao Guo, Zijian Hu, Ruilin Luo, Ruize Chen, Songtao Jiang, Peng Wang, Wei Shen, Junyang Lin, Xiaokang Yang
Abstract:
When MLLMs fail at Science, Technology, Engineering, and Mathematics (STEM) visual reasoning, a fundamental question arises: is it due to perceptual deficiencies or reasoning limitations? Through systematic scaling analysis that independently scales perception and reasoning components, we uncover a critical insight: scaling perception consistently outperforms scaling reasoning. This reveals perception as the true lever limiting current STEM visual reasoning. Motivated by this insight, our work focuses on systematically enhancing the perception capabilities of MLLMs by establishing code as a powerful perceptual medium‑‑executable code provides precise semantics that naturally align with the structured nature of STEM visuals. Specifically, we construct ICC‑1M, a large‑scale dataset comprising 1M Image‑Caption‑Code triplets that materializes this code‑as‑perception paradigm through two complementary approaches: (1) Code‑Grounded Caption Generation treats executable code as ground truth for image captions, eliminating the hallucinations inherent in existing knowledge distillation methods; (2) STEM Image‑to‑Code Translation prompts models to generate reconstruction code, mitigating the ambiguity of natural language for perception enhancement. To validate this paradigm, we further introduce STEM2Code‑Eval, a novel benchmark that directly evaluates visual perception in STEM domains. Unlike existing work relying on problem‑solving accuracy as a proxy that only measures problem‑relevant understanding, our benchmark requires comprehensive visual comprehension through executable code generation for image reconstruction, providing deterministic and verifiable assessment. Code is available at https://github.com/TongkunGuan/Qwen‑CodePercept.
Authors:Wenhao Sun, Ji Li, Zhaoqiang Liu
Abstract:
Diffusion Transformers have established a new state‑of‑the‑art in image synthesis, but the high computational cost of iterative sampling severely hampers their practical deployment. While existing acceleration methods often focus on the temporal domain, they overlook the substantial spatial redundancy inherent in the generative process, where global structures emerge long before fine‑grained details are formed. The uniform computational treatment of all spatial regions represents a critical inefficiency. In this paper, we introduce Just‑in‑Time (JiT), a novel training‑free framework that addresses this challenge by acceleration in the spatial domain. JiT formulates a spatially approximated generative ordinary differential equation (ODE) that drives the full latent state evolution based on computations from a dynamically selected, sparse subset of anchor tokens. To ensure seamless transitions as new tokens are incorporated to expand the dimensions of the latent state, we propose a deterministic micro‑flow, a simple and effective finite‑time ODE that maintains both structural coherence and statistical correctness. Extensive experiments on the state‑of‑the‑art FLUX.1‑dev model demonstrate that JiT achieves up to a 7x speedup with nearly lossless performance, significantly outperforming existing acceleration methods and establishing a new and superior trade‑off between inference speed and generation fidelity.
Authors:Yu Zhang, Zhicheng Zhao, Ze Luo, Chenglong Li, Jin Tang
Abstract:
Traffic scene understanding from unmanned aerial vehicle (UAV) platforms is crucial for intelligent transportation systems due to its flexible deployment and wide‑area monitoring capabilities. However, existing methods face significant challenges in real‑world surveillance, as their heavy reliance on optical imagery leads to severe performance degradation under adverse illumination conditions like nighttime and fog. Furthermore, current Visual Question Answering (VQA) models are restricted to elementary perception tasks, lacking the domain‑specific regulatory knowledge required to assess complex traffic behaviors. To address these limitations, we propose a novel Multi‑modal Traffic Cognition Network (MTCNet) for robust UAV traffic scene understanding. Specifically, we design a Prototype‑Guided Knowledge Embedding (PGKE) module that leverages high‑level semantic prototypes from an external Traffic Regulation Memory (TRM) to anchor domain‑specific knowledge into visual representations, enabling the model to comprehend complex behaviors and distinguish fine‑grained traffic violations. Moreover, we develop a Quality‑Aware Spectral Compensation (QASC) module that exploits the complementary characteristics of optical and thermal modalities to perform bidirectional context exchange, effectively compensating for degraded features to ensure robust representation in complex environments. In addition, we construct Traffic‑VQA, the first large‑scale optical‑thermal infrared benchmark for cognitive UAV traffic understanding, comprising 8,180 aligned image pairs and 1.3 million question‑answer pairs across 31 diverse types. Extensive experiments demonstrate that MTCNet significantly outperforms state‑of‑the‑art methods in both cognition and perception scenarios. The dataset is available at https://github.com/YuZhang‑2004/UAV‑traffic‑scene‑understanding.
Authors:Rafi Ibn Sultan, Hui Zhu, Xiangyu Zhou, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
Abstract:
Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision‑Language Models (LVLMs) struggle to meet. Although these models can describe visual content, their lack of explicit grounding leads to object hallucinations and unreliable depth reasoning, limiting their usefulness for accessibility guidance. We introduce WalkGPT, a pixel‑grounded LVLM for the new task of Grounded Navigation Guide, unifying language reasoning and segmentation within a single architecture for depth‑aware accessibility guidance. Given a pedestrian‑view image and a navigation query, WalkGPT generates a conversational response with segmentation masks that delineate accessible and harmful features, along with relative depth estimation. The model incorporates a Multi‑Scale Query Projector (MSQP) that shapes the final image tokens by aggregating them along text tokens across spatial hierarchies, and a Calibrated Text Projector (CTP), guided by a proposed Region Alignment Loss, that maps language embeddings into segmentation‑aware representations. These components enable fine‑grained grounding and depth inference without user‑provided cues or anchor points, allowing the model to generate complete and realistic navigation guidance. We also introduce PAVE, a large‑scale benchmark of 41k pedestrian‑view images paired with accessibility‑aware questions and depth‑grounded answers. Experiments show that WalkGPT achieves strong grounded reasoning and segmentation performance. The source code and dataset are available on the \hrefhttps://sites.google.com/view/walkgpt‑26/homeproject website.
Authors:Jeonghyeok Do, Yun Chen, Geunhyuk Youk, Munchurl Kim
Abstract:
The landscape of skeleton‑based action representation learning has evolved from Contrastive Learning (CL) to Masked Auto‑Encoder (MAE) architectures. However, each paradigm faces inherent limitations: CL often overlooks fine‑grained local details, while MAE is burdened by computationally heavy decoders. Moreover, MAE suffers from severe computational asymmetry ‑‑ benefiting from efficient masking during pre‑training but requiring exhaustive full‑sequence processing for downstream tasks. To resolve these bottlenecks, we propose SLiM (Skeleton Less is More), a novel unified framework that harmonizes masked modeling with contrastive learning via a shared encoder. By eschewing the reconstruction decoder, SLiM not only eliminates computational redundancy but also compels the encoder to capture discriminative features directly. SLiM is the first framework with decoder‑free masked modeling of representative learning. Crucially, to prevent trivial reconstruction arising from high skeletal‑temporal correlation, we introduce semantic tube masking, alongside skeletal‑aware augmentations designed to ensure anatomical consistency across diverse temporal granularities. Extensive experiments demonstrate that SLiM consistently achieves state‑of‑the‑art performance across all downstream protocols. Notably, our method delivers this superior accuracy with exceptional efficiency, reducing inference computational cost by 7.89x compared to existing MAE methods.
Authors:Stefanos Pasios, Nikos Nikolaidis
Abstract:
Generative models are widely employed to enhance the photorealism of visual synthetic data for training computer vision algorithms. However, they often introduce visual artifacts that degrade the accuracy of these algorithms and require high computational resources, limiting their applicability in real‑time training or evaluation scenarios. In this paper, we propose Hybrid Patch Enhanced Realism Generative Adversarial Network (HyPER‑GAN), a lightweight image‑to‑image translation method based on a U‑Net‑style generator designed for real‑time inference. The model is trained using paired synthetic and photorealism‑enhanced images, complemented by a hybrid training strategy that incorporates matched patches from real‑world images to improve visual realism and semantic consistency. Experimental results demonstrate that HyPER‑GAN outperforms state‑of‑the‑art lightweight paired image‑to‑image translation methods in terms of inference latency, visual realism, and semantic robustness. Moreover, it is illustrated that the proposed hybrid training strategy indeed improves visual quality and semantic consistency compared to training the model solely with paired synthetic and photorealism‑enhanced images. Code and pretrained models are publicly available for download at: https://github.com/stefanos50/HyPER‑GAN
Authors:Yawen Yang, Feng Li, Shuqi Kong, Yunfeng Diao, Xinjian Gao, Zenglin Shi, Meng Wang
Abstract:
Recent rapid advancement of generative models has significantly improved the fidelity and accessibility of AI‑generated synthetic images. While enabling various innovative applications, the unprecedented realism of these synthetics makes them increasingly indistinguishable from authentic photographs, posing serious security risks, such as media credibility and content manipulation. Although extensive efforts have been dedicated to detecting synthetic images, most existing approaches suffer from poor generalization to unseen data due to their reliance on model‑specific artifacts or low‑level statistical cues. In this work, we identify a previously unexplored distinction that real images maintain consistent semantic attention and structural coherence in their latent representations, exhibiting more stable feature transitions across network layers, whereas synthetic ones present discernible distinct patterns. Therefore, we propose a novel approach termed latent transition discrepancy (LTD), which captures the inter‑layer consistency differences of real and synthetic images. LTD adaptively identifies the most discriminative layers and assesses the transition discrepancies across layers. Benefiting from the proposed inter‑layer discriminative modeling, our approach exceeds the base model by 14.35% in mean Acc across three datasets containing diverse GANs and DMs. Extensive experiments demonstrate that LTD outperforms recent state‑of‑the‑art methods, achieving superior detection accuracy, generalizability, and robustness. The code is available at https://github.com/yywencs/LTD
Authors:Jakub Gregorek, Paraskevas Pegios, Nando Metzger, Konrad Schindler, Theodora Kontogianni, Lazaros Nalpantidis
Abstract:
We introduce Marigold‑SSD, a single‑step, late‑fusion depth completion framework that leverages strong diffusion priors while eliminating the costly test‑time optimization typically associated with diffusion‑based methods. By shifting computational burden from inference to finetuning, our approach enables efficient and robust 3D perception under real‑world latency constraints. Marigold‑SSD achieves significantly faster inference with a training cost of only 4.5 GPU days. We evaluate our method across four indoor and two outdoor benchmarks, demonstrating strong cross‑domain generalization and zero‑shot performance compared to existing depth completion approaches. Our approach significantly narrows the efficiency gap between diffusion‑based and discriminative models. Finally, we challenge common evaluation protocols by analyzing performance under varying input sparsity levels. Page: https://dtu‑pas.github.io/marigold‑ssd/
Authors:Hongsong Wang, Renxi Cheng, Chaolei Han, Jie Gui
Abstract:
With the rapid advancement of AIGC technologies, image forensics will encounter unprecedented challenges. Traditional methods are incapable of dealing with increasingly realistic images generated by rapidly evolving image generation techniques. To facilitate the identification of AI‑generated images and the attribution of their source models, generative image watermarking and AI‑generated image attribution have emerged as key research focuses in recent years. However, existing methods are model‑dependent, requiring access to the generative models and lacking generality and scalability to new and unseen generators. To address these limitations, this work presents a new paradigm for AI‑generated image attribution by formulating it as an instance retrieval problem instead of a conventional image classification problem. We propose an efficient model‑agnostic framework, called Low‑bIt‑plane‑based Deepfake Attribution (LIDA). The input to LIDA is produced by Low‑Bit Fingerprint Generation module, while the training involves Unsupervised Pre‑Training followed by subsequent Few‑Shot Attribution Adaptation. Comprehensive experiments demonstrate that LIDA achieves state‑of‑the‑art performance for both Deepfake detection and image attribution under zero‑ and few‑shot settings. The code is at https://github.com/hongsong‑wang/LIDA
Authors:Yuan Mei, Lang Nie, Kang Liao, Yunqiu Xu, Chunyu Lin, Bin Xiao
Abstract:
Traditional image stitching methods estimate warps from hand‑crafted geometric features, whereas recent learning‑based solutions leverage semantic features from neural networks instead. These two lines of research have largely diverged along separate evolution, with virtually no meaningful convergence to date. In this paper, we take a pioneering step to bridge this gap by unifying semantic and geometric features with UniStitch, a unified image stitching framework from multimodal features. To align discrete geometric features (i.e., keypoint) with continuous semantic feature maps, we present a Neural Point Transformer (NPT) module, which transforms unordered, sparse 1D geometric keypoints into ordered, dense 2D semantic maps. Then, to integrate the advantages of both representations, an Adaptive Mixture of Experts (AMoE) module is designed to fuse geometric and semantic representations. It dynamically shifts focus toward more reliable features during the fusion process, allowing the model to handle complex scenes, especially when either modality might be compromised. The fused representation can be adopted into common deep stitching pipelines, delivering significant performance gains over any single feature. Experiments show that UniStitch outperforms existing state‑of‑the‑art methods with a large margin, paving the way for a unified paradigm between traditional and learning‑based image stitching.
Authors:Longan Wang, Yuang Shi, Wei Tsang Ooi
Abstract:
Gaussian splatting has emerged as a competitive explicit representation for image and video reconstruction. In this work, we present P‑GSVC, the first layered progressive 2D Gaussian splatting framework that provides a unified solution for scalable Gaussian representation in both images and videos. P‑GSVC organizes 2D Gaussian splats into a base layer and successive enhancement layers, enabling coarse‑to‑fine reconstructions. To effectively optimize this layered representation, we propose a joint training strategy that simultaneously updates Gaussians across layers, aligning their optimization trajectories to ensure inter‑layer compatibility and a stable progressive reconstruction. P‑GSVC supports scalability in terms of both quality and resolution. Our experiments show that the joint training strategy can gain up to 1.9 dB improvement in PSNR for video and 2.6 dB improvement in PSNR for image when compared to methods that perform sequential layer‑wise training. Project page: https://longanwang‑cs.github.io/PGSVC‑webpage/
Authors:Caroline Magg, Maaike A. ter Wee, Johannes G. G. Dobbe, Geert J. Streekstra, Leendert Blankevoort, Clara I. Sánchez, Hoel Kervadec
Abstract:
Promptable Foundation Models (FMs), initially introduced for natural image segmentation, have also revolutionized medical image segmentation. The increasing number of models, along with evaluations varying in datasets, metrics, and compared models, makes direct performance comparison between models difficult and complicates the selection of the most suitable model for specific clinical tasks. In our study, 11 promptable FMs are tested using non‑iterative 2D and 3D prompting strategies on a private and public dataset focusing on bone and implant segmentation in four anatomical regions (wrist, shoulder, hip and lower leg). The Pareto‑optimal models are identified and further analyzed using human prompts collected through a dedicated observer study. Our findings are: 1) The segmentation performance varies a lot between FMs and prompting strategies; 2) The Pareto‑optimal models in 2D are SAM and SAM2.1, in 3D nnInteractive and Med‑SAM2; 3) Localization accuracy and rater consistency vary with anatomical structures, with higher consistency for simple structures (wrist bones) and lower consistency for complex structures (pelvis, tibia, implants); 4) The segmentation performance drops using human prompts, suggesting that performance reported on "ideal" prompts extracted from reference labels might overestimate the performance in a human‑driven setting; 5) All models were sensitive to prompt variations. While two models demonstrated intra‑rater robustness, it did not scale to inter‑rater settings. We conclude that the selection of the most optimal FM for a human‑driven setting remains challenging, with even high‑performing FMs being sensitive to variations in human input prompts. Our code base for prompt extraction and model inference is available: https://github.com/CarolineMagg/segmentation‑FM‑benchmark/
Authors:Pei Liu, Xiangxiang Zeng, Tengfei Ma, Yucheng Xing, Xuanbai Ren, Yiping Liu
Abstract:
Whole‑Slide Images (WSIs) are widely used for estimating the prognosis of cancer patients. Current studies generally follow a cancer‑specific learning paradigm. However, the available training samples for one cancer type are usually scarce in pathology. Consequently, the model often struggles to learn generalizable knowledge, thus performing worse on the tumor samples with inherent high heterogeneity. Although multi‑cancer joint learning and knowledge transfer approaches have been explored recently to address it, they either rely on large‑scale joint training or extensive inference across multiple models, posing new challenges in computational efficiency. To this end, this paper proposes a new scheme, Sparse Task Vector Mixup with Hypernetworks (STEPH). Unlike previous ones, it efficiently absorbs generalizable knowledge from other cancers for the target via model merging: i) applying task vector mixup to each source‑target pair and then ii) sparsely aggregating task vector mixtures to obtain an improved target model, driven by hypernetworks. Extensive experiments on 13 cancer datasets show that STEPH improves over cancer‑specific learning and an existing knowledge transfer baseline by 5.14% and 2.01%, respectively. Moreover, it is a more efficient solution for learning prognostic knowledge from other cancers, without requiring large‑scale joint training or extensive multi‑model inference. Code is publicly available at https://github.com/liupei101/STEPH.
Authors:Xin Huang, Junjie Liang, Qingshan Hou, Peng Cao, Jinzhu Yang, Xiaoli Liu, Osmar R. Zaiane
Abstract:
Medical image synthesis is crucial for alleviating data scarcity and privacy constraints. However, fine‑tuning general text‑to‑image (T2I) models remains challenging, mainly due to the significant modality gap between complex visual details and abstract clinical text. In addition, semantic entanglement persists, where coarse‑grained text embeddings blur the boundary between anatomical structures and imaging styles, thus weakening controllability during generation. To address this, we propose a Visually‑Guided Text Disentanglement framework. We introduce a cross‑modal latent alignment mechanism that leverages visual priors to explicitly disentangle unstructured text into independent semantic representations. Subsequently, a Hybrid Feature Fusion Module (HFFM) injects these features into a Diffusion Transformer (DiT) through separated channels, enabling fine‑grained structural control. Experimental results in three datasets demonstrate that our method outperforms existing approaches in terms of generation quality and significantly improves performance on downstream classification tasks. The source code is available at https://github.com/hx111/VG‑MedGen.
Authors:Hamidreza Dastmalchi, Aijun An, Ali Cheraghian, Hamed Barzamini
Abstract:
While large vision‑language models (LVLMs) achieve strong performance on multimodal tasks, they frequently generate hallucinations ‑‑ unfaithful outputs misaligned with the visual input. To address this issue, we introduce CIPHER (Counterfactual Image Perturbations for Hallucination Extraction and Removal), a training‑free method that suppresses vision‑induced hallucinations via lightweight feature‑level correction. Unlike prior training‑free approaches that primarily focus on text‑induced hallucinations, CIPHER explicitly targets hallucinations arising from the visual modality. CIPHER operates in two phases. In the offline phase, we construct OHC‑25K (Object‑Hallucinated Counterfactuals, 25,000 samples), a counterfactual dataset consisting of diffusion‑edited images that intentionally contradict the original ground‑truth captions. We pair these edited images with the unchanged ground‑truth captions and process them through an LVLM to extract hallucination‑related representations. Contrasting these representations with those from authentic (image, caption) pairs reveals structured, systematic shifts spanning a low‑rank subspace characterizing vision‑induced hallucination. In the inference phase, CIPHER suppresses hallucinations by projecting intermediate hidden states away from this subspace. Experiments across multiple benchmarks show that CIPHER significantly reduces hallucination rates while preserving task performance, demonstrating the effectiveness of counterfactual visual perturbations for improving LVLM faithfulness. Code and additional materials are available at https://hamidreza‑dastmalchi.github.io/cipher‑cvpr2026/.
Authors:Dengdi Sun, Jie Chen, Xiao Wang, Jin Tang
Abstract:
Physics‑Informed Neural Networks (PINNs) have shown promise in solving incompressible Navier‑Stokes equations, yet existing approaches are predominantly designed for single‑flow settings. When extended to multi‑flow scenarios, these methods face three key challenges: (1) difficulty in simultaneously capturing both shared physical principles and flow‑specific characteristics, (2) susceptibility to inter‑task negative transfer that degrades prediction accuracy, and (3) unstable training dynamics caused by disparate loss magnitudes across heterogeneous flow regimes. To address these limitations, we propose UniPINN, a unified multi‑flow PINN framework that integrates three complementary components: a shared‑specialized architecture that disentangles universal physical laws from flow‑specific features, a cross‑flow attention mechanism that selectively reinforces relevant patterns while suppressing task‑irrelevant interference, and a dynamic weight allocation strategy that adaptively balances loss contributions to stabilize multi‑objective optimization. Extensive experiments on three canonical flows demonstrate that UniPINN effectively unifies multi‑flow learning, achieving superior prediction accuracy and balanced performance across heterogeneous regimes while successfully mitigating negative transfer. The source code of this paper will be released on https://github.com/Event‑AHU/OpenFusion
Authors:Yijie Li, Xi Zhu, Junyi Wang, Ye Wu, Lauren J. O'Donnell, Fan Zhang
Abstract:
Diffusion MRI tractography enables in vivo reconstruction of white matter (WM) pathways. Two key tasks in tractography analysis include: 1) tractogram registration that aligns streamlines across individuals, and 2) streamline clustering that groups streamlines into compact fiber bundles. Although both tasks share the goal of capturing geometrically similar structures to characterize consistent WM organization, they are typically performed independently. In this work, we propose TractoRC, a unified probabilistic framework that jointly performs tractogram registration and streamline clustering within a single optimization scheme, enabling the two tasks to leverage complementary information. TractoRC learns a latent embedding space for streamline points, which serves as a shared representation for both tasks. Within this space, both tasks are formulated as probabilistic inference over structural representations: registration learns the distribution of anatomical landmarks as probabilistic keypoints to align tractograms across subjects, and clustering learns streamline structural prototypes that capture geometric similarity to form coherent streamline clusters. To support effective learning of this shared space, we introduce a transformation‑equivariant self‑supervised strategy to learn geometry‑aware and transformation‑invariant embeddings. Experiments demonstrate that jointly optimizing registration and clustering significantly improves performance in both tasks over state‑of‑the‑art methods that treat them independently. Code will be made publicly available at https://github.com/yishengpoxiao/TractoRC .
Authors:Tianshuo Xu, Zhifei Chen, Leyi Wu, Hao Lu, Ying-cong Chen
Abstract:
The ultimate goal of video generation is to satisfy a fundamental trilemma: achieving high visual quality, maintaining rigorous physical consistency, and enabling precise controllability. While recent models can maintain this balance in simple, isolated scenarios, we observe that this equilibrium is fragile and often breaks down as scene complexity increases (e.g., involving collisions or dense traffic). To address this, we introduce Motion Forcing, a framework designed to stabilize this trilemma even in complex generative tasks. Our key insight is to explicitly decouple physical reasoning from visual synthesis via a hierarchical ``Point‑Shape‑Appearance'' paradigm. This approach decomposes generation into verifiable stages: modeling complex dynamics as sparse geometric anchors (Point), expanding them into dynamic depth maps that explicitly resolve 3D geometry (Shape), and finally rendering high‑fidelity textures (Appearance). Furthermore, to foster robust physical understanding, we employ a Masked Point Recovery strategy. By randomly masking input anchors during training and enforcing the reconstruction of complete dynamic depth, the model is compelled to move beyond passive pattern matching and learn latent physical laws (e.g., inertia) to infer missing trajectories. Extensive experiments on autonomous driving benchmarks show that Motion Forcing significantly outperforms state‑of‑the‑art baselines, maintaining trilemma stability across complex scenes. Evaluations on physics and robotics further confirm our framework's generality.
Authors:Hangyu Liu, Jianyong Wang, Yutao Sun
Abstract:
Latent diffusion models have established a new state‑of‑the‑art in high‑resolution visual generation. Integrating Vision Foundation Model priors improves generative efficiency, yet existing latent designs remain largely heuristic. These approaches often struggle to unify semantic discriminability, reconstruction fidelity, and latent compactness. In this paper, we propose Geometric Autoencoder (GAE), a principled framework that systematically addresses these challenges. By analyzing various alignment paradigms, GAE constructs an optimized low‑dimensional semantic supervision target from VFMs to provide guidance for the autoencoder. Furthermore, we leverage latent normalization that replaces the restrictive KL‑divergence of standard VAEs, enabling a more stable latent manifold specifically optimized for diffusion learning. To ensure robust reconstruction under high‑intensity noise, GAE incorporates a dynamic noise sampling mechanism. Empirically, GAE achieves compelling performance on the ImageNet‑1K 256 × 256 benchmark, reaching a gFID of 1.82 at only 80 epochs and 1.31 at 800 epochs without Classifier‑Free Guidance, significantly surpassing existing state‑of‑the‑art methods. Beyond generative quality, GAE establishes a superior equilibrium between compression, semantic depth and robust reconstruction stability. These results validate our design considerations, offering a promising paradigm for latent diffusion modeling. Code and models are publicly available at https://github.com/sii‑research/GAE.
Authors:Ke Zhang, Xiangchen Zhao, Yunjie Tian, Jiayu Zheng, Vishal M. Patel, Di Fu
Abstract:
Conventional video classification models, acting as effective imitators, excel in scenarios with homogeneous data distributions. However, real‑world applications often present an open‑instance challenge, where intra‑class variations are vast and complex, beyond existing benchmarks. While traditional video encoder models struggle to fit these diverse distributions, vision‑language models (VLMs) offer superior generalization but have not fully leveraged their reasoning capabilities (intuition) for such tasks. In this paper, we bridge this gap with an intrinsic reasoning framework that evolves open‑instance video classification from imitation to intuition. Our approach, namely DeepIntuit, begins with a cold‑start supervised alignment to initialize reasoning capability, followed by refinement using Group Relative Policy Optimization (GRPO) to enhance reasoning coherence through reinforcement learning. Crucially, to translate this reasoning into accurate classification, DeepIntuit then introduces an intuitive calibration stage. In this stage, a classifier is trained on this intrinsic reasoning traces generated by the refined VLM, ensuring stable knowledge transfer without distribution mismatch. Extensive experiments demonstrate that for open‑instance video classification, DeepIntuit benefits significantly from transcending simple feature imitation and evolving toward intrinsic reasoning. Our project is available at https://bwgzk‑keke.github.io/DeepIntuit/.
Authors:Xiaoyan Zhang, Jiangpeng He
Abstract:
Class‑incremental learning (CIL) aims to acquire new classes over time while retaining prior knowledge, yet most setups and methods assume balanced task streams. In practice, the number of classes per task often varies significantly. We refer to this as step imbalance, where large tasks that contain more classes dominate learning and small tasks inject unstable updates. Existing CIL methods assume balanced tasks and therefore treat all tasks uniformly, producing imbalanced updates that degrade overall learning performance. To address this challenge, we propose One‑A, a unified and imbalance‑aware framework that incrementally merges task updates into a single adapter, maintaining constant inference cost. One‑A performs asymmetric subspace alignment to preserve dominant subspaces learned from large tasks while constraining low‑information updates within them. An information‑adaptive weighting balances the contribution between base and new adapters, and a directional gating mechanism selectively fuses updates along each singular direction, maintaining stability in head directions and plasticity in tail ones. Across multiple benchmarks and step‑imbalanced streams, One‑A achieves competitive accuracy with significantly low inference overhead, showing that a single, asymmetrically fused adapter can remain both adaptive to dynamic task sizes and efficient at deployment.
Authors:Shuaiyu Chen, Ming Yin, Peng Ren, Chunbo Luo, Zeyu Fu
Abstract:
Segmenting oil spills from Synthetic Aperture Radar (SAR) imagery remains challenging due to severe appearance variability, scale heterogeneity, and the absence of temporal continuity in real world monitoring scenarios. While foundation models such as Segment Anything (SAM) enable prompt driven segmentation, existing SAM based approaches operate on single images and cannot effectively reuse information across scenes. Memory augmented variants (e.g., SAM2) further assume temporal coherence, making them prone to semantic drift when applied to unordered SAR image collections. We propose OilSAM2, a memory augmented segmentation framework tailored for unordered SAR oil spill monitoring. OilSAM2 introduces a hierarchical feature aware multi scale memory bank that explicitly models texture, structure, and semantic level representations, enabling robust cross image information reuse. To mitigate memory drift, we further propose a structure semantic consistent memory update strategy that selectively refreshes memory based on semantic discrepancy and structural variation.Experiments on two public SAR oil spill datasets demonstrate that OilSAM2 achieves state of the art segmentation performance, delivering stable and accurate results under noisy SAR monitoring scenarios. The source code is available at https://github.com/Chenshuaiyu1120/OILSAM2.
Authors:Feng Li, Ziyuan Li, Zhongliang Jiang, Nassir Navab, Yuan Bi
Abstract:
Intraoperative Cone Beam Computed Tomography (CBCT) provides a reliable 3D anatomical context essential for interventional planning. However, its static nature fails to provide continuous monitoring of soft‑tissue deformations induced by respiration, probe pressure, and surgical manipulation, leading to navigation discrepancies. We propose a deformation‑aware CBCT updating framework that leverages robotic ultrasound as a dynamic proxy to infer tissue motion and update static CBCT slices in real time. Starting from calibration‑initialized alignment with linear correlation of linear combination (LC2)‑based rigid refinement, our method establishes accurate multimodal correspondence. To capture intraoperative dynamics, we introduce the ultrasound correlation UNet (USCorUNet), a lightweight network trained with optical flow‑guided supervision to learn deformation‑aware correlation representations, enabling accurate, real‑time dense deformation field estimation from ultrasound streams. The inferred deformation is spatially regularized and transferred to the CBCT reference to produce deformation‑consistent visualizations without repeated radiation exposure. We validate the proposed approach through deformation estimation and ultrasound‑guided CBCT updating experiments. Results demonstrate real‑time end‑to‑end CBCT slice updating and physically plausible deformation estimation, enabling dynamic refinement of static CBCT guidance during robotic ultrasound‑assisted interventions. The source code is publicly available at https://github.com/anonymous‑codebase/us‑cbct‑demo.
Authors:Chujie Chang, Shoko Miyauchi, Ken'ichi Morooka, Ryo Kurazume, Oscar Martinez Mozos
Abstract:
Cardiac magnetic resonance (CMR) imaging is widely used to visualise cardiac motion and diagnose heart disease. However, standard CMR imaging requires patients to lie still in a confined space inside a loud machine for 40‑60 min, which increases patient discomfort. In addition, shorter scan times decrease either or both the temporal and spatial resolutions of cardiac motion, and thus, the diagnostic accuracy of the procedure. Of these, we focus on reduced temporal resolution and propose a neural network called FusionNet to obtain four‑dimensional (4D) cardiac motion with high temporal resolution from CMR images captured in a short period of time. The model estimates intermediate 3D heart shapes based on adjacent shapes. The results of an experimental evaluation of the proposed FusionNet model showed that it achieved a performance of over 0.897 in terms of the Dice coefficient, confirming that it can recover shapes more precisely than existing methods. This code is available at: https://github.com/smiyauchi199/FusionNet.git
Authors:Yinjie Wang, Xuyang Chen, Xiaolong Jin, Mengdi Wang, Ling Yang
Abstract:
Every agent interaction generates a next‑state signal, namely the user reply, tool output, terminal or GUI state change that follows each action, yet no existing agentic RL system recovers it as a live, online learning source. We present OpenClaw‑RL, a framework that employs next‑state signals to optimize personal agents online through infrastructure and methodology innovations. On the infrastructure side, we extend existing RL systems to a server‑client architecture where the RL server hosts the policy behind an inference API and user terminals stream interaction data back over HTTP. From each observed next state, the system extracts two complementary training signals, evaluative and directive, via a separate asynchronous server so that neither signal extraction nor optimization blocks inference. On the methodology side, we introduce a hybrid RL objective that unifies both signal types in a single update: directive signals provide richer, token‑level supervision but are sparser, while evaluative signals are more broadly available. To stabilize distillation under teacher‑student mismatch, we propose overlap‑guided hint selection, which picks the hint whose induced teacher distribution maximally overlaps with the student's top‑k tokens, together with a log‑probability‑difference clip that bounds per‑token advantages. Applied to personal agents, OpenClaw‑RL enables an agent to improve simply by being used, recovering conversational signals from user re‑queries, corrections, and explicit feedback. Applied to general agents, OpenClaw‑RL is the first RL framework to unify real‑world agent settings spanning terminal, GUI, SWE, and tool‑call environments, where we additionally demonstrate the utility of next‑state signals in long‑horizon settings.
Authors:Daichao Zhao, Qiupu Chen, Feng He, Xin Ning, Qiankun Li
Abstract:
Lane detection is a crucial task in autonomous driving, as it helps ensure the safe operation of vehicles. However, existing datasets such as CULane and TuSimple contain relatively limited data under extreme weather conditions, including rain, snow, and fog. As a result, detection models trained on these datasets often become unreliable in such environments, which may lead to serious safety‑critical failures on the road. To address this issue, we propose HG‑Lane, a High‑fidelity Generation framework for Lane Scenes under adverse weather and lighting conditions without requiring re‑annotation. Based on this framework, we further construct a benchmark that includes adverse weather and lighting scenarios, containing 30,000 images. Experimental results demonstrate that our method consistently and significantly improves the performance of existing lane detection networks. For example, using the state‑of‑the‑art CLRNet, the overall mF1 score on our benchmark increases by 20.87 percent. The F1@50 score for the overall, normal, snow, rain, fog, night, and dusk categories increases by 19.75 percent, 8.63 percent, 38.8 percent, 14.96 percent, 26.84 percent, 21.5 percent, and 12.04 percent, respectively. The code and dataset are available at: https://github.com/zdc233/HG‑Lane.
Authors:Jin Lyu, Liang An, Pujin Cheng, Yebin Liu, Xiaoying Tang
Abstract:
4D reconstruction of equine family (e.g. horses) from monocular video is important for animal welfare. Previous mainstream 4D animal reconstruction methods require joint optimization of motion and appearance over a whole video, which is time‑consuming and sensitive to incomplete observation. In this work, we propose a novel framework called 4DEquine by disentangling the 4D reconstruction problem into two sub‑problems: dynamic motion reconstruction and static appearance reconstruction. For motion, we introduce a simple yet effective spatio‑temporal transformer with a post‑optimization stage to regress smooth and pixel‑aligned pose and shape sequences from video. For appearance, we design a novel feed‑forward network that reconstructs a high‑fidelity, animatable 3D Gaussian avatar from as few as a single image. To assist training, we create a large‑scale synthetic motion dataset, VarenPoser, which features high‑quality surface motions and diverse camera trajectories, as well as a synthetic appearance dataset, VarenTex, comprising realistic multi‑view images generated through multi‑view diffusion. While training only on synthetic datasets, 4DEquine achieves state‑of‑the‑art performance on real‑world APT36K and AiM datasets, demonstrating the superiority of 4DEquine and our new datasets for both geometry and appearance reconstruction. Comprehensive ablation studies validate the effectiveness of both the motion and appearance reconstruction network. Project page: https://luoxue‑star.github.io/4DEquine_Project_Page/.
Authors:Lucas Prieto, Edward Stevinson, Melih Barsbey, Tolga Birdal, Pedro A. M. Mediano
Abstract:
A central idea in mechanistic interpretability is that neural networks represent more features than they have dimensions, arranging them in superposition to form an over‑complete basis. This framing has been influential, motivating dictionary learning approaches such as sparse autoencoders. However, superposition has mostly been studied in idealized settings where features are sparse and uncorrelated. In these settings, superposition is typically understood as introducing interference that must be minimized geometrically and filtered out by non‑linearities such as ReLUs, yielding local structures like regular polytopes. We show that this account is incomplete for realistic data by introducing Bag‑of‑Words Superposition (BOWS), a controlled setting to encode binary bag‑of‑words representations of internet text in superposition. Using BOWS, we find that when features are correlated, interference can be constructive rather than just noise to be filtered out. This is achieved by arranging features according to their co‑activation patterns, making interference between active features constructive, while still using ReLUs to avoid false positives. We show that this kind of arrangement is more prevalent in models trained with weight decay and naturally gives rise to semantic clusters and cyclical structures which have been observed in real language models yet were not explained by the standard picture of superposition. Code for this paper can be found at https://github.com/LucasPrietoAl/correlations‑feature‑geometry.
Authors:Xinyu Gao, Gang Chen, Javier Alonso-Mora
Abstract:
Language‑conditioned local navigation requires a robot to infer a nearby traversable target location from its current observation and an open‑vocabulary, relational instruction. Existing vision‑language spatial grounding methods usually rely on vision‑language models (VLMs) to reason in image space, producing 2D predictions tied to visible pixels. As a result, they struggle to infer target locations in occluded regions, typically caused by furniture or moving humans. To address this issue, we propose BEACON, which predicts an ego‑centric Bird's‑Eye View (BEV) affordance heatmap over a bounded local region including occluded areas. Given an instruction and surround‑view RGB‑D observations from four directions around the robot, BEACON predicts the BEV heatmap by injecting spatial cues into a VLM and fusing the VLM's output with depth‑derived BEV features. Using an occlusion‑aware dataset built in the Habitat simulator, we conduct detailed experimental analysis to validate both our BEV space formulation and the design choices of each module. Our method improves the accuracy averaged across geodesic thresholds by 22.74 percentage points over the state‑of‑the‑art image‑space baseline on the validation subset with occluded target locations. Our project page is: https://xin‑yu‑gao.github.io/beacon.
Authors:Rong Zhou, Houliang Zhou, Yao Su, Brian Y. Chen, Yu Zhang, Lifang He, Alzheimer's Disease Neuroimaging Initiative
Abstract:
Multimodal neuroimaging provides complementary insights for Alzheimer's disease diagnosis, yet clinical datasets frequently suffer from missing modalities. We propose ACADiff, a framework that synthesizes missing brain imaging modalities through adaptive clinical‑aware diffusion. ACADiff learns mappings between incomplete multimodal observations and target modalities by progressively denoising latent representations while attending to available imaging data and clinical metadata. The framework employs adaptive fusion that dynamically reconfigures based on input availability, coupled with semantic clinical guidance via GPT‑4o‑encoded prompts. Three specialized generators enable bidirectional synthesis among sMRI, FDG‑PET, and AV45‑PET. Evaluated on ADNI subjects, ACADiff achieves superior generation quality and maintains robust diagnostic performance even under extreme 80% missing scenarios, outperforming all existing baselines. To promote reproducibility, code is available at https://github.com/rongzhou7/ACADiff
Authors:Shan Ning, Longtian Qiu, Jiaxuan Sun, Xuming He
Abstract:
Open‑domain visual entity recognition (VER) seeks to associate images with entities in encyclopedic knowledge bases such as Wikipedia. Recent generative methods tailored for VER demonstrate strong performance but incur high computational costs, limiting their scalability and practical deployment. In this work, we revisit the contrastive paradigm for VER and introduce WikiCLIP, a simple yet effective framework that establishes a strong and efficient baseline for open‑domain VER. WikiCLIP leverages large language model embeddings as knowledge‑rich entity representations and enhances them with a Vision‑Guided Knowledge Adaptor (VGKA) that aligns textual semantics with visual cues at the patch level. To further encourage fine‑grained discrimination, a Hard Negative Synthesis Mechanism generates visually similar but semantically distinct negatives during training. Experimental results on popular open‑domain VER benchmarks, such as OVEN, demonstrate that WikiCLIP significantly outperforms strong baselines. Specifically, WikiCLIP achieves a 16% improvement on the challenging OVEN unseen set, while reducing inference latency by nearly 100 times compared with the leading generative model, AutoVER. The project page is available at https://artanic30.github.io/project_pages/WikiCLIP/
Authors:Jiazhi Guan, Quanwei Yang, Luying Huang, Junhao Liang, Borong Liang, Haocheng Feng, Wei He, Kaisiyuan Wang, Hang Zhou, Jingdong Wang
Abstract:
Human‑centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human‑Object Interaction (HOI) videos. Existing works rely on dense control signals, template videos, or carefully crafted text prompts, which limit flexibility and generalization to novel objects. We introduce a framework, namely DISPLAY, guided by Sparse Motion Guidance, composed only of wrist joint coordinates and a shape‑agnostic object bounding box. This lightweight guidance alleviates the imbalance between human and object representations and enables intuitive user control. To enhance fidelity under such sparse conditions, we propose an Object‑Stressed Attention mechanism that improves object robustness. To address the scarcity of high‑quality HOI data, we further develop a Multi‑Task Auxiliary Training strategy with a dedicated data curation pipeline, allowing the model to benefit from both reliable HOI samples and auxiliary tasks. Comprehensive experiments show that our method achieves high‑fidelity, controllable HOI generation across diverse tasks. The project page can be found at \hrefhttps://mumuwei.github.io/DISPLAY/.
Authors:Changyao Tian, Danni Yang, Guanzhou Chen, Erfei Cui, Zhaokai Wang, Yuchen Duan, Penghao Yin, Sitao Chen, Ganlin Yang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Leyao Gu, Haomin Wang, Qi Wei, Jinhui Yin, Xue Yang, Zhihang Zhong, Qi Qin, Yi Xin, Bin Fu, Yihao Liu, Jiaye Ge, Qipeng Guo, Gen Luo, Hongsheng Li, Yu Qiao, Kai Chen, Hongjie Zhang
Abstract:
Unified multimodal models (UMMs) that integrate understanding, reasoning, generation, and editing face inherent trade‑offs between maintaining strong semantic comprehension and acquiring powerful generation capabilities. In this report, we present InternVL‑U, a lightweight 4B‑parameter UMM that democratizes these capabilities within a unified framework. Guided by the principles of unified contextual modeling and modality‑specific modular design with decoupled visual representations, InternVL‑U integrates a state‑of‑the‑art Multimodal Large Language Model (MLLM) with a specialized MMDiT‑based visual generation head. To further bridge the gap between aesthetic generation and high‑level intelligence, we construct a comprehensive data synthesis pipeline targeting high‑semantic‑density tasks, such as text rendering and scientific reasoning, under a reasoning‑centric paradigm that leverages Chain‑of‑Thought (CoT) to better align abstract user intent with fine‑grained visual generation details. Extensive experiments demonstrate that InternVL‑U achieves a superior performance ‑ efficiency balance. Despite using only 4B parameters, it consistently outperforms unified baseline models with over 3x larger scales such as BAGEL (14B) on various generation and editing tasks, while retaining strong multimodal understanding and reasoning capabilities.
Authors:Shuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao, Qilin Zhang, Benjamin Busam, Xieyuanli Chen, Yun Liu
Abstract:
Text‑to‑point‑cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and communicate spatial layouts through language. However, existing methods largely rely on shallow text‑point cloud correspondence without effective spatial reasoning, limiting their accuracy in complex environments. To address this limitation, we propose VLM‑Loc, a framework that leverages the spatial reasoning capability of large vision‑language models (VLMs) for T2P localization. Specifically, we transform point clouds into bird's‑eye‑view (BEV) images and scene graphs that jointly encode geometric and semantic context, providing structured inputs for the VLM to learn cross‑modal representations bridging linguistic and spatial semantics. On top of these representations, we introduce a partial node assignment mechanism that explicitly associates textual cues with scene graph nodes, enabling interpretable spatial reasoning for accurate localization. To facilitate systematic evaluation across diverse scenes, we present CityLoc, a benchmark built from multi‑source point clouds for fine‑grained T2P localization. Experiments on CityLoc demonstrate VLM‑Loc achieves superior accuracy and robustness compared to state‑of‑the‑art methods. Our code, model, and dataset are available at \hrefhttps://github.com/MCG‑NKU/nku‑3d‑visionrepository.
Authors:Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu, Fanman Meng, Lili Pan, Hongliang Li
Abstract:
Efficient adaptation between Egocentric (Ego) and Exocentric (Exo) views is crucial for applications such as human‑robot cooperation. However, the success of most existing Ego‑Exo adaptation methods relies heavily on target‑view data for training, thereby increasing computational and data collection costs. In this paper, we make the first exploration of a Test‑time Ego‑Exo Adaptation for Action Anticipation (TE^2A^3) task, which aims to adjust the source‑view‑trained model online during test time to anticipate target‑view actions. It is challenging for existing Test‑Time Adaptation (TTA) methods to address this task due to the multi‑action candidates and significant temporal‑spatial inter‑view gap. Hence, we propose a novel Dual‑Clue enhanced Prototype Growing Network (DCPGN), which accumulates multi‑label knowledge and integrates cross‑modality clues for effective test‑time Ego‑Exo adaptation and action anticipation. Specifically, we propose a Multi‑Label Prototype Growing Module (ML‑PGM) to balance multiple positive classes via multi‑label assignment and confidence‑based reweighting for class‑wise memory banks, which are updated by an entropy priority queue strategy. Then, the Dual‑Clue Consistency Module (DCCM) introduces a lightweight narrator to generate textual clues indicating action progressions, which complement the visual clues containing various objects. Moreover, we constrain the inferred textual and visual logits to construct dual‑clue consistency for temporally and spatially bridging Ego and Exo views. Extensive experiments on the newly proposed EgoMe‑anti and the existing EgoExoLearn benchmarks show the effectiveness of our method, which outperforms related state‑of‑the‑art methods by a large margin. Code is available at \hrefhttps://github.com/ZhaofengSHI/DCPGNhttps://github.com/ZhaofengSHI/DCPGN.
Authors:Guoliang Zhu, Wanjun Jia, Caoyang Shao, Yuheng Zhang, Zhiyong Li, Kailun Yang
Abstract:
Global perception is essential for embodied agents in 360° spaces, yet current affordance grounding remains largely object‑centric and restricted to perspective views. To bridge this gap, we introduce a novel task: Holistic Affordance Grounding in 360° Indoor Environments. This task faces unique challenges, including severe geometric distortions from Equirectangular Projection (ERP), semantic dispersion, and cross‑scale alignment difficulties. We propose PanoAffordanceNet, an end‑to‑end framework featuring a Distortion‑Aware Spectral Modulator (DASM) for latitude‑dependent calibration and an Omni‑Spherical Densification Head (OSDH) to restore topological continuity from sparse activations. By integrating multi‑level constraints comprising pixel‑wise, distributional, and region‑text contrastive objectives, our framework effectively suppresses semantic drift under low supervision. Furthermore, we construct 360‑AGD, the first high‑quality panoramic affordance grounding dataset. Extensive experiments demonstrate that PanoAffordanceNet significantly outperforms existing methods, establishing a solid baseline for scene‑level perception in embodied intelligence. The source code and benchmark dataset will be made publicly available at https://github.com/GL‑ZHU925/PanoAffordanceNet.
Authors:Francesco Ragusa, Rosario Leonardi, Michele Mazzamuto, Daniele Di Mauro, Camillo Quattrocchi, Alessandro Passanisi, Irene D'Ambra, Antonino Furnari, Giovanni Maria Farinella
Abstract:
Understanding human behavior from complementary egocentric (ego) and exocentric (exo) points of view enables the development of systems that can support workers in industrial environments and enhance their safety. However, progress in this area is hindered by the lack of datasets capturing both views in realistic industrial scenarios. To address this gap, we propose ENIGMA‑360, a new ego‑exo dataset acquired in a real industrial scenario. The dataset is composed of 180 egocentric and 180 exocentric procedural videos temporally synchronized offering complementary information of the same scene. The 360 videos have been labeled with temporal and spatial annotations, enabling the study of different aspects of human behavior in industrial domain. We provide baseline experiments for 3 foundational tasks for human behavior understanding: 1) Temporal Action Segmentation, 2) Keystep Recognition and 3) Egocentric Human‑Object Interaction Detection, showing the limits of state‑of‑the‑art approaches on this challenging scenario. These results highlight the need for new models capable of robust ego‑exo understanding in real‑world environments. We publicly release the dataset and its annotations at https://fpv‑iplab.github.io/ENIGMA‑360/.
Authors:Kaixin Lin, Kunyu Peng, Di Wen, Yufan Chen, Ruiping Liu, Kailun Yang
Abstract:
Semantic occupancy prediction enables dense 3D geometric and semantic understanding for autonomous driving. However, existing camera‑based approaches implicitly assume complete surround‑view observations, an assumption that rarely holds in real‑world deployment due to occlusion, hardware malfunction, or communication failures. We study semantic occupancy prediction under incomplete multi‑camera inputs and introduce M^2‑Occ, a framework designed to preserve geometric structure and semantic coherence when views are missing. M^2‑Occ addresses two complementary challenges. First, a Multi‑view Masked Reconstruction (MMR) module leverages the spatial overlap among neighboring cameras to recover missing‑view representations directly in the feature space. Second, a Feature Memory Module (FMM) introduces a learnable memory bank that stores class‑level semantic prototypes. By retrieving and integrating these global priors, the FMM refines ambiguous voxel features, ensuring semantic consistency even when observational evidence is incomplete. We introduce a systematic missing‑view evaluation protocol on the nuScenes‑based SurroundOcc benchmark, encompassing both deterministic single‑view failures and stochastic multi‑view dropout scenarios. Under the safety‑critical missing back‑view setting, M^2‑Occ improves the IoU by 4.93%. As the number of missing cameras increases, the robustness gap further widens; for instance, under the setting with five missing views, our method boosts the IoU by 5.01%. These gains are achieved without compromising full‑view performance. The source code will be publicly released at https://github.com/qixi7up/M2‑Occ.
Authors:Minh Khoa Le, Kien Do, Duc Thanh Nguyen, Truyen Tran
Abstract:
High‑fidelity video generation remains challenging for diffusion models due to the difficulty of modeling complex spatio‑temporal dynamics efficiently. Recent video diffusion methods typically represent a video as a sequence of spatio‑temporal tokens which can be modeled using Diffusion Transformers (DiTs). However, this approach faces a trade‑off between the strong but expensive Full 3D Attention and the efficient but temporally limited Local Factorized Attention. To resolve this trade‑off, we propose Matrix Attention, a frame‑level temporal attention mechanism that processes an entire frame as a matrix and generates query, key, and value matrices via matrix‑native operations. By attending across frames rather than tokens, Matrix Attention effectively preserves global spatio‑temporal structure and adapts to significant motion. We build FrameDiT‑G, a DiT architecture based on MatrixAttention, and further introduce FrameDiT‑H, which integrates Matrix Attention with Local Factorized Attention to capture both large and small motion. Extensive experiments show that FrameDiT‑H achieves state‑of‑the‑art results across multiple video generation benchmarks, offering improved temporal coherence and video quality while maintaining efficiency comparable to Local Factorized Attention.
Authors:Luca Carlini, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque
Abstract:
Surgical Video Question Answering (VideoQA) requires accurate temporal grounding while remaining robust to natural variation in how clinicians phrase questions, where linguistic bias can arise. Standard Parameter Efficient Fine Tuning (PEFT) methods adapt pretrained projections without explicitly modeling frame‑to‑frame interactions within the adaptation pathway, limiting their ability to exploit sparse temporal evidence. We introduce TemporalDoRA, a video‑specific PEFT formulation that extends Weight‑Decomposed Low‑Rank Adaptation by (i) inserting lightweight temporal Multi‑Head Attention (MHA) inside the low‑rank bottleneck of the vision encoder and (ii) selectively applying weight decomposition only to the trainable low‑rank branch rather than the full adapted weight. This design enables temporally‑aware updates while preserving a frozen backbone and stable scaling. By mixing information across frames within the adaptation subspace, TemporalDoRA steers updates toward temporally consistent visual cues and improves robustness with minimal parameter overhead. To benchmark this setting, we present REAL‑Colon‑VQA, a colonoscopy VideoQA dataset with 6,424 clip‑‑question pairs, including paired rephrased Out‑of‑Template questions to evaluate sensitivity to linguistic variation. TemporalDoRA improves Out‑of‑Template performance, and ablation studies confirm that temporal mixing inside the low‑rank branch is the primary driver of these gains. We also validate on EndoVis18‑VQA adapted to short clips and observe consistent improvements on the Out‑of‑Template split. Code and dataset available at~\hrefhttps://anonymous.4open.science/r/TemporalDoRA‑BFC8/Anonymous GitHub.
Authors:Yuanhang Lei, Boming Zhao, Zesong Yang, Xingxuan Li, Tao Cheng, Haocheng Peng, Ru Zhang, Yang Yang, Siyuan Huang, Yujun Shen, Ruizhen Hu, Hujun Bao, Zhaopeng Cui
Abstract:
Modeling wind‑driven object dynamics from video observations is highly challenging due to the invisibility and spatio‑temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics‑informed differentiable framework that unifies wind‑object interaction modeling, video‑based reconstruction, and forward simulation. Specifically, we represent wind as a grid‑based physical field and objects as particle systems derived from 3D Gaussian Splatting, with their interaction modeled by the Material Point Method (MPM). To recover wind‑driven object dynamics, we introduce a reconstruction framework that jointly optimizes the spatio‑temporal wind force field and object motion through differentiable rendering and simulation. To ensure physical validity, we incorporate the Lattice Boltzmann Method (LBM) as a physics‑informed constraint, enforcing compliance with fluid dynamics laws. Beyond reconstruction, our method naturally supports forward simulation under novel wind conditions and enables new applications such as wind retargeting. We further introduce WD‑Objects, a dataset of synthetic and real‑world wind‑driven scenes. Extensive experiments demonstrate that our method significantly outperforms prior dynamic scene modeling approaches in both reconstruction accuracy and simulation fidelity, opening a new avenue for video‑based wind‑object interaction modeling.
Authors:KunHo Heo, SuYeon Kim, Yonghyun Gwon, Youngbin Kim, MyeongAh Cho
Abstract:
Text‑to‑motion synthesis aims to generate natural and expressive human motions from textual descriptions. While existing approaches primarily focus on generating holistic motions from text descriptions, they struggle to accurately reflect actions involving specific body parts. Recent part‑wise motion generation methods attempt to resolve this but face two critical limitations: (i) they lack explicit mechanisms for aligning textual semantics with individual body parts, and (ii) they often generate incoherent full‑body motions due to integrating independently generated part motions. To overcome these issues and resolve the fundamental trade‑off in existing methods, we propose ParTY, a novel framework that enhances part expressiveness while generating coherent full‑body motions. ParTY comprises: (1) Part‑Guided Network, which first generates part motions to obtain part guidance, then uses it to generate holistic motions; (2) Part‑aware Text Grounding, which diversely transforms text embeddings and appropriately aligns them with each body part; and (3) Holistic‑Part Fusion, which adaptively fuses holistic motions and part motions. Extensive experiments, including part‑level and coherence‑level evaluations, demonstrate that ParTY achieves substantial improvements over previous methods.
Authors:Chaodong Xiao, Zhengqiang Zhang, Lei Zhang
Abstract:
Transformers have achieved widespread and remarkable success, while the computational complexity of their attention modules remains a major bottleneck for vision tasks. Existing methods mainly employ 8‑bit or 4‑bit quantization to balance efficiency and accuracy. In this paper, with theoretical justification, we indicate that binarization of attention preserves the essential similarity relationships, and propose BinaryAttention, an effective method for fast and accurate 1‑bit qk‑attention. Specifically, we retain only the sign of queries and keys in computing the attention, and replace the floating dot products with bit‑wise operations, significantly reducing the computational cost. We mitigate the inherent information loss under 1‑bit quantization by incorporating a learnable bias, and enable end‑to‑end acceleration. To maintain the accuracy of attention, we adopt quantization‑aware training and self‑distillation techniques, mitigating quantization errors while ensuring sign‑aligned similarity. BinaryAttention is more than 2x faster than FlashAttention2 on A100 GPUs. Extensive experiments on vision transformer and diffusion transformer benchmarks demonstrate that BinaryAttention matches or even exceeds full‑precision attention, validating its effectiveness. Our work provides a highly efficient and effective alternative to full‑precision attention, pushing the frontier of low‑bit vision and diffusion transformers. The codes and models can be found at https://github.com/EdwardChasel/BinaryAttention.
Authors:Weijia Fan, Ruiping Liu, Jiale Wei, Yufan Chen, Junwei Zheng, Zichao Zeng, Jiaming Zhang, Qiufu Li, Linlin Shen, Rainer Stiefelhagen
Abstract:
Existing vision‑language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field‑of‑view inputs to piece together a complete omni‑scene understanding. Yet, such multi‑view perception overlooks the holistic spatial and contextual relationships that a single panorama inherently preserves. In this work, we introduce the Panorama‑Language Modeling (PLM)paradigm, a unified 360^\circ vision‑language reasoning that is more than the sum of its pinhole counterparts. Besides, we present PanoVQA, a large‑scale panoramic VQA dataset that involves adverse omni‑scenes, enabling comprehensive reasoning under object occlusions and driving accidents. To establish a foundation for PLM, we develop a plug‑and‑play panoramic sparse attention module that allows existing pinhole‑based VLMs to process equirectangular panoramas without retraining. Extensive experiments demonstrate that our PLM achieves superior robustness and holistic reasoning under challenging omni‑scenes, yielding understanding greater than the sum of its narrow parts. Project page: https://github.com/InSAI‑Lab/PanoVQA.
Authors:Lang Sun, Ronghao Fu, Zhuoran Duan, Haoran Liu, Xueyan Liu, Bo Yang
Abstract:
While Vision‑Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step‑by‑step reasoning remains highly challenging. Recent efforts to introduce Chain‑of‑Thought (CoT) reasoning to this domain have shown promise, yet ensuring the visual faithfulness of these intermediate steps remains a critical bottleneck. To address this, we introduce GeoSolver, a novel framework that transitions remote sensing reasoning toward verifiable, process‑supervised reinforcement learning. We first construct Geo‑PRM‑2M, a large‑scale, token‑level process supervision dataset synthesized via entropy‑guided Monte Carlo Tree Search (MCTS) and targeted visual hallucination injection. Building upon this dataset, we train GeoPRM, a token‑level process reward model (PRM) that provides granular faithfulness feedback. To effectively leverage these verification signals, we propose Process‑Aware Tree‑GRPO, a reinforcement learning algorithm that integrates tree‑structured exploration with a faithfulness‑weighted reward mechanism to precisely assign credit to intermediate steps. Extensive experiments demonstrate that our resulting model, GeoSolver‑9B, achieves state‑of‑the‑art performance across diverse remote sensing benchmarks. Crucially, GeoPRM unlocks robust Test‑Time Scaling (TTS). Serving as a universal geospatial verifier, it seamlessly scales the performance of GeoSolver‑9B and directly enhances general‑purpose VLMs, highlighting its remarkable cross‑model generalization.
Authors:Won Shik Jang, Ue-Hwan Kim
Abstract:
Text‑goal instance navigation (TGIN) asks an agent to resolve a single, free‑form description into actions that reach the correct object instance among same‑category distractors. We present Context‑Nav, which elevates long, contextual captions from a local matching cue to a global exploration prior and verifies candidates through 3D spatial reasoning. First, we compute dense text‑image alignments for a value map that ranks frontiers ‑‑ guiding exploration toward regions consistent with the entire description rather than early detections. Second, upon observing a candidate, we perform a viewpoint‑aware relation check: the agent samples plausible observer poses, aligns local frames, and accepts a target only if the spatial relations can be satisfied from at least one viewpoint. The pipeline requires no task‑specific training or fine‑tuning; we attain state‑of‑the‑art performance on InstanceNav and CoIN‑Bench. Ablations show that (i) encoding full captions into the value map avoids wasted motion and (ii) explicit, viewpoint‑aware 3D verification prevents semantically plausible but incorrect stops. This suggests that geometry‑grounded spatial reasoning is a scalable alternative to heavy policy training or human‑in‑the‑loop interaction for fine‑grained instance disambiguation in cluttered 3D scenes.
Authors:Zhengyao Fang, Pengyuan Lyu, Chengquan Zhang, Guangming Lu, Jun Yu, Wenjie Pei
Abstract:
Vision‑language models (VLMs) face significant computational inefficiencies caused by excessive generation of visual tokens. While prior work shows that a large fraction of visual tokens are redundant, existing compression methods struggle to balance importance preservation and information diversity. To address this, we propose PruneSID, a training‑free Synergistic Importance‑Diversity approach featuring a two‑stage pipeline: (1) Principal Semantic Components Analysis (PSCA) for clustering tokens into semantically coherent groups, ensuring comprehensive concept coverage, and (2) Intra‑group Non‑Maximum Suppression (NMS) for pruning redundant tokens while preserving key representative tokens within each group. Additionally, PruneSID incorporates an information‑aware dynamic compression ratio mechanism that optimizes token compression rates based on image complexity, enabling more effective average information preservation across diverse scenes. Extensive experiments demonstrate state‑of‑the‑art performance, achieving 96.3% accuracy on LLaVA‑1.5 with only 11.1% token retention, and 92.8% accuracy at extreme compression rates (5.6%) on LLaVA‑NeXT, outperforming prior methods by 2.5% with 7.8 × faster prefilling speed compared to the original model. Our framework generalizes across diverse VLMs and both image and video modalities, showcasing strong cross‑modal versatility. Code is available at https://github.com/ZhengyaoFang/PruneSID.
Authors:Jiajun Cao, Xiaoan Zhang, Xiaobao Wei, Liyuqiu Huang, Zijian Wang, Hanzhen Zhang, Zhengyu Jia, Wei Mao, Hao Wang, Xianming Liu, Shuchang Zhou, Yang Wang, Shanghang Zhang
Abstract:
Vision‑Language‑Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumulated instability in long‑term planning. To address these challenges, we propose EvoDriveVLA‑a novel collaborative perception‑planning distillation framework that integrates self‑anchored perceptual constraints and future‑informed trajectory optimization. Specifically, self‑anchored visual distillation leverages self‑anchor teacher to deliver visual anchoring constraints, regularizing student representations via trajectory‑guided key‑region awareness. In parallel, future‑informed trajectory distillation employs a future‑aware oracle teacher with coarse‑to‑fine trajectory refinement and Monte Carlo dropout sampling to synthesize reasoning trajectories that model future evolutions, enabling the student model to internalize the future‑aware insights of the teacher. EvoDriveVLA achieves SOTA performance in nuScenes open‑loop evaluation and significantly enhances performance in NAVSIM closed‑loop evaluation. Our code is available at: https://github.com/hey‑cjj/EvoDriveVLA.
Authors:Bohao Li, Zhicheng Cao, Huixian Li, Yangming Guo
Abstract:
State‑of‑the‑art whole‑body pose estimators often lack robustness, producing anatomically implausible predictions in challenging scenes. We posit this failure stems from spurious correlations learned from visual context, a problem we formalize using a Structural Causal Model (SCM). The SCM identifies visual context as a confounder that creates a non‑causal backdoor path, corrupting the model's reasoning. We introduce the Causal Intervention Graph Pose (CIGPose) framework to address this by approximating the true causal effect between visual evidence and pose. The core of CIGPose is a novel Causal Intervention Module: it first identifies confounded keypoint representations via predictive uncertainty and then replaces them with learned, context‑invariant canonical embeddings. These deconfounded embeddings are processed by a hierarchical graph neural network that reasons over the human skeleton at both local and global semantic levels to enforce anatomical plausibility. Extensive experiments show CIGPose achieves a new state‑of‑the‑art on COCO‑WholeBody. Notably, our CIGPose‑x model achieves 67.0% AP, surpassing prior methods that rely on extra training data. With the additional UBody dataset, CIGPose‑x is further boosted to 67.5% AP, demonstrating superior robustness and data efficiency. The codes and models are publicly available at https://github.com/53mins/CIGPose.
Authors:Zirui Zhang, Yaping Zhang, Lu Xiang, Yang Zhao, Feifei Zhai, Yu Zhou, Chengqing Zong
Abstract:
Document Layout Analysis (DLA) is crucial for document artificial intelligence and has recently received increasing attention, resulting in an influx of large‑scale public DLA datasets. Existing work often combines data from various domains in recent public DLA datasets to improve the generalization of DLA. However, directly merging these datasets for training often results in suboptimal model performance, as it overlooks the different layout structures inherent to various domains. These variations include different labeling styles, document types, and languages. This paper introduces PromptDLA, a domain‑aware Prompter for Document Layout Analysis that effectively leverages descriptive knowledge as cues to integrate domain priors into DLA. The innovative PromptDLA features a unique domain‑aware prompter that customizes prompts based on the specific attributes of the data domain. These prompts then serve as cues that direct the DLA toward critical features and structures within the data, enhancing the model's ability to generalize across varied domains. Extensive experiments show that our proposal achieves state‑of‑the‑art performance among DocLayNet, PubLayNet, M6Doc, and D^4LA. Our code is available at https://github.com/Zirui00/PromptDLA.
Authors:Taesung Kwon, Lorenzo Bianchi, Lennart Wittke, Felix Watine, Fabio Carrara, Jong Chul Ye, Romann Weber, Vinicius Azevedo
Abstract:
Recent diffusion models increasingly favor Transformer backbones, motivated by the remarkable scalability of fully attentional architectures. Yet the locality bias, parameter efficiency, and hardware friendliness‑‑the attributes that established ConvNets as the efficient vision backbone‑‑have seen limited exploration in modern generative modeling. Here we introduce the fully convolutional diffusion model (FCDM), a model having a backbone similar to ConvNeXt, but designed for conditional diffusion modeling. We find that using only 50% of the FLOPs of DiT‑XL/2, FCDM‑XL achieves competitive performance with 7× and 7.5× fewer training steps at 256×256 and 512×512 resolutions, respectively. Remarkably, FCDM‑XL can be trained on a 4‑GPU system, highlighting the exceptional training efficiency of our architecture. Our results demonstrate that modern convolutional designs provide a competitive and highly efficient alternative for scaling diffusion models, reviving ConvNeXt as a simple yet powerful building block for efficient generative modeling.
Authors:Zhe Li, Xiaoyu Ding, Jiaxin Zheng, Yongtao Wang
Abstract:
Neural Architecture Search (NAS) for object detection is severely bottlenecked by high evaluation cost, as fully training each candidate YOLO architecture on COCO demands days of GPU time. Meanwhile, existing NAS benchmarks largely target image classification, leaving the detection community without a comparable benchmark for NAS evaluation. To address this gap, we introduce YOLO‑NAS‑Bench, the first surrogate benchmark tailored to YOLO‑style detectors. YOLO‑NAS‑Bench defines a search space spanning channel width, block depth, and operator type across both backbone and neck, covering the core modules of YOLOv8 through YOLO12. We sample 1,000 architectures via random, stratified, and Latin Hypercube strategies, train them on COCO‑mini, and build a LightGBM surrogate predictor. To sharpen the predictor in the high‑performance regime most relevant to NAS, we propose a Self‑Evolving Mechanism that progressively aligns the predictor's training distribution with the high‑performance frontier, by using the predictor itself to discover and evaluate informative architectures in each iteration. This method grows the pool to 1,500 architectures and raises the ensemble predictor's R2 from 0.770 to 0.815 and Sparse Kendall Tau from 0.694 to 0.752, demonstrating strong predictive accuracy and ranking consistency. Using the final predictor as the fitness function for evolutionary search, we discover architectures that surpass all official YOLOv8‑YOLO12 baselines at comparable latency on COCO‑mini, confirming the predictor's discriminative power for top‑performing detection architectures. The code is available at https://github.com/VDIGPKU/YOLO‑NAS‑Bench.
Authors:Tengjin Weng, Wenhao Jiang, Jingyi Wang, Ming Li, Lin Ma, Zhong Ming
Abstract:
Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision language tasks. However, their ability in low‑level visual perception, particularly in detecting fine‑grained visual discrepancies, remains underexplored and lacks systematic analysis. In this work, we introduce OddGridBench, a controllable benchmark for evaluating the visual discrepancy sensitivity of MLLMs. OddGridBench comprises over 1,400 grid‑based images, where a single element differs from all others by one or multiple visual attributes such as color, size, rotation, or position. Experiments reveal that all evaluated MLLMs, including open‑source families such as Qwen3‑VL and InternVL3.5, and proprietary systems like Gemini‑2.5‑Pro and GPT‑5, perform far below human levels in visual discrepancy detection. We further propose OddGrid‑GRPO, a reinforcement learning framework that integrates curriculum learning and distance‑aware reward. By progressively controlling the difficulty of training samples and incorporating spatial proximity constraints into the reward design, OddGrid‑GRPO significantly enhances the model's fine‑grained visual discrimination ability. We hope OddGridBench and OddGrid‑GRPO will lay the groundwork for advancing perceptual grounding and visual discrepancy sensitivity in multimodal intelligence. Code and dataset are available at https://wwwtttjjj.github.io/OddGridBench/.
Authors:Aodi Wu, Jianhong Zuo, Zeyuan Zhao, Xubo Luo, Ruisuo Wang, Xue Wan
Abstract:
Autonomous space operations such as on‑orbit servicing and active debris removal demand robust part‑level semantic understanding and precise relative navigation of target spacecraft, yet collecting large‑scale real data in orbit remains impractical due to cost and access constraints. Existing synthetic datasets, moreover, suffer from limited target diversity, single‑modality sensing, and incomplete ground‑truth annotations. We present SpaceSense‑Bench, a large‑scale multi‑modal benchmark for spacecraft perception encompassing 136~satellite models with approximately 70~GB of data. Each frame provides time‑synchronized 1024×1024 RGB images, millimeter‑precision depth maps, and 256‑beam LiDAR point clouds, together with dense 7‑class part‑level semantic labels at both the pixel and point level as well as accurate 6‑DoF pose ground truth. The dataset is generated through a high‑fidelity space simulation built in Unreal Engine~5 and a fully automated pipeline covering data acquisition, multi‑stage quality control, and conversion to mainstream formats. We benchmark five representative tasks (object detection, 2D semantic segmentation, RGB‑‑LiDAR fusion‑based 3D point cloud segmentation, monocular depth estimation, and orientation estimation) and identify two key findings: (i)~perceiving small‑scale components (\emphe.g., thrusters and omni‑antennas) and generalizing to entirely unseen spacecraft in a zero‑shot setting remain critical bottlenecks for current methods, and (ii)~scaling up the number of training satellites yields substantial performance gains on novel targets, underscoring the value of large‑scale, diverse datasets for space perception research. The dataset, code, and toolkit are publicly available at https://github.com/wuaodi/SpaceSense‑Bench.
Authors:Jiagao Hu, Yuxuan Chen, Fuhao Li, Zepeng Wang, Fei Wang, Daiguo Zhou, Jian Luan
Abstract:
Removing objects from videos remains difficult in the presence of real‑world imperfections such as shadows, abrupt motion, and defective masks. Existing diffusion‑based video inpainting models often struggle to maintain temporal stability and visual consistency under these challenges. We propose Stable Video Object Removal (SVOR), a robust framework that achieves shadow‑free, flicker‑free, and mask‑defect‑tolerant removal through three key designs: (1) Mask Union for Stable Erasure (MUSE), a windowed union strategy applied during temporal mask downsampling to preserve all target regions observed within each window, effectively handling abrupt motion and reducing missed removals; (2) Denoising‑Aware Segmentation (DA‑Seg), a lightweight segmentation head on a decoupled side branch equipped with Denoising‑Aware AdaLN and trained with mask degradation to provide an internal diffusion‑aware localization prior without affecting content generation; and (3) Curriculum Two‑Stage Training: where Stage I performs self‑supervised pretraining on unpaired real‑background videos with online random masks to learn realistic background and temporal priors, and Stage II refines on synthetic pairs using mask degradation and side‑effect‑weighted losses, jointly removing objects and their associated shadows/reflections while improving cross‑domain robustness. Extensive experiments show that SVOR attains new state‑of‑the‑art results across multiple datasets and degraded‑mask benchmarks, advancing video object removal from ideal settings toward real‑world applications. Project page: https://xiaomi‑research.github.io/svor/.
Authors:Jiaqi Liu, Zhizhong Han
Abstract:
3D Gaussian splatting (3DGS) has become a vital tool for learning a radiance field from multiple posed images. Although 3DGS shows great advantages over NeRF in terms of rendering quality and efficiency, it remains a research challenge to further improve the efficiency of learning 3D Gaussians. To overcome this challenge, we propose novel training strategies and losses to shorten each Gaussian list used to render a pixel, which speeds up the splatting by involving fewer Gaussians along a ray. Specifically, we shrink the size of each Gaussian by resetting their scales regularly, encouraging smaller Gaussians to cover fewer nearby pixels, which shortens the Gaussian lists of pixels. Additionally, we introduce an entropy constraint on the alpha blending procedure to sharpen the weight distribution of Gaussians along each ray, which drives dominant weights larger while making minor weights smaller. As a result, each Gaussian becomes more focused on the pixels where it is dominant, which reduces its impact on nearby pixels, leading to even shorter Gaussian lists. Eventually, we integrate our method into a rendering resolution scheduler which further improves efficiency through progressive resolution increase. We evaluate our method by comparing it with state‑of‑the‑art methods on widely used benchmarks. Our results show significant advantages over others in efficiency without sacrificing rendering quality.
Authors:Junhao Cai, Deyu Zeng, Junhao Pang, Lini Li, Zongze Wu, Xiaopin Zhong
Abstract:
Current text‑to‑3D generation methods excel in natural scenes but struggle with industrial applications due to two critical limitations: domain adaptation challenges where conventional LoRA fusion causes knowledge interference across categories, and geometric reasoning deficiencies where pairwise consistency constraints fail to capture higher‑order structural dependencies essential for precision manufacturing. We propose a novel framework named ForgeDreamer addressing both challenges through two key innovations. First, we introduce a Multi‑Expert LoRA Ensemble mechanism that consolidates multiple category‑specific LoRA models into a unified representation, achieving superior cross‑category generalization while eliminating knowledge interference. Second, building on enhanced semantic understanding, we develop a Cross‑View Hypergraph Geometric Enhancement approach that captures structural dependencies spanning multiple viewpoints simultaneously. These components work synergistically improved semantic understanding, enables more effective geometric reasoning, while hypergraph modeling ensures manufacturing‑level consistency. Extensive experiments on a custom industrial dataset demonstrate superior semantic generalization and enhanced geometric fidelity compared to state‑of‑the‑art approaches. Code is available at https://github.com/Junhaocai27/ForgeDreamer
Authors:Mingkun Zhang, Wangtian Shen, Fan Zhang, Haijian Qin, Zihao Pei, Ziyang Meng
Abstract:
Visual navigation requires agents to reach goals in complex environments through perception and planning. World models address this task by simulating action‑conditioned state transitions to predict future observations. Current navigation world models typically learn state evolution under actions within the compressed latent space of a Variational Autoencoder, where spatial compression often discards fine‑grained structural information and hinders precise control. To better understand the propagation characteristics of different representations, we conduct a linear dynamics probe and observe that dense DINOv2 features exhibit stronger linear predictability for action‑conditioned transitions. Motivated by this observation, we propose the Representation Autoencoder‑based Navigation World Model (RAE‑NWM), which models navigation dynamics in a dense visual representation space. We employ a Conditional Diffusion Transformer with Decoupled Diffusion Transformer head (CDiT‑DH) to model continuous transitions, and introduce a separate time‑driven gating module for dynamics conditioning to regulate action injection strength during generation. Extensive evaluations show that modeling sequential rollouts in this space improves structural stability and action accuracy, benefiting downstream planning and navigation.
Authors:Kunyu Tan, Mingjian Liang
Abstract:
RGB‑Thermal (RGB‑T) semantic segmentation is essential for robotic systems operating in low‑light or dark environments. However, traditional approaches often overemphasize modality balance, resulting in limited robustness and severe performance degradation when sensor signals are partially missing. Recent advances such as cross‑modal knowledge distillation and modality‑adaptive fine‑tuning attempt to enhance cross‑modal interaction, but they typically decouple modality fusion and modality adaptation, requiring multi‑stage training with frozen models or teacher‑student frameworks. We present RTFDNet, a three‑branch encoder‑decoder that unifies fusion and decoupling for robust RGB‑T segmentation. Synergistic Feature Fusion (SFF) performs channel‑wise gated exchange and lightweight spatial attention to inject complementary cues. Cross‑Modal Decouple Regularization (CMDR) isolates modality‑specific components from the fused representation and supervises unimodal decoders via stop‑gradient targets. Region Decouple Regularization (RDR) enforces class‑selective prediction consistency in confident regions while blocking gradients to the fusion branch. This feedback loop strengthens unimodal paths without degrading the fused stream, enabling efficient standalone inference at test time. Extensive experiments demonstrate the effectiveness of RTFDNet, showing consistent performance across varying modality conditions. Our implementation will be released to facilitate further research. Our source code are publicly available at https://github.com/curapima/RTFDNet.
Authors:Zhongchen Zhao, Qi Xie, Keyu Huang, Lei Zhang, Deyu Meng, Zongben Xu
Abstract:
Rotation equivariance constitutes one of the most general and crucial structural priors for visual data, yet it remains notably absent from current Mamba‑based vision architectures. Despite the success of Mamba in natural language processing and its growing adoption in computer vision, existing visual Mamba models fail to account for rotational symmetry in their design. This omission renders them inherently sensitive to image rotations, thereby constraining their robustness and cross‑task generalization. To address this limitation, we incorporate rotation symmetry, a universal and fundamental geometric prior in images, into Mamba‑based architectures. Specifically, we introduce EQ‑VMamba, the first rotation equivariant visual Mamba architecture for vision tasks. The core components of EQ‑VMamba include a carefully designed rotation equivariant cross‑scan strategy and group Mamba blocks. Moreover, we provide a rigorous theoretical analysis of the intrinsic equivariance error, demonstrating that the proposed architecture enforces end‑to‑end rotation equivariance throughout the network. Extensive experiments across multiple benchmarks ‑‑ including high‑level image classification, mid‑level semantic segmentation, and low‑level image super‑resolution ‑‑ demonstrate that EQ‑VMamba consistently improves rotation robustness and achieves superior or competitive performance compared to non‑equivariant baselines, while requiring approximately 50% fewer parameters. These results indicate that embedding rotation equivariance not only effectively bolsters the robustness of visual Mamba models against rotation transformations, but also enhances overall performance with significantly improved parameter efficiency. Code is available at https://github.com/zhongchenzhao/EQ‑VMamba.
Authors:Junjie Yin, Jiaju Li, Hanfa Xing
Abstract:
Diffusion‑based image super‑resolution (ISR) has shown strong potential, but it still struggles in real‑world scenarios where degradations are unknown and spatially non‑uniform, often resulting in lost details or visual artifacts. To address this challenge, we propose a novel super‑resolution diffusion model, QUSR, which integrates a Quality‑Aware Prior (QAP) with an Uncertainty‑Guided Noise Generation (UNG) module. The UNG module adaptively adjusts the noise injection intensity, applying stronger perturbations to high‑uncertainty regions (e.g., edges and textures) to reconstruct complex details, while minimizing noise in low‑uncertainty regions (e.g., flat areas) to preserve original information. Concurrently, the QAP leverages an advanced Multimodal Large Language Model (MLLM) to generate reliable quality descriptions, providing an effective and interpretable quality prior for the restoration process. Experimental results confirm that QUSR can produce high‑fidelity and high‑realism images in real‑world scenarios. The source code is available at https://github.com/oTvTog/QUSR.
Authors:Zixuan Wang, Ziqin Zhou, Feng Chen, Duo Peng, Yixin Hu, Changsheng Li, Yinjie Lei
Abstract:
Compositional video generation aims to synthesize multiple instances with diverse appearance and motion. However, current approaches mainly focus on binding semantics, neglecting to understand diverse motion categories specified in prompts. In this paper, we propose a motion factorization framework that decomposes complex motion into three primary categories: motionlessness, rigid motion, and non‑rigid motion. Specifically, our framework follows a planning before generation paradigm. (1) During planning, we reason about motion laws on the motion graph to obtain frame‑wise changes in the shape and position of each instance. This alleviates semantic ambiguities in the user prompt by organizing it into a structured representation of instances and their interactions. (2) During generation, we modulate the synthesis of distinct motion categories in a disentangled manner. Conditioned on the motion cues, guidance branches stabilize appearance in motionless regions, preserve rigid‑body geometry, and regularize local non‑rigid deformations. Crucially, our two modules are model‑agnostic, which can be seamlessly incorporated into various diffusion model architectures. Extensive experiments demonstrate that our framework achieves impressive performance in motion synthesis on real‑world benchmarks. Code is available at https://github.com/ZixuanWang0525/MF‑CVG.
Authors:Chenran Zhang, Ruiqi Wu, Tao Zhou, Yi Zhou
Abstract:
Medical vision‑language pretraining (VLP) models have recently been investigated for their generalization to diverse downstream tasks. However, current medical VLP methods typically force the model to learn simple and complex concepts simultaneously. This anti‑cognitive process leads to suboptimal feature representations, especially under distribution shift. To address this limitation, we propose a Knowledge‑driven Cognitive Orchestration for Medical VLP (MedKCO) that involves both the ordering of the pretraining data and the learning objective of vision‑language contrast. Specifically, we design a two level curriculum by incorporating diagnostic sensitivity and intra‑class sample representativeness for the ordering of the pretraining data. Moreover, considering the inter‑class similarity of medical images, we introduce a self‑paced asymmetric contrastive loss to dynamically adjust the participation of the pretraining objective. We evaluate the proposed pretraining method on three medical imaging scenarios in multiple vision‑language downstream tasks, and compare it with several curriculum learning methods. Extensive experiments show that our method significantly surpasses all baselines. https://github.com/Mr‑Talon/MedKCO.
Authors:Zixuan Wang, Yixin Hu, Haolan Wang, Feng Chen, Yan Liu, Wen Li, Yinjie Lei
Abstract:
Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real‑world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diffusion models. Current approaches leverage commonsense reasoning capability of large language models to embed physical concepts into prompts. However, generation models often render physical phenomena as a single moment defined by prompts, due to the lack of conditioning mechanisms for modeling causal progression. In this paper, we view PPVG as generating a sequence of causally connected and dynamically evolving events. To realize this paradigm, we design two key modules: (1) Physics‑driven Event Chain Reasoning. This module decomposes the physical phenomena described in prompts into multiple elementary event units, leveraging chain‑of‑thought reasoning. To mitigate causal ambiguity, we embed physical formulas as constraints to impose deterministic causal dependencies during reasoning. (2) Transition‑aware Cross‑modal Prompting (TCP). To maintain continuity between events, this module transforms causal event units into temporally aligned vision‑language prompts. It summarizes discrete event descriptions to obtain causally consistent narratives, while progressively synthesizing visual keyframes of individual events by interactive editing. Comprehensive experiments on PhyGenBench and VideoPhy benchmarks demonstrate that our framework achieves superior performance in generating physically plausible videos across diverse physical domains. Code is available at https://github.com/ZixuanWang0525/CoECT.
Authors:Lixiang Lin, Siyuan Jin, Jinshan Zhang
Abstract:
Lip synchronization and audio‑visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, virtual avatars, and telepresence. Despite recent progress, most existing methods for lip synchronization and audio‑visual editing depend on supervised fine‑tuning of pre‑trained models, leading to considerable computational overhead and data requirements. In this paper, we present OmniEdit, a training‑free framework designed for both lip synchronization and audio‑visual editing. Our approach reformulates the editing paradigm by substituting the edit sequence in FlowEdit with the target sequence, yielding an unbiased estimation of the desired output. Moreover, by removing stochastic elements from the generation process, we establish a smooth and stable editing trajectory. Extensive experimental results validate the effectiveness and robustness of the proposed framework. Code is available at https://github.com/l1346792580123/OmniEdit.
Authors:Pranav Mantini, Shishir K. Shah
Abstract:
Recent advances in vision‑language models (VLMs) have demonstrated remarkable zero‑shot capabilities, yet adapting these models to specialized domains remains a significant challenge. Building on recent theoretical insights suggesting that independently trained VLMs are related by a canonical transformation, we extend this understanding to the concept of domains. We hypothesize that image features across disparate domains are related by a canonicalized geometric transformation that can be recovered using a small set of anchors. Few‑shot classification provides a natural setting for this alignment, as the limited labeled samples serve as the anchors required to estimate this transformation. Motivated by this hypothesis, we introduce BiCLIP, a framework that applies a targeted transformation to multimodal features to enhance cross‑modal alignment. Our approach is characterized by its extreme simplicity and low parameter footprint. Extensive evaluations across 11 standard benchmarks, including EuroSAT, DTD, and FGVCAircraft, demonstrate that BiCLIP consistently achieves state‑of‑the‑art results. Furthermore, we provide empirical verification of existing geometric findings by analyzing the orthogonality and angular distribution of the learned transformations, confirming that structured alignment is the key to robust domain adaptation. Code is available at https://github.com/QuantitativeImagingLaboratory/BilinearCLIP
Authors:Brian Isett, Rebekah Dadey, Aofei Li, Ryan C. Augustin, Kate Smith, Aatur D. Singhi, Qiangqiang Gu, Riyue Bao
Abstract:
Accurate localization of tumor regions from hematoxylin and eosin‑stained whole‑slide images is fundamental for translational research including spatial analysis, molecular profiling, and tissue architecture investigation. However, deep learning‑based tumor detection trained within specific cancers may exhibit reduced robustness when applied across different tumor types. We investigated whether balanced training across cancers at modest scale can achieve high performance and generalize to unseen tumor types. A multi‑cancer tumor localization model (MuCTaL) was trained on 79,984 non‑overlapping tiles from four cancers (melanoma, hepatocellular carcinoma, colorectal cancer, and non‑small cell lung cancer) using transfer learning with DenseNet169. The model achieved a tile‑level ROC‑AUC of 0.97 in validation data from the four training cancers, and 0.71 on an independent pancreatic ductal adenocarcinoma cohort. A scalable inference workflow was built to generate spatial tumor probability heatmaps compatible with existing digital pathology tools. Code and models are publicly available at https://github.com/AivaraX‑AI/MuCTaL.
Authors:Soumik Mukhopadhyay, Prateksha Udhayanan, Abhinav Shrivastava
Abstract:
Diffusion models degrade images through noise, and reversing this process reveals an information hierarchy across timesteps. Scale‑space theory exhibits a similar hierarchy via low‑pass filtering. We formalize this connection and show that highly noisy diffusion states contain no more information than small, downsampled images ‑ raising the question of why they must be processed at full resolution. To address this, we fuse scale spaces into the diffusion process by formulating a family of diffusion models with generalized linear degradations and practical implementations. Using downsampling as the degradation yields our proposed Scale Space Diffusion. To support Scale Space Diffusion, we introduce Flexi‑UNet, a UNet variant that performs resolution‑preserving and resolution‑increasing denoising using only the necessary parts of the network. We evaluate our framework on CelebA and ImageNet and analyze its scaling behavior across resolutions and network depths. Our project website ( https://prateksha.github.io/projects/scale‑space‑diffusion/ ) is available publicly.
Authors:Haoyang Li, Liang Wang, Siyu Zhou, Jiacheng Sun, Jing Jiang, Chao Wang, Guodong Long, Yan Peng
Abstract:
CLIP‑based prompt tuning enables pretrained Vision‑Language Models (VLMs) to efficiently adapt to downstream tasks. Although existing studies have made significant progress, they pay limited attention to changes in the internal attention representations of VLMs during the tuning process. In this paper, we attribute the failure modes of prompt tuning predictions to shifts in foreground attention of the visual encoder, and propose Foreground View‑Guided Prompt Tuning (FVG‑PT), an adaptive plug‑and‑play foreground attention guidance module, to alleviate the shifts. Concretely, FVG‑PT introduces a learnable Foreground Reliability Gate to automatically enhance the foreground view quality, applies a Foreground Distillation Compensation module to guide visual attention toward the foreground, and further introduces a Prior Calibration module to mitigate generalization degradation caused by excessive focus on the foreground. Experiments on multiple backbone models and datasets show the effectiveness and compatibility of FVG‑PT. Codes are available at: https://github.com/JREion/FVG‑PT
Authors:Kai Zou, Dian Zheng, Hongbo Liu, Tiankai Hang, Bin Liu, Nenghai Yu
Abstract:
Autoregressive (AR) diffusion offers a promising framework for generating videos of theoretically infinite length. However, a major challenge is maintaining temporal continuity while preventing the progressive quality degradation caused by error accumulation. To ensure continuity, existing methods typically condition on highly denoised contexts; yet, this practice propagates prediction errors with high certainty, thereby exacerbating degradation. In this paper, we argue that a highly clean context is unnecessary. Drawing inspiration from bidirectional diffusion models, which denoise frames at a shared noise level while maintaining coherence, we propose that conditioning on context at the same noise level as the current block provides sufficient signal for temporal consistency while effectively mitigating error propagation. Building on this insight, we propose HiAR, a hierarchical denoising framework that reverses the conventional generation order: instead of completing each block sequentially, it performs causal generation across all blocks at every denoising step, so that each block is always conditioned on context at the same noise level. This hierarchy naturally admits pipelined parallel inference, yielding a 1.8 wall‑clock speedup in our 4‑step setting. We further observe that self‑rollout distillation under this paradigm amplifies a low‑motion shortcut inherent to the mode‑seeking reverse‑KL objective. To counteract this, we introduce a forward‑KL regulariser in bidirectional‑attention mode, which preserves motion diversity for causal inference without interfering with the distillation loss. On VBench (20s generation), HiAR achieves the best overall score and the lowest temporal drift among all compared methods.
Authors:Jordi Muñoz Vicente
Abstract:
Recent advancements in 3D Gaussian Splatting (3DGS) have shifted the focus toward balancing reconstruction fidelity with computational efficiency. In this work, we propose ImprovedGS+, a high‑performance, low‑level reinvention of the ImprovedGS strategy, implemented natively within the LichtFeld‑Studio framework. By transitioning from high‑level Python logic to hardware‑optimized C++/CUDA kernels, we achieve a significant reduction in host‑device synchronization and training latency. Our implementation introduces a Long‑Axis‑Split (LAS) CUDA kernel, custom Laplacian‑based importance kernels with Non‑Maximum Suppression (NMS) for edge scores, and an adaptive Exponential Scale Scheduler. Experimental results on the Mip‑NeRF360 dataset demonstrate that ImprovedGS+ establishes a new Pareto‑optimal front for scene reconstruction. Our 1M‑budget variant outperforms the state‑of‑the‑art MCMC baseline by achieving a 26.8% reduction in training time (saving 17 minutes per session) and utilizing 13.3% fewer Gaussians while maintaining superior visual quality. Furthermore, our full variant demonstrates a 1.28 dB PSNR increase over the ADC baseline with a 38.4% reduction in parametric complexity. These results validate ImprovedGS+ as a scalable, high‑speed solution that upholds the core pillars of Speed, Quality, and Usability within the LichtFeld‑Studio ecosystem.
Authors:Chao Wang, Zijin Yang, Yaofei Wang, Yuang Qi, Weiming Zhang, Nenghai Yu, Kejiang Chen
Abstract:
Recent advancements in video generation technologies have been significant, resulting in their widespread application across multiple domains. However, concerns have been mounting over the potential misuse of generated content. Tracing the origin of generated videos has become crucial to mitigate potential misuse and identify responsible parties. Existing video attribution methods require additional operations or the training of source attribution models, which may degrade video quality or necessitate large amounts of training samples. To address these challenges, we define for the first time the "few‑shot training‑free generated video attribution" task and propose SWIFT, which is tightly integrated with the temporal characteristics of the video. By leveraging the "Pixel Frames(many) to Latent Frame(one)" temporal mapping within each video chunk, SWIFT applies a fixed‑length sliding window to perform two distinct reconstructions: normal and corrupted. The variation in the losses between two reconstructions is then used as an attribution signal. We conducted an extensive evaluation of five state‑of‑the‑art (SOTA) video generation models. Experimental results show that SWIFT achieves over 90% average attribution accuracy with merely 20 video samples across all models and even enables zero‑shot attribution for HunyuanVideo, EasyAnimate, and Wan2.2. Our source code is available at https://github.com/wangchao0708/SWIFT.
Authors:Yongzhi Lin, Kai Luo, Yuanfan Zheng, Hao Shi, Mengfei Duan, Yang Liu, Kailun Yang
Abstract:
Understanding dynamic 3D environments in a spatially continuous and temporally consistent manner is fundamental for robotics and autonomous driving. While recent advances in occupancy prediction provide a unified representation of scene geometry and semantics, progress in 4D panoptic occupancy tracking remains limited by the lack of benchmarks that support surround‑view fisheye sensing, long temporal sequences, and instance‑level voxel tracking. To address this gap, we present OccTrack360, a new benchmark for 4D panoptic occupancy tracking from surround‑view fisheye cameras. OccTrack360 provides substantially longer and more diverse sequences (174~2234 frames) than prior benchmarks, together with principled voxel visibility annotations, including an all‑direction occlusion mask and an MEI‑based fisheye field‑of‑view mask. To establish a strong fisheye‑oriented baseline, we further propose Focus on Sphere Occ (FoSOcc), a framework that addresses two core challenges in fisheye occupancy tracking: distorted spherical projection and inaccurate voxel‑space localization. FoSOcc includes a Center Focusing Module (CFM) to enhance instance‑aware spatial localization through supervised focus guidance, and a Spherical Lift Module (SLM) that extends perspective lifting to fisheye imaging under the Unified Projection Model. Extensive experiments on Occ3D‑Waymo and OccTrack360 show that our method improves occupancy tracking quality with notable gains on geometrically regular categories, and establishes a strong baseline for future research on surround‑view fisheye 4D occupancy tracking. The benchmark and source code will be made publicly available at https://github.com/YouthZest‑Lin/OccTrack360.
Authors:Zhe Yang, Guoqiang Zhao, Sheng Wu, Kai Luo, Kailun Yang
Abstract:
Omnidirectional images are increasingly used in robotics and vision due to their wide field of view. However, extending 3D Gaussian Splatting (3DGS) to panoramic camera models remains challenging, as existing formulations are designed for perspective projections and naive adaptations often introduce distortion and geometric inconsistencies. We present Spherical‑GOF, an omnidirectional Gaussian rendering framework built upon Gaussian Opacity Fields (GOF). Unlike projection‑based rasterization, Spherical‑GOF performs GOF ray sampling directly on the unit sphere in spherical ray space, enabling consistent ray‑Gaussian interactions for panoramic rendering. To make the spherical ray casting efficient and robust, we derive a conservative spherical bounding rule for fast ray‑Gaussian culling and introduce a spherical filtering scheme that adapts Gaussian footprints to distortion‑varying panoramic pixel sampling. Extensive experiments on standard panoramic benchmarks (OmniBlender and OmniPhotos) demonstrate competitive photometric quality and substantially improved geometric consistency. Compared with the strongest baseline, Spherical‑GOF reduces depth reprojection error by 57% and improves cycle inlier ratio by 21%. Qualitative results show cleaner depth and more coherent normal maps, with strong robustness to global panorama rotations. We further validate generalization on OmniRob, a real‑world robotic omnidirectional dataset introduced in this work, featuring UAV and quadruped platforms. The source code and the OmniRob dataset will be released at https://github.com/1170632760/Spherical‑GOF.
Authors:Yu Yang, Yue Liao, Jianbiao Mei, Baisen Wang, Xuemeng Yang, Licheng Wen, Jiangning Zhang, Xiangtai Li, Liang Lv, Hanlin Chen, Botian Shi, Yong Liu, Shuicheng Yan, Gim Hee Lee
Abstract:
Long‑horizon action‑conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended horizons, requiring procedural ordering, persistent action execution, and scene consistency beyond conventional TI2V's short‑term fidelity. Existing single‑shot video generation models typically operate in an open‑loop manner, leading to incomplete action execution, hallucinated motions, and temporal drift. To address this, we propose SPIRAL, a closed‑loop framework that performs sequential planning and iterative reflection for action‑conditioned long‑horizon video generation. Specifically, SPIRAL instantiates a think‑act‑reflect process: a PlanAgent decomposes high‑level goals into sub‑actions, which condition a VideoGenerator to synthesize each segment alongside a memory context, while a CriticAgent evaluates intermediate video segments to provide corrective feedback for iterative refinement. This closed‑loop design further supports self‑evolution by utilizing PlanAgent‑proposed actions and CriticAgent‑derived rewards for GRPO‑based post‑training to enhance the video generator's long‑horizon consistency. Moreover, we introduce ActVideoGen‑Dataset for task‑specific training, and establish ActVideoGen‑Bench as a dedicated evaluation suite for measuring action quality and temporal coherence. Experiments across multiple TI2V backbones alongside the self‑evolving strategy show consistent gains on ActVideoGen‑Bench and VBench, demonstrating the effectiveness of SPIRAL.
Authors:Yijie Zhu, Jie He, Rui Shao, Kaishen Yuan, Tao Tan, Xiaochen Yuan, Zitong Yu
Abstract:
Recent vision‑language‑action (VLA) models have significantly advanced robotic manipulation by unifying perception, reasoning, and control. To achieve such integration, recent studies adopt a predictive paradigm that models future visual states or world knowledge to guide action generation. However, these models emphasize forecasting outcomes rather than reasoning about the underlying process of change, which is essential for determining how to act. To address this, we propose ΔVLA, a prior‑guided framework that models world‑knowledge variations relative to an explicit current‑world knowledge prior for action generation, rather than regressing absolute future world states. Specifically, 1) to construct the current world knowledge prior, we propose the Prior‑Guided WorldKnowledge Extractor (PWKE). It extracts manipulable regions, spatial relations, and semantic cues from the visual input, guided by auxiliary heads and prior pseudo labels, thus reducing redundancy. 2) Building upon this, to represent how world knowledge evolves under actions, we introduce the Latent World Variation Quantization (LWVQ). It learns a discrete latent space via a VQ‑VAE objective to encode world knowledge variations, shifting prediction from full modalities to compact latent. 3)Moreover, to mitigate interference during variation modeling, we design the Conditional Variation Attention (CV‑Atten), whichpromotes disentangled learning and preserves the independence of knowledge representations. Extensive experiments on both simulated benchmarks and real‑world robotic tasks demonstrate ΔVLA achieves state‑of‑the‑art performance while improving efficiency. Code and real‑world execution videos are available at https://github.com/JiuTian‑VL/DeltaVLA.
Authors:Deniz Kizaroğlu, Ülku Tuncer Küçüktas, Emre Çakmakyurdu, Alptekin Temizel
Abstract:
Few‑shot adaptation of vision‑language models (VLMs) like CLIP typically relies on learning textual prompts matched to global image embeddings. Recent works extend this paradigm by incorporating local image‑text alignment to capture fine‑grained visual cues, yet these approaches often select local regions independently for each prompt, leading to redundant local feature usage and prompt overlap. We propose SOT‑GLP, which introduces a shared sparse patch support and balanced optimal transport allocation to explicitly partition salient visual regions among class‑specific local prompts while preserving global alignment. Our method learns shared global prompts and class‑specific local prompts. The global branch maintains standard image‑text matching for robust category‑level alignment. The local branch constructs a class‑conditioned sparse patch set using V‑V attention and aligns it to multiple class‑specific prompts via balanced entropic optimal transport, yielding a soft partition of patches that prevents prompt overlap and collapse. We evaluate our method on two complementary objectives: (i) few‑shot classification accuracy on 11 standard benchmarks and (ii) out‑of‑distribution (OOD) detection. On the standard 11‑dataset benchmark with 16‑shot ViT‑B/16, SOT‑GLP achieves 85.1% average accuracy, outperforming prior prompt‑learning methods. We identify a distinct accuracy‑robustness trade‑off in prompt learning: while learnable projections optimize in‑distribution fit, they alter the foundational feature space. We demonstrate that a projection‑free local alignment preserves the native geometry of the CLIP manifold, yielding state‑of‑the‑art OOD detection performance (94.2% AUC) that surpasses fully adapted models. Implementation available at: https://github.com/Deniz2304988/SOT‑GLP
Authors:Mina Jamshidi Idaji, Julius Hense, Tom Neuhäuser, Augustin Krause, Yanqing Luo, Oliver Eberle, Thomas Schnake, Laure Ciernik, Farnoush Rezaei Jafari, Reza Vahidimajd, Jonas Dippel, Christoph Walz, Frederick Klauschen, Andreas Mock, Klaus-Robert Müller
Abstract:
Multiple instance learning (MIL) has enabled substantial progress in computational histopathology, where a large amount of patches from gigapixel whole slide images are aggregated into slide‑level predictions. Heatmaps are widely used to validate MIL models and to discover tissue biomarkers. Yet, the validity of these heatmaps has barely been investigated. In this work, we introduce a general framework for evaluating the quality of MIL heatmaps without requiring additional labels. We conduct a large‑scale benchmark experiment to assess six explanation methods across histopathology task types (classification, regression, survival), MIL model architectures (Attention‑, Transformer‑, Mamba‑based), and patch encoder backbones (UNI2, Virchow2). Our results show that explanation quality mostly depends on MIL model architecture and task type, with perturbation ("Single"), layer‑wise relevance propagation (LRP), and integrated gradients (IG) consistently outperforming attention‑based and gradient‑based saliency heatmaps, which often fail to reflect model decision mechanisms. We further demonstrate the advanced capabilities of the best‑performing explanation methods: (i) We provide a proof‑of‑concept that MIL heatmaps of a bulk gene expression prediction model can be correlated with spatial transcriptomics for biological validation, and (ii) showcase the discovery of distinct model strategies for predicting human papillomavirus (HPV) infection from head and neck cancer slides. Our work highlights the importance of validating MIL heatmaps and establishes that improved explainability can enable more reliable model validation and yield biological insights, making a case for a broader adoption of explainable AI in digital pathology. Our code is provided in a public GitHub repository: https://github.com/bifold‑pathomics/xMIL/tree/xmil‑journal
Authors:Junxian Li, Tu Lan, Haozhen Tan, Yan Meng, Haojin Zhu
Abstract:
Modern vision‑language‑model (VLM) based graphical user interface (GUI) agents are expected not only to execute actions accurately but also to respond to user instructions with low latency. While existing research on GUI‑agent security mainly focuses on manipulating action correctness, the security risks related to response efficiency remain largely unexplored. In this paper, we introduce SlowBA, a novel backdoor attack that targets the responsiveness of VLM‑based GUI agents. The key idea is to manipulate response latency by inducing excessively long reasoning chains under specific trigger patterns. To achieve this, we propose a two‑stage reward‑level backdoor injection (RBI) strategy that first aligns the long‑response format and then learns trigger‑aware activation through reinforcement learning. In addition, we design realistic pop‑up windows as triggers that naturally appear in GUI environments, improving the stealthiness of the attack. Extensive experiments across multiple datasets and baselines demonstrate that SlowBA can significantly increase response length and latency while largely preserving task accuracy. The attack remains effective even with a small poisoning ratio and under several defense settings. These findings reveal a previously overlooked security vulnerability in GUI agents and highlight the need for defenses that consider both action correctness and response efficiency. Code can be found in https://github.com/tu‑tuing/SlowBA.
Authors:Shin Dong-Yeon, Kim Jun-Seong, Kwon Byung-Ki, Tae-Hyun Oh
Abstract:
Radiance of real‑world scenes typically spans a much wider dynamic range than what standard cameras can capture. While conventional HDR methods merge alternating‑exposure frames, these approaches are inherently constrained to 2D pixel‑level alignment, often leading to ghosting artifacts and temporal inconsistency in dynamic scenes. To address these limitations, we present HDR‑NSFF, a paradigm shift from 2D‑based merging to 4D spatio‑temporal modeling. Our framework reconstructs dynamic HDR radiance fields from alternating‑exposure monocular videos by representing the scene as a continuous function of space and time, and is compatible with both neural radiance field and 4D Gaussian Splatting (4DGS) based dynamic representations. This unified end‑to‑end pipeline explicitly models HDR radiance, 3D scene flow, geometry, and tone‑mapping, ensuring physical plausibility and global coherence. We further enhance robustness by (i) extending semantic‑based optical flow with DINO features to achieve exposure‑invariant motion estimation, and (ii) incorporating a generative prior as a regularizer to compensate for limited observation in monocular captures and saturation‑induced information loss. To evaluate HDR space‑time view synthesis, we present the first real‑world HDR‑GoPro dataset specifically designed for dynamic HDR scenes. Experiments demonstrate that HDR‑NSFF recovers fine radiance details and coherent dynamics even under challenging exposure variations, thereby achieving state‑of‑the‑art performance in novel space‑time view synthesis. Project page: https://shin‑dong‑yeon.github.io/HDR‑NSFF/
Authors:Yehonatan Elisha, Oren Barkan, Noam Koenigstein
Abstract:
Vision Transformers (ViTs) often degrade under distribution shifts because they rely on spurious correlations, such as background cues, rather than semantically meaningful features. Existing regularization methods, typically relying on simple foreground‑background masks, which fail to capture the fine‑grained semantic concepts that define an object (e.g., ``long beak'' and ``wings'' for a ``bird''). As a result, these methods provide limited robustness to distribution shifts. To address this limitation, we introduce a novel finetuning framework that steers model reasoning toward concept‑level semantics. Our approach optimizes the model's internal relevance maps to align with spatially grounded concept masks. These masks are generated automatically, without manual annotation: class‑relevant concepts are first proposed using an LLM‑based, label‑free method, and then segmented using a VLM. The finetuning objective aligns relevance with these concept regions while simultaneously suppressing focus on spurious background areas. Notably, this process requires only a minimal set of images and uses half of the dataset classes. Extensive experiments on five out‑of‑distribution benchmarks demonstrate that our method improves robustness across multiple ViT‑based models. Furthermore, we show that the resulting relevance maps exhibit stronger alignment with semantic object parts, offering a scalable path toward more robust and interpretable vision models. Finally, we confirm that concept‑guided masks provide more effective supervision for model robustness than conventional segmentation maps, supporting our central hypothesis.
Authors:Lei Wang, Yang Cheng, Senmao Li, Ge Wu, Yaxing Wang, Jian Yang
Abstract:
Despite the impressive performance of diffusion models such as Stable Diffusion (SD) in image generation, their slow inference limits practical deployment. Recent works accelerate inference by distilling multi‑step diffusion into one‑step generators. To better understand the distillation mechanism, we analyze U‑Net/DiT weight changes between one‑step students and their multi‑step teacher counterparts. Our analysis reveals that changes in weight direction significantly exceed those in weight norm, highlighting it as the key factor during distillation. Motivated by this insight, we propose the Low‑rank Rotation of weight Direction (LoRaD), a parameter‑efficient adapter tailored to one‑step diffusion distillation. LoRaD is designed to model these structured directional changes using learnable low‑rank rotation matrices. We further integrate LoRaD into Variational Score Distillation (VSD), resulting in Weight Direction‑aware Distillation (WaDi)‑a novel one‑step distillation framework. WaDi achieves state‑of‑the‑art FID scores on COCO 2014 and COCO 2017 while using only approximately 10% of the trainable parameters of the U‑Net/DiT. Furthermore, the distilled one‑step model demonstrates strong versatility and scalability, generalizing well to various downstream tasks such as controllable generation, relation inversion, and high‑resolution synthesis.
Authors:Jiageng Wen, Shengjie Zhao, Bing Li, Jiafeng Huang, Kenan Ye, Hao Deng
Abstract:
Collaborative perception integrates multi‑agent perspectives to enhance the sensing range and overcome occlusion issues. While existing multimodal approaches leverage complementary sensors to improve performance, they are highly prone to failure‑‑especially when a key sensor like LiDAR is unavailable. The root cause is that feature fusion leads to semantic mismatches between single‑modality features and the downstream modules. This paper addresses this challenge for the first time in the field of collaborative perception, introducing Single‑Modality‑Operable Multimodal Collaborative Perception (SiMO). By adopting the proposed Length‑Adaptive Multi‑Modal Fusion (LAMMA), SiMO can adaptively handle remaining modal features during modal failures while maintaining consistency of the semantic space. Additionally, leveraging the innovative "Pretrain‑Align‑Fuse‑RD" training strategy, SiMO addresses the issue of modality competition‑‑generally overlooked by existing methods‑‑ensuring the independence of each individual modality branch. Experiments demonstrate that SiMO effectively aligns multimodal features while simultaneously preserving modality‑specific features, enabling it to maintain optimal performance across all individual modalities. The implementation details can be found in https://github.com/dempsey‑wen/SiMO.
Authors:Michael Kösel, Marcel Schreiber, Michael Ulrich, Claudius Gläser, Klaus Dietmayer
Abstract:
LiDAR‑based 3D object detection plays a critical role for reliable and safe autonomous driving systems. However, existing detectors often produce overly confident predictions for objects not belonging to known categories, posing significant safety risks. This is caused by so‑called out‑of‑distribution (OOD) objects, which were not part of the training data, resulting in incorrect predictions. To address this challenge, we propose ALOOD (Aligned LiDAR representations for Out‑Of‑Distribution Detection), a novel approach that incorporates language representations from a vision‑language model (VLM). By aligning the object features from the object detector to the feature space of the VLM, we can treat the detection of OOD objects as a zero‑shot classification task. We demonstrate competitive performance on the nuScenes OOD benchmark, establishing a novel approach to OOD object detection in LiDAR using language representations. The source code is available at https://github.com/uulm‑mrm/mmood3d.
Authors:Şebnem Sarıözkan, Hürkan Şahin, Olaya Álvarez-Tuñón, Erdal Kayacan
Abstract:
Conventional visual simultaneous localization and mapping (SLAM) algorithms often fail under rapid motion, low illumination, or abrupt lighting transitions due to motion blur and limited dynamic range. Event cameras mitigate these issues with high temporal resolution and high dynamic range (HDR), but their sparse, asynchronous outputs complicate feature extraction and integration with other sensors; e.g. inertial measurement units (IMUs) and standard cameras. We present Edged USLAM, a hybrid visual‑inertial system that extends Ultimate SLAM (USLAM) with an edge‑aware front‑end and a lightweight depth module. The frontend enhances event frames for robust feature tracking and nonlinear motion compensation, while the depth module provides coarse, region‑of‑interest (ROI)‑based scene depth to improve motion compensation and scale consistency. Evaluations across public benchmarks and real‑world unmanned air vehicle (UAV) flights demonstrate that performance varies significantly by scenario. For instance, event‑only methods like point‑line event‑based visual‑inertial odometry (PL‑EVIO) or learning‑based pipelines such as deep event‑based visual odometry (DEVO) excel in highly aggressive or extreme HDR conditions. In contrast, Edged USLAM provides superior stability and minimal drift in slow or structured trajectories, ensuring consistently accurate localization on real flights under challenging illumination. These findings highlight the complementary strengths of event‑only, learning‑based, and hybrid approaches, while positioning Edged USLAM as a robust solution for diverse aerial navigation tasks.
Authors:Hunor Laczkó, Libang Jia, Loc-Phat Truong, Diego Hernández, Sergio Escalera, Jordi Gonzalez, Meysam Madadi
Abstract:
Existing 4D human datasets fall short for fashion‑specific research, lacking either realistic garment dynamics or task‑specific annotations. Synthetic datasets suffer from a realism gap, whereas real‑world captures lack the detailed annotations and paired data required for virtual try‑on (VTON) and size estimation tasks. To bridge this gap, we introduce MV‑Fashion, a large‑scale, multi‑view video dataset engineered for domain‑specific fashion analysis. MV‑Fashion features 3,273 sequences (72.5 million frames) from 80 diverse subjects wearing 3‑10 outfits each. It is designed to capture complex, real‑world garment dynamics, including multiple layers and varied styling (e.g. rolled sleeves, tucked shirt). A core contribution is a rich data representation that includes pixel‑level semantic annotations, ground‑truth material properties like elasticity, and 3D point clouds. Crucially for VTON applications, MV‑Fashion provides paired data: multi‑view synchronized captures of worn garments alongside their corresponding flat, catalogue images. We leverage this dataset to establish baselines for fashion‑centric tasks, including virtual try‑on, clothing size estimation, and novel view synthesis. The dataset is available at https://hunorlaczko.github.io/MV‑Fashion .
Authors:Chengchao Shen
Abstract:
Large vision transformers present impressive scalability, as their performance can be well improved with increased model capacity. Nevertheless, their cumbersome parameters results in exorbitant computational and memory demands. By analyzing prevalent transformer structures, we find that multilayer perceptron (MLP) modules constitute the largest share of the model's parameters. In this paper, we propose an Adaptive MLP Pruning (AMP) method to substantially reduce the parameters of large vision transformers without obvious performance degradation. First, we adopt Taylor based method to evaluate neuron importance of MLP. However, the importance computation using one‑hot cross entropy loss ignores the potential predictions on other categories, thus degrading the quality of the evaluated importance scores. To address this issue, we introduce label‑free information entropy criterion to fully model the predictions of the original model for more accurate importance evaluation. Second, we rank the hidden neurons of MLP by the above importance scores and apply binary search algorithm to adaptively prune the ranked neurons according to the redundancy of different MLP modules, thereby avoiding the predefined compression ratio. Experimental results on several state‑of‑the‑art large vision transformers, including CLIP and DINOv2, demonstrate that our method achieves roughly 40% parameter and FLOPs reduction in a near lossless manner. Moreover, when the models are not finetuned after pruning, our method outperforms other pruning methods by significantly large margin. The source code and trained weights are available at https://github.com/visresearch/AMP.
Authors:Bryce Grant, Aryeh Rothenberg, Atri Banerjee, Peng Wang
Abstract:
Localizing objects and parts from natural language in 3D space is essential for robotics, AR, and embodied AI, yet existing methods face a trade‑off between the accuracy and geometric consistency of per‑scene optimization and the efficiency of feed‑forward inference. We present TrianguLang, a feed‑forward framework for 3D localization that requires no camera calibration at inference. Unlike prior methods that treat views independently, we introduce Geometry‑Aware Semantic Attention (GASA), which utilizes predicted geometry to gate cross‑view feature correspondence, suppressing semantically‑plausible but geometrically‑inconsistent matches without requiring ground‑truth poses. Validated on five benchmarks including ScanNet++ and uCO3D, TrianguLang achieves state‑of‑the‑art feed‑forward text‑guided segmentation and localization, reducing user effort from O(N) clicks to a single text query. The model processes each frame at 1008x1008 resolution in ~57ms (~18 FPS) without optimization, enabling practical deployment for interactive robotics and AR applications. Code and checkpoints are available at https://cwru‑aism.github.io/triangulang/.
Authors:Yanan Wu, Yuhan Yan, Tailai Chen, Zhixiang Chi, ZiZhang Wu, Yi Jin, Yang Wang, Zhenbo Li
Abstract:
On‑the‑fly category discovery (OCD) aims to recognize known categories while simultaneously discovering novel ones from an unlabeled online stream, using a model trained only on labeled data. Existing approaches freeze the feature extractor trained offline and employ a hash‑based framework that quantizes features into binary codes as class prototypes. However, discovering novel categories with a fixed knowledge base is counterintuitive, as the learning potential of incoming data is entirely neglected. In addition, feature quantization introduces information loss, diminishes representational expressiveness, and amplifies intra‑class variance. It often results in category explosion, where a single class is fragmented into multiple pseudo‑classes. To overcome these limitations, we propose a test‑time adaptation framework that enables learning through discovery. It incorporates two complementary strategies: a semantic‑aware prototype update and a stable test‑time encoder update. The former dynamically refines class prototypes to enhance classification, whereas the latter integrates new information directly into the parameter space. Together, these components allow the model to continuously expand its knowledge base with newly encountered samples. Furthermore, we introduce a margin‑aware logit calibration in the offline stage to enlarge inter‑class margins and improve intra‑class compactness, thereby reserving embedding space for future class discovery. Experiments on standard OCD benchmarks demonstrate that our method substantially outperforms existing hash‑based state‑of‑the‑art approaches, yielding notable improvements in novel‑class accuracy and effectively mitigating category explosion. The code is publicly available at \textcolorbluehttps://github.com/ynanwu/TALON.
Authors:Zexi Jia, Pengcheng Luo, Yijia Zhong, Jinchao Zhang, Jie Zhou
Abstract:
Most evaluations of generative models rely on feature‑distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the space of discrete visual tokens, where modern 1D image tokenizers compactly encode both semantic and perceptual information and quality manifests as predictable token statistics. We introduce Codebook Histogram Distance (CHD), a training‑free distribution metric in token space, and Code Mixture Model Score (CMMS), a no‑reference quality metric learned from synthetic degradations of token sequences. To stress‑test metrics under broad distribution shifts, we further propose VisForm, a benchmark of 210K images spanning 62 visual forms and 12 generative models with expert annotations. Across AGIQA, HPDv2/3, and VisForm, our token‑based metrics achieve state‑of‑the‑art correlation with human judgments. We will release all code and datasets to facilitate future research, with the code publicly available at https://github.com/zexiJia/1d‑Distance.
Authors:Weining Ren, Xiao Tan, Kai Han
Abstract:
While recent feed‑forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a prohibitive computational bottleneck that severely limits inference speed. To resolve this, we introduce Speed3R, an end‑to‑end trainable model inspired by the core principle of Structure‑from‑Motion: that a sparse set of keypoints is sufficient for robust pose estimation. Speed3R features a dual‑branch attention mechanism where a compression branch creates a coarse contextual prior to guide a selection branch, which performs fine‑grained attention only on the most informative image tokens. This strategy mimics the efficiency of traditional keypoint matching, achieving a remarkable 12.4x inference speedup on 1000‑view sequences, while introducing a minimal, controlled trade‑off in geometric accuracy. Validated on standard benchmarks with both VGGT and π^3 backbones, our method delivers high‑quality reconstructions at a fraction of computational cost, paving the way for efficient large‑scale scene modeling.
Authors:Sangjune Park, Inhyeok Choi, Donghyeon Soon, Youngwoo Jeon, Kyungdon Joo
Abstract:
Dance is a form of human motion characterized by emotional expression and communication, playing a role in various fields such as music, virtual reality, and content creation. Existing methods for dance generation often fail to adequately capture the inherently sequential, rhythmical, and music‑synchronized characteristics of dance. In this paper, we propose \emphMambaDance, a new dance generation approach that leverages a Mamba‑based diffusion model. Mamba, well‑suited to handling long and autoregressive sequences, is integrated into our two‑stage diffusion architecture, substituting off‑the‑shelf Transformer. Additionally, considering the critical role of musical beats in dance choreography, we propose a Gaussian‑based beat representation to explicitly guide the decoding of dance sequences. Experiments on AIST++ and FineDance datasets for each sequence length show that our proposed method effectively generates plausible dance movements while reflecting essential characteristics, consistently from short to long dances, compared to the previous methods. Additional qualitative results and demo videos are available at \smallhttps://vision3d‑lab.github.io/mambadance.
Authors:Yafei Zhang, Meng Ma, Huafeng Li, Yu Liu
Abstract:
Infrared‑visible (IR‑VIS) image fusion is vital for perception and security, yet most methods rely on the availability of both modalities during training and inference. When the infrared modality is absent, pixel‑space generative substitutes become hard to control and inherently lack interpretability. We address missing‑IR fusion by proposing a dictionary‑guided, coefficient‑domain framework built upon a shared convolutional dictionary. The pipeline comprises three key components: (1) Joint Shared‑dictionary Representation Learning (JSRL) learns a unified and interpretable atom space shared by both IR and VIS modalities; (2) VIS‑Guided IR Inference (VGII) transfers VIS coefficients to pseudo‑IR coefficients in the coefficient domain and performs a one‑step closed‑loop refinement guided by a frozen large language model as a weak semantic prior; and (3) Adaptive Fusion via Representation Inference (AFRI) merges VIS structures and inferred IR cues at the atom level through window attention and convolutional mixing, followed by reconstruction with the shared dictionary. This encode‑transfer‑fuse‑reconstruct pipeline avoids uncontrolled pixel‑space generation while ensuring prior preservation within interpretable dictionary‑coefficient representation. Experiments under missing‑IR settings demonstrate consistent improvements in perceptual quality and downstream detection performance. To our knowledge, this represents the first framework that jointly learns a shared dictionary and performs coefficient‑domain inference‑fusion to tackle missing‑IR fusion. The source code is publicly available at https://github.com/harukiv/DCMIF.
Authors:Stefan Lionar, Gim Hee Lee
Abstract:
Physics‑based humanoid control has achieved remarkable progress in enabling realistic and high‑performing single‑agent behaviors, yet extending these capabilities to cooperative human‑object interaction (HOI) remains challenging. We present TeamHOI, a framework that enables a single decentralized policy to handle cooperative HOIs across any number of cooperating agents. Each agent operates using local observations while attending to other teammates through a Transformer‑based policy network with teammate tokens, allowing scalable coordination across variable team sizes. To enforce motion realism while addressing the scarcity of cooperative HOI data, we further introduce a masked Adversarial Motion Prior (AMP) strategy that uses single‑human reference motions while masking object‑interacting body parts during training. The masked regions are then guided through task rewards to produce diverse and physically plausible cooperative behaviors. We evaluate TeamHOI on a challenging cooperative carrying task involving two to eight humanoid agents and varied object geometries. Finally, to promote stable carrying, we design a team‑size‑ and shape‑agnostic formation reward. TeamHOI achieves high success rates and demonstrates coherent cooperation across diverse configurations with a single policy.
Authors:Zanming Huang, Jinsu Yoo, Sooyoung Jeon, Zhenzhen Liu, Mark Campbell, Kilian Q Weinberger, Bharath Hariharan, Wei-Lun Chao, Katie Z Luo
Abstract:
LiDAR‑based 3D object detectors typically rely on proposal heads with hand‑crafted components like anchor assignment and non‑maximum suppression (NMS), complicating training and limiting extensibility. We present AutoReg3D, an autoregressive 3D detector that casts detection as sequence generation. Given point‑cloud features, AutoReg3D emits objects in a range‑causal (near‑to‑far) order and encodes each object as a short, discrete‑token sequence consisting of its center, size, orientation, velocity, and class. This near‑to‑far ordering mirrors LiDAR geometry‑‑near objects occlude far ones but not vice versa‑‑enabling straightforward teacher forcing during training and autoregressive decoding at test time. AutoReg3D is compatible across diverse point‑cloud or backbones and attains competitive nuScenes performance without anchors or NMS. Beyond parity, the sequential formulation unlocks language‑model advances for 3D perception, including GRPO‑style reinforcement learning for task‑aligned objectives. These results position autoregressive decoding as a viable, flexible alternative for LiDAR‑based detection and open a path to importing modern sequence‑modeling tools into 3D perception.
Authors:Yanning Hou, Peiyuan Li, Zirui Liu, Yitong Wang, Yanran Ruan, Jianfeng Qiu, Ke Xu
Abstract:
Zero‑shot anomaly detection (ZSAD) requires detecting and localizing anomalies without access to target‑class anomaly samples. Mainstream methods rely on vision‑language models (VLMs) such as CLIP: they build hand‑crafted or learned prompt sets for normal and abnormal semantics, then compute image‑text similarities for open‑set discrimination. While effective, this paradigm depends on a text encoder and cross‑modal alignment, which can lead to training instability and parameter redundancy. This work revisits the necessity of the text branch in ZSAD and presents VisualAD, a purely visual framework built on Vision Transformers. We introduce two learnable tokens within a frozen backbone to directly encode normality and abnormality. Through multi‑layer self‑attention, these tokens interact with patch tokens, gradually acquiring high‑level notions of normality and anomaly while guiding patches to highlight anomaly‑related cues. Additionally, we incorporate a Spatial‑Aware Cross‑Attention (SCA) module and a lightweight Self‑Alignment Function (SAF): SCA injects fine‑grained spatial information into the tokens, and SAF recalibrates patch features before anomaly scoring. VisualAD achieves state‑of‑the‑art performance on 13 zero‑shot anomaly detection benchmarks spanning industrial and medical domains, and adapts seamlessly to pretrained vision backbones such as the CLIP image encoder and DINOv2. Code: https://github.com/7HHHHH/VisualAD
Authors:Sunghyun Baek, Jaemyung Yu, Seunghee Koh, Minsu Kim, Hyeonseong Jeon, Junmo Kim
Abstract:
Test‑time adaptation (TTA) has been widely explored to prevent performance degradation when test data differ from the training distribution. However, fully leveraging the rich representations of large pretrained models with minimal parameter updates remains underexplored. In this paper, we propose Intrinsic Mixture of Spectral Experts (IMSE) that leverages the spectral experts inherently embedded in Vision Transformers. We decompose each linear layer via singular value decomposition (SVD) and adapt only the singular values, while keeping the singular vectors fixed. We further identify a key limitation of entropy minimization in TTA: it often induces feature collapse, causing the model to rely on domain‑specific features rather than class‑discriminative features. To address this, we propose a diversity maximization loss based on expert‑input alignment, which encourages diverse utilization of spectral experts during adaptation. In the continual test‑time adaptation (CTTA) scenario, beyond preserving pretrained knowledge, it is crucial to retain and reuse knowledge from previously observed domains. We introduce Domain‑Aware Spectral Code Retrieval, which estimates input distributions to detect domain shifts, and retrieves adapted singular values for rapid adaptation. Consequently, our method achieves state‑of‑the‑art performance on various distribution‑shift benchmarks under the TTA setting. In CTTA and Gradual CTTA, it further improves accuracy by 3.4 percentage points (pp) and 2.4 pp, respectively, while requiring 385 times fewer trainable parameters. Our code is available at https://github.com/baek85/IMSE.
Authors:Yingkai Zhang, Tao Zhang, Jing Nie, Ying Fu
Abstract:
Unregistered hyperspectral image (HSI) super‑resolution (SR) typically aims to enhance a low‑resolution HSI using an unregistered high‑resolution reference image. In this paper, we propose an unmixing‑based fusion framework that decouples spatial‑spectral information to simultaneously mitigate the impact of unregistered fusion and enhance the learnability of SR models. Specifically, we first utilize singular value decomposition for initial spectral unmixing, preserving the original endmembers while dedicating the subsequent network to enhancing the initial abundance map. To leverage the spatial texture of the unregistered reference, we introduce a coarse‑to‑fine deformable aggregation module, which first estimates a pixel‑level flow and a similarity map using a coarse pyramid predictor. It further performs fine sub‑pixel refinement to achieve deformable aggregation of the reference features. The aggregative features are then refined via a series of spatial‑channel abundance cross‑attention blocks. Furthermore, a spatial‑channel modulated fusion module is presented to merge encoder‑decoder features using dynamic gating weights, yielding a high‑quality, high‑resolution HSI. Experimental results on simulated and real datasets confirm that our proposed method achieves state‑of‑the‑art super‑resolution performance. The code will be available at https://github.com/yingkai‑zhang/UAFL.
Authors:Hao Wei, Yanhui Zhou, Chenyang Ge
Abstract:
Although learned video compression methods have exhibited outstanding performance, most of them typically follow a hybrid coding paradigm that requires explicit motion estimation and compensation, resulting in a complex solution for video compression. In contrast, we introduce a streamlined yet effective video compression framework founded on a direct transform strategy, i.e., nonlinear transform, quantization, and entropy coding. We first develop a cascaded Mamba module (CMM) with different embedded geometric transformations to effectively explore both long‑range spatial and temporal dependencies. To improve local spatial representation, we introduce a locality refinement feed‑forward network (LRFFN) that incorporates a hybrid convolution block based on difference convolutions. We integrate the proposed CMM and LRFFN into the encoder and decoder of our compression framework. Moreover, we present a conditional channel‑wise entropy model that effectively utilizes conditional temporal priors to accurately estimate the probability distributions of current latent features. Extensive experiments demonstrate that our method outperforms state‑of‑the‑art video compression approaches in terms of perceptual quality and temporal consistency under low‑bitrate constraints. Our source codes and models will be available at https://github.com/cshw2021/GTEM‑LVC.
Authors:Hui Liu, Kecheng Chen, Jialiang Wang, Xianming Liu, Wenya Wang, Haoliang Li
Abstract:
Vision‑Language Models (VLMs), such as CLIP, have significantly advanced zero‑shot image recognition. However, their performance remains limited by suboptimal prompt engineering and poor adaptability to target classes. While recent methods attempt to improve prompts through diverse class descriptions, they often rely on heuristic designs, lack versatility, and are vulnerable to outlier prompts. This paper enhances prompt by incorporating class‑specific concepts. By treating concepts as latent variables, we rethink zero‑shot image classification from a Bayesian perspective, casting prediction as marginalization over the concept space, where each concept is weighted by a prior and a test‑image conditioned likelihood. This formulation underscores the importance of both a well‑structured concept proposal distribution and the refinement of concept priors. To construct an expressive and efficient proposal distribution, we introduce a multi‑stage concept synthesis pipeline driven by LLMs to generate discriminative and compositional concepts, followed by a Determinantal Point Process to enforce diversity. To mitigate the influence of outlier concepts, we propose a training‑free, adaptive soft‑trim likelihood, which attenuates their impact in a single forward pass. We further provide robustness guarantees and derive multi‑class excess risk bounds for our framework. Extensive experiments demonstrate that our method consistently outperforms state‑of‑the‑art approaches, validating its effectiveness in zero‑shot image classification. Our code is available at https://github.com/less‑and‑less‑bugs/CGBC.
Authors:Gil Shapira, Ishay Goldin, Evgeny Artyomov, Donghoon Kim, Yosi Keller, Niv Zehngut
Abstract:
Gaze estimation is instrumental in modern virtual reality (VR) systems. Despite significant progress in remote‑camera gaze estimation, VR gaze research remains constrained by data scarcity, particularly the lack of large‑scale, accurately labeled datasets captured with the off‑axis camera configurations typical of modern headsets. Gaze annotation is difficult since fixation on intended targets cannot be guaranteed. To address these challenges, we introduce VRGaze, the first large‑scale off‑axis gaze estimation dataset for VR, comprising 2.1 million near‑eye infrared images collected from 68 participants. We further propose GazeShift, an attention‑guided unsupervised framework for learning gaze representations without labeled data. Unlike prior redirection‑based methods that rely on multi‑view or 3D geometry, GazeShift is tailored to near‑eye imagery, achieving effective gaze‑appearance disentanglement in a compact, real‑time model. GazeShift embeddings can be optionally adapted to individual users via lightweight few‑shot calibration, achieving a 1.84° mean error on VRGaze. On the remote‑camera MPIIGaze dataset, the model achieves a 7.15° person‑agnostic error, doing so with 10x fewer parameters and 35x fewer FLOPs than baseline methods. Deployed natively on a VR headset GPU, inference takes only 5 ms. Combined with demonstrated robustness to illumination changes, these results highlight GazeShift as a label‑efficient, real‑time solution for VR gaze tracking. Project code and the VRGaze dataset are released at https://github.com/gazeshift3/gazeshift
Authors:Han Yan, Zishang Xiang, Zeyu Zhang, Hao Tang
Abstract:
World models enable planning in imagined future predicted space, offering a promising framework for embodied navigation. However, existing navigation world models often lack action‑conditioned consistency, so visually plausible predictions can still drift under multi‑step rollout and degrade planning. Moreover, efficient deployment requires few‑step diffusion inference, but existing distillation methods do not explicitly preserve rollout consistency, creating a training‑inference mismatch. To address these challenges, we propose MWM, a mobile world model for planning‑based image‑goal navigation. Specifically, we introduce a two‑stage training framework that combines structure pretraining with Action‑Conditioned Consistency (ACC) post‑training to improve action‑conditioned rollout consistency. We further introduce Inference‑Consistent State Distillation (ICSD) for few‑step diffusion distillation with improved rollout consistency. Our experiments on benchmark and real‑world tasks demonstrate consistent gains in visual fidelity, trajectory accuracy, planning success, and inference efficiency. Code: https://github.com/AIGeeksGroup/MWM. Website: https://aigeeksgroup.github.io/MWM.
Authors:Zixuan Pan, Kaiyuan Tang, Jun Xia, Yifan Qin, Lin Gu, Chaoli Wang, Jianxu Chen, Yiyu Shi
Abstract:
2D Gaussian Splatting has emerged as a novel image representation technique that can support efficient rendering on low‑end devices. However, scaling to high‑resolution images requires optimizing and storing millions of unstructured Gaussian primitives independently, leading to slow convergence and redundant parameters. To address this, we propose Structured Gaussian Image (SGI), a compact and efficient framework for representing high‑resolution images. SGI decomposes a complex image into multi‑scale local spaces defined by a set of seeds. Each seed corresponds to a spatially coherent region and, together with lightweight multi‑layer perceptrons (MLPs), generates structured implicit 2D neural Gaussians. This seed‑based formulation imposes structural regularity on otherwise unstructured Gaussian primitives, which facilitates entropy‑based compression at the seed level to reduce the total storage. However, optimizing seed parameters directly on high‑resolution images is a challenging and non‑trivial task. Therefore, we designed a multi‑scale fitting strategy that refines the seed representation in a coarse‑to‑fine manner, substantially accelerating convergence. Quantitative and qualitative evaluations demonstrate that SGI achieves up to 7.5x compression over prior non‑quantized 2D Gaussian methods and 1.6x over quantized ones, while also delivering 1.6x and 6.5x faster optimization, respectively, without degrading, and often improving, image fidelity. Code is available at https://github.com/zx‑pan/SGI.
Authors:Yihong Luo, Tianyang Hu, Weijian Luo, Jing Tang
Abstract:
While few‑step generative models have enabled powerful image and video generation at significantly lower cost, generic reinforcement learning (RL) paradigms for few‑step models remain an unsolved problem. Existing RL approaches for few‑step diffusion models strongly rely on back‑propagating through differentiable reward models, thereby excluding the majority of important real‑world reward signals, e.g., non‑differentiable rewards such as humans' binary likeness, object counts, etc. To properly incorporate non‑differentiable rewards to improve few‑step generative models, we introduce TDM‑R1, a novel reinforcement learning paradigm built upon a leading few‑step model, Trajectory Distribution Matching (TDM). TDM‑R1 decouples the learning process into surrogate reward learning and generator learning. Furthermore, we developed practical methods to obtain per‑step reward signals along the deterministic generation trajectory of TDM, resulting in a unified RL post‑training method that significantly improves few‑step models' ability with generic rewards. We conduct extensive experiments ranging from text‑rendering, visual quality, and preference alignment. All results demonstrate that TDM‑R1 is a powerful reinforcement learning paradigm for few‑step text‑to‑image models, achieving state‑of‑the‑art reinforcement learning performances on both in‑domain and out‑of‑domain metrics. Furthermore, TDM‑R1 also scales effectively to the recent strong Z‑Image model, consistently outperforming both its 100‑NFE and few‑step variants with only 4 NFEs. Project page: https://github.com/Luo‑Yihong/TDM‑R1
Authors:Junkun Jiang, Jie Chen, Ho Yin Au, Jingyu Xiang
Abstract:
Vision‑based motion capture solutions often struggle with occlusions, which result in the loss of critical joint information and hinder accurate 3D motion reconstruction. Other wearable alternatives also suffer from noisy or unstable data, often requiring extensive manual cleaning and correction to achieve reliable results. To address these challenges, we introduce the Masked Motion Diffusion Model (MMDM), a diffusion‑based generative reconstruction framework that enhances incomplete or low‑confidence motion data using partially available high‑quality reconstructions within a Masked Autoencoder architecture. Central to our design is the Kinematic Attention Aggregation (KAA) mechanism, which enables efficient, deep, and iterative encoding of both joint‑level and pose‑level features, capturing structural and temporal motion patterns essential for task‑specific reconstruction. We focus on learning context‑adaptive motion priors, specialized structural and temporal features extracted by the same reusable architecture, where each learned prior emphasizes different aspects of motion dynamics and is specifically efficient for its corresponding task. This enables the architecture to adaptively specialize without altering its structure. Such versatility allows MMDM to efficiently learn motion priors tailored to scenarios such as motion refinement, completion, and in‑betweening. Extensive evaluations on public benchmarks demonstrate that MMDM achieves strong performance across diverse masking strategies and task settings. The source code is available at https://github.com/jjkislele/MMDM.
Authors:Yuhang Wang, Hai Li, Shujuan Hou, Zhetao Dong, Xiaoyao Yang
Abstract:
In bandwidth‑limited online video streaming, videos are usually downsampled and compressed. Although recent online video super‑resolution (online VSR) approaches achieve promising results, they are still compute‑intensive and fall short of real‑time processing at higher resolutions, due to complex motion estimation for alignment and redundant processing of consecutive frames. To address these issues, we propose a compressed‑domain‑aware network (CDA‑VSR) for online VSR, which utilizes compressed‑domain information, including motion vectors, residual maps, and frame types to balance quality and efficiency. Specifically, we propose a motion‑vector‑guided deformable alignment module that uses motion vectors for coarse warping and learns only local residual offsets for fine‑tuned adjustments, thereby maintaining accuracy while reducing computation. Then, we utilize a residual map gated fusion module to derive spatial weights from residual maps, suppressing mismatched regions and emphasizing reliable details. Further, we design a frame‑type‑aware reconstruction module for adaptive compute allocation across frame types, balancing accuracy and efficiency. On the REDS4 dataset, our CDA‑VSR surpasses the state‑of‑the‑art method TMP, with a maximum PSNR improvement of 0.13 dB while delivering more than double the inference speed. The code will be released at https://github.com/sspBIT/CDA‑VSR.
Authors:Congcong Bian, Haolong Ma, Hui Li, Zhongwei Shen, Xiaoqing Luo, Xiaoning Song, Xiao-Jun Wu
Abstract:
Spatial registration across different visual modalities is a critical but formidable step in multi‑modality image fusion for real‑world perception. Although several methods are proposed to address this issue, the existing registration‑based fusion methods typically require extensive pre‑registration operations, limiting their efficiency. To overcome these limitations, a general cross‑modality registration method guided by visual priors is proposed for infrared and visible image fusion task, termed FusionRegister. Firstly, FusionRegister achieves robustness by learning cross‑modality misregistration representations rather than forcing alignment of all differences, ensuring stable outputs even under challenging input conditions. Moreover, FusionRegister demonstrates strong generality by operating directly on fused results, where misregistration is explicitly represented and effectively handled, enabling seamless integration with diverse fusion methods while preserving their intrinsic properties. In addition, its efficiency is further enhanced by serving the backbone fusion method as a natural visual prior provider, which guides the registration process to focus only on mismatch regions, thereby avoiding redundant operations. Extensive experiments on three datasets demonstrate that FusionRegister not only inherits the fusion quality of state‑of‑the‑art methods, but also delivers superior detail alignment and robustness, making it highly suitable for infrared and visible image fusion method. The code will be available at https://github.com/bociic/FusionRegister.
Authors:Yuanyuan Gao, Hao Li, Yifei Liu, Xinhao Ji, Yuning Gong, Yuanjun Liao, Fangfu Liu, Manyuan Zhang, Yuchen Yang, Dan Xu, Xue Yang, Huaxi Huang, Hongjie Zhang, Ziwei Liu, Xiao Sun, Dingwen Zhang, Zhihang Zhong
Abstract:
The pursuit of spatial intelligence fundamentally relies on access to large‑scale, fine‑grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question‑answer (QA) pairs from a limited number of manually annotated datasets, rather than systematically annotating new large‑scale 3D scenes from raw web data. As a result, their scalability is severely constrained, and model performance is further hindered by domain gaps inherent in these narrowly curated datasets.
In this work, we propose Holi‑Spatial, the first fully automated, large‑scale, spatially‑aware multimodal dataset, constructed from raw video inputs without human intervention, using the proposed data curation pipeline. Holi‑Spatial supports multi‑level spatial supervision, ranging from geometrically accurate 3D Gaussian Splatting (3DGS) reconstructions with rendered depth maps to object‑level and relational semantic annotations, together with corresponding spatial Question‑Answer (QA) pairs.
Following a principled and systematic pipeline, we further construct Holi‑Spatial‑4M, the first large‑scale, high‑quality 3D semantic dataset, containing 12K optimized 3DGS scenes, 1.3M 2D masks, 320K 3D bounding boxes, 320K instance captions, 1.2M 3D grounding instances, and 1.2M spatial QA pairs spanning diverse geometric, relational, and semantic reasoning tasks.
Holi‑Spatial demonstrates exceptional performance in data curation quality, significantly outperforming existing feed‑forward and per‑scene optimized methods on datasets such as ScanNet, ScanNet++, and DL3DV. Furthermore, fine‑tuning Vision‑Language Models (VLMs) on spatial reasoning tasks using this dataset has also led to substantial improvements in model performance.
Authors:Kaihua Tang, Jiaxin Qi, Jinli Ou, Yuhua Zheng, Jianqiang Huang
Abstract:
The emergence of Large Language Models (LLMs) has driven rapid progress in multi‑modal learning, particularly in the development of Large Vision‑Language Models (LVLMs). However, existing LVLM training paradigms place excessive reliance on the LLM component, giving rise to two critical robustness challenges: language bias and language sensitivity. To address both issues simultaneously, we propose a novel Self‑Critical Inference (SCI) framework that extends Visual Contrastive Decoding by conducting multi‑round counterfactual reasoning through both textual and visual perturbations. This process further introduces a new strategy for improving robustness by scaling the number of counterfactual rounds. Moreover, we also observe that failure cases of LVLMs differ significantly across models, indicating that fixed robustness benchmarks may not be able to capture the true reliability of LVLMs. To this end, we propose the Dynamic Robustness Benchmark (DRBench), a model‑specific evaluation framework targeting both language bias and sensitivity issues. Extensive experiments show that SCI consistently outperforms baseline methods on DRBench, and that increasing the number of inference rounds further boosts robustness beyond existing single‑step counterfactual reasoning methods.
Authors:Likui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen, Zisheng Chen, Jianhua Han, Jiangtong Zhu, Pei Xu, Hang Xu, Hefeng Wu, Liang Lin, Xiaodan Liang
Abstract:
Recent advances in Visual‑Language‑Action (VLA) models have shown promising potential for robotic manipulation tasks. However, real‑world robotic tasks often involve long‑horizon, multi‑step problem‑solving and require generalization for continual skill acquisition, extending beyond single actions or skills. These challenges present significant barriers for existing VLA models, which use monolithic action decoders trained on aggregated data, resulting in poor scalability. To address these challenges, we propose AtomicVLA, a unified planning‑and‑execution framework that jointly generates task‑level plans, atomic skill abstractions, and fine‑grained actions. AtomicVLA constructs a scalable atomic skill library through a Skill‑Guided Mixture‑of‑Experts (SG‑MoE), where each expert specializes in mastering generic yet precise atomic skills. Furthermore, we introduce a flexible routing encoder that automatically assigns dedicated atomic experts to new skills, enabling continual learning. We validate our approach through extensive experiments. In simulation, AtomicVLA outperforms π_0 by 2.4% on LIBERO, 10% on LIBERO‑LONG, and outperforms π_0 and π_0.5 by 0.22 and 0.25 in average task length on CALVIN. Additionally, our AtomicVLA consistently surpasses baselines by 18.3% and 21% in real‑world long‑horizon tasks and continual learning. These results highlight the effectiveness of atomic skill abstraction and dynamic expert composition for long‑horizon and lifelong robotic tasks. The project page is \hrefhttps://zhanglk9.github.io/atomicvla‑web/here.
Authors:Shumeng Li, Jintao Guo, Jian Zhang, Yulin Zhou, Luyang Cao, Yinghuan Shi
Abstract:
Cross‑subject visual decoding aims to reconstruct visual experiences from brain activity across individuals, enabling more scalable and practical brain‑computer interfaces. However, existing methods often suffer from degraded performance when adapting to new subjects with limited data, as they struggle to preserve both the semantic consistency of stimuli and the alignment of brain responses. To address these challenges, we propose Duala, a dual‑level alignment framework designed to achieve stimulus‑level consistency and subject‑level alignment in fMRI‑based cross‑subject visual decoding. (1) At the stimulus level, Duala introduces a semantic alignment and relational consistency strategy that preserves intra‑class similarity and inter‑class separability, maintaining clear semantic boundaries during adaptation. (2) At the subject level, a distribution‑based feature perturbation mechanism is developed to capture both global and subject‑specific variations, enabling adaptation to individual neural representations without overfitting. Experiments on the Natural Scenes Dataset (NSD) demonstrate that Duala effectively improves alignment across subjects. Remarkably, even when fine‑tuned with only about one hour of fMRI data, Duala achieves over 81.1% image‑to‑brain retrieval accuracy and consistently outperforms existing fine‑tuning strategies in both retrieval and reconstruction. Our code is available at https://github.com/ShumengLI/Duala.
Authors:Zixiao Wen, Zhen Yang, Jiawei Li, Xiantai Xiang, Guangyao Zhou, Yuxin Hu, Yuhan Liu
Abstract:
Single object tracking in satellite videos is inherently challenged by small target, blurred background, large aspect ratio changes, and frequent visual occlusions. These constraints often cause appearance‑based trackers to accumulate errors and lose targets irreversibly. To systematically mitigate both spatial ambiguities and temporal information loss, we propose SiamGM, a novel geometry‑aware and motion‑guided Siamese network. From a spatial perspective, we introduce an Inter‑Frame Graph Attention (IFGA) module, closely integrated with an Aspect Ratio‑Constrained Label Assignment (LA) method, establishing fine‑grained topological correspondences and explicitly preventing surrounding background noise. From a temporal perspective, we introduce the Motion Vector‑Guided Online Tracking Optimization method. By adopting the Normalized Peak‑to‑Sidelobe Ratio (nPSR) as a dynamic confidence indicator, we propose an Online Motion Model Refinement (OMMR) strategy to utilize historical trajectory information. Evaluations on two challenging SatSOT and SV248S benchmarks confirm that SiamGM outperforms most state‑of‑the‑art trackers in both precision and success metrics. Notably, the proposed components of SiamGM introduce virtually no computational overhead, enabling real‑time tracking at 130 frames per second (FPS). Codes and tracking results are available at https://github.com/wenzx18/SiamGM.
Authors:Chenhui Wang, Boyun Zheng, Liuxin Bao, Zhihao Peng, Peter Y. M. Woo, Hongming Shan, Yixuan Yuan
Abstract:
Precise prognostic modeling of glioblastoma (GBM) under varying treatment interventions is essential for optimizing clinical outcomes. While generative AI has shown promise in simulating GBM evolution, existing methods typically treat interventions as static conditional inputs rather than dynamic decision variables. Consequently, they fail to capture the complex, reciprocal interplay between tumor evolution and treatment response. To bridge this gap, we present Brain‑WM, a pioneering brain GBM world model that unifies next‑step treatment prediction and future MRI generation, thereby capturing the co‑evolutionary dynamics between tumor and treatment. Specifically, Brain‑WM encodes spatiotemporal dynamics into a shared latent space for joint autoregressive treatment prediction and flow‑based future MRI generation. Then, instead of a conventional monolithic framework, Brain‑WM adopts a novel Y‑shaped Mixture‑of‑Transformers (MoT) architecture. This design structurally disentangles heterogeneous objectives, successfully leveraging cross‑task synergies while preventing feature collapse. Finally, a synergistic multi‑timepoint mask alignment objective explicitly anchors latent representations to anatomically grounded tumor structures and progression‑aware semantics. Extensive validation on internal and external multi‑institutional cohorts demonstrates the superiority of Brain‑WM, achieving 91.5% accuracy in treatment planning and SSIMs of 0.8524, 0.8581, and 0.8404 for FLAIR, T1CE, and T2W sequences, respectively. Ultimately, Brain‑WM offers a robust clinical sandbox for optimizing patient healthcare. The source code is made available at https://github.com/thibault‑wch/Brain‑GBM‑world‑model.
Authors:Zhichao Liao, Xiaole Xian, Qingyu Li, Wenyu Qin, Meng Wang, Weicheng Xie, Siyang Song, Pingfa Feng, Long Zeng, Liang Pan
Abstract:
Existing concept customization methods have achieved remarkable outcomes in high‑fidelity and multi‑concept customization. However, they often neglect the influence on the original model's behavior and capabilities when learning new personalized concepts. To address this issue, we propose PureCC. PureCC introduces a novel decoupled learning objective for concept customization, which combines the implicit guidance of the target concept with the original conditional prediction. This separated form enables PureCC to substantially focus on the original model during training. Moreover, based on this objective, PureCC designs a dual‑branch training pipeline that includes a frozen extractor providing purified target concept representations as implicit guidance and a trainable flow model producing the original conditional prediction, jointly achieving pure learning for personalized concepts. Furthermore, PureCC introduces a novel adaptive guidance scale λ^\star to dynamically adjust the guidance strength of the target concept, balancing customization fidelity and model preservation. Extensive experiments show that PureCC achieves state‑of‑the‑art performance in preserving the original behavior and capabilities while enabling high‑fidelity concept customization. The code is available at https://github.com/lzc‑sg/PureCC.
Authors:Guoqing Zhang, Jingyun Yang, Siqi Chen, Anping Zhang, Yang Li
Abstract:
Anatomy shape modeling is a fundamental problem in medical data analysis. However, the geometric complexity and topological variability of anatomical structures pose significant challenges to accurate anatomical shape generation. In this work, we propose a skeletal latent diffusion framework that explicitly incorporates structural priors for efficient and high‑fidelity medical shape generation. We introduce a shape auto‑encoder in which the encoder captures global geometric information through a differentiable skeletonization module and aggregates local surface features into shape latents, while the decoder predicts the corresponding implicit fields over sparsely sampled coordinates. New shapes are generated via a latent‑space diffusion model, followed by neural implicit decoding and mesh extraction. To address the limited availability of medical shape data, we construct a large‑scale dataset, MedSDF, comprising surface point clouds and corresponding signed distance fields across multiple anatomical categories. Extensive experiments on MedSDF and vessel datasets demonstrate that the proposed method achieves superior reconstruction and generation quality while maintaining a higher computational efficiency compared with existing approaches. Code is available at: https://github.com/wlsdzyzl/meshage.
Authors:Wenqi Cai, Yawen Zou, Guang Li, Chunzhi Gu, Chao Zhang
Abstract:
Dataset distillation (DD) aims to synthesize compact training sets that enable models to achieve high accuracy with significantly fewer samples. Recent diffusion‑based DD methods commonly introduce semantic guidance through late‑stage cross‑attention, where textual prompts tend to dominate the generative process. Although this strategy enforces label relevance, it diminishes the contribution of visual latents, resulting in over‑corrected samples that mirror prompt patterns rather than reflecting intrinsic visual features. To solve this problem, we introduce an Early Vision‑Language Fusion (EVLF) method that aligns textual and visual embeddings at the transition between the encoder and the generative backbone. By incorporating a lightweight cross‑attention module at this transition, the early representations simultaneously encode local textures and global semantic directions across the denoising process. Importantly, EVLF is plug‑and‑play and can be easily integrated into any diffusion‑based dataset distillation pipeline with an encoder. It works across different denoiser architectures and sampling schedules without any task‑specific modifications. Extensive experiments demonstrate that EVLF generates semantically faithful and visually coherent synthetic data, yielding consistent improvements in downstream classification accuracy across varied settings. Source code is available at https://github.com/wenqi‑cai297/earlyfusion‑for‑dd/.
Authors:Xiaokang Zhang, Xuran Xiong, Jianzhong Huang, Lefei Zhang
Abstract:
Remote sensing image segmentation (RSIS) in federated environments has gained increasing attention because it enables collaborative model training across distributed datasets without sharing raw imagery or annotations. Federated RSIS combined with parameter‑efficient fine‑tuning (PEFT) can unleash the generalization power of pretrained foundation models for real‑world applications, with minimal parameter aggregation and communication overhead. However, the dynamic adaptation of pretrained models to heterogeneous client data inevitably increases update uncertainty and compromises the reliability of collaborative optimization due to the lack of uncertainty estimation for each local model. To bridge this gap, we present FedEU, a federated optimization framework for fine‑tuning RSIS models driven by evidential uncertainty. Specifically, personalized evidential uncertainty modeling is introduced to quantify epistemic variations of local models and identify high‑risk areas under local data distributions. Furthermore, the client‑specific feature embedding (CFE) is exploited to enhance channel‑aware feature representation while preserving client‑specific properties through personalized attention and an element‑aware parameter update approach. These uncertainty estimates are uploaded to the server to enable adaptive global aggregation via a Top‑k uncertainty‑guided weighting (TUW) strategy, which mitigates the impact of distribution shifts and unreliable updates. Extensive experiments on three large‑scale heterogeneous datasets demonstrate the superior performance of FedEU. More importantly, FedEU enables balanced model adaptation across diverse clients by explicitly reducing prediction uncertainty, resulting in more robust and reliable federated outcomes. The source codes will be available at https://github.com/zxk688/FedEU.
Authors:Xiaokang Zhang, Bo Li, Chufeng Zhou, Weikang Yu, Lefei Zhang
Abstract:
Pretraining and fine‑tuning have emerged as a new paradigm in remote sensing image interpretation. Among them, Masked Autoencoder (MAE)‑based pretraining stands out for its strong capability to learn general feature representations via reconstructing masked image regions. However, applying MAE to multispectral remote sensing images remains challenging due to complex backgrounds, indistinct targets, and the lack of semantic guidance during masking, which hinders the learning of underlying structures and meaningful spatial‑spectral features. To address this, we propose a simple yet effective approach, Spectral Index‑Guided MAE (SIGMAE), for multispectral image pretraining. The core idea is to incorporate domain‑specific spectral indices as prior knowledge to guide dynamic token masking toward informative regions. SIGMAE introduces Semantic Saliency‑Guided Dynamic Token Masking (SSDTM), a curriculum‑style strategy that quantifies each patch's semantic richness and internal heterogeneity to adaptively select the most informative tokens during training. By prioritizing semantically salient regions and progressively increasing sample difficulty, SSDTM enhances spectrally rich and structurally aware representation learning, mitigates overfitting, and reduces redundant computation compared with random masking. Extensive experiments on five widely used datasets covering various downstream tasks, including scene classification, semantic segmentation, object extraction and change detection, demonstrate that SIGMAE outperforms other pretrained geospatial foundation models. Moreover, it exhibits strong spatial‑spectral reconstruction capability, even with a 90% mask ratio, and improves complex target recognition under limited labeled data. The source codes and model weights will be released at https://github.com/zxk688/SIGMAE.
Authors:Mohammad Saeid, Amir Salarpour, Pedram MohajerAnsari, Mert D. Pesé
Abstract:
We present SLNet, a lightweight backbone for 3D point cloud recognition designed to achieve strong performance without the computational cost of many recent attention, graph, and deep MLP based models. The model is built on two simple ideas: NAPE (Nonparametric Adaptive Point Embedding), which captures spatial structure using a combination of Gaussian RBF and cosine bases with input adaptive bandwidth and blending, and GMU (Geometric Modulation Unit), a per channel affine modulator that adds only 2D learnable parameters. These components are used within a four stage hierarchical encoder with FPS+kNN grouping, nonparametric normalization, and shared residual MLPs. In experiments, SLNet shows that a very small model can still remain highly competitive across several 3D recognition tasks. On ModelNet40, SLNet‑S with 0.14M parameters and 0.31 GFLOPs achieves 93.64% overall accuracy, outperforming PointMLP‑elite with 5x fewer parameters, while SLNet‑M with 0.55M parameters and 1.22 GFLOPs reaches 93.92%, exceeding PointMLP with 24x fewer parameters. On ScanObjectNN, SLNet‑M achieves 84.25% overall accuracy within 1.2 percentage points of PointMLP while using 28x fewer parameters. For large scale scene segmentation, SLNet‑T extends the backbone with local Point Transformer attention and reaches 58.2% mIoU on S3DIS Area 5 with only 2.5M parameters, more than 17x fewer than Point Transformer V3. We also introduce NetScore+, which extends NetScore by incorporating latency and peak memory so that efficiency can be evaluated in a more deployment oriented way. Across multiple benchmarks and hardware settings, SLNet delivers a strong overall balance between accuracy and efficiency. Code is available at: https://github.com/m‑saeid/SLNet.
Authors:Suorong Yang, Fangjian Su, Hai Gan, Ziqi Ye, Jie Li, Baile Xu, Furao Shen, Soujanya Poria
Abstract:
Dynamic Data selection aims to accelerate training by prioritizing informative samples during online training. However, existing methods typically rely on task‑specific handcrafted metrics or static/snapshot‑based criteria to estimate sample importance, limiting scalability across learning paradigms and making it difficult to capture the evolving utility of data throughout training. To address this challenge, we propose Data Agent, an end‑to‑end dynamic data selection framework that formulates data selection as a training‑aware sequential decision‑making problem. The agent learns a sample‑wise selection policy that co‑evolves with model optimization, guided by a composite reward that integrates loss‑based difficulty and confidence‑based uncertainty signals. The reward signals capture complementary objectives of optimization impact and information gain, together with a tuning‑free adaptive weighting mechanism that balances these signals over training. Extensive experiments across a wide range of datasets and architectures demonstrate that Data Agent consistently accelerates training while preserving or improving performance, e.g., reducing costs by over 50% on ImageNet‑1k and MMLU with lossless performance. Moreover, its dataset‑agnostic formulation and modular reward make it plug‑and‑play across tasks and scenarios, e.g., robustness to noisy datasets, highlighting its potential in real‑world scenarios. Code is available at https://github.com/Jackbrocp/Data‑Agent.
Authors:Li Gu, Zihuan Jiang, Zhixiang Chi, Huan Liu, Ziqiang Wang, Yuanhao Yu, Glen Berseth, Yang Wang
Abstract:
Graphical user interface (GUI)‑based mobile agents automate digital tasks on mobile devices by interpreting natural‑language instructions and interacting with the screen. While recent methods apply reinforcement learning (RL) to train vision‑language‑model(VLM) agents in interactive environments with a primary focus on performance, generalization remains underexplored due to the lack of standardized benchmarks and open‑source RL systems. In this work, we formalize the problem as a Contextual Markov Decision Process (CMDP) and introduce AndroidWorld‑Generalization, a benchmark with three increasingly challenging regimes for evaluating zero‑shot generalization to unseen task instances, templates, and applications. We further propose an RL training system that integrates Group Relative Policy Optimization (GRPO) with a scalable rollout collection system, consisting of containerized infrastructure and asynchronous execution % , and error recovery to support reliable and efficient training. Experiments on AndroidWorld‑Generalization show that RL enables a 7B‑parameter VLM agent to surpass supervised fine‑tuning baselines, yielding a 26.1% improvement on unseen instances but only limited gains on unseen templates (15.7%) and apps (8.3%), underscoring the challenges of generalization. As a preliminary step, we demonstrate that few‑shot adaptation at test‑time improves performance on unseen apps, motivating future research in this direction. To support reproducibility and fair comparison, we open‑source the full RL training system, including the environment, task suite, models, prompt configurations, and the underlying infrastructure \footnotehttps://github.com/zihuanjiang/AndroidWorld‑Generalization.
Authors:Shanshan Wan, Lai Kang, Yingmei Wei, Tianrui Shen, Haixuan Wang, Chao Zuo
Abstract:
Visual place recognition (VPR) aiming at predicting the location of an image based solely on its visual features is a fundamental task in robotics and autonomous systems. Domain variation remains one of the main challenges in VPR and is relatively unexplored. Existing VPR models attempt to achieve domain agnosticism either by training on large‑scale datasets that inherently contain some domain variations, or by being specifically adapted to particular target domains. In practice, the former lacks explicit domain supervision, while the latter generalizes poorly to unseen domain shifts. This paper proposes a novel query‑based domain‑agnostic VPR model called QdaVPR. First, a dual‑level adversarial learning framework is designed to encourage domain invariance for both the query features forming the global descriptor and the image features from which these query features are derived. Then, a triplet supervision based on query combinations is designed to enhance the discriminative power of the global descriptors. To support the learning process, we augment a large‑scale VPR dataset using style transfer methods, generating various synthetic domains with corresponding domain labels as auxiliary supervision. Extensive experiments show that QdaVPR achieves state‑of‑the‑art performance on multiple VPR benchmarks with significant domain variations. Specifically, it attains the best Recall@1 and Recall@10 on nearly all test scenarios: 93.5%/98.6% on Nordland (seasonal changes), 97.5%/99.0% on Tokyo24/7 (day‑night transitions), and the highest Recall@1 across almost all weather conditions on the SVOX dataset. Our code will be released at https://github.com/shuimushan/QdaVPR.
Authors:Abbas Mammadov, So Takao, Bohan Chen, Ricardo Baptista, Morteza Mardani, Yee Whye Teh, Julius Berner
Abstract:
Flow maps enable high‑quality image generation in a single forward pass. However, unlike iterative diffusion models, their lack of an explicit sampling trajectory impedes incorporating external constraints for conditional generation and solving inverse problems. We put forth Variational Flow Maps, a framework for conditional sampling that shifts the perspective of conditioning from "guiding a sampling path", to that of "learning the proper initial noise". Specifically, given an observation, we seek to learn a noise adapter model that outputs a noise distribution, so that after mapping to the data space via flow map, the samples respect the observation and data prior. To this end, we develop a principled variational objective that jointly trains the noise adapter and the flow map, improving noise‑data alignment, such that sampling from complex data posterior is achieved with a simple adapter. Experiments on various inverse problems show that VFMs produce well‑calibrated conditional samples in a single (or few) steps. For ImageNet, VFM attains competitive fidelity while accelerating the sampling by orders of magnitude compared to alternative iterative diffusion/flow models. Code is available at https://github.com/abbasmammadov/VFM
Authors:Reo Fukunaga, Soh Yoshida, Mitsuji Muneyasu
Abstract:
Deep neural networks are prone to memorizing incorrect labels during training, which degrades their generalizability. Although recent methods have combined sample selection with semi‑supervised learning (SSL) to exploit the memorization effect ‑‑ where networks learn from clean data before noisy data ‑‑ they cannot correct selection errors once a sample is misclassified. To overcome this, we propose asymmetric co‑teaching with different architectures (ACD)‑U, an asymmetric co‑teaching framework that uses different model architectures and incorporates machine unlearning. ACD‑U addresses this limitation through two core mechanisms. First, its asymmetric co‑teaching pairs a contrastive language‑image pretraining (CLIP)‑pretrained vision Transformer with a convolutional neural network (CNN), leveraging their complementary learning behaviors: the pretrained model provides stable predictions, whereas the CNN adapts throughout training. This asymmetry, where the vision Transformer is trained only on clean samples and the CNN is trained through SSL, effectively mitigates confirmation bias. Second, selective unlearning enables post‑hoc error correction by identifying incorrectly memorized samples through loss trajectory analysis and CLIP consistency checks, and then removing their influence via Kullback‑‑Leibler divergence‑based forgetting. This approach shifts the learning paradigm from passive error avoidance to active error correction. Experiments on synthetic and real‑world noisy datasets, including CIFAR‑10/100, CIFAR‑N, WebVision, Clothing1M, and Red Mini‑ImageNet, demonstrate state‑of‑the‑art performance, particularly in high‑noise regimes and under instance‑dependent noise. The code is publicly available at https://github.com/meruemon/ACD‑U.
Authors:Zicheng Duan, Jiatong Xia, Zeyu Zhang, Wenbo Zhang, Gengze Zhou, Chenhui Gou, Yefei He, Feng Chen, Xinyu Zhang, Lingqiao Liu
Abstract:
Recent generative video world models aim to simulate visual environment evolution, allowing an observer to interactively explore the scene via camera control. However, they implicitly assume that the world only evolves within the observer's field of view. Once an object leaves the observer's view, its state is "frozen" in memory, and revisiting the same region later often fails to reflect events that should have occurred in the meantime. In this work, we identify and formalize this overlooked limitation as the "out‑of‑sight dynamics" problem, which impedes video world models from representing a continuously evolving world. To address this issue, we propose LiveWorld, a novel framework that extends video world models to support persistent world evolution. Instead of treating the world as static observational memory, LiveWorld models a persistent global state composed of a static 3D background and dynamic entities that continue evolving even when unobserved. To maintain these unseen dynamics, LiveWorld introduces a monitor‑based mechanism that autonomously simulates the temporal progression of active entities and synchronizes their evolved states upon revisiting, ensuring spatially coherent rendering. For evaluation, we further introduce LiveBench, a dedicated benchmark for the task of maintaining out‑of‑sight dynamics. Extensive experiments show that LiveWorld enables persistent event evolution and long‑term scene consistency, bridging the gap between existing 2D observation‑based memory and true 4D dynamic world simulation. The baseline and benchmark will be publicly available at https://zichengduan.github.io/LiveWorld/index.html.
Authors:Li Jin, Yuchen Yang, Weikai Chen, Yujie Wang, Dehao Hao, Tanghui Jia, Yingda Yin, Zeyu Hu, Runze Zhang, Keyang Luo, Li Yuan, Long Quan, Xin Wang, Xueying Qin
Abstract:
3D learning systems implicitly assume that objects occupy a coherent reference frame. Nonetheless, in practice, every asset arrives with an arbitrary global rotation, and models are left to resolve directional ambiguity on their own. This persistent misalignment suppresses pose‑consistent generation, and blocks the emergence of stable directional semantics. To address this issue, we construct \methodName, a massive canonical 3D dataset of 320K objects over 1,156 categories ‑‑ an order‑of‑magnitude increase over prior work. At this scale, directional semantics become statistically learnable: Canoverse improves 3D generation stability, enables precise cross‑modal 3D shape retrieval, and unlocks zero‑shot point‑cloud orientation estimation even for out‑of‑distribution data. This is achieved by a new canonicalization framework that reduces alignment from minutes to seconds per object via compact hypothesis generation and lightweight human discrimination, transforming canonicalization from manual curation into a high‑throughput data generation pipeline. The Canoverse dataset will be publicly released upon acceptance. Project page: https://github.com/123321456‑gif/Canoverse
Authors:Xijun Lu, Hongying Liu, Fanhua Shang, Yanming Hui, Liang Wan
Abstract:
Medical image anomaly detection faces unique challenges due to subtle, heterogeneous anomalies embedded in complex anatomical structures. Through systematic Grad‑CAM analysis, we reveal that discriminative activation maps fail on medical data, unlike their success on industrial datasets, motivating the need for manifold‑level modeling. We propose PDD (Manifold‑Prior Diverse Distillation), a framework that unifies dual‑teacher priors into a shared high‑dimensional manifold and distills this knowledge into dual students with complementary behaviors. Specifically, frozen VMamba‑Tiny and wide‑ResNet50 encoders provide global contextual and local structural priors, respectively. Their features are unified through a Manifold Matching and Unification (MMU) module, while an Inter‑Level Feature Adaption (InA) module enriches intermediate representations. The unified manifold is distilled into two students: one performs layer‑wise distillation via InA for local consistency, while the other receives skip‑projected representations through a Manifold Prior Affine (MPA) module to capture cross‑layer dependencies. A diversity loss prevents representation collapse while maintaining detection sensitivity. Extensive experiments on multiple medical datasets demonstrate that PDD significantly outperforms existing state‑of‑the‑art methods, achieving improvements of up to 11.8%, 5.1%, and 8.5% in AUROC on HeadCT, BrainMRI, and ZhangLab datasets, respectively, and 3.4% in F1 max on the Uni‑Medical dataset, establishing new state‑of‑the‑art performance in medical image anomaly detection. The implementation will be released at https://github.com/OxygenLu/PDD
Authors:Landi He, Xiaoyu Yang, Lijian Xu
Abstract:
Visual tokens dominate inference cost in vision‑language models (VLMs), yet many carry redundant information. Existing pruning methods alleviate this but typically rely on attention magnitude or similarity scores. We reformulate visual token pruning as capacity constrained communication: given a fixed budget K, the model must allocate limited bandwidth to maximally preserve visual information. We propose AutoSelect, which attaches a lightweight Scorer and Denoiser to a frozen VLM and trains with only the standard next token prediction loss, without auxiliary objectives or extra annotations. During training, a variance preserving noise gate modulates each token's information flow according to its predicted importance so that gradients propagate through all tokens; a diagonal attention Denoiser then recovers the perturbed representations. At inference, only the Scorer and a hard top‑K selection remain, adding negligible latency. On ten VLM benchmarks, AutoSelect retains 96.5% of full model accuracy while accelerating LLM prefill by 2.85x with only 0.69 ms overhead, and transfers to different VLM backbones without architecture‑specific tuning. Code is available at https://github.com/MedHK23/AutoSelect.
Authors:Trong-Thang Pham, Loc Nguyen, Anh Nguyen, Hien Nguyen, Ngan Le
Abstract:
Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training data. Re‑prompting rerolls the entire generation trajectory, altering anatomy, texture, and background. Inversion‑based editing methods introduce reconstruction error that causes structural drift. We propose MedSteer, a training‑free activation‑steering framework for endoscopic synthesis. MedSteer identifies a pathology vector for each contrastive prompt pair in the cross‑attention layers of a diffusion transformer. At inference time, it steers image activations along this vector, generating counterfactual pairs from scratch where the only difference is the steered concept. All other structure is preserved by construction. We evaluate MedSteer across three experiments on Kvasir v3 and HyperKvasir. On counterfactual generation across three clinical concept pairs, MedSteer achieves flip rates of 0.800, 0.925, and 0.950, outperforming the best inversion‑based baseline in both concept flip rate and structural preservation. On dye disentanglement, MedSteer achieves 75% dye removal against 20% (PnP) and 10% (h‑Edit). On downstream polyp detection, augmenting with MedSteer counterfactual pairs achieves ViT AUC of 0.9755 versus 0.9083 for quantity‑matched re‑prompting, confirming that counterfactual structure drives the gain. Code is at link https://github.com/phamtrongthang123/medsteer
Authors:Tong Shao, Yusen Fu, Guoying Sun, Jingde Kong, Zhuotao Tian, Jingyong Su
Abstract:
Diffusion Transformers have become a dominant paradigm in visual generation, yet their low inference efficiency remains a key bottleneck hindering further advancement. Among common training‑free techniques, caching offers high acceleration efficiency but often compromises fidelity, whereas pruning shows the opposite trade‑off. Integrating caching with pruning achieves a balance between acceleration and generation quality. However, existing methods typically employ fixed and heuristic schemes to configure caching and pruning strategies. While they roughly follow the overall sensitivity trend of generation models to acceleration, they fail to capture fine‑grained and complex variations, inevitably skipping highly sensitive computations and leading to quality degradation. Furthermore, such manually designed strategies exhibit poor generalization. To address these issues, we propose SODA, a Sensitivity‑Oriented Dynamic Acceleration method that adaptively performs caching and pruning based on fine‑grained sensitivity. SODA builds an offline sensitivity error modeling framework across timesteps, layers, and modules to capture the sensitivity to different acceleration operations. The cache intervals are optimized via dynamic programming with sensitivity error as the cost function, minimizing the impact of caching on model sensitivity. During pruning and cache reuse, SODA adaptively determines the pruning timing and rate to preserve computations of highly sensitive tokens, significantly enhancing generation fidelity. Extensive experiments on DiT‑XL/2, PixArt‑α, and OpenSora demonstrate that SODA achieves state‑of‑the‑art generation fidelity under controllable acceleration ratios. Our code is released publicly at: https://github.com/leaves162/SODA.
Authors:Leilei Wang, Longfei Liu, Xi Shen, Xuanlong Yu, Ying Tiffany He, Fei Richard Yu, Yingyi Chen
Abstract:
Real‑time open‑vocabulary object detection (OVOD) is essential for practical deployment in dynamic environments, where models must recognize a large and evolving set of categories under strict latency constraints. Current real‑time OVOD methods are predominantly built upon YOLO‑style models. In contrast, real‑time DETR‑based methods still lag behind in terms of inference latency, model lightweightness, and overall performance. In this work, we present OV‑DEIM, an end‑to‑end DETR‑style open‑vocabulary detector built upon the recent DEIMv2 framework with integrated vision‑language modeling for efficient open‑vocabulary inference. We further introduce a simple query supplement strategy that improves Fixed AP without compromising inference speed. Beyond architectural improvements, we introduce GridSynthetic, a simple yet effective data augmentation strategy that composes multiple training samples into structured image grids. By exposing the model to richer object co‑occurrence patterns and spatial layouts within a single forward pass, GridSynthetic mitigates the negative impact of noisy localization signals on the classification loss and improves semantic discrimination, particularly for rare categories. Extensive experiments demonstrate that OV‑DEIM achieves state‑of‑the‑art performance on open‑vocabulary detection benchmarks, delivering superior efficiency and notable improvements on challenging rare categories. Code and pretrained models are available at https://github.com/wleilei/OV‑DEIM.
Authors:Zanlin Ni, Yulin Wang, Yeguo Hua, Renping Zhou, Jiayi Guo, Jun Song, Bo Zheng, Gao Huang
Abstract:
Recent advances in image synthesis have been propelled by powerful generative models, such as Masked Generative Transformers (MaskGIT), autoregressive models, diffusion models, and rectified flow models. A common principle behind their success is the decomposition of synthesis into multiple steps. However, this introduces a proliferation of step‑specific parameters (e.g., noise level or temperature at each step). Existing approaches typically rely on manually‑designed rules to manage this complexity, demanding expert knowledge and trial‑and‑error. Furthermore, these static schedules lack the flexibility to adapt to the unique characteristics of each sample, yielding sub‑optimal performance. To address this issue, we present AdaGen, a general, learnable, and sample‑adaptive framework for scheduling the iterative generation process. Specifically, we formulate the scheduling problem as a Markov Decision Process, where a lightweight policy network determines suitable parameters given the current generation state, and can be trained through reinforcement learning. Importantly, we demonstrate that simple reward designs, such as FID or pre‑trained reward models, can be easily hacked and may not reliably guarantee the desired quality or diversity of generated samples. Therefore, we propose an adversarial reward design to guide the training of the policy networks. Finally, we introduce an inference‑time refinement strategy and a controllable fidelity‑diversity trade‑off mechanism to further enhance the performance and flexibility of AdaGen. Comprehensive experiments on four generative paradigms validate the superiority of AdaGen. For example, AdaGen achieves better performance on DiT‑XL with 3 times lower inference cost and improves the FID of VAR from 1.92 to 1.59 with negligible computational overhead.
Authors:Kaiyuan Xu, Fangzhou Hong, Daniel Elson, Baoru Huang
Abstract:
Reconstructing surgical scenes from monocular endoscopic video is critical for advancing robotic‑assisted surgery. However, the application of state‑of‑the‑art general‑purpose reconstruction models is constrained by two key challenges: the lack of supervised training data and performance degradation over long video sequences. To overcome these limitations, we propose SurgCUT3R, a systematic framework that adapts unified 3D reconstruction models to the surgical domain. Our contributions are threefold. First, we develop a data generation pipeline that exploits public stereo surgical datasets to produce large‑scale, metric‑scale pseudo‑ground‑truth depth maps, effectively bridging the data gap. Second, we propose a hybrid supervision strategy that couples our pseudo‑ground‑truth with geometric self‑correction to enhance robustness against inherent data imperfections. Third, we introduce a hierarchical inference framework that employs two specialized models to effectively mitigate accumulated pose drift over long surgical videos: one for global stability and one for local accuracy. Experiments on the SCARED and StereoMIS datasets demonstrate that our method achieves a competitive balance between accuracy and efficiency, delivering near state‑of‑the‑art but substantially faster pose estimation and offering a practical and effective solution for robust reconstruction in surgical environments. Project page: https://chumo‑xu.github.io/SurgCUT3R‑ICRA26/.
Authors:Sarah S. L. Chow, Rui Wang, Robert B. Serafin, Yujie Zhao, Elena Baraznenok, Xavier Farré, Jennifer Salguero-Lopez, Gan Gao, Huai-Ching Hsieh, Lawrence D. True, Priti Lal, Anant Madabhushi, Jonathan T. C. Liu
Abstract:
Diagnostic grading of prostate cancer (PCa) relies on the examination of 2D histology sections. However, the limited sampling of specimens afforded by 2D histopathology, and ambiguities when viewing 2D cross‑sections, can lead to suboptimal treatment decisions. Recent studies have shown that 3D histomorphometric analysis of glands and nuclei can improve PCa risk assessment compared to analogous 2D features. Here, we expand on these efforts by developing an analytical pipeline to extract 3D features related to perineural invasion (PNI) and lymphovascular invasion (LVI), which correlate with poor prognosis for a variety of cancers. A 3D segmentation model (nnU‑Net) was trained to segment nerves and vessels in 3D datasets of archived prostatectomy specimens that were optically cleared, labeled with a fluorescent analog of H&E, and imaged with open‑top light‑sheet (OTLS) microscopy. PNI‑ and LVI‑related features, including metrics describing cancer‑nerve and cancer‑vessel proximity, were then extracted based on the 3D nerve/vessel segmentation masks in conjunction with 3D masks of cancer‑enriched regions. As a preliminary exploration of the prognostic value of these features, we trained a supervised machine learning classifier to predict 5‑year biochemical recurrence (BCR) outcomes, finding that 3D PNI‑related features are moderately prognostic and outperform 2D PNI‑related features (AUC = 0.71 vs. 0.52). Source code is available at https://github.com/sarahrahsl/SegCIA.git.
Authors:Hang Zhou, Xinxin Zuo, Sen Wang, Li Cheng
Abstract:
Despite strong single‑turn performance, diffusion‑based image compositing often struggles to preserve coherent spatial relations in pairwise or sequential edits, where subsequent insertions may overwrite previously generated content and disrupt physical consistency. We introduce PICS, a self‑supervised composition‑by‑decomposition paradigm that composes objects in parallel while explicitly modeling the compositional interactions among (fully‑/partially‑)visible objects and background. At its core, an Interaction Transformer employs mask‑guided Mixture‑of‑Experts to route background, exclusive, and overlap regions to dedicated experts, with an adaptive α‑blending strategy that infers a compatibility‑aware fusion of overlapping objects while preserving boundary fidelity. To further enhance robustness to geometric variations, we incorporate geometry‑aware augmentations covering both out‑of‑plane and in‑plane pose changes of objects. Our method delivers superior pairwise compositing quality and substantially improved stability, with extensive evaluations across virtual try‑on, indoor, and street scene settings showing consistent gains over state‑of‑the‑art baselines. Code and data are available at https://github.com/RyanHangZhou/PICS
Authors:Weronika Smolak-Dyżewska, Joanna Kaleta, Diego Dall'Alba, Przemysław Spurek
Abstract:
Accurate 3D reconstruction of colonoscopy data, accounting for complex peristaltic movements, is crucial for advanced surgical navigation and retrospective diagnostics. While recent novel view synthesis and 3D reconstruction methods have demonstrated remarkable success in general endoscopic scenarios, they struggle in the highly constrained environment of the colon. Due to the limited field of view of a camera moving through an actively deforming tubular structure, existing endoscopic methods reconstruct the colon appearance only for initial camera trajectory. However, the underlying anatomy remains largely static; instead of updating Gaussians' spatial coordinates (xyz), these methods encode deformation through either rotation, scale or opacity adjustments. In this paper, we first present a benchmark analysis of state‑of‑the‑art dynamic endoscopic methods for realistic colonoscopic scenes, showing that they fail to model true anatomical motion. To enable rigorous evaluation of global reconstruction quality, we introduce DynamicColon, a synthetic dataset with ground‑truth point clouds at every timestep. Building on these insights, we propose ColonSplat, a dynamic Gaussian Splatting framework that captures peristaltic‑like motion while preserving global geometric consistency, achieving superior geometric fidelity on C3VDv2 and DynamicColon datasets. Project page: https://wmito.github.io/ColonSplat
Authors:Zhenyuan Chen, Guanyuan Shen, Feng Zhang
Abstract:
Cross‑modal image‑to‑image translation among Electro‑Optical (EO), Infrared (IR), and Synthetic Aperture Radar (SAR) sensors is essential for comprehensive multi‑modal aerial‑view analysis. However, translating between these modalities is notoriously difficult due to their distinct electromagnetic signatures and geometric characteristics. This paper presents EarthBridge, a high‑fidelity translation framework developed for the 4th Multi‑modal Aerial View Image Challenge ‑‑ Translation (MAVIC‑T). We explore two distinct methodologies: Diffusion Bridge Implicit Models (DBIM), which we generalize using non‑Markovian bridge processes for high‑quality deterministic sampling, and Contrastive Unpaired Translation (CUT), which utilizes contrastive learning for structural consistency. Our EarthBridge framework employs a channel‑concatenated UNet denoiser trained with Karras‑weighted bridge scalings and a specialized "booting noise" initialization to handle the inherent ambiguity in cross‑modal mappings. We evaluate these methods across all four challenge tasks (SAR\rightarrowEO, SAR\rightarrowRGB, SAR\rightarrowIR, RGB\rightarrowIR), achieving superior spatial detail and spectral accuracy. Our solution achieved a composite score of 0.38, securing the second position on the MAVIC‑T leaderboard. Code is available at https://github.com/Bili‑Sakura/EarthBridge‑Preview.
Authors:Neil Tripathi
Abstract:
We present VB, a benchmark that tests whether vision‑language models can determine what is and is not visible in a photograph, and abstain when a human viewer cannot reliably answer. Each item pairs a single photo with a short yes/no visibility claim; the model must output VISIBLY_TRUE, VISIBLY_FALSE, or ABSTAIN, together with a confidence score. Items are organized into 100 families using a 2x2 design that crosses a minimal image edit with a minimal text edit, yielding 300 headline evaluation cells. Unlike prior unanswerable‑VQA benchmarks, VB tests not only whether a question is unanswerable but why (via reason codes tied to specific visibility factors), and uses controlled minimal edits to verify that model judgments change when and only when the underlying evidence changes. We score models on confidence‑aware accuracy with abstention (CAA), minimal‑edit flip rate (MEFR), confidence‑ranked selective prediction (SelRank), and second‑order perspective reasoning (ToMAcc); all headline numbers are computed on the strict XOR subset (three cells per family, 300 scored items per model). We evaluate nine models spanning flagship and prior‑generation closed‑source systems, and open‑source models from 8B to 12B parameters. GPT‑4o and Gemini 3.1 Pro effectively tie for the best composite score (0.728 and 0.727), followed by Gemini 2.5 Pro (0.678). The best open‑source model, Gemma 3 12B (0.505), surpasses one prior‑generation closed‑source system. Text‑flip robustness exceeds image‑flip robustness for six of nine models, and confidence calibration varies substantially: GPT‑4o and Gemini 2.5 Pro achieve similar accuracy yet differ sharply in selective prediction quality.
Authors:Zhen Lin, Qiujie Xie, Minjun Zhu, Shichen Li, Qiyao Sun, Enhao Gu, Yiran Ding, Ke Sun, Fang Guo, Panzhong Lu, Zhiyuan Ning, Yixuan Weng, Yue Zhang
Abstract:
High‑quality scientific illustrations are essential for communicating complex scientific and technical concepts, yet existing automated systems remain limited in editability, stylistic controllability, and efficiency. We present AutoFigure‑Edit, an end‑to‑end system that generates fully editable scientific illustrations from long‑form scientific text while enabling flexible style adaptation through user‑provided reference images. By combining long‑context understanding, reference‑guided styling, and native SVG editing, it enables efficient creation and refinement of high‑quality scientific illustrations. To facilitate further progress in this field, we release the video at https://youtu.be/10IH8SyJjAQ, full codebase at https://github.com/ResearAI/AutoFigure‑Edit and provide a website for easy access and interactive use at https://deepscientist.cc/.
Authors:Yuan Wu, Zongxian Yang, Jiayu Qian, Songpan Gao, Guanxing Chen, Qiankun Li, Yu-An Huang, Zhi-An Huang
Abstract:
Large vision‑language models (VLMs) often benefit from chain‑of‑thought (CoT) prompting in general domains, yet its efficacy in medical vision‑language tasks remains underexplored. We report a counter‑intuitive trend: on medical visual question answering, CoT frequently underperforms direct answering (DirA) across general‑purpose and medical‑specific models. We attribute this to a \emphmedical perception bottleneck: subtle, domain‑specific cues can weaken visual grounding, and CoT may compound early perceptual uncertainty rather than correct it. To probe this hypothesis, we introduce two training‑free, inference‑time grounding interventions: (i) \emphperception anchoring via region‑of‑interest cues and (ii) \emphdescription grounding via high‑quality textual guidance. Across multiple benchmarks and model families, these interventions improve accuracy, mitigate CoT degradation, and in several settings reverse the CoT‑‑DirA inversion. Our findings suggest that reliable clinical VLMs require robust visual grounding and cross‑modal alignment, beyond extending text‑driven reasoning chains. Code is available \hrefhttps://github.com/TianYin123/Better_Eyes_Better_Thoughtshere.
Authors:Linfeng Ye, Shayan Mohajer Hamidi, Zhixiang Chi, Guang Li, Mert Pilanci, Takahiro Ogawa, Miki Haseyama, Konstantinos N. Plataniotis
Abstract:
Attention‑based multiple instance learning (MIL) has emerged as a powerful framework for whole slide image (WSI) diagnosis, leveraging attention to aggregate instance‑level features into bag‑level predictions. Despite this success, we find that such methods exhibit a new failure mode: unstable attention dynamics. Across four representative attention‑based MIL methods and two public WSI datasets, we observe that attention distributions oscillate across epochs rather than converging to a consistent pattern, degrading performance. This instability adds to two previously reported challenges: overfitting and over‑concentrated attention distribution. To simultaneously overcome these three limitations, we introduce attention‑stabilized multiple instance learning (ASMIL), a novel unified framework. ASMIL uses an anchor model to stabilize attention, replaces softmax with a normalized sigmoid function in the anchor to prevent over‑concentration, and applies token random dropping to mitigate overfitting. Extensive experiments demonstrate that ASMIL achieves up to a 6.49% F1 score improvement over state‑of‑the‑art methods. Moreover, integrating the anchor model and normalized sigmoid into existing attention‑based MIL methods consistently boosts their performance, with F1 score gains up to 10.73%. All code and data are publicly available at https://github.com/Linfeng‑Ye/ASMIL.
Authors:Kuan Zhang, Dongchen Liu, Qiyue Zhao, Jinkun Hou, Xinran Zhang, Qinlei Xie, Miao Liu, Yiming Li
Abstract:
Human gameplay is a visually grounded interaction loop in which players act, reflect on failures, and watch tutorials to refine strategies. Can Vision‑Language Models (VLMs) also learn from video‑based reflection? We present GameVerse, a comprehensive video game benchmark that enables a reflective visual interaction loop. Moving beyond traditional fire‑and‑forget evaluations, it uses a novel reflect‑and‑retry paradigm to assess how VLMs internalize visual experience and improve policies. To facilitate systematic and scalable evaluation, we also introduce a cognitive hierarchical taxonomy spanning 15 globally popular games, dual action space for both semantic and GUI control, and milestone evaluation using advanced VLMs to quantify progress. Our experiments show that VLMs benefit from video‑based reflection in varied settings, and perform best by combining failure trajectories and expert tutorials‑a training‑free analogue to reinforcement learning (RL) plus supervised fine‑tuning (SFT).Our project page is available at https://gameverse‑bench.github.io/ . Our code is available at https://github.com/THUSI‑Lab/GameVerse .
Authors:Nikita Kisel, Illia Volkov, Klara Janouskova, Jiri Matas
Abstract:
Multimodal Large Language Models (MLLM) classification performance depends critically on evaluation protocol and ground truth quality. Studies comparing MLLMs with supervised and vision‑language models report conflicting conclusions, and we show these conflicts stem from protocols that either inflate or underestimate performance. Across the most common evaluation protocols, we identify and fix key issues: model outputs that fall outside the provided class list and are discarded, inflated results from weak multiple‑choice distractors, and an open‑world setting that underperforms only due to poor output mapping. We additionally quantify the impact of commonly overlooked design choices ‑ batch size, image ordering, and text encoder selection ‑ showing they substantially affect accuracy. Evaluating on ReGT, our multilabel reannotation of 625 ImageNet‑1k classes, reveals that MLLMs benefit most from corrected labels (up to +10.8%), substantially narrowing the perceived gap with supervised models. Much of the reported MLLMs underperformance on classification is thus an artifact of noisy ground truth and flawed evaluation protocol rather than genuine model deficiency. Models less reliant on supervised training signals prove most sensitive to annotation quality. Finally, we show that MLLMs can assist human annotators: in a controlled case study, annotators confirmed or integrated MLLMs predictions in approximately 50% of difficult cases, demonstrating their potential for large‑scale dataset curation. This work is part of the Aiming for Perfect ImageNet‑1k project, see https://klarajanouskova.github.io/ImageNet/.
Authors:Vishal Thengane, Zhaochong An, Tianjin Huang, Son Lam Phung, Abdesselam Bouzerdoum, Lu Yin, Na Zhao, Xiatian Zhu
Abstract:
Incremental Few‑Shot (IFS) segmentation aims to learn new categories over time from only a few annotations. Although widely studied in 2D, it remains underexplored for 3D point clouds. Existing methods suffer from catastrophic forgetting or fail to learn discriminative prototypes under sparse supervision, and often overlook a key cue: novel categories frequently appear as unlabelled background in base‑training scenes. We introduce SCOPE (Scene‑COntextualised Prototype Enrichment), a plug‑and‑play background‑guided prototype enrichment framework that integrates with any prototype‑based 3D segmentation method. After base training, a class‑agnostic segmentation model extracts high‑confidence pseudo‑instances from background regions to build a prototype pool. When novel classes arrive with few labelled samples, relevant background prototypes are retrieved and fused with few‑shot prototypes to form enriched representations without retraining the backbone or adding parameters. Experiments on ScanNet and S3DIS show that SCOPE achieves SOTA performance, improving novel‑class IoU by up to 6.98% and 3.61%, and mean IoU by 2.25% and 1.70%, respectively, while maintaining low forgetting. Code is available https://github.com/Surrey‑UP‑Lab/SCOPE.
Authors:Boqiang Zhang, Lei Ke, Ruihan Yang, Qi Gao, Tianyuan Qu, Rossell Chen, Dong Yu, Leoweiliang
Abstract:
Vision Language Model (VLM) development has largely relied on scaling model size, which hinders deployment on compute‑constrained mobile and edge devices such as smartphones and robots. In this work, we explore the performance limits of compact (e.g., 2B and 8B) VLMs. We challenge the prevailing practice that state‑of‑the‑art VLMs must rely on vision encoders initialized via massive contrastive pretraining (e.g., CLIP/SigLIP). We identify an objective mismatch: contrastive learning, optimized for discrimination, enforces coarse and category‑level invariances that suppress fine‑grained visual cues needed for dense captioning and complex VLM reasoning. To address this issue, we present Penguin‑VL, whose vision encoder is initialized from a text‑only LLM. Our experiments reveal that Penguin‑Encoder serves as a superior alternative to traditional contrastive pretraining, unlocking a higher degree of visual fidelity and data efficiency for multimodal understanding. Across various image and video benchmarks, Penguin‑VL achieves performance comparable to leading VLMs (e.g., Qwen3‑VL) in mathematical reasoning and surpasses them in tasks such as document understanding, visual knowledge, and multi‑perspective video understanding. Notably, these gains are achieved with a lightweight architecture, demonstrating that improved visual representation rather than model scaling is the primary driver of performance. Our ablations show that Penguin‑Encoder consistently outperforms contrastive‑pretrained encoders, preserving fine‑grained spatial and temporal cues that are critical for dense perception and complex reasoning. This makes it a strong drop‑in alternative for compute‑efficient VLMs and enables high performance in resource‑constrained settings. Code: https://github.com/tencent‑ailab/Penguin‑VL
Authors:Yuhan Zhou, Mehri Sattari, Haihua Chen, Kewei Sha
Abstract:
Next‑generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal (M^2) data to support real‑time decision‑making. In practice, data quality (DQ) varies across sources and modalities due to environmental conditions and sensor limitations, yet AV research has largely prioritized algorithm design over DQ analysis. This work focuses on redundancy as a fundamental but underexplored DQ issue in AV datasets. Using the nuScenes and Argoverse 2 (AV2) datasets, we model and measure redundancy in multisource camera data and multimodal image‑LiDAR data, and evaluate how removing redundant labels affects the YOLOv8 object detection task. Experimental results show that selectively removing redundant multisource image object labels from cameras with shared fields of view improves detection. In nuScenes, mAP50 gains from 0.66 to 0.70, 0.64 to 0.67, and from 0.53 to 0.55, on three representative overlap regions, while detection on other overlapping camera pairs remains at the baseline even under stronger pruning. In AV2, 4.1‑8.6% of labels are removed, and mAP50 stays near the 0.64 baseline. Multimodal analysis also reveals substantial redundancy between image and LiDAR data. These findings demonstrate that redundancy is a measurable and actionable DQ factor with direct implications for AV performance. This work highlights the role of redundancy as a data quality factor in AV perception and motivates a data‑centric perspective for evaluating and improving AV datasets. Code, data, and implementation details are publicly available at: https://github.com/yhZHOU515/RedundancyAD
Authors:Ashkan Shahbazi, Elaheh Akbari, Kyvia Pereira, Jon S. Heiselman, Annie C. Benson, Garrison L. H. Johnston, Jie Ying Wu, Nabil Simaan, Michael I. Miga, Soheil Kolouri
Abstract:
We introduce SurgFormer, a multiresolution gated transformer for data driven soft tissue simulation on volumetric meshes. High fidelity biomechanical solvers are often too costly for interactive use, so we train SurgFormer on solver generated data to predict nodewise displacement fields at near real time rates. SurgFormer builds a fixed mesh hierarchy and applies repeated multibranch blocks that combine local message passing, coarse global self attention, and pointwise feedforward updates, fused by learned per node, per channel gates to adaptively integrate local and long range information while remaining scalable on large meshes. For cut conditioned simulation, resection information is encoded as a learned cut embedding and provided as an additional input, enabling a unified model for both standard deformation prediction and topology altering cases. We also introduce two surgical simulation datasets generated under a unified protocol with XFEM based supervision: a cholecystectomy resection dataset and an appendectomy manipulation and resection dataset with cut and uncut cases. To our knowledge, this is the first learned volumetric surrogate setting to study XFEM supervised cut conditioned deformation within the same volumetric pipeline as standard deformation prediction. Across diverse baselines, SurgFormer achieves strong accuracy with favorable efficiency, making it a practical backbone for both tasks. Code, data, and project page: \hrefhttps://mint‑vu.github.io/SurgFormer/available here
Authors:Yitong Chen, Zuxuan Wu, Xipeng Qiu, Yu-Gang Jiang
Abstract:
Autoregressive (AR) language models rely on causal tokenization, but extending this paradigm to vision remains non‑trivial. Current visual tokenizers either flatten 2D patches into non‑causal sequences or enforce heuristic orderings that misalign with the "next‑token prediction" pattern. Recent diffusion autoencoders similarly fall short: conditioning the decoder on all tokens lacks causality, while applying nested dropout mechanism introduces imbalance. To address these challenges, we present CaTok, a 1D causal image tokenizer with a MeanFlow decoder. By selecting tokens over time intervals and binding them to the MeanFlow objective, as illustrated in Fig. 1, CaTok learns causal 1D representations that support both fast one‑step generation and high‑fidelity multi‑step sampling, while naturally capturing diverse visual concepts across token intervals. To further stabilize and accelerate training, we propose a straightforward regularization REPA‑A, which aligns encoder features with Vision Foundation Models (VFMs). Experiments demonstrate that CaTok achieves state‑of‑the‑art results on ImageNet reconstruction, reaching 0.75 FID, 22.53 PSNR and 0.674 SSIM with fewer training epochs, and the AR model attains performance comparable to leading approaches.
Authors:Maëlic Neau, Zoe Falomir
Abstract:
Scene Graph Generation (SGG) is a task that encodes visual relationships between objects in images as graph structures. SGG shows significant promise as a foundational component for downstream tasks, such as reasoning for embodied agents. To enable real‑time applications, SGG must address the trade‑off between performance and inference speed. However, current methods tend to focus on one of the following: (1) improving relation prediction accuracy, (2) enhancing object detection accuracy, or (3) reducing latency, without aiming to balance all three objectives simultaneously. To address this limitation, we build on the powerful Real‑time Efficiency and Accuracy Compromise for Tradeoffs in Scene Graph Generation (REACT) architecture and propose REACT++, a new state‑of‑the‑art model for real‑time SGG. By leveraging efficient feature extraction and subject‑to‑object cross‑attention within the prototype space, REACT++ balances latency and representational power. REACT++ achieves the highest inference speed among existing SGG models, improving relation prediction accuracy without sacrificing object detection performance. Compared to the previous REACT version, REACT++ is 20% faster with a gain of 10% in relation prediction accuracy on average. The code is available at https://github.com/Maelic/SGG‑Benchmark.
Authors:Zhen Wang, Youcan Xu, Jun Xiao, Long Chen
Abstract:
Video motion transfer aims to generate a target video that inherits motion patterns from a source video while rendering new scenes. Existing training‑free approaches focus on constructing motion guidance based on the intermediate outputs of pre‑trained T2V models, which results in heavy computational overhead and limited flexibility. In this paper, we present FlowMotion, a novel training‑free framework that enables efficient and flexible motion transfer by directly leveraging the predicted outputs of flow‑based T2V models. Our key insight is that early latent predictions inherently encode rich temporal information. Motivated by this, we propose flow guidance, which extracts motion representations based on latent predictions to align motion patterns between source and generated videos. We further introduce a velocity regularization strategy to stabilize optimization and ensure smooth motion evolution. By operating purely on model predictions, FlowMotion achieves superior time and resource efficiency as well as competitive performance compared with state‑of‑the‑art methods.
Authors:Wenxin Li, Kunyu Peng, Di Wen, Junwei Zheng, Jiale Wei, Mengfei Duan, Yuheng Zhang, Rui Fan, Kailun Yang
Abstract:
3D semantic occupancy prediction is a cornerstone of robotic perception, yet real‑world voxel annotations are inherently corrupted by structural artifacts and dynamic trailing effects. This raises a critical but underexplored question: can autonomous systems safely rely on such unreliable occupancy supervision? To systematically investigate this issue, we establish OccNL, the first benchmark dedicated to 3D occupancy under occupancy‑asymmetric and dynamic trailing noise. Our analysis reveals a fundamental domain gap: state‑of‑the‑art 2D label noise learning strategies collapse catastrophically in sparse 3D voxel spaces, exposing a critical vulnerability in existing paradigms. To address this challenge, we propose DPR‑Occ, a principled label noise‑robust framework that constructs reliable supervision through dual‑source partial label reasoning. By synergizing temporal model memory with representation‑level structural affinity, DPR‑Occ dynamically expands and prunes candidate label sets to preserve true semantics while suppressing noise propagation. Extensive experiments on SemanticKITTI demonstrate that DPR‑Occ prevents geometric and semantic collapse under extreme corruption. Notably, even at 90% label noise, our method achieves significant performance gains (up to 2.57% mIoU and 13.91% IoU) over existing label noise learning baselines adapted to the 3D occupancy prediction task. By bridging label noise learning and 3D perception, OccNL and DPR‑Occ provide a reliable foundation for safety‑critical robotic perception in dynamic environments. The benchmark and source code will be made publicly available at https://github.com/mylwx/OccNL.
Authors:Jingkai Wang, Yixin Tang, Jue Gong, Jiatong Li, Shu Li, Libo Liu, Jianliang Lan, Yutong Liu, Yulun Zhang
Abstract:
Diffusion transformer (DiT) architectures show great potential for real‑world image super‑resolution (Real‑ISR). However, their computationally expensive iterative sampling necessitates one‑step distillation. Existing one‑step distillation methods struggle with Real‑ISR on DiT. They suffer from fundamental trajectory mismatch and generate severe grid‑like periodic artifacts. To tackle these challenges, we propose StrSR, a novel one‑step adversarial distillation framework featuring spectral and trajectory regularization. Specifically, we propose an asymmetric discriminative distillation architecture to bridge the trajectory gap. Additionally, we design a frequency distribution matching strategy to effectively suppress DiT‑specific periodic artifacts caused by high‑frequency spectral leakage. Extensive experiments demonstrate that StrSR achieves state‑of‑the‑art performance in Real‑ISR, across both quantitative metrics and visual perception. The code and models will be released at https://github.com/jkwang28/StrSR .
Authors:Kai Luo, Xu Wang, Rui Fan, Kailun Yang
Abstract:
Generalizing across unknown targets is critical for open‑world perception, yet existing 3D Multi‑Object Tracking (3D MOT) pipelines remain limited by closed‑set assumptions and ``semantic‑blind'' heuristics. To address this, we propose Next‑step Open‑Vocabulary Autoregression (NOVA), an innovative paradigm that shifts 3D tracking from traditional fragmented distance‑based matching toward generative spatio‑temporal semantic modeling. NOVA reformulates 3D trajectories as structured spatio‑temporal semantic sequences, enabling the simultaneous encoding of physical motion continuity and deep linguistic priors. By leveraging the autoregressive capabilities of Large Language Models (LLMs), we transform the tracking task into a principled process of next‑step sequence completion. This mechanism allows the model to explicitly utilize the hierarchical structure of language space to resolve fine‑grained semantic ambiguities and maintain identity consistency across complex long‑range sequences through high‑level commonsense reasoning. Extensive experiments on nuScenes, V2X‑Seq‑SPD, and KITTI demonstrate the superior performance of NOVA. Notably, on the nuScenes dataset, NOVA achieves an AMOTA of 22.41% for Novel categories, yielding a significant 20.21% absolute improvement over the baseline. These gains are realized through a compact 0.5B autoregressive model. Code will be available at https://github.com/xifen523/NOVA.
Authors:Han-Chen Zhang, Zi-Hao Zhou, Mao-Lin Luo, Shimin Di, Min-Ling Zhang, Tong Wei
Abstract:
Model merging aims to integrate multiple task‑adapted models into a unified model that preserves the knowledge of each task. In this paper, we identify that the key to this knowledge retention lies in maintaining the directional consistency of singular spaces between merged multi‑task vector and individual task vectors. However, this consistency is frequently compromised by two issues: i) an imbalanced energy distribution within task vectors, where a small fraction of singular values dominate the total energy, leading to the neglect of semantically important but weaker components upon merging, and ii) the geometric inconsistency of task vectors in parameter space, which causes direct merging to distort their underlying directional geometry. To address these challenges, we propose DC‑Merge, a method for directional‑consistent model merging. It first balances the energy distribution of each task vector by smoothing its singular values, ensuring all knowledge components are adequately represented. These energy‑balanced vectors are then projected onto a shared orthogonal subspace to align their directional geometries with minimal reconstruction error. Finally, the aligned vectors are aggregated in the shared orthogonal subspace and projected back to the original parameter space. Extensive experiments on vision and vision‑language benchmarks show that DC‑Merge consistently achieves state‑of‑the‑art performance in both full fine‑tuning and LoRA settings. The implementation code is available at https://github.com/Tobeginwith/DC‑Merge.
Authors:Mingyu Fan, Yi Liu, Hao Zhou, Deheng Qian, Mohammad Haziq Khan, Matthias Raetsch
Abstract:
Trajectory prediction is essential for autonomous driving, enabling vehicles to anticipate the motion of surrounding agents to support safe planning. However, most existing predictors assume fixed‑length histories and suffer substantial performance degradation when observations are variable or extremely short in real‑world settings (e.g., due to occlusion or a limited sensing range). We propose TaPD (Temporal‑adaptive Progressive Distillation), a unified plug‑and‑play framework for observation‑adaptive trajectory forecasting under variable history lengths. TaPD comprises two cooperative modules: an Observation‑Adaptive Forecaster (OAF) for future prediction and a Temporal Backfilling Module (TBM) for explicit reconstruction of the past. OAF is built on progressive knowledge distillation (PKD), which transfers motion pattern knowledge from long‑horizon "teachers" to short‑horizon "students" via hierarchical feature regression, enabling short observations to recover richer motion context. We further introduce a cosine‑annealed distillation weighting scheme to balance forecasting supervision and feature alignment, improving optimization stability and cross‑length consistency. For extremely short histories where implicit alignment is insufficient, TBM backfills missing historical segments conditioned on scene evolution, producing context‑rich trajectories that strengthen PKD and thereby improve OAF. We employ a decoupled pretrain‑reconstruct‑finetune protocol to preserve real‑motion priors while adapting to backfilled inputs. Extensive experiments on Argoverse 1 and Argoverse 2 show that TaPD consistently outperforms strong baselines across all observation lengths, delivers especially large gains under very short inputs, and improves other predictors (e.g., HiVT) in a plug‑and‑play manner. Code will be available at https://github.com/zhouhao94/TaPD.
Authors:Xiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang, Jun Yu
Abstract:
Multimodal Summarization (MMS) aims to generate concise textual summaries by understanding and integrating information across videos, transcripts, and images. However, existing approaches still suffer from three main challenges: (1) reliance on domain‑specific supervision, (2) implicit fusion with weak cross‑modal grounding, and (3) flat temporal modeling without event transitions. To address these issues, we introduce CoE, a training‑free MMS framework that performs structured reasoning through a Chain‑of‑Events guided by a Hierarchical Event Graph (HEG). The HEG encodes textual semantics into an explicit event hierarchy that scaffolds cross‑modal grounding and temporal reasoning. Guided by this structure, CoE localizes key visual cues, models event evolution and causal transitions, and refines outputs via lightweight style adaptation for domain alignment. Extensive experiments on eight diverse datasets demonstrate that CoE consistently outperforms state‑of‑the‑art video CoT baselines, achieving average gains of +3.04 ROUGE, +9.51 CIDEr, and +1.88 BERTScore, highlighting its robustness, interpretability, and cross‑domain generalization. Our code is available at https://github.com/youxiaoxing/CoE.
Authors:Siyan Fang, Yuntao Wang, Jinpu Zhang, Ziwen Li, Yuehuan Wang
Abstract:
Existing image reflection removal methods struggle to handle complex reflections. Accurate language descriptions can help the model understand the image content to remove complex reflections. However, due to blurred and distorted interferences in reflected images, machine‑generated language descriptions of the image content are often inaccurate, which harms the performance of language‑guided reflection removal. To address this, we propose the Adaptive Language‑Aware Network (ALANet) to remove reflections even with inaccurate language inputs. Specifically, ALANet integrates both filtering and optimization strategies. The filtering strategy reduces the negative effects of language while preserving its benefits, whereas the optimization strategy enhances the alignment between language and visual features. ALANet also utilizes language cues to decouple specific layer content from feature maps, improving its ability to handle complex reflections. To evaluate the model's performance under complex reflections and varying levels of language accuracy, we introduce the Complex Reflection and Language Accuracy Variance (CRLAV) dataset. Experimental results demonstrate that ALANet surpasses state‑of‑the‑art methods for image reflection removal. The code and dataset are available at https://github.com/fashyon/ALANet.
Authors:Mohammed Baharoon, Thibault Heintz, Siavash Raissi, Mahmoud Alabbad, Mona Alhammad, Hassan AlOmaish, Sung Eun Kim, Oishi Banerjee, Pranav Rajpurkar
Abstract:
We introduce CRIMSON, a clinically grounded evaluation framework for chest X‑ray report generation that assesses reports based on diagnostic correctness, contextual relevance, and patient safety. Unlike prior metrics, CRIMSON incorporates full clinical context, including patient age, indication, and guideline‑based decision rules, and prevents normal or clinically insignificant findings from exerting disproportionate influence on the overall score. The framework categorizes errors into a comprehensive taxonomy covering false findings, missing findings, and eight attribute‑level errors (e.g., location, severity, measurement, and diagnostic overinterpretation). Each finding is assigned a clinical significance level (urgent, actionable non‑urgent, non‑actionable, or expected/benign), based on a guideline developed in collaboration with attending cardiothoracic radiologists, enabling severity‑aware weighting that prioritizes clinically consequential mistakes over benign discrepancies. CRIMSON is validated through strong alignment with clinically significant error counts annotated by six board‑certified radiologists in ReXVal (Kendalls tau = 0.61‑0.71; Pearsons r = 0.71‑0.84), and through two additional benchmarks that we introduce. In RadJudge, a targeted suite of clinically challenging pass‑fail scenarios, CRIMSON shows consistent agreement with expert judgment. In RadPref, a larger radiologist preference benchmark of over 100 pairwise cases with structured error categorization, severity modeling, and 1‑5 overall quality ratings from three cardiothoracic radiologists, CRIMSON achieves the strongest alignment with radiologist preferences. We release the metric, the evaluation benchmarks, RadJudge and RadPref, and a fine‑tuned MedGemma model to enable reproducible evaluation of report generation, all available at https://github.com/rajpurkarlab/CRIMSON.
Authors:Benyuan Meng, Qianqian Xu, Zitai Wang, Xiaochun Cao, Longtao Huang, Qingming Huang
Abstract:
As powerful generative models, text‑to‑image diffusion models have recently been explored for discriminative tasks. A line of research focuses on adapting a pre‑trained diffusion model to semantic segmentation without any further training, leading to training‑free diffusion segmentors. These methods typically rely on cross‑attention maps from the model's attention layers, which are assumed to capture semantic relationships between image pixels and text tokens. Ideally, such approaches should benefit from more powerful diffusion models, i.e., stronger generative capability should lead to better segmentation. However, we observe that existing methods often fail to scale accordingly. To understand this issue, we identify two underlying gaps: (i) cross‑attention is computed across multiple heads and layers, but there exists a discrepancy between these individual attention maps and a unified global representation. (ii) Even when a global map is available, it does not directly translate to accurate semantic correlation for segmentation, due to score imbalances among different text tokens. To bridge these gaps, we propose two techniques: auto aggregation and per‑pixel rescaling, which together enable training‑free segmentation to better leverage generative capability. We evaluate our approach on standard semantic segmentation benchmarks and further integrate it into a generative technique, demonstrating both improved performance broad applicability. Codes are at https://github.com/Darkbblue/goca.
Authors:Bohai Gu, Taiyi Wu, Dazhao Du, Jian Liu, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo
Abstract:
Modern video editing techniques have achieved high visual fidelity when inserting video objects. However, they focus on optimizing visual fidelity rather than physical causality, leading to edits that are physically inconsistent with their environment. In this work, we present Place‑it‑R1, an end‑to‑end framework for video object insertion that unlocks the environment‑aware reasoning potential of Multimodal Large Language Models (MLLMs). Our framework leverages the Chain‑of‑Thought (CoT) reasoning of MLLMs to orchestrate video diffusion, following a Think‑then‑Place paradigm. To bridge cognitive reasoning and generative execution, we introduce three key innovations: First, MLLM performs physical scene understanding and interaction reasoning, generating environment‑aware chain‑of‑thought tokens and inferring valid insertion regions to explicitly guide the diffusion toward physically plausible insertion. Then, we introduce MLLM‑guided Spatial Direct Preference Optimization (DPO), where diffusion outputs are fed back to the MLLM for scoring, enabling visual naturalness. During inference, the MLLM iteratively triggers refinement cycles and elicits adaptive adjustments from the diffusion model, forming a closed‑loop that progressively enhances editing quality. Furthermore, we provide two user‑selectable modes: a plausibility‑oriented flexible mode that permits environment modifications (\eg, generating support structures) to enhance physical plausibility, and a fidelity‑oriented standard mode that preserves scene integrity for maximum fidelity, offering users explicit control over the plausibility‑fidelity trade‑off. Extensive experiments demonstrate Place‑it‑R1 achieves physically‑coherent video object insertion compared with state‑of‑the‑art solutions and commercial models.
Authors:Soumya Mazumdar, Vineet Kumar Rakesh
Abstract:
Diffusion models have recently advanced photorealistic human synthesis, although practical talking‑head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect audio‑visual alignment under challenging speech conditions. This paper introduces TempoSyncDiff, a reference‑conditioned latent diffusion framework that explores few‑step inference for efficient audio‑driven talking‑head generation. The approach adopts a teacher‑student distillation formulation in which a diffusion teacher trained with a standard noise prediction objective guides a lightweight student denoiser capable of operating with significantly fewer inference steps to improve generation stability. The framework incorporates identity anchoring and temporal regularization designed to mitigate identity drift and frame‑to‑frame flicker during synthesis, while viseme‑based audio conditioning provides coarse lip motion control. Experiments on the LRS3 dataset report denoising‑stage component‑level metrics relative to VAE reconstructions and preliminary latency characterization, including CPU‑only and edge computing measurements and feasibility estimates for edge deployment. The results suggest that distilled diffusion models can retain much of the reconstruction behaviour of a stronger teacher while enabling substantially lower latency inference. The study is positioned as an initial step toward practical diffusion‑based talking‑head generation under constrained computational settings. GitHub: https://mazumdarsoumya.github.io/TempoSyncDiff
Authors:Canyu Chen, Yuguang Yang, Zhewen Tan, Yizhi Wang, Ruiyi Zhan, Haiyan Liu, Xuanyao Mao, Jason Bao, Xinyue Tang, Linlin Yang, Bingchuan Sun, Yan Wang, Baochang Zhang
Abstract:
We identify a fundamental Narrow Policy limitation undermining the performance of autonomous VLA models, where driving Imitation Learning (IL) tends to collapse exploration and limit the potential of subsequent Reinforcement Learning (RL) stages, which often saturate prematurely due to insufficient feedback diversity. Thereby, we propose Curious‑VLA, a framework that alleviates the exploit‑explore dilemma through a two‑stage design. During IL, we introduce a Feasible Trajectory Expansion (FTE) strategy to generate multiple physically valid trajectories and a step‑wise normalized trajectory representation to adapt this diverse data. In the RL stage, we present Adaptive Diversity‑Aware Sampling (ADAS) that prioritizes high‑diversity samples and introduce Spanning Driving Reward (SDR) with a focal style weighting to amplify reward's value span for improving sensitivity to driving quality. On the Navsim benchmark, Curious‑VLA achieves SoTA results (PDMS 90.3, EPDMS 85.4) and a Best‑of‑N PDMS of 94.8, demonstrating its effectiveness in unlocking the exploratory potential of VLA models. Code: https://github.com/Mashiroln/curious_vla.git.
Authors:Xuan Huang, Mochu Xiang, Zhelun Shen, Jinbo Wu, Chenming Wu, Chen Zhao, Kaisiyuan Wang, Hang Zhou, Shanshan Liu, Haocheng Feng, Wei He, Jingdong Wang
Abstract:
Hand‑Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across frames. Although recent HOI reenactment approaches have achieved progress, they are typically trained and evaluated in‑domain and fail to generalize to complex, in‑the‑wild scenarios. In contrast, all‑in‑one video editing models exhibit broader robustness but still struggle with HOI‑specific issues such as inconsistent object appearance. In this paper, we present GenHOI, a lightweight augmentation to pretrained video generation models that injects reference‑object information in a temporally balanced and spatially selective manner. For temporal balancing, we propose Head‑Sliding RoPE, which assigns head‑specific temporal offsets to reference tokens, distributing their influence evenly across frames and mitigating the temporal decay of 3D RoPE to improve long‑range object consistency. For spatial selectivity, we design a two‑level spatial attention gate that concentrates object‑conditioned attention on HOI regions and adaptively scales its strength, preserving background realism while enhancing interaction fidelity. Extensive qualitative and quantitative evaluations on unseen, in‑the‑wild scenes demonstrate that GenHOI significantly outperforms state‑of‑the‑art HOI reenactment and all‑in‑one video editing methods. Project page: https://xuanhuang0.github.io/GenHOI/
Authors:Xia Xin, Yuki Endo, Yoshihiro Kanamori
Abstract:
Recent text‑to‑image models can generate high‑quality images from natural‑language prompts, yet controlling typography remains challenging: requested typographic appearance is often ignored or only weakly followed. We address this limitation with a data‑centric approach that trains image generation models using targeted supervision derived from a structured annotation pipeline specialized for typography. Our pipeline constructs a large‑scale typography‑focused dataset, FontUse, consisting of about 70K images annotated with user‑friendly prompts, text‑region locations, and OCR‑recognized strings. The annotations are automatically produced using segmentation models and multimodal large language models (MLLMs). The prompts explicitly combine font styles (e.g., serif, script, elegant) and use cases (e.g., wedding invitations, coffee‑shop menus), enabling intuitive specification even for novice users. Fine‑tuning existing generators with these annotations allows them to consistently interpret style and use‑case conditions as textual prompts without architectural modification. For evaluation, we introduce a Long‑CLIP‑based metric that measures alignment between generated typography and requested attributes. Experiments across diverse prompts and layouts show that models trained with our pipeline produce text renderings more consistent with prompts than competitive baselines. The source code for our annotation pipeline is available at https://github.com/xiaxinz/FontUSE.
Authors:Jiayang Sun, Zixin Guo, Min Cao, Guibo Zhu, Jorma Laaksonen
Abstract:
Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring the rich temporal dynamics of the change procedure, which is the key to understand not only what has changed but also how it occurs. We introduce ProCap, a novel framework that reformulates change modeling from static image comparison to dynamic procedure modeling. ProCap features a two‑stage design: The first stage trains a procedure encoder to learn the change procedure from a sparse set of keyframes. These keyframes are obtained by automatically generating intermediate frames to make the implicit procedural dynamics explicit and then sampling them to mitigate redundancy. Then the encoder learns to capture the latent dynamics of these keyframes via a caption‑conditioned, masked reconstruction task. The second stage integrates this trained encoder within an encoder‑decoder model for captioning. Instead of relying on explicit frames from the previous stage ‑‑ a process incurring computational overhead and sensitivity to visual noise ‑‑ we introduce learnable procedure queries to prompt the encoder for inferring the latent procedure representation, which the decoder then translates into text. The entire model is then trained end‑to‑end with a captioning loss, ensuring the encoder's output is both temporally coherent and captioning‑aligned. Experiments on three datasets demonstrate the effectiveness of ProCap. Code and pre‑trained models are available at https://github.com/BlueberryOreo/ProCap
Authors:Si-Yu Lu, Po-Ting Chen, Hui-Che Hsu, Sin-Ye Jhong, Wen-Huang Cheng, Yung-Yao Chen
Abstract:
Reconstructing 3D geometry from streaming video requires continuous inference under bounded resources. Recent geometric foundation models achieve impressive reconstruction quality through all‑to‑all attention, yet their quadratic cost confines them to short, offline sequences. Causal‑attention variants such as StreamVGGT enable single‑pass streaming but accumulate an ever‑growing KV cache, exhausting GPU memory within hundreds of frames and precluding the long‑horizon deployment that motivates streaming inference in the first place. We present OVGGT, a training‑free framework that bounds both memory and compute to a fixed budget regardless of sequence length. Our approach combines Self‑Selective Caching, which leverages FFN residual magnitudes to compress the KV cache while remaining fully compatible with FlashAttention, with Dynamic Anchor Protection, which shields coordinate‑critical tokens from eviction to suppress geometric drift over extended trajectories. Extensive experiments on indoor, outdoor, and ultra‑long sequence benchmarks demonstrate that OVGGT processes arbitrarily long videos within a constant VRAM envelope while achieving state‑of‑the‑art 3D geometric accuracy. Project page: https://vaisr.github.io/OVGGT/ Code: https://github.com/VAISR/OVGGT
Authors:Hongli Liu, Yu Wang, Shengjie Zhao
Abstract:
Few‑shot segmentation (FSS) has gained significant attention for its ability to generalize to novel classes with limited supervision, yet remains challenged by structural misalignment and cross‑view inconsistency under large appearance or viewpoint variations. This paper tackles these challenges by introducing VINE (View‑Informed NEtwork), a unified framework that jointly models structural consistency and foreground discrimination to refine class‑specific prototypes. Specifically, VINE introduces a spatial‑view graph on backbone features, where the spatial graph captures local geometric topology and the view graph connects features from different perspectives to propagate view‑invariant structural semantics. To further alleviate foreground ambiguity, we derive a discriminative prior from the support‑query feature discrepancy to capture category‑specific contrast, which reweights SAM features by emphasizing salient regions and recalibrates backbone activations for improved structural focus. The foreground‑enhanced SAM features and structurally enriched ResNet features are progressively integrated through masked cross‑attention, yielding class‑consistent prototypes used as adaptive prompts for the SAM decoder to generate accurate masks. Extensive experiments on multiple FSS benchmarks validate the effectiveness and robustness of VINE, particularly under challenging scenarios with viewpoint shifts and complex structures. The code is available at https://github.com/HongliLiu1/VINE‑main.
Authors:Luan Pham, Phu Hao Hoang, Xuan Toan Mai, Tuan Anh Tran
Abstract:
Skew estimation is one of the vital tasks in document processing systems, especially for scanned document images, because its performance impacts subsequent steps directly. Over the years, an enormous number of researches focus on this challenging problem in the rise of digitization age. In this research, we first propose a novel skew estimation method that extracts the dominant skew angle of the given document image by applying an Adaptive Radial Projection on the 2D Discrete Fourier Magnitude spectrum. Second, we introduce a high quality skew estimation dataset DISE‑2021 to assess the performance of different estimators. Finally, we provide comprehensive analyses that focus on multiple improvement aspects of Fourier‑based methods. Our results show that the proposed method is robust, reliable, and outperforms all compared methods. The source code is available at https://github.com/phamquiluan/jdeskew.
Authors:Luan Pham, The Huynh Vu, Tuan Anh Tran
Abstract:
Automatic facial expression recognition (FER) has gained much attention due to its applications in human‑computer interaction. Among the approaches to improve FER tasks, this paper focuses on deep architecture with the attention mechanism. We propose a novel Masking idea to boost the performance of CNN in facial expression task. It uses a segmentation network to refine feature maps, enabling the network to focus on relevant information to make correct decisions. In experiments, we combine the ubiquitous Deep Residual Network and Unet‑like architecture to produce a Residual Masking Network. The proposed method holds state‑of‑the‑art (SOTA) accuracy on the well‑known FER2013 and private VEMO datasets. The source code is available at https://github.com/phamquiluan/ResidualMaskingNetwork.
Authors:Hongwei Fang, Jiahang Cai, Xun Wang, Wenwu Yang
Abstract:
Vision Transformers (ViTs) have recently achieved state‑of‑the‑art performance in 2D human pose estimation due to their strong global modeling capability. However, existing ViT‑based pose estimators are designed for static images and process each frame independently, thereby ignoring the temporal coherence that exists in video sequences. This limitation often results in unstable predictions, especially in challenging scenes involving motion blur, occlusion, or defocus. In this paper, we propose TAR‑ViTPose, a novel Temporal Aggregate‑and‑Restore Vision Transformer tailored for video‑based 2D human pose estimation. TAR‑ViTPose enhances static ViT representations by aggregating temporal cues across frames in a plug‑and‑play manner, leading to more robust and accurate pose estimation. To effectively aggregate joint‑specific features that are temporally aligned across frames, we introduce a joint‑centric temporal aggregation (JTA) that assigns each joint a learnable query token to selectively attend to its corresponding regions from neighboring frames. Furthermore, we develop a global restoring attention (GRA) to restore the aggregated temporal features back into the token sequence of the current frame, enriching its pose representation while fully preserving global context for precise keypoint localization. Extensive experiments demonstrate that TAR‑ViTPose substantially improves upon the single‑frame baseline ViTPose, achieving a +2.3 mAP gain on the PoseTrack2017 benchmark. Moreover, our approach outperforms existing state‑of‑the‑art video‑based methods, while also achieving a noticeably higher runtime frame rate in real‑world applications. Project page: https://github.com/zgspose/TARViTPose.
Authors:Sen Fang, Yalin Feng, Yanxin Zhang, Dimitris N. Metaxas
Abstract:
In this paper, we propose a Rectified Flow Auto Coder (RAC) inspired by Rectified Flow to replace the traditional VAE: 1. It achieves multi‑step decoding by applying the decoder to flow timesteps. Its decoding path is straight and correctable, enabling step‑by‑step refinement. 2. The model inherently supports bidirectional inference, where the decoder serves as the encoder through time reversal (hence Coder rather than encoder or decoder), reducing parameter count by nearly 41%. 3. This generative decoding method improves generation quality since the model can correct latent variables along the path, partially addressing the reconstruction‑‑generation gap. Experiments show that RAC surpasses SOTA VAEs in both reconstruction and generation with approximately 70% lower computational cost.
Authors:Feiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang, Xilin Zhao, Xiaochun Cao, Qingming Huang
Abstract:
This paper investigates the challenging task of detecting backdoored text‑to‑image models under black‑box settings and introduces a novel detection framework BlackMirror. Existing approaches typically rely on analyzing image‑level similarity, under the assumption that backdoor‑triggered generations exhibit strong consistency across samples. However, they struggle to generalize to recently emerging backdoor attacks, where backdoored generations can appear visually diverse. BlackMirror is motivated by an observation: across backdoor attacks, only partial semantic patterns within the generated image are steadily manipulated, while the rest of the content remains diverse or benign. Accordingly, BlackMirror consists of two components: MirrorMatch, which aligns visual patterns with the corresponding instructions to detect semantic deviations; and MirrorVerify, which evaluates the stability of these deviations across varied prompts to distinguish true backdoor behavior from benign responses. BlackMirror is a general, training‑free framework that can be deployed as a plug‑and‑play module in Model‑as‑a‑Service (MaaS) applications. Comprehensive experiments demonstrate that BlackMirror achieves accurate detection across a wide range of attacks. Code is available at https://github.com/Ferry‑Li/BlackMirror.
Authors:Yuxin Xie, Yuming Chen, Yishan Yang, Yi Zhou, Tao Zhou, Zhen Zhao, Jiacheng Liu, Huazhu Fu
Abstract:
Medical image segmentation is undergoing a paradigm shift from conventional visual pattern matching to cognitive reasoning analysis. Although Multimodal Large Language Models (MLLMs) have shown promise in integrating linguistic and visual knowledge, significant gaps remain: existing general MLLMs possess broad common sense but lack the specialized visual reasoning required for complex lesions, whereas traditional segmentation models excel at pixel‑level segmentation but lack logical interpretability. In this paper, we introduce ComLesion‑14K, the first diverse Chain‑of‑Thought (CoT) benchmark for reasoning‑driven complex lesion segmentation. To accomplish this task, we propose CORE‑Seg, an end‑to‑end framework integrating reasoning with segmentation through a Semantic‑Guided Prompt Adapter. We design a progressive training strategy from SFT to GRPO, equipped with an adaptive dual‑granularity reward mechanism to mitigate reward sparsity. Our Method achieves state‑of‑the‑art results with a mean Dice of 37.06% (14.89% higher than the second‑best baseline), while reducing the failure rate to 18.42%. Project Page: https://xyxl024.github.io/CORE‑Seg.github.io/
Authors:Zidian Qiu, Ancong Wu
Abstract:
Current compositional image‑to‑3D scene generation approaches construct 3D scenes by time‑consuming iterative layout optimization or inflexible joint object‑layout generation. Moreover, most methods rely on limited field‑of‑view perspective images, hindering the creation of complete 360‑degree environments. To address these limitations, we design Pano3DComposer, an efficient feed‑forward framework for panoramic images. To decouple object generation from layout estimation, we propose a plug‑and‑play Object‑World Transformation Predictor. This module converts the 3D objects generated by off‑the‑shelf image‑to‑3D models from local to world coordinates. To achieve this, we adapt the VGGT architecture to Alignment‑VGGT by using target object crop, multi‑view object renderings and camera parameters to predict the transformation. The predictor is trained using pseudo‑geometric supervision to address the shape discrepancy between generated and ground‑truth objects. For input images from unseen domains, we further introduce a Coarse‑to‑Fine (C2F) alignment mechanism for Pano3DComposer that iteratively refines geometric consistency with feedback of scene rendering. Our method achieves superior geometric accuracy for image/text‑to‑3D tasks on synthetic and real‑world datasets. It can generate a high‑fidelity 3D scene in approximately 20 seconds on an RTX 4090 GPU. Project page: https://qiuzidian.github.io/pano3dcomposer‑page/.
Authors:Xuecheng Bai, Yuxiang Wang, Chuanzhi Xu, Boyu Hu, Kang Han, Ruijie Pan, Xiaowei Niu, Xiaotian Guan, Liqiang Fu, Pengfei Ye
Abstract:
Small object detection in unmanned aerial vehicle (UAV) imagery is challenging, mainly due to scale variation, structural detail degradation, and limited computational resources. In high‑altitude scenarios, fine‑grained features are further weakened during hierarchical downsampling and cross‑scale fusion, resulting in unstable localization and reduced robustness. To address this issue, we propose CollabOD, a lightweight collaborative detection framework that explicitly preserves structural details and aligns heterogeneous feature streams before multi‑scale fusion. The framework integrates Structural Detail Preservation, Cross‑Path Feature Alignment, and Localization‑Aware Lightweight Design strategies. From the perspectives of image processing, channel structure, and lightweight design, it optimizes the architecture of conventional UAV perception models. The proposed design enhances representation stability while maintaining efficient inference. A unified detail‑aware detection head further improves regression robustness without introducing additional deployment overhead. The code is available at: https://github.com/Bai‑Xuecheng/CollabOD.
Authors:Xiang Zhang, Sohyun Yoo, Hongrui Wu, Chuan Li, Jianwen Xie, Zhuowen Tu
Abstract:
We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post‑hoc layout optimization, PixARMesh jointly predicts object layout and geometry within a unified model, producing coherent and artist‑ready meshes in a single forward pass. Building on recent advances in mesh generative models, we augment a point‑cloud encoder with pixel‑aligned image features and global scene context via cross‑attention, enabling accurate spatial reasoning from a single image. Scenes are generated autoregressively from a unified token stream containing context, pose, and mesh, yielding compact meshes with high‑fidelity geometry. Experiments on synthetic and real‑world datasets show that PixARMesh achieves state‑of‑the‑art reconstruction quality while producing lightweight, high‑quality meshes ready for downstream applications.
Authors:Sijing Li, Zhongwei Qiu, Jiang Liu, Wenqiao Zhang, Tianwei Lin, Yihan Xie, Jianxiang An, Boxiang Yun, Chenglin Yang, Jun Xiao, Guangyu Guo, Jiawen Yao, Wei Liu, Yuan Gao, Ke Yan, Weiwei Cao, Zhilin Zheng, Tony C. W. Mok, Kai Cao, Yu Shi, Jiuyu Zhang, Jian Zhou, Beng Chin Ooi, Yingda Xia, Ling Zhang
Abstract:
Accurate tumor analysis is central to clinical radiology and precision oncology, where early detection, reliable lesion characterization, and pathology‑level risk assessment guide diagnosis and treatment planning. Chain‑of‑Thought (CoT) reasoning is particularly important in this setting because it enables step‑by‑step interpretation from imaging findings to clinical impressions and pathology conclusions, improving traceability and reducing diagnostic errors. Here, we target the clinical tumor analysis task and build a large‑scale benchmark that operationalizes a multimodal reasoning pipeline, spanning findings, impressions, and pathology predictions. We curate TumorCoT, a large‑scale dataset of 1.5M CoT‑labeled VQA instructions paired with 3D CT scans, with step‑aligned rationales and cross‑modal alignments along the trajectory from findings to impression to pathology, enabling evaluation of both answer accuracy and reasoning consistency. We further propose TumorChain, a multimodal interleaved reasoning framework that tightly couples 3D imaging encoders, clinical text understanding, and organ‑level vision‑language alignment. Through cross‑modal alignment and iterative interleaved causal reasoning, TumorChain grounds visual evidence, aggregates conclusions, and issues pathology predictions after multiple rounds of self‑refinement, improving traceability and reducing hallucination risk. Experiments show consistent improvements over strong baselines in lesion detection, impression generation, and pathology classification, and demonstrate strong generalization on the DeepTumorVQA benchmark. These results highlight the potential of multimodal reasoning for reliable and interpretable tumor analysis in clinical practice. Detailed information about our project can be found on our project homepage at https://github.com/ZJU4HealthCare/TumorChain.
Authors:Junyu Chen, Md Yousuf Harun, Christopher Kanan
Abstract:
The original ImageNet benchmark enforces a single‑label assumption, despite many images depicting multiple objects. This leads to label noise and limits the richness of the learning signal. Multi‑label annotations more accurately reflect real‑world visual scenes, where multiple objects co‑occur and contribute to semantic understanding, enabling models to learn richer and more robust representations. While prior efforts (e.g., ReaL, ImageNetv2) have improved the validation set, there has not yet been a scalable, high‑quality multi‑label annotation for the training set. To this end, we present an automated pipeline to convert the ImageNet training set into a multi‑label dataset, without human annotations. Using self‑supervised Vision Transformers, we perform unsupervised object discovery, select regions aligned with original labels to train a lightweight classifier, and apply it to all regions to generate coherent multi‑label annotations across the dataset. Our labels show strong alignment with human judgment in qualitative evaluations and consistently improve performance across quantitative benchmarks. Compared to traditional single‑label scheme, models trained with our multi‑label supervision achieve consistently better in‑domain accuracy across architectures (up to +2.0 top‑1 accuracy on ReaL and +1.5 on ImageNet‑V2) and exhibit stronger transferability to downstream tasks (up to +4.2 and +2.3 mAP on COCO and VOC, respectively). These results underscore the importance of accurate multi‑label annotations for enhancing both classification performance and representation learning. Project code and the generated multi‑label annotations are available at https://github.com/jchen175/MultiLabel‑ImageNet.
Authors:Zhiyuan Zhou, Ruofeng Liu, Taichi Liu, Weijian Zuo, Shanshan Wang, Zhiqing Hong, Desheng Zhang
Abstract:
Accurate, dense depth estimation is crucial for robotic perception, but commodity sensors often yield sparse or incomplete measurements due to hardware limitations. Existing RGBD‑fused depth completion methods learn priors jointly conditioned on training RGB distribution and specific depth patterns, limiting domain generalization and robustness to various depth patterns. Recent efforts leverage monocular depth estimation (MDE) models to introduce domain‑general geometric priors, but current two‑stage integration strategies relying on explicit relative‑to‑metric alignment incur additional computation and introduce structured distortions. To this end, we present Any2Full, a one‑stage, domain‑general, and pattern‑agnostic framework that reformulates completion as a scale‑prompting adaptation of a pretrained MDE model. To address varying depth sparsity levels and irregular spatial distributions, we design a Scale‑Aware Prompt Encoder. It distills scale cues from sparse inputs into unified scale prompts, guiding the MDE model toward globally scale‑consistent predictions while preserving its geometric priors. Extensive experiments demonstrate that Any2Full achieves superior robustness and efficiency. It outperforms OMNI‑DC by 32.2% in average AbsREL and delivers a 1.4× speedup over PriorDA with the same MDE backbone, establishing a new paradigm for universal depth completion. Codes and checkpoints are available at https://github.com/zhiyuandaily/Any2Full.
Authors:Tongda Xu, Mingwei He, Shady Abu-Hussein, Jose Miguel Hernandez-Lobato, Chunhang Zheng, Kai Zhao, Chao Zhou, Ya-Qin Zhang, Yan Wang
Abstract:
It is well known that the reconstruction FID (rFID) of a VAE is poorly correlated with the generation FID (gFID) of a latent diffusion model. We propose interpolated FID (iFID), a simple variant of rFID that exhibits a strong correlation with gFID. Specifically, for each dataset element, we retrieve its nearest neighbor in latent space, interpolate between their latent representations, decode the interpolated latent, and compute the FID between the decoded samples and the original dataset. We provide an intuitive explanation for why iFID correlates well with gFID, and why reconstruction metrics can be negatively correlated with gFID, by connecting iFID to recent results on diffusion generalization and hallucination. Theoretically, we show that iFID evaluates decoded interpolations aligned with the ridge set around which diffusion samples concentrate, thereby measuring a quantity closely related to diffusion sample quality. Empirically, iFID is the first metric shown to strongly correlate with diffusion gFID across diverse VAEs, achieving Pearson and Spearman correlations of approximately 0.85. The project page is available at https://tongdaxu.github.io/pages/ifid.html.
Authors:Jieneng Chen, Wenxin Ma, Ruisheng Yuan, Yunzhi Zhang, Jiajun Wu, Alan Yuille
Abstract:
We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical‑world visual question answering. We highlight the empirical finding that our proposed spatial encoder can parse videos into structured spatial code with explicit 3D oriented bounding boxes and semantic labels, enabling large language models (LLMs) to reason directly over explicit spatial variables. Specifically, we propose the spatial encoder that encodes image and geometric features by unifying 6D object parsing and tracking backbones with geometric prediction, and we further finetuning LLMs with reinforcement learning using a spatial rubric reward that encourages perspective‑aware, geometrically grounded inference. As a result, our model outperforms proprietary vision‑language models on VSI‑Bench, setting a new state‑of‑the‑art. Code is available at https://github.com/Beckschen/spatialcode.
Authors:Leif Van Holland, Domenic Zingsheim, Mana Takhsha, Hannah Dröge, Patrick Stotko, Markus Plack, Reinhard Klein
Abstract:
High‑quality 3D streaming from multiple cameras is crucial for immersive experiences in many AR/VR applications. The limited number of views ‑ often due to real‑time constraints ‑ leads to missing information and incomplete surfaces in the rendered images. Existing approaches typically rely on simple heuristics for the hole filling, which can result in inconsistencies or visual artifacts. We propose to complete the missing textures using a novel, application‑targeted inpainting method independent of the underlying representation as an image‑based post‑processing step after the novel view rendering. The method is designed as a standalone module compatible with any calibrated multi‑camera system. For this we introduce a multi‑view aware, transformer‑based network architecture using spatio‑temporal embeddings to ensure consistency across frames while preserving fine details. Additionally, our resolution‑independent design allows adaptation to different camera setups, while an adaptive patch selection strategy balances inference speed and quality, allowing real‑time performance. We evaluate our approach against state‑of‑the‑art inpainting techniques under the same real‑time constraints and demonstrate that our model achieves the best trade‑off between quality and speed, outperforming competitors in both image and video‑based metrics.
Authors:Weijie Lyu, Ming-Hsuan Yang, Zhixin Shu
Abstract:
We introduce FaceCam, a system that generates video under customizable camera trajectories for monocular human portrait video input. Recent camera control approaches based on large video‑generation models have shown promising progress but often exhibit geometric distortions and visual artifacts on portrait videos due to scale‑ambiguous camera representations or 3D reconstruction errors. To overcome these limitations, we propose a face‑tailored scale‑aware representation for camera transformations that provides deterministic conditioning without relying on 3D priors. We train a video generation model on both multi‑view studio captures and in‑the‑wild monocular videos, and introduce two camera‑control data generation strategies: synthetic camera motion and multi‑shot stitching, to exploit stationary training cameras while generalizing to dynamic, continuous camera trajectories at inference time. Experiments on Ava‑256 dataset and diverse in‑the‑wild videos demonstrate that FaceCam achieves superior performance in camera controllability, visual quality, identity and motion preservation.
Authors:Scout Jarman, Zigfried Hampel-Arias, Adra Carr, Kevin R. Moon
Abstract:
Hyperspectral images (HSI) have many applications, ranging from environmental monitoring to national security, and can be used for material detection and identification. Longwave infrared (LWIR) HSI can be used for gas plume detection and analysis. Oftentimes, only a few images of a scene of interest are available and are analyzed individually. The ability to combine information from multiple images into a single, cohesive representation could enhance analysis by providing more context on the scene's geometry and spectral properties. Neural radiance fields (NeRFs) create a latent neural representation of volumetric scene properties that enable novel‑view rendering and geometry reconstruction, offering a promising avenue for hyperspectral 3D scene reconstruction. We explore the possibility of using NeRFs to create 3D scene reconstructions from LWIR HSI and demonstrate that the model can be used for the basic downstream analysis task of gas plume detection. The physics‑based DIRSIG software suite was used to generate a synthetic multi‑view LWIR HSI dataset of a simple facility with a strong sulfur hexafluoride gas plume. Our method, built on the standard Mip‑NeRF architecture, combines state‑of‑the‑art methods for hyperspectral NeRFs and sparse‑view NeRFs, along with a novel adaptive weighted MSE loss. Our final NeRF method requires around 50% fewer training images than the standard Mip‑NeRF and achieves an average PSNR of 39.8 dB with as few as 30 training images. Gas plume detection applied to NeRF‑rendered test images using the adaptive coherence estimator achieves an average AUC of 0.821 when compared with detection masks generated from ground‑truth test images.
Authors:Wei Liu, Ziyu Chen, Zizhang Li, Yue Wang, Hong-Xing Yu, Jiajun Wu
Abstract:
Current video generation models cannot simulate physical consequences of 3D actions like forces and robotic manipulations, as they lack structural understanding of how actions affect 3D scenes. We present RealWonder, the first real‑time system for action‑conditioned video generation from a single image. Our key insight is using physics simulation as an intermediate bridge: instead of directly encoding continuous actions, we translate them through physics simulation into visual representations (optical flow and RGB) that video models can process. RealWonder integrates three components: 3D reconstruction from single images, physics simulation, and a distilled video generator requiring only 4 diffusion steps. Our system achieves 13.2 FPS at 480x832 resolution, enabling interactive exploration of forces, robot actions, and camera controls on rigid objects, deformable bodies, fluids, and granular materials. We envision RealWonder opens new opportunities to apply video models in immersive experiences, AR/VR, and robot learning. Our code and model weights are publicly available in our project website: https://liuwei283.github.io/RealWonder/
Authors:Jiayin Zhu, Guoji Fu, Xiaolu Liu, Qiyuan He, Yicong Li, Angela Yao
Abstract:
Image‑to‑3D generation faces inherent semantic ambiguity under occlusion, where partial observation alone is often insufficient to determine object category. In this work, we formalize text‑driven amodal 3D generation, where text prompts steer the completion of unseen regions while strictly preserving input observation. Crucially, we identify that these objectives demand distinct control granularities: rigid control for the observation versus relaxed structural control for the prompt. To this end, we propose RelaxFlow, a training‑free dual‑branch framework that decouples control granularity via a Multi‑Prior Consensus Module and a Relaxation Mechanism. Theoretically, we prove that our relaxation is equivalent to applying a low‑pass filter on the generative vector field, which suppresses high‑frequency instance details to isolate geometric structure that accommodates the observation. To facilitate evaluation, we introduce two diagnostic benchmarks, ExtremeOcc‑3D and AmbiSem‑3D. Extensive experiments demonstrate that RelaxFlow successfully steers the generation of unseen regions to match the prompt intent without compromising visual fidelity.
Authors:Sijia Chen, Zihan Zhou, Yanqiu Yu, En Yu, Wenbing Tao
Abstract:
Multi‑Object Tracking (MOT) is a fundamental task in computer vision, aiming to track targets across video frames. Existing MOT methods perform well in general visual scenes, but face significant challenges and limitations when extended to visual‑language settings. To bridge this gap, the task of Referring Multi‑Object Tracking (RMOT) has recently been proposed, which aims to track objects that correspond to language descriptions. However, current RMOT methods are primarily developed on datasets captured by conventional cameras, which suffer from limited field of view. This constraint often causes targets to move out of the frame, leading to fragmented tracking and loss of contextual information. In this work, we propose a novel task, called Omnidirectional Referring Multi‑Object Tracking (ORMOT), which extends RMOT to omnidirectional imagery, aiming to overcome the field‑of‑view (FoV) limitation of conventional datasets and improve the model's ability to understand long‑horizon language descriptions. To advance the ORMOT task, we construct ORSet, an Omnidirectional Referring Multi‑Object Tracking dataset, which contains 27 diverse omnidirectional scenes, 848 language descriptions, and 3,401 annotated objects, providing rich visual, temporal, and language information. Furthermore, we propose ORTrack, a Large Vision‑Language Model (LVLM)‑driven framework tailored for Omnidirectional Referring Multi‑Object Tracking. Extensive experiments on the ORSet dataset demonstrate the effectiveness of our ORTrack framework. The dataset and code will be open‑sourced at https://github.com/chen‑si‑jia/ORMOT.
Authors:Shan Ning, Longtian Qiu, Xuming He
Abstract:
Knowledge‑Based Visual Question Answering (KB‑VQA) requires models to answer questions about an image by integrating external knowledge, posing significant challenges due to noisy retrieval and the structured, encyclopedic nature of the knowledge base. These characteristics create a distributional gap from pretrained multimodal large language models (MLLMs), making effective reasoning and domain adaptation difficult in the post‑training stage. In this work, we propose Wiki‑R1, a data‑generation‑based curriculum reinforcement learning framework that systematically incentivizes reasoning in MLLMs for KB‑VQA. Wiki‑R1 constructs a sequence of training distributions aligned with the model's evolving capability, bridging the gap from pretraining to the KB‑VQA target distribution. We introduce controllable curriculum data generation, which manipulates the retriever to produce samples at desired difficulty levels, and a curriculum sampling strategy that selects informative samples likely to yield non‑zero advantages during RL updates. Sample difficulty is estimated using observed rewards and propagated to unobserved samples to guide learning. Experiments on two KB‑VQA benchmarks, Encyclopedic VQA and InfoSeek, demonstrate that Wiki‑R1 achieves new state‑of‑the‑art results, improving accuracy from 35.5% to 37.1% on Encyclopedic VQA and from 40.1% to 44.1% on InfoSeek. The project page is available at https://artanic30.github.io/project_pages/WikiR1/.
Authors:Yingxue Su, Yiheng Zhong, Keying Zhu, Zimu Zhang, Zhuoru Zhang, Yifang Wang, Yuxin Zhang, Jingxin Liu
Abstract:
Medical image segmentation is critical for computer‑aided diagnosis. However, dense pixel‑level annotation is time‑consuming and expensive, and medical datasets often exhibit severe class imbalance. Such imbalance causes minority structures to be overwhelmed by dominant classes in feature representations, hindering the learning of discriminative features and making reliable segmentation particularly challenging. To address this, we propose the Semantic Class Distribution Learning (SCDL) framework, a plug‑and‑play module that mitigates supervision and representation biases by learning structured class‑conditional feature distributions. SCDL integrates Class Distribution Bidirectional Alignment (CDBA) to align embeddings with learnable class proxies and leverages Semantic Anchor Constraints (SAC) to guide proxies using labeled data. Experiments on the Synapse and AMOS datasets demonstrate that SCDL significantly improves segmentation performance across both overall and class‑level metrics, with particularly strong gains on minority classes, achieving state‑of‑the‑art results. Our code is released at https://github.com/Zyh55555/SCDL.
Authors:Muhammad Zarar, MingZheng Zhang, Xiaowang Zhang, Zhiyong Feng, Sofonias Yitagesu, Kawsar Farooq
Abstract:
Patient Activity Recognition (PAR) in clinical settings uses activity data to improve safety and quality of care. Although significant progress has been made, current models mainly identify which activity is occurring. They often spatially compose sub‑sparse visual cues using global and local attention mechanisms, yet only learn logically implicit patterns due to their neural‑pipeline. Advancing clinical safety requires methods that can infer why a set of visual cues implies a risk, and how these can be compositionally reasoned through explicit logic beyond mere classification. To address this, we proposed Logi‑PAR, the first Logic‑Infused Patient Activity Recognition Framework that integrates contextual fact fusion as a multi‑view primitive extractor and injects neural‑guided differentiable rules. Our method automatically learns rules from visual cues, optimizing them end‑to‑end while enabling the implicit emergence patterns to be explicitly labelled during training. To the best of our knowledge, Logi‑PAR is the first framework to recognize patient activity by applying learnable logic rules to symbolic mappings. It produces auditable why explanations as rule traces and supports counterfactual interventions (e.g., risk would decrease by 65% if assistance were present). Extensive evaluation on clinical benchmarks (VAST and OmniFall) demonstrates state‑of‑the‑art performance, significantly outperforming Vision‑Language Models and transformer baselines. The code is available via: https://github.com/zararkhan985/Logi‑PAR.git
Authors:Yuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu, Qiaoyu Tan
Abstract:
Recent advances in large language models (LLMs) have opened new avenues for multimodal reasoning. Yet, most existing methods still rely on pretrained vision‑language models (VLMs) to encode image‑text pairs in isolation, ignoring the relational structure that real‑world multimodal data naturally form. This motivates reasoning on multimodal graphs (MMGs), where each node has textual and visual attributes and edges provide structural cues. Enabling LLM‑based reasoning on such heterogeneous multimodal signals while preserving graph topology introduces two key challenges: resolving weak cross‑modal consistency and handling heterogeneous modality preference. To address this, we propose Mario, a unified framework that simultaneously resolves the two above challenges and enables effective LLM‑based reasoning over MMGs. Mario consists of two innovative stages. Firstly, a graph‑conditioned VLM design that jointly refines textual and visual features through fine‑grained cross‑modal contrastive learning guided by graph topology. Secondly, a modality‑adaptive graph instruction tuning mechanism that organizes aligned multimodal features into graph‑aware instruction views and employs a learnable router to surface, for each node and its neighborhood, the most informative modality configuration to the LLM. Extensive experiments across diverse MMG benchmarks demonstrate that Mario consistently outperforms state‑of‑the‑art graph models in both supervised and zero‑shot scenarios for node classification and link prediction. The code will be made available at https://github.com/sunyuanfu/Mario.
Authors:Ningjing Fan, Yiqun Wang
Abstract:
In recent years, 3D Gaussian splatting (3DGS) has achieved remarkable progress in novel view synthesis. However, accurately reconstructing glossy surfaces under complex illumination remains challenging, particularly in scenes with strong specular reflections and multi‑surface interreflections. To address this issue, we propose SSR‑GS, a specular reflection modeling framework for glossy surface reconstruction. Specifically, we introduce a prefiltered Mip‑Cubemap to model direct specular reflections efficiently, and propose an IndiASG module to capture indirect specular reflections.
Furthermore, we design Visual Geometry Priors (VGP) that couple a reflection‑aware visual prior via a reflection score (RS) to downweight the photometric loss contribution of reflection‑dominated regions, with geometry priors derived from VGGT, including progressively decayed depth supervision and transformed normal constraints. Extensive experiments on both synthetic and real‑world datasets demonstrate that SSR‑GS achieves state‑of‑the‑art performance in glossy surface reconstruction.
Authors:Minghe Xu, Rouying Wu, Jiarui Xu, Minhao Sun, Zikang Yan, Xiao Wang, ChiaWei Chu, Yu Li
Abstract:
Pedestrian Attribute Recognition is a foundational computer vision task that provides essential support for downstream applications, including person retrieval in video surveillance and intelligent retail analytics. However, existing research is frequently constrained by the ``one‑model‑per‑dataset" paradigm and struggles to handle significant discrepancies across domains in terms of modalities, attribute definitions, and environmental scenarios. To address these challenges, we propose UniPAR, a unified Transformer‑based framework for PAR. By incorporating a unified data scheduling strategy and a dynamic classification head, UniPAR enables a single model to simultaneously process diverse datasets from heterogeneous modalities, including RGB images, video sequences, and event streams. We also introduce an innovative phased fusion encoder that explicitly aligns visual features with textual attribute queries through a late deep fusion strategy. Experimental results on the widely used benchmark datasets, including MSP60K, DukeMTMC, and EventPAR, demonstrate that UniPAR achieves performance comparable to specialized SOTA methods. Furthermore, multi‑dataset joint training significantly enhances the model's cross‑domain generalization and recognition robustness in extreme environments characterized by low light and motion blur. The source code of this paper will be released on https://github.com/Event‑AHU/OpenPAR
Authors:Cenwei Zhang, Lin Zhu, Manxi Lin, Lei You
Abstract:
Feature attributions often hide a critical modeling choice: they explain a prediction along a counterfactual path from a reference state to an input. Different baselines, interpolations, and generative trajectories define different paths and can therefor produce different explanations. We study this path ambiguity as a modeling problem. Our central question is whether the path can be chosen by the data‑generating transport process, rather than by a hand‑designed interpolation or by the sensitivity geometry of the model being explained. We separate attribution into fixed‑path credit allocation and path selection. For a fixed path, we prove that the Aumann‑Shapley line integral is the unique attribution rule under standard fixed‑path axioms and explicit coordinate‑trace regularity. For path selection, we minimize kinetic action over flows that transport a reference distribution to the data distribution, yielding a transport‑geodesic attribution principle. We approximate this ideal with Rectified Flow and Reflow and derive stability bounds linking vector‑field error to attribution error. Experiments show that lower‑action, transport‑consistent paths produce more stable and structured explanations, preserving competitive deletion faithfulness, without claiming data‑manifold membership. Our code is available at https://github.com/cenweizhang/OTFlowSHAP.
Authors:Juntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di, Xuancheng Zhang, Chengmin Yang, Yu-Shen Liu
Abstract:
Reconstructing dynamic 4D scenes remains challenging due to the presence of moving objects that corrupt camera pose estimation. Existing optimization methods alleviate this issue with additional supervision, but they are mostly computationally expensive and impractical in real‑time applications. To address these limitations, we propose MoRe, a feedforward 4D reconstruction network that efficiently recovers dynamic 3D scenes from monocular videos. Built upon a strong static reconstruction backbone, MoRe employs an attention‑forcing strategy to disentangle dynamic motion from static structure. To further enhance robustness, we fine‑tune the model on large‑scale, diverse datasets encompassing both dynamic and static scenes. Moreover, our grouped causal attention captures temporal dependencies and adapts to varying token lengths across frames, ensuring temporally coherent geometry reconstruction. Extensive experiments on multiple benchmarks demonstrate that MoRe achieves high‑quality dynamic reconstructions with exceptional efficiency.
Authors:Yanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang, Yiran Zhao, Haodong Li, Congyue Zhou, Weijie Zheng, Yushen Yan, Shengqiong Wu, Wei Ji, Lei Cui, Furu Wei, Hao Fei, Mong-Li Lee, Wynne Hsu
Abstract:
In real‑world multimodal applications, systems usually need to comprehend arbitrarily combined and interleaved multimodal inputs from users, while also generating outputs in any interleaved multimedia form. This capability defines the goal of any‑to‑any interleaved multimodal learning under a unified paradigm of understanding and generation, posing new challenges and opportunities for advancing Multimodal Large Language Models (MLLMs). To foster and benchmark this capability, this paper introduces the UniM benchmark, the first Unified Any‑to‑Any Interleaved Multimodal dataset. UniM contains 31K high‑quality instances across 30 domains and 7 representative modalities: text, image, audio, video, document, code, and 3D, each requiring multiple intertwined reasoning and generation capabilities. We further introduce the UniM Evaluation Suite, which assesses models along three dimensions: Semantic Correctness & Generation Quality, Response Structure Integrity, and Interleaved Coherence. In addition, we propose UniMA, an agentic baseline model equipped with traceable reasoning for structured interleaved generation. Comprehensive experiments demonstrate the difficulty of UniM and highlight key challenges and directions for advancing unified any‑to‑any multimodal intelligence. The project page is https://any2any‑mllm.github.io/unim.
Authors:Nian Liu, Jin Gao, Shubo Lin, Yutong Kou, Sikui Zhang, Fudong Ge, Zhiqiang Pu, Liang Li, Gang Wang, Yizheng Wang, Weiming Hu
Abstract:
Infrared small target detection (ISTD) is challenging because tiny, low‑contrast targets are easily obscured by complex and dynamic backgrounds. Conventional multi‑frame approaches typically learn motion implicitly through deep neural networks, often requiring additional motion supervision or explicit alignment modules. We propose Motion Integration DETR (MI‑DETR), a bio‑inspired dual‑pathway detector that processes one infrared frame per time step while explicitly modeling motion. First, a retina‑inspired cellular automaton (RCA) converts raw frame sequences into a motion map defined on the same pixel grid as the appearance image, enabling parvocellular‑like appearance and magnocellular‑like motion pathways to be supervised by a single set of bounding boxes without extra motion labels or alignment operations. Second, a Parvocellular‑Magnocellular Interconnection (PMI) Block facilitates bidirectional feature interaction between the two pathways, providing a biologically motivated intermediate interconnection mechanism. Finally, a RT‑DETR decoder operates on features from the two pathways to produce detection results. Surprisingly, our proposed simple yet effective approach yields strong performance on three commonly used ISTD benchmarks. MI‑DETR achieves 70.3% mAP@50 and 72.7% F1 on IRDST‑H (+26.35 mAP@50 over the best multi‑frame baseline), 98.0% mAP@50 on DAUB‑R, and 88.3% mAP@50 on ITSDT‑15K, demonstrating the effectiveness of biologically inspired motion‑appearance integration. Code is available at https://github.com/nliu‑25/MI‑DETR.
Authors:Yulong Shi, Shijie Li, Ziyi Li, Lin Qi
Abstract:
Source Free Unsupervised Domain Adaptation (SFUDA) is critical for deploying deep learning models across diverse clinical settings. However, existing methods are typically designed for low‑gap, specific domain shifts and cannot generalize into a unified, multi‑modalities, and multi‑target framework, which presents a major barrier to real‑world application. To overcome this issue, we introduce Tell2Adapt, a novel SFUDA framework that harnesses the vast, generalizable knowledge of the Vision Foundation Model (VFM). Our approach ensures high‑fidelity VFM prompts through Context‑Aware Prompts Regularization (CAPR), which robustly translates varied text prompts into canonical instructions. This enables the generation of high‑quality pseudo‑labels for efficiently adapting the lightweight student model to target domain. To guarantee clinical reliability, the framework incorporates Visual Plausibility Refinement (VPR), which leverages the VFM's anatomical knowledge to re‑ground the adapted model's predictions in target image's low‑level visual features, effectively removing noise and false positives. We conduct one of the most extensive SFUDA evaluations to date, validating our framework across 10 domain adaptation directions and 22 anatomical targets, including brain, cardiac, polyp, and abdominal targets. Our results demonstrate that Tell2Adapt consistently outperforms existing approaches, achieving SOTA for a unified SFUDA framework in medical image segmentation. Code are avaliable at https://github.com/derekshiii/Tell2Adapt.
Authors:Jie Zhu, Hanghang Ma, Jia Wang, Yayong Guan, Yanbing Zeng, Lishuai Gao, Junqiang Wu, Jie Hu, Leye Wang
Abstract:
In this work, we introduce Wallaroo, a simple autoregressive baseline that leverages next‑token prediction to unify multi‑modal understanding, image generation, and editing at the same time. Moreover, Wallaroo supports multi‑resolution image input and output, as well as bilingual support for both Chinese and English. We decouple the visual encoding into separate pathways and apply a four‑stage training strategy to reshape the model's capabilities. Experiments are conducted on various benchmarks where Wallaroo produces competitive performance or exceeds other unified models, suggesting the great potential of autoregressive models in unifying multi‑modality understanding and generation. Our code is available at https://github.com/JiePKU/Wallaroo.
Authors:Zheng Wang, Haoran Chen, Haoxuan Qin, Zhipeng Wei, Tianwen Qian, Cong Bai
Abstract:
Long video understanding is challenging due to dense visual redundancy, long‑range temporal dependencies, and the tendency of chain‑of‑thought and retrieval‑based agents to accumulate semantic drift and correlation‑driven errors. We argue that long‑video reasoning should begin not with reactive retrieval, but with deliberate task formulation: the model must first articulate what must be true in the video for each candidate answer to hold. This thinking‑before‑finding principle motivates VideoHV‑Agent, a framework that reformulates video question answering as a structured hypothesis‑verification process. Based on video summaries, a Thinker rewrites answer candidates into testable hypotheses, a Judge derives a discriminative clue specifying what evidence must be checked, a Verifier grounds and tests the clue using localized, fine‑grained video content, and an Answer agent integrates validated evidence to produce the final answer. Experiments on three long‑video understanding benchmarks show that VideoHV‑Agent achieves state‑of‑the‑art accuracy while providing enhanced interpretability, improved logical soundness, and lower computational cost. We make our code publicly available at: https://github.com/Haorane/VideoHV‑Agent.
Authors:Zishu Yao, Xiang-Xiang Su, Shengning Zhou, Guang-Yong Chen, Guodong Fan, Xing Chen
Abstract:
Event cameras, with their high dynamic range, show great promise for Low‑light Image Enhancement (LLIE). Existing works primarily focus on designing effective modal fusion strategies. However, a key challenge is the dual degradation from intrinsic background activity (BA) noise in events and low signal‑to‑noise ratio (SNR) in images, which causes severe noise coupling during modal fusion, creating a critical performance bottleneck. We therefore posit that precise event denoising is the prerequisite to unlocking the full potential of event‑based fusion. To this end, we propose BiEvLight, a hierarchical and task‑aware framework that collaboratively optimizes enhancement and denoising by exploiting their intrinsic interdependence. Specifically, BiEvLight exploits the strong gradient correlation between images and events to build a gradient‑guided event denoising prior that alleviates insufficient denoising in heavily noisy regions. Moreover, instead of treating event denoising as a static pre‑processing stage‑which inevitably incurs a trade‑off between over‑ and under‑denoising and cannot adapt to the requirements of a specific enhancement objective‑we recast it as a bilevel optimization problem constrained by the enhancement task. Through cross‑task interaction, the upper‑level denoising problem learns event representations tailored to the lower‑level enhancement objective, thereby substantially improving overall enhancement quality. Extensive experiments on the Real‑world noise Dataset SDE demonstrate that our method significantly outperforms state‑of‑the‑art (SOTA) approaches, with average improvements of 1.30dB in PSNR, 2.03dB in PSNR and 0.047 in SSIM, respectively. The code will be publicly available at https://github.com/iijjlk/BiEvlight.
Authors:Toby Chong, Ryota Nakajima
Abstract:
We introduce a novel camera model for monocular 3D Morphable Model (3DMM) regression methods that effectively captures the perspective distortion effect commonly seen in close‑up facial images.
Fitting 3D morphable models to video is a key technique in content creation. In particular, regression‑based approaches have produced fast and accurate results by matching the rendered output of the morphable model to the target image. These methods typically achieve stable performance with orthographic projection, which eliminates the ambiguity between focal length and object distance. However, this simplification makes them unsuitable for close‑up footage, such as that captured with head‑mounted cameras.
We extend orthographic projection with a new shrinkage parameter, incorporating a pseudo‑perspective effect while preserving the stability of the original projection. We present several techniques that allow finetuning of existing models, and demonstrate the effectiveness of our modification through both quantitative and qualitative comparisons using a custom dataset recorded with head‑mounted cameras.
Authors:Nilusha Jayawickrama, Henrik Toikka, Risto Ojala
Abstract:
This paper investigates person detection and tracking in an industrial indoor workspace using a LiDAR mounted on an overhead crane. The overhead viewpoint introduces a strong domain shift from common vehicle‑centric LiDAR benchmarks, and limited availability of suitable public training data. Henceforth, we curate a site‑specific overhead LiDAR dataset with 3D human bounding‑box annotations and adapt selected candidate 3D detectors under a unified training and evaluation protocol. We further integrate lightweight tracking‑by‑detection using AB3DMOT and SimpleTrack to maintain person identities over time. Detection performance is reported with distance‑sliced evaluation to quantify the practical operating envelope of the sensing setup. The best adapted detector configurations achieve average precision (AP) up to 0.84 within a 5.0 m horizontal radius, increasing to 0.97 at 1.0 m, with VoxelNeXt and SECOND emerging as the most reliable backbones across this range. The acquired results contribute in bridging the domain gap between standard driving datasets and overhead sensing for person detection and tracking. We also report latency measurements, highlighting practical real‑time feasibility. Finally, we release our dataset and implementations in GitHub to support further research
Authors:Chanmi Lee, Minsung Yoon, Woojae Kim, Sebin Lee, Sung-eui Yoon
Abstract:
Neural network‑based visuomotor policies enable robots to perform manipulation tasks but remain susceptible to perceptual attacks. For example, conventional 2D adversarial patches are effective under fixed‑camera setups, where appearance is relatively consistent; however, their efficacy often diminishes under dynamic viewpoints from moving cameras, such as wrist‑mounted setups, due to perspective distortions. To proactively investigate potential vulnerabilities beyond 2D patches, this work proposes a viewpoint‑consistent adversarial texture optimization method for 3D objects through differentiable rendering. As optimization strategies, we employ Expectation over Transformation (EOT) with a Coarse‑to‑Fine (C2F) curriculum, exploiting distance‑dependent frequency characteristics to induce textures effective across varying camera‑object distances. We further integrate saliency‑guided perturbations to redirect policy attention and design a targeted loss that persistently drives robots toward adversarial objects. Our comprehensive experiments show that the proposed method is effective under various environmental conditions, while confirming its black‑box transferability and real‑world applicability.
Authors:Sina Hajimiri, Farzad Beizaee, Fereshteh Shakeri, Christian Desrosiers, Ismail Ben Ayed, Jose Dolz
Abstract:
Vision transformers have demonstrated remarkable success in classification by leveraging global self‑attention to capture long‑range dependencies. However, this same mechanism can obscure fine‑grained spatial details crucial for tasks such as segmentation. In this work, we seek to enhance segmentation performance of vision transformers after standard image‑level classification training. More specifically, we present a simple yet effective add‑on that improves performance on segmentation tasks while retaining vision transformers' image‑level recognition capabilities. In our approach, we modulate the self‑attention with a learnable Gaussian kernel that biases the attention toward neighboring patches. We further refine the patch representations to learn better embeddings at patch positions. These modifications encourage tokens to focus on local surroundings and ensure meaningful representations at spatial positions, while still preserving the model's ability to incorporate global information. Experiments demonstrate the effectiveness of our modifications, evidenced by substantial segmentation gains on three benchmarks (e.g., over 6% and 4% on ADE20K for ViT Tiny and Base), without changing the training regime or sacrificing classification performance. The code is available at https://github.com/sinahmr/LocAtViT/.
Authors:Sicheng Li, Zaiwang Gu, Jie Zhang, Qing Guo, Xudong Jiang, Jun Cheng
Abstract:
Establishing reliable image correspondences is essential for many robotic vision problems. However, existing methods often struggle in challenging scenarios with large viewpoint changes or textureless regions, where incorrect cor‑ respondences may still receive high similarity scores. This is mainly because conventional models rely solely on fea‑ ture similarity, lacking an explicit mechanism to estimate the reliability of predicted matches, leading to overconfident errors. To address this issue, we propose SURE, a Semi‑ dense Uncertainty‑REfined matching framework that jointly predicts correspondences and their confidence by modeling both aleatoric and epistemic uncertainties. Our approach in‑ troduces a novel evidential head for trustworthy coordinate regression, along with a lightweight spatial fusion module that enhances local feature precision with minimal overhead. We evaluated our method on multiple standard benchmarks, where it consistently outperforms existing state‑of‑the‑art semi‑dense matching models in both accuracy and efficiency. our code will be available on https://github.com/LSC‑ALAN/SURE.
Authors:Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler
Abstract:
The rapid progress of Multi‑Modal Large Language Models (MLLMs) has significantly advanced downstream applications. However, this progress also exposes serious transferable adversarial vulnerabilities. In general, existing adversarial attacks against MLLMs typically rely on surrogate models trained within a single learning paradigm and perform independent optimisation in their respective feature spaces. This straightforward setting naturally restricts the richness of feature representations, delivering limits on the search space and thus impeding the diversity of adversarial perturbations. To address this, we propose a novel Multi‑Paradigm Collaborative Attack (MPCAttack) framework to boost the transferability of adversarial examples against MLLMs. In principle, MPCAttack aggregates semantic representations, from both visual images and language texts, to facilitate joint adversarial optimisation on the aggregated features through a Multi‑Paradigm Collaborative Optimisation (MPCO) strategy. By performing contrastive matching on multi‑paradigm features, MPCO adaptively balances the importance of different paradigm representations and guides the global perturbation optimisation, effectively alleviating the representation bias. Extensive experimental results on multiple benchmarks demonstrate the superiority of MPCAttack, indicating that our solution consistently outperforms state‑of‑the‑art methods in both targeted and untargeted attacks on open‑source and closed‑source MLLMs. The code is released at https://github.com/LiYuanBoJNU/MPCAttack.
Authors:Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler
Abstract:
With the rapid advancement and widespread application of vision‑language pre‑training (VLP) models, their vulnerability to adversarial attacks has become a critical concern. In general, the adversarial examples can typically be designed to exhibit transferable power, attacking not only different models but also across diverse tasks. However, existing attacks on language‑vision models mainly rely on static cross‑modal interactions and focus solely on disrupting positive image‑text pairs, resulting in limited cross‑modal disruption and poor transferability. To address this issue, we propose a Semantic‑Augmented Dynamic Contrastive Attack (SADCA) that enhances adversarial transferability through progressive and semantically guided perturbation. SADCA progressively disrupts cross‑modal alignment through dynamic interactions between adversarial images and texts. This is accomplished by SADCA establishing a contrastive learning mechanism involving adversarial, positive and negative samples, to reinforce the semantic inconsistency of the obtained perturbations. Moreover, we empirically find that input transformations commonly used in traditional transfer‑based attacks also benefit VLPs, which motivates a semantic augmentation module that increases the diversity and generalization of adversarial examples. Extensive experiments on multiple datasets and models demonstrate that SADCA significantly improves adversarial transferability and consistently surpasses state‑of‑the‑art methods. The code is released at https://github.com/LiYuanBoJNU/SADCA.
Authors:Rui Zhao, Bin Shi, Kai Sun, Bo Dong
Abstract:
Partial label learning is a prominent weakly supervised classification task, where each training instance is ambiguously labeled with a set of candidate labels. In real‑world scenarios, candidate labels are often influenced by instance features, leading to the emergence of instance‑dependent PLL (ID‑PLL), a setting that more accurately reflects this relationship. A significant challenge in ID‑PLL is instance entanglement, where instances from similar classes share overlapping features and candidate labels, resulting in increased class confusion. To address this issue, we propose a novel Class‑specific Augmentation based Disentanglement (CAD) framework, which tackles instance entanglement by both intra‑ and inter‑class regulations. For intra‑class regulation, CAD amplifies class‑specific features to generate class‑wise augmentations and aligns same‑class augmentations across instances. For inter‑class regulation, CAD introduces a weighted penalty loss function that applies stronger penalties to more ambiguous labels, encouraging larger inter‑class distances. By jointly applying intra‑ and inter‑class regulations, CAD improves the clarity of class boundaries and reduces class confusion caused by entanglement. Extensive experimental results demonstrate the effectiveness of CAD in mitigating the entanglement problem and enhancing ID‑PLL performance. The code is available at https://github.com/RyanZhaoIc/CAD.git.
Authors:Boyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang, Ruochen Cui, Xilin Zhao, Qingming Huang
Abstract:
The limited understanding capacity of the visual encoder in Contrastive Language‑Image Pre‑training (CLIP) has become a key bottleneck for downstream performance. This capacity includes both Discriminative Ability (D‑Ability), which reflects class separability, and Detail Perceptual Ability (P‑Ability), which focuses on fine‑grained visual cues. Recent solutions use diffusion models to enhance representations by conditioning image reconstruction on CLIP visual tokens. We argue that such paradigms may compromise D‑Ability and therefore fail to effectively address CLIP's representation limitations. To address this, we integrate contrastive signals into diffusion‑based reconstruction to pursue more comprehensive visual representations. We begin with a straightforward design that augments the diffusion process with contrastive learning on input images. However, empirical results show that the naive combination suffers from gradient conflict and yields suboptimal performance. To balance the optimization, we introduce the Diffusion Contrastive Reconstruction (DCR), which unifies the learning objective. The key idea is to inject contrastive signals derived from each reconstructed image, rather than from the original input, into the diffusion process. Our theoretical analysis shows that the DCR loss can jointly optimize D‑Ability and P‑Ability. Extensive experiments across various benchmarks and multi‑modal large language models validate the effectiveness of our method. The code is available at https://github.com/boyuh/DCR.
Authors:Lulu Hu, Wenhu Xiao, Xin Chen, Xinhua Xu, Bowen Xu, Kun Li, Yongliang Tao
Abstract:
Post‑training quantization (PTQ) with computational invariance for Large Language Models~(LLMs) have demonstrated remarkable advances, however, their application to Multimodal Large Language Models~(MLLMs) presents substantial challenges. In this paper, we analyze SmoothQuant as a case study and identify two critical issues: Smoothing Misalignment and Cross‑Modal Computational Invariance. To address these issues, we propose Modality‑Aware Smoothing Quantization (MASQuant), a novel framework that introduces (1) Modality‑Aware Smoothing (MAS), which learns separate, modality‑specific smoothing factors to prevent Smoothing Misalignment, and (2) Cross‑Modal Compensation (CMC), which addresses Cross‑modal Computational Invariance by using SVD whitening to transform multi‑modal activation differences into low‑rank forms, enabling unified quantization across modalities. MASQuant demonstrates stable quantization performance across both dual‑modal and tri‑modal MLLMs. Experimental results show that MASQuant is competitive among the state‑of‑the‑art PTQ algorithms. Source code: https://github.com/alibaba/EfficientAI.
Authors:Feng Liu, Bingyu Nan, Xuezhong Qian, Xiaolan Fu
Abstract:
Existing manual labeling of micro‑expressions is subject to errors in accuracy, especially in cross‑cultural scenarios where deviation in labeling of key frames is more prominent. To address this issue, this paper presents a novel Global Anti‑Monotonic Differential Selection Strategy (GAMDSS) architecture for enhancing the effectiveness of spatio‑temporal modeling of micro‑expressions through keyframe re‑selection. Specifically, the method identifies Onset and Apex frames, which are characterized by significant micro‑expression variation, from complete micro‑expression action sequences via a dynamic frame reselection mechanism. It then uses these to determine Offset frames and construct a rich spatio‑temporal dynamic representation. A two‑branch structure with shared parameters is then used to efficiently extract spatio‑temporal features. Extensive experiments are conducted on seven widely recognized micro‑expression datasets. The results demonstrate that GAMDSS effectively reduces subjective errors caused by human factors in multicultural datasets such as SAMM and 4DME. Furthermore, quantitative analyses confirm that offset‑frame annotations in multicultural datasets are more uncertain, providing theoretical justification for standardizing micro‑expression annotations. These findings directly support our argument for reconsidering the validity and generalizability of dataset annotation paradigms. Notably, this design can be integrated into existing models without increasing the number of parameters, offering a new approach to enhancing micro‑expression recognition performance. The source code is available on GitHub[https://github.com/Cross‑Innovation‑Lab/GAMDSS].
Authors:Yang Zou, Jun Ma, Zhidong Jiao, Xingyuan Li, Zhiying Jiang, Jinyuan Liu
Abstract:
Infrared image super‑resolution (IISR) under real‑world conditions is a practically significant yet rarely addressed task. Pioneering works are often trained and evaluated on simulated datasets or neglect the intrinsic differences between infrared and visible imaging. In practice, however, real infrared images are affected by coupled optical and sensing degradations that jointly deteriorate both structural sharpness and thermal fidelity. To address these challenges, we propose Real‑IISR, a unified autoregressive framework for real‑world IISR that progressively reconstructs fine‑grained thermal structures and clear backgrounds in a scale‑by‑scale manner via thermal‑structural guided visual autoregression. Specifically, a Thermal‑Structural Guidance module encodes thermal priors to mitigate the mismatch between thermal radiation and structural edges. Since non‑uniform degradations typically induce quantization bias, Real‑IISR adopts a Condition‑Adaptive Codebook that dynamically modulates discrete representations based on degradation‑aware thermal priors. Also, a Thermal Order Consistency Loss enforces a monotonic relation between temperature and pixel intensity, ensuring relative brightness order rather than absolute values to maintain physical consistency under spatial misalignment and thermal drift. We build FLIR‑IISR, a real‑world IISR dataset with paired LR‑HR infrared images acquired via automated focus variation and motion‑induced blur. Extensive experiments demonstrate the promising performance of Real‑IISR, providing a unified foundation for real‑world IISR and benchmarking. The dataset and code are available at: https://github.com/JZD151/Real‑IISR.
Authors:Junlong Tong, Zilong Wang, YuJie Ren, Peiran Yin, Hao Wu, Wei Zhang, Xiaoyu Shen
Abstract:
Standard Large Language Models (LLMs) are predominantly designed for static inference with pre‑defined inputs, which limits their applicability in dynamic, real‑time scenarios. To address this gap, the streaming LLM paradigm has emerged. However, existing definitions of streaming LLMs remain fragmented, conflating streaming generation, streaming inputs, and interactive streaming architectures, while a systematic taxonomy is still lacking. This paper provides a comprehensive overview and analysis of streaming LLMs. First, we establish a unified definition of streaming LLMs based on data flow and dynamic interaction to clarify existing ambiguities. Building on this definition, we propose a systematic taxonomy of current streaming LLMs and conduct an in‑depth discussion on their underlying methodologies. Furthermore, we explore the applications of streaming LLMs in real‑world scenarios and outline promising research directions to support ongoing advances in streaming intelligence. We maintain a continuously updated repository of relevant papers at https://github.com/EIT‑NLP/Awesome‑Streaming‑LLMs.
Authors:Ancymol Thomas, Jaya Sreevalsan-Nair
Abstract:
Local Climate Zones (LCZs) give a zoning map to study urban structures and land use and analyze the impact of urbanization on local climate. Multimodal remote sensing enables LCZ classification, for which data fusion is significant for improving accuracy owing to the data complexity. However, there is a gap in a comprehensive analysis of the fusion mechanisms used in their deep learning (DL) classifier architectures. This study analyzes different fusion strategies in the multi‑class LCZ classification models for multimodal data and grouping strategies based on inherent data characteristics. The different models involving Convolutional Neural Networks (CNNs) include: (i) baseline hybrid fusion (FM1), (ii) with self‑ and cross‑attention mechanisms (FM2), (iii) with the multi‑scale Gaussian filtered images (FM3), and (iv) weighted decision‑level fusion (FM4). Ablation experiments are conducted to study the pixel‑, feature‑, and decision‑level fusion effects in the model performance. Grouping strategies include band grouping (BG) within the data modalities and label merging (LM) in the ground truth. Our analysis is exclusively done on the So2Sat LCZ42 dataset, which consists of Synthetic Aperture Radar (SAR) and Multispectral Imaging (MSI) image pairs. Our results show that FM1 consistently outperforms simple fusion methods. FM1 with BG and LM is found to be the most effective approach among all fusion strategies, giving an overall accuracy of 76.6%. Importantly, our study highlights the effect of these strategies in improving prediction accuracy for the underrepresented classes. Our code and processed datasets are available at https://github.com/GVCL/LCZC‑MultiModalHybridFusion
Authors:Ruobing Zheng, Tianqi Li, Jianing Li, Qingpei Guo, Yi Yuan, Jingdong Chen
Abstract:
Reasoning post‑training improves Large Language Models (LLMs) on complex tasks such as mathematics and coding, but its benefits across diverse multimodal tasks remains uncertain. The trend of releasing parallel "Instruct" and "Thinking" models by leading teams is both resource‑intensive and user‑unfriendly. Prior work finds that the gains from reasoning training are influenced by multiple factors, such as base model capabilities, task characteristics, and Chain‑of‑Thought (CoT) data quality. However, principled criteria for determining when reasoning post‑training is beneficial and which data should support it are still lacking. In this paper, we propose Dual Tuning, a reasoning efficacy‑driven data curation framework for multimodal LLMs training. Given a target task and a base model, Dual Tuning jointly evaluates whether the training data is beneficial and whether reasoning training with current CoT content yields positive gains over non‑reasoning alternatives. We apply Dual Tuning across spatial, mathematical, and multi‑disciplinary tasks, and further analyze how reinforcement learning and thinking patterns affect reasoning efficacy. The Dual Tuning results guide data curation by identifying data that benefit reasoning training, data better suited to direct‑answer training, and data that are detrimental under both training modes. Our work provides quantitative criteria for selecting appropriate training data and matching post‑training strategies.
Authors:Ekansh Arora
Abstract:
Foundation models are increasingly applied to computational pathology, yet their behavior under cross‑cancer and cross‑species transfer remains unspecified. This study investigated how fine‑tuning CPath‑CLIP affects cancer detection under same‑cancer, cross‑cancer, and cross‑species conditions using whole‑slide image patches from canine and human histopathology. Performance was measured using area under the receiver operating characteristic curve (AUC). Few‑shot fine‑tuning improved same‑cancer (64.9% to 72.6% AUC) and cross‑cancer performance (56.84% to 66.31% AUC). Cross‑species evaluation revealed that while tissue matching enables meaningful transfer, performance remains below state‑of‑the‑art benchmarks (H‑optimus‑0: 84.97% AUC), indicating that standard vision‑language alignment is suboptimal for cross‑species generalization. Embedding space analysis revealed extremely high cosine similarity (greater than 0.99) between tumor and normal prototypes. Grad‑CAM shows prototype‑based models remain domain‑locked, while language‑guided models attend to conserved tumor morphology. To address this, we introduce Semantic Anchoring, which uses language to provide a stable coordinate system for visual features. Ablation studies reveal that benefits stem from the text‑alignment mechanism itself, regardless of text encoder complexity. Benchmarking against H‑optimus‑0 shows that CPath‑CLIP's failure stems from intrinsic embedding collapse, which text alignment effectively circumvents. Additional gains were observed in same‑cancer (8.52%) and cross‑cancer classification (5.67%). We identified a previously uncharacterized failure mode: semantic collapse driven by species‑dominated alignment rather than missing visual information. These results demonstrate that language acts as a control mechanism, enabling semantic re‑interpretation without retraining.
Authors:Haian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao, Jonathan T. Barron, Noah Snavely, Aleksander Holynski
Abstract:
Feed‑forward transformer models have driven rapid progress in 3D vision, but state‑of‑the‑art methods such as VGGT and π^3 have a computational cost that scales quadratically with the number of input images, making them inefficient when applied to large image collections. Sequential‑reconstruction approaches reduce this cost but sacrifice reconstruction quality. We introduce ZipMap, a stateful feed‑forward model that achieves linear‑time, bidirectional 3D reconstruction while matching or surpassing the accuracy of quadratic‑time methods. ZipMap employs test‑time training layers to zip an entire image collection into a compact hidden scene state in a single forward pass, enabling reconstruction of over 700 frames in under 10 seconds on a single H100 GPU, more than 20× faster than state‑of‑the‑art methods such as VGGT. Moreover, we demonstrate the benefits of having a stateful representation in real‑time scene‑state querying and its extension to sequential streaming reconstruction.
Authors:Chris Vorster, Mayug Maniparambil, Noel E. O'Connor, Noel Murphy, Derek Molloy
Abstract:
Large‑scale Vision‑Language Foundation Models (VLFMs), such as CLIP, now underpin a wide range of computer vision research and applications. VLFMs are often adapted to various domain‑specific tasks. However, VLFM performance on novel, specialised, or underrepresented domains remains inconsistent. Evaluating VLFMs typically requires labelled test sets, which are often unavailable for niche domains of interest, particularly those from the Global South. We address this gap by proposing a highly data‑efficient method to predict a VLFM's zero‑shot accuracy on a target domain using only a single labelled image per class. Our approach uses a Large Language Model to generate plausible counterfactual descriptions of a given image. By measuring the VLFM's ability to distinguish the correct description from these hard negatives, we engineer features that capture the VLFM's discriminative power in its shared embedding space. A linear regressor trained on these similarity scores estimates the VLFM's zero‑shot test accuracy across various visual domains with a Pearson‑r correlation of 0.96. We demonstrate our method's performance across five diverse datasets, including standard benchmark datasets and underrepresented datasets from Africa. Our work provides a low‑cost, reliable tool for probing VLFMs, enabling researchers and practitioners to make informed decisions about data annotation efforts before committing significant resources. The model training code, generated captions and counterfactuals are released here: https://github.com/chris‑vorster/PreLabellingProbe.
Authors:Chris Vorster, Mayug Maniparambil, Noel E. O'Connor, Noel Murphy, Derek Molloy
Abstract:
In many CLIP adaptation methods, a blending ratio hyperparameter controls the trade‑off between general pretrained CLIP knowledge and the limited, dataset‑specific supervision from the few‑shot cases. Most few‑shot CLIP adaptation techniques report results by ablation of the blending ratio on the test set or require additional validation sets to select the blending ratio per dataset, and thus are not strictly few‑shot. We present a simple, validation‑free method for learning the blending ratio in CLIP adaptation. Hold‑One‑Shot‑Out (HOSO) presents a novel approach for CLIP‑Adapter‑style methods to compete in the newly established validation‑free setting. CLIP‑Adapter with HOSO (HOSO‑Adapter) learns the blending ratio using a one‑shot, hold‑out set, while the adapter trains on the remaining few‑shot support examples. Under the validation‑free few‑shot protocol, HOSO‑Adapter outperforms the CLIP‑Adapter baseline by more than 4 percentage points on average across 11 standard few‑shot datasets. Interestingly, in the 8‑ and 16‑shot settings, HOSO‑Adapter outperforms CLIP‑Adapter even with the optimal blending ratio selected on the test set. Ablation studies validate the use of a one‑shot hold‑out mechanism, decoupled training, and improvements over the naively learnt blending ratio baseline. Code is released here: https://github.com/chris‑vorster/HOSO‑Adapter
Authors:William Grolleau, Achraf Chaouch, Astrid Sabourin, Guillaume Lapouge, Catherine Achard
Abstract:
Animal re‑identification (ReID) faces critical challenges due to viewpoint variations, particularly in Aerial‑Ground (AG‑ReID) settings where models must match individuals across drastic elevation changes. However, existing datasets lack the precise angular annotations required to systematically analyze these geometric variations. To address this, we introduce the Multi‑view Oriented Observation (MOO) dataset, a large‑scale synthetic AG‑ReID dataset of 1,000 cattle individuals captured from 128 uniformly sampled viewpoints (128,000 annotated images). Using this controlled dataset, we quantify the influence of elevation and identify a critical elevation threshold, above which models generalize significantly better to unseen views. Finally, we validate the transferability to real‑world applications in both zero‑shot and supervised settings, demonstrating performance gains across four real‑world cattle datasets and confirming that synthetic geometric priors effectively bridge the domain gap. Collectively, this dataset and analysis lay the foundation for future model development in cross‑view animal ReID. MOO is publicly available at https://github.com/TurtleSmoke/MOO.
Authors:Lingen Li, Guangzhi Wang, Xiaoyu Li, Zhaoyang Zhang, Qi Dou, Jinwei Gu, Tianfan Xue, Ying Shan
Abstract:
Generating high‑quality 360° panoramic videos from perspective input is one of the crucial applications for virtual reality (VR), whereby high‑resolution videos are especially important for immersive experience. Existing methods are constrained by computational limitations of vanilla diffusion models, only supporting \leq 1K resolution native generation and relying on suboptimal post super‑resolution to increase resolution. We introduce CubeComposer, a novel spatio‑temporal autoregressive diffusion model that natively generates 4K‑resolution 360° videos. By decomposing videos into cubemap representations with six faces, CubeComposer autoregressively synthesizes content in a well‑planned spatio‑temporal order, reducing memory demands while enabling high‑resolution output. Specifically, to address challenges in multi‑dimensional autoregression, we propose: (1) a spatio‑temporal autoregressive strategy that orchestrates 360° video generation across cube faces and time windows for coherent synthesis; (2) a cube face context management mechanism, equipped with a sparse context attention design to improve efficiency; and (3) continuity‑aware techniques, including cube‑aware positional encoding, padding, and blending to eliminate boundary seams. Extensive experiments on benchmark datasets demonstrate that CubeComposer outperforms state‑of‑the‑art methods in native resolution and visual quality, supporting practical VR application scenarios. Project page: https://lg‑li.github.io/project/cubecomposer
Authors:Seungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee Lee
Abstract:
Understanding a 3D scene immediately with its exploration is essential for embodied tasks, where an agent must construct and comprehend the 3D scene in an online and nearly real‑time manner. In this study, we propose EmbodiedSplat, an online feed‑forward 3DGS for open‑vocabulary scene understanding that enables simultaneous online 3D reconstruction and 3D semantic understanding from the streaming images. Unlike existing open‑vocabulary 3DGS methods which are typically restricted to either offline or per‑scene optimization setting, our objectives are two‑fold: 1) Reconstructs the semantic‑embedded 3DGS of the entire scene from over 300 streaming images in an online manner. 2) Highly generalizable to novel scenes with feed‑forward design and supports nearly real‑time 3D semantic reconstruction when combined with real‑time 2D models. To achieve these objectives, we propose an Online Sparse Coefficients Field with a CLIP Global Codebook where it binds the 2D CLIP embeddings to each 3D Gaussian while minimizing memory consumption and preserving the full semantic generalizability of CLIP. Furthermore, we generate 3D geometric‑aware CLIP features by aggregating the partial point cloud of 3DGS through 3D U‑Net to compensate the 3D geometric prior to 2D‑oriented language embeddings. Extensive experiments on diverse indoor datasets, including ScanNet, ScanNet++, and Replica, demonstrate both the effectiveness and efficiency of our method. Check out our project page in https://0nandon.github.io/EmbodiedSplat/.
Authors:Zijiang Yang, Chen Kuang, Dongmei Fu
Abstract:
Pathology Foundation Models (FMs) have shown strong performance across a wide range of pathology image representation and diagnostic tasks. However, FMs do not exhibit the expected performance advantage over traditional specialized models in Nuclei Detection and Classification (NDC). In this work, we reveal that jointly optimizing nuclei detection and classification leads to severe representation degradation in FMs. Moreover, we identify that the substantial intrinsic disparity in task difficulty between nuclei detection and nuclei classification renders joint NDC optimization unnecessarily computationally burdensome for the detection stage. To address these challenges, we propose DeNuC, a simple yet effective method designed to break through existing bottlenecks by Decoupling Nuclei detection and Classification. DeNuC employs a lightweight model for accurate nuclei localization, subsequently leveraging a pathology FM to encode input images and query nucleus‑specific features based on the detected coordinates for classification. Extensive experiments on three widely used benchmarks demonstrate that DeNuC effectively unlocks the representational potential of FMs for NDC and significantly outperforms state‑of‑the‑art methods. Notably, DeNuC improves F1 scores by 4.2% and 3.6% (or higher) on the BRCAM2C and PUMA datasets, respectively, while using only 16% (or fewer) trainable parameters compared to other methods. Code is available at https://github.com/ZijiangY1116/DeNuC.
Authors:Mengping Yang, Zhiyu Tan, Binglei Li, Xiaomeng Yang, Hesen Chen, Hao Li
Abstract:
Recent breakthroughs in Diffusion Transformers (DiTs) have revolutionized the field of visual synthesis due to their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs are not well understood. To this end, we first systematically investigate the representation dynamics of DiTs. Through analyzing the evolution and influence of internal representations under various settings, we reveal that representation diversity across blocks is a crucial factor for effective learning. Based on this key insight, we propose DiverseDiT, a novel framework that explicitly promotes representation diversity. DiverseDiT incorporates long residual connections to diversify input representations across blocks and a representation diversity loss to encourage blocks to learn distinct features. Extensive experiments on ImageNet 256x256 and 512x512 demonstrate that our DiverseDiT yields consistent performance gains and convergence acceleration when applied to different backbones with various sizes, even when tested on the challenging one‑step generation setting. Furthermore, we show that DiverseDiT is complementary to existing representation learning techniques, leading to further performance gains. Our work provides valuable insights into the representation learning dynamics of DiTs and offers a practical approach for enhancing their performance.
Authors:Weirong Chen, Chuanxia Zheng, Ganlin Zhang, Andrea Vedaldi, Daniel Cremers
Abstract:
We present NOVA3R, an effective approach for non‑pixel‑aligned 3D reconstruction from a set of unposed images in a feed‑forward manner. Unlike pixel‑aligned methods that tie geometry to per‑ray predictions, our formulation learns a global, view‑agnostic scene representation that decouples reconstruction from pixel alignment. This addresses two key limitations in pixel‑aligned 3D: (1) it recovers both visible and invisible points with a complete scene representation, and (2) it produces physically plausible geometry with fewer duplicated structures in overlapping regions. To achieve this, we introduce a scene‑token mechanism that aggregates information across unposed images and a diffusion‑based 3D decoder that reconstructs complete, non‑pixel‑aligned point clouds. Extensive experiments on both scene‑level and object‑level datasets demonstrate that NOVA3R outperforms state‑of‑the‑art methods in terms of reconstruction accuracy and completeness.
Authors:Yinghong Yu, Guangyuan Li, Jiancheng Yang
Abstract:
Large‑scale 2D foundation models exhibit strong transferable representations, yet extending them to 3D volumetric data typically requires retraining, adapters, or architectural redesign. We introduce PlaneCycle, a training‑free, adapter‑free operator for architecture‑agnostic 2D‑to‑3D lifting of foundation models. PlaneCycle reuses the original pretrained 2D backbone by cyclically distributing spatial aggregation across orthogonal HW, DW, and DH planes throughout network depth, enabling progressive 3D fusion while preserving pretrained inductive biases. The method introduces no additional parameters and is applicable to arbitrary 2D networks. Using pretrained DINOv3 models, we evaluate PlaneCycle on six 3D classification and three 3D segmentation benchmarks. Without any training, the lifted models exhibit intrinsic 3D fusion capability and, under linear probing, outperform slice‑wise 2D baselines and strong 3D counterparts, approaching the performance of fully trained models. With full fine‑tuning, PlaneCycle matches standard 3D architectures, highlighting its potential as a seamless and practical 2D‑to‑3D lifting operator. These results demonstrate that 3D capability can be unlocked from pretrained 2D foundation models without structural modification or retraining. Code is available at https://github.com/HINTLab/PlaneCycle.
Authors:Stefano Berti, Giulia Pasquale, Lorenzo Natale
Abstract:
Few‑Shot Action Recognition (FS‑AR) has shown promising results but is often limited by a closed‑set assumption that fails in real‑world open‑set scenarios. While Few‑Shot Open‑Set (FSOS) recognition is well‑established for images, its extension to spatio‑temporal video data remains underexplored. To address this, we propose an architectural extension based on a Feature‑Residual Discriminator (FR‑Disc), adapting previous work on skeletal data to the more complex video domain. Extensive experiments on five datasets demonstrate that while common open‑set techniques provide only marginal gains, our FR‑Disc significantly enhances unknown rejection capabilities without compromising closed‑set accuracy, setting a new state‑of‑the‑art for FSOS‑AR. The project website, code, and benchmark are available at: https://hsp‑iit.github.io/fsosar/.
Authors:Haoyang Chen, Jing Zhang, Hebaixu Wang, Shiqin Wang, Pohsun Huang, Jiayuan Li, Haonan Guo, Di Wang, Zheng Wang, Bo Du
Abstract:
Multi‑modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross‑modal translation methods treat each modality pair as an independent task, resulting in quadratic complexity and limited generalization to unseen modality combinations. We formulate Any‑to‑Any translation as inference over a shared latent representation of the scene, where different modalities correspond to partial observations of the same underlying semantics. Based on this formulation, we propose Any2Any, a unified latent diffusion framework that projects heterogeneous inputs into a geometrically aligned latent space. Such structure performs anchored latent regression with a shared backbone, decoupling modality‑specific representation learning from semantic mapping. Moreover, lightweight target‑specific residual adapters are used to correct systematic latent mismatches without increasing inference complexity. To support learning under sparse but connected supervision, we introduce RST‑1M, the first million‑scale remote sensing dataset with paired observations across five sensing modalities, providing supervision anchors for any‑to‑any translation. Experiments across 14 translation tasks show that Any2Any consistently outperforms pairwise translation methods and exhibits strong zero‑shot generalization to unseen modality pairs. Code and models are available at https://github.com/MiliLab/Any2Any.
Authors:Yanmei Zou, Hongshan Yu, Yaonan Wang, Zhengeng Yang, Xieyuanli Chen, Kailun Yang, Naveed Akhtar
Abstract:
Multi‑Layer Perceptron (MLP) models are the foundation of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength and limit the application of these models. In this article, we develop a two‑stage abstraction and refinement (ABS‑REF) view for modular feature extraction in point cloud processing. This view elucidates that whereas the early models focused on ABS stages, the more recent techniques devise sophisticated REF stages to attain performance advantages. Then, we propose a High‑dimensional Positional Encoding (HPE) module to explicitly utilize intrinsic positional information, extending the ``positional encoding'' concept from Transformer literature. HPE can be readily deployed in MLP‑based architectures and is compatible with transformer‑based methods. Within our ABS‑REF view, we rethink local aggregation in MLP‑based methods and propose replacing time‑consuming local MLP operations, which are used to capture local relationships among neighbors. Instead, we use non‑local MLPs for efficient non‑local information updates, combined with the proposed HPE for effective local information representation. We leverage our modules to develop HPENets, a suite of MLP networks that follow the ABS‑REF paradigm, incorporating a scalable HPE‑based REF stage. Extensive experiments on seven public datasets across four different tasks show that HPENets deliver a strong balance between efficiency and effectiveness. Notably, HPENet surpasses PointNeXt, a strong MLP‑based counterpart, by 1.1% mAcc, 4.0% mIoU, 1.8% mIoU, and 0.2% Cls. mIoU, with only 50.0%, 21.5%, 23.1%, 44.4% of FLOPs on ScanObjectNN, S3DIS, ScanNet, and ShapeNetPart, respectively. Source code is available at https://github.com/zouyanmei/HPENet_v2.git.
Authors:Simon Warmers, Muhammad Zawish, Fayaz Ali Dharejo, Steven Davy, Radu Timofte
Abstract:
Modeling plant growth dynamics plays a central role in modern agricultural research. However, learning robust predictors from multi‑view plant imagery remains challenging due to strong viewpoint redundancy and viewpoint‑dependent appearance changes. We propose a level‑aware vision language framework that jointly predicts plant age and leaf count using a single multi‑task model built on CLIP embeddings. Our method aggregates rotational views into angle‑invariant representations and conditions visual features on lightweight text priors encoding viewpoint level for stable prediction under incomplete or unordered inputs. On the GroMo25 benchmark, our approach reduces mean age MAE from 7.74 to 3.91 and mean leaf‑count MAE from 5.52 to 3.08 compared to the GroMo baseline, corresponding to improvements of 49.5% and 44.2%, respectively. The unified formulation simplifies the pipeline by replacing the conventional dual‑model setup while improving robustness to missing views. The models and code is available at: https://github.com/SimonWarmers/CLIP‑MVP
Authors:Valentin Biller, Niklas Bubeck, Lucas Zimmer, Ayhan Can Erdur, Sandeep Nagar, Anke Meyer-Baese, Daniel Rückert, Benedikt Wiestler, Jonas Weidner
Abstract:
Glioblastoma exhibits diverse, infiltrative, and patient‑specific growth patterns that are only partially visible on routine MRI, making it difficult to reliably assess true tumor extent and personalize treatment planning and follow‑up. We present a biophysically‑conditioned generative framework that synthesizes biologically realistic 3D brain MRI volumes from estimated, spatially continuous tumor‑concentration fields. Our approach combines a generative model with tumor‑infiltration maps that can be propagated through time using a biophysical growth model, enabling fine‑grained control over tumor shape and growth while preserving patient anatomy. This enables us to synthesize consistent tumor growth trajectories directly in the space of real patients, providing interpretable, controllable estimation of tumor infiltration and progression beyond what is explicitly observed in imaging. We evaluate the framework on longitudinal glioblastoma cases and demonstrate that it can generate temporally coherent sequences with realistic changes in tumor appearance and surrounding tissue response. These results suggest that integrating mechanistic tumor growth priors with modern generative modeling can provide a practical tool for patient‑specific progression visualization and for generating controlled synthetic data to support downstream neuro‑oncology workflows. In longitudinal extrapolation, we achieve a consistent 75% Dice overlap with the biophysical model while maintaining a constant PSNR of 25 in the surrounding tissue. Our code is available at: https://github.com/valentin‑biller/lgm.git
Authors:Tao Yang, Qing Zhou, Yanliang Li, Qi Wang
Abstract:
Reasoning segmentation increasingly employs reinforcement learning to generate explanatory reasoning chains that guide Multimodal Large Language Models. While these geometric rewards are primarily confined to guiding the final localization, they are incapable of discriminating whether the reasoning process remains anchored on the referred region or strays into irrelevant context. Lacking this discriminative guidance, the model's reasoning often devolves into unfocused and verbose chains that ultimately fail to disambiguate and perceive the target in complex scenes. This suggests a need to complement the RL objective with Discriminative Perception, an ability to actively distinguish a target from its context. To realize this, we propose DPAD to compel the model to generate a descriptive caption of the referred object, which is then used to explicitly discriminate by contrasting the caption's semantic relevance to the referred object against the wider context. By optimizing for this discriminative capability, the model is forced to focus on the unique attributes of the target, leading to a more converged and efficient reasoning chain. The descriptive caption also serves as an interpretability rationale that aligns with the segmentation. Experiments on the benchmarks confirm the validity of our approach, delivering substantial performance gains, with the cIoU on ReasonSeg increasing by 3.09% and the reasoning chain length decreasing by approximately 42%. Code is available at https://github.com/mrazhou/DPAD
Authors:Yansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng, Yi Wang, Limin Wang
Abstract:
The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real‑time interactivity. Addressing this gap, we introduce the Real‑tIme Video intERaction Bench (RIVER Bench), designed for evaluating online video comprehension. RIVER Bench introduces a novel framework comprising Retrospective Memory, Live‑Perception, and Proactive Anticipation tasks, closely mimicking interactive dialogues rather than responding to entire videos at once. We conducted detailed annotations using videos from diverse sources and varying lengths, and precisely defined the real‑time interactive format. Evaluations across various model categories reveal that while offline models perform well in single question‑answering tasks, they struggle with real‑time processing. Addressing the limitations of existing models in online video interaction, especially their deficiencies in long‑term memory and future perception, we proposed a general improvement method that enables models to interact with users more flexibly in real time. We believe this work will significantly advance the development of real‑time interactive video understanding models and inspire future research in this emerging field. Datasets and code are publicly available at https://github.com/OpenGVLab/RIVER.
Authors:Qianfeng Yang, Qiyuan Guan, Xiang Chen, Jiyu Jin, Guiyue Jin, Jiangxin Dong
Abstract:
Despite significant progress has been made in image deraining, we note that most existing methods are often developed for only specific types of rain degradation and fail to generalize across diverse real‑world rainy scenes. How to effectively model different rain degradations within a universal framework is important for real‑world image deraining. In this paper, we propose UniRain, an effective unified image deraining framework capable of restoring images degraded by rain streak and raindrop under both daytime and nighttime conditions. To better enhance unified model generalization, we construct an intelligent retrieval augmented generation (RAG)‑based dataset distillation pipeline that selects high‑quality training samples from all public deraining datasets for better mixed training. Furthermore, we incorporate a simple yet effective multi‑objective reweighted optimization strategy into the asymmetric mixture‑of‑experts (MoE) architecture to facilitate consistent performance and improve robustness across diverse scenes. Extensive experiments show that our framework performs favorably against the state‑of‑the‑art models on our proposed benchmarks and multiple public datasets.
Authors:Jaewon Lee, Jaeseok Heo, Gunmin Lee, Howoong Jun, Jeongwoo Oh, Songhwai Oh
Abstract:
Safe visual navigation is critical for indoor mobile robots operating in cluttered environments. Existing benchmarks, however, often neglect collisions or are designed for outdoor scenarios, making them unsuitable for indoor visual navigation. To address this limitation, we introduce the reactive visual navigation benchmark (RVN‑Bench), a collision‑aware benchmark for indoor mobile robots. In RVN‑Bench, an agent must reach sequential goal positions in previously unseen environments using only visual observations and no prior map, while avoiding collisions. Built on the Habitat 2.0 simulator and leveraging high‑fidelity HM3D scenes, RVN‑Bench provides large‑scale, diverse indoor environments, defines a collision‑aware navigation task and evaluation metrics, and offers tools for standardized training and benchmarking. RVN‑Bench supports both online and offline learning by offering an environment for online reinforcement learning, a trajectory image dataset generator, and tools for producing negative trajectory image datasets that capture collision events. Experiments show that policies trained on RVN‑Bench generalize effectively to unseen environments, demonstrating its value as a standardized benchmark for safe and robust visual navigation. Code and additional materials are available at: https://rvn‑bench.github.io/.
Authors:Yanguang Zhao, Jie Yang, Shengqiong Wu, Shutong Hu, Hongbo Qiu, Yu Wang, Guijia Zhang, Tan Kai Ze, Hao Fei, Chia-Wen Lin, Mong-Li Lee, Wynne Hsu
Abstract:
Spatial reasoning, the ability to understand spatial relations, causality, and dynamic evolution, is central to human intelligence and essential for real‑world applications such as autonomous driving and robotics. Existing studies, however, primarily assess models on visible spatio‑temporal understanding, overlooking their ability to infer unseen past or future spatial states. In this work, we introduce Spatial Causal Prediction (SCP), a new task paradigm that challenges models to reason beyond observation and predict spatial causal outcomes. We further construct SCP‑Bench, a benchmark comprising 2,500 QA pairs across 1,181 videos spanning diverse viewpoints, scenes, and causal directions, to support systematic evaluation. Through comprehensive experiments on 23 state‑of‑the‑art models, we reveal substantial gaps between human and model performance, limited temporal extrapolation, and weak causal grounding. We further analyze key factors influencing performance and propose perception‑enhancement and reasoning‑guided strategies toward advancing spatial causal intelligence. The project page is https://guangstrip.github.io/SCP‑Bench.
Authors:Radia Daci, Vito Renò, Cosimo Patruno, Angelo Cardellicchio, Abdelmalik Taleb-Ahmed, Marco Leo, Cosimo Distante
Abstract:
Multimodal industrial anomaly detection benefits from integrating RGB appearance with 3D surface geometry, yet existing \emphunsupervised approaches commonly rely on memory banks, teacher‑student architectures, or fragile fusion schemes, limiting robustness under noisy depth, weak texture, or missing modalities. This paper introduces CMDR‑IAD, a lightweight and modality‑flexible unsupervised framework for reliable anomaly detection in 2D+3D multimodal as well as single‑modality (2D‑only or 3D‑only) settings. CMDR‑IAD combines bidirectional 2D\leftrightarrow3D cross‑modal mapping to model appearance‑geometry consistency with dual‑branch reconstruction that independently captures normal texture and geometric structure. A two‑part fusion strategy integrates these cues: a reliability‑gated mapping anomaly highlights spatially consistent texture‑geometry discrepancies, while a confidence‑weighted reconstruction anomaly adaptively balances appearance and geometric deviations, yielding stable and precise anomaly localization even in depth‑sparse or low‑texture regions. On the MVTec 3D‑AD benchmark, CMDR‑IAD achieves state‑of‑the‑art performance while operating without memory banks, reaching 97.3% image‑level AUROC (I‑AUROC), 99.6% pixel‑level AUROC (P‑AUROC), and 97.6% AUPRO. On a real‑world polyurethane cutting dataset, the 3D‑only variant attains 92.6% I‑AUROC and 92.5% P‑AUROC, demonstrating strong effectiveness under practical industrial conditions. These results highlight the framework's robustness, modality flexibility, and the effectiveness of the proposed fusion strategies for industrial visual inspection. Our source code is available at https://github.com/ECGAI‑Research/CMDR‑IAD/
Authors:Felix Igelbrink, Lennart Niecksch, Martin Atzmueller, Joachim Hertzberg
Abstract:
Open‑set semantic mapping enables language‑driven robotic perception, but current instance‑centric approaches are bottlenecked by context‑depriving and computationally expensive crop‑based feature extraction. To overcome this fundamental limitation, we introduce DISC (Dense Integrated Semantic Context), featuring a novel single‑pass, distance‑weighted extraction mechanism. By deriving high‑fidelity CLIP embeddings directly from the vision transformer's intermediate layers, our approach eliminates the latency and domain‑shift artifacts of traditional image cropping, yielding pure, mask‑aligned semantic representations. To fully leverage these features in large‑scale continuous mapping, DISC is built upon a fully GPU‑accelerated architecture that replaces periodic offline processing with precise, on‑the‑fly voxel‑level instance refinement. We evaluate our approach on standard benchmarks (Replica, ScanNet) and a newly generated large‑scale‑mapping dataset based on Habitat‑Matterport 3D (HM3DSEM) to assess scalability across complex scenes in multi‑story buildings. Extensive evaluations demonstrate that DISC significantly surpasses current state‑of‑the‑art zero‑shot methods in both semantic accuracy and query retrieval, providing a robust, real‑time capable framework for robotic deployment. The full source code, data generation and evaluation pipelines will be made available at https://github.com/DFKI‑NI/DISC.
Authors:Yang Li, Youyang Sha, Yinzhi Wang, Timothy Hospedales, Xi Shen, Shell Xu Hu, Xuanlong Yu
Abstract:
Building reliable classifiers is a fundamental challenge for deploying machine learning in real‑world applications. A reliable system should not only detect out‑of‑distribution (OOD) inputs but also anticipate in‑distribution (ID) errors by assigning low confidence to potentially misclassified samples. Yet, most prior work treats OOD detection and failure prediction as separated problems, overlooking their closed connection. We argue that reliability requires evaluating them jointly. To this end, we propose a unified evaluation framework that integrates OOD detection and failure prediction, quantified by our new metrics DS‑F1 and DS‑AURC, where DS denotes double scoring functions. Experiments on the OpenOOD benchmark show that double scoring functions yield classifiers that are substantially more reliable than traditional single scoring approaches. Our analysis further reveals that OOD‑based approaches provide notable gains under simple or far‑OOD shifts, but only marginal benefits under more challenging near‑OOD conditions. Beyond evaluation, we extend the reliable classifier SURE and introduce SURE+, a new approach that significantly improves reliability across diverse scenarios. Together, our framework, metrics, and method establish a new benchmark for trustworthy classification and offer practical guidance for deploying robust models in real‑world settings. The source code is publicly available at https://github.com/Intellindust‑AI‑Lab/SUREPlus.
Authors:Jinyuan Liu, Xingyuan Li, Qingyun Mei, Haoyuan Xu, Zhiying Jiang, Long Ma, Risheng Liu, Xin Fan
Abstract:
Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and objective metrics, often resulting in fusion outcomes that do not align with human visual preferences. This challenge is further exacerbated by the ill‑posed nature of IVIF, which severely limits its effectiveness in human perceptual environments such as security surveillance and driver assistance systems. To address these limitations, we propose a feedback reinforcement framework that bridges human evaluation to infrared and visible image fusion. To address the lack of human‑centric evaluation metrics and data, we introduce the first large‑scale human feedback dataset for IVIF, containing multidimensional subjective scores and artifact annotations, and enriched by a fine‑tuned large language model with expert review. Based on this dataset, we design a domain‑specific reward function and train a reward model to quantify perceptual quality. Guided by this reward, we fine‑tune the fusion network through Group Relative Policy Optimization, achieving state‑of‑the‑art performance that better aligns fused images with human aesthetics. Code is available at https://github.com/ALKA‑Wind/EVAFusion.
Authors:Ruilin Luo, Chufan Shi, Yizhen Zhang, Cheng Yang, Songtao Jiang, Tongkun Guan, Ruizhe Chen, Ruihang Chu, Peng Wang, Mingkun Yang, Yujiu Yang, Junyang Lin, Zhibo Yang
Abstract:
The cold‑start initialization stage plays a pivotal role in training Multimodal Large Reasoning Models (MLRMs), yet its mechanisms remain insufficiently understood. To analyze this stage, we introduce the Visual Attention Score (VAS), an attention‑based metric that quantifies how much a model attends to visual tokens. We find that reasoning performance is strongly correlated with VAS (r=0.9616): models with higher VAS achieve substantially stronger multimodal reasoning. Surprisingly, multimodal cold‑start fails to elevate VAS, resulting in attention distributions close to the base model, whereas text‑only cold‑start leads to a clear increase. We term this counter‑intuitive phenomenon Lazy Attention Localization. To validate its causal role, we design training‑free interventions that directly modulate attention allocation during inference, performance gains of 1‑2% without any retraining. Building on these insights, we further propose Attention‑Guided Visual Anchoring and Reflection (AVAR), a comprehensive cold‑start framework that integrates visual‑anchored data synthesis, attention‑guided objectives, and visual‑anchored reward shaping. Applied to Qwen2.5‑VL‑7B, AVAR achieves an average gain of 7.0% across 7 multimodal reasoning benchmarks. Ablation studies further confirm that each component of AVAR contributes step‑wise to the overall gains. The code, data, and models are available at https://github.com/lrlbbzl/Qwen‑AVAR.
Authors:Taejun Lim, Joong-Won Hwang, Kibok Lee
Abstract:
When continual test‑time adaptation (TTA) persists over the long term, errors accumulate in the model and further cause it to predict only a few classes for all inputs, a phenomenon known as model collapse. Recent studies have explored reset strategies that completely erase these accumulated errors. However, their periodic resets lead to suboptimal adaptation, as they occur independently of the actual risk of collapse. Moreover, their full resets cause catastrophic loss of knowledge acquired over time, even though such knowledge could be beneficial in the future. To this end, we propose (1) an Adaptive and Selective Reset (ASR) scheme that dynamically determines when and where to reset, (2) an importance‑aware regularizer to recover essential knowledge lost due to reset, and (3) an on‑the‑fly adaptation adjustment scheme to enhance adaptability under challenging domain shifts. Extensive experiments across long‑term TTA benchmarks demonstrate the effectiveness of our approach, particularly under challenging conditions. Our code is available at https://github.com/YonseiML/asr.
Authors:Tuan Duc Ngo, Jiahui Huang, Seoung Wug Oh, Kevin Blackburn-Matzen, Evangelos Kalogerakis, Chuang Gan, Joon-Young Lee
Abstract:
Estimating accurate, view‑consistent geometry and camera poses from uncalibrated multi‑view/video inputs remains challenging ‑ especially at high spatial resolutions and over long sequences. We present DAGE, a dual‑stream transformer whose main novelty is to disentangle global coherence from fine detail. A low‑resolution stream operates on aggressively downsampled frames with alternating frame/global attention to build a view‑consistent representation and estimate cameras efficiently, while a high‑resolution stream processes the original images per‑frame to preserve sharp boundaries and small structures. A lightweight adapter fuses these streams via cross‑attention, injecting global context without disturbing the pretrained single‑frame pathway. This design scales resolution and clip length independently, supports inputs up to 2K, and maintains practical inference cost. DAGE delivers sharp depth/pointmaps, strong cross‑view consistency, and accurate poses, establishing new state‑of‑the‑art results for video geometry estimation and multi‑view reconstruction.
Authors:Risto Ojala, Tristan Ellison, Mo Chen
Abstract:
Glass surface segmentation from RGB images is a challenging task, since glass as a transparent material distinctly lacks visual characteristics. However, glass segmentation is critical for scene understanding and robotics, as transparent glass surfaces must be identified as solid material. This paper presents a novel architecture for glass segmentation, deploying a dual‑backbone producing general visual features as well as task‑specific learned visual features. General visual features are produced by a frozen DINOv3 vision foundation model, and the task‑specific features are generated with a Swin model trained in a supervised manner. Resulting multi‑scale feature representations are downsampled with residual Squeeze‑and‑Excitation Channel Reduction, and fed into a Mask2Former Decoder, producing the final segmentation masks. The architecture was evaluated on four commonly used glass segmentation datasets, achieving state‑of‑the‑art results on several accuracy metrics. The model also has a competitive inference speed compared to the previous state‑of‑the‑art method, and surpasses it when using a lighter DINOv3 backbone variant. The implementation source code and model weights are available at: https://github.com/ojalar/lgnet
Authors:Inho Kong, Sojin Lee, Youngjoon Hong, Hyunwoo J. Kim
Abstract:
Classifier‑Free Guidance (CFG) has established the foundation for guidance mechanisms in diffusion models, showing that well‑designed guidance proxies significantly improve conditional generation and sample quality. Autoguidance (AG) has extended this idea, but it relies on an auxiliary network and leaves solver‑induced errors unaddressed. In stiff regions, the ODE trajectory changes sharply, where local truncation error (LTE) becomes a critical factor that deteriorates sample quality. Our key observation is that these errors align with the dominant eigenvector, motivating us to leverage the solver‑induced error as a guidance signal. We propose Embedded Runge‑Kutta Guidance (ERK‑Guid), which exploits detected stiffness to reduce LTE and stabilize sampling. We theoretically and empirically analyze stiffness and eigenvector estimators with solver errors to motivate the design of ERK‑Guid. Our experiments on both synthetic datasets and the popular benchmark dataset, ImageNet, demonstrate that ERK‑Guid consistently outperforms state‑of‑the‑art methods. Code is available at https://github.com/mlvlab/ERK‑Guid.
Authors:Zhiqiang Sheng, Xumeng Han, Zhiwei Zhang, Zenghui Xiong, Yifan Ding, Aoxiang Ping, Xiang Li, Tong Guo, Yao Mao
Abstract:
Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically does not extend to complex scenarios requiring dynamic reasoning, leaving them ill‑equipped to model the coherent, intermediate logical pathways that constitute a multi‑step evolution from an initial state to a final one. This capacity is crucial for unlocking a deeper level of procedural and causal understanding in visual manipulation. To systematically measure this critical limitation, we introduce InEdit‑Bench, the first evaluation benchmark dedicated to reasoning over intermediate pathways in image editing. InEdit‑Bench comprises meticulously annotated test cases covering four fundamental task categories: state transition, dynamic process, temporal sequence, and scientific simulation. Additionally, to enable fine‑grained evaluation, we propose a set of assessment criteria to evaluate the logical coherence and visual naturalness of the generated pathways, as well as the model's fidelity to specified path constraints. Our comprehensive evaluation of 14 representative image editing models on InEdit‑Bench reveals significant and widespread shortcomings in this domain. By providing a standardized and challenging benchmark, we aim for InEdit‑Bench to catalyze research and steer development towards more dynamic, reason‑aware, and intelligent multimodal generative models.
Authors:Hao Li, Yuhao Wang, Wenning Hao, Pingping Zhang, Dong Wang, Huchuan Lu
Abstract:
RGB‑Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing RGBT trackers rely solely on initial‑frame visual information for target modeling, failing to adapt to appearance variations due to the absence of language guidance. Furthermore, current methods suffer from redundant search regions and heterogeneous modality gaps, causing background distraction. To address these issues, we first introduce textual descriptions into RGBT tracking benchmarks. This is accomplished through a pipeline that leverages Multi‑modal Large Language Models (MLLMs) to automatically produce texual annotations. Afterwards, we propose RAGTrack, a novel Retrieval‑Augmented Generation framework for robust RGBT tracking. To this end, we introduce a Multi‑modal Transformer Encoder (MTE) for unified visual‑language modeling. Then, we design an Adaptive Token Fusion (ATF) to select target‑relevant tokens and perform channel exchanges based on cross‑modal correlations, mitigating search redundancies and modality gaps. Finally, we propose a Context‑aware Reasoning Module (CRM) to maintain a dynamic knowledge base and employ a Retrieval‑Augmented Generation (RAG) to enable temporal linguistic reasoning for robust target modeling. Extensive experiments on four RGBT benchmarks demonstrate that our framework achieves state‑of‑the‑art performance across various challenging scenarios. The source code is available https://github.com/IdolLab/RAGTrack.
Authors:Yan Tian, Pengcheng Xue, Weiping Ding, Mahmoud Hassaballah, Karen Egiazarian, Aura Conci, Abdulkadir Sengur, Leszek Rutkowski
Abstract:
The automatic design of a 3D tooth model plays a crucial role in dental digitization. However, current approaches face challenges in compositional 3D tooth generation because both the layouts and shapes of missing teeth need to be optimized.In addition, collision conflicts are often omitted in 3D Gaussian‑based compositional 3D generation, where objects may intersect with each other due to the absence of explicit geometric information on the object surfaces. Motivated by graph generation through diffusion models and collision detection using 3D Gaussians, we propose an approach named DM‑CFO for compositional tooth generation, where the layout of missing teeth is progressively restored during the denoising phase under both text and graph constraints. Then, the Gaussian parameters of each layout‑guided tooth and the entire jaw are alternately updated using score distillation sampling (SDS). Furthermore, a regularization term based on the distances between the 3D Gaussians of neighboring teeth and the anchor tooth is introduced to penalize tooth intersections. Experimental results on three tooth‑design datasets demonstrate that our approach significantly improves the multiview consistency and realism of the generated teeth compared with existing methods. Project page: https://amateurc.github.io/CF‑3DTeeth/.
Authors:Xu Yao, Lei Kang
Abstract:
Scene text recognition (STR) and handwritten text recognition (HTR) face significant challenges in accurately transcribing textual content from images into machine‑readable formats. Conventional OCR models often predict transcriptions directly, which limits detailed reasoning about text structure. We propose a VQA‑inspired data augmentation framework that strengthens OCR training through structured question‑answering tasks. For each image‑text pair, we generate natural‑language questions probing character‑level attributes such as presence, position, and frequency, with answers derived from ground‑truth text. These auxiliary tasks encourage finer‑grained reasoning, and the OCR model aligns visual features with textual queries to jointly reason over images and questions. Experiments on WordArt and Esposalles datasets show consistent improvements over baseline models, with significant reductions in both CER and WER. Our code is publicly available at https://github.com/xuyaooo/DataAugOCR.
Authors:Qifan Zhang, Sai Haneesh Allu, Jikai Wang, Yangxiao Lu, Yu Xiang
Abstract:
Detecting and segmenting novel object instances in open‑world environments is a fundamental problem in robotic perception. Given only a small set of template images, a robot must locate and segment a specific object instance in a cluttered, previously unseen scene. Existing proposal‑based approaches are highly sensitive to proposal quality and often fail under occlusion and background clutter. We propose L2G‑Det, a local‑to‑global instance detection framework that bypasses explicit object proposals by leveraging dense patch‑level matching between templates and the query image. Locally matched patches generate candidate points, which are refined through a candidate selection module to suppress false positives. The filtered points are then used to prompt an augmented Segment Anything Model (SAM) with instance‑specific object tokens, enabling reliable reconstruction of complete instance masks. Experiments demonstrate improved performance over proposal‑based methods in challenging open‑world settings.
Authors:Yimin Zhu, Zack Dewis, Quinn Ledingham, Saeid Taleghanidoozdoozan, Mabel Heffring, Zhengsen Xu, Motasem Alkayid, Megan Greenwood, Lincoln Linlin Xu
Abstract:
Recently, DeepSeek has invented the manifold‑constrained hyper‑connection (mHC) approach which has demonstrated significant improvements over the traditional residual connection in deep learning models \citexie2026mhc. Nevertheless, this approach has not been tailor‑designed for improving hyperspectral image (HSI) classification. This paper presents a clustering‑guided mHC Mamba model (mHC‑HSI) for enhanced HSI classification, with the following contributions. First, to improve spatial‑spectral feature learning, we design a novel clustering‑guided Mamba module, based on the mHC framework, that explicitly learns both spatial and spectral information in HSI. Second, to decompose the complex and heterogeneous HSI into smaller clusters, we design a new implementation of the residual matrix in mHC, which can be treated as soft cluster membership maps, leading to improved explainability of the mHC approach. Third, to leverage the physical spectral knowledge, we divide the spectral bands into physically‑meaningful groups and use them as the "parallel streams" in mHC, leading to a physically‑meaningful approach with enhanced interpretability. The proposed approach is tested on benchmark datasets in comparison with the state‑of‑the‑art methods, and the results suggest that the proposed model not only improves the accuracy but also enhances the model explainability. Code is available here: https://github.com/GSIL‑UCalgary/mHC_HyperSpectral
Authors:Yujia Zhang, Xiaoyang Wu, Yunhan Yang, Xianzhe Fan, Han Li, Yuechen Zhang, Zehao Huang, Naiyan Wang, Hengshuang Zhao
Abstract:
We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self‑supervised point transformer encoder across diverse domains, spanning remote sensing, outdoor LiDAR, indoor RGB‑D sequences, object‑centric CAD models, and point clouds lifted from RGB‑only videos. Despite their distinct sensing geometries, densities, and priors, Utonia learns a consistent representation space that transfers across domains. This unification improves perception capability while revealing intriguing emergent behaviors that arise only when domains are trained jointly. Beyond perception, we observe that Utonia representations can also benefit embodied and multimodal reasoning: conditioning vision‑language‑action policies on Utonia features improves robotic manipulation, and integrating them into vision‑language models yields gains on spatial reasoning. We hope Utonia can serve as a step toward foundation models for sparse 3D data, and support downstream applications in AR/VR, robotics, and autonomous driving.
Authors:Hanyang Wang, Yiyang Liu, Jiawei Chi, Fangfu Liu, Ran Xue, Yueqi Duan
Abstract:
Classifier‑Free Guidance (CFG) has emerged as a central approach for enhancing semantic alignment in flow‑based diffusion models. In this paper, we explore a unified framework called CFG‑Ctrl, which reinterprets CFG as a control applied to the first‑order continuous‑time generative flow, using the conditional‑unconditional discrepancy as an error signal to adjust the velocity field. From this perspective, we summarize vanilla CFG as a proportional controller (P‑control) with fixed gain, and typical follow‑up variants develop extended control‑law designs derived from it. However, existing methods mainly rely on linear control, inherently leading to instability, overshooting, and degraded semantic fidelity especially on large guidance scales. To address this, we introduce Sliding Mode Control CFG (SMC‑CFG), which enforces the generative flow toward a rapidly convergent sliding manifold. Specifically, we define an exponential sliding mode surface over the semantic prediction error and introduce a switching control term to establish nonlinear feedback‑guided correction. Moreover, we provide a Lyapunov stability analysis to theoretically support finite‑time convergence. Experiments across text‑to‑image generation models including Stable Diffusion 3.5, Flux, and Qwen‑Image demonstrate that SMC‑CFG outperforms standard CFG in semantic alignment and enhances robustness across a wide range of guidance scales. Project Page: https://hanyang‑21.github.io/CFG‑Ctrl
Authors:Toru Lin, Shuying Deng, Zhao-Heng Yin, Pieter Abbeel, Jitendra Malik
Abstract:
Many essential manipulation tasks ‑ such as food preparation, surgery, and craftsmanship ‑ remain intractable for autonomous robots. These tasks are characterized not only by contact‑rich, force‑sensitive dynamics, but also by their "implicit" success criteria: unlike pick‑and‑place, task quality in these domains is continuous and subjective (e.g. how well a potato is peeled), making quantitative evaluation and reward engineering difficult. We present a learning framework for such tasks, using peeling with a knife as a representative example. Our approach follows a two‑stage pipeline: first, we learn a robust initial policy via force‑aware data collection and imitation learning, enabling generalization across object variations; second, we refine the policy through preference‑based finetuning using a learned reward model that combines quantitative task metrics with qualitative human feedback, aligning policy behavior with human notions of task quality. Using only 50‑200 peeling trajectories, our system achieves over 90% average success rates on challenging produce including cucumbers, apples, and potatoes, with performance improving by up to 40% through preference‑based finetuning. Remarkably, policies trained on a single produce category exhibit strong zero‑shot generalization to unseen in‑category instances and to out‑of‑distribution produce from different categories while maintaining over 90% success rates.
Authors:Yufu Wang, Evonne Ng, Soyong Shin, Rawal Khirodkar, Yuan Dong, Zhaoen Su, Jinhyung Park, Kris Kitani, Alexander Richard, Fabian Prada, Michael Zollhofer
Abstract:
We present DuoMo, a generative method that recovers human motion in world‑space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade‑off: generalizing from diverse and noisy video inputs while maintaining global motion consistency. Our approach addresses this problem by factorizing motion learning into two diffusion models. The camera‑space model first estimates motion from videos in camera coordinates. The world‑space model then lifts this initial estimate into world coordinates and refines it to be globally consistent. Together, the two models can reconstruct motion across diverse scenes and trajectories, even from highly noisy or incomplete observations. Moreover, our formulation is general, generating the motion of mesh vertices directly and bypassing parametric models. DuoMo achieves state‑of‑the‑art performance. On EMDB, our method obtains a 16% reduction in world‑space reconstruction error while maintaining low foot skating. On RICH, it obtains a 30% reduction in world‑space error. Project page: https://yufu‑wang.github.io/duomo/
Authors:Miguel Espinosa, Eva Gmelich Meijling, Valerio Marsocci, Elliot J. Crowley, Mikolaj Czerkawski
Abstract:
Earth observation applications increasingly rely on data from multiple sensors, including optical, radar, elevation, and land‑cover. Relationships between modalities are fundamental for data integration but are inherently non‑injective: identical conditioning information can correspond to multiple physically plausible observations, and should be parametrised as conditional distributions. Deterministic models, by contrast, collapse toward conditional means and fail to represent the uncertainty and variability required for tasks such as data completion and cross‑sensor translation. We introduce COP‑GEN, a multimodal latent diffusion transformer that models the joint distribution of heterogeneous EO modalities at their native spatial resolutions. By parameterising cross‑modal mappings as conditional distributions, COP‑GEN enables flexible any‑to‑any conditional generation, including zero‑shot modality translation without task‑specific retraining. Experiments show that COP‑GEN generates diverse yet physically consistent realisations while maintaining strong peak fidelity across optical, radar, and elevation modalities. Qualitative and quantitative analyses demonstrate that the model captures meaningful cross‑modal structure and adapts its output uncertainty as conditioning information increases. We release a stochastic benchmark built from multi‑temporal Sentinel‑2 observations that enables distribution‑level comparison of generative EO models. On this benchmark, COP‑GEN covers 90% of the real observation manifold and 63% of its per‑band reflectance range, while the strongest competing method collapses to 2.8% and 18%, respectively. These results highlight the importance of stochastic generative modeling for EO and motivate evaluation protocols beyond single‑reference, pointwise metrics. Website: https://miquel‑espinosa.github.io/cop‑gen
Authors:Ziyang Gong, Zehang Luo, Anke Tang, Zhe Liu, Shi Fu, Zhi Hou, Ganlin Yang, Weiyun Wang, Xiaofeng Wang, Jianbo Liu, Gen Luo, Haolan Kang, Shuang Luo, Yue Zhou, Yong Luo, Li Shen, Xiaosong Jia, Yao Mu, Xue Yang, Chunxiao Liu, Junchi Yan, Hengshuang Zhao, Dacheng Tao, Xiaogang Wang
Abstract:
Universal embodied intelligence demands robust generalization across heterogeneous embodiments, such as autonomous driving, robotics, and unmanned aerial vehicles (UAVs). However, existing embodied brain in training a unified model over diverse embodiments frequently triggers long‑tail data, gradient interference, and catastrophic forgetting, making it notoriously difficult to balance universal generalization with domain‑specific proficiency. In this report, we introduce ACE‑Brain‑0, a generalist foundation brain that unifies spatial reasoning, autonomous driving, and embodied manipulation within a single multimodal large language model~(MLLM). Our key insight is that spatial intelligence serves as a universal scaffold across diverse physical embodiments: although vehicles, robots, and UAVs differ drastically in morphology, they share a common need for modeling 3D mental space, making spatial cognition a natural, domain‑agnostic foundation for cross‑embodiment transfer. Building on this insight, we propose the Scaffold‑Specialize‑Reconcile~(SSR) paradigm, which first establishes a shared spatial foundation, then cultivates domain‑specialized experts, and finally harmonizes them through data‑free model merging. Furthermore, we adopt Group Relative Policy Optimization~(GRPO) to strengthen the model's comprehensive capability. Extensive experiments demonstrate that ACE‑Brain‑0 achieves competitive and even state‑of‑the‑art performance across 24 spatial and embodiment‑related benchmarks.
Authors:Samuele Angheben, Davide Berasi, Alessandro Conti, Elisa Ricci, Yiming Wang
Abstract:
Classifying fine‑grained visual concepts under open‑world settings, i.e., without a predefined label set, demands models to be both accurate and specific. Recent reasoning Large Multimodal Models (LMMs) exhibit strong visual understanding capability but tend to produce overly generic predictions when performing fine‑grained image classification. Our preliminary analysis reveals that models do possess the intrinsic fine‑grained domain knowledge. However, promoting more specific predictions (specificity) without compromising correct ones (correctness) remains a non‑trivial and understudied challenge. In this work, we investigate how to steer reasoning LMMs toward predictions that are both correct and specific. We propose a novel specificity‑aware reinforcement learning framework, SpeciaRL, to fine‑tune reasoning LMMs on fine‑grained image classification under the open‑world setting. SpeciaRL introduces a dynamic, verifier‑based reward signal anchored to the best predictions within online rollouts, promoting specificity while respecting the model's capabilities to prevent incorrect predictions. Our out‑of‑domain experiments show that SpeciaRL delivers the best trade‑off between correctness and specificity across extensive fine‑grained benchmarks, surpassing existing methods and advancing open‑world fine‑grained image classification. Code and model are publicly available at https://github.com/s‑angheben/SpeciaRL.
Authors:Fuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang, Lei Fan, Hao Li, Chen Wei, Tonghua Su, Baorui Ma
Abstract:
Vision‑Language‑Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal‑causal structure underlying visual dynamics. World‑model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. Latent‑action VLAs encode frame‑to‑frame transitions compactly, but lack temporally continuous dynamic modeling and world knowledge. To overcome these limitations, we introduce CoWVLA (Chain‑of‑World VLA), a new "Chain of World" paradigm that unifies world‑model temporal reasoning with a disentangled latent motion representation. First, a pretrained video VAE serves as a latent motion extractor, explicitly factorizing video segments into structure and motion latents. Then, during pre‑training, the VLA learns from an instruction and an initial frame to infer a continuous latent motion chain and predict the segment's terminal frame. Finally, during co‑fine‑tuning, this latent dynamic is aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder. This design preserves the world‑model benefits of temporal reasoning and world knowledge while retaining the compactness and interpretability of latent actions, enabling efficient visuomotor learning. Extensive experiments on robotic simulation benchmarks show that CoWVLA outperforms existing world‑model and latent‑action approaches and achieves moderate computational efficiency, highlighting its potential as a more effective VLA pretraining paradigm. The project website can be found at https://fx‑hit.github.io/cowvla‑io.
Authors:Chun-Wun Cheng, Yanqi Cheng, Peiyuan Jing, Guang Yang, Javier A. Montoya-Zegarra, Carola-Bibiane Schönlieb, Angelica I. Aviles-Rivero
Abstract:
Medical image segmentation commonly relies on U‑shaped encoder‑decoder architectures such as U‑Net, where skip connections preserve fine spatial detail by injecting high‑resolution encoder features into the decoder. However, these skip pathways also propagate low‑level textures, background clutter, and acquisition noise, allowing irrelevant information to bypass deeper semantic filtering ‑‑ an issue that is particularly detrimental in low‑contrast clinical imaging. Although attention gates have been introduced to address this limitation, they typically produce dense sigmoid masks that softly reweight features rather than explicitly removing irrelevant activations. We propose ProSMA‑UNet (Proximal‑Sparse Multi‑Scale Attention U‑Net), which reformulates skip gating as a decoder‑conditioned sparse feature selection problem. ProSMA constructs a multi‑scale compatibility field using lightweight depthwise dilated convolutions to capture relevance across local and contextual scales, then enforces explicit sparsity via an \ell_1 proximal operator with learnable per‑channel thresholds, yielding a closed‑form soft‑thresholding gate that can remove noisy responses. To further suppress semantically irrelevant channels, ProSMA incorporates decoder‑conditioned channel gating driven by global decoder context. Extensive experiments on challenging 2D and 3D benchmarks demonstrate state‑of‑the‑art performance, with particularly large gains (\approx20%) on difficult 3D segmentation tasks. Project page: https://math‑ml‑x.github.io/ProSMA‑UNet/
Authors:Jun Yeong Park, JunYoung Seo, Minji Kang, Yu Rang Park
Abstract:
The CLIP model's outstanding generalization has driven recent success in Zero‑Shot Anomaly Detection (ZSAD) for detecting anomalies in unseen categories. The core challenge in ZSAD is to specialize the model for anomaly detection tasks while preserving CLIP's powerful generalization capability. Existing approaches attempting to solve this challenge share the fundamental limitation of a patch‑agnostic design that processes all patches monolithically without regard for their unique characteristics. To address this limitation, we propose MoECLIP, a Mixture‑of‑Experts (MoE) architecture for the ZSAD task, which achieves patch‑level adaptation by dynamically routing each image patch to a specialized Low‑Rank Adaptation (LoRA) expert based on its unique characteristics. Furthermore, to prevent functional redundancy among the LoRA experts, we introduce (1) Frozen Orthogonal Feature Separation (FOFS), which orthogonally separates the input feature space to force experts to focus on distinct information, and (2) a simplex equiangular tight frame (ETF) loss to regulate the expert outputs to form maximally equiangular representations. Comprehensive experimental results across 14 benchmark datasets spanning industrial and medical domains demonstrate that MoECLIP outperforms existing state‑of‑the‑art methods. The code is available at https://github.com/CoCoRessa/MoECLIP.
Authors:Baoliang Chen, Xinlong Bu, Hanwei Zhu, Lingyu Zhu, Jieyu Zhan
Abstract:
Existing AI‑generated video quality assessment (AIGVQA) methods mainly focus on global perceptual realism and coarse text‑video alignment, while overlooking a critical requirement in educational scenarios: concept correctness. In early mathematics education, subtle errors in numerical quantities, geometric relations, or spatial configurations may fundamentally alter the conveyed knowledge despite visually plausible generation. To address this problem, we introduce EduAVQABench, the first benchmark for concept‑aware educational AIGV assessment, containing 1,130 videos generated by ten state‑of‑the‑art T2V models together with over 310,650 fine‑grained human annotations spanning perceptual quality and semantic alignment. Built upon this benchmark, we further propose EduVQA, a concept‑aware AIGVQA framework equipped with a Structured 2D Mixture‑of‑Experts (S2D‑MoE) architecture. By jointly modeling fine‑grained concept assessment and overall quality prediction through shared experts and adaptive two‑dimensional routing, EduVQA effectively captures subtle concept‑level inconsistencies overlooked by conventional global scoring methods. Extensive experiments demonstrate that EduVQA consistently outperforms existing AIGVQA approaches across both perceptual and semantic evaluation tasks while exhibiting strong generalization capability on unseen benchmarks. Code and dataset will be publicly available at: https://github.com/EduVQA/EduVQA.
Authors:Wenqing Cui, Zhenyu Li, Mykola Lavreniuk, Jian Shi, Ramzi Idoughi, Xiangjun Tang, Peter Wonka
Abstract:
Joint estimation of surface normals and depth is essential for holistic 3D scene understanding, yet high‑resolution prediction remains difficult due to the trade‑off between preserving fine local detail and maintaining global consistency. To address this challenge, we propose the Ultra Resolution Geometry Transformer (URGT), which adapts the Visual Geometry Grounded Transformer (VGGT) into a unified multi‑patch transformer for monocular high‑resolution depth‑‑normal estimation. A single high‑resolution image is partitioned into patches that are augmented with coarse depth and normal priors from pre‑trained models, and jointly processed in a single forward pass to predict refined geometric outputs. Global coherence is enforced through cross‑patch attention, which enables long‑range geometric reasoning and seamless propagation of information across patches within a shared backbone. To further enhance spatial robustness, we introduce a GridMix patch sampling strategy that probabilistically samples grid configurations during training, improving inter‑patch consistency and generalization. Our method achieves state‑of‑the‑art results on UnrealStereo4K, jointly improving depth and normal estimation, reducing AbsRel from 0.0582 to 0.0291, RMSE from 2.17 to 1.31, and lowering mean angular error from 23.36 degrees to 18.51 degrees, while producing sharper and more stable geometry. The proposed multi‑patch framework also demonstrates strong zero‑shot and cross‑domain generalization and scales effectively to very high resolutions, offering an efficient and extensible solution for high‑quality geometry refinement.
Authors:Ertunc Erdil, Nico Schulthess, Guney Tombak, Ender Konukoglu
Abstract:
DINO models provide rich patch‑level representations that have recently enabled strong performance in unsupervised anomaly detection (UAD). Most existing methods extract patch embeddings from ``normal'' images and model them independently, ignoring spatial and neighborhood relationships between patches. This implicitly assumes that self‑attention and positional encodings sufficiently encode contextual information within each patch embedding. In addition, the normative distribution is often modeled as memory banks or prototype‑based representations, which require storing large numbers of features and performing costly comparisons at inference time, leading to substantial memory and computational overhead. In this work, we address both limitations by proposing a simple and efficient framework that explicitly models spatial and contextual dependencies between patch embeddings using a 2D autoregressive (AR) model. Instead of storing embeddings or clustering prototypes, our approach learns a compact parametric model of the normative distribution via an AR convolutional neural network (CNN). At test time, anomaly detection reduces to a single forward pass through the network and enables fast and memory‑efficient inference. We evaluate our method on the BMAD benchmark, which comprises three medical imaging datasets, and compare it against existing work including recent DINO‑based methods. Experimental results demonstrate that explicitly modeling spatial dependencies achieves competitive anomaly detection performance while substantially reducing inference time and memory requirements. Code is available at the project page: https://eerdil.github.io/spatial‑ar‑dinov3‑uad/.
Authors:Jiaxing Liu, Zexi Zhang, Xiaoyan Li, Boyue Wang, Yongli Hu, Baocai Yin
Abstract:
Vision‑Language Navigation (VLN) presents a unique challenge for Large Vision‑Language Models (VLMs) due to their inherent architectural mismatch: VLMs are primarily pretrained on static, disembodied vision‑language tasks, which fundamentally clash with the dynamic, embodied, and spatially‑structured nature of navigation. Existing large‑model‑based methods often resort to converting rich visual and spatial information into text, forcing models to implicitly infer complex visual‑topological relationships or limiting their global action capabilities. To bridge this gap, we propose TagaVLM (Topology‑Aware Global Action reasoning), an end‑to‑end framework that explicitly injects topological structures into the VLM backbone. To introduce topological edge information, Spatial Topology Aware Residual Attention (STAR‑Att) directly integrates it into the VLM's self‑attention mechanism, enabling intrinsic spatial reasoning while preserving pretrained knowledge. To enhance topological node information, an Interleaved Navigation Prompt strengthens node‑level visual‑text alignment. Finally, with the embedded topological graph, the model is capable of global action reasoning, allowing for robust path correction. On the R2R benchmark, TagaVLM achieves state‑of‑the‑art performance among large‑model‑based methods, with a Success Rate (SR) of 51.09% and SPL of 47.18 in unseen environments, outperforming prior work by 3.39% in SR and 9.08 in SPL. This demonstrates that, for embodied spatial reasoning, targeted enhancements on smaller open‑source VLMs can be more effective than brute‑force model scaling. The code can be found on our project page: https://apex‑bjut.github.io/Taga‑VLM
Authors:Julio Silva-Rodríguez, Ender Konukoglu
Abstract:
Vision‑language models (VLMs) pre‑trained on large, heterogeneous data sources are becoming increasingly popular, providing rich multi‑modal embeddings that enable efficient transfer to new tasks. A particularly relevant application is few‑shot adaptation, where only a handful of annotated examples are available to adapt the model through multi‑modal linear probes. In medical imaging, specialized VLMs have shown promising performance in zero‑ and few‑shot image classification, which is valuable for mitigating the high cost of expert annotations. However, challenges remain in extremely low‑shot regimes: the inherent class imbalances in medical tasks often lead to underrepresented categories, penalizing overall model performance. To address this limitation, we propose leveraging unlabeled data by introducing an efficient semi‑supervised solver that propagates text‑informed pseudo‑labels during few‑shot adaptation. The proposed method enables lower‑budget annotation pipelines for adapting VLMs, reducing labeling effort by >50% in low‑shot regimes.
Authors:Hao Zhang, Yiqun Wang, Qinran Lin, Runze Fan, Yong Li
Abstract:
Despite the growing interest in open‑vocabulary object detection in recent years, most existing methods rely heavily on manually curated fine‑grained training datasets as well as resource‑intensive layer‑wise cross‑modal feature extraction. In this paper, we propose HDINO, a concise yet efficient open‑vocabulary object detector that eliminates the dependence on these components. Specifically, we propose a two‑stage training strategy built upon the transformer‑based DINO model. In the first stage, noisy samples are treated as additional positive object instances to construct a One‑to‑Many Semantic Alignment Mechanism(O2M) between the visual and textual modalities, thereby facilitating semantic alignment. A Difficulty Weighted Classification Loss (DWCL) is also designed based on initial detection difficulty to mine hard examples and further improve model performance. In the second stage, a lightweight feature fusion module is applied to the aligned representations to enhance sensitivity to linguistic semantics. Under the Swin Transformer‑T setting, HDINO‑T achieves 49.2 mAP on COCO using 2.2M training images from two publicly available detection datasets, without any manual data curation and the use of grounding data, surpassing Grounding DINO‑T and T‑Rex2 by 0.8 mAP and 2.8 mAP, respectively, which are trained on 5.4M and 6.5M images. After fine‑tuning on COCO, HDINO‑T and HDINO‑L further achieve 56.4 mAP and 59.2 mAP, highlighting the effectiveness and scalability of our approach. Code and models are available at https://github.com/HaoZ416/HDINO.
Authors:Hao Ai, Wenjie Chang, Jianbo Jiao, Ales Leonardis, Ofek Eyal
Abstract:
Articulated objects are ubiquitous in daily life. Our goal is to achieve a high‑quality reconstruction, segmentation of independent moving parts, and analysis of articulation. Recent methods analyse two different articulation states and perform per‑point part segmentation, optimising per‑part articulation using cross‑state correspondences, given a priori knowledge of the number of parts. Such assumptions greatly limit their applications and performance. Their robustness is reduced when objects cannot be clearly visible in both states. To address these issues, in this paper, we present a new framework, Articulation in Motion (AiM). We infer part‑level decomposition, articulation kinematics, and reconstruct an interactive 3D digital replica from a user‑object interaction video and a start‑state scan. We propose a dual‑Gaussian scene representation that is learned from an initial 3DGS scan of the object and a video that shows the movement of separate parts. It uses motion cues to segment the object into parts and assign articulation joints. Subsequently, a robust, sequential RANSAC is employed to achieve part mobility analysis without any part‑level structural priors, which clusters moving primitives into rigid parts and estimates kinematics while automatically determining the number of parts. The proposed approach separates the object into parts, each represented as a 3D Gaussian set, enabling high‑quality rendering. Our approach yields higher quality part segmentation than previous methods, without prior knowledge. Extensive experimental analysis on both simple and complex objects validates the effectiveness and strong generalisation ability of our approach. Project page: https://haoai‑1997.github.io/AiM/.
Authors:Xinjie Zhu, Zijing Zhao, Hui Jin, Qingxiao Guo, Yilong Ma, Yunhao Wang, Xiaobing Guo, Weifeng Zhang
Abstract:
Artificial Intelligence Generated Content (AIGC), particularly video generation with diffusion models, has been advanced rapidly. Invisible watermarking is a key technology for protecting AI‑generated videos and tracing harmful content, and thus plays a crucial role in AI safety. Beyond post‑processing watermarks which inevitably degrade video quality, recent studies have proposed distortion‑free in‑generation watermarking for video diffusion models. However, existing in‑generation approaches are non‑blind: they require maintaining all the message‑key pairs and performing template‑based matching during extraction, which incurs prohibitive computational costs at scale. Moreover, when applied to modern video diffusion models with causal 3D Variational Autoencoders (VAEs), their robustness against temporal disturbance becomes extremely weak. To overcome these challenges, we propose SIGMark, a Scalable In‑Generation watermarking framework with blind extraction for video diffusion. To achieve blind‑extraction, we propose to generate watermarked initial noise using a Global set of Frame‑wise PseudoRandom Coding keys (GF‑PRC), reducing the cost of storing large‑scale information while preserving noise distribution and diversity for distortion‑free watermarking. To enhance robustness, we further design a Segment Group‑Ordering module (SGO) tailored to causal 3D VAEs, ensuring robust watermark inversion during extraction under temporal disturbance. Comprehensive experiments on modern diffusion models show that SIGMark achieves very high bit‑accuracy during extraction under both temporal and spatial disturbances with minimal overhead, demonstrating its scalability and robustness. Our project is available at https://jeremyzhao1998.github.io/SIGMark‑release/.
Authors:Jialiang Zhang, Junlong Tong, Junyan Lin, Hao Wu, Yirong Sun, Yunpu Ma, Xiaoyu Shen
Abstract:
Large Vision Language Models (LVLMs) exhibit strong Chain‑of‑Thought (CoT) capabilities, yet most existing paradigms assume full‑video availability before inference, a batch‑style process misaligned with real‑world video streams where information arrives sequentially. Motivated by the streaming nature of video data, we investigate two streaming reasoning paradigms for LVLMs. The first, an interleaved paradigm, alternates between receiving frames and producing partial reasoning but remains constrained by strictly ordered cache updates. To better match streaming inputs, we propose Think‑as‑You‑See (TaYS), a unified framework enabling true concurrent reasoning. TaYS integrates parallelized CoT generation, stream‑constrained training, and stream‑parallel inference. It further employs temporally aligned reasoning units, streaming attention masks and positional encodings, and a dual KV‑cache that decouples visual encoding from textual reasoning. We evaluate all paradigms on the Qwen2.5‑VL family across representative video CoT tasks, including event dynamics analysis, causal reasoning, and thematic understanding. Experiments show that TaYS consistently outperforms both batch and interleaved baselines, improving reasoning performance while substantially reducing time‑to‑first‑token (TTFT) and overall reasoning delay. These results demonstrate the effectiveness of data‑aligned streaming reasoning in enabling efficient and responsive video understanding for LVLMs. We release our code at https://github.com/EIT‑NLP/StreamingLLM/tree/main/TaYS
Authors:Huanlei Guo, Hongxin Wei, Bingyi Jing
Abstract:
Recent text‑to‑image (T2I) diffusion and flow‑matching models can produce highly realistic images from natural language prompts. In practical scenarios, T2I systems are often run in a ``generate‑‑then‑‑select'' mode: many seeds are sampled and only a few images are kept for use. However, this pipeline is highly resource‑intensive since each candidate requires tens to hundreds of denoising steps, and evaluation metrics such as CLIPScore and ImageReward are post‑hoc. In this work, we address this inefficiency by introducing Probe‑Select, a plug‑in module that enables efficient evaluation of image quality within the generation process. We observe that certain intermediate denoiser activations, even at early timesteps, encode a stable coarse structure, object layout and spatial arrangement‑‑that strongly correlates with final image fidelity. Probe‑Select exploits this property by predicting final quality scores directly from early activations, allowing unpromising seeds to be terminated early. Across diffusion and flow‑matching backbones, our experiments show that early evaluation at only 20% of the trajectory accurately ranks candidate seeds and enables selective continuation. This strategy reduces sampling cost by over 60% while improving the quality of the retained images, demonstrating that early structural signals can effectively guide selective generation without altering the underlying generative model. Code is available at https://github.com/Guhuary/ProbeSelect.
Authors:Yi Liu, Jing Zhang, Di Wang, Xiaoyu Tian, Haonan Guo, Bo Du
Abstract:
Multimodal large language models (MLLMs) suffer from pronounced hallucinations in remote sensing visual question‑answering (RS‑VQA), primarily caused by visual grounding failures in large‑scale scenes or misinterpretation of fine‑grained small targets. To systematically analyze these issues, we introduce RSHBench, a protocol‑based benchmark for fine‑grained diagnosis of factual and logical hallucinations. To mitigate grounding‑induced factual hallucinations, we further propose Relative Attention‑Driven Actively Reasoning (RADAR), a training‑free inference method that leverages intrinsic attention in MLLMs to guide progressive localization and fine‑grained local reasoning at test time. Extensive experiments across diverse MLLMs demonstrate that RADAR consistently improves RS‑VQA performance and reduces both factual and logical hallucinations. Code and data will be publicly available at: https://github.com/MiliLab/RADAR
Authors:Hongbo Zheng, Afshin Bozorgpour, Dorit Merhof, Minjia Zhang
Abstract:
Medical image segmentation requires models that preserve fine anatomical boundaries while remaining practical for clinical deployment. Transformers capture long‑range dependencies but incur quadratic attention cost, whereas CNNs are efficient but less effective at global reasoning. Linear attention offers \(\mathcalO(N)\) scaling, but often produces diffuse feature aggregation that weakens boundary‑sensitive prediction. We introduce a gated differential linear‑attention mixer for medical image segmentation. Its global path, Gated Differential Linear Attention (GDLA), performs differential subtraction between two kernelized attention branches over complementary query/key subspaces to suppress redundant responses, and employs a data‑dependent gate for token refinement. A parallel local token‑mixing branch with depthwise convolution strengthens neighborhood interactions for better refinement, and the two branches are fused while preserving \(\mathcalO(N)\) complexity. When instantiated in a pretrained Pyramid Vision Transformer (PVT)‑based encoder‑‑decoder model, \name achieves state‑of‑the‑art results on the evaluated 2D medical segmentation benchmarks spanning CT, MRI, ultrasound, and dermoscopy, with a favorable accuracy‑‑efficiency trade‑off over closely related baselines. The code is publicly available at \hrefhttps://github.com/xmindflow/gdlahttps://github.com/xmindflow/gdla.
Authors:Hongying Zhang, ShuaiShuai Ma
Abstract:
Cross‑view geo‑localization (CVGL) aims to establish spatial correspondences between images captured from significantly different viewpoints and constitutes a fundamental technique for visual localization in GNSS‑denied environments. Nevertheless, CVGL remains challenging due to severe geometric asymmetry, texture inconsistency across imaging domains, and the progressive degradation of discriminative local information. Existing methods predominantly rely on spatial domain feature alignment, which is inherently sensitive to large scale viewpoint variations and local disturbances. To alleviate these limitations, this paper proposes the Spatial and Frequency Domain Enhancement Network (SFDE), which leverages complementary representations from spatial and frequency domains. SFDE adopts a three branch parallel architecture to model global semantic context, local geometric structure, and statistical stability in the frequency domain, respectively, thereby characterizing consistency across domains from the perspectives of scene topology, multiscale structural patterns, and frequency invariance. The resulting complementary features are jointly optimized in a unified embedding space via progressive enhancement and coupled constraints, enabling the learning of cross‑view representations with consistency across multiple granularities. Comprehensive experiments show that SFDE achieves competitive performance and in many cases even surpasses state‑of‑the‑art methods, while maintaining a lightweight and computationally efficient design. Our code is available at https://github.com/Mashuaishuai669/SFDE
Authors:Lingshun Kong, Jiawei Zhang, Zhengpeng Duan, Xiaohe Wu, Yueqi Yang, Xiaotao Wang, Dongqing Zou, Lei Lei, Jinshan Pan
Abstract:
All‑in‑one image restoration is challenging because different degradation types, such as haze, blur, noise, and low‑light, impose diverse requirements on restoration strategies, making it difficult for a single model to handle them effectively. In this paper, we propose a unified image restoration framework that integrates a dual‑level Mixture‑of‑Experts (MoE) architecture with a pretrained diffusion model. The framework operates at two levels: the Inter‑MoE layer adaptively combines expert groups to handle major degradation types, while the Intra‑MoE layer further selects specialized sub‑experts to address fine‑grained variations within each type. This design enables the model to achieve coarse‑grained adaptation across diverse degradation categories while performing fine‑grained modulation for specific intra‑class variations, ensuring both high specialization in handling complex, real‑world corruptions. Extensive experiments demonstrate that the proposed method performs favorably against the state‑of‑the‑art approaches on multiple image restoration task.
Authors:Aro Kim, Myeongjin Jang, Chaewon Moon, Youngjin Shin, Jinwoo Jeong, Sang-hyo Park
Abstract:
Diffusion‑based approaches have recently driven remarkable progress in real‑world image super‑resolution (SR). However, existing methods still struggle to simultaneously preserve fine details and ensure high‑fidelity reconstruction, often resulting in suboptimal visual quality. In this paper, we propose FiDeSR, a high‑fidelity and detail‑preserving one‑step diffusion super‑resolution framework. During training, we introduce a detail‑aware weighting strategy that adaptively emphasizes regions where the model exhibits higher prediction errors. During inference, low‑ and high‑frequency adaptive enhancers further refine the reconstruction without requiring model retraining, enabling flexible enhancement control. To further improve the reconstruction accuracy, FiDeSR incorporates a residual‑in‑residual noise refinement, which corrects prediction errors in the diffusion noise and enhances fine detail recovery. FiDeSR achieves superior real‑world SR performance compared to existing diffusion‑based methods, producing outputs with both high perceptual quality and faithful content restoration. The source code will be released at: https://github.com/Ar0Kim/FiDeSR.
Authors:Seunguk Do, Minwoo Huh, Joonghyuk Shin, Jaesik Park
Abstract:
Single‑view 3D human reconstruction has achieved remarkable progress through the adoption of multi‑view diffusion models, yet the recovered 3D humans often exhibit unnatural poses. This phenomenon becomes pronounced when reconstructing 3D humans with dynamic or challenging poses, which we attribute to the limited scale of available 3D human datasets with diverse poses. To address this limitation, we introduce DrPose, Direct Reward fine‑tuning algorithm on Poses, which enables post‑training of a multi‑view diffusion model on diverse poses without requiring expensive 3D human assets. DrPose trains a model using only human poses paired with single‑view images, employing a direct reward fine‑tuning to maximize PoseScore, which is our proposed differentiable reward that quantifies consistency between a generated multi‑view latent image and a ground‑truth human pose. This optimization is conducted on DrPose15K, a novel dataset that was constructed from an existing human motion dataset and a pose‑conditioned video generative model. Constructed from abundant human pose sequence data, DrPose15K exhibits a broader pose distribution compared to existing 3D human datasets. We validate our approach through evaluation on conventional benchmark datasets, in‑the‑wild images, and a newly constructed benchmark, with a particular focus on assessing performance on challenging human poses. Our results demonstrate consistent qualitative and quantitative improvements across all benchmarks. Project page: https://seunguk‑do.github.io/drpose.
Authors:Jiahao Lu, Jiayi Xu, Wenbo Hu, Ruijie Zhu, Chengfeng Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
Abstract:
Estimating the 3D trajectory of every pixel from a monocular video is crucial and promising for a comprehensive understanding of the 3D dynamics of videos. Recent monocular 3D tracking works demonstrate impressive performance, but are limited to either tracking sparse points on the first frame or a slow optimization‑based framework for dense tracking. In this paper, we propose a feedforward model, called Track4World, enabling an efficient holistic 3D tracking of every pixel in the world‑centric coordinate system. Built on the global 3D scene representation encoded by a VGGT‑style ViT, Track4World applies a novel 3D correlation scheme to simultaneously estimate the pixel‑wise 2D and 3D dense flow between arbitrary frame pairs. The estimated scene flow, along with the reconstructed 3D geometry, enables subsequent efficient 3D tracking of every pixel of this video. Extensive experiments on multiple benchmarks demonstrate that our approach consistently outperforms existing methods in 2D/3D flow estimation and 3D tracking, highlighting its robustness and scalability for real‑world 4D reconstruction tasks.
Authors:Huichun Liu, Xiaosong Li, Zhuangfan Huang, Tao Ye, Yang Liu, Haishu Tan
Abstract:
Multimodal Image Fusion (MMIF) integrates complementary information from various modalities to produce clearer and more informative fused images. MMIF under adverse weather is particularly crucial in autonomous driving and UAV monitoring applications. However, existing adverse weather fusion methods generally only tackle single types of degradation such as haze, rain, or snow, and fail when multiple degradations coexist (e.g., haze+rain, rain+snow). To address this challenge, we propose Compound Adverse Weather Mamba (CAWM‑Mamba), the first end‑to‑end framework that jointly performs image fusion and compound weather restoration with unified shared weights. Our network contains three key components: (1) a Weather‑Aware Preprocess Module (WAPM) to enhance degraded visible features and extracts global weather embeddings; (2) a Cross‑modal Feature Interaction Module (CFIM) to facilitate the alignment of heterogeneous modalities and exchange of complementary features across modalities; and (3) a Wavelet Space State Block (WSSB) that leverages wavelet‑domain decomposition to decouple multi‑frequency degradations. WSSB includes Freq‑SSM, a module that models anisotropic high‑frequency degradation without redundancy, and a unified degradation representation mechanism to further improve generalization across complex compound weather conditions. Extensive experiments on the AWMM‑100K benchmark and three standard fusion datasets demonstrate that CAWM‑Mamba consistently outperforms state‑of‑the‑art methods in both compound and single‑weather scenarios. In addition, our fusion results excel in downstream tasks covering semantic segmentation and object detection, confirming the practical value in real‑world adverse weather perception. The source code will be available at https://github.com/Feecuin/CAWM‑Mamba.
Authors:Maoyuan Shao, Yutong Gao, Xinyang Huang, Chuang Zhu, Lijuan Sun, Guoshun Nan
Abstract:
Vision‑language models like CLIP have achieved remarkable progress in cross‑modal representation learning, yet suffer from systematic misclassifications among visually and semantically similar categories. We observe that such confusion patterns are not random but persistently occur between specific category pairs, revealing the model's intrinsic bias and limited fine‑grained discriminative ability. To address this, we propose CAPT, a Confusion‑Aware Prompt Tuning framework that enables models to learn from their own misalignment. Specifically, we construct a Confusion Bank to explicitly model stable confusion relationships across categories and misclassified samples. On this basis, we introduce a Semantic Confusion Miner (SEM) to capture global inter‑class confusion through semantic difference and commonality prompts, and a Sample Confusion Miner (SAM) to retrieve representative misclassified instances from the bank and capture sample‑level cues through a Diff‑Manner Adapter that integrates global and local contexts. To further unify confusion information across different granularities, a Multi‑Granularity Difference Expert (MGDE) module is designed to jointly leverage semantic‑ and sample‑level experts for more robust confusion‑aware reasoning. Extensive experiments on 11 benchmark datasets demonstrate that our method significantly reduces confusion‑induced errors while enhancing the discriminability and generalization of both base and novel classes, successfully resolving 50.72 percent of confusable sample pairs. Code will be released at https://github.com/greatest‑gourmet/CAPT.
Authors:Zhiyu Pan, Yizheng Wu, Jiashen Hua, Junyi Feng, Shaotian Yan, Bing Deng, Zhiguo Cao, Jieping Ye
Abstract:
Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self‑improving techniques that refine reasoning paths for subsequent finetuning. However, extending these language‑based self‑improving approaches to vision language models (VLMs) presents a unique challenge:~visual hallucinations in reasoning paths cannot be effectively verified or rectified. Our solution starts with a key observation about visual contrast: when presented with a contrastive VQA pair, i.e., two visually similar images with synonymous questions, VLMs identify relevant visual cues more precisely. Motivated by this observation, we propose Visual Contrastive Self‑Taught Reasoner (VC‑STaR), a novel self‑improving framework that leverages visual contrast to mitigate hallucinations in model‑generated rationales. We collect a diverse suite of VQA datasets, curate contrastive pairs according to multi‑modal similarity, and generate rationales using VC‑STaR. Consequently, we obtain a new visual reasoning dataset, VisCoR‑55K, which is then used to boost the reasoning capability of various VLMs through supervised finetuning. Extensive experiments show that VC‑STaR not only outperforms existing self‑improving approaches but also surpasses models finetuned on the SoTA visual reasoning datasets, demonstrating that the inherent contrastive ability of VLMs can bootstrap their own visual reasoning. Project at: https://github.com/zhiyupan42/VC‑STaR.
Authors:Chonghua Lv, Dong Zhao, Shuang Wang, Dou Quan, Ning Huyan, Nicu Sebe, Zhun Zhong
Abstract:
Knowledge distillation (KD) has been widely applied in semantic segmentation to compress large models, but conventional approaches primarily preserve in‑domain accuracy while neglecting out‑of‑domain generalization, which is essential under distribution shifts. This limitation becomes more severe with the emergence of vision foundation models (VFMs): although VFMs exhibit strong robustness on unseen data, distilling them with conventional KD often compromises this ability. We propose Generalizable Knowledge Distillation (GKD), a multi‑stage framework that explicitly enhances generalization. GKD decouples representation learning from task learning. In the first stage, the student acquires domain‑agnostic representations through selective feature distillation, and in the second stage, these representations are frozen for task adaptation, thereby mitigating overfitting to visible domains. To further support transfer, we introduce a query‑based soft distillation mechanism, where student features act as queries to teacher representations to selectively retrieve transferable spatial knowledge from VFMs. Extensive experiments on five domain generalization benchmarks demonstrate that GKD consistently outperforms existing KD methods, achieving average gains of +1.9% in foundation‑to‑foundation (F2F) and +10.6% in foundation‑to‑local (F2L) distillation. The code will be available at https://github.com/Younger‑hua/GKD.
Authors:Xuejin Luo, Shiquan Sun, Runshi Zhang, Ruizhi Zhang, Junchen Wang
Abstract:
During surgery, scrub nurses are required to frequently deliver surgical instruments to surgeons, which can lead to physical fatigue and decreased focus. Robotic scrub nurses provide a promising solution that can replace repetitive tasks and enhance efficiency. Existing research on robotic scrub nurses relies on predefined paths for instrument delivery, which limits their generalizability and poses safety risks in dynamic environments. To address these challenges, we present a collision‑free dual‑arm surgical assistive robot capable of performing instrument delivery. A vision‑language model is utilized to automatically generate the robot's grasping and delivery trajectories in a zero‑shot manner based on surgeons' instructions. A real‑time obstacle minimum distance perception method is proposed and integrated into a unified quadratic programming framework. This framework ensures reactive obstacle avoidance and self‑collision prevention during the dual‑arm robot's autonomous movement in dynamic environments. Extensive experimental validations demonstrate that the proposed robotic system achieves an 83.33% success rate in surgical instrument delivery while maintaining smooth, collision‑free movement throughout all trials. The project page and source code are available at https://give‑me‑scissors.github.io/.
Authors:Kang Yang, Peng Wang, Lantao Li, Tianci Bu, Chen Sun, Deying Li, Yongcai Wang
Abstract:
Multi‑modal collaborative perception calls for great attention to enhancing the safety of autonomous driving. However, current multi‑modal approaches remain a ``local fusion to communication'' sequence, which fuses multi‑modal data locally and needs high bandwidth to transmit an individual's feature data before collaborative fusion. EIMC innovatively proposes an early collaborative paradigm. It injects lightweight collaborative voxels, transmitted by neighbor agents, into the ego's local modality‑fusion step, yielding compact yet informative 3D collaborative priors that tighten cross‑modal alignment. Next, a heatmap‑driven consensus protocol identifies exactly where cooperation is needed by computing per‑pixel confidence heatmaps. Only the Top‑K instance vectors located in these low‑confidence, high‑discrepancy regions are queried from peers, then fused via cross‑attention for completion. Afterwards, we apply a refinement fusion that involves collecting the top‑K most confident instances from each agent and enhancing their features using self‑attention. The above instance‑centric messaging reduces redundancy while guaranteeing that critical occluded objects are recovered. Evaluated on OPV2V and DAIR‑V2X, EIMC attains 73.01% AP@0.5 while reducing byte bandwidth usage by 87.98% compared with the best published multi‑modal collaborative detector. Code publicly released at https://github.com/sidiangongyuan/EIMC.
Authors:David Pujol-Perich, Albert Clapés, Dima Damen, Sergio Escalera, Michael Wray
Abstract:
In this work, we investigate the degradation of existing VMR methods, particularly of DETR architectures, when trained on caption‑based queries but evaluated on search queries. For this, we introduce three benchmarks by modifying the textual queries in three public VMR datasets ‑‑ i.e., HD‑EPIC, YouCook2 and ActivityNet‑Captions. Our analysis reveals two key generalization challenges: (i) A language gap, arising from the linguistic under‑specification of search queries, and (ii) a multi‑moment gap, caused by the shift from single‑moment to multi‑moment queries. We also identify a critical issue in these architectures ‑‑ an active decoder‑query collapse ‑‑ as a primary cause of the poor generalization to multi‑moment instances. We mitigate this issue with architectural modifications that effectively increase the number of active decoder queries. Extensive experiments demonstrate that our approach improves performance on search queries by up to 14.82% mAP_m, and up to 21.83% mAP_m on multi‑moment search queries. The code, models and data are available in the project webpage: https://davidpujol.github.io/beyond‑vmr/
Authors:Leo Kaixuan Cheng, Abdus Shaikh, Ruofan Liang, Zhijie Wu, Yushi Guan, Nandita Vijaykumar
Abstract:
Recent advancements in neural visual geometry, including transformer‑based models such as VGGT and Pi3, have achieved impressive accuracy on 3D reconstruction tasks. However, their reliance on full attention makes them fundamentally limited by GPU memory capacity, preventing them from scaling to large, unordered image collections. We introduce MERG3R, a training‑free divide‑and‑conquer framework that enables geometric foundation models to operate far beyond their native memory limits. MERG3R first reorders and partitions unordered images into overlapping, geometrically diverse subsets that can be reconstructed independently. It then merges the resulting local reconstructions through an efficient global alignment and confidence‑weighted bundle adjustment procedure, producing a globally consistent 3D model. Our framework is model‑agnostic and can be paired with existing neural geometry models. Across large‑scale datasets, including 7‑Scenes, NRGBD, Tanks & Temples, and Cambridge Landmarks, MERG3R consistently improves reconstruction accuracy, memory efficiency, and scalability, enabling high‑quality reconstruction when the dataset exceeds memory capacity limits.
Authors:Lei Yao, Yong Chen, Yuejiao Su, Yi Wang, Moyun Liu, Lap-Pui Chau
Abstract:
Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. Inspired by this principle, we advocate for a novel framework that leverages emerging multimodal large language models (MLLMs) for interaction intention‑driven 3D affordance grounding, namely HAMMER. Instead of generating explicit object attribute descriptions or relying on off‑the‑shelf 2D segmenters, we alternatively aggregate the interaction intention depicted in the image into a contact‑aware embedding and guide the model to infer textual affordance labels, ensuring it thoroughly excavates object semantics and contextual cues. We further devise a hierarchical cross‑modal integration mechanism to fully exploit the complementary information from the MLLM for 3D representation refinement and introduce a multi‑granular geometry lifting module that infuses spatial characteristics into the extracted intention embedding, thus facilitating accurate 3D affordance localization. Extensive experiments on public datasets and our newly constructed corrupted benchmark demonstrate the superiority and robustness of HAMMER compared to existing approaches. All code and weights are publicly available.
Authors:Nikhileswara Rao Sulake
Abstract:
Long‑tailed class distributions pose a significant challenge for multi‑label chest X‑ray (CXR) classification, where rare but clinically important findings are severely underrepresented. In this work, we present a systematic empirical evaluation of loss functions, CNN backbone architectures and post‑training strategies on the CXR‑LT 2026 benchmark, comprising approximately 143K images with 30 disease labels from PadChest. Our experiments demonstrate that LDAM with deferred re‑weighting (LDAM‑DRW) consistently outperforms standard BCE and asymmetric losses for rare class recognition. Amongst the architectures evaluated, ConvNeXt‑Large achieves the best single‑model performance with 0.5220 mAP and 0.3765 F1 on our development set, whilst classifier re‑training and test‑time augmentation further improve ranking metrics. On the official test leaderboard, our submission achieved 0.3950 mAP, ranking 5th amongst all 68 participating teams with total of 1528 submissions. We provide a candid analysis of the development‑to‑test performance gap and discuss practical insights for handling class imbalance in clinical imaging settings. Code is available at https://github.com/Nikhil‑Rao20/Long_Tail.
Authors:Yaoteng Zhang, Zhou Qing, Junyu Gao, Qi Wang
Abstract:
Incremental Object Detection (IOD) aims to continuously learn new object categories without forgetting previously learned ones. Recently, prompt‑based methods have gained popularity for their replay‑free design and parameter efficiency. However, due to prompt coupling and prompt drift, these methods often suffer from prompt degradation during continual adaptation. To address these issues, we propose a novel prompt‑decoupled framework called PDP. PDP innovatively designs a dual‑pool prompt decoupling paradigm, which consists of a shared pool used to capture task‑general knowledge for forward transfer, and a private pool used to learn task‑specific discriminative features. This paradigm explicitly separates task‑general and task‑specific prompts, preventing interference between prompts and mitigating prompt coupling. In addition, to counteract prompt drift resulting from inconsistent supervision where old foreground objects are treated as background in subsequent tasks, PDP introduces a Prototypical Pseudo‑Label Generation (PPG) module. PPG can dynamically update the class prototype space during training and use the class prototypes to further filter valuable pseudo‑labels, maintaining supervisory signal consistency throughout the incremental process. PDP achieves state‑of‑the‑art performance on MS‑COCO (with a 9.2% AP improvement) and PASCAL VOC (with a 3.3% AP improvement) benchmarks, highlighting its potential in balancing stability and plasticity. The code and dataset are released at: https://github.com/zyt95579/PDP\_IOD/tree/main
Authors:Yichen Liu, Donghao Zhou, Jie Wang, Xin Gao, Guisheng Liu, Jiatong Li, Quanwei Zhang, Qiang Lyu, Lanqing Guo, Shilei Wen, Weiqiang Wang, Pheng-Ann Heng
Abstract:
Human‑product images, which showcase the integration of humans and products, play a vital role in advertising, e‑commerce, and digital marketing. The essential challenge of generating such images lies in ensuring the high‑fidelity preservation of product details. Among existing paradigms, reference‑based inpainting offers a targeted solution by leveraging product reference images to guide the inpainting process. However, limitations remain in three key aspects: the lack of diverse large‑scale training data, the struggle of current models to focus on product detail preservation, and the inability of coarse supervision for achieving precise guidance. To address these issues, we propose HiFi‑Inpaint, a novel high‑fidelity reference‑based inpainting framework tailored for generating human‑product images. HiFi‑Inpaint introduces Shared Enhancement Attention (SEA) to refine fine‑grained product features and Detail‑Aware Loss (DAL) to enforce precise pixel‑level supervision using high‑frequency maps. Additionally, we construct a new dataset, HP‑Image‑40K, with samples curated from self‑synthesis data and processed with automatic filtering. Experimental results show that HiFi‑Inpaint achieves state‑of‑the‑art performance, delivering detail‑preserving human‑product images.
Authors:Moru Liu, Hao Dong, Olga Fink, Mario Trapp
Abstract:
The deployment of multimodal models in high‑stakes domains, such as self‑driving vehicles and medical diagnostics, demands not only strong predictive performance but also reliable mechanisms for detecting failures. In this work, we address the largely unexplored problem of failure detection in multimodal contexts. We propose Adaptive Confidence Regularization (ACR), a novel framework specifically designed to detect multimodal failures. Our approach is driven by a key observation: in most failure cases, the confidence of the multimodal prediction is significantly lower than that of at least one unimodal branch, a phenomenon we term confidence degradation. To mitigate this, we introduce an Adaptive Confidence Loss that penalizes such degradations during training. In addition, we propose Multimodal Feature Swapping, a novel outlier synthesis technique that generates challenging, failure‑aware training examples. By training with these synthetic failures, ACR learns to more effectively recognize and reject uncertain predictions, thereby improving overall reliability. Extensive experiments across four datasets, three modalities, and multiple evaluation settings demonstrate that ACR achieves consistent and robust gains. The source code will be available at https://github.com/mona4399/ACR.
Authors:Yiqi Lin, Guoqiang Liang, Ziyun Zeng, Zechen Bai, Yanzhe Chen, Mike Zheng Shou
Abstract:
Instruction‑based video editing has witnessed rapid progress, yet current methods often struggle with precise visual control, as natural language is inherently limited in describing complex visual nuances. Although reference‑guided editing offers a robust solution, its potential is currently bottlenecked by the scarcity of high‑quality paired training data. To bridge this gap, we introduce a scalable data generation pipeline that transforms existing video editing pairs into high‑fidelity training quadruplets, leveraging image generative models to create synthesized reference scaffolds. Using this pipeline, we construct RefVIE, a large‑scale dataset tailored for instruction‑reference‑following tasks, and establish RefVIE‑Bench for comprehensive evaluation. Furthermore, we propose a unified editing architecture, Kiwi‑Edit, that synergizes learnable queries and latent visual features for reference semantic guidance. Our model achieves significant gains in instruction following and reference fidelity via a progressive multi‑stage training curriculum. Extensive experiments demonstrate that our data and architecture establish a new state‑of‑the‑art in controllable video editing. All datasets, models, and code is released at https://github.com/showlab/Kiwi‑Edit.
Authors:Yiying Yang, Wei Cheng, Sijin Chen, Honghao Fu, Xianfang Zeng, Yujun Cai, Gang Yu, Xingjun Ma
Abstract:
OmniLottie is a versatile framework that generates high quality vector animations from multi‑modal instructions. For flexible motion and visual content control, we focus on Lottie, a light weight JSON formatting for both shapes and animation behaviors representation. However, the raw Lottie JSON files contain extensive invariant structural metadata and formatting tokens, posing significant challenges for learning vector animation generation. Therefore, we introduce a well designed Lottie tokenizer that transforms JSON files into structured sequences of commands and parameters representing shapes, animation functions and control parameters. Such tokenizer enables us to build OmniLottie upon pretrained vision language models to follow multi‑modal interleaved instructions and generate high quality vector animations. To further advance research in vector animation generation, we curate MMLottie‑2M, a large scale dataset of professionally designed vector animations paired with textual and visual annotations. With extensive experiments, we validate that OmniLottie can produce vivid and semantically aligned vector animations that adhere closely to multi modal human instructions.
Authors:Chong Xia, Fangfu Liu, Yule Wang, Yize Pang, Yueqi Duan
Abstract:
Recent advances in generalizable 3D Gaussian Splatting (3DGS) have enabled rapid 3D scene reconstruction within seconds, eliminating the need for per‑scene optimization. However, existing methods primarily follow an offline reconstruction paradigm, lacking the capacity for continuous reconstruction, which limits their applicability to online scenarios such as robotics and VR/AR. In this paper, we introduce OnlineX, a feed‑forward framework that reconstructs both 3D visual appearance and language fields in an online manner using only streaming images. A key challenge in online formulation is the cumulative drift issue, which is rooted in the fundamental conflict between two opposing roles of the memory state: an active role that constantly refreshes to capture high‑frequency local geometry, and a stable role that conservatively accumulates and preserves the long‑term global structure. To address this, we introduce a decoupled active‑to‑stable state evolution paradigm. Our framework decouples the memory state into a dedicated active state and a persistent stable state, and then cohesively fuses the information from the former into the latter to achieve both fidelity and stability. Moreover, we jointly model visual appearance and language fields and incorporate an implicit Gaussian fusion module to enhance reconstruction quality. Experiments on mainstream datasets demonstrate that our method consistently outperforms prior work in novel view synthesis and semantic understanding, showcasing robust performance across input sequences of varying lengths with real‑time inference speed.
Authors:Chong Xia, Kai Zhu, Zizhuo Wang, Fangfu Liu, Zhizheng Zhang, Yueqi Duan
Abstract:
Compositional scene reconstruction seeks to create object‑centric representations rather than holistic scenes from real‑world videos, which is natively applicable for simulation and interaction. Conventional compositional reconstruction approaches primarily emphasize on visual appearance and show limited generalization ability to real‑world scenarios. In this paper, we propose SimRecon, a framework that realizes a "Perception‑Generation‑Simulation" pipeline towards cluttered scene reconstruction, which first conducts scene‑level semantic reconstruction from video input, then performs single‑object generation, and finally assembles these assets in the simulator. However, naively combining these three stages leads to visual infidelity of generated assets and physical implausibility of the final scene, a problem particularly severe for complex scenes. Thus, we further propose two bridging modules between the three stages to address this problem. To be specific, for the transition from Perception to Generation, critical for visual fidelity, we introduce Active Viewpoint Optimization, which actively searches in 3D space to acquire optimal projected images as conditions for single‑object completion. Moreover, for the transition from Generation to Simulation, essential for physical plausibility, we propose a Scene Graph Synthesizer, which guides the construction from scratch in 3D simulators, mirroring the native, constructive principle of the real world. Extensive experiments on the ScanNet dataset validate our method's superior performance over previous state‑of‑the‑art approaches.
Authors:Jiahao Huang, Fengyan Lin, Xuechao Yang, Chen Feng, Kexin Zhu, Xu Yang, Zhide Chen
Abstract:
The development of affective multimodal language models (MLMs) has long been constrained by a gap between low‑level perception and high‑level interaction, leading to fragmented affective capabilities and limited generalization. To bridge this gap, we propose a cognitively inspired three‑level hierarchy that organizes affective tasks according to their cognitive depth‑perception, understanding, and interaction‑and provides a unified conceptual foundation for advancing affective modeling. Guided by this hierarchy, we introduce Nano‑EmoX, a small‑scale multitask MLM, and P2E (Perception‑to‑Empathy), a curriculum‑based training framework. Nano‑EmoX integrates a suite of omni‑modal encoders, including an enhanced facial encoder and a fusion encoder, to capture key multimodal affective cues and improve cross‑task transferability. The outputs are projected into a unified language space via heterogeneous adapters, empowering a lightweight language model to tackle diverse affective tasks. Concurrently, P2E progressively cultivates emotional intelligence by aligning rapid perception with chain‑of‑thought‑driven empathy. To the best of our knowledge, Nano‑EmoX is the first compact MLM (2.2B) to unify six core affective tasks across all three hierarchy levels, achieving state‑of‑the‑art or highly competitive performance across multiple benchmarks, demonstrating excellent efficiency and generalization. The code is available at https://github.com/waHAHJIAHAO/Nano‑EmoX.
Authors:Chuong Huynh, Manh Luong, Abhinav Shrivastava
Abstract:
Multimodal retrieval is the task of aggregating information from queries across heterogeneous modalities to retrieve desired targets. State‑of‑the‑art multimodal retrieval models can understand complex queries, yet they are typically limited to two modalities: text and vision. This limitation impedes the development of universal retrieval systems capable of comprehending queries that combine more than two modalities. To advance toward this goal, we present OmniRet, the first retrieval model capable of handling complex, composed queries spanning three key modalities: text, vision, and audio. Our OmniRet model addresses two critical challenges for universal retrieval: computational efficiency and representation fidelity. First, feeding massive token sequences from modality‑specific encoders to Large Language Models (LLMs) is computationally inefficient. We therefore introduce an attention‑based resampling mechanism to generate compact, fixed‑size representations from these sequences. Second, compressing rich omni‑modal data into a single embedding vector inevitably causes information loss and discards fine‑grained details. We propose Attention Sliced Wasserstein Pooling to preserve these fine‑grained details, leading to improved omni‑modal representations. OmniRet is trained on an aggregation of approximately 6 million query‑target pairs spanning 30 datasets. We benchmark our model on 13 retrieval tasks and a MMEBv2 subset. Our model demonstrates significant improvements on composed query, audio and video retrieval tasks, while achieving on‑par performance with state‑of‑the‑art models on others. Furthermore, we curate a new Audio‑Centric Multimodal Benchmark (ACM). This new benchmark introduces two critical, previously missing tasks‑composed audio retrieval and audio‑visual retrieval to more comprehensively evaluate a model's omni‑modal embedding capacity.
Authors:Harikrishnan Unnikrishnan, Rita Patel
Abstract:
We present a fully automated, two‑stage modular glottal area segmentation framework for high‑speed videoendoscopy (HSV) designed for accuracy, generalizability, and real‑time playback. Our detection‑gated pipeline combines a YOLOv8n glottis localizer with a U‑Net segmenter; the localizer defines a tight crop to ensure a consistent field of view and gates the output to reduce spurious segmentations during glottal closure. The models were trained on the GIRAFE (N=600) and BAGLS (N=55,750) datasets. Cross‑dataset portability was evaluated by benchmarking GIRAFE‑trained models on the BAGLS test set without fine‑tuning. In these evaluations, the pipeline achieved a Dice Similarity Coefficient (DSC) of 0.745 (87% of the in‑domain ceiling). On in‑distribution test sets, the system achieved DSCs of 0.81 (GIRAFE) and 0.856 (BAGLS), outperforming or competing with state‑of‑the‑art methods. An exploratory clinical study of 40 subjects demonstrated that the glottal area Coefficient of Variation (CV distinguished healthy from pathological function (p=0.006). The system processes ~35 frames per second on commodity hardware, enabling interactive clinical review. This design supports uniform extraction of laryngeal kinematic measures across varying acquisition settings. Code, weights, and software are available at https://github.com/hari‑krishnan/openglottal.
Authors:Joël Küchler, Ellen van Maren, Vaiva Vasiliauskaitė, Katarina Vulić, Reza Abbasi-Asl, Stephan J. Ihle
Abstract:
Although data generation is often straightforward, extracting information from data is more difficult. Object‑centric representation learning can extract information from images in an unsupervised manner. It does so by segmenting an image into its subcomponents: the objects. Each object is then represented in a low‑dimensional latent space that can be used for downstream processing. Object‑centric representation learning is dominated by autoencoder architectures (AEs). Here, we present ORGAN, a novel approach for object‑centric representation learning, which is based on cycle‑consistent Generative Adversarial Networks instead. We show that it performs similarly to other state‑of‑the‑art approaches on synthetic datasets, while at the same time being the only approach tested here capable of handling more challenging real‑world datasets with many objects and low visual contrast. Complementing these results, ORGAN creates expressive latent space representations that allow for object manipulation. Finally, we show that ORGAN scales well both with respect to the number of objects and the size of the images, giving it a unique edge over current state‑of‑the‑art approaches.
Authors:Fabian Schmidt, Karol Fedurko, Markus Enzweiler, Abhinav Valada
Abstract:
While multimodal large language models (MLLMs) provide advanced reasoning for autonomous driving, translating their discrete semantic knowledge into continuous trajectories remains a fundamental challenge. Existing methods often rely on unimodal planning heads that inherently limit their ability to represent multimodal driving behavior. Furthermore, most generative approaches frequently condition on one‑hot encoded actions, discarding the nuanced navigational uncertainty critical for complex scenarios. To resolve these limitations, we introduce LAD‑Drive, a generative framework that structurally disentangles high‑level intention from low‑level spatial planning. LAD‑Drive employs an action decoder to infer a probabilistic meta‑action distribution, establishing an explicit belief state that preserves the nuanced intent typically lost by one‑hot encodings. This distribution, fused with the vehicle's kinematic state, conditions an action‑aware diffusion decoder that utilizes a truncated denoising process to refine learned motion anchors into safe, kinematically feasible trajectories. Extensive evaluations on the LangAuto benchmark demonstrate that LAD‑Drive achieves state‑of‑the‑art results, outperforming competitive baselines by up to 59% in Driving Score while significantly reducing route deviations and collisions. We will publicly release the code and models on https://github.com/iis‑esslingen/lad‑drive.
Authors:Jingbiao Mei, Jinghong Chen, Guangyu Yang, Xinyu Hou, Margaret Li, Bill Byrne
Abstract:
Personalized AI assistants must recall and reason over long‑term user memory, which naturally spans multiple modalities and sources such as images, videos, and emails. However, existing Long‑term Memory benchmarks focus primarily on dialogue history, failing to capture realistic personalized references grounded in lived experience. We introduce ATM‑Bench, the first benchmark for multimodal, multi‑source personalized referential Memory QA. ATM‑Bench contains approximately four years of privacy‑preserving personal memory data and human‑annotated question‑answer pairs with ground‑truth memory evidence, including queries that require resolving personal references, multi‑evidence reasoning from multi‑source and handling conflicting evidence. We propose Schema‑Guided Memory (SGM) to structurally represent memory items originated from different sources. In experiments, we implement 5 state‑of‑the‑art memory systems along with a standard RAG baseline and evaluate variants with different memory ingestion, retrieval, and answer generation techniques. We find poor performance (under 20% accuracy) on the ATM‑Bench‑Hard set, and that SGM improves performance over Descriptive Memory commonly adopted in prior works. Code available at: https://github.com/JingbiaoMei/ATM‑Bench
Authors:Pengyuan Wu, Pingrui Zhang, Zhigang Wang, Dong Wang, Bin Zhao, Xuelong Li
Abstract:
Diffusion‑based policies have achieved remarkable results in robotic manipulation but often struggle to adapt rapidly in dynamic scenarios, leading to delayed responses or task failures. We present DCDP, a Dynamic Closed‑Loop Diffusion Policy framework that integrates chunk‑based action generation with real‑time correction. DCDP integrates a self‑supervised dynamic feature encoder, cross‑attention fusion, and an asymmetric action encoder‑decoder to inject environmental dynamics before action execution, achieving real‑time closed‑loop action correction and enhancing the system's adaptability in dynamic scenarios. In dynamic PushT simulations, DCDP improves adaptability by 19% without retraining while requiring only 5% additional computation. Its modular design enables plug‑and‑play integration, achieving both temporal coherence and real‑time responsiveness in dynamic robotic scenarios, including real‑world manipulation tasks. The project page is at: https://github.com/wupengyuan/dcdp
Authors:Dinh Nam Pham, Leonard Prokisch, Bennet Meyer, Jonas Thumbs
Abstract:
Smartphone clip‑on microscopes turn everyday devices into low‑cost, portable imaging systems that can even reveal fungal structures at the microscopic level, enabling mold inspection beyond unaided visual checks. In this paper, we introduce MobileMold, an open smartphone‑based microscopy dataset for food mold detection and food classification. MobileMold contains 4,941 handheld microscopy images spanning 11 food types, 4 smartphones, 3 microscopes, and diverse real‑world conditions. Beyond the dataset release, we establish baselines for (i) mold detection and (ii) food‑type classification, including a multi‑task setting that predicts both attributes. Across multiple pretrained deep learning architectures and augmentation strategies, we obtain near‑ceiling performance (accuracy = 0.9954, F1 = 0.9954, MCC = 0.9907), validating the utility of our dataset for detecting food spoilage. To increase transparency, we complement our evaluation with saliency‑based visual explanations highlighting mold regions associated with the model's predictions. MobileMold aims to contribute to research on accessible food‑safety sensing, mobile imaging, and exploring the potential of smartphones enhanced with attachments.
Authors:Zijin Yin, Tiankai Hang, Yiji Cheng, Shiyi Zhang, Runze He, Yu Xu, Chunyu Wang, Bing Li, Zheng Chang, Kongming Liang, Qinglin Lu, Zhanyu Ma
Abstract:
Existing image editing methods struggle to perceive where to edit, especially under complex scenes and nuanced spatial instructions. To address this issue, we propose Generative Visual Chain‑of‑Thought (GVCoT), a unified framework that performs native visual reasoning by first generating spatial cues to localize the target region and then executing the edit. Unlike prior text‑only CoT or tool‑dependent visual CoT paradigms, GVCoT jointly optimizes visual tokens generated during the reasoning and editing phases in an end‑to‑end manner. This way fosters the emergence of innate spatial reasoning ability and enables more effective utilization of visual‑domain cues. The main challenge of training GCVoT lies in the scarcity of large‑scale editing data with precise edit region annotations; to this end, we construct GVCoT‑Edit‑Instruct, a dataset of 1.8M high‑quality samples spanning 19 tasks. We adopt a progressive training strategy: supervised fine‑tuning to build foundational localization ability in reasoning trace before final editing, followed by reinforcement learning to further improve reasoning and editing quality. Finally, we introduce SREdit‑Bench, a new benchmark designed to comprehensively stress‑test models under sophisticated scenes and fine‑grained referring expressions. Experiments demonstrate that GVCoT consistently outperforms state‑of‑the‑art models on SREdit‑Bench and ImgEdit. We hope our GVCoT will inspire future research toward interpretable and precise image editing.
Authors:Yiheng Li, Zichang Tan, Guoqing Xu, Yijun Ye, Yang Yang, Zhen Lei
Abstract:
With the rapid development of generative AI in medical imaging, synthetic Computed Tomography (CT) images have demonstrated great potential in applications such as data augmentation and clinical diagnosis, but they also introduce serious security risks. Despite the increasing security concerns, existing studies on CT forgery detection are still limited and fail to adequately address real‑world challenges. These limitations are mainly reflected in two aspects: the absence of datasets that can effectively evaluate model generalization to reflect the real‑world application requirements, and the reliance on detection methods designed for natural images that are insensitive to CT‑specific forgery artifacts. In this view, we propose CTForensics, a comprehensive dataset designed to systematically evaluate the generalization capability of CT forgery detection methods, which includes ten diverse CT generative methods. Moreover, we introduce the Enhanced Spatial‑Frequency CT Forgery Detector (ESF‑CTFD), an efficient CNN‑based neural network that captures forgery cues across the wavelet, spatial, and frequency domains. First, it transforms the input CT image into three scales and extracts features at each scale via the Wavelet‑Enhanced Central Stem. Then, starting from the largest‑scale features, the Spatial Process Block gradually performs feature fusion with the smaller‑scale ones. Finally, the Frequency Process Block learns frequency‑domain information for predicting the final results. Experiments demonstrate that ESF‑CTFD consistently outperforms existing methods and exhibits superior generalization across different CT generative models.
Authors:Alexander Prutsch, David Schinagl, Horst Possegger
Abstract:
Future trajectories of neighboring traffic agents have a significant influence on the path planning and decision‑making of autonomous vehicles. While trajectory forecasting is a well‑studied field, research mainly focuses on snapshot‑based prediction, where each scenario is treated independently of its global temporal context. However, real‑world autonomous driving systems need to operate in a continuous setting, requiring real‑time processing of data streams with low latency and consistent predictions over successive timesteps. We leverage this continuous setting to propose a lightweight yet highly accurate streaming‑based trajectory forecasting approach. We integrate valuable information from previous predictions with a novel endpoint‑aware modeling scheme. Our temporal context propagation uses the trajectory endpoints of the previous forecasts as anchors to extract targeted scenario context encodings. Our approach efficiently guides its scene encoder to extract highly relevant context information without needing refinement iterations or segment‑wise decoding. Our experiments highlight that our approach effectively relays information across consecutive timesteps. Unlike methods using multi‑stage refinement processing, our approach significantly reduces inference latency, making it well‑suited for real‑world deployment. We achieve state‑of‑the‑art streaming trajectory prediction results on the Argoverse~2 multi‑agent and single‑agent benchmarks, while requiring substantially fewer resources.
Authors:Yutong Yang, Katarina Popović, Julian Wiederer, Markus Braun, Vasileios Belagiannis, Bin Yang
Abstract:
Detection Transformer (DETR) and its variants show strong performance on object detection, a key task for autonomous systems. However, a critical limitation of these models is that their confidence scores only reflect semantic uncertainty, failing to capture the equally important spatial uncertainty. This results in an incomplete assessment of the detection reliability. On the other hand, Deep Ensembles can tackle this by providing high‑quality spatial uncertainty estimates. However, their immense memory consumption makes them impractical for real‑world applications. A cheaper alternative, Monte Carlo (MC) Dropout, suffers from high latency due to the need of multiple forward passes during inference to estimate uncertainty.
To address these limitations, we introduce GroupEnsemble, an efficient and effective uncertainty estimation method for DETR‑like models. GroupEnsemble simultaneously predicts multiple individual detection sets by feeding additional diverse groups of object queries to the transformer decoder during inference. Each query group is transformed by the shared decoder in isolation and predicts a complete detection set for the same input. An attention mask is applied to the decoder to prevent inter‑group query interactions, ensuring each group detects independently to achieve reliable ensemble‑based uncertainty estimation. By leveraging the decoder's inherent parallelism, GroupEnsemble efficiently estimates uncertainty in a single forward pass without sequential repetition. We validated our method under autonomous driving scenes and common daily scenes using the Cityscapes and COCO datasets, respectively. The results show that a hybrid approach combining MC‑Dropout and GroupEnsemble outperforms Deep Ensembles on several metrics at a fraction of the cost. The code is available at https://github.com/yutongy98/GroupEnsemble.
Authors:Bosen Lin, Feng Gao, Yanwei Yu, Junyu Dong, Qian Du
Abstract:
In real underwater environments, downstream image recognition tasks such as semantic segmentation and object detection often face challenges posed by problems like blurring and color inconsistencies. Underwater image enhancement (UIE) has emerged as a promising preprocessing approach, aiming to improve the recognizability of targets in underwater images. However, most existing UIE methods mainly focus on enhancing images for human visual perception, frequently failing to reconstruct high‑frequency details that are critical for task‑specific recognition. To address this issue, we propose a Downstream Task‑Inspired Underwater Image Enhancement (DTI‑UIE) framework, which leverages human visual perception model to enhance images effectively for underwater vision tasks. Specifically, we design an efficient two‑branch network with task‑aware attention module for feature mixing. The network benefits from a multi‑stage training framework and a task‑driven perceptual loss. Additionally, inspired by human perception, we automatically construct a Task‑Inspired UIE Dataset (TI‑UIED) using various task‑specific networks. Experimental results demonstrate that DTI‑UIE significantly improves task performance by generating preprocessed images that are beneficial for downstream tasks such as semantic segmentation, object detection, and instance segmentation. The codes are publicly available at https://github.com/oucailab/DTIUIE.
Authors:Yuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng, Xiang Li, Jian Yang
Abstract:
Heterogeneous multi‑modal remote sensing object detection aims to accurately detect objects from diverse sensors (e.g., RGB, SAR, Infrared). Existing approaches largely adopt a late alignment paradigm, in which modality alignment and task‑specific optimization are entangled during downstream fine‑tuning. This tight coupling complicates optimization and often results in unstable training and suboptimal generalization. To address these limitations, we propose BabelRS, a unified language‑pivoted pretraining framework that explicitly decouples modality alignment from downstream task learning. BabelRS comprises two key components: Concept‑Shared Instruction Aligning (CSIA) and Layerwise Visual‑Semantic Annealing (LVSA). CSIA aligns each sensor modality to a shared set of linguistic concepts, using language as a semantic pivot to bridge heterogeneous visual representations. To further mitigate the granularity mismatch between high‑level language representations and dense detection objectives, LVSA progressively aggregates multi‑scale visual features to provide fine‑grained semantic guidance. Extensive experiments demonstrate that BabelRS stabilizes training and consistently outperforms state‑of‑the‑art methods without bells and whistles. Code: https://github.com/zcablii/SM3Det.
Authors:Guanglu Dong, Chunlei Li, Chao Ren, Jingliang Hu, Yilei Shi, Xiao Xiang Zhu, Lichao Mou
Abstract:
Recently, significant breakthroughs have been made in all‑in‑one image restoration (AiOIR), which can handle multiple restoration tasks with a single model. However, existing methods typically focus on a specific image domain, such as natural scene, medical imaging, or remote sensing. In this work, we aim to extend AiOIR to multiple domains and propose the first multi‑domain all‑in‑one image restoration method, DATPRL‑IR, based on our proposed Domain‑Aware Task Prompt Representation Learning. Specifically, we first construct a task prompt pool containing multiple task prompts, in which task‑related knowledge is implicitly encoded. For each input image, the model adaptively selects the most relevant task prompts and composes them into an instance‑level task representation via a prompt composition mechanism (PCM). Furthermore, to endow the model with domain awareness, we introduce another domain prompt pool and distill domain priors from multimodal large language models into the domain prompts. PCM is utilized to combine the adaptively selected domain prompts into a domain representation for each input image. Finally, the two representations are fused to form a domain‑aware task prompt representation which can make full use of both specific and shared knowledge across tasks and domains to guide the subsequent restoration process. Extensive experiments demonstrate that our DATPRL‑IR significantly outperforms existing SOTA image restoration methods, while exhibiting strong generalization capabilities. Code is available at https://github.com/GuangluDong0728/DATPRL‑IR.
Authors:Ruize Cui, Jialun Pei, Haiqiao Wang, Jun Zhou, Jeremy Yuen-Chun Teoh, Pheng-Ann Heng, Jing Qin
Abstract:
In laparoscopic liver surgery, augmented reality technology enhances intraoperative anatomical guidance by overlaying 3D liver models from preoperative CT/MRI onto laparoscopic 2D views. However, existing registration methods lack explicit modeling of reliable 2D‑3D geometric correspondences supported by latent evidence, leading to limited interpretability and potentially unstable alignment in clinical scenarios. In this work, we introduce Land‑Reg, a correspondence‑driven deformable registration framework that explicitly learns latent‑grounded 2D‑3D landmark correspondences as an interpretable intermediate representation to bridge cross‑modal alignment. For rigid registration, Land‑Reg embraces a Cross‑modal Latent Alignment module to map multi‑modal features into a unified latent space. Further, an Uncertainty‑enhanced Overlap Landmark Detector with similarity matching is proposed to robustly estimate explicit 2D‑3D landmark correspondences. For non‑rigid registration, we design a novel shape‑constrained supervision strategy that anchors shape deformation to matched landmarks through reprojection consistency and incorporates local‑isometric regularization to alleviate inherent 2D‑3D depth ambiguity, while a rendered‑mask alignment enforces global shape consistency. Experimental results on the P2ILF dataset demonstrate the superiority of our method on both rigid pose estimation and non‑rigid deformation. Our code will be available at https://github.com/cuiruize/Land‑Reg.
Authors:Le Dong, Qinzhong Tan, Chunlei Li, Jingliang Hu, Yilei Shi, Weisheng Dong, Xiao Xiang Zhu, Lichao Mou
Abstract:
Anomaly detection is a critical task in computer vision with profound implications for medical imaging, where identifying pathologies early can directly impact patient outcomes. While recent unsupervised anomaly detection approaches show promise, they require substantial normal training data and struggle to generalize across anatomical contexts. We introduce D^24FAD, a novel dual distillation framework for few‑shot anomaly detection that identifies anomalies in previously unseen tasks using only a small number of normal reference images. Our approach leverages a pre‑trained encoder as a teacher network to extract multi‑scale features from both support and query images, while a student decoder learns to distill knowledge from the teacher on query images and self‑distill on support images. We further propose a learn‑to‑weight mechanism that dynamically assesses the reference value of each support image conditioned on the query, optimizing anomaly detection performance. To evaluate our method, we curate a comprehensive benchmark dataset comprising 13,084 images across four organs, four imaging modalities, and five disease categories. Extensive experiments demonstrate that D^24FAD significantly outperforms existing approaches, establishing a new state‑of‑the‑art in few‑shot medical anomaly detection. Code is available at https://github.com/ttttqz/D24FAD.
Authors:Xiangyang He, Lin Wan
Abstract:
Cloth‑Changing Person Re‑Identification (CC‑ReID) aims to match the same individual across cameras under varying clothing conditions. Existing approaches often remove apparel and focus on the head region to reduce clothing bias. However, treating the head holistically without distinguishing between face and hair leads to over‑reliance on volatile hairstyle cues, causing performance degradation under hairstyle changes. To address this issue, we propose the Mitigating Hairstyle Distraction and Structural Preservation (MSP) framework. Specifically, MSP introduces Hairstyle‑Oriented Augmentation (HSOA), which generates intra‑identity hairstyle diversity to reduce hairstyle dependence and enhance attention to stable facial and body cues. To prevent the loss of structural information, we design Cloth‑Preserved Random Erasing (CPRE), which performs ratio‑controlled erasing within clothing regions to suppress texture bias while retaining body shape and context. Furthermore, we employ Region‑based Parsing Attention (RPA) to incorporate parsing‑guided priors that highlight face and limb regions while suppressing hair features. Extensive experiments on multiple CC‑ReID benchmarks demonstrate that MSP achieves state‑of‑the‑art performance, providing a robust and practical solution for long‑term person re‑identification.
Authors:Bo Ma, Jinsong Wu, Weiqi Yan, Catherine Shi, Minh Nguyen
Abstract:
Dashcam videos collected by autonomous or assisted‑driving systems are increasingly shared for safety auditing and model improvement. Even when explicit GPS metadata are removed, an attacker can still infer the recording location by matching background visual cues (e.g., buildings and road layouts) against large‑scale street‑view imagery. This paper studies location‑privacy leakage under a background‑based retrieval attacker, and proposes PPEDCRF, a privacy‑preserving enhanced dynamic conditional random field framework that injects calibrated perturbations only into inferred location‑sensitive background regions while preserving foreground detection utility. PPEDCRF consists of three components: (i) a dynamic CRF that enforces temporal consistency to discover and track location sensitive regions across frames, (ii) a normalized control penalty (NCP) that allocates perturbation strength according to a hierarchical sensitivity model, and (iii) a utility‑preserving noise injection module that minimizes interference to object detection and segmentation. Experiments on public driving datasets demonstrate that PPEDCRF significantly reduces location‑retrieval attack success (e.g., Top‑k retrieval accuracy) while maintaining competitive detection performance (e.g., mAP and segmentation metrics) compared with common baselines such as global noise, white‑noise masking, and feature‑based anonymization. The source code is in https://github.com/mabo1215/PPEDCRF.git
Authors:Jisoo Kim, Jungbin Cho, Sanghyeok Chu, Ananya Bal, Jinhyung Kim, Gunhee Lee, Sihaeng Lee, Seung Hwan Kim, Bohyung Han, Hyunmin Lee, Laszlo A. Jeni, Seungryong Kim
Abstract:
Humans learn not only how their bodies move, but also how the surrounding world responds to their actions. In contrast, while recent Vision‑Language‑Action (VLA) models exhibit impressive semantic understanding, they often fail to capture the spatiotemporal dynamics governing physical interaction. In this paper, we introduce Pri4R, a simple yet effective approach that endows VLA models with an implicit understanding of world dynamics by leveraging privileged 4D information during training. Specifically, Pri4R augments VLAs with a lightweight point track head that predicts 3D point tracks. By injecting VLA features into this head to jointly predict future 3D trajectories, the model learns to incorporate evolving scene geometry within its shared representation space, enabling more physically aware context for precise control. Due to its architectural simplicity, Pri4R is compatible with dominant VLA design patterns with minimal changes. During inference, we run the model using the original VLA architecture unchanged; Pri4R adds no extra inputs, outputs, or computational overhead. Across simulation and real‑world evaluations, Pri4R significantly improves performance on challenging manipulation tasks, including a +10% gain on LIBERO‑Long and a +40% gain on RoboCasa. We further show that 3D point track prediction is an effective supervision target for learning action‑world dynamics, and validate our design choices through extensive ablations. Project page: https://jiiiisoo.github.io/Pri4R/
Authors:Yutian Zhang, Zhongyi Pei, Yi Mao, Chen Wang, Lin Liu, Jianmin Wang
Abstract:
The widespread adoption of AI in industry is often hampered by its limited robustness when faced with scenarios absent from training data, leading to prediction bias and vulnerabilities. To address this, we propose a novel streaming inference pipeline that enhances data‑driven models by explicitly incorporating prior knowledge. This paper presents the work on an industrial AI application that automatically counts excavator workloads from surveillance videos. Our approach integrates an object detection model with a Finite State Machine (FSM), which encodes knowledge of operational scenarios to guide and correct the AI's predictions on streaming data. In experiments on a real‑world dataset of over 7,000 images from 12 site videos, encompassing more than 300 excavator workloads, our method demonstrates superior performance and greater robustness compared to the original solution based on manual heuristic rules. We will release the code at https://github.com/thulab/video‑streamling‑inference‑pipeline.
Authors:Jianqiang Ren, Lin Liu, Steven Hoi
Abstract:
We propose OMG‑Avatar, a novel One‑shot method that leverages a Multi‑LOD (Level‑of‑Detail) Gaussian representation for animatable 3D head reconstruction from a single image in 0.2s. Our method enables LOD head avatar modeling using a unified model that accommodates diverse hardware capabilities and inference speed requirements. To capture both global and local facial characteristics, we employ a transformer‑based architecture for global feature extraction and projection‑based sampling for local feature acquisition. These features are effectively fused under the guidance of a depth buffer, ensuring occlusion plausibility. We further introduce a coarse‑to‑fine learning paradigm to support Level‑of‑Detail functionality and enhance the perception of hierarchical details. To address the limitations of 3DMMs in modeling non‑head regions such as the shoulders, we introduce a multi‑region decomposition scheme in which the head and shoulders are predicted separately and then integrated through cross‑region combination. Extensive experiments demonstrate that OMG‑Avatar outperforms state‑of‑the‑art methods in reconstruction quality, reenactment performance, and computational efficiency. The project homepage is https://human3daigc.github.io/OMGAvatar_project_page/ .
Authors:Brian Cheong, Letian Wang, Sandro Papais, Steven L. Waslander
Abstract:
LiDAR‑based tracking‑by‑attention (TBA) frameworks inherently suffer from high false negative errors, leading to a significant performance gap compared to traditional LiDAR‑based tracking‑by‑detection (TBD) methods. This paper introduces SCATR, a novel LiDAR‑based TBA model designed to address this fundamental challenge systematically. SCATR leverages recent progress in vision‑based tracking and incorporates targeted training strategies specifically adapted for LiDAR.
Our work's core innovations are two architecture‑agnostic training strategies for TBA methods: Second Chance Assignment and Track Query Dropout.
Second Chance Assignment is a novel ground truth assignment that concatenates unassigned track queries to the proposal queries before bipartite matching, giving these track queries a second chance to be assigned to a ground truth object and effectively mitigating the conflict between detection and tracking tasks inherent in tracking‑by‑attention.
Track Query Dropout is a training method that diversifies supervised object query configurations to efficiently train the decoder to handle different track query sets, enhancing robustness to missing or newborn tracks.
Experiments on the nuScenes tracking benchmark demonstrate that SCATR achieves state‑of‑the‑art performance among LiDAR‑based TBA methods, outperforming previous works by 7.6% AMOTA and successfully bridging the long‑standing performance gap between LiDAR‑based TBA and TBD methods.
Ablation studies further validate the effectiveness and generalization of Second Chance Assignment and Track Query Dropout.
Code can be found at the following link: \hrefhttps://github.com/TRAILab/SCATRhttps://github.com/TRAILab/SCATR
Authors:Joshua Knights, Joseph Reid, Kaushik Roy, David Hall, Mark Cox, Peyman Moghadam
Abstract:
Recent years have seen a significant increase in demand for robotic solutions in unstructured natural environments, alongside growing interest in bridging 2D and 3D scene understanding. However, existing robotics datasets are predominantly captured in structured urban environments, making them inadequate for addressing the challenges posed by complex, unstructured natural settings. To address this gap, we propose WildCross, a cross‑modal benchmark for place recognition and metric depth estimation in large‑scale natural environments. WildCross comprises over 476K sequential RGB frames with semi‑dense depth and surface normal annotations, each aligned with accurate 6DoF poses and synchronized dense lidar submaps. We conduct comprehensive experiments on visual, lidar, and cross‑modal place recognition, as well as metric depth estimation, demonstrating the value of WildCross as a challenging benchmark for multi‑modal robotic perception tasks. We provide access to the code repository and dataset at https://csiro‑robotics.github.io/WildCross.
Authors:Niu Lian, Yuting Wang, Hanshu Yao, Jinpeng Wang, Bin Chen, Yaowei Wang, Min Zhang, Shu-Tao Xia
Abstract:
While multimodal large language models have demonstrated impressive short‑term reasoning, they struggle with long‑horizon video understanding due to limited context windows and static memory mechanisms that fail to mirror human cognitive efficiency. Existing paradigms typically fall into two extremes: vision‑centric methods that incur high latency and redundancy through dense visual accumulation, or text‑centric approaches that suffer from detail loss and hallucination via aggressive captioning. To bridge this gap, we propose MM‑Mem, a pyramidal multimodal memory architecture grounded in Fuzzy‑Trace Theory. MM‑Mem structures memory hierarchically into a Sensory Buffer, Episodic Stream, and Symbolic Schema, enabling the progressive distillation of fine‑grained perceptual traces (verbatim) into high‑level semantic schemas (gist). Furthermore, to govern the dynamic construction of memory, we derive a Semantic Information Bottleneck objective and introduce SIB‑GRPO to optimize the trade‑off between memory compression and task‑relevant information retention. In inference, we design an entropy‑driven top‑down memory retrieval strategy. Extensive experiments across 4 benchmarks confirm that MM‑Mem achieves state‑of‑the‑art performance on both offline and streaming tasks, demonstrating robust generalization and validating the effectiveness of cognition‑inspired memory organization. Code and associated configurations are publicly available at https://github.com/EliSpectre/MM‑Mem.
Authors:Jianfeng Liao, Yichen Wei, Raymond Chan Ching Bon, Shulan Wang, Kam-Pui Chow, Kwok-Yan Lam
Abstract:
The rapid advancement of deepfake generation techniques poses significant threats to public safety and causes societal harm through the creation of highly realistic synthetic facial media. While existing detection methods demonstrate limitations in generalizing to emerging forgery patterns, this paper presents Deepfake Forensics Adapter (DFA), a novel dual‑stream framework that synergizes vision‑language foundation models with targeted forensics analysis. Our approach integrates a pre‑trained CLIP model with three core components to achieve specialized deepfake detection by leveraging the powerful general capabilities of CLIP without changing CLIP parameters: 1) A Global Feature Adapter is used to identify global inconsistencies in image content that may indicate forgery, 2) A Local Anomaly Stream enhances the model's ability to perceive local facial forgery cues by explicitly leveraging facial structure priors, and 3) An Interactive Fusion Classifier promotes deep interaction and fusion between global and local features using a transformer encoder. Extensive evaluations of frame‑level and video‑level benchmarks demonstrate the superior generalization capabilities of DFA, particularly achieving state‑of‑the‑art performance in the challenging DFDC dataset with frame‑level AUC/EER of 0.816/0.256 and video‑level AUC/EER of 0.836/0.251, representing a 4.8% video AUC improvement over previous methods. Our framework not only demonstrates state‑of‑the‑art performance, but also points out a feasible and effective direction for developing a robust deepfake detection system with enhanced generalization capabilities against the evolving deepfake threats. Our code is available at https://github.com/Liao330/DFA.git
Authors:Ben Kang, Jie Zhao, Xin Chen, Wanting Geng, Bin Zhang, Lu Zhang, Dong Wang, Huchuan Lu
Abstract:
With growing real‑world demands, efficient tracking has received increasing attention. However, most existing methods are limited to RGB inputs and struggle in multi‑modal scenarios. Moreover, current multi‑modal tracking approaches typically use complex designs, making them too heavy and slow for resource‑constrained deployment. To tackle these limitations, we propose UETrack, an efficient framework for single object tracking. UETrack demonstrates high practicality and versatility, efficiently handling multiple modalities including RGB, Depth, Thermal, Event, and Language, and addresses the gap in efficient multi‑modal tracking. It introduces two key components: a Token‑Pooling‑based Mixture‑of‑Experts mechanism that enhances modeling capacity through feature aggregation and expert specialization, and a Target‑aware Adaptive Distillation strategy that selectively performs distillation based on sample characteristics, reducing redundant supervision and improving performance. Extensive experiments on 12 benchmarks across 3 hardware platforms show that UETrack achieves a superior speed‑accuracy trade‑off compared to previous methods. For instance, UETrack‑B achieves 69.2% AUC on LaSOT and runs at 163/56/60 FPS on GPU/CPU/AGX, demonstrating strong practicality and versatility. Code is available at https://github.com/kangben258/UETrack.
Authors:Jinlong Li, Liyuan Jiang, Haonan Zhang, Nicu Sebe
Abstract:
Video Large Language Models (VLLMs) demonstrate strong video understanding but suffer from inefficiency due to redundant visual tokens. Existing pruning primary targets intra‑frame spatial redundancy or prunes inside the LLM with shallow‑layer overhead, yielding suboptimal spatiotemporal reduction and underutilizing long‑context compressibility. All of them often discard subtle yet informative context from merged or pruned tokens. In this paper, we propose a new perspective that elaborates token Anchors within intra‑frame and inter‑frame to comprehensively aggregate the informative contexts via local‑global Optimal Transport (AOT). Specifically, we first establish local‑ and global‑aware token anchors within each frame under the attention guidance, which then optimal transport aggregates the informative contexts from pruned tokens, constructing intra‑frame token anchors. Then, building on the temporal frame clips, the first frame within each clip will be considered as the keyframe anchors to ensemble similar information from consecutive frames through optimal transport, while keeping distinct tokens to represent temporal dynamics, leading to efficient token reduction in a training‑free manner. Extensive evaluations show that our proposed AOT obtains competitive performances across various short‑ and long‑video benchmarks on leading video LLMs, obtaining substantial computational efficiency while preserving temporal and visual fidelity. Project webpage: https://tyroneli.github.io/AOT.
Authors:Zilong Zhao, Zhengming Ding, Pei Niu, Wenhao Sun, Feng Guo
Abstract:
Feature encoders play a key role in pixel‑level crack segmentation by shaping the representation of fine textures and thin structures. Existing CNN‑, Transformer‑, and Mamba‑based models each capture only part of the required spatial or structural information, leaving clear gaps in modeling complex crack patterns. To address this, we present MixerCSeg, a mixer architecture designed like a coordinated team of specialists, where CNN‑like pathways focus on local textures, Transformer‑style paths capture global dependencies, and Mamba‑inspired flows model sequential context within a single encoder. At the core of MixerCSeg is the TransMixer, which explores Mamba's latent attention behavior while establishing dedicated pathways that naturally express both locality and global awareness. To further enhance structural fidelity, we introduce a spatial block processing strategy and a Direction‑guided Edge Gated Convolution (DEGConv) that strengthens edge sensitivity under irregular crack geometries with minimal computational overhead. A Spatial Refinement Multi‑Level Fusion (SRF) module is then employed to refine multi‑scale details without increasing complexity. Extensive experiments on multiple crack segmentation benchmarks show that MixerCSeg achieves state‑of‑the‑art performance with only 2.05 GFLOPs and 2.54 M parameters, demonstrating both efficiency and strong representational capability. The code is available at https://github.com/spiderforest/MixerCSeg.
Authors:Andrew Wang, Mike Davies
Abstract:
Multispectral demosaicing is crucial to reconstruct full‑resolution spectral images from snapshot mosaiced measurements, enabling real‑time imaging from neurosurgery to autonomous driving. Classical methods are blurry, while supervised learning requires costly ground truth (GT) obtained from slow line‑scanning systems. We propose Perspective‑Equivariant Fine‑tuning for Demosaicing (PEFD), a framework that learns multispectral demosaicing from mosaiced measurements alone. PEFD a) exploits the projective geometry of camera‑based imaging systems to leverage a richer group structure than previous demosaicing methods to recover more null‑space information, and b) learns efficiently without GT by adapting pretrained foundation models designed for 1‑3 channel imaging. On surgical and automotive datasets, PEFD recovers fine details such as blood vessels and preserves spectral fidelity, substantially outperforming recent approaches, nearing supervised performance. Furthermore, the performance of PEFD is demonstrated on raw, unprocessed data from a commercial multispectral sensor. Code is at https://github.com/Andrewwango/pefd.
Authors:Abdullah Al Shafi, Md Kawsar Mahmud Khan Zunayed, Safin Ahmmed, Sk Imran Hossain, Engelbert Mephu Nguifo
Abstract:
Breast ultrasound interpretation requires simultaneous lesion segmentation and tissue classification. However, conventional multi‑task learning approaches suffer from task interference and rigid coordination strategies that fail to adapt to instance‑specific prediction difficulty. We propose a multi‑task framework addressing these limitations through multi‑level decoder interaction and uncertainty‑aware adaptive coordination. Task Interaction Modules operate at all decoder levels, establishing bidirectional segmentation‑classification communication during spatial reconstruction through attention weighted pooling and multiplicative modulation. Unlike prior single‑level or encoder‑only approaches, this multi‑level design captures scale specific task synergies across semantic‑to‑spatial scales, producing complementary task interaction streams. Uncertainty‑Proxy Attention adaptively weights base versus enhanced features at each level using feature activation variance, enabling per‑level and per‑sample task balancing without heuristic tuning. To support instance‑adaptive prediction, multi‑scale context fusion captures morphological cues across varying lesion sizes. Evaluation on multiple publicly available breast ultrasound datasets demonstrates competitive performance, including 74.5% lesion IoU and 90.6% classification accuracy on BUSI dataset. Ablation studies confirm that multi‑level task interaction provides significant performance gains, validating that decoder‑level bidirectional communication is more effective than conventional encoder‑only parameter sharing. The code is available at: https://github.com/C‑loud‑Nine/Uncertainty‑Aware‑Multi‑Level‑Decoder‑Interaction.
Authors:Changwoo Baek, Jouwon Song, Sohyeon Kim, Kyeongbo Kong
Abstract:
Large Vision‑Language Models (LVLMs) have adopted visual token pruning strategies to mitigate substantial computational overhead incurred by extensive visual token sequences. While prior works primarily focus on either attention‑based or diversity‑based pruning methods, in‑depth analysis of these approaches' characteristics and limitations remains largely unexplored. In this work, we conduct thorough empirical analysis using effective rank (erank) as a measure of feature diversity and attention score entropy to investigate visual token processing mechanisms and analyze the strengths and weaknesses of each approach. Our analysis reveals two insights: (1) Our erank‑based quantitative analysis shows that many diversity‑oriented pruning methods preserve substantially less feature diversity than intended; moreover, analysis using the CHAIR dataset reveals that the diversity they do retain is closely tied to increased hallucination frequency compared to attention‑based pruning. (2) We further observe that attention‑based approaches are more effective on simple images where visual evidence is concentrated, while diversity‑based methods better handle complex images with distributed features. Building on these empirical insights, we show that incorporating image‑aware adjustments into existing hybrid pruning strategies consistently improves their performance. We also provide a minimal instantiation of our empirical findings through a simple adaptive pruning mechanism, which achieves strong and reliable performance across standard benchmarks as well as hallucination‑specific evaluations. Our project page available at https://cvsp‑lab.github.io/AgilePruner.
Authors:Mochu Xiang, Zhelun Shen, Xuesong Li, Jiahui Ren, Jing Zhang, Chen Zhao, Shanshan Liu, Haocheng Feng, Jingdong Wang, Yuchao Dai
Abstract:
Human perceive the 3D world through 2D observations from limited viewpoints. While recent feed‑forward generalizable 3D reconstruction models excel at recovering 3D structures from sparse images, their representations are often confined to observed regions, leaving unseen geometry un‑modeled. This raises a key, fundamental challenge: Can we infer a complete 3D structure from partial 2D observations? We present RnG (Reconstruction and Generation), a novel feed‑forward Transformer that unifies these two tasks by predicting an implicit, complete 3D representation. At the core of RnG, we propose a reconstruction‑guided causal attention mechanism that separates reconstruction and generation at the attention level, and treats the KV‑cache as an implicit 3D representation. Then, arbitrary poses can efficiently query this cache to render high‑fidelity, novel‑view RGBD outputs. As a result, RnG not only accurately reconstructs visible geometry but also generates plausible, coherent unseen geometry and appearance. Our method achieves state‑of‑the‑art performance in both generalizable 3D reconstruction and novel view generation, while operating efficiently enough for real‑time interactive applications. Project page: https://npucvr.github.io/RnG
Authors:Sumin Kim, Hyemin Jeong, Mingu Kang, Yejin Kim, Yoori Oh, Joonseok Lee
Abstract:
The exponential growth of video content necessitates effective video summarization to efficiently extract key information from long videos. However, current approaches struggle to fully comprehend complex videos, primarily because they employ static or modality‑agnostic fusion strategies. These methods fail to account for the dynamic, frame‑dependent variations in modality saliency inherent in video data. To overcome these limitations, we propose TripleSumm, a novel architecture that adaptively weights and fuses the contributions of visual, text, and audio modalities at the frame level. Furthermore, a significant bottleneck for research into multimodal video summarization has been the lack of comprehensive benchmarks. Addressing this bottleneck, we introduce MoSu (Most Replayed Multimodal Video Summarization), the first large‑scale benchmark that provides all three modalities. Extensive experiments demonstrate that TripleSumm achieves state‑of‑the‑art performance, outperforming existing methods by a significant margin on four benchmarks, including MoSu. Our code and dataset are available at https://github.com/smkim37/TripleSumm.
Authors:Maomao Li, Yunfei Liu, Yu Li
Abstract:
Image‑driven video editing aims to propagate edit contents from the modified first frame to the remaining frames. Existing methods usually invert the source video to noise using a pre‑trained image‑to‑video (I2V) model and then guide the sampling process using the edited first frame. Generally, a popular choice for maintaining motion and layout from the source video is intervening in the denoising process by injecting attention during reconstruction. However, such injection often leads to unsatisfactory results, where excessive injection leads to conflicting semantics with the source video while insufficient injection brings limited source representation. Recognizing this, we propose an Editing‑awaRE (REE) injection method to modulate the injection intensity of each token. Specifically, we first compute the pixel difference between the source and edited first frame to form a corresponding editing mask. Next, we track the editing area throughout the entire video by using optical flow to warp the first‑frame mask. Then, editing‑aware feature injection intensity for each token is generated accordingly, where injection is not conducted in editing areas. Building upon REE injection, we further propose a zero‑shot image‑driven video editing framework with recent‑emerging rectified‑Flow models, dubbed FREE‑Edit. Without fine‑tuning or training, our FREE‑Edit demonstrates effectiveness in various image‑driven video editing scenarios, showing its capability to produce higher‑quality outputs compared with existing techniques. Project page: https://free‑edit.github.io/page/.
Authors:Durgesh Ameta, Ujjwal Mishra, Praful Hambarde, Amit Shukla
Abstract:
Change detection (CD) in remote sensing aims to identify semantic differences between satellite images captured at different times. While deep learning has significantly advanced this field, existing approaches based on convolutional neural networks (CNNs), transformers and Selective State Space Models (SSMs) still struggle to precisely delineate change regions. In particular, traditional transformer‑based methods suffer from quadratic computational complexity when applied to very high‑resolution (VHR) satellite images and often perform poorly with limited training data, leading to under‑utilization of the rich spatial information available in VHR imagery. We present GRAD‑Former, a novel framework that enhances contextual understanding while maintaining efficiency through reduced model size. The proposed framework consists of a novel encoder with Adaptive Feature Relevance and Refinement (AFRAR) module, fusion and decoder blocks. AFRAR integrates global‑local contextual awareness through two proposed components: the Selective Embedding Amplification (SEA) module and the Global‑Local Feature Refinement (GLFR) module. SEA and GLFR leverage gating mechanisms and differential attention, respectively, which generates multiple softmax heaps to capture important features while minimizing the captured irreverent features. Multiple experiments across three challenging CD datasets (LEVIR‑CD, CDD, DSIFN‑CD) demonstrate GRAD‑Former's superior performance compared to existing approaches. Notably, GRAD‑Former outperforms the current state‑of‑the‑art models across all the metrics and all the datasets while using fewer parameters. Our framework establishes a new benchmark for remote sensing change detection performance. Our code will be released at: https://github.com/Ujjwal238/GRAD‑Former
Authors:Penghao Wang, Siyuan Xie, Hongyu Yan, Xianghui Yang, Jingwei Huang, Chunchao Guo, Jiayuan Gu
Abstract:
Creating interactive digital environments for gaming, robotics, and simulation relies on articulated 3D objects whose functionality emerges from their part geometry and kinematic structure. However, existing approaches remain fundamentally limited: optimization‑based reconstruction methods require slow, per‑object joint fitting and typically handle only simple, single‑joint objects, while retrieval‑based methods assemble parts from a fixed library, leading to repetitive geometry and poor generalization. To address these challenges, we introduce ArtLLM, a novel framework for generating high‑quality articulated assets directly from complete 3D meshes. At its core is a 3D multimodal large language model trained on a large‑scale articulation dataset curated from both existing articulation datasets and procedurally generated objects. Unlike prior work, ArtLLM autoregressively predicts a variable number of parts and joints, inferring their kinematic structure in a unified manner from the object's point cloud. This articulation‑aware layout then conditions a 3D generative model to synthesize high‑fidelity part geometries. Experiments on the PartNet‑Mobility dataset show that ArtLLM significantly outperforms state‑of‑the‑art methods in both part layout accuracy and joint prediction, while generalizing robustly to real‑world objects. Finally, we demonstrate its utility in constructing digital twins, highlighting its potential for scalable robot learning.
Authors:Zhuonan Liang, Wei Guo, Jie Gan, Yaxuan Song, Runnan Chen, Hang Chang, Weidong Cai
Abstract:
Foundation vision models are increasingly adopted in medical image analysis. Due to domain shift, these pretrained models misalign with medical image segmentation needs without being fully fine‑tuned or lightly adapted. We introduce GuiDINO, a framework that repositions native foundation model to acting as a visual guidance generator for downstream segmentation. GuiDINO extracts visual feature representation from DINOv3 and converts them into a spatial guide mask via a lightweight TokenBook mechanism, which aggregates token‑prototype similarities. This guide mask gates feature activations in multiple segmentation backbones, thereby injecting foundation‑model priors while preserving the inductive biases and efficiency of medical dedicated architectures. Training relies on a guide supervision objective loss that aligns the guide mask to ground‑truth regions, optionally augmented by a boundary‑focused hinge loss to sharpen fine structures. GuiDINO also supports parameter‑efficient adaptation through LoRA on the DINOv3 guide backbone. Across diverse medical datasets and nnUNet‑style inference, GuiDINO consistently improves segmentation quality and boundary robustness, suggesting a practical alternative to fine‑tuning and offering a new perspective on how foundation models can best serve medical vision. Code is available at https://github.com/Hi‑FishU/GuiDINO
Authors:Tajamul Ashraf, Abrar Ul Riyaz, Wasif Tak, Tavaheed Tariq, Sonia Yadav, Moloud Abdar, Janibul Bashir
Abstract:
Clinically reliable perception of surgical scenes is essential for advancing intelligent, context‑aware intraoperative assistance such as instrument handoff guidance, collision avoidance, and workflow‑aware robotic support. Existing surgical tool benchmarks primarily evaluate category‑level segmentation, requiring models to detect all instances of predefined instrument classes. However, real‑world clinical decisions often require resolving references to a specific instrument instance based on its functional role, spatial relation, or anatomical interaction capabilities not captured by current evaluation paradigms. We introduce GroundedSurg, the first language‑conditioned, instance‑level surgical grounding benchmark. Each instance pairs a surgical image with a natural‑language description targeting a single instrument, accompanied by structured spatial grounding annotations including bounding boxes and point‑level anchors. The dataset spans ophthalmic, laparoscopic, robotic, and open procedures, encompassing diverse instrument types, imaging conditions, and operative complexities. By jointly evaluating linguistic reference resolution and pixel‑level localization, GroundedSurg enables a systematic and realistic evaluation of vision‑language models in clinically realistic multi‑instrument scenes. Extensive experiments demonstrate substantial performance gaps across modern segmentation and VLMs, highlighting the urgent need for clinically grounded vision‑language reasoning in surgical AI systems. Code and data are publicly available at https://github.com/gaash‑lab/GroundedSurg
Authors:Arctanx An, Shizhao Sun, Danqing Huang, Mingxi Cheng, Yan Gao, Ji Li, Yu Qiao, Jiang Bian
Abstract:
Assessing the aesthetic quality of graphic design is central to visual communication, yet remains underexplored in vision language models (VLMs). We investigate whether VLMs can evaluate design aesthetics in ways comparable to humans. Prior work faces three key limitations: benchmarks restricted to narrow principles and coarse evaluation protocols, a lack of systematic VLM comparisons, and limited training data for model improvement. In this work, we introduce AesEval‑Bench, a comprehensive benchmark spanning four dimensions, twelve indicators, and three fully quantifiable tasks: aesthetic judgment, region selection, and precise localization. Then, we systematically evaluate proprietary, open‑source, and reasoning‑augmented VLMs, revealing clear performance gaps against the nuanced demands of aesthetic assessment. Moreover, we construct a training dataset to fine‑tune VLMs for this domain, leveraging human‑guided VLM labeling to produce task labels at scale and indicator‑grounded reasoning to tie abstract indicators to concrete design regions.Together, our work establishes the first systematic framework for aesthetic quality assessment in graphic design. Our code and dataset will be released at: \hrefhttps://github.com/arctanxarc/AesEval‑Benchhttps://github.com/arctanxarc/AesEval‑Bench
Authors:Xuan Lu, Kangle Li, Haohang Huang, Rui Meng, Wenjun Zeng, Xiaoyu Shen
Abstract:
Recent advances in multimodal large language models (MLLMs) have substantially expanded the capabilities of multimodal retrieval, enabling systems to align and retrieve information across visual and textual modalities. Yet, existing benchmarks largely focus on coarse‑grained or single‑condition alignment, overlooking real‑world scenarios where user queries specify multiple interdependent constraints across modalities. To bridge this gap, we introduce MCMR (Multi‑Conditional Multimodal Retrieval): a large‑scale benchmark designed to evaluate fine‑grained, multi‑condition cross‑modal retrieval under natural‑language queries. MCMR spans five product domains: upper and bottom clothing, jewelry, shoes, and furniture. It also preserves rich long‑form metadata essential for compositional matching. Each query integrates complementary visual and textual attributes, requiring models to jointly satisfy all specified conditions for relevance. We benchmark a diverse suite of MLLM‑based multimodal retrievers and vision‑language rerankers to assess their condition‑aware reasoning abilities. Experimental results reveal: (i) distinct modality asymmetries across models; (ii) visual cues dominate early‑rank precision, while textual metadata stabilizes long‑tail ordering; and (iii) MLLM‑based pointwise rerankers markedly improve fine‑grained matching by explicitly verifying query‑candidate consistency. Overall, MCMR establishes a challenging and diagnostic benchmark for advancing multimodal retrieval toward compositional, constraint‑aware, and interpretable understanding. Our code and dataset is available at https://github.com/EIT‑NLP/MCMR
Authors:Yunguan Fu, Wenjia Bai, Wen Yan, Matthew J Clarkson, Rhodri Huw Davies, Yipeng Hu
Abstract:
Diffusion‑based unsupervised image registration has been explored for cardiac cine MR, but expensive multi‑step inference limits practical use. We propose FlowReg, a flow‑matching framework in displacement field space that achieves strong registration in as few as two steps and supports further refinement with more steps. FlowReg uses warmup‑reflow training: a single‑step network first acts as a teacher, then a student learns to refine from arbitrary intermediate states, removing the need for a pre‑trained model as in existing methods. An Initial Guess strategy feeds back the model prediction as the next starting point, improving refinement from step two onward. On ACDC and MM2 across six tasks (including cross‑dataset generalization), FlowReg outperforms the state of the art on five tasks (+0.6% mean Dice score on average), with the largest gain in the left ventricle (+1.09%), and reduces LVEF estimation error on all six tasks (‑2.58 percentage points), using only 0.7% extra parameters and no segmentation labels. Code is available at https://github.com/mathpluscode/FlowReg.
Authors:Zebin You, Xiaolu Zhang, Jun Zhou, Chongxuan Li, Ji-Rong Wen
Abstract:
We present LLaDA‑o, an effective and length‑adaptive omni diffusion model for multimodal understanding and generation. LLaDA‑o is built on a Mixture of Diffusion (MoD) framework that decouples discrete masked diffusion for text understanding and continuous diffusion for visual generation, while coupling them through a shared, simple, and efficient attention backbone that reduces redundant computation for fixed conditions. Building on MoD, we further introduce a data‑centric length adaptation strategy that enables flexible‑length decoding in multimodal settings without architectural changes. Extensive experiments show that LLaDA‑o achieves state‑of‑the‑art performance among omni‑diffusion models on multimodal understanding and generation benchmarks, and reaches 87.04 on DPG‑Bench for text‑to‑image generation, supporting the effectiveness of unified omni diffusion modeling. Code is available at https://github.com/ML‑GSAI/LLaDA‑o.
Authors:Huanjin Yao, Qixiang Yin, Min Yang, Ziwang Zhao, Yibo Wang, Haotian Luo, Jingyi Zhang, Jiaxing Huang
Abstract:
We aim to develop a multimodal research agent capable of explicit reasoning and planning, multi‑tool invocation, and cross‑modal information synthesis, enabling it to conduct deep research tasks. However, we observe three main challenges in developing such agents: (1) scarcity of search‑intensive multimodal QA data, (2) lack of effective search trajectories, and (3) prohibitive cost of training with online search APIs. To tackle them, we first propose Hyper‑Search, a hypergraph‑based QA generation method that models and connects visual and textual nodes within and across modalities, enabling to generate search‑intensive multimodal QA pairs that require invoking various search tools to solve. Second, we introduce DR‑TTS, which first decomposes search‑involved tasks into several categories according to search tool types, and respectively optimize specialized search tool experts for each tool. It then recomposes tool experts to jointly explore search trajectories via tree search, producing trajectories that successfully solve complex tasks using various search tools. Third, we build an offline search engine supporting multiple search tools, enabling agentic reinforcement learning without using costly online search APIs. With the three designs, we develop MM‑DeepResearch, a powerful multimodal deep research agent, and extensive results shows its superiority across benchmarks. Code is available at https://github.com/HJYao00/MM‑DeepResearch
Authors:Yangyang Xu, Junbo Ke, You-Wei Wen, Chao Wang
Abstract:
Tensor Ring (TR) decomposition is a powerful tool for high‑order data modeling, but is inherently restricted to discrete forms defined on fixed meshgrids. In this work, we propose a TR functional decomposition for both meshgrid and non‑meshgrid data, where factors are parameterized by Implicit Neural Representations (INRs). However, optimizing this continuous framework to capture fine‑scale details is intrinsically difficult. Through a frequency‑domain analysis, we demonstrate that the spectral structure of TR factors determines the frequency composition of the reconstructed tensor and limits the high‑frequency modeling capacity. To mitigate this, we propose a reparameterized TR functional decomposition, in which each TR factor is a structured combination of a learnable latent tensor and a fixed basis. This reparameterization is theoretically shown to improve the training dynamics of TR factor learning. We further derive a principled initialization scheme for the fixed basis and prove the Lipschitz continuity of our proposed model. Extensive experiments on image inpainting, denoising, super‑resolution, and point cloud recovery demonstrate that our method achieves consistently superior performance over existing approaches. Code is available at https://github.com/YangyangXu2002/RepTRFD.
Authors:Zhuolin He, Jiacheng Tang, Jian Pu, Xiangyang Xue
Abstract:
Safe autonomous systems in complex environments require robust road anomaly segmentation to identify unknown obstacles. However, existing approaches often rely on pixel‑level statistics to determine whether a region appears anomalous. This reliance leads to high false‑positive rates on semantically normal background regions such as sky or vegetation, and poor recall of true Out‑of‑distribution (OOD) instances, thereby posing safety risks for robotic perception and decision‑making. To address these challenges, we propose VL‑Anomaly, a vision‑language anomaly segmentation framework that incorporates semantic priors from pre‑trained Vision‑Language Models (VLMs). Specifically, we design a prompt learning‑driven alignment module that adapts Mask2Forme's visual features to CLIP text embeddings of known categories, effectively suppressing spurious anomaly responses in background regions. At inference time, we further introduce a multi‑source inference strategy that integrates text‑guided similarity, CLIP‑based image‑text similarity and detector confidence, enabling more reliable anomaly prediction by leveraging complementary information sources. Extensive experiments demonstrate that VL‑Anomaly achieves state‑of‑the‑art performance on benchmark datasets including RoadAnomaly, SMIYC and Fishyscapes.Code is released on https://github.com/NickHezhuolin/VL‑aligner‑Road‑anomaly‑segment.
Authors:Junbo Ke, Yangyang Xu, You-Wei Wen, Chao Wang
Abstract:
Implicit Neural Representations (INRs) have emerged as a powerful paradigm for various signal processing tasks, but their inherent spectral bias limits the ability to capture high‑frequency details. Existing methods partially mitigate this issue by using Fourier‑based features, which usually rely on fixed frequency bases. This forces multi‑layer perceptrons (MLPs) to inefficiently compose the required frequencies, thereby constraining their representational capacity. To address this limitation, we propose Content‑Aware Frequency Encoding (CAFE), which builds upon Fourier features through multiple parallel linear layers combined via a Hadamard product. CAFE can explicitly and efficiently synthesize a broader range of frequency bases, while the learned weights enable the selection of task‑relevant frequencies. Furthermore, we extend this framework to CAFE+, which incorporates Chebyshev features as a complementary component to Fourier bases. This combination provides a stronger and more stable frequency representation. Extensive experiments across multiple benchmarks validate the effectiveness and efficiency of our approach, consistently achieving superior performance over existing methods. Our code is available at https://github.com/JunboKe0619/CAFE.
Authors:Xuqin Wang, Tao Wu, Yanfeng Zhang, Lu Liu, Mingwei Sun, Yongliang Wang, Niclas Zeller, Daniel Cremers
Abstract:
Recent advances in generative modeling have substantially enhanced novel view synthesis, yet maintaining consistency across viewpoints remains challenging. Diffusion‑based models rely on stochastic noise‑to‑data transitions, which obscure deterministic structures and yield inconsistent view predictions. We advocate a Data‑to‑Data Flow Matching framework that learns deterministic transformations between paired views, enhancing view‑consistent synthesis through explicit data coupling. Building on this, we propose Probability Density Geodesic Flow Matching (PDG‑FM), which aligns interpolation trajectories with density‑based geodesics of a data manifold. To enable tractable geodesic estimation, we employ a teacher‑student framework that distills density‑based geodesic interpolants into an efficient ambient‑space predictor. Empirically, our method surpasses diffusion‑based baselines on Objaverse and GSO30 datasets, demonstrating improved structural coherence and smoother transitions across views. These results highlight the advantages of incorporating data‑dependent geometric regularization into deterministic flow matching for consistent novel view generation.
Authors:Yuze Li, Dong Gong, Xiao Cao, Junchao Yuan, Dongsheng Li, Lei Zhou, Yun Sing Koh, Cheng Yan, Xinyu Zhang
Abstract:
Motion transfer has emerged as a promising direction for controllable video generation, yet existing methods largely focus on single‑object scenarios and struggle when multiple objects require distinct motion patterns. In this work, we present FlexiMMT, the first implicit image‑to‑video (I2V) motion transfer framework that explicitly enables multi‑object, multi‑motion transfer. Given a static multi‑object image and multiple reference videos, FlexiMMT independently extracts motion representations and accurately assigns them to different objects, supporting flexible recombination and arbitrary motion‑to‑object mappings. To address the core challenge of cross‑object motion entanglement, we introduce a Motion Decoupled Mask Attention Mechanism that uses object‑specific masks to constrain attention, ensuring that motion and text tokens only influence their designated regions. We further propose a Differentiated Mask Propagation Mechanism that derives object‑specific masks directly from diffusion attention and progressively propagates them across frames efficiently. Extensive experiments demonstrate that FlexiMMT achieves precise, compositional, and state‑of‑the‑art performance in I2V‑based multi‑object multi‑motion transfer. Our project page is: https://ethan‑li123.github.io/FlexiMMT_page/
Authors:Wenxiang Jiang, Yujun Lan, Shuo Zhao, Yuanshan Liu, Mingzhu Zhou, Jinxin Wang
Abstract:
Recently, Instant Neural Graphics Primitives (Instant‑NGP) has achieved significant success in rapid 3D scene reconstruction, but securely embedding high‑capacity hidden data, such as an entire 3D scene, remains a challenge. Existing methods rely on external decoders, require architectural modifications, and suffer from limited capacity, which makes them easily detectable. We propose a novel parameter‑free 3D Cryptographic Steganography using Instant‑NGP (StegoNGP), which leverages the Instant‑NGP hash encoding function as a key‑controlled scene switcher. By associating a default key with a cover scene and a secret key with a hidden scene, our method trains a single model to interweave both representations within the same network weights. The resulting model is indistinguishable from a standard Instant‑NGP in architecture and parameter count. We also introduce an enhanced Multi‑Key scheme, which assigns multiple independent keys across hash levels, dramatically expanding the key space and providing high robustness against partial key disclosure attacks. Experimental results demonstrated that StegoNGP can hide a complete high‑quality 3D scene with strong imperceptibility and security, providing a new paradigm for high‑capacity, undetectable information hiding in neural fields. The code can be found at https://github.com/jiang‑wenxiang/StegoNGP.
Authors:Zhenchen Wan, Ce Chen, Runqi Lin, Jiaxin Huang, Tianxi Chen, Yanwu Xu, Tongliang Liu, Mingming Gong
Abstract:
Virtual try‑on (VTON) has recently achieved impressive visual fidelity, but most existing systems require uploading personal photos to cloud‑based GPUs, raising privacy concerns and limiting on‑device deployment. To address this, we present Mobile‑VTON, a high‑quality, privacy‑preserving framework that enables fully offline virtual try‑on on commodity mobile devices using only a single user image and a garment image. Mobile‑VTON introduces a modular TeacherNet‑GarmentNet‑TryonNet (TGT) architecture that integrates knowledge distillation, garment‑conditioned generation, and garment alignment into a unified pipeline optimized for on‑device efficiency. Within this framework, we propose a Feature‑Guided Adversarial (FGA) Distillation strategy that combines teacher supervision with adversarial learning to better match real‑world image distributions. GarmentNet is trained with a trajectory‑consistency loss to preserve garment semantics across diffusion steps, while TryonNet uses latent concatenation and lightweight cross‑modal conditioning to enable robust garment‑to‑person alignment without large‑scale pretraining. By combining these components, Mobile‑VTON achieves high‑fidelity generation with low computational overhead. Experiments on VITON‑HD and DressCode at 1024 x 768 show that it matches or outperforms strong server‑based baselines while running entirely offline. These results demonstrate that high‑quality VTON is not only feasible but also practical on‑device, offering a secure solution for real‑world applications. Code and project page are available at https://zhenchenwan.github.io/Mobile‑VTON/.
Authors:Zhiye Wang, Yanbo Jiang, Rui Zhou, Bo Zhang, Fang Zhang, Zhenhua Xu, Yaqin Zhang, Jianqiang Wang
Abstract:
Large language models (LLMs) have shown great promise for autonomous driving. However, discretizing numbers into tokens limits precise numerical reasoning, fails to reflect the positional significance of digits in the training objective, and makes it difficult to achieve both decoding efficiency and numerical precision. These limitations affect both the processing of sensor measurements and the generation of precise control commands, creating a fundamental barrier for deploying LLM‑based autonomous driving systems. In this paper, we introduce DriveCode, a novel numerical encoding method that represents numbers as dedicated embeddings rather than discrete text tokens. DriveCode employs a number projector to map numbers into the language model's hidden space, enabling seamless integration with visual and textual features in a unified multimodal sequence. Evaluated on OmniDrive, DriveGPT4, and DriveGPT4‑V2 datasets, DriveCode demonstrates superior performance in trajectory prediction and control signal generation, confirming its effectiveness for LLM‑based autonomous driving systems.
Authors:Seungwook Kim, Minsu Cho
Abstract:
Text‑to‑image generation powers content creation across design, media, and data augmentation. Post‑training of text‑to‑image generative models is a promising path to improve human preference alignment, factuality, and aesthetics. We introduce SOLACE (Self‑Originating LAtent Confidence Estimation), a post‑training framework that replaces external reward supervision with an internal self‑confidence signal: we re‑noise the model's own outputs and measure how accurately it recovers the injected noise, treating low reconstruction error as high self‑confidence. SOLACE converts this intrinsic signal into scalar rewards for reinforcement learning, requiring no external reward models, annotators, or preference data. By reinforcing high‑confidence generations, SOLACE delivers consistent gains in compositional generation, text rendering, and text‑image alignment. Integrating SOLACE with external rewards yields complementary improvements while alleviating reward hacking.
Authors:Yang Cao, Feize Wu, Dave Zhenyu Chen, Yingji Zhong, Lanqing Hong, Dan Xu
Abstract:
Current multi‑view indoor 3D object detectors rely on sensor geometry that is costly to obtain (i.e., precisely calibrated multi‑view camera poses) to fuse multi‑view information into a global scene representation, limiting deployment in real‑world scenes. We target a more practical setting: Sensor‑Geometry‑Free (SG‑Free) multi‑view indoor 3D object detection, where there are no sensor‑provided geometric inputs (multi‑view poses or depth). Recent Visual Geometry Grounded Transformer (VGGT) shows that strong 3D cues can be inferred directly from images. Building on this insight, we present VGGT‑Det, the first framework tailored for SG‑Free multi‑view indoor 3D object detection. Rather than merely consuming VGGT predictions, our method integrates VGGT encoder into a transformer‑based pipeline. To effectively leverage both the semantic and geometric priors from inside VGGT, we introduce two novel key components: (i) Attention‑Guided Query Generation (AG): exploits VGGT attention maps as semantic priors to initialize object queries, improving localization by focusing on object regions while preserving global spatial structure; (ii) Query‑Driven Feature Aggregation (QD): a learnable See‑Query interacts with object queries to 'see' what they need, and then dynamically aggregates multi‑level geometric features across VGGT layers that progressively lift 2D features into 3D. Experiments show that VGGT‑Det significantly surpasses the best‑performing method in the SG‑Free setting by 4.4 and 8.6 mAP@0.25 on ScanNet and ARKitScenes, respectively. Ablation study shows that VGGT's internally learned semantic and geometric priors can be effectively leveraged by our AG and QD.
Authors:Puyun Wang, Kaimin Yu, Huayang He, Feng Huang, Xianyu Wu, Yating Chen
Abstract:
Underwater optical imaging is severely hindered by scattering, but polarization imaging offers the unique dual advantages of descattering and shape‑from‑polarization (SfP) 3D reconstruction. To exploit these advantages, this paper proposes UD‑SfPNet, an underwater descattering shape‑from‑polarization network that leverages polarization cues for improved 3D surface normal prediction. The framework jointly models polarization‑based image descattering and SfP normal estimation in a unified pipeline, avoiding error accumulation from sequential processing and enabling global optimization across both tasks. UD‑SfPNet further incorporates a novel color embedding module to enhance geometric consistency by exploiting the relationship between color encodings and surface orientation. A detail enhancement convolution module is also included to better preserve high‑frequency geometric details that are lost under scattering. Experiments on the MuS‑Polar3D dataset show that the proposed method significantly improves reconstruction accuracy, achieving a mean surface normal angular error of 15.12^\circ (the lowest among compared methods). These results confirm the efficacy of combining descattering with polarization‑based shape inference, and highlight the practical significance and potential applications of UD‑SfPNet for optical 3D imaging in challenging underwater environments. The code is available at https://github.com/WangPuyun/UD‑SfPNet.
Authors:Xiaolong Zeng, Yitong Yu, Shiyao Xiong, Jinhua Hao, Ming Sun, Chao Zhou, Bin Wang
Abstract:
Look‑Up Table based methods have emerged as a promising direction for efficient image restoration tasks. Recent LUT‑based methods focus on improving their performance by expanding the receptive field. However, they inevitably introduce extra computational and storage overhead, which hinders their deployment in edge devices. To address this issue, we propose ShiftLUT, a novel framework that attains the largest receptive field among all LUT‑based methods while maintaining high efficiency. Our key insight lies in three complementary components. First, Learnable Spatial Shift module (LSS) is introduced to expand the receptive field by applying learnable, channel‑wise spatial offsets on feature maps. Second, we propose an asymmetric dual‑branch architecture that allocates more computation to the information‑dense branch, substantially reducing inference latency without compromising restoration quality. Finally, we incorporate a feature‑level LUT compression strategy called Error‑bounded Adaptive Sampling (EAS) to minimize the storage overhead. Compared to the previous state‑of‑the‑art method TinyLUT, ShiftLUT achieves a 3.8× larger receptive field and improves an average PSNR by over 0.21 dB across multiple standard benchmarks, while maintaining a small storage size and inference time. The code is available at: https://github.com/Sailor‑t/ShiftLUT .
Authors:Longmi Gao, Pan Gao
Abstract:
Volume Electron Microscopy (VEM) is crucial for 3D tissue imaging but often produces anisotropic data with poor axial resolution, hindering visualization and downstream analysis. Existing methods for isotropic reconstruction often suffer from neglecting abundant axial information and employing simple downsampling to simulate anisotropic data. To address these limitations, we propose VEMamba, an efficient framework for isotropic reconstruction. The core of VEMamba is a novel 3D Dependency Reordering paradigm, implemented via two key components: an Axial‑Lateral Chunking Selective Scan Module (ALCSSM), which intelligently re‑maps complex 3D spatial dependencies (both axial and lateral) into optimized 1D sequences for efficient Mamba‑based modeling, explicitly enforcing axial‑lateral consistency; and a Dynamic Weights Aggregation Module (DWAM) to adaptively aggregate these reordered sequence outputs for enhanced representational power. Furthermore, we introduce a realistic degradation simulation and then leverage Momentum Contrast (MoCo) to integrate this degradation‑aware knowledge into the network for superior reconstruction. Extensive experiments on both simulated and real‑world anisotropic VEM datasets demonstrate that VEMamba achieves highly competitive performance across various metrics while maintaining a lower computational footprint. The source code is available on GitHub: https://github.com/I2‑Multimedia‑Lab/VEMamba
Authors:Yu Luo, Guangyu Wei, Yangfan Li, Jieyu He, Yueming Lyu
Abstract:
Segmentation of the main coronary artery from X‑ray coronary angiography (XCA) sequences is crucial for the diagnosis of coronary artery diseases. However, this task is challenging due to issues such as blurred boundaries, inconsistent radiation contrast, complex motion patterns, and a lack of annotated images for training. Although Semi‑Supervised Learning (SSL) can alleviate the annotation burden, conventional methods struggle with complicated temporal dynamics and unreliable uncertainty quantification. To address these challenges, we propose SAM3‑based Teacher‑student framework with Motion‑Aware consistency and Progressive Confidence Regularization (SMART), a semi‑supervised vessel segmentation approach for X‑ray angiography videos. First, our method utilizes SAM3's unique promptable concept segmentation design and innovates a SAM3‑based teacher‑student framework to maximize the performance potential of both the teacher and the student. Second, we enhance segmentation by integrating the vessel mask warping technique and motion consistency loss to model complex vessel dynamics. To address the issue of unreliable teacher predictions caused by blurred boundaries and minimal contrast, we further propose a progressive confidence‑aware consistency regularization to mitigate the risk of unreliable outputs. Extensive experiments on three datasets of XCA sequences from different institutions demonstrate that SMART achieves state‑of‑the‑art performance while requiring significantly fewer annotations, making it particularly valuable for real‑world clinical applications where labeled data is scarce. Our code is available at: https://github.com/qimingfan10/SMART.
Authors:Cong Wang, Jinshan Pan, Liyan Wang, Wei Wang, Yang Yang
Abstract:
We propose a simple yet effective UHDPromer, a neural discrimination‑prompted Transformer, for Ultra‑High‑Definition (UHD) image restoration and enhancement. Our UHDPromer is inspired by an interesting observation that there implicitly exist neural differences between high‑resolution and low‑resolution features, and exploring such differences can facilitate low‑resolution feature representation. To this end, we first introduce Neural Discrimination Priors (NDP) to measure the differences and then integrate NDP into the proposed Neural Discrimination‑Prompted Attention (NDPA) and Neural Discrimination‑Prompted Network (NDPN). The proposed NDPA re‑formulates the attention by incorporating NDP to globally perceive useful discrimination information, while the NDPN explores a continuous gating mechanism guided by NDP to selectively permit the passage of beneficial content. To enhance the quality of restored images, we propose a super‑resolution‑guided reconstruction approach, which is guided by super‑resolving low‑resolution features to facilitate final UHD image restoration. Experiments show that UHDPromer achieves the best computational efficiency while still maintaining state‑of‑the‑art performance on 3 UHD image restoration and enhancement tasks, including low‑light image enhancement, image dehazing, and image deblurring. The source codes and pre‑trained models will be made available at https://github.com/supersupercong/uhdpromer.
Authors:Amir Belder, Ayellet Tal
Abstract:
In recent years, various methods have been proposed for mesh analysis, each offering distinct advantages and often excelling on different object classes. We present a novel Mixture of Experts (MoE) framework designed to harness the complementary strengths of these diverse approaches. We propose a new gate architecture that encourages each expert to specialise in the classes it excels in. Our design is guided by two key ideas: (1) random walks over the mesh surface effectively capture the regions that individual experts attend to, and (2) an attention mechanism that enables the gate to focus on the areas most informative for each expert's decision‑making. To further enhance performance, we introduce a dynamic loss balancing scheme that adjusts a trade‑off between diversity and similarity losses throughout the training, where diversity prompts expert specialization, and similarity enables knowledge sharing among the experts. Our framework achieves state‑of‑the‑art results in mesh classification, retrieval, and semantic segmentation tasks. Our code is available at: https://github.com/amirbelder/MME‑Mixture‑of‑Mesh‑Experts.
Authors:Seemandhar Jain, Keshav Gupta, Kunal Gupta, Manmohan Chandraker
Abstract:
The proliferation of neural radiance field (NeRF) research requires significant efforts to reimplement papers before building upon them. We introduce NERFIFY, a multi‑agent framework that reliably converts NeRF research papers into trainable Nerfstudio plugins, in contrast to generic paper‑to‑code methods and frontier models like GPT‑5 that usually fail to produce runnable code. NERFIFY achieves domain‑specific executability through six key innovations: (1) Context‑free grammar (CFG): LLM synthesis is constrained by Nerfstudio formalized as a CFG, ensuring generated code satisfies architectural invariants. (2) Graph‑of‑Thought code synthesis: Specialized multi‑file‑agents generate repositories in topological dependency order, validating contracts and errors at each node. (3) Compositional citation recovery: Agents automatically retrieve and integrate components (samplers, encoders, proposal networks) from citation graphs of references. (4) Visual feedback: Artifacts are diagnosed through PSNR‑minima ROI analysis, cross‑view geometric validation, and VLM‑guided patching to iteratively improve quality. (5) Knowledge enhancement: Beyond reproduction, methods can be improved with novel optimizations. (6) Benchmarking: An evaluation framework is designed for NeRF paper‑to‑code synthesis across 30 diverse papers. On papers without public implementations, NERFIFY achieves visual quality matching expert human code (+/‑0.5 dB PSNR, +/‑0.2 SSIM) while reducing implementation time from weeks to minutes. NERFIFY demonstrates that a domain‑aware design enables code translation for complex vision papers, potentiating accelerated and democratized reproducible research. Code, data and implementations will be publicly released.
Authors:Zikang Xu, Ruinan Jin, Xiaoxiao Li
Abstract:
Fairness in medical agents is becoming critical as tool‑using clinical AI systems orchestrate specialized vision and language modules for tasks such as chest X‑ray question answering. While these medical AI agents can improve flexibility, their added pipeline complexity also creates new pathways for demographic bias beyond standalone models. We present DUCK, Decomposing Unfairness in Chest X‑ray agents, a systematic audit of fairness in tool‑using chest X‑ray agents instantiated with MedRAX. To localize where disparities arise, we introduce a stage‑wise fairness decomposition that separates end‑to‑end bias from three agent‑specific sources: tool exposure bias, or utility gaps conditioned on tool presence; tool transition bias, or subgroup differences in tool‑routing patterns; and model reasoning bias, or subgroup differences in synthesis behaviors. Extensive experiments on tool‑using agentic frameworks across five driver backbones reveal that demographic gaps persist in end‑to‑end performance, with equalized odds up to 20.79% and the lowest fairness‑utility tradeoff down to 28.65%. Intermediate behaviors, including tool usage, transition patterns, and reasoning traces, exhibit distinct subgroup disparities that are not predictable from end‑to‑end evaluation alone. For example, conditioned on segmentation‑tool availability, the subgroup utility gap reaches as high as 50%. Our findings underscore the need for process‑level fairness auditing and debiasing to ensure the equitable deployment of clinical agentic systems. Code: https://github.com/Nanboy‑Ronan/DUCK.
Authors:Zhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya Wang
Abstract:
Planning physically feasible dexterous hand manipulation is a central challenge in robotic manipulation and Embodied AI. Prior work typically relies on object‑centric cues or precise hand‑object interaction sequences, foregoing the rich, compositional guidance of open‑vocabulary instruction. We introduce UniHM, the first framework for unified dexterous hand manipulation guided by free‑form language commands. We propose a Unified Hand‑Dexterous Tokenizer that maps heterogeneous dexterous‑hand morphologies into a single shared codebook, improving cross‑dexterous hand generalization and scalability to new morphologies. Our vision language action model is trained solely on human‑object interaction data, eliminating the need for massive real‑world teleoperation datasets, and demonstrates strong generalizability in producing human‑like manipulation sequences from open‑ended language instructions. To ensure physical realism, we introduce a physics‑guided dynamic refinement module that performs segment‑wise joint optimization under generative and temporal priors, yielding smooth and physically feasible manipulation sequences. Across multiple datasets and real‑world evaluations, UniHM attains state‑of‑the‑art results on both seen and unseen objects and trajectories, demonstrating strong generalization and high physical feasibility. Our project page at \hrefhttps://unihm.github.io/https://unihm.github.io/.
Authors:Yihui Li, Chengxin Lv, Zichen Tang, Hongyu Yang, Di Huang
Abstract:
We present TokenSplat, a feed‑forward framework for joint 3D Gaussian reconstruction and camera pose estimation from unposed multi‑view images. At its core, TokenSplat introduces a Token‑aligned Gaussian Prediction module that aligns semantically corresponding information across views directly in the feature space. Guided by coarse token positions and fusion confidence, it aggregates multi‑scale contextual features to enable long‑range cross‑view reasoning and reduce redundancy from overlapping Gaussians. To further enhance pose robustness and disentangle viewpoint cues from scene semantics, TokenSplat employs learnable camera tokens and an Asymmetric Dual‑Flow Decoder (ADF‑Decoder) that enforces directionally constrained communication between camera and image tokens. This maintains clean factorization within a feed‑forward architecture, enabling coherent reconstruction and stable pose estimation without iterative refinement. Extensive experiments demonstrate that TokenSplat achieves higher reconstruction fidelity and novel‑view synthesis quality in pose‑free settings, and significantly improves pose estimation accuracy compared to prior pose‑free methods. Project page: https://kidleyh.github.io/tokensplat/.
Authors:Guoquan Wei, Liu Shi, Shaoyu Wang, Mohan Li, Cunfeng Wei, Qiegen Liu
Abstract:
Noise and artifacts during computed tomography (CT) scans are a fundamental challenge affecting disease diagnosis. However, current methods either involve excessively long reconstruction times or rely on data‑driven models for optimization, failing to adequately consider the valuable information inherent in the data itself, especially medical 3D data. This work proposes a reconstruction method under ultra‑low raw data conditions, requiring no external data and avoiding lengthy pre‑training processes. By leveraging spatial nonlocal similarity and the conjugate properties of the projection domain to generate pseudo‑3D data for self‑supervised training, high‑fidelity results can be achieved in a very short time. Extensive experiments demonstrate that this method not only mitigates detector‑induced ring artifacts but also exhibits unprecedented capabilities in detail recovery. This method provides a new paradigm for research using unlabeled raw projection data. Code is available at https://github.com/yqx7150/SCOUT.
Authors:Yushan Han, Hui Zhang, Qiming Xia, Yi Jin, Yidong Li
Abstract:
Collaborative perception empowers autonomous agents to share complementary information and overcome perception limitations. While early fusion offers more perceptual complementarity and is inherently robust to model heterogeneity, its high communication cost has limited its practical deployment, prompting most existing works to favor intermediate or late fusion. To address this, we propose a communication‑efficient early Collaborative perception framework that incorporates LiDAR Completion to restore scene completeness under sparse transmission, dubbed as CoLC. Specifically, the CoLC integrates three complementary designs. First, each neighbor agent applies Foreground‑Aware Point Sampling (FAPS) to selectively transmit informative points that retain essential structural and contextual cues under bandwidth constraints. The ego agent then employs Completion‑Enhanced Early Fusion (CEEF) to reconstruct dense pillars from the received sparse inputs and adaptively fuse them with its own observations, thereby restoring spatial completeness. Finally, the Dense‑Guided Dual Alignment (DGDA) strategy enforces semantic and geometric consistency between the enhanced and dense pillars during training, ensuring consistent and robust feature learning. Experiments on both simulated and real‑world datasets demonstrate that CoLC achieves superior perception‑communication trade‑offs and remains robust under heterogeneous model settings. The code is available at https://github.com/CatOneTwo/CoLC.
Authors:Ying Liu, Yudong Han, Kean Shi, Liyuan Pan
Abstract:
Multimodal Large Language Models (MLLMs) have achieved remarkable performance by aligning pretrained visual representations with the linguistic knowledge embedded in Large Language Models (LLMs). However, existing approaches typically rely on final‑layer visual features or learnable multi‑layer fusion, which often fail to sufficiently exploit hierarchical visual cues without explicit cross‑layer interaction design. In this work, we propose a Memory‑Augmented Adapter (Mema) within the vision encoder. Specifically, Mema maintains a stateful memory that accumulates hierarchical visual representations across layers, with its evolution conditioned on both query embeddings and step‑wise visual features. A portion of this memory is selectively injected into token representations via a feedback mechanism, thereby mitigating the attenuation of fine‑grained visual cues from shallow layers. Designed as a lightweight and plug‑and‑play module, Mema integrates seamlessly into pretrained vision encoders without modifying the vanilla backbone architecture. Only a minimal set of additional parameters requires training, enabling adaptive visual feature refinement while reducing training overhead. Extensive experiments across multiple benchmarks demonstrate that Mema consistently improves performance, validating its effectiveness in complex multimodal reasoning tasks. The code have been released at https://github.com/Sisiliu312/Mema.
Authors:Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Zhiqiang Shen
Abstract:
Dataset pruning has been widely studied for 2D images to remove redundancy and accelerate training, while particular pruning methods for 3D data remain largely unexplored. In this work, we study dataset pruning for 3D data, where its observed common long‑tail class distribution nature make optimization under conventional evaluation metrics Overall Accuracy (OA) and Mean Accuracy (mAcc) inherently conflicting, and further make pruning particularly challenging. To address this, we formulate pruning as approximating the full‑data expected risk with a weighted subset, which reveals two key errors: coverage error from insufficient representativeness and prior‑mismatch bias from inconsistency between subset‑induced class weights and target metrics. We propose representation‑aware subset selection with per‑class retention quotas for long‑tail coverage, and prior‑invariant teacher supervision using calibrated soft labels and embedding‑geometry distillation. The retention quota also serves as a switch to control the OA‑mAcc trade‑off. Extensive experiments on 3D datasets show that our method can improve both metrics across multiple settings while adapting to different downstream preferences. Our code is available at https://github.com/XiaohanZhao123/3D‑Dataset‑Pruning.
Authors:Zhanwang Liu, Yuting Li, Haoyuan Gao, Yexin Li, Linghe Kong, Lichao Sun, Weiran Huang
Abstract:
Catastrophic forgetting, the tendency of neural networks to forget previously learned knowledge when learning new tasks, has been a major challenge in continual learning (CL). To tackle this challenge, CL methods have been proposed and shown to reduce forgetting. Furthermore, CL models deployed in mission‑critical settings can benefit from uncertainty awareness by calibrating their predictions to reliably assess their confidences. However, existing uncertainty‑aware continual learning methods suffer from high computational overhead and incompatibility with mainstream replay methods. To address this, we propose idempotent experience replay (IDER), a novel approach based on the idempotent property where repeated function applications yield the same output. Specifically, we first adapt the training loss to make model idempotent on current data streams. In addition, we introduce an idempotence distillation loss. We feed the output of the current model back into the old checkpoint and then minimize the distance between this reprocessed output and the original output of the current model. This yields a simple and effective new baseline for building reliable continual learners, which can be seamlessly integrated with other CL approaches. Extensive experiments on different CL benchmarks demonstrate that IDER consistently improves prediction reliability while simultaneously boosting accuracy and reducing forgetting. Our results suggest the potential of idempotence as a promising principle for deploying efficient and trustworthy continual learning systems in real‑world applications.Our code is available at https://github.com/YutingLi0606/Idempotent‑Continual‑Learning.
Authors:Lijing Cai, Zhan Shi, Chenglong Huang, Jinyao Wu, Qiping Li, Zikang Huo, Linsen Chen, Chongde Zi, Xun Cao
Abstract:
Recently, Spectral Compressive Imaging (SCI) has achieved remarkable success, unlocking significant potential for dynamic spectral vision. However, existing reconstruction methods, primarily image‑based, suffer from two limitations: (i) Encoding process masks spatial‑spectral features, leading to uncertainty in reconstructing missing information from single compressed measurements, and (ii) The frame‑by‑frame reconstruction paradigm fails to ensure temporal consistency, which is crucial in the video perception. To address these challenges, this paper seeks to advance spectral reconstruction from the image level to the video level, leveraging the complementary features and temporal continuity across adjacent frames in dynamic scenes. Initially, we construct the first high‑quality dynamic hyperspectral image dataset (DynaSpec), comprising 30 sequences obtained through frame‑scanning acquisition. Subsequently, we propose the Propagation‑Guided Spectral Video Reconstruction Transformer (PG‑SVRT), which employs a spatial‑then‑temporal attention to effectively reconstruct spectral features from abundant video information, while using a bridged token to reduce computational complexity. Finally, we conduct simulation experiments to assess the performance of four SCI systems, and construct a DD‑CASSI prototype for real‑world data collection and benchmarking. Extensive experiments demonstrate that PG‑SVRT achieves superior performance in reconstruction quality, spectral fidelity, and temporal consistency, while maintaining minimal FLOPs. Project page: https://github.com/nju‑cite/DynaSpec
Authors:Changxing Liu, Zichen Chao, Siheng Chen
Abstract:
Collaborative perception leverages data exchange among multiple agents to enhance overall perception capabilities. However, heterogeneity across agents introduces domain gaps that hinder collaboration, and this is further exacerbated by an underexplored issue: modality isolation. It arises when multiple agents with different modalities never co‑occur in any training data frame, enlarging cross‑modal domain gaps. Existing alignment methods rely on supervision from spatially overlapping observations, thus fail to handle modality isolation. To address this challenge, we propose CodeAlign, the first efficient, co‑occurrence‑free alignment framework that smoothly aligns modalities via cross‑modal feature‑code‑feature(FCF) translation. The key idea is to explicitly identify the representation consistency through codebook, and directly learn mappings between modality‑specific feature spaces, thereby eliminating the need for spatial correspondence. Codebooks regularize feature spaces into code spaces, providing compact yet expressive representations. With a prepared code space for each modality, CodeAlign learns FCF translations that map features to the corresponding codes of other modalities, which are then decoded back into features in the target code space, enabling effective alignment. Experiments show that, when integrating three modalities, CodeAlign requires only 8% of the training parameters of prior alignment methods, reduces communication load by 1024x, and achieves state‑of‑the‑art perception performance on both OPV2V and DAIR‑V2X dataset. Code will be released on https://github.com/cxliu0314/CodeAlign.
Authors:Keiller Nogueira, Codrut-Andrei Diaconu, Dávid Kerekes, Jakob Gawlikowski, Cédric Léonard, Nassim Ait Ali Braham, June Moh Goo, Zichao Zeng, Zhipeng Liu, Pallavi Jain, Andrea Nascetti, Ronny Hänsch
Abstract:
High‑quality pixel‑level annotations are essential for the semantic segmentation of remote sensing imagery. However, such labels are expensive to obtain and often affected by noise due to the labor‑intensive and time‑consuming nature of pixel‑wise annotation, which makes it challenging for human annotators to label every pixel accurately. Annotation errors can significantly degrade the performance and robustness of modern segmentation models, motivating the need for reliable mechanisms to identify and quantify noisy training samples. This paper introduces a novel Data‑Centric benchmark, together with a novel, publicly available dataset and two techniques for identifying, quantifying, and ranking training samples according to their level of label noise in remote sensing semantic segmentation. Such proposed methods leverage complementary strategies based on model uncertainty, prediction consistency, and representation analysis, and consistently outperform established baselines across a range of experimental settings. The outcomes of this work are publicly available at https://github.com/keillernogueira/label_noise_segmentation.
Authors:Yuchen Hou, Lin Zhao
Abstract:
Vision‑Language‑Action (VLA) models achieve over 95% success on standard benchmarks. However, through systematic experiments, we find that current state‑of‑the‑art VLA models largely ignore language instructions. Prior work lacks: (1) systematic semantic perturbation diagnostics, (2) a benchmark that forces language understanding by design, and (3) linguistically diverse training data.
This paper constructs the LangGap benchmark, based on a four‑dimensional semantic perturbation method ‑‑ varying instruction semantics while keeping the tabletop layout fixed ‑‑ revealing language understanding deficits in π0.5. Existing benchmarks like LIBERO assign only one task per layout, underutilizing available objects and target locations; LangGap fully diversifies pick‑and‑place tasks under identical layouts, forcing models to truly understand language.
Experiments show that targeted data augmentation can partially close the language gap ‑‑ success rate improves from 0% to 90% with single‑task training, and 0% to 28% with multi‑task training. However, as semantic diversity of extended tasks increases, model learning capacity proves severely insufficient; even trained tasks perform poorly. This reveals a fundamental challenge for VLA models in understanding diverse language instructions ‑‑ precisely the long‑term value of LangGap.
Authors:Rongsheng Wang, Minghao Wu, Hongru Zhou, Zhihan Yu, Zhenyang Cai, Junying Chen, Benyou Wang
Abstract:
Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phenomena remains largely unexplored. Microscale simulation holds great promise for biomedical applications such as drug discovery, organ‑on‑chip systems, and disease mechanism studies, while also showing potential in education and interactive visualization. In this work, we introduce MicroWorldBench, a multi‑level rubric‑based benchmark for microscale simulation tasks. MicroWorldBench enables systematic, rubric‑based evaluation through 459 unique expert‑annotated criteria spanning multiple microscale simulation task (e.g., organ‑level processes, cellular dynamics, and subcellular molecular interactions) and evaluation dimensions (e.g., scientific fidelity, visual quality, instruction following). MicroWorldBench reveals that current SOTA video generation models fail in microscale simulation, showing violations of physical laws, temporal inconsistency, and misalignment with expert criteria. To address these limitations, we construct MicroSim‑10K, a high‑quality, expert‑verified simulation dataset. Leveraging this dataset, we train MicroVerse, a video generation model tailored for microscale simulation. MicroVerse can accurately reproduce complex microscale mechanism. Our work first introduce the concept of Micro‑World Simulation and present a proof of concept, paving the way for applications in biology, education, and scientific visualization. Our work demonstrates the potential of educational microscale simulations of biological mechanisms. Our data and code are publicly available at https://github.com/FreedomIntelligence/MicroVerse
Authors:Yilian Liu, Xiaojun Jia, Guoshun Nan, Jiuyang Lyu, Zhican Chen, Tao Guan, Shuyuan Luo, Zhongyi Zhai, Yang Liu
Abstract:
Multimodal Large Language Models (MLLMs) have achieved remarkable performance but remain vulnerable to jailbreak attacks that can induce harmful content and undermine their secure deployment. Previous studies have shown that introducing additional inference steps, which disrupt security attention, can make MLLMs more susceptible to being misled into generating malicious content. However, these methods rely on single‑image masking or isolated visual cues, which only modestly extend reasoning paths and thus achieve limited effectiveness, particularly against strongly aligned commercial closed‑source models. To address this problem, in this paper, we propose Multi‑Image Dispersion and Semantic Reconstruction (MIDAS), a multimodal jailbreak framework that decomposes harmful semantics into risk‑bearing subunits, disperses them across multiple visual clues, and leverages cross‑image reasoning to gradually reconstruct the malicious intent, thereby bypassing existing safety mechanisms. The proposed MIDAS enforces longer and more structured multi‑image chained reasoning, substantially increases the model's reliance on visual cues while delaying the exposure of malicious semantics and significantly reducing the model's security attention, thereby improving the performance of jailbreak against advanced MLLMs. Extensive experiments across different datasets and MLLMs demonstrate that the proposed MIDAS outperforms state‑of‑the‑art jailbreak attacks for MLLMs and achieves an average attack success rate of 81.46% across 4 closed‑source MLLMs. Our code is available at this [link](https://github.com/Winnie‑Lian/MIDAS).
Authors:Ke Cao, Xuanhua He, Xueheng Li, Lingting Zhu, Yingying Wang, Ao Ma, Zhanjie Zhang, Man Zhou, Chengjun Xie, Jie Zhang
Abstract:
Pansharpening aims to generate high‑resolution multi‑spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low‑resolution MS data. However, most existing methods are evaluated under limited, low‑resolution settings, limiting their generalization to real‑world, high‑resolution scenarios. To bridge this gap, we systematically investigate the data, algorithmic, and computational challenges of cross‑scale pansharpening. We first introduce PanScale, the first large‑scale, cross‑scale pansharpening dataset, accompanied by PanScale‑Bench, a comprehensive benchmark for evaluating generalization across varying resolutions and scales. To realize scale generalization, we propose ScaleFormer, a novel architecture designed for multi‑scale pansharpening. ScaleFormer reframes generalization across image resolutions as generalization across sequence lengths: it tokenizes images into patch sequences of the same resolution but variable length proportional to image scale. A Scale‑Aware Patchify module enables training for such variations from fixed‑size crops. ScaleFormer then decouples intra‑patch spatial feature learning from inter‑patch sequential dependency modeling, incorporating Rotary Positional Encoding to enhance extrapolation to unseen scales. Extensive experiments show that our approach outperforms SOTA methods in fusion quality and cross‑scale generalization. The datasets and source code are available at https://github.com/caoke‑963/ScaleFormer.
Authors:Xianhao Zhou, Jianghao Wu, Lanfeng Zhong, Ku Zhao, Jinlong He, Shaoting Zhang, Guotai Wang
Abstract:
Cone‑beam CT (CBCT) is routinely acquired in radiotherapy but suffers from severe artifacts and unreliable Hounsfield Unit (HU) values, limiting its direct use for dose calculation. Synthetic CT (sCT) generation from CBCT is therefore an important task, yet paired CBCT‑‑CT data are often unavailable or unreliable due to temporal gaps, anatomical variation, and registration errors. In this work, we introduce rectified flow (RF) into unpaired CBCT‑to‑CT translation in medical imaging. Although RF is theoretically compatible with unpaired learning through distribution‑level coupling and deterministic transport, its practical effectiveness under small medical datasets and limited batch sizes remains underexplored. Direct application with random or batch‑local pseudo pairing can produce unstable supervision due to semantically mismatched endpoint samples. To address this challenge, we propose Retrieval‑Augmented Flow Matching (RAFM), which adapts RF to the medical setting by constructing retrieval‑guided pseudo pairs using a frozen DINOv3 encoder and a global CT memory bank. This strategy improves empirical coupling quality and stabilizes unpaired flow‑based training. Experiments on SynthRAD2023 under a strict subject‑level true‑unpaired protocol show that RAFM outperforms existing methods across FID, MAE, SSIM, PSNR, and SegScore. The code is available at https://github.com/HiLab‑git/RAFM.git.
Authors:Yuyang Chen, Linqian Zeng, Yijin ZHou, Hengjie Li, Jidong Zhai
Abstract:
Diffusion models have achieved remarkable success in generative AI, yet their computational efficiency remains a significant challenge, particularly for Diffusion Transformers (DiTs) requiring intensive full‑attention computation. While existing acceleration approaches focus on content‑agnostic uniform optimization strategies, we observe that different regions in generated content exhibit heterogeneous convergence patterns during the denoising process. We present Jano, a training‑free framework that leverages this insight for efficient region‑aware generation. Jano introduces an early‑stage complexity recognition algorithm that accurately identifies regional convergence requirements within initial denoising steps, coupled with an adaptive token scheduling runtime that optimizes computational resource allocation. Through comprehensive evaluation on state‑of‑the‑art models, Jano achieves substantial acceleration (average 2.0 times speedup, up to 2.4 times) while preserving generation quality. Our work challenges conventional uniform processing assumptions and provides a practical solution for accelerating large‑scale content generation. The source code of our implementation is available at https://github.com/chen‑yy20/Jano.
Authors:Xingyilang Yin, Chengzhengxu Li, Jiahao Chang, Chi-Man Pun, Xiaodong Cun
Abstract:
Humans are born with vision‑based 4D spatial‑temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLMs). To tackle this challenge, we introduce MLLM‑4D, a comprehensive framework designed to bridge the gaps in training data curation and model post‑training for spatiotemporal understanding and reasoning. On the data front, we develop a cost‑efficient data curation pipeline that repurposes existing stereo video datasets into high‑quality 4D spatiotemporal instructional data. This results in the MLLM4D‑2M and MLLM4D‑R1‑30k datasets for Supervised Fine‑Tuning (SFT) and Reinforcement Fine‑Tuning (RFT), alongside MLLM4D‑Bench for comprehensive evaluation. Regarding model training, our post‑training strategy establishes a foundational 4D understanding via SFT and further catalyzes 4D reasoning capabilities by employing Group Relative Policy Optimization (GRPO) with specialized Spatiotemporal Chain of Thought (ST‑CoT) prompting and Spatiotemporal reward functions (ST‑reward) without involving the modification of architecture. Extensive experiments demonstrate that MLLM‑4D achieves state‑of‑the‑art spatial‑temporal understanding and reasoning capabilities from purely 2D RGB inputs. Project page: https://github.com/GVCLab/MLLM‑4D.
Authors:Wang Chen, Yuhui Zeng, Yongdong Luo, Tianyu Xie, Luojun Lin, Jiayi Ji, Yan Zhang, Xiawu Zheng
Abstract:
Frame selection is crucial due to high frame redundancy and limited context windows when applying Large Vision‑Language Models (LVLMs) to long videos. Current methods typically select frames with high relevance to a given query, resulting in a disjointed set of frames that disregard the narrative structure of video. In this paper, we introduce Wavelet‑based Frame Selection by Detecting Semantic Boundary (WFS‑SB), a training‑free framework that presents a new perspective: effective video understanding hinges not only on high relevance but, more importantly, on capturing semantic shifts ‑ pivotal moments of narrative change that are essential to comprehending the holistic storyline of video. However, direct detection of abrupt changes in the query‑frame similarity signal is often unreliable due to high‑frequency noise arising from model uncertainty and transient visual variations. To address this, we leverage the wavelet transform, which provides an ideal solution through its multi‑resolution analysis in both time and frequency domains. By applying this transform, we decompose the noisy signal into multiple scales and extract a clean semantic change signal from the coarsest scale. We identify the local extrema of this signal as semantic boundaries, which segment the video into coherent clips. Building on this, WFS‑SB comprises a two‑stage strategy: first, adaptively allocating a frame budget to each clip based on a composite importance score; and second, within each clip, employing the Maximal Marginal Relevance approach to select a diverse yet relevant set of frames. Extensive experiments show that WFS‑SB significantly boosts LVLM performance, e.g., improving accuracy by 5.5% on VideoMME, 9.5% on MLVU, and 6.2% on LongVideoBench, consistently outperforming state‑of‑the‑art methods. Our code is available at https://github.com/MAC‑AutoML/WFS‑SB.
Authors:Yingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu Shen
Abstract:
Multimodal large language models (MLLMs) project visual tokens into the embedding space of language models, yet the internal structuring and processing of visual semantics remain poorly understood. In this work, we introduce a two‑fold analytical framework featuring a novel probing tool, EmbedLens, to conduct a fine‑grained analysis. We uncover a pronounced semantic sparsity at the input level: visual tokens consistently partition into sink, dead, and alive categories. Remarkably, only the alive tokens, comprising \approx60% of the total input, carry image‑specific meaning. Furthermore, using a targeted patch‑compression benchmark, we demonstrate that these alive tokens already encode rich, fine‑grained cues (e.g., objects, colors, and OCR) prior to entering the LLM. Internal visual computations (such as visual attention and feed‑forward networks) are redundant for most standard tasks. For the small subset of highly vision‑centric tasks that actually benefit from internal processing, we reveal that alive tokens naturally align with intermediate LLM layers rather than the initial embedding space, indicating that shallow‑layer processing is unnecessary and that direct mid‑layer injection is both sufficient. Ultimately, our findings provide a unified mechanistic view of visual token processing, paving the way for more efficient and interpretable MLLM architectures through selective token pruning, minimized visual computation, and mid‑layer injection. The code is released at: https://github.com/EIT‑NLP/EmbedLens.
Authors:Qihang Fan, Yuang Ai, Huaibo Huang, Ran He
Abstract:
Since Transformers are introduced into vision architectures, their quadratic complexity has always been a significant issue that many research efforts aim to address. A representative approach involves grouping tokens, performing self‑attention calculations within each group, or pooling the tokens within each group into a single token. To this end, various carefully designed grouping strategies have been proposed to enhance the performance of Vision Transformers. Here, we pose the following questions: Are these carefully designed grouping methods truly necessary? Is there a simpler and more unified token grouping method that can replace these diverse methods? Therefore, we propose the random grouping strategy, which involves a simple and fast random grouping strategy for vision tokens. We validate this approach on multiple baselines, and experiments show that random grouping almost outperforms all other grouping methods. When transferred to downstream tasks, such as object detection, random grouping demonstrates even more pronounced advantages. In response to this phenomenon, we conduct a detailed analysis of the advantages of random grouping from multiple perspectives and identify several crucial elements for the design of grouping strategies: positional information, head feature diversity, global receptive field, and fixed grouping pattern. We demonstrate that as long as these four conditions are met, vision tokens require only an extremely simple grouping strategy to efficiently and effectively handle various visual tasks. We also validate the effectiveness of our proposed random method across multiple modalities, including visual tasks, point cloud processing, and vision‑language models. Code will be available at https://github.com/qhfan/random.
Authors:Liyao Jiang, Ruichen Chen, Chao Gao, Di Niu
Abstract:
Recent text‑to‑image (T2I) diffusion models achieve remarkable realism, yet faithful prompt‑image alignment remains challenging, particularly for complex prompts with multiple objects, relations, and fine‑grained attributes. Existing training‑free inference‑time scaling methods rely on fixed iteration budgets that cannot adapt to prompt difficulty, while reflection‑tuned models require carefully curated reflection datasets and extensive joint fine‑tuning of diffusion and vision‑language models, often overfitting to reflection paths data and lacking transferability across models. We introduce RAISE (Requirement‑Adaptive Self‑Improving Evolution), a training‑free, requirement‑driven evolutionary framework for adaptive T2I generation. RAISE formulates image generation as a requirement‑driven adaptive scaling process, evolving a population of candidates at inference time through a diverse set of refinement actions‑including prompt rewriting, noise resampling, and instructional editing. Each generation is verified against a structured checklist of requirements, enabling the system to dynamically identify unsatisfied items and allocate further computation only where needed. This achieves adaptive test‑time scaling that aligns computational effort with semantic query complexity. On GenEval and DrawBench, RAISE attains state‑of‑the‑art alignment (0.94 overall GenEval) while incurring fewer generated samples (reduced by 30‑40%) and VLM calls (reduced by 80%) than prior scaling and reflection‑tuned baselines, demonstrating efficient, generalizable, and model‑agnostic multi‑round self‑improvement. Code is available at https://github.com/LiyaoJiang1998/RAISE.
Authors:Pengcheng Shi, Minghui Zhang, Kehan Song, Jiaqi Liu, Yun Gu, Xinglin Zhang
Abstract:
Automated radiology report generation is key for reducing radiologist workload and improving diagnostic consistency, yet generating accurate reports for 3D medical imaging remains challenging. Existing vision‑language models face two limitations: they do not leverage segmentation‑pretrained encoders, and they inject visual features only at the input layer of language models, losing multi‑scale information. We propose U‑VLM, which enables hierarchical vision‑language modeling in both training and architecture: (1) progressive training from segmentation to classification to report generation, and (2) multi‑layer visual injection that routes U‑Net encoder features to corresponding language model layers. Each training stage can leverage different datasets without unified annotations. U‑VLM achieves state‑of‑the‑art performance on CT‑RATE (F1: 0.414 vs 0.258, BLEU‑mean: 0.349 vs 0.305) and AbdomenAtlas 3.0 (F1: 0.624 vs 0.518 for segmentation‑based detection) using only a 0.1B decoder trained from scratch, demonstrating that well‑designed vision encoder pretraining outweighs the benefits of 7B+ pre‑trained language models. Ablation studies show that progressive pretraining significantly improves F1, while multi‑layer injection improves BLEU‑mean. Code is available at https://github.com/yinghemedical/U‑VLM.
Authors:Xu Luo, Ji Zhang, Lianli Gao, Heng Tao Shen, Jingkuan Song
Abstract:
Few‑shot transfer has been revolutionized by stronger pre‑trained models and improved adaptation algorithms.However, there lacks a unified, rigorous evaluation protocol that is both challenging and realistic for real‑world usage. In this work, we establish FEWTRANS, a comprehensive benchmark containing 10 diverse datasets, and propose the Hyperparameter Ensemble (HPE) protocol to overcome the "validation set illusion" in data‑scarce regimes. Our empirical findings demonstrate that the choice of pre‑trained model is the dominant factor for performance, while many sophisticated transfer methods offer negligible practical advantages over a simple full‑parameter fine‑tuning baseline. To explain this surprising effectiveness, we provide an in‑depth mechanistic analysis showing that full fine‑tuning succeeds via distributed micro‑adjustments and more flexible reshaping of high‑level semantic presentations without suffering from overfitting. Additionally, we quantify the performance collapse of multimodal models in specialized domains as a result of linguistic rarity using adjusted Zipf frequency scores. By releasing FEWTRANS, we aim to provide a rigorous "ruler" to streamline reproducible advances in few‑shot transfer learning research. We make the FEWTRANS benchmark publicly available at https://github.com/Frankluox/FewTrans.
Authors:Boming Tan, Xiangdong Zhang, Ning Liao, Yuqing Zhang, Shaofeng Zhang, Xue Yang, Qi Fan, Yanyong Zhang
Abstract:
Despite impressive progress in video generation, existing models remain limited to surface‑level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a single form of world‑related knowledge or rely on rigid alignment strategies to introduce additional knowledge. However, aligning the single world knowledge is insufficient to constitute a world model that requires jointly modeling multiple heterogeneous dimensions (e.g., physical commonsense, 3D and temporal consistency). To address this limitation, we introduce DreamWorld, a unified framework that integrates complementary world knowledge into video generators via a Joint World Modeling Paradigm, jointly predicting video pixels and features from foundation models to capture temporal dynamics, spatial geometry, and semantic consistency. However, naively optimizing these heterogeneous objectives can lead to visual instability and temporal flickering. To mitigate this issue, we propose Consistent Constraint Annealing (CCA) to progressively regulate world‑level constraints during training, and Multi‑Source Inner‑Guidance to enforce learned world priors at inference. Extensive evaluations show that DreamWorld improves world consistency, outperforming Wan2.1 by 2.26 points on VBench. Code will be made publicly available at \hrefhttps://github.com/ABU121111/DreamWorld\textcolormypinkGithub.
Authors:Xueyang Li, Yunzhong Lou, Yu Song, Xiangdong Zhou
Abstract:
Computer‑Aided Design (CAD) generative modeling has a strong and long‑term application in the industry. Recently, the parametric CAD sequence as the design logic of an object has been widely mined by sequence models. However, the industrial CAD models, especially in component objects, are fine‑grained and complex, requiring a longer parametric CAD sequence to define. To address the problem, we introduce Mamba‑CAD, a self‑supervised generative modeling for complex CAD models in the industry, which can model on a longer parametric CAD sequence. Specifically, we first design an encoder‑decoder framework based on a Mamba architecture and pair it with a CAD reconstruction task for pre‑training to model the latent representation of CAD models; and then we utilize the learned representation to guide a generative adversarial network to produce the fake representation of CAD models, which would be finally recovered into parametric CAD sequences via the decoder of MambaCAD. To train Mamba‑CAD, we further create a new dataset consisting of 77,078 CAD models with longer parametric CAD sequences. Comprehensive experiments are conducted to demonstrate the effectiveness of our model under various evaluation metrics, especially in the generation length of valid parametric CAD sequences. The code and dataset can be achieved from https://github.com/Sunny‑Hack/Code‑for‑Mamba‑CAD‑AAAI‑2025‑.
Authors:Hui Wan, Libin Lan
Abstract:
Executing multiple tasks simultaneously in medical image analysis, including segmentation, classification, detection, and regression, often introduces significant challenges regarding model generalizability and the optimization of shared feature representations. While Vision Foundation Models (VFMs) provide powerful general representations, full fine‑tuning on limited medical data is prone to overfitting and incurs high computational costs. Moreover, existing parameter‑efficient fine‑tuning approaches typically adopt task‑agnostic adaptation protocols, overlooking both task‑specific mechanisms and the varying sensitivity of model layers during fine‑tuning. In this work, we propose Task‑Aware Prompting and Selective Layer Fine‑Tuning (TAP‑SLF), a unified framework for multi‑task ultrasound image analysis. TAP‑SLF incorporates task‑aware soft prompts to encode task‑specific priors into the input token sequence and applies LoRA to selected specific top layers of the encoder. This strategy updates only a small fraction of the VFM parameters while keeping the pre‑trained backbone frozen. By combining task‑aware prompts with selective high‑layer fine‑tuning, TAP‑SLF enables efficient VFM adaptation to diverse medical tasks within a shared backbone. Results on the FMC_UIA 2026 Challenge test set, where TAP‑SLF wins fifth place, combined with evaluations on the officially released training dataset using an 8:2 train‑test split, demonstrate that task‑aware prompting and selective layer tuning are effective strategies for efficient VFM adaptation.
Authors:Hulingxiao He, Zhi Tan, Yuxin Peng
Abstract:
A high‑performing, general‑purpose visual understanding model should map visual inputs to a taxonomic tree of labels, identify novel categories beyond the training set for which few or no publicly available images exist. Large Multimodal Models (LMMs) have achieved remarkable progress in fine‑grained visual recognition (FGVR) for known categories. However, they remain limited in hierarchical visual recognition (HVR) that aims at predicting consistent label paths from coarse to fine categories, especially for novel categories. To tackle these challenges, we propose Taxonomy‑Aware Representation Alignment (TARA), a simple yet effective strategy to inject taxonomic knowledge into LMMs. TARA leverages representations from biology foundation models (BFMs) that encode rich biological relationships through hierarchical contrastive learning. By aligning the intermediate representations of visual features with those of BFMs, LMMs are encouraged to extract discriminative visual cues well structured in the taxonomy tree. Additionally, we align the representations of the first answer token with the ground‑truth label, flexibly bridging the gap between contextualized visual features and categories of varying granularity according to user intent. Experiments demonstrate that TARA consistently enhances LMMs' hierarchical consistency and leaf node accuracy, enabling reliable recognition of both known and novel categories within complex biological taxonomies. Code is available at https://github.com/PKU‑ICST‑MIPL/TARA_CVPR2026.
Authors:Changpu Li, Shuang Wu, Songlin Tang, Guangming Lu, Jun Yu, Wenjie Pei
Abstract:
Reconstructing transparent objects from a set of multi‑view images is a challenging task due to the complicated nature and indeterminate behavior of light propagation. Typical methods are primarily tailored to specific scenarios, such as objects following a uniform topology, exhibiting ideal transparency and surface specular reflections, or with only surface materials, which substantially constrains their practical applicability in real‑world settings. In this work, we propose a differentiable rendering framework for transparent objects, dubbed DiffTrans, which allows for efficient decomposition and reconstruction of the geometry and materials of transparent objects, thereby reconstructing transparent objects accurately in intricate scenes with diverse topology and complex texture. Specifically, we first utilize FlexiCubes with dilation and smoothness regularization as the iso‑surface representation to reconstruct an initial geometry efficiently from the multi‑view object silhouette. Meanwhile, we employ the environment light radiance field to recover the environment of the scene. Then we devise a recursive differentiable ray tracer to further optimize the geometry, index of refraction and absorption rate simultaneously in a unified and end‑to‑end manner, leading to high‑quality reconstruction of transparent objects in intricate scenes. A prominent advantage of the designed ray tracer is that it can be implemented in CUDA, enabling a significantly reduced computational cost. Extensive experiments on multiple benchmarks demonstrate the superior reconstruction performance of our DiffTrans compared with other methods, especially in intricate scenes involving transparent objects with diverse topology and complex texture. The code is available at https://github.com/lcp29/DiffTrans.
Authors:Yuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi Fan
Abstract:
The development of 3D Vision‑Language Models (VLMs), crucial for applications in robotics, autonomous driving, and augmented reality, is severely constrained by the scarcity of paired 3D‑text data. Existing methods rely solely on next‑token prediction loss, using only language tokens for supervision. This results in inefficient utilization of limited 3D data and leads to a significant degradation and loss of valuable geometric information in intermediate representations. To address these limitations, we propose \mname, a novel feature‑level alignment regularization method. \mname explicitly supervises intermediate point cloud tokens to preserve fine‑grained 3D geometric‑semantic information throughout the language modeling process. Specifically, we constrain the intermediate point cloud tokens within the LLM to align with visual input tokens via a consistency loss. By training only a lightweight alignment projector and LoRA adapters, \mname achieves explicit feature‑level supervision with minimal computational overhead, effectively preventing geometric degradation. Extensive experiments on ModelNet40 and Objaverse datasets demonstrate that our method achieves 2.08 pp improvement on average for classification tasks, with a substantial 7.50 pp gain on the challenging open‑vocabulary Objaverse classification task and 4.88 pp improvement on 3D object captioning evaluated by Qwen2‑72B‑Instruct, validating the effectiveness of \mname. Code is publicly available at \hrefhttps://github.com/yharoldsu0627/PointAlignhttps://github.com/yharoldsu0627/PointAlign.
Authors:Xuanshuo Fu, Lei Kang, Javier Vazquez-Corral
Abstract:
Low‑light images often suffer from low contrast, noise, and color distortion, degrading visual quality and impairing downstream vision tasks. We propose a novel conditional diffusion framework for low‑light image enhancement that incorporates a Structured Control Embedding Module (SCEM). SCEM decomposes a low‑light image into four informative components including illumination, illumination‑invariant features, shadow priors, and color‑invariant cues. These components serve as control signals that condition a U‑Net‑based diffusion model trained with a simplified noise‑prediction loss. Thus, the proposed SCEM equipped Diffusion method enforces structured enhancement guided by physical priors. In experiments, our model is trained only on the LOLv1 dataset and evaluated without fine‑tuning on LOLv2‑real, LSRW, DICM, MEF, and LIME. The method achieves state‑of‑the‑art performance in quantitative and perceptual metrics, demonstrating strong generalization across benchmarks. https://casted.github.io/scem/.
Authors:Jiayang Shi, Lincen Yang, Zhong Li, Tristan Van Leeuwen, Daniel M. Pelt, K. Joost Batenburg
Abstract:
Generative models, particularly Diffusion Models (DM), have shown strong potential for Computed Tomography (CT) reconstruction serving as expressive priors for solving ill‑posed inverse problems. However, diffusion‑based reconstruction relies on Stochastic Differential Equations (SDEs) for forward diffusion and reverse denoising, where such stochasticity can interfere with repeated data consistency corrections in CT reconstruction. Since CT reconstruction is often time‑critical in clinical and interventional scenarios, improving reconstruction efficiency is essential. In contrast, Flow Matching (FM) models sampling as a deterministic Ordinary Differential Equation (ODE), yielding smooth trajectories without stochastic noise injection. This deterministic formulation is naturally compatible with repeated data consistency operations. Furthermore, we observe that FM‑predicted velocity fields exhibit strong correlations across adjacent steps. Motivated by this, we propose an FM‑based CT reconstruction framework (FMCT) and an efficient variant (EFMCT) that reuses previously predicted velocity fields over consecutive steps to substantially reduce the number of Neural network Function Evaluations (NFEs), thereby improving inference efficiency. We provide theoretical analysis showing that the error introduced by velocity reuse is bounded when combined with data consistency operations. Extensive experiments demonstrate that FMCT/EFMCT achieve competitive reconstruction quality while significantly improving computational efficiency compared with diffusion‑based methods. The codebase is open‑sourced at https://github.com/EFMCT/EFMCT.
Authors:Lingfeng He, De Cheng, Huaijie Wang, Xi Yang, Nannan Wang, Xinbo Gao
Abstract:
Continual Learning (CL) requires models to sequentially adapt to new tasks without forgetting old knowledge. Recently, Low‑Rank Adaptation (LoRA), a representative Parameter‑Efficient Fine‑Tuning (PEFT) method, has gained increasing attention in CL. Several LoRA‑based CL methods reduce interference across tasks by separating their update spaces, typically building the new space from the estimated null space of past tasks. However, they (i) overlook task‑shared directions, which suppresses knowledge transfer, and (ii) fail to capture truly effective task‑specific directions since these ``null bases" of old tasks can remain nearly inactive for new task under correlated tasks. To address this, we study LoRA learning capability from a projection energy perspective, and propose Low‑rank Decomposition and Adaptation (LoDA). It performs a task‑driven decomposition to build general and truly task‑specific LoRA subspaces by solving two energy‑based objectives, decoupling directions for knowledge sharing and isolation. LoDA fixes LoRA down‑projections on two subspaces and learns robust up‑projections via a Gradient‑Aligned Optimization (GAO) approach. After each task, before integrating the LoRA updates into the backbone, LoDA derives a closed‑form recalibration for the general update, approximating a feature‑level joint optimum along this task‑shared direction. Experiments indicate that LoDA outperforms existing CL methods. Our code is available at https://github.com/HHHLF/LoDA_ICML2026.
Authors:Wenxin Tang, Jingyu Xiao, Yanpei Gong, Fengyuan Ran, Tongchuan Xia, Junliang Liu, Man Ho Lam, Wenxuan Wang, Michael R. Lyu
Abstract:
Automated academic poster generation aims to distill lengthy research papers into concise, visually coherent presentations. Existing Multimodal Large Language Models (MLLMs) based approaches, however, suffer from three critical limitations: low information density in full‑paper inputs, excessive token consumption, and unreliable layout verification. We present EfficientPosterGen, an end‑to‑end framework that addresses these challenges through semantic‑aware retrieval and token‑efficient multimodal generation. EfficientPosterGen introduces three core innovations: (1) Semantic‑aware Key Information Retrieval (SKIR), which constructs a semantic contribution graph to model inter‑segment relationships and selectively preserves important content; (2) Visual‑based Context Compression (VCC), which renders selected text segments into images to shift textual information into the visual modality, significantly reducing token usage while generating poster‑ready bullet points; and (3) Agentless Layout Violation Detection (ALVD), a deterministic color‑gradient‑based algorithm that reliably detects content overflow and spatial sparsity without auxiliary MLLMs. Extensive experiments demonstrate that EfficientPosterGen achieves substantial improvements in token efficiency and layout reliability while maintaining high poster quality, offering a scalable solution for automated academic poster generation. Our code is available at https://github.com/vinsontang1/EfficientPosterGen‑Code.
Authors:Haoxiang Sun, Tao Wang, Chenwei Tang, Li Yuan, Jiancheng Lv
Abstract:
Following the success of Group Relative Policy Optimization (GRPO) in foundation LLMs, an increasing number of works have sought to adapt GRPO to Visual Large Language Models (VLLMs) for visual perception tasks (e.g., detection and segmentation). However, much of this line of research rests on a long‑standing yet unexamined assumption: training paradigms developed for language reasoning can be transferred seamlessly to visual perception. Our experiments show that this assumption is not valid, revealing intrinsic differences between reasoning‑oriented and perception‑oriented settings. Using reasoning segmentation as a representative case, we surface two overlooked factors: (i) the need for a broader output space, and (ii) the importance of fine‑grained, stable rewards. Building on these observations, we propose Dr.~Seg, a simple, plug‑and‑play GRPO‑based framework consisting of a Look‑to‑Confirm mechanism and a Distribution‑Ranked Reward module, requiring no architectural modifications and integrating seamlessly with existing GRPO‑based VLLMs. Extensive experiments demonstrate that Dr.~Seg improves performance in complex visual scenarios while maintaining strong generalization. Code, models, and datasets are available at https://github.com/eVI‑group‑SCU/Dr‑Seg.
Authors:Zihang Zou, Boqing Gong, Liqiang Wang
Abstract:
In this paper, we highlight a critical threat posed by emerging neural models: data plagiarism. We demonstrate how modern neural models (e.g., diffusion models) can replicate copyrighted images, even when protected by advanced watermarking techniques. To expose vulnerabilities in copyright protection and facilitate future research, we propose a general approach to neural plagiarism that can either forge replicas of copyrighted data or introduce copyright ambiguity. Our method, based on "anchors and shims", employs inverse latents as anchors and finds shim perturbations that gradually deviate the anchor latents, thereby evading watermark or copyright detection. By applying perturbations to the cross‑attention mechanism at different timesteps, our approach induces varying degrees of semantic modification in copyrighted images, enabling it to bypass protections ranging from visible trademarks and signatures to invisible watermarks. Notably, our method is a purely gradient‑based search that requires no additional training or fine‑tuning. Experiments on MS‑COCO and real‑world copyrighted images show that diffusion models can replicate copyrighted images, underscoring the urgent need for countermeasures against neural plagiarism.
Authors:Zhihao Li, Shengwei Dong, Chuang Yi, Junxuan Gao, Zhilu Lai, Zhiqiang Liu, Wei Wang, Guangtao Zhang
Abstract:
Existing image SR and generic diffusion models transfer poorly to fluid SR: they are sampling‑intensive, ignore physical constraints, and often yield spectral mismatch and spurious divergence. We address fluid super‑resolution (SR) with ReMD (\underlineResidual‑\underlineMultigrid \underlineDiffusion), a physics‑consistent diffusion framework. At each reverse step, ReMD performs a \emphmultigrid residual correction: the update direction is obtained by coupling data consistency with lightweight physics cues and then correcting the residual across scales; the multiscale hierarchy is instantiated with a \emphmulti‑wavelet basis to capture both large structures and fine vortical details. This coarse‑to‑fine design accelerates convergence and preserves fine structures while remaining equation‑free. Across atmospheric and oceanic benchmarks, ReMD improves accuracy and spectral fidelity, reduces divergence, and reaches comparable quality with markedly fewer sampling steps than diffusion baselines. Our results show that enforcing physics consistency \emphinside the diffusion process via multigrid residual correction and multi‑wavelet multiscale modeling is an effective route to efficient fluid SR. Our code are available on https://github.com/lizhihao2022/ReMD.
Authors:Sevda Öğüt, Cédric Vincent-Cuaz, Natalia Dubljevic, Carlos Hurtado, Vaishnavi Subramanian, Pascal Frossard, Dorina Thanou
Abstract:
Self‑supervised vision models have achieved notable success in digital pathology. However, their domain‑agnostic transformer architectures are not originally designed to account for fundamental biological elements of histopathology images, namely cells and their complex interactions. In this work, we hypothesize that a biologically‑informed modeling of tissues as cell graphs offers a more efficient representation learning. Thus, we introduce GrapHist, a novel graph‑based self‑supervised learning framework for histopathology, which learns generalizable and structurally‑informed embeddings that enable diverse downstream tasks. GrapHist integrates masked autoencoders and heterophilic graph neural networks that are explicitly designed to capture the heterogeneity of tumor microenvironments. We pre‑train GrapHist on a large collection of 11 million cell graphs derived from breast tissues and evaluate its transferability across in‑ and out‑of‑domain benchmarks. Our results show that GrapHist achieves competitive performance compared to its vision‑based counterparts in slide‑, region‑, and cell‑level tasks, while requiring four times fewer parameters. It also drastically outperforms fully‑supervised graph models on cancer subtyping tasks. Finally, we also release five graph‑based digital pathology datasets used in our study at https://huggingface.co/ogutsevda/datasets , establishing the first large‑scale graph benchmark in this field. Our code is available at https://github.com/ogutsevda/graphist .
Authors:Sathwik Karnik, Juyeop Kim, Sanmi Koyejo, Jong-Seok Lee, Somil Bansal
Abstract:
Text‑to‑image diffusion models often memorize training data, revealing a fundamental failure to generalize beyond the training set. Current mitigation strategies typically sacrifice image quality or prompt alignment to reduce memorization. To address this, we propose Reachability‑Aware Diffusion Steering (RADS), an inference‑time framework that prevents memorization while preserving generation fidelity. RADS models the diffusion denoising process as a dynamical system and applies concepts from reachability analysis to approximate the "backward reachable tube"‑‑the set of intermediate states that inevitably evolve into memorized samples. We then formulate mitigation as a constrained reinforcement learning (RL) problem, where a policy learns to steer the trajectory away from memorization via minimal perturbations in the caption embedding space. Empirical evaluations show that RADS achieves a superior Pareto frontier between generation diversity (SSCD), quality (FID), and alignment (CLIP) compared to state‑of‑the‑art baselines. Crucially, RADS provides robust mitigation without modifying the diffusion backbone, offering a plug‑and‑play solution for safe generation. Our website is available at: https://s‑karnik.github.io/rads‑memorization‑project‑page/.
Authors:Shengqu Cai, Weili Nie, Chao Liu, Julius Berner, Lvmin Zhang, Nanye Ma, Hansheng Chen, Maneesh Agrawala, Leonidas Guibas, Gordon Wetzstein, Arash Vahdat
Abstract:
Scaling video generation from seconds to minutes faces a critical bottleneck: while short‑video data is abundant and high‑fidelity, coherent long‑form data is scarce and limited to narrow domains. To address this, we propose a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long‑term coherence based on a unified representation via a Decoupled Diffusion Transformer. Our approach utilizes a global Flow Matching head trained via supervised learning on long videos to capture narrative structure, while simultaneously employing a local Distribution Matching head that aligns sliding windows to a frozen short‑video teacher via a mode‑seeking reverse‑KL divergence. This strategy enables the synthesis of minute‑scale videos that learns long‑range coherence and motions from limited long videos via supervised flow matching, while inheriting local realism by aligning every sliding‑window segment of the student to a frozen short‑video teacher, resulting in a few‑step fast long video generator. Evaluations show that our method effectively closes the fidelity‑horizon gap by jointly improving local sharpness, motion and long‑range consistency. Project website: https://primecai.github.io/mmm/.
Authors:Arnas Uselis, Andrea Dittadi, Seong Joon Oh
Abstract:
Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Although modern models are trained on massive datasets, they still cover only a tiny fraction of the combinatorial space of possible inputs, raising the question of what structure representations must have to support generalization to unseen combinations. We formalize three desiderata for compositional generalization under standard training (divisibility, transferability, stability) and show they impose necessary geometric constraints: representations must decompose linearly into per‑concept components, and these components must be orthogonal across concepts. This provides theoretical grounding for the Linear Representation Hypothesis: the linear structure widely observed in neural representations is a necessary consequence of compositional generalization. We further derive dimension bounds linking the number of composable concepts to the embedding geometry. Empirically, we evaluate these predictions across modern vision models (CLIP, SigLIP, DINO) and find that representations exhibit partial linear factorization with low‑rank, near‑orthogonal per‑concept factors, and that the degree of this structure correlates with compositional generalization on unseen combinations. As models continue to scale, these conditions predict the representational geometry they may converge to. Code is available at https://github.com/oshapio/necessary‑compositionality.
Authors:Chengyan Deng, Zhangquan Chen, Li Yu, Kai Zhang, Xue Zhou, Wang Zhang
Abstract:
Diffusion‑based Real‑World Image Super‑Resolution (Real‑ISR) achieves impressive perceptual quality but suffers from high computational costs due to iterative sampling. While recent distillation approaches leveraging large‑scale Text‑to‑Image (T2I) priors have enabled one‑step generation, they are typically hindered by prohibitive parameter counts and the inherent capability bounds imposed by teacher models. As a lightweight alternative, Consistency Models offer efficient inference but struggle with two critical limitations: the accumulation of consistency drift inherent to transitive training, and a phenomenon we term "Geometric Decoupling" ‑ where the generative trajectory achieves pixel‑wise alignment yet fails to preserve structural coherence. To address these challenges, we propose GTASR (Geometric Trajectory Alignment Super‑Resolution), a simple yet effective consistency training paradigm for Real‑ISR. Specifically, we introduce a Trajectory Alignment (TA) strategy to rectify the tangent vector field via full‑path projection, and a Dual‑Reference Structural Rectification (DRSR) mechanism to enforce strict structural constraints. Extensive experiments verify that GTASR delivers superior performance over representative baselines while maintaining minimal latency. The code and model will be released at https://github.com/Blazedengcy/GTASR.
Authors:Zhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu, Xiaojie Li, Rui Wang, Yunpeng Chen, Daquan Zhou
Abstract:
Recent progress in text‑to‑image generation has greatly advanced visual fidelity and creativity, but it has also imposed higher demands on prompt complexity‑particularly in encoding intricate spatial relationships. In such cases, achieving satisfactory results often requires multiple sampling attempts. To address this challenge, we introduce a novel method that strengthens the spatial understanding of current image generation models. We first construct the SpatialReward‑Dataset with over 80k preference pairs. Building on this dataset, we build SpatialScore, a reward model designed to evaluate the accuracy of spatial relationships in text‑to‑image generation, achieving performance that even surpasses leading proprietary models on spatial evaluation. We further demonstrate that this reward model effectively enables online reinforcement learning for the complex spatial generation. Extensive experiments across multiple benchmarks show that our specialized reward model yields significant and consistent gains in spatial understanding for image generation.
Authors:Zhengren Wang, Dongsheng Ma, Huaping Zhong, Jiayu Li, Wentao Zhang, Bin Wang, Conghui He
Abstract:
The expansion of retrieval‑augmented generation (RAG) into multimodal domains has intensified the challenge for processing complex visual documents, such as financial reports. While page‑level chunking and retrieval is a natural starting point, it creates a critical bottleneck: delivering entire pages to the generator introduces excessive extraneous context. This not only overloads the generator's attention mechanism but also dilutes the most salient evidence. Moreover, compressing these information‑rich pages into a limited visual token budget further increases the risk of hallucinations. To address this, we introduce AgenticOCR, a dynamic parsing paradigm that transforms optical character recognition (OCR) from a static, full‑text process into a query‑driven, on‑demand extraction system. By autonomously analyzing document layout in a "thinking with images" manner, AgenticOCR identifies and selectively recognizes regions of interest. This approach performs on‑demand decompression of visual tokens precisely where needed, effectively decoupling retrieval granularity from rigid page‑level chunking. AgenticOCR has the potential to serve as the "third building block" of the visual document RAG stack, operating alongside and enhancing standard Embedding and Reranking modules. Experimental results demonstrate that AgenticOCR improves both the efficiency and accuracy of visual RAG systems, achieving expert‑level performance in long document understanding. Code and models are available at https://github.com/OpenDataLab/AgenticOCR.
Authors:Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin, Jiayang Li, Yu Qiao, Jinjin Gu, Yihao Liu
Abstract:
Masked Image Generation Models (MIGMs) have achieved great success, yet their efficiency is hampered by the multiple steps of bi‑directional attention. In fact, there exists notable redundancy in their computation: when sampling discrete tokens, the rich semantics contained in the continuous features are lost. Some existing works attempt to cache the features to approximate future features. However, they exhibit considerable approximation error under aggressive acceleration rates. We attribute this to their limited expressivity and the failure to account for sampling information. To fill this gap, we propose to learn a lightweight model that incorporates both previous features and sampled tokens, and regresses the average velocity field of feature evolution. The model has moderate complexity that suffices to capture the subtle dynamics while keeping lightweight compared to the original base model. We apply our method, MIGM‑Shortcut, to two representative MIGM architectures and tasks. In particular, on the state‑of‑the‑art Lumina‑DiMOO, it achieves over 4x acceleration of text‑to‑image generation while maintaining quality, significantly pushing the Pareto frontier of masked image generation. The code and model weights are available at https://github.com/Kaiwen‑Zhu/MIGM‑Shortcut.
Authors:Tianxiang Du, Hulingxiao He, Yuxin Peng
Abstract:
The widespread use of smartphones has made photography ubiquitous, yet a clear gap remains between ordinary users and professional photographers, who can identify aesthetic issues and provide actionable shooting guidance during capture. We define this capability as aesthetic guidance (AG) ‑‑ an essential but largely underexplored domain in computational aesthetics. Existing multimodal large language models (MLLMs) primarily offer overly positive feedback, failing to identify issues or provide actionable guidance. Without AG capability, they cannot effectively identify distracting regions or optimize compositional balance, thus also struggling in aesthetic cropping, which aims to refine photo composition through reframing after capture. To address this, we introduce AesGuide, the first large‑scale AG dataset and benchmark with 10,748 photos annotated with aesthetic scores, analyses, and guidance. Building upon it, we propose Venus, a two‑stage framework that first empowers MLLMs with AG capability through progressively complex aesthetic questions and then activates their aesthetic cropping power via CoT‑based rationales. Extensive experiments show that Venus substantially improves AG capability and achieves state‑of‑the‑art (SOTA) performance in aesthetic cropping, enabling interpretable and interactive aesthetic refinement across both stages of photo creation. Code is available at https://github.com/PKU‑ICST‑MIPL/Venus_CVPR2026.
Authors:Qiuyang Zhang, Jiujun Cheng, Qichao Mao, Cong Liu, Yu Fang, Yuhong Li, Mengying Ge, Shangce Gao
Abstract:
Spiking Neural Networks (SNNs) promise energy‑efficient vision, but applying them to RGB visual tracking remains difficult: Existing SNN tracking frameworks either do not fully align with spike‑driven computation or do not fully leverage neurons' spatiotemporal dynamics, leading to a trade‑off between efficiency and accuracy. To address this, we introduce SpikeTrack, a spike‑driven framework for energy‑efficient RGB object tracking. SpikeTrack employs a novel asymmetric design that uses asymmetric timestep expansion and unidirectional information flow, harnessing spatiotemporal dynamics while cutting computation. To ensure effective unidirectional information transfer between branches, we design a memory‑retrieval module inspired by neural inference mechanisms. This module recurrently queries a compact memory initialized by the template to retrieve target cues and sharpen target perception over time. Extensive experiments demonstrate that SpikeTrack achieves the state‑of‑the‑art among SNN‑based trackers and remains competitive with advanced ANN trackers. Notably, it surpasses TransT on LaSOT dataset while consuming only 1/26 of its energy. To our knowledge, SpikeTrack is the first spike‑driven framework to make RGB tracking both accurate and energy efficient. The code and models are available at https://github.com/faicaiwawa/SpikeTrack.
Authors:Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu, Zhongqi Yue, Hanwang Zhang
Abstract:
Recent multimodal large language models (MLLMs) increasingly rely on visual chain‑of‑thought to perform region‑grounded reasoning over images. However, existing approaches ground regions via either textified coordinates‑causing modality mismatch and semantic fragmentation or fixed‑granularity patches that both limit precise region selection and often require non‑trivial architectural changes. In this paper, we propose Numerical Visual Chain‑of‑Thought (NV‑CoT), a framework that enables MLLMs to reason over images using continuous numerical coordinates. NV‑CoT expands the MLLM action space from discrete vocabulary tokens to a continuous Euclidean space, allowing models to directly generate bounding‑box coordinates as actions with only minimal architectural modification. The framework supports both supervised fine‑tuning and reinforcement learning. In particular, we replace categorical token policies with a Gaussian (or Laplace) policy over coordinates and introduce stochasticity via reparameterized sampling, making NV‑CoT fully compatible with GRPO‑style policy optimization. Extensive experiments on three benchmarks against eight representative visual reasoning baselines demonstrate that NV‑CoT significantly improves localization precision and final answer accuracy, while also accelerating training convergence, validating the effectiveness of continuous‑action visual reasoning in MLLMs. The code is available in https://github.com/kesenzhao/NV‑CoT.
Authors:Yuyang Hong, Jiaqi Gu, Yujin Lou, Lubin Fan, Qi Yang, Ying Wang, Kun Ding, Yue Wu, Shiming Xiang, Jieping Ye
Abstract:
Knowledge‑based visual question answering (KB‑VQA) demonstrates significant potential for handling knowledge‑intensive tasks. However, conflicts arise between static parametric knowledge in vision language models (VLMs) and dynamically retrieved information due to the static model knowledge from pre‑training. The outputs either ignore retrieved contexts or exhibit inconsistent integration with parametric knowledge, posing substantial challenges for KB‑VQA. Current knowledge conflict mitigation methods primarily adapted from language‑based approaches, focusing on context‑level conflicts through engineered prompting strategies or context‑aware decoding mechanisms. However, these methods neglect the critical role of visual information in conflicts and suffer from redundant retrieved contexts, which impair accurate conflict identification and effective mitigation. To address these limitations, we propose CC‑VQA: a novel training‑free, conflict‑ and correlation‑aware method for KB‑VQA. Our method comprises two core components: (1) Vision‑Centric Contextual Conflict Reasoning, which performs visual‑semantic conflict analysis across internal and external knowledge contexts; and (2) Correlation‑Guided Encoding and Decoding, featuring positional encoding compression for low‑correlation statements and adaptive decoding using correlation‑weighted conflict scoring. Extensive evaluations on E‑VQA, InfoSeek, and OK‑VQA benchmarks demonstrate that CC‑VQA achieves state‑of‑the‑art performance, yielding absolute accuracy improvements of 3.3% to 6.4% compared to existing methods. Code is available at https://github.com/cqu‑student/CC‑VQA.
Authors:Qiyu Feng, Jiwei Shan, Shing Shin Cheng, Hesheng Wang
Abstract:
Neural implicit surface reconstruction with signed distance function has made significant progress, but recovering fine details such as thin structures and complex geometries remains challenging due to unreliable or noisy geometric priors. Existing approaches rely on implicit uncertainty that arises during optimization to filter these priors, which is indirect and inefficient, and masking supervision in high‑uncertainty regions further leads to under‑constrained optimization. To address these issues, we propose GPU‑SDF, a neural implicit framework for indoor surface reconstruction that leverages geometric prior uncertainty and complementary constraints. We introduce a self‑supervised module that explicitly estimates prior uncertainty without auxiliary networks. Based on this estimation, we design an uncertainty‑guided loss that modulates prior influence rather than discarding it, thereby retaining weak but informative cues. To address regions with high prior uncertainty, GPU‑SDF further incorporates two complementary constraints: an edge distance field that strengthens boundary supervision and a multi‑view consistency regularization that enforces geometric coherence. Extensive experiments confirm that GPU‑SDF improves the reconstruction of fine details and serves as a plug‑and‑play enhancement for existing frameworks. Source code will be available at https://github.com/IRMVLab/GPU‑SDF
Authors:Bora Kargi, Arnas Uselis, Seong Joon Oh
Abstract:
When a text description is extended with an additional detail, image‑text similarity should drop if that detail is wrong. We show that CLIP‑style dual encoders often violate this intuition: appending a plausible but incorrect object or relation to an otherwise correct description can increase the similarity score. We call such cases half‑truths. On COCO, CLIP prefers the correct shorter description only 40.6% of the time, and performance drops to 32.9% when the added detail is a relation. We trace this vulnerability to weak supervision on caption parts: contrastive training aligns full sentences but does not explicitly enforce that individual entities and relations are grounded. We propose CS‑CLIP (Component‑Supervised CLIP), which decomposes captions into entity and relation units, constructs a minimally edited foil for each unit, and fine‑tunes the model to score the correct unit above its foil while preserving standard dual‑encoder inference. CS‑CLIP raises half‑truth accuracy to 69.3% and improves average performance on established compositional benchmarks by 5.7 points, suggesting that reducing half‑truth errors aligns with broader gains in compositional understanding. Code is publicly available at: https://github.com/kargibora/CS‑CLIP
Authors:Andrei-Alexandru Bunea, Dan-Matei Popovici, Radu Tudor Ionescu
Abstract:
State‑of‑the‑art models for medical image segmentation achieve excellent accuracy but require substantial computational resources, limiting deployment in resource‑constrained clinical settings. We present SegMate, an efficient 2.5D framework that achieves state‑of‑the‑art accuracy, while considerably reducing computational requirements. Our efficient design is the result of meticulously integrating asymmetric architectures, attention mechanisms, multi‑scale feature fusion, slice‑based positional conditioning, and multi‑task optimization. We demonstrate the efficiency‑accuracy trade‑off of our framework across three modern backbones (EfficientNetV2‑M, MambaOut‑Tiny, FastViT‑T12). We perform experiments on three datasets: TotalSegmentator, SegTHOR and AMOS22. Compared with the vanilla models, SegMate reduces computation (GFLOPs) by up to 2.5x and memory footprint (VRAM) by up to 2.1x, while generally registering performance gains of around 1%. On TotalSegmentator, we achieve a Dice score of 93.51% with only 295MB peak GPU memory. Zero‑shot cross‑dataset evaluations on SegTHOR and AMOS22 demonstrate strong generalization, with Dice scores of up to 86.85% and 89.35%, respectively. We release our open‑source code at https://github.com/andreibunea99/SegMate.
Authors:Fan Yang, Peiguang Jing, Kaihua Qu, Ningyuan Zhao, Yuting Su
Abstract:
Robotic manipulation requires policies that are smooth and responsive to evolving observations. However, synchronous inference in the raw action space introduces several challenges, including intra‑chunk jitter, inter‑chunk discontinuities, and stop‑and‑go execution. These issues undermine a policy's smoothness and its responsiveness to environmental changes. We propose ABPolicy, an asynchronous flow‑matching policy that operates in a B‑spline control‑point action space. First, the B‑spline representation ensures intra‑chunk smoothness. Second, we introduce bidirectional action prediction coupled with refitting optimization to enforce inter‑chunk continuity. Finally, by leveraging asynchronous inference, ABPolicy delivers real‑time, continuous updates. We evaluate ABPolicy across seven tasks encompassing both static settings and dynamic settings with moving objects. Empirical results indicate that ABPolicy reduces trajectory jerk, leading to smoother motion and improved performance. Project website: https://teee000.github.io/ABPolicy/.
Authors:Xiaoyan Lei, Wenlong Zhang, Biao Luo, Hui Liang, Weifeng Cao, Qiuting Lin
Abstract:
Multimodal large models have shown excellent ability in addressing image super‑resolution in real‑world scenarios by leveraging language class as condition information, yet their abilities in degraded images remain limited. In this paper, we first revisit the capabilities of the Recognize Anything Model (RAM) for degraded images by calculating text similarity. We find that directly using contrastive learning to fine‑tune RAM in the degraded space is difficult to achieve acceptable results. To address this issue, we employ a degradation selection strategy to propose a Real Embedding Extractor (REE), which achieves significant recognition performance gain on degraded image content through contrastive learning. Furthermore, we use a Conditional Feature Modulator (CFM) to incorporate the high‑level information of REE for a powerful Mamba‑based network, which can leverage effective pixel information to restore image textures and produce visually pleasing results. Extensive experiments demonstrate that the REE can effectively help image super‑resolution networks balance fidelity and perceptual quality, highlighting the great potential of Mamba in real‑world applications. The source code of this work will be made publicly available at: https://github.com/nathan66666/DACESR.git
Authors:Xiaoyu Guo, Arkaitz Zubiaga
Abstract:
With the aim of detecting AI‑generated images and identifying the specific models responsible for their generation, we propose a multi‑modal multi‑task model. The model leverages pre‑trained BERT and CLIP Vision encoders for text and image feature extraction, respectively, and employs cross‑modal feature fusion with a tailored multi‑task loss function. Additionally, a pseudo‑labeling‑based data augmentation strategy was utilized to expand the training dataset with high‑confidence samples. The model achieved fifth place in both Tasks A and B of the `CT2: AI‑Generated Image Detection' competition, with F1 scores of 83.16% and 48.88%, respectively. These findings highlight the effectiveness of the proposed architecture and its potential for advancing AI‑generated content detection in real‑world scenarios. The source code for our method is published on https://github.com/xxxxxxxxy/AIGeneratedImageDetection.
Authors:Chongyang Xu, Haipeng Li, Shen Cheng, Jingyu Hu, Haoqiang Fan, Ziliang Feng, Shuaicheng Liu
Abstract:
Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial awareness, or require explicit point clouds that are difficult to obtain reliably in real‑world settings. At the same time, recent 3D geometric foundation models show that accurate and diverse 3D structure can be reconstructed directly from RGB images in a fast and robust manner. We leverage this opportunity and propose a framework that builds bimanual manipulation directly on a pre‑trained 3D geometric foundation model. Our policy fuses geometry‑aware latents, 2D semantic features, and proprioception into a unified state representation, and uses diffusion model to jointly predict a future action chunk and a future 3D latent that decodes into a dense pointmap. By explicitly predicting how the 3D scene will evolve together with the action sequence, the policy gains strong spatial understanding and predictive capability using only RGB observations. We evaluate our method both in simulation on the RoboTwin benchmark and in real‑world robot executions. Our approach consistently outperforms 2D‑based and point‑cloud‑based baselines, achieving state‑of‑the‑art performance in manipulation success, inter‑arm coordination, and 3D spatial prediction accuracy. Code is available at https://github.com/Chongyang‑99/GAP.git.
Authors:Hyejin Park, Jiwon Yoon, Sumin Park, Suree Kim, Sinae Jang, Eunsoo Lee, Dongmin Kang, Dongbo Min
Abstract:
Accurate focus quality assessment (FQA) in fluorescence microscopy is challenging due to stain‑dependent optical variations that induce heterogeneous focus behavior across images. Existing methods, however, treat focus quality as a stain‑agnostic problem, assuming a shared global ordering. We formulate stain‑aware FQA for fluorescence microscopy, showing that focus‑rank relationships vary substantially across stains due to stain‑dependent imaging characteristics and invalidate this assumption. To support this formulation, we introduce FluoMix, the first dataset for stain‑aware FQA spanning multiple tissues, fluorescent stains, and focus levels. We further propose FluoCLIP, a two‑stage vision‑language framework that grounds stain semantics and enables stain‑conditioned ordinal reasoning for focus prediction, effectively decoupling stain representation from ordinal structure. By explicitly modeling stain‑dependent focus behavior, FluoCLIP consistently outperforms both conventional FQA methods and recent vision‑language baselines, demonstrating strong generalization across diverse fluorescence microscopy conditions. Code and dataset are publicly available at https://fluoclip.github.io/.
Authors:Changyu Gu, Linwei Chen, Lin Gu, Ying Fu
Abstract:
In remote sensing rotated object detection, mainstream methods suffer from two bottlenecks, directional incoherence at detector neck and task conflict at detecting head. Ulitising fourier rotation equivariance, we introduce Fourier Angle Alignment, which analyses angle information through frequency spectrum and aligns the main direction to a certain orientation. Then we propose two plug and play modules : FAAFusion and FAA Head. FAAFusion works at the detector neck, aligning the main direction of higher‑level features to the lower‑level features and then fusing them. FAA Head serves as a new detection head, which pre‑aligns RoI features to a canonical angle and adds them to the original features before classification and regression. Experiments on DOTA‑v1.0, DOTA‑v1.5 and HRSC2016 show that our method can greatly improve previous work. Particularly, our method achieves new state‑of‑the‑art results of 78.72% mAP on DOTA‑v1.0 and 72.28% mAP on DOTA‑v1.5 datasets with single scale training and testing, validating the efficacy of our approach in remote sensing object detection. The code is made publicly available at https://github.com/gcy0423/Fourier‑Angle‑Alignment .
Authors:Hao Wu, Xudong Wang, Jialiang Zhang, Junlong Tong, Xinghao Chen, Junyan Lin, Yunpu Ma, Xiaoyu Shen
Abstract:
One‑stream Transformer‑based trackers achieve advanced performance in visual object tracking but suffer from significant computational overhead that hinders real‑time deployment. While token pruning offers a path to efficiency, existing methods are fragmented. They typically prune the search region, dynamic template, and static template in isolation, overlooking critical inter‑component dependencies, which yields suboptimal pruning and degraded accuracy. To address this, we introduce UTPTrack, a simple and Unified Token Pruning framework that, for the first time, jointly compresses all three components. UTPTrack employs an attention‑guided, token type‑aware strategy to holistically model redundancy, a design that seamlessly supports unified tracking across multimodal and language‑guided tasks within a single model. Extensive evaluations on 10 benchmarks demonstrate that UTPTrack achieves a new state‑of‑the‑art in the accuracy‑efficiency trade‑off for pruning‑based trackers, pruning 65.4% of vision tokens in RGB‑based tracking and 67.5% in unified tracking while preserving 99.7% and 100.5% of baseline performance, respectively. This strong performance across both RGB and multimodal scenarios underlines its potential as a robust foundation for future research in efficient visual tracking. Code will be released at https://github.com/EIT‑NLP/UTPTrack.
Authors:Hao Wu, Yingqi Fan, Jinyang Dai, Junlong Tong, Yunpu Ma, Xiaoyu Shen
Abstract:
The quadratic computational cost of processing vision tokens in Multimodal Large Language Models (MLLMs) hinders their widespread adoption. While progressive vision token pruning offers a promising solution, current methods misinterpret shallow layer functions and use rigid schedules, which fail to unlock the full efficiency potential. To address these issues, we propose HiDrop, a framework that aligns token pruning with the true hierarchical function of MLLM layers. HiDrop features two key innovations: (1) Late Injection, which bypasses passive shallow layers to introduce visual tokens exactly where active fusion begins; and (2) Concave Pyramid Pruning with an Early Exit mechanism to dynamically adjust pruning rates across middle and deep layers. This process is optimized via an inter‑layer similarity measure and a differentiable top‑k operator. To ensure practical efficiency, HiDrop further incorporates persistent positional encoding, FlashAttention‑compatible token selection, and parallel decoupling of vision computation to eliminate hidden overhead associated with dynamic token reduction. Extensive experiments show that HiDrop compresses about 90% visual tokens while matching the original performance and accelerating training by 1.72 times. Our work not only sets a new state‑of‑the‑art for efficient MLLM training and inference but also provides valuable insights into the hierarchical nature of multimodal fusion. The code is released at https://github.com/EIT‑NLP/HiDrop.
Authors:Dingqi Ye, Daniel Kiv, Wei Hu, Jimeng Shi, Shaowen Wang
Abstract:
The remote sensing community is witnessing a rapid growth of foundation models, which provide powerful embeddings for a wide range of downstream tasks. However, practical adoption and fair comparison remain challenging due to substantial heterogeneity in model release formats, platforms and interfaces, and input data specifications. These inconsistencies significantly increase the cost of obtaining, using, and benchmarking embeddings across models. To address this issue, we propose rs‑embed, a Python library that offers a unified, region of interst (ROI) centric interface: with a single line of code, users can retrieve embeddings from any supported model for any location and any time range. The library also provides efficient batch processing to enable large‑scale embedding generation and evaluation. The code is available at: https://github.com/cybergis/rs‑embed
Authors:Wei Luo, Yangfan Ou, Jin Deng, Zeshuai Deng, Xiquan Yan, Zhiquan Wen, Mingkui Tan
Abstract:
Large‑scale Vision‑Language Models (VLMs) exhibit strong zero‑shot recognition, yet their real‑world deployment is challenged by distribution shifts. While Test‑Time Adaptation (TTA) can mitigate this, existing VLM‑based TTA methods operate under a closed‑set assumption, failing in open‑set scenarios where test streams contain both covariate‑shifted in‑distribution (csID) and out‑of‑distribution (csOOD) data. This leads to a critical difficulty: the model must discriminate unknown csOOD samples to avoid interference while simultaneously adapting to known csID classes for accuracy. Current open‑set TTA (OSTTA) methods rely on hard thresholds for separation and entropy minimization for adaptation. These strategies are brittle, often misclassifying ambiguous csOOD samples and inducing overconfident predictions, and their parameter‑update mechanism is computationally prohibitive for VLMs. To address these limitations, we propose Prototype‑based Double‑Check Separation (ProtoDCS), a robust framework for OSTTA that effectively separates csID and csOOD samples, enabling safe and efficient adaptation of VLMs to csID data. Our main contributions are: (1) a novel double‑check separation mechanism employing probabilistic Gaussian Mixture Model (GMM) verification to replace brittle thresholding; and (2) an evidence‑driven adaptation strategy utilizing uncertainty‑aware loss and efficient prototype‑level updates, mitigating overconfidence and reducing computational overhead. Extensive experiments on CIFAR‑10/100‑C and Tiny‑ImageNet‑C demonstrate that ProtoDCS achieves state‑of‑the‑art performance, significantly boosting both known‑class accuracy and OOD detection metrics. Code will be available at https://github.com/O‑YangF/ProtoDCS.
Authors:Haowen Zhu, Ning Yin, Xiaogen Zhou
Abstract:
Vision‑language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi‑organ medical imaging introduces two principal challenges: (1) modality‑specific vision‑language alignment and (2) cross‑modal feature fusion. In this work, we propose MedMAP, a Medical Modality‑Aware Pretraining framework that enhances vision‑language representation learning in 3D MRI. MedMAP comprises a modality‑aware vision‑language alignment stage and a fine‑tuning stage for multi‑organ abnormality detection. During the pre‑training stage, the modality‑aware encoders implicitly capture the joint modality distribution and improve alignment between visual and textual representations. We then fine‑tune the pre‑trained vision encoders (while keeping the text encoder frozen) for downstream tasks. To this end, we curated MedMoM‑MRI3D, comprising 7,392 3D MRI volume‑report pairs spanning twelve MRI modalities and nine abnormalities tailored for various 3D medical analysis tasks. Extensive experiments on MedMoM‑MRI3D demonstrate that MedMAP significantly outperforms existing VLMs in 3D MRI‑based multi‑organ abnormality detection. Our code is available at https://github.com/RomantiDr/MedMAP.
Authors:Abhishek Dalvi, Vasant Honavar
Abstract:
Large unimodal foundation models for vision and language encode rich semantic structures, yet aligning them typically requires computationally intensive multimodal fine‑tuning. Such approaches depend on large‑scale parameter updates, are resource intensive, and can perturb pretrained representations. Emerging evidence suggests, however, that independently trained foundation models may already exhibit latent semantic compatibility, reflecting shared structures in the data they model. This raises a fundamental question: can cross‑modal alignment be achieved without modifying the models themselves? Here we introduce HDFLIM (HyperDimensional computing with Frozen Language and Image Models), a framework that establishes cross‑modal mappings while keeping pretrained vision and language models fully frozen. HDFLIM projects unimodal embeddings into a shared hyperdimensional space and leverages lightweight symbolic operations ‑‑ binding, bundling, and similarity‑based retrieval to construct associative cross‑modal representations in a single pass over the data. Caption generation emerges from high‑dimensional memory retrieval rather than iterative gradient‑based optimization. We show that HDFLIM achieves performance comparable to end‑to‑end vision‑language training methods and produces captions that are more semantically grounded than zero‑shot baselines. By decoupling alignment from parameter tuning, our results suggest that semantic mapping across foundation models can be realized through symbolic operations on hyperdimensional encodings of the respective embeddings. More broadly, this work points toward an alternative paradigm for foundation model alignment in which frozen models are integrated through structured representational mappings rather than through large‑scale retraining. The codebase for our implementation can be found at https://github.com/Abhishek‑Dalvi410/HDFLIM.
Authors:Jeongbin Hong, Dooseop Choi, Taeg-Hyun An, Kyounghwan An, Kyoung-Wook Min
Abstract:
Transforming image features from perspective view (PV) space to bird's‑eye‑view (BEV) space remains challenging in autonomous driving due to depth ambiguity and occlusion. Although several view transformation (VT) paradigms have been proposed, the challenge still remains. In this paper, we propose a new regularization framework, dubbed CycleBEV, that enhances existing VT models for BEV semantic segmentation. Inspired by cycle consistency, widely used in image distribution modeling, we devise an inverse view transformation (IVT) network that maps BEV segmentation maps back to PV segmentation maps and use it to regularize VT networks during training through cycle consistency losses, enabling them to capture richer semantic and geometric information from input PV images. To further exploit the capacity of the IVT network, we introduce two novel ideas that extend cycle consistency into geometric and representation spaces. We evaluate CycleBEV on four representative VT models covering three major paradigms using the large‑scale nuScenes dataset. Experimental results show consistent improvements ‑‑ with gains of up to 0.74, 4.86, and 3.74 mIoU for drivable area, vehicle, and pedestrian classes, respectively ‑‑ without increasing inference complexity, since the IVT network is used only during training. The implementation code is available at https://github.com/JeongbinHong/CycleBEV.
Authors:Ruxiao Duan, Alex Wong
Abstract:
Understanding sources of uncertainty is fundamental to trustworthy three‑dimensional scene modeling. While recent advances in neural radiance fields (NeRFs) achieve impressive accuracy in scene reconstruction and novel view synthesis, the lack of uncertainty estimation significantly limits their deployment in safety‑critical settings. Existing uncertainty quantification methods for NeRFs fail to separately capture both aleatoric and epistemic uncertainties. Among those that do quantify one or the other, many of them either compromise rendering quality or incur significant computational overhead to obtain uncertainty estimates. To address these issues, we introduce Evidential Neural Radiance Fields, a probabilistic approach that seamlessly integrates with the NeRF rendering process, enabling direct quantification of both aleatoric and epistemic uncertainties from a single forward pass. We compare multiple uncertainty quantification methods on three standardized benchmarks, where our approach demonstrates state‑of‑the‑art scene reconstruction fidelity and uncertainty estimation quality. Code is available at https://github.com/KerryDRX/EvidentialNeRF.
Authors:Cho-Ying Wu, Zixun Huang, Xinyu Huang, Liu Ren
Abstract:
We present the first study of cross‑sensor view synthesis across different modalities. We examine a practical, fundamental, yet widely overlooked problem: getting aligned RGB‑X data, where most RGB‑X prior work assumes such pairs exist and focuses on modality fusion, but it empirically requires huge engineering effort in calibration. We propose a match‑densify‑consolidate method. First, we perform RGB‑X image matching followed by guided point densification. Using the proposed confidence‑aware densification and self‑matching filtering, we attain better view synthesis and later consolidate them in 3D Gaussian Splatting (3DGS). Our method uses no 3D priors for X‑sensor and only assumes nearly no‑cost COLMAP for RGB. We aim to remove the cumbersome calibration for various RGB‑X sensors and advance the popularity of cross‑sensor learning by a scalable solution that breaks through the bottleneck in large‑scale real‑world RGB‑X data collection.
Authors:Junjiang Wu, Liejun Wang, Zhiqing Guo
Abstract:
With the rapid advancement of deepfake technology, malicious face manipulations pose a significant threat to personal privacy and social security. However, existing proactive forensics methods typically treat deepfake detection, tampering localization, and source tracing as independent tasks, lacking a unified framework to address them jointly. To bridge this gap, we propose a unified proactive forensics framework that jointly addresses these three core tasks. Our core framework adopts an innovative 152‑dimensional landmark‑identity watermark termed LIDMark, which structurally interweaves facial landmarks with a unique source identifier. To robustly extract the LIDMark, we design a novel Factorized‑Head Decoder (FHD). Its architecture factorizes the shared backbone features into two specialized heads (i.e., regression and classification), robustly reconstructing the embedded landmarks and identifier, respectively, even when subjected to severe distortion or tampering. This design realizes an "all‑in‑one" trifunctional forensic solution: the regression head underlies an "intrinsic‑extrinsic" consistency check for detection and localization, while the classification head robustly decodes the source identifier for tracing. Extensive experiments show that the proposed LIDMark framework provides a unified, robust, and imperceptible solution for the detection, localization, and tracing of deepfake content. The code is available at https://github.com/vpsg‑research/LIDMark.
Authors:Bo Shi, Wei-ping Zhu, M. N. S. Swamy
Abstract:
Spatially variant dynamic convolution provides a principled approach of integrating spatial adaptivity into deep neural networks. However, mainstream designs in medical segmentation commonly generate dynamic kernels through average pooling, which implicitly collapses high‑frequency spatial details into a coarse, spatially‑compressed representation, leading to over‑smoothed predictions that degrade the fidelity of fine‑grained clinical structures. To address this limitation, we propose a novel Structure‑Guided Dynamic Convolution (SGDC) mechanism, which leverages an explicitly supervised structure‑extraction branch to guide the generation of dynamic kernels and gating signals for structure‑aware feature modulation. Specifically, the high‑fidelity boundary information from this auxiliary branch is fused with semantic features to enable spatially‑precise feature modulation. By replacing context aggregation with pixel‑wise structural guidance, the proposed design effectively prevents the information loss introduced by average pooling. Experimental results show that SGDC achieves state‑of‑the‑art performance on ISIC 2016, PH2, ISIC 2018, and CoNIC datasets, delivering superior boundary fidelity by reducing the Hausdorff Distance (HD95) by 2.05, and providing consistent IoU gains of 0.99%‑1.49% over pooling‑based baselines. Moreover, the mechanism exhibits strong potential for extension to other fine‑grained, structure‑sensitive vision tasks, such as small‑object detection, offering a principled solution for preserving structural integrity in medical image analysis. To facilitate reproducibility and encourage further research, the implementation code for both our SGE and SGDC modules has been is publicly released at https://github.com/solstice0621/SGDC.
Authors:Shuang Li, Yibing Wang, Yu Zhang, Changhui Li
Abstract:
Here we present a comprehensive derivation of the analytical expression for the spatiotemporal acoustic pressure generated by photoacoustic sources with spherically symmetric initial pressure distributions. Starting from the fundamental photoacoustic wave equation, we derive a unified analytical solution applicable to arbitrary spherically symmetric initial distributions. Specific expressions are provided for several common distributions including uniform spherical sources, Gaussian distributions, exponential distributions, and power‑law distributions. Far‑field approximations are also discussed. The derived expressions provide valuable tools for photoacoustic imaging system design and signal analysis. We provide codes for ultrafast forward simulation using the general analytical spherically symmetric model, the implementation is available in the GitHub repository: \hrefhttps://github.com/JaegerCQ/SlingBAG_Ultra.
Authors:Yiran Guan, Sifan Tu, Dingkang Liang, Linghao Zhu, Jianzhong Ju, Zhenbo Luo, Jian Luan, Yuliang Liu, Xiang Bai
Abstract:
Omni‑modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni‑modal large language models (OLLM) excel at perceiving diverse modalities, they lack the complex reasoning abilities of recent large reasoning models (LRM). However, enhancing the reasoning ability of OLLMs through additional training presents significant challenges, including the need for high‑quality data, task‑specific adaptation, and substantial computational costs. To address these limitations, we propose ThinkOmni, a training‑free and data‑free framework that lifts textual reasoning to omni‑modal scenarios. ThinkOmni introduces two key components: 1) LRM‑as‑a‑Guide, which leverages off‑the‑shelf LRMs to guide the OLLM decoding process; 2) Stepwise Contrastive Scaling, which adaptively balances perception and reasoning signals without manual hyperparameter tuning. Experiments on six multi‑modal reasoning benchmarks demonstrate that ThinkOmni consistently delivers performance improvements, with main results achieving 70.2 on MathVista and 75.5 on MMAU. Overall, ThinkOmni offers a flexible and generalizable solution for omni‑modal reasoning and provides new insights into the generalization and application of reasoning capabilities. Code is publicly available at https://github.com/1ranGuan/thinkomni
Authors:Thomas Woergaard, Raghavendra Selvan
Abstract:
Compressing neural networks by quantizing model parameters offers useful trade‑off between performance and efficiency. Methods like quantization‑aware training and post‑training quantization strive to maintain the downstream performance of compressed models compared to the full precision models. However, these techniques do not explicitly consider the impact on algorithmic fairness. In this work, we study fairness‑aware mixed‑precision quantization schemes for medical image classification under explicit bit budgets. We introduce FairQuant, a framework that combines group‑aware importance analysis, budgeted mixed‑precision allocation, and a learnable Bit‑Aware Quantization (BAQ) mode that jointly optimizes weights and per‑unit bit allocations under bitrate and fairness regularization. We evaluate the method on Fitzpatrick17k and ISIC2019 across ResNet18/50, DeiT‑Tiny, and TinyViT. Results show that FairQuant configurations with average precision near 4‑6 bits recover much of the Uniform 8‑bit accuracy while improving worst‑group performance relative to Uniform 4‑ and 8‑bit baselines, with comparable fairness metrics under shared budgets.
Authors:Zhaochen Su, Jincheng Gao, Hangyu Guo, Zhenhua Liu, Lueyang Zhang, Xinyu Geng, Shijue Huang, Peng Xia, Guanyu Jiang, Cheng Wang, Yue Zhang, Yi R. Fung, Junxian He
Abstract:
Real‑world multimodal agents solve multi‑step workflows grounded in visual evidence. For example, an agent can troubleshoot a device by linking a wiring photo to a schematic and validating the fix with online documentation, or plan a trip by interpreting a transit map and checking schedules under routing constraints. However, existing multimodal benchmarks mainly evaluate single‑turn visual reasoning or specific tool skills, and they do not fully capture the realism, visual subtlety, and long‑horizon tool use that practical agents require. We introduce AgentVista, a benchmark for generalist multimodal agents that spans 25 sub‑domains across 7 categories, pairing realistic and detail‑rich visual scenarios with natural hybrid tool use. Tasks require long‑horizon tool interactions across modalities, including web search, image search, page navigation, and code‑based operations for both image processing and general programming. Comprehensive evaluation of state‑of‑the‑art models exposes significant gaps in their ability to carry out long‑horizon multimodal tool use. Even the best model in our evaluation, Gemini‑3‑Pro with tools, achieves only 27.3% overall accuracy, and hard instances can require more than 25 tool‑calling turns. We expect AgentVista to accelerate the development of more capable and reliable multimodal agents for realistic and ultra‑challenging problem solving.
Authors:Guofeng Mei, Wei Lin, Luigi Riz, Yujiao Wu, Yiming Wang, Fabio Poiesi
Abstract:
Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre‑trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate such encoders for efficiency and scalability, extending this paradigm to 3D remains challenging due to the unordered and large‑scale nature of point clouds. This leaves a critical unanswered question: How can we design an LMM that tokenizes unordered 3D data effectively and efficiently without a cumbersome encoder? We propose Fase3D, the first efficient encoder‑free Fourier‑based 3D scene LMM. Fase3D tackles the challenges of scalability and permutation invariance with a novel tokenizer that combines point cloud serialization and the Fast Fourier Transform (FFT) to approximate self‑attention. This design enables an effective and computationally minimal architecture, built upon three key innovations: First, we represent large scenes compactly via structured superpoints. Second, our space‑filling curve serialization followed by an FFT enables efficient global context modeling and graph‑based token merging. Lastly, our Fourier‑augmented LoRA adapters inject global frequency‑aware interactions into the LLMs at a negligible cost. Fase3D achieves performance comparable to encoder‑based 3D LMMs while being significantly more efficient in computation and parameters. Project website: https://tev‑fbk.github.io/Fase3D.
Authors:Xiaosen Wang, Zhijin Ge, Bohan Liu, Zheng Fang, Fengfan Zhou, Ruixuan Zhang, Shaokang Wang, Yuyang Luo
Abstract:
Adversarial transferability refers to the capacity of adversarial examples generated on the surrogate model to deceive alternate, unexposed victim models. This property eliminates the need for direct access to the victim model during an attack, thereby raising considerable security concerns in practical applications and attracting substantial research attention recently. In this work, we discern a lack of a standardized framework and criteria for evaluating transfer‑based attacks, leading to potentially biased assessments of existing approaches. To rectify this gap, we have conducted an exhaustive review of hundreds of related works, organizing various transfer‑based attacks into six distinct categories. Subsequently, we propose a comprehensive framework designed to serve as a benchmark for evaluating these attacks. In addition, we delineate common strategies that enhance adversarial transferability and highlight prevalent issues that could lead to unfair comparisons. Finally, we provide a brief review of transfer‑based attacks beyond image classification.
Authors:Xudong Yan, Songhe Feng, Jiaxin Wang, Xin Su, Yi Jin
Abstract:
Compositional Zero‑Shot Learning (CZSL) aims to recognize novel attribute‑object compositions based on the knowledge learned from seen ones. Existing methods suffer from performance degradation caused by the distribution shift of label space at test time, which stems from the inclusion of unseen compositions recombined from attributes and objects. To overcome the challenge, we propose a novel approach that accumulates comprehensive knowledge in both textual and visual modalities from unsupervised data to update multimodal prototypes at test time. Building on this, we further design an adaptive update weight to control the degree of prototype adjustment, enabling the model to flexibly adapt to distribution shift during testing. Moreover, a dynamic priority queue is introduced that stores high‑confidence images to acquire visual prototypes from historical images for inference. Since the model tends to favor compositions already stored in the queue during testing, we warm‑start the queue by initializing it with training images for visual prototypes of seen compositions and generating unseen visual prototypes using the mapping learned between seen and unseen textual prototypes. Considering the semantic consistency of multimodal knowledge, we align textual and visual prototypes by multimodal collaborative representation learning. To provide a more reliable evaluation for CZSL, we introduce a new benchmark dataset, C‑Fashion, and refine the widely used but noisy MIT‑States dataset. Extensive experiments indicate that our approach achieves state‑of‑the‑art performance on four benchmark datasets under both closed‑world and open‑world settings. The source code and datasets are available at https://github.com/xud‑yan/WARM‑CAT .
Authors:Zeyu Zhang, Danning Li, Ian Reid, Richard Hartley
Abstract:
Energy‑based predictive world models provide a powerful approach for multi‑step visual planning by reasoning over latent energy landscapes rather than generating pixels. However, existing approaches face two major challenges: (i) their latent representations are typically learned in Euclidean space, neglecting the underlying geometric and hierarchical structure among states, and (ii) they struggle with long‑horizon prediction, which leads to rapid degradation across extended rollouts. To address these challenges, we introduce GeoWorld, a geometric world model that preserves geometric structure and hierarchical relations through a Hyperbolic JEPA, which maps latent representations from Euclidean space onto hyperbolic manifolds. We further introduce Geometric Reinforcement Learning for energy‑based optimization, enabling stable multi‑step planning in hyperbolic latent space. Extensive experiments on CrossTask and COIN demonstrate around 3% SR improvement in 3‑step planning and 2% SR improvement in 4‑step planning compared to the state‑of‑the‑art V‑JEPA 2. Project website: https://steve‑zeyu‑zhang.github.io/GeoWorld.
Authors:Argo Saakyan, Dmitry Solntsev
Abstract:
Transformer‑based real‑time object detectors achieve strong accuracy‑latency trade‑offs, and D‑FINE is among the top‑performing recent architectures. However, real‑time instance segmentation with transformers is still less common. We present D‑FINE‑seg, an instance segmentation extension of D‑FINE that adds: a lightweight mask head, segmentation‑aware training, including box cropped BCE and dice mask losses, auxiliary and denoising mask supervision, and adapted Hungarian matching cost. On the TACO dataset, D‑FINE‑seg improves F1‑score over Ultralytics YOLO26 under a unified TensorRT FP16 end‑to‑end benchmarking protocol, while maintaining competitive latency. Second contribution is an end‑to‑end pipeline for training, exporting, and optimized inference across ONNX, TensorRT, OpenVINO for both object detection and instance segmentation tasks. This framework is released as open‑source under the Apache‑2.0 license. GitHub repository ‑ https://github.com/ArgoHA/D‑FINE‑seg.
Authors:Tianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao, Yifan Xu, Haiyun Guo, Jinqiao Wang
Abstract:
Zero‑Shot Composed Image Retrieval (ZS‑CIR) aims to retrieve target images given a multimodal query (comprising a reference image and a modification text), without training on annotated triplets. Existing methods typically convert the multimodal query into a single modality‑either as an edited caption for Text‑to‑Image retrieval (T2I) or as an edited image for Image‑to‑Image retrieval (I2I). However, each paradigm has inherent limitations: T2I often loses fine‑grained visual details, while I2I struggles with complex semantic modifications. To effectively leverage their complementary strengths under diverse query intents, we propose WISER, a training‑free framework that unifies T2I and I2I via a "retrieve‑verify‑refine" pipeline, explicitly modeling intent awareness and uncertainty awareness. Specifically, WISER first performs Wider Search by generating both edited captions and images for parallel retrieval to broaden the candidate pool. Then, it conducts Adaptive Fusion with a verifier to assess retrieval confidence, triggering refinement for uncertain retrievals, and dynamically fusing the dual‑path for reliable ones. For uncertain retrievals, WISER generates refinement suggestions through structured self‑reflection to guide the next retrieval round toward Deeper Thinking. Extensive experiments demonstrate that WISER significantly outperforms previous methods across multiple benchmarks, achieving relative improvements of 45% on CIRCO (mAP@5) and 57% on CIRR (Recall@1) over existing training‑free methods. Notably, it even surpasses many training‑dependent methods, highlighting its superiority and generalization under diverse scenarios. Code will be released at https://github.com/Physicsmile/WISER.
Authors:Xinglong Luo, Ao Luo, Zhengning Wang, Yueqi Yang, Chaoyu Feng, Lei Lei, Bing Zeng, Shuaicheng Liu
Abstract:
Image alignment is a fundamental task in computer vision with broad applications. Existing methods predominantly employ optical flow‑based image warping. However, this technique is susceptible to common challenges such as occlusions and illumination variations, leading to degraded alignment visual quality and compromised accuracy in downstream tasks. In this paper, we present DMAligner, a diffusion‑based framework for image alignment through alignment‑oriented view synthesis. DMAligner is crafted to tackle the challenges in image alignment from a new perspective, employing a generation‑based solution that showcases strong capabilities and avoids the problems associated with flow‑based image warping. Specifically, we propose a Dynamics‑aware Diffusion Training approach for learning conditional image generation, synthesizing a novel view for image alignment. This incorporates a Dynamics‑aware Mask Producing (DMP) module to adaptively distinguish dynamic foreground regions from static backgrounds, enabling the diffusion model to more effectively handle challenges that classical methods struggle to solve. Furthermore, we develop the Dynamic Scene Image Alignment (DSIA) dataset using Blender, which includes 1,033 indoor and outdoor scenes with over 30K image pairs tailored for image alignment. Extensive experimental results demonstrate the superiority of the proposed approach on DSIA benchmarks, as well as on a series of widely‑used video datasets for qualitative comparisons. Our code is available at https://github.com/boomluo02/DMAligner.
Authors:Camile Lendering, Erkut Akdag, Egor Bondarev
Abstract:
Detecting visual anomalies in industrial inspection often requires training with only a few normal images per category. Recent few‑shot methods achieve strong results employing foundation‑model features, but typically rely on memory banks, auxiliary datasets, or multi‑modal tuning of vision‑language models. We therefore question whether such complexity is necessary given the feature representations of vision foundation models. To answer this question, we introduce SubspaceAD, a training‑free method, that operates in two simple stages. First, patch‑level features are extracted from a small set of normal images by a frozen DINOv2 backbone. Second, a Principal Component Analysis (PCA) model is fit to these features to estimate the low‑dimensional subspace of normal variations. At inference, anomalies are detected via the reconstruction residual with respect to this subspace, producing interpretable and statistically grounded anomaly scores. Despite its simplicity, SubspaceAD achieves state‑of‑the‑art performance across one‑shot and few‑shot settings without training, prompt tuning, or memory banks. In the one‑shot anomaly detection setting, SubspaceAD achieves image‑level and pixel‑level AUROC of 97.1% and 97.5% on the MVTec‑AD dataset, and 93.2% and 98.2% on the VisA dataset, respectively, surpassing prior state‑of‑the‑art results. Code and demo are available at https://github.com/CLendering/SubspaceAD.
Authors:Gorkem Yildiz
Abstract:
We present Helmlab, a family of two purpose‑built color spaces for UI design systems sharing a common 11‑stage analytical structure: MetricSpace, a 72‑parameter space optimized for color‑difference prediction, and GenSpace, a 44‑parameter space optimized for gradient and palette generation. The forward transform maps CIE XYZ to a perceptually‑organized Lab representation through learned matrices, per‑channel power compression, Fourier hue correction, and embedded Helmholtz‑Kohlrausch lightness adjustment. A post‑pipeline neutral correction holds gray‑axis chroma below 1e‑5 on a 21‑step ramp, and a rigid rotation of the chromatic plane improves hue‑angle alignment without affecting the distance metric (which is invariant under isometries).
On COMBVD (3,813 color pairs), MetricSpace v21 achieves STRESS 22.48, a 23 percent reduction from CIEDE2000 (29.20). On the held‑out MacAdam 1974 dataset it scores 19.51 (CIEDE2000: 22.13; CAM16‑UCS leads at 18.71). On a self‑collected 3,552‑judgement screen‑condition set it scores 23.26 vs 62.54 for CIEDE2000. On academic He et al. 2022 (82 3D‑printed pairs) MetricSpace scores 35.9 vs CIEDE2000 32.6, a regression we own. Averaging the three primary datasets, MetricSpace scores 21.75 vs the next‑best baseline CIECAM02‑UCS at 35.98.
GenSpace v0.11.1 trades distance accuracy for generation quality: on a 90‑metric, 3,038‑pair gradient/palette benchmark across sRGB, P3, and Rec.2020, it wins 65 of 90 vs OKLab. The transform is invertible with round‑trip errors below 1e‑13. Production implementations ship on PyPI, npm, Color.js (PR 722, merged), and as a PostCSS plugin.
Authors:Tianxing Xu, Zixuan Wang, Guangyuan Wang, Li Hu, Zhongyi Zhang, Peng Zhang, Bang Zhang, Song-Hai Zhang
Abstract:
World models based on video generation demonstrate remarkable potential for simulating interactive environments but face persistent difficulties in two key areas: maintaining long‑term content consistency when scenes are revisited and enabling precise camera control from user‑provided inputs. Existing methods based on explicit 3D reconstruction often compromise flexibility in unbounded scenarios and fine‑grained structures. Alternative methods rely directly on previously generated frames without establishing explicit spatial correspondence, thereby constraining controllability and consistency. To address these limitations, we present UCM, a novel framework that unifies long‑term memory and precise camera control via a time‑aware positional encoding warping mechanism. To reduce computational overhead, we design an efficient dual‑stream diffusion transformer for high‑fidelity generation. Moreover, we introduce a scalable data curation strategy utilizing point‑cloud‑based rendering to simulate scene revisiting, facilitating training on over 500K monocular videos. Extensive experiments on real‑world and synthetic benchmarks demonstrate that UCM significantly outperforms state‑of‑the‑art methods in long‑term scene consistency, while also achieving precise camera controllability in high‑fidelity video generation.
Authors:Zihao Zhao, Frederik Hauke, Juliana De Castilhos, Sven Nebelung, Daniel Truhn
Abstract:
The rapid progress of multimodal large language models (MLLMs) has led to increasing interest in agent‑based systems. While most prior work in medical imaging concentrates on automating routine clinical workflows, we study an underexplored yet clinically significant setting: distinguishing visually hard‑to‑separate diseases in a zero‑shot setting. We benchmark representative agents on two imaging‑only proxy diagnostic tasks, (1) melanoma vs. atypical nevus and (2) pulmonary edema vs. pneumonia, where visual features are highly confounded despite substantial differences in clinical management. We introduce a multi‑agent framework based on contrastive adjudication. Experimental results show improved diagnostic performance (an 11‑percentage‑point gain in accuracy on dermoscopy data) and reduced unsupported claims on qualitative samples, although overall performance remains insufficient for clinical deployment. We acknowledge the inherent uncertainty in human annotations and the absence of clinical context, which further limit the translation to real‑world settings. Within this controlled setting, this pilot study provides preliminary insights into zero‑shot agent performance in visually confounded scenarios.
Authors:Feng Guo, Jiaxiang Liu, Yang Li, Qianqian Shi, Mingkun Xu
Abstract:
Accurate brain tumor diagnosis requires models to not only detect lesions but also generate clinically interpretable reasoning grounded in imaging manifestations, yet existing public datasets remain limited in annotation richness and diagnostic semantics. To bridge this gap, we introduce MM‑NeuroOnco, a large‑scale multimodal benchmark and instruction‑tuning dataset for brain tumor MRI understanding, consisting of 24,726 MRI slices from 20 data sources paired with approximately 200,000 semantically enriched multimodal instructions spanning diverse tumor subtypes and imaging modalities. To mitigate the scarcity and high cost of diagnostic semantic annotations, we develop a multi‑model collaborative pipeline for automated medical information completion and quality control, enabling the generation of diagnosis‑related semantics beyond mask‑only annotations. Building upon this dataset, we further construct MM‑NeuroOnco‑Bench, a manually annotated evaluation benchmark with a rejection‑aware setting to reduce biases inherent in closed‑ended question formats. Evaluation across ten representative models shows that even the strongest baseline, Gemini 3 Flash, achieves only 41.88% accuracy on diagnosis‑related questions, highlighting the substantial challenges of multimodal brain tumor diagnostic understanding. Leveraging MM‑NeuroOnco, we further propose NeuroOnco‑GPT, which achieves a 27% absolute accuracy improvement on diagnostic questions following fine‑tuning. This result demonstrates the effectiveness of our dataset and benchmark in advancing clinically grounded multimodal diagnostic reasoning. Code and dataset are publicly available at: https://github.com/gfnnnb/MM‑NeuroOnco
Authors:Junuk Cha, Jihyeon Kim, Han-Mu Park
Abstract:
Fingerspelling is a component of sign languages in which words are spelled out letter by letter using specific hand poses. Automatic fingerspelling recognition plays a crucial role in bridging the communication gap between Deaf and hearing communities, yet it remains challenging due to the signing‑hand ambiguity issue, the lack of appropriate training losses, and the out‑of‑vocabulary (OOV) problem. Prior fingerspelling recognition methods rely on explicit signing‑hand detection, which often leads to recognition failures, and on a connectionist temporal classification (CTC) loss, which exhibits the peaky behavior problem. To address these issues, we develop OpenFS, an open‑source approach for fingerspelling recognition and synthesis. We propose a multi‑hand‑capable fingerspelling recognizer that supports both single‑ and multi‑hand inputs and performs implicit signing‑hand detection by incorporating a dual‑level positional encoding and a signing‑hand focus (SF) loss. The SF loss encourages cross‑attention to focus on the signing hand, enabling implicit signing‑hand detection during recognition. Furthermore, without relying on the CTC loss, we introduce a monotonic alignment (MA) loss that enforces the output letter sequence to follow the temporal order of the input pose sequence through cross‑attention regularization. In addition, we propose a frame‑wise letter‑conditioned generator that synthesizes realistic fingerspelling pose sequences for OOV words. This generator enables the construction of a new synthetic benchmark, called FSNeo. Through comprehensive experiments, we demonstrate that our approach achieves state‑of‑the‑art performance in recognition and validate the effectiveness of the proposed recognizer and generator. Codes and data are available in: https://github.com/AIRC‑KETI/OpenFS.
Authors:Hongzhao Li, Hao Dong, Hualei Wan, Shupan Li, Mingliang Xu, Muhammad Haris Khan
Abstract:
Multimodal models ideally should generalize to unseen domains while remaining data‑efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi‑Supervised Multimodal Domain Generalization (SSMDG), which aims to learn robust multimodal models from multi‑source data with few labeled samples. We observe that existing approaches fail to address this setting effectively: multimodal domain generalization methods cannot exploit unlabeled data, semi‑supervised multimodal learning methods ignore domain shifts, and semi‑supervised domain generalization methods are confined to single‑modality inputs. To overcome these limitations, we propose a unified framework featuring three key components: Consensus‑Driven Consistency Regularization, which obtains reliable pseudo‑labels through confident fused‑unimodal consensus; Disagreement‑Aware Regularization, which effectively utilizes ambiguous non‑consensus samples; and Cross‑Modal Prototype Alignment, which enforces domain‑ and modality‑invariant representations while promoting robustness under missing modalities via cross‑modal translation. We further establish the first SSMDG benchmarks, on which our method consistently outperforms strong baselines in both standard and missing‑modality scenarios. Our benchmarks and code are available at https://github.com/lihongzhao99/SSMDG.
Authors:Hongrui Jia, Chaoya Jiang, Yongrui Heng, Shikun Zhang, Wei Ye
Abstract:
As Large Multimodal Models (LMMs) scale up and reinforcement learning (RL) methods mature, LMMs have made notable progress in complex reasoning and decision making. Yet training still relies on static data and fixed recipes, making it difficult to diagnose capability blind spots or provide dynamic, targeted reinforcement. Motivated by findings that test driven error exposure and feedback based correction outperform repetitive practice, we propose Diagnostic‑driven Progressive Evolution (DPE), a spiral loop where diagnosis steers data generation and reinforcement, and each iteration re‑diagnoses the updated model to drive the next round of targeted improvement. DPE has two key components. First, multiple agents annotate and quality control massive unlabeled multimodal data, using tools such as web search and image editing to produce diverse, realistic samples. Second, DPE attributes failures to specific weaknesses, dynamically adjusts the data mixture, and guides agents to generate weakness focused data for targeted reinforcement. Experiments on Qwen3‑VL‑8B‑Instruct and Qwen2.5‑VL‑7B‑Instruct show stable, continual gains across eleven benchmarks, indicating DPE as a scalable paradigm for continual LMM training under open task distributions. Our code, models, and data are publicly available at https://github.com/hongruijia/DPE.
Authors:Mingde Yao, Zhiyuan You, King-Man Tam, Menglu Wang, Tianfan Xue
Abstract:
With the recent fast development of generative models, instruction‑based image editing has shown great potential in generating high‑quality images. However, the quality of editing highly depends on carefully designed instructions, placing the burden of task decomposition and sequencing entirely on the user. To achieve autonomous image editing, we present PhotoAgent, a system that advances image editing through explicit aesthetic planning. Specifically, PhotoAgent formulates autonomous image editing as a long‑horizon decision‑making problem. It reasons over user aesthetic intent, plans multi‑step editing actions via tree search, and iteratively refines results through closed‑loop execution with memory and visual feedback, without requiring step‑by‑step user prompts. To support reliable evaluation in real‑world scenarios, we introduce UGC‑Edit, an aesthetic evaluation benchmark consisting of 7,000 photos and a learned aesthetic reward model. We also construct a test set containing 1,017 photos to systematically assess autonomous photo editing performance. Extensive experiments demonstrate that PhotoAgent consistently improves both instruction adherence and visual quality compared with baseline methods. The project page is https://mdyao.github.io/PhotoAgent/.
Authors:Hanliang Du, Zhangji Lu, Zewei Cai, Qijian Tang, Qifeng Yu, Xiaoli Liu
Abstract:
Atmospheric turbulence causes significant image degradation due to pixel displacement (tilt) and blur, particularly in long‑range imaging applications. In this paper, we propose a novel framework for atmospheric turbulence mitigation, GSTurb, which integrates optical flow‑guided tilt correction and Gaussian splatting for modeling non‑isoplanatic blur. The framework employs Gaussian parameters to represent tilt and blur, and optimizes them across multiple frames to enhance restoration. Experimental results on the ATSyn‑static dataset demonstrate the effectiveness of our method, achieving a peak PSNR of 27.67 dB and SSIM of 0.8735. Compared to the state‑of‑the‑art method, GSTurb improves PSNR by 1.3 dB (a 4.5% increase) and SSIM by 0.048 (a 5.8% increase). Additionally, on real datasets, including the TSRWGAN Real‑World and CLEAR datasets, GSTurb outperforms existing methods, showing significant improvements in both qualitative and quantitative performance. These results highlight that combining optical flow‑guided tilt correction with Gaussian splatting effectively enhances image restoration under both synthetic and real‑world turbulence conditions. The code for this method will be available at https://github.com/DuhlLiamz/3DGS_turbulence/tree/main.
Authors:Ling Wang, Hao-Xiang Guo, Xinzhou Wang, Fuchun Sun, Kai Sun, Pengkun Liu, Hang Xiao, Zhong Wang, Guangyuan Fu, Eric Li, Yang Liu, Yikai Wang
Abstract:
We introduce SceneTransporter, an end‑to‑end framework for structured 3D scene generation from a single image. While existing methods generate part‑level 3D objects, they often fail to organize these parts into distinct instances in open‑world scenes. Through a debiased clustering probe, we reveal a critical insight: this failure stems from the lack of structural constraints within the model's internal assignment mechanism. Based on this finding, we reframe the task of structured 3D scene generation as a global correlation assignment problem. To solve this, SceneTransporter formulates and solves an entropic Optimal Transport (OT) objective within the denoising loop of the compositional DiT model. This formulation imposes two powerful structural constraints. First, the resulting transport plan gates cross‑attention to enforce an exclusive, one‑to‑one routing of image patches to part‑level 3D latents, preventing entanglement. Second, the competitive nature of the transport encourages the grouping of similar patches, a process that is further regularized by an edge‑based cost, to form coherent objects and prevent fragmentation. Extensive experiments show that SceneTransporter outperforms existing methods on open‑world scene generation, significantly improving instance‑level coherence and geometric fidelity. Code and models will be publicly available at https://2019epwl.github.io/SceneTransporter/.
Authors:Fengming Liu, Tat-Jen Cham, Chuanxia Zheng
Abstract:
Most text‑to‑video (T2V) generators prioritize aesthetic quality, but often ignoring the spatial constraints in the generated videos. In this work, we present SPATIALALIGN, a self‑improvement framework that enhances T2V models capabilities to depict Dynamic Spatial Relationships (DSR) specified in text prompts. We present a zeroth‑order regularized Direct Preference Optimization (DPO) to fine‑tune T2V models towards better alignment with DSR. Specifically, we design DSR‑SCORE, a geometry‑based metric that quantitatively measures the alignment between generated videos and the specified DSRs in prompts, which is a step forward from prior works that rely on VLM for evaluation. We also conduct a dataset of text‑video pairs with diverse DSRs to facilitate the study. Extensive experiments demonstrate that our fine‑tuned model significantly out performs the baseline in spatial relationships. The code will be released in Link. Project page: https://fengming001ntu.github.io/SpatialAlign/
Authors:Tongfei Chen, Shuo Yang, Yuguang Yang, Linlin Yang, Runtang Guo, Changbai Li, He Long, Chunyu Xie, Dawei Leng, Baochang Zhang
Abstract:
Referring Image Segmentation (RIS) aims to segment the object in an image uniquely referred to by a natural language expression. However, RIS training often contains hard‑to‑align and instance‑specific visual signals; optimizing on such pixels injects misleading gradients and drives the model in the wrong direction. By explicitly estimating pixel‑level vision‑language alignment, the learner can suppress low‑alignment regions, concentrate on reliable cues, and acquire more generalizable alignment features.
In this paper, we propose Alignment‑Aware Masked Learning (AML), a simple yet effective training strategy that quantifies region‑referent alignment (PMME) and filters out unreliable pixels during optimization (AFM). Specifically, each sample first computes a similarity map between visual and textual features, and then masks out pixels falling below an adaptive similarity threshold, thereby excluding poorly aligned regions from the training process. AML does not require architectural changes and incurs no inference overhead, directing attention to the areas aligned with the textual description. Experiments on the RefCOCO (vanilla/+/g) datasets show that AML achieves state‑of‑the‑art results across all 8 splits, and beyond improving RIS performance, AML also enhances the model's robustness to diverse descriptions and scenarios. Code is available at https://github.com/pipashu1/AMLRIS.
Authors:Muzi Tao, Chufan Shi, Huijuan Wang, Shengbang Tong, Xuezhe Ma
Abstract:
In this work, we study idiosyncrasies in the caption models and their downstream impact on text‑to‑image models. We design a systematic analysis: given either a generated caption or the corresponding image, we train neural networks to predict the originating caption model. Our results show that text classification yields very high accuracy (99.70%), indicating that captioning models embed distinctive stylistic signatures. In contrast, these signatures largely disappear in the generated images, with classification accuracy dropping to at most 50% even for the state‑of‑the‑art Flux model. To better understand this cross‑modal discrepancy, we further analyze the data and find that the generated images fail to preserve key variations present in captions, such as differences in the level of detail, emphasis on color and texture, and the distribution of objects within a scene. Overall, our classification‑based framework provides a novel methodology for quantifying both the stylistic idiosyncrasies of caption models and the prompt‑following ability of text‑to‑image systems.
Authors:Changqing Zhou, Yueru Luo, Han Zhang, Zeyu Jiang, Changhao Chen
Abstract:
Open‑vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abundant and evolve beyond fixed taxonomies. While recent work has explored open‑vocabulary occupancy in outdoor driving scenarios, such methods transfer poorly indoors, where geometry is denser, layouts are more intricate, and semantics are far more fine‑grained. To address these challenges, we adopt a geometry‑only supervision paradigm that uses only binary occupancy labels (occupied vs free). Our framework builds upon 3D Language‑Embedded Gaussians, which serve as a unified intermediate representation coupling fine‑grained 3D geometry with a language‑aligned semantic embedding. On the geometry side, we find that existing Gaussian‑to‑Occupancy operators fail to converge under such weak supervision, and we introduce an opacity‑aware, Poisson‑based approach that stabilizes volumetric aggregation. On the semantic side, direct alignment between rendered features and open‑vocabulary segmentation features suffers from feature mixing; we therefore propose a Progressive Temperature Decay schedule that gradually sharpens opacities during splatting, strengthening Gaussian‑language alignment. On Occ‑ScanNet, our framework achieves 59.50 IoU and 21.05 mIoU in the open‑vocabulary setting, surpassing all existing occupancy methods in IoU and outperforming prior open‑vocabulary approaches by a large margin in mIoU. Code will be released at https://github.com/JuIvyy/LegoOcc.
Authors:Renyu Yang, Jian Jin, Lili Meng, Meiqin Liu, Yilin Wang, Balu Adsumilli, Weisi Lin
Abstract:
Audio‑visual quality assessment (AVQA) research has been stalled by limitations of existing datasets: they are typically small in scale, with insufficient diversity in content and quality, and annotated only with overall scores. These shortcomings provide limited support for model development and multimodal perception research. We propose a practical approach for AVQA dataset construction. First, we design a crowdsourced subjective experiment framework for AVQA, breaks the constraints of in‑lab settings and achieves reliable annotation across varied environments. Second, a systematic data preparation strategy is further employed to ensure broad coverage of both quality levels and semantic scenarios. Third, we extend the dataset with additional annotations, enabling research on multimodal perception mechanisms and their relation to content. Finally, we validate this approach through YT‑NTU‑AVQ, the largest and most diverse AVQA dataset to date, consisting of 1,620 user‑generated audio and video (A/V) sequences. The dataset and platform code are available at https://github.com/renyu12/YT‑NTU‑AVQ
Authors:Bowen Cui, Yuanbin Wang, Huajiang Xu, Biaolong Chen, Aixi Zhang, Hao Jiang, Zhengzheng Jin, Xu Liu, Pipei Huang
Abstract:
Diffusion models have demonstrated remarkable success in image and video generation, yet their practical deployment remains hindered by the substantial computational overhead of multi‑step iterative sampling. Among acceleration strategies, caching‑based methods offer a training‑free and effective solution by reusing or predicting features across timesteps. However, existing approaches rely on fixed or locally adaptive schedules without considering the global structure of the denoising trajectory, often leading to error accumulation and visual artifacts. To overcome this limitation, we propose DPCache, a novel training‑free acceleration framework that formulates diffusion sampling acceleration as a global path planning problem. DPCache constructs a Path‑Aware Cost Tensor from a small calibration set to quantify the path‑dependent error of skipping timesteps conditioned on the preceding key timestep. Leveraging this tensor, DPCache employs dynamic programming to select an optimal sequence of key timesteps that minimizes the total path cost while preserving trajectory fidelity. During inference, the model performs full computations only at these key timesteps, while intermediate outputs are efficiently predicted using cached features. Extensive experiments on DiT, FLUX, and HunyuanVideo demonstrate that DPCache achieves strong acceleration with minimal quality loss, outperforming prior acceleration methods by +0.031 ImageReward at 4.87× speedup and even surpassing the full‑step baseline by +0.028 ImageReward at 3.54× speedup on FLUX, validating the effectiveness of our path‑aware global scheduling framework. Code is available at https://github.com/argsss/DPCache.
Authors:Woojae Hong, Jong Ha Hwang, Jiyong Chung, Joongyeon Choi, Hyunngun Kim, Yong Hwy Kim
Abstract:
Interactive Medical‑SAM2 GUI is an open‑source desktop application for semi‑automatic annotation of 2D and 3D medical images. Built on the Napari multi‑dimensional viewer, box/point prompting is integrated with SAM2‑style propagation by treating a 3D volume as a slice sequence, enabling mask propagation from sparse prompts using Medical‑SAM2 on top of SAM2. Voxel‑level annotation remains essential for developing and validating medical imaging algorithms, yet manual labeling is slow and expensive for 3D scans, and existing integrations frequently emphasize per‑slice interaction without providing a unified, cohort‑oriented workflow for navigation, propagation, interactive correction, and quantitative export in a single local pipeline. To address this practical limitation, a local‑first Napari workflow is provided for efficient 3D annotation across multiple studies using standard DICOM series and/or NIfTI volumes. Users can annotate cases sequentially under a single root folder with explicit proceed/skip actions, initialize objects via box‑first prompting (including first/last‑slice initialization for single‑object propagation), refine predictions with point prompts, and finalize labels through prompt‑first correction prior to saving. During export, per‑object volumetry and 3D volume rendering are supported, and image geometry is preserved via SimpleITK. The GUI is implemented in Python using Napari and PyTorch, with optional N4 bias‑field correction, and is intended exclusively for research annotation workflows. The code is released on the project page: https://github.com/SKKU‑IBE/Medical‑SAM2GUI/.
Authors:Boyang Dai, Zeng Fan, Zihao Qi, Meng Lou, Yizhou Yu
Abstract:
Source‑Free Domain Adaptive Object Detection (SF‑DAOD) aims to adapt a detector trained on a labeled source domain to an unlabeled target domain without retaining any source data. Despite recent progress, most popular approaches focus on tuning pseudo‑label thresholds or refining the teacher‑student framework, while overlooking object‑level structural cues within cross‑domain data. In this work, we present CGSA, the first framework that brings Object‑Centric Learning (OCL) into SF‑DAOD by integrating slot‑aware adaptation into the DETR‑based detector. Specifically, our approach integrates a Hierarchical Slot Awareness (HSA) module into the detector to progressively disentangle images into slot representations that act as visual priors. These slots are then guided toward class semantics via a Class‑Guided Slot Contrast (CGSC) module, maintaining semantic consistency and prompting domain‑invariant adaptation. Extensive experiments on multiple cross‑domain datasets demonstrate that our approach outperforms previous SF‑DAOD methods, with theoretical derivations and experimental analysis further demonstrating the effectiveness of the proposed components and the framework, thereby indicating the promise of object‑centric design in privacy‑sensitive adaptation scenarios. Code is released at https://github.com/Michael‑McQueen/CGSA.
Authors:Minh Kha Do, Wei Xiang, Kang Han, Di Wu, Khoa Phan, Yi-Ping Phoebe Chen, Gaowen Liu, Ramana Rao Kompella
Abstract:
Vision‑language foundation models (VLFMs) promise zero‑shot and retrieval understanding for Earth observation. While operational satellite systems often lack full multi‑spectral coverage, making RGB‑only inference highly desirable for scalable deployment, the adoption of VLFMs for satellite imagery remains hindered by two factors: (1) multi‑spectral inputs are informative but difficult to exploit consistently due to band redundancy and misalignment; and (2) CLIP‑style text encoders limit semantic expressiveness and weaken fine‑grained alignment. We present SATtxt, a spectrum‑aware VLFM that operates with RGB inputs only at inference while retaining spectral cues learned during training. Our framework comprises two stages. First, Spectral Representation Distillation transfers spectral priors from a frozen multi‑spectral teacher to an RGB student via a lightweight projector. Second, Spectrally Grounded Alignment with Instruction‑Augmented LLMs bridges the distilled visual space and an expressive LLM embedding space. Across EuroSAT, BigEarthNet, and ForestNet, SATtxt improves zero‑shot classification on average by 4.2%, retrieval by 5.9%, and linear probing by 2.7% over baselines, showing an efficient path toward spectrum‑aware vision‑language learning for Earth observation. Project page: https://ikhado.github.io/sattxt/
Authors:Yibo Peng, Peng Xia, Ding Zhong, Kaide Zeng, Siwei Han, Yiyang Zhou, Jiaqi Liu, Ruiyi Zhang, Huaxiu Yao
Abstract:
Despite the rapid advancements in Multimodal Large Language Models (MLLMs), a critical question regarding their visual grounding mechanism remains unanswered: do these models genuinely ``read'' text embedded in images, or do they merely rely on parametric shortcuts in the text prompt? In this work, we diagnose this issue by introducing the Visualized‑Question (VQ) setting, where text queries are rendered directly onto images to structurally mandate visual engagement. Our diagnostic experiments on Qwen2.5‑VL reveal a startling capability‑utilization gap: despite possessing strong OCR capabilities, models suffer a performance degradation of up to 12.7% in the VQ setting, exposing a deep‑seated ``modality laziness.'' To bridge this gap, we propose SimpleOCR, a plug‑and‑play training strategy that imposes a structural constraint on the learning process. By transforming training samples into the VQ format with randomized styles, SimpleOCR effectively invalidates text‑based shortcuts, compelling the model to activate and optimize its visual text extraction pathways. Empirically, SimpleOCR yields robust gains without architectural modifications. On four representative OOD benchmarks, it surpasses the base model by 5.4% and GRPO based on original images by 2.7%, while exhibiting extreme data efficiency, achieving superior performance with 30x fewer samples (8.5K) than recent RL‑based methods. Furthermore, its plug‑and‑play nature allows seamless integration with advanced RL strategies like NoisyRollout to yield complementary improvements. Code is available at https://github.com/aiming‑lab/SimpleOCR.
Authors:Julian Kaltheuner, Hannah Dröge, Markus Plack, Patrick Stotko, Reinhard Klein
Abstract:
Temporally consistent surface reconstruction of dynamic 3D objects from unstructured point cloud data remains challenging, especially for very long sequences. Existing methods either optimize deformations incrementally, risking drift and requiring long runtimes, or rely on complex learned models that demand category‑specific training. We present Neu‑PiG, a fast deformation optimization method based on a novel preconditioned latent‑grid encoding that distributes spatial features parameterized on the position and normal direction of a keyframe surface. Our method encodes entire deformations across all time steps at various spatial scales into a multi‑resolution latent grid, parameterized by the position and normal direction of a reference surface from a single keyframe. This latent representation is then augmented for time modulation and decoded into per‑frame 6‑DoF deformations via a lightweight multilayer perceptron (MLP). To achieve high‑fidelity, drift‑free surface reconstructions in seconds, we employ Sobolev preconditioning during gradient‑based training of the latent space, completely avoiding the need for any explicit correspondences or further priors. Experiments across diverse human and animal datasets demonstrate that Neu‑PiG outperforms state‑the‑art approaches, offering both superior accuracy and scalability to long sequences while running at least 60x faster than existing training‑free methods and achieving inference speeds on the same order as heavy pretrained models.
Authors:Yufei Ye, Jiaman Li, Ryan Rong, C. Karen Liu
Abstract:
Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. Current methods typically focus on recovering either hand or object pose in isolation, but both struggle during interactions and fail to handle out‑of‑sight cases. Moreover, their independent predictions often lead to inconsistent hand‑object relations. We introduce WHOLE, a method that holistically reconstructs hand and object motion in world space from egocentric videos given object templates. Our key insight is to learn a generative prior over hand‑object motion to jointly reason about their interactions. At test time, the pretrained prior is guided to generate trajectories that conform to the video observations. This joint generative reconstruction substantially outperforms approaches that process hands and objects separately followed by post‑processing. WHOLE achieves state‑of‑the‑art performance on hand motion estimation, 6D object pose estimation, and their relative interaction reconstruction. Project website: https://judyye.github.io/whole‑www
Authors:Xavier Pleimling, Sifat Muhammad Abdullah, Gunjan Balde, Peng Gao, Mainack Mondal, Murtuza Jadliwala, Bimal Viswanath
Abstract:
Advances in Generative AI (GenAI) have led to the development of various protection strategies to prevent the unauthorized use of images. These methods rely on adding imperceptible protective perturbations to images to thwart misuse such as style mimicry or deepfake manipulations. Although previous attacks on these protections required specialized, purpose‑built methods, we demonstrate that this is no longer necessary. We show that off‑the‑shelf image‑to‑image GenAI models can be repurposed as generic ``denoisers" using a simple text prompt, effectively removing a wide range of protective perturbations. Across 8 case studies spanning 6 diverse protection schemes, our general‑purpose attack not only circumvents these defenses but also outperforms existing specialized attacks while preserving the image's utility for the adversary. Our findings reveal a critical and widespread vulnerability in the current landscape of image protection, indicating that many schemes provide a false sense of security. We stress the urgent need to develop robust defenses and establish that any future protection mechanism must be benchmarked against attacks from off‑the‑shelf GenAI models. Code is available in this repository: https://github.com/mlsecviswanath/img2imgdenoiser
Authors:Lingfeng Ren, Weihao Yu, Runpeng Yu, Xinchao Wang
Abstract:
Object hallucination is a critical issue in Large Vision‑Language Models (LVLMs), where outputs include objects that do not appear in the input image. A natural question arises from this phenomenon: Which component of the LVLM pipeline primarily contributes to object hallucinations? The vision encoder to perceive visual information, or the language decoder to generate text responses? In this work, we strive to answer this question through designing a systematic experiment to analyze the roles of the vision encoder and the language decoder in hallucination generation. Our observations reveal that object hallucinations are predominantly associated with the strong priors from the language decoder. Based on this finding, we propose a simple and training‑free framework, No‑Language‑Hallucination Decoding, NoLan, which refines the output distribution by dynamically suppressing language priors, modulated based on the output distribution difference between multimodal and text‑only inputs. Experimental results demonstrate that NoLan effectively reduces object hallucinations across various LVLMs on different tasks. For instance, NoLan achieves substantial improvements on POPE, enhancing the accuracy of LLaVA‑1.5 7B and Qwen‑VL 7B by up to 6.45 and 7.21, respectively. The code is publicly available at: https://github.com/lingfengren/NoLan.
Authors:Yuetan Chu, Xinhua Ma, Xinran Jin, Gongning Luo, Xin Gao
Abstract:
Medical vision‑language pretraining increasingly relies on medical reports as large‑scale supervisory signals; however, raw reports often exhibit substantial stylistic heterogeneity, variable length, and a considerable amount of image‑irrelevant content. Although text normalization is frequently adopted as a preprocessing step in prior work, its design principles and empirical impact on vision‑language pretraining remain insufficiently and systematically examined. In this study, we present MedTri, a deployable normalization framework for medical vision‑language pretraining that converts free‑text reports into a unified [Anatomical Entity: Radiologic Description + Diagnosis Category] triplet. This structured, anatomy‑grounded normalization preserves essential morphological and spatial information while removing stylistic noise and image‑irrelevant content, providing consistent and image‑grounded textual supervision at scale. Across multiple datasets spanning both X‑ray and computed tomography (CT) modalities, we demonstrate that structured, anatomy‑grounded text normalization is an important factor in medical vision‑language pretraining quality, yielding consistent improvements over raw reports and existing normalization baselines. In addition, we illustrate how this normalization can easily support modular text‑level augmentation strategies, including knowledge enrichment and anatomy‑grounded counterfactual supervision, which provide complementary gains in robustness and generalization without altering the core normalization process. Together, our results position structured text normalization as a critical and generalizable preprocessing component for medical vision‑language learning, while MedTri provides this normalization platform. Code and data will be released at https://github.com/Arturia‑Pendragon‑Iris/MedTri.
Authors:Yulin Zhang, Cheng Shi, Sibei Yang
Abstract:
Recent advances in Multimodal Large Language Models have greatly improved visual understanding and reasoning, yet their quadratic attention and offline training protocols make them ill‑suited for streaming settings where frames arrive sequentially and future observations are inaccessible. We diagnose a core limitation of current Video‑LLMs, namely Time‑Agnosticism, in which videos are treated as an unordered bag of evidence rather than a causally ordered sequence, yielding two failures in streams: temporal order ambiguity, in which the model cannot follow or reason over the correct chronological order, and past‑current focus blindness where it fails to distinguish present observations from accumulated history. We present WeaveTime, a simple, efficient, and model agnostic framework that first teaches order and then uses order. We introduce a lightweight Temporal Reconstruction objective‑our Streaming Order Perception enhancement‑that instills order aware representations with minimal finetuning and no specialized streaming data. At inference, a Past‑Current Dynamic Focus Cache performs uncertainty triggered, coarse‑to‑fine retrieval, expanding history only when needed. Plugged into exsiting Video‑LLM without architectural changes, WeaveTime delivers consistent gains on representative streaming benchmarks, improving accuracy while reducing latency. These results establish WeaveTime as a practical path toward time aware stream Video‑LLMs under strict online, time causal constraints. Code and weights will be made publicly available. Project Page: https://zhangyl4.github.io/publications/weavetime/
Authors:Abhipsa Basu, Mohana Singh, Shashank Agnihotri, Margret Keuper, R. Venkatesh Babu
Abstract:
Text‑to‑image (T2I) models are rapidly gaining popularity, yet their outputs often lack geographical diversity, reinforce stereotypes, and misrepresent regions. Given their broad reach, it is critical to rigorously evaluate how these models portray the world. Existing diversity metrics either rely on curated datasets or focus on surface‑level visual similarity, limiting interpretability. We introduce GeoDiv, a framework leveraging large language and vision‑language models to assess geographical diversity along two complementary axes: the Socio‑Economic Visual Index (SEVI), capturing economic and condition‑related cues, and the Visual Diversity Index (VDI), measuring variation in primary entities and backgrounds. Applied to images generated by models such as Stable Diffusion and FLUX.1‑dev across 10 entities and 16 countries, GeoDiv reveals a consistent lack of diversity and identifies fine‑grained attributes where models default to biased portrayals. Strikingly, depictions of countries like India, Nigeria, and Colombia are disproportionately impoverished and worn, reflecting underlying socio‑economic biases. These results highlight the need for greater geographical nuance in generative models. GeoDiv provides the first systematic, interpretable framework for measuring such biases, marking a step toward fairer and more inclusive generative systems. Project page: https://abhipsabasu.github.io/geodiv
Authors:Wenhua Wu, Huai Guan, Zhe Liu, Hesheng Wang
Abstract:
Editable high‑fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end‑to‑end training and closed‑loop simulation. However, existing reconstruction methods are primarily limited to replicating observed scenes and lack the capability for diverse weather simulation. While image‑level weather editing methods tend to introduce scene artifacts and offer poor controllability over the weather effects. To address these limitations, we propose WeatherCity, a novel framework for 4D urban scene reconstruction and weather editing. Specifically, we leverage a text‑guided image editing model to achieve flexible editing of image weather backgrounds. To tackle the challenge of multi‑weather modeling, we introduce a novel weather Gaussian representation based on shared scene features and dedicated weather‑specific decoders. This representation is further enhanced with a content consistency optimization, ensuring coherent modeling across different weather conditions. Additionally, we design a physics‑driven model that simulates dynamic weather effects through particles and motion patterns. Extensive experiments on multiple datasets and various scenes demonstrate that WeatherCity achieves flexible controllability, high fidelity, and temporal consistency in 4D reconstruction and weather editing. Our framework not only enables fine‑grained control over weather conditions (e.g., light rain and heavy snow) but also supports object‑level manipulation within the scene. Codes are released at https://github.com/IRMVLab/WeatherCity.
Authors:Artur Xarles, Sergio Escalera, Thomas B. Moeslund, Albert Clapés
Abstract:
Precise Event Spotting aims to localize fast‑paced actions or events in videos with high temporal precision, a key task for applications in sports analytics, robotics, and autonomous systems. Existing methods typically process all frames uniformly, overlooking the inherent spatio‑temporal redundancy in video data. This leads to redundant computation on non‑informative regions while limiting overall efficiency. To remain tractable, they often spatially downsample inputs, losing fine‑grained details crucial for precise localization. To address these limitations, we propose AdaSpot, a simple yet effective framework that processes low‑resolution videos to extract global task‑relevant features while adaptively selecting the most informative region‑of‑interest in each frame for high‑resolution processing. The selection is performed via an unsupervised, task‑aware strategy that maintains spatio‑temporal consistency across frames and avoids the training instability of learnable alternatives. This design preserves essential fine‑grained visual cues with a marginal computational overhead compared to low‑resolution‑only baselines, while remaining far more efficient than uniform high‑resolution processing. Experiments on standard PES benchmarks demonstrate that AdaSpot achieves state‑of‑the‑art performance under strict evaluation metrics (\eg, +3.96 and +2.26 mAP@0 frames on Tennis and FineDiving), while also maintaining strong results under looser metrics. Code is available at: \hrefhttps://github.com/arturxe2/AdaSpothttps://github.com/arturxe2/AdaSpot.
Authors:Xiaoyu Xian, Shiao Wang, Xiao Wang, Daxin Tian, Yan Tian
Abstract:
Metro trains often operate in highly complex environments, characterized by illumination variations, high‑speed motion, and adverse weather conditions. These factors pose significant challenges for visual perception systems, especially those relying solely on conventional RGB cameras. To tackle these difficulties, we explore the integration of event cameras into the perception system, leveraging their advantages in low‑light conditions, high‑speed scenarios, and low power consumption. Specifically, we focus on Kilometer Marker Recognition (KMR), a critical task for autonomous metro localization under GNSS‑denied conditions. In this context, we propose a robust baseline method based on a pre‑trained RGB OCR foundation model, enhanced through multi‑modal adaptation. Furthermore, we construct the first large‑scale RGB‑Event dataset, EvMetro5K, containing 5,599 pairs of synchronized RGB‑Event samples, split into 4,479 training and 1,120 testing samples. Extensive experiments on EvMetro5K and other widely used benchmarks demonstrate the effectiveness of our approach for KMR. Both the dataset and source code will be released on https://github.com/Event‑AHU/EvMetro5K_benchmark
Authors:Yue Su, Sijin Chen, Haixin Shi, Mingyu Liu, Zhengshen Zhang, Ningyuan Huang, Weiheng Zhong, Zhengbang Zhu, Yuxiao Liu, Xihui Liu
Abstract:
Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision‑Language‑Action (VLA) models. However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine‑grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. The VLA is then trained to simultaneously predict these compressed conditions alongside future actions, thereby achieving effective world modeling within the condition space for action inference. We demonstrate that modeling and predicting this condition space not only facilitates fine‑grained action generation but also exhibits superior generalization capabilities. Moreover, it learns effectively from substantial human manipulation videos. Extensive experiments across both simulation and real‑world environments validate that our method significantly outperforms existing methods based on future prediction. Project page is available at: https://selen‑suyue.github.io/WoGNet/
Authors:Tong Wei, Giorgos Tolias, Jiri Matas, Daniel Barath
Abstract:
The pose graph is a core component of Structure‑from‑Motion (SfM), where images act as nodes and edges encode relative poses. Since geometric verification is expensive, SfM pipelines restrict the pose graph to a sparse set of candidate edges, making initialization critical. Existing methods rely on image retrieval to connect each image to its k nearest neighbors, treating pairs independently and ignoring global consistency. We address this limitation through the concept of edge prioritization, ranking candidate edges by their utility for SfM. Our approach has three components: (1) a GNN trained with SfM‑derived supervision to predict globally consistent edge reliability; (2) multi‑minimal‑spanning‑tree‑based pose graph construction guided by these ranks; and (3) connectivity‑aware score modulation that reinforces weak regions and reduces graph diameter. This globally informed initialization yields more reliable and compact pose graphs, improving reconstruction accuracy in sparse and high‑speed settings and outperforming SOTA retrieval methods on ambiguous scenes. The ode and trained models are available at https://github.com/weitong8591/global_edge_prior.
Authors:Lingjun Zhang, Yujian Yuan, Changjie Wu, Xinyuan Chang, Xin Cai, Shuang Zeng, Linzhe Shi, Sijin Wang, Hang Zhang, Mu Xu
Abstract:
Vision‑Language Models (VLM) exhibit strong reasoning capabilities, showing promise for end‑to‑end autonomous driving systems. Chain‑of‑Thought (CoT), as VLM's widely used reasoning strategy, is facing critical challenges. Existing textual CoT has a large gap between text semantic space and trajectory physical space. Although the recent approach utilizes future image to replace text as CoT process, it lacks clear planning‑oriented objective guidance to generate images with accurate scene evolution. To address these, we innovatively propose MindDriver, a progressive multimodal reasoning framework that enables VLM to imitate human‑like progressive thinking for autonomous driving. MindDriver presents semantic understanding, semantic‑to‑physical space imagination, and physical‑space trajectory planning. To achieve aligned reasoning processes in MindDriver, we develop a feedback‑guided automatic data annotation pipeline to generate aligned multimodal reasoning training data. Furthermore, we develop a progressive reinforcement fine‑tuning method to optimize the alignment through progressive high‑ level reward‑based learning. MindDriver demonstrates superior performance in both nuScences open‑loop and Bench2Drive closed‑loop evaluation. Codes are available at https://github.com/hotdogcheesewhite/MindDriver.
Authors:Huangwei Chen, Junhao Jia, Ruocheng Li, Cunyuan Yang, Wu Li, Xiaotao Pang, Yifei Chen, Haishuai Wang, Jiajun Bu, Lei Wu
Abstract:
Diabetic Retinopathy (DR) progresses as a continuous and irreversible deterioration of the retina, following a well‑defined clinical trajectory from mild to severe stages. However, most existing ordinal regression approaches model DR severity as a set of static, symmetric ranks, capturing relative order while ignoring the inherent unidirectional nature of disease progression. As a result, the learned feature representations may violate biological plausibility, allowing implausible proximity between non‑consecutive stages or even reverse transitions. To bridge this gap, we propose Directed Ordinal Diffusion Regularization (D‑ODR), which explicitly models the feature space as a directed flow by constructing a progression‑constrained directed graph that strictly enforces forward disease evolution. By performing multi‑scale diffusion on this directed structure, D‑ODR imposes penalties on score inversions along valid progression paths, thereby effectively preventing the model from learning biologically inconsistent reverse transitions. This mechanism aligns the feature representation with the natural trajectory of DR worsening. Extensive experiments demonstrate that D‑ODR yields superior grading performance compared to state‑of‑the‑art ordinal regression and DR‑specific grading methods, offering a more clinically reliable assessment of disease severity. Our code is available on https://github.com/HovChen/D‑ODR.
Authors:Cuong Anh Pham, Praneeth Vepakomma, Samuel Horváth
Abstract:
Alleviating catastrophic forgetting while enabling further learning is a primary challenge in continual learning (CL). Orthogonal‑based training methods have gained attention for their efficiency and strong theoretical properties, and many existing approaches enforce orthogonality through gradient projection. In this paper, we revisit orthogonality and exploit the fact that small singular values correspond to directions that are nearly orthogonal to the input space of previous tasks. Building on this principle, we introduce NESS (Null‑space Estimated from Small Singular values), a CL method that applies orthogonality directly in the weight space rather than through gradient manipulation. Specifically, NESS constructs an approximate null space using the smallest singular values of each layer's input representation and parameterizes task‑specific updates via a compact low‑rank adaptation (LoRA‑style) formulation constrained to this subspace. The subspace basis is fixed to preserve the null‑space constraint, and only a single trainable matrix is learned for each task. This design ensures that the resulting updates remain approximately in the null space of previous inputs while enabling adaptation to new tasks. Our theoretical analysis and experiments on three benchmark datasets demonstrate competitive performance, low forgetting, and stable accuracy across tasks, highlighting the role of small singular values in continual learning. The code is available at https://github.com/pacman‑ctm/NESS.
Authors:Yinheng Lin, Yiming Huang, Beilei Cui, Long Bai, Huxin Gao, Hongliang Ren, Jiewen Lai
Abstract:
Accurate depth estimation plays a critical role in the navigation of endoscopic surgical robots, forming the foundation for 3D reconstruction and safe instrument guidance. Fine‑tuning pretrained models heavily relies on endoscopic surgical datasets with precise depth annotations. While existing self‑supervised depth estimation techniques eliminate the need for accurate depth annotations, their performance degrades in environments with weak textures and variable lighting, leading to sparse reconstruction with invalid depth estimation. Depth completion using sparse depth maps can mitigate these issues and improve accuracy. Despite the advances in depth completion techniques in general fields, their application in endoscopy remains limited. To overcome these limitations, we propose EndoDDC, an endoscopy depth completion method that integrates images, sparse depth information with depth gradient features, and optimizes depth maps through a diffusion model, addressing the issues of weak texture and light reflection in endoscopic environments. Extensive experiments on two publicly available endoscopy datasets show that our approach outperforms state‑of‑the‑art models in both depth accuracy and robustness. This demonstrates the potential of our method to reduce visual errors in complex endoscopic environments. Our code will be released at https://github.com/yinheng‑lin/EndoDDC.
Authors:Francesco Laiti, Davide Talon, Jacopo Staiano, Elisa Ricci
Abstract:
Image memorability, i.e., how likely an image is to be remembered, has traditionally been studied in computer vision either as a passive prediction task, with models regressing a scalar score, or with generative methods altering the visual input to boost the image likelihood of being remembered. Yet, none of these paradigms supports users at capture time, when the crucial question is how to improve a photo memorability. We introduce the task of Memorability Feedback (MemFeed), where an automated model should provide actionable, human‑interpretable guidance to users with the goal to enhance an image future recall. We also present MemCoach, the first approach designed to provide concrete suggestions in natural language for memorability improvement (e.g., "emphasize facial expression," "bring the subject forward"). Our method, based on Multimodal Large Language Models (MLLMs), is training‑free and employs a teacher‑student steering strategy, aligning the model internal activations toward more memorable patterns learned from a teacher model progressing along least‑to‑most memorable samples. To enable systematic evaluation on this novel task, we further introduce MemBench, a new benchmark featuring sequence‑aligned photoshoots with annotated memorability scores. Our experiments, considering multiple MLLMs, demonstrate the effectiveness of MemCoach, showing consistently improved performance over several zero‑shot models. The results indicate that memorability can not only be predicted but also taught and instructed, shifting the focus from mere prediction to actionable feedback for human creators.
Authors:Xiankang He, Peile Lin, Ying Cui, Dongyan Guo, Chunhua Shen, Xiaoqin Zhang
Abstract:
Motion segmentation in dynamic scenes is highly challenging, as conventional methods heavily rely on estimating camera poses and point correspondences from inherently noisy motion cues. Existing statistical inference or iterative optimization techniques that struggle to mitigate the cumulative errors in multi‑stage pipelines often lead to limited performance or high computational cost. In contrast, we propose a fully learning‑based approach that directly infers moving objects from latent feature representations via attention mechanisms, thus enabling end‑to‑end feed‑forward motion segmentation. Our key insight is to bypass explicit correspondence estimation and instead let the model learn to implicitly disentangle object and camera motion. Supported by recent advances in 4D scene geometry reconstruction (e.g., π^3), the proposed method leverages reliable camera poses and rich spatial‑temporal priors, which ensure stable training and robust inference for the model. Extensive experiments demonstrate that by eliminating complex pre‑processing and iterative refinement, our approach achieves state‑of‑the‑art motion segmentation performance with high efficiency. The code is available at:https://github.com/zjutcvg/GeoMotion.
Authors:Zunhai Su, Weihao Ye, Hansen Feng, Keyu Fan, Jing Zhang, Dahai Yu, Zhengwu Liu, Ngai Wong
Abstract:
Learning‑based 3D visual geometry models have significantly advanced with the advent of large‑scale transformers. Among these, StreamVGGT leverages frame‑wise causal attention to deliver robust and efficient streaming 3D reconstruction. However, it suffers from unbounded growth in the Key‑Value (KV) cache due to the massive influx of vision tokens from multi‑image and long‑video inputs, leading to increased memory consumption and inference latency as input frames accumulate. This ultimately limits its scalability for long‑horizon applications. To address this gap, we propose XStreamVGGT, a tuning‑free approach that seamlessly integrates pruning and quantization to systematically compress the KV cache, enabling extremely memory‑efficient streaming inference. Specifically, redundant KVs generated from multi‑frame inputs are initially pruned to conform to a fixed KV memory budget using an efficient token‑importance identification mechanism that maintains full compatibility with high‑performance attention kernels (e.g., FlashAttention). Additionally, leveraging the inherent distribution patterns of KV tensors, we apply dimension‑adaptive KV quantization within the pruning pipeline to further minimize memory overhead while preserving numerical accuracy. Extensive evaluations show that XStreamVGGT achieves mostly negligible performance degradation while substantially reducing memory usage by 4.42× and accelerating inference by 5.48×, enabling practical and scalable streaming 3D applications. The code is available at https://github.com/ywh187/XStreamVGGT/.
Authors:Liangbing Zhao, Le Zhuo, Sayak Paul, Hongsheng Li, Mohamed Elhoseiny
Abstract:
Instruction‑based image editing has achieved remarkable success in semantic alignment, yet state‑of‑the‑art models frequently fail to render physically plausible results when editing involves complex causal dynamics, such as refraction or material deformation. We attribute this limitation to the dominant paradigm that treats editing as a discrete mapping between image pairs, which provides only boundary conditions and leaves transition dynamics underspecified. To address this, we reformulate physics‑aware editing as predictive physical state transitions and introduce PhysicTran38K, a large‑scale video‑based dataset comprising 38K transition trajectories across five physical domains, constructed via a two‑stage filtering and constraint‑aware annotation pipeline. Building on this supervision, we propose PhysicEdit, an end‑to‑end framework equipped with a textual‑visual dual‑thinking mechanism. It combines a frozen Qwen2.5‑VL for physically grounded reasoning with learnable transition queries that provide timestep‑adaptive visual guidance to a diffusion backbone. Experiments show that PhysicEdit improves over Qwen‑Image‑Edit by 5.9% in physical realism and 10.1% in knowledge‑grounded editing, setting a new state‑of‑the‑art for open‑source methods, while remaining competitive with leading proprietary models.
Authors:Euisoo Jung, Byunghyun Kim, Hyunjin Kim, Seonghye Cho, Jae-Gil Lee
Abstract:
Diffusion models have achieved remarkable progress in high‑fidelity image, video, and audio generation, yet inference remains computationally expensive. Nevertheless, current diffusion acceleration methods based on distributed parallelism suffer from noticeable generation artifacts and fail to achieve substantial acceleration proportional to the number of GPUs. Therefore, we propose a hybrid parallelism framework that combines a novel data parallel strategy, condition‑based partitioning, with an optimal pipeline scheduling method, adaptive parallelism switching, to reduce generation latency and achieve high generation quality in conditional diffusion models. The key ideas are to (i) leverage the conditional and unconditional denoising paths as a new data‑partitioning perspective and (ii) adaptively enable optimal pipeline parallelism according to the denoising discrepancy between these two paths. Our framework achieves 2.31× and 2.07× latency reductions on SDXL and SD3, respectively, using two NVIDIA RTX~3090 GPUs, while preserving image quality. This result confirms the generality of our approach across U‑Net‑based diffusion models and DiT‑based flow‑matching architectures. Our approach also outperforms existing methods in acceleration under high‑resolution synthesis settings. Code is available at https://github.com/kaist‑dmlab/Hybridiff.
Authors:Juan Yang, Yuyan Zhang, Han Jia, Bing Hu, Wanzhong Song
Abstract:
Monocular depth estimation (MDE) for colonoscopy is hampered by the domain gap between simulated and real‑world images. Existing image‑to‑image translation methods, which use depth as a posterior constraint, often produce structural distortions and specular highlights by failing to balance realism with structure consistency. To address this, we propose a Structure‑to‑Image paradigm that transforms the depth map from a passive constraint into an active generative foundation. We are the first to introduce phase congruency to colonoscopic domain adaptation and design a cross‑level structure constraint to co‑optimize geometric structures and fine‑grained details like vascular textures. In zero‑shot evaluations conducted on a publicly available phantom dataset, the MDE model that was fine‑tuned on our generated data achieved a maximum reduction of 44.18% in RMSE compared to competing methods. Our code is available at https://github.com/YyangJJuan/PC‑S2I.git.
Authors:Guanyi Qin, Xiaozhen Wang, Zhu Zhuo, Chang Han Low, Yuancan Xiao, Yibing Fu, Haofeng Liu, Kai Wang, Chunjiang Li, Yueming Jin
Abstract:
Minimally invasive surgery has dramatically improved patient operative outcomes, yet identifying safe operative zones remains challenging in critical phases, requiring surgeons to integrate visual cues, procedural phase, and anatomical context under high cognitive load. Existing AI systems offer binary safety verification or static detection, ignoring the phase‑dependent nature of intraoperative reasoning. We introduce ResGo, a benchmark of laparoscopic frames annotated with Go Zone bounding boxes and clinician‑authored rationales covering phase, exposure quality reasoning, next action and risk reminder. We introduce evaluation metrics that treat correct grounding under incorrect phase as failures, revealing that most vision‑language models cannot handle such tasks and perform poorly. We then present SurGo‑R1, a model optimized via RLHF with a multi‑turn phase‑then‑go architecture where the model first identifies the surgical phase, then generates reasoning and Go Zone coordinates conditioned on that context. On unseen procedures, SurGo‑R1 achieves 76.6% phase accuracy, 32.7 mIoU, and 54.8% hardcore accuracy, a 6.6× improvement over the mainstream generalist VLMs. Code, model and benchmark will be available at https://github.com/jinlab‑imvr/SurGo‑R1
Authors:Meiqi Sun, Mingyu Li, Junxiong Zhu
Abstract:
Generative AI is widely used to create commercial posters. However, rapid advances in generation have outpaced automated quality assessment. Existing models emphasize generic esthetics or low level distortions and lack the functional criteria required for e‑commerce design. It is especially challenging for Chinese content, where complex characters often produce subtle but critical textual artifacts that are overlooked by existing methods. To address this, we introduce E‑comIQ‑ZH, a framework for evaluating Chinese e‑commerce posters. We build the first dataset E‑comIQ‑18k to feature multi dimensional scores and expert calibrated Chain of Thought (CoT) rationales. Using this dataset, we train E‑comIQ‑M, a specialized evaluation model that aligns with human expert judgment. Our framework enables E‑comIQ‑Bench, the first automated and scalable benchmark for the generation of Chinese e‑commerce posters. Extensive experiments show our E‑comIQ‑M aligns more closely with expert standards and enables scalable automated assessment of e‑commerce posters. All datasets, models, and evaluation tools will be released to support future research in this area.Code will be available at https://github.com/4mm7/E‑comIQ‑ZH.
Authors:Junmyeong Lee, Hoseung Choi, Minsu Cho
Abstract:
Forecasting dynamic scenes remains a fundamental challenge in computer vision, as limited observations make it difficult to capture coherent object‑level motion and long‑term temporal evolution. We present Motion Group‑aware Gaussian Forecasting (MoGaF), a framework for long‑term scene extrapolation built upon the 4D Gaussian Splatting representation. MoGaF introduces motion‑aware Gaussian grouping and group‑wise optimization to enforce physically consistent motion across both rigid and non‑rigid regions, yielding spatially coherent dynamic representations. Leveraging this structured space‑time representation, a lightweight forecasting module predicts future motion, enabling realistic and temporally stable scene evolution. Experiments on synthetic and real‑world datasets demonstrate that MoGaF consistently outperforms existing baselines in rendering quality, motion plausibility, and long‑term forecasting stability. Our project page is available at https://slime0519.github.io/mogaf
Authors:Shaoxuan Wu, Jingkun Chen, Chong Ma, Cong Shen, Xiao Zhang, Jun Feng
Abstract:
Computer‑aided diagnosis (CAD) has significantly advanced automated chest X‑ray diagnosis but remains isolated from clinical workflows and lacks reliable decision support and interpretability. Human‑AI collaboration seeks to enhance the reliability of diagnostic models by integrating the behaviors of controllable radiologists. However, the absence of interactive tools seamlessly embedded within diagnostic routines impedes collaboration, while the semantic gap between radiologists' decision‑making patterns and model representations further limits clinical adoption. To overcome these limitations, we propose a visual cognition‑guided collaborative network (VCC‑Net) to achieve the cooperative diagnostic paradigm. VCC‑Net centers on visual cognition (VC) and employs clinically compatible interfaces, such as eye‑tracking or the mouse, to capture radiologists' visual search traces and attention patterns during diagnosis. VCC‑Net employs VC as a spatial cognition guide, learning hierarchical visual search strategies to localize diagnostically key regions. A cognition‑graph co‑editing module subsequently integrates radiologist VC with model inference to construct a disease‑aware graph. The module captures dependencies among anatomical regions and aligns model representations with VC‑driven features, mitigating radiologist bias and facilitating complementary, transparent decision‑making. Experiments on the public datasets SIIM‑ACR, EGD‑CXR, and self‑constructed TB‑Mouse dataset achieved classification accuracies of 88.40%, 85.05%, and 92.41%, respectively. The attention maps produced by VCC‑Net exhibit strong concordance with radiologists' gaze distributions, demonstrating a mutual reinforcement of radiologist and model inference. The code is available at https://github.com/IPMI‑NWU/VCC‑Net.
Authors:Chenyv Liu, Wentao Tan, Lei Zhu, Fengling Li, Jingjing Li, Guoli Yang, Heng Tao Shen
Abstract:
Standard vision‑language‑action (VLA) models rely on fitting statistical data priors, limiting their robust understanding of underlying physical dynamics. Reinforcement learning enhances physical grounding through exploration yet typically relies on external reward signals that remain isolated from the agent's internal states. World action models have emerged as a promising paradigm that integrates imagination and control to enable predictive planning. However, they rely on implicit context modeling, lacking explicit mechanisms for self‑improvement. To solve these problems, we propose Self‑Correcting VLA (SC‑VLA), which achieve self‑improvement by intrinsically guiding action refinement through sparse imagination. We first design sparse world imagination by integrating auxiliary predictive heads to forecast current task progress and future trajectory trends, thereby constraining the policy to encode short‑term physical evolution. Then we introduce the online action refinement module to reshape progress‑dependent dense rewards, adjusting trajectory orientation based on the predicted sparse future states. Evaluations on challenging robot manipulation tasks from simulation benchmarks and real‑world settings demonstrate that SC‑VLA achieve state‑of‑the‑art performance, yielding the highest task throughput with 16% fewer steps and a 9% higher success rate than the best‑performing baselines, alongside a 14% gain in real‑world experiments. Code is available at https://github.com/Kisaragi0/SC‑VLA.
Authors:Abhineet Singh, Justin Rozeboom, Nilanjan Ray
Abstract:
This paper presents a new unified approach to semantic segmentation in both images and videos by using language modeling to output the masks as sequences of discrete tokens. We use run length encoding (RLE) to discretize the segmentation masks, and adapt the Pix2Seq framework to learn autoregressive models to output these tokens. We propose novel tokenization strategies to compress the lengths of the token sequences to make it practicable to extend this approach to videos. We also show how instance information can be incorporated into the tokenization process to perform panoptic segmentation. We evaluate our models on two domain‑specific datasets to demonstrate their competitiveness with the state of the art in certain scenarios, in spite of being severely bottlenecked by our limited computational resources. We supplement these analyses by proposing several promising approaches to foster future competitiveness in general‑purpose applications, and facilitate this by making our code and models publicly available.
Authors:Yingcheng Hu, Haowen Gong, Chuanguang Yang, Zhulin An, Yongjun Xu, Songhua Liu
Abstract:
Pose‑guided human image animation aims to synthesize realistic videos of a reference character driven by a sequence of poses. While diffusion‑based methods have achieved remarkable success, most existing approaches are limited to single‑character animation. We observe that naively extending these methods to multi‑character scenarios often leads to identity confusion and implausible occlusions between characters. To address these challenges, in this paper, we propose an extensible multi‑character image animation framework built upon modern Diffusion Transformers (DiTs) for video generation. At its core, our framework introduces two novel components‑Identifier Assigner and Identifier Adapter ‑ which collaboratively capture per‑person positional cues and inter‑person spatial relationships. This mask‑driven scheme, along with a scalable training strategy, not only enhances flexibility but also enables generalization to scenarios with more characters than those seen during training. Remarkably, trained on only a two‑character dataset, our model generalizes to multi‑character animation while maintaining compatibility with single‑character cases. Extensive experiments demonstrate that our approach achieves state‑of‑the‑art performance in multi‑character image animation, surpassing existing diffusion‑based baselines.
Authors:Changqing Zhou, Yueru Luo, Changhao Chen
Abstract:
Accurate 3D scene understanding is essential for embodied intelligence, with occupancy prediction emerging as a key task for reasoning about both objects and free space. Existing approaches largely rely on depth priors (e.g., DepthAnything) but make only limited use of 3D cues, restricting performance and generalization. Recently, visual geometry models such as VGGT have shown strong capability in providing rich 3D priors, but similar to monocular depth foundation models, they still operate at the level of visible surfaces rather than volumetric interiors, motivating us to explore how to more effectively leverage these increasingly powerful geometry priors for 3D occupancy prediction. We present GPOcc, a framework that leverages generalizable visual geometry priors (GPs) for monocular occupancy prediction. Our method extends surface points inward along camera rays to generate volumetric samples, which are represented as Gaussian primitives for probabilistic occupancy inference. To handle streaming input, we further design a training‑free incremental update strategy that fuses per‑frame Gaussians into a unified global representation. Experiments on Occ‑ScanNet and EmbodiedOcc‑ScanNet demonstrate significant gains: GPOcc improves mIoU by +9.99 in the monocular setting and +11.79 in the streaming setting over prior state of the art. Under the same depth prior, it achieves +6.73 mIoU while running 2.65× faster. These results highlight that GPOcc leverages geometry priors more effectively and efficiently. Code will be released at https://github.com/JuIvyy/GPOcc.
Authors:Chaojie Shen, Jingjun Gu, Zihao Zhao, Ruocheng Li, Cunyuan Yang, Jiajun Bu, Lei Wu
Abstract:
Accurate Couinaud liver segmentation is critical for preoperative surgical planning and tumor localization.However, existing methods primarily rely on image intensity and spatial location cues, without explicitly modeling vascular topology. As a result, they often produce indistinct boundaries near vessels and show limited generalization under anatomical variability.We propose VasGuideNet, the first Couinaud segmentation framework explicitly guided by vascular topology. Specifically, skeletonized vessels, Euclidean distance transform (EDT)‑‑derived geometry, and k‑nearest neighbor (kNN) connectivity are encoded into topology features using Graph Convolutional Networks (GCNs). These features are then injected into a 3D encoder‑‑decoder backbone via a cross‑attention fusion module. To further improve inter‑class separability and anatomical consistency, we introduce a Structural Contrastive Loss (SCL) with a global memory bank.On Task08_HepaticVessel and our private LASSD dataset, VasGuideNet achieves Dice scores of 83.68% and 76.65% with RVDs of 1.68 and 7.08, respectively. It consistently outperforms representative baselines including UNETR, Swin UNETR, and G‑UNETR++, delivering higher Dice/mIoU and lower RVD across datasets, demonstrating its effectiveness for anatomically consistent segmentation. Code is available at https://github.com/Qacket/VasGuideNet.git.
Authors:Pengli Zhu, Yitao Zhu, Haowen Pang, Anqi Qiu
Abstract:
Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challenges, we introduce IHF‑Harmony, a unified invertible hierarchy flow framework for multi‑modality harmonization using unpaired data. By decomposing the translation process into reversible feature transformations, IHF‑Harmony guarantees bijective mapping and lossless reconstruction to prevent anatomical distortion. Specifically, an invertible hierarchy flow (IHF) performs hierarchical subtractive coupling to progressively remove artefact‑related features, while an artefact‑aware normalization (AAN) employs anatomy‑fixed feature modulation to accurately transfer target characteristics. Combined with anatomy and artefact consistency loss objectives, IHF‑Harmony achieves high‑fidelity harmonization that retains source anatomy. Experiments across multiple MRI modalities demonstrate that IHF‑Harmony outperforms existing methods in both anatomical fidelity and downstream task performance, facilitating robust harmonization for large‑scale multi‑site imaging studies. Code is available at https://github.com/Idea89560041/IHF‑Harmony.
Authors:Yue Yang, Shuo Cheng, Yu Fang, Homanga Bharadhwaj, Mingyu Ding, Gedas Bertasius, Daniel Szafir
Abstract:
General‑purpose robots must master long‑horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision‑Language‑Action (VLA) models offer the potential to master diverse atomic skills, they struggle with the combinatorial complexity of sequencing them and are prone to cascading failures due to environmental sensitivity. To address these challenges, we propose LiLo‑VLA (Linked Local VLA), a modular framework capable of zero‑shot generalization to novel long‑horizon tasks without ever being trained on them. Our approach decouples transport from interaction: a Reaching Module handles global motion, while an Interaction Module employs an object‑centric VLA to process isolated objects of interest, ensuring robustness against irrelevant visual features and invariance to spatial configurations. Crucially, this modularity facilitates robust failure recovery through dynamic replanning and skill reuse, effectively mitigating the cascading errors common in end‑to‑end approaches. We introduce a 21‑task simulation benchmark consisting of two challenging suites: LIBERO‑Long++ and Ultra‑Long. In these simulations, LiLo‑VLA achieves a 69% average success rate, outperforming Pi0.5 by 41% and OpenVLA‑OFT by 67%. Furthermore, real‑world evaluations across 8 long‑horizon tasks demonstrate an average success rate of 85%. Project page: https://yy‑gx.github.io/LiLo‑VLA/.
Authors:Shimin Hu, Yuanyi Wei, Fei Zha, Yudong Guo, Juyong Zhang
Abstract:
Existing 3D editing methods rely on computationally intensive scene‑by‑scene iterative optimization and suffer from multi‑view inconsistency. We propose an effective and feed‑forward 3D editing framework based on the TRELLIS generative backbone, capable of modifying 3D models from a single editing view. Our framework addresses two key issues: adapting training‑free 2D editing to structured 3D representations, and overcoming the bottleneck of appearance fidelity in compressed 3D features. To ensure geometric consistency, we introduce Voxel FlowEdit, an edit‑driven flow in the sparse voxel latent space that achieves globally consistent 3D deformation in a single pass. To restore high‑fidelity details, we develop a normal‑guided single to multi‑view generation module as an external appearance prior, successfully recovering high‑frequency textures. Experiments demonstrate that our method enables fast, globally consistent, and high‑fidelity 3D model editing.
Authors:Yushen He, Lei Zhao, Weidong Chen
Abstract:
3D object detection is essential for autonomous driving and robotic perception, yet its reliance on large‑scale manually annotated data limits scalability and adaptability. To reduce annotation dependency, unsupervised and sparsely‑supervised paradigms have emerged. However, they face intertwined challenges: low‑quality pseudo‑labels, unstable feature mining, and a lack of a unified training framework. This paper proposes SPL, a unified training framework for both unsupervised and sparsely‑supervised 3D object detection via \underlineSemantic \underlinePseudo‑labeling and prototype \underlineLearning. SPL first generates high‑quality pseudo‑labels by integrating image semantics, point cloud geometry, and temporal cues, producing both 3D bounding boxes for dense objects and 3D point labels for sparse ones. These pseudo‑labels are not used directly but as probabilistic priors within a novel, multi‑stage prototype learning strategy. This strategy stabilizes feature representation learning through memory‑based initialization and momentum‑based prototype updating, effectively mining features from both labeled and unlabeled data. Extensive experiments on KITTI and nuScenes datasets demonstrate that SPL significantly outperforms state‑of‑the‑art methods in both settings. Our work provides a robust and generalizable solution for learning 3D object detectors with minimal or no manual annotations. Our code is available at https://github.com/TossherO/SPL.
Authors:Wei Zhou, Yixiao Li, Hadi Amirpour, Xiaoshuai Hao, Jiang Liu, Peng Wang, Hantao Liu
Abstract:
Single‑image super‑resolution (SR) has achieved remarkable progress with deep learning, yet most approaches rely on distortion‑oriented losses or heuristic perceptual priors, which often lead to a trade‑off between fidelity and visual quality. To address this issue, we propose an Efficient Perceptual Bi‑directional Attention Network (Efficient‑PBAN) that explicitly optimizes SR towards human‑preferred quality. Unlike patch‑based quality models, Efficient‑PBAN avoids extensive patch sampling and enables efficient image‑level perception. The proposed framework is trained on our self‑constructed SR quality dataset that covers a wide range of state‑of‑the‑art SR methods with corresponding human opinion scores. Using this dataset, Efficient‑PBAN learns to predict perceptual quality in a way that correlates strongly with subjective judgments. The learned metric is further integrated into SR training as a differentiable perceptual loss, enabling closed‑loop alignment between reconstruction and perceptual assessment. Extensive experiments demonstrate that our approach delivers superior perceptual quality. Code is publicly available at https://github.com/Lighting‑YXLI/Efficient‑PBAN.
Authors:Jan Pauls, Karsten Schrödter, Sven Ligensa, Martin Schwartz, Berkant Turan, Max Zimmer, Sassan Saatchi, Sebastian Pokutta, Philippe Ciais, Fabian Gieseke
Abstract:
Forest monitoring is critical for climate change mitigation. However, existing global tree height maps provide only static snapshots and do not capture temporal forest dynamics, which are essential for accurate carbon accounting. We introduce ECHOSAT, a global and temporally consistent tree height map at 10 m resolution spanning multiple years. To this end, we resort to multi‑sensor satellite data to train a specialized vision transformer model, which performs pixel‑level temporal regression. A self‑supervised growth loss regularizes the predictions to follow growth curves that are in line with natural tree development, including gradual height increases over time, but also abrupt declines due to forest loss events such as fires. Our experimental evaluation shows that our model improves state‑of‑the‑art accuracies in the context of single‑year predictions. We also provide the first global‑scale height map that accurately quantifies tree growth and disturbances over time. We expect ECHOSAT to advance global efforts in carbon monitoring and disturbance assessment. The maps can be accessed at https://github.com/ai4forest/echosat.
Authors:Alina Devkota, Jacob Thrasher, Donald Adjeroh, Binod Bhattarai, Prashnna K. Gyawali
Abstract:
Federated Learning (FL) enables collaborative model training across multiple clients without sharing their private data. However, data heterogeneity across clients leads to client drift, which degrades the overall generalization performance of the model. This effect is further compounded by overemphasis on poorly performing clients. To address this problem, we propose FedVG, a novel gradient‑based federated aggregation framework that leverages a global validation set to guide the optimization process. Such a global validation set can be established using readily available public datasets, ensuring accessibility and consistency across clients without compromising privacy. In contrast to conventional approaches that prioritize client dataset volume, FedVG assesses the generalization ability of client models by measuring the magnitude of validation gradients across layers. Specifically, we compute layerwise gradient norms to derive a client‑specific score that reflects how much each client needs to adjust for improved generalization on the global validation set, thereby enabling more informed and adaptive federated aggregation. Extensive experiments on both natural and medical image benchmarking datasets, across diverse model architectures, demonstrate that FedVG consistently improves performance, particularly in highly heterogeneous settings. Moreover, FedVG is modular and can be seamlessly integrated with various state‑of‑the‑art FL algorithms, often further improving their results. Our code is available at https://github.com/alinadevkota/FedVG.
Authors:Yongxin Guo, Hao Lu, Onur C. Koyun, Zhengjie Zhu, Muhammet Fatih Demir, Metin Nafi Gurcan
Abstract:
Multimodal learning that integrates genomics and histopathology has shown strong potential in cancer diagnosis, yet its clinical translation is hindered by the limited availability of paired histology‑genomics data. Knowledge distillation (KD) offers a practical solution by transferring genomic supervision into histopathology models, enabling accurate inference using histology alone. However, existing KD methods rely on batch‑local alignment, which introduces instability due to limited within‑batch comparisons and ultimately degrades performance.
To address these limitations, we propose Momentum Memory Knowledge Distillation (MoMKD), a cross‑modal distillation framework driven by a momentum‑updated memory. This memory aggregates genomic and histopathology information across batches, effectively enlarging the supervisory context available to each mini‑batch. Furthermore, we decouple the gradients of the genomics and histology branches, preventing genomic signals from dominating histology feature learning during training and eliminating the modality‑gap issue at inference time.
Extensive experiments on the TCGA‑BRCA benchmark (HER2, PR, and ODX classification tasks) and an independent in‑house testing dataset demonstrate that MoMKD consistently outperforms state‑of‑the‑art MIL and multimodal KD baselines, delivering strong performance and generalization under histology‑only inference. Overall, MoMKD establishes a robust and generalizable knowledge distillation paradigm for computational pathology.
Authors:Abdulaziz Almuzairee, Henrik I. Christensen
Abstract:
Visual reinforcement learning is appealing for robotics but expensive ‑‑ off‑policy methods are sample‑efficient yet slow; on‑policy methods parallelize well but waste samples. Recent work has shown that off‑policy methods can train faster than on‑policy methods in wall‑clock time for state‑based control. Extending this to vision remains challenging, where high‑dimensional input images complicate training dynamics and introduce substantial storage and encoding overhead. To address these challenges, we introduce Squint, a visual Soft Actor Critic method that achieves faster wall‑clock training than prior visual off‑policy and on‑policy methods. Squint achieves this via parallel simulation, a distributional critic, resolution squinting, layer normalization, a tuned update‑to‑data ratio, and an optimized implementation. We evaluate on the SO‑101 Task Set, a new suite of eight manipulation tasks in ManiSkill3 with heavy domain randomization, and demonstrate sim‑to‑real transfer to a real SO‑101 robot. We train policies for 15 minutes on a single RTX 3090 GPU, with most tasks converging in under 6 minutes.
Authors:Haoyi Jiang, Liu Liu, Xinjie Wang, Yonghao He, Wei Sui, Zhizhong Su, Wenyu Liu, Xinggang Wang
Abstract:
While Vision‑Language Models (VLMs) exhibit exceptional 2D visual understanding, their ability to comprehend and reason about 3D space‑‑a cornerstone of spatial intelligence‑‑remains superficial. Current methodologies attempt to bridge this domain gap either by relying on explicit 3D modalities or by augmenting VLMs with partial, view‑conditioned geometric priors. However, such approaches hinder scalability and ultimately burden the language model with the ill‑posed task of implicitly reconstructing holistic 3D geometry from sparse cues. In this paper, we argue that spatial intelligence can emerge inherently from 2D vision alone, rather than being imposed via explicit spatial instruction tuning. To this end, we introduce Spa3R, a self‑supervised framework that learns a unified, view‑invariant spatial representation directly from unposed multi‑view images. Spa3R is built upon the proposed Predictive Spatial Field Modeling (PSFM) paradigm, where Spa3R learns to synthesize feature fields for arbitrary unseen views conditioned on a compact latent representation, thereby internalizing a holistic and coherent understanding of the underlying 3D scene. We further integrate the pre‑trained Spa3R Encoder into existing VLMs via a lightweight adapter to form Spa3‑VLM, effectively grounding language reasoning in a global spatial context. Experiments on the challenging VSI‑Bench demonstrate that Spa3‑VLM achieves state‑of‑the‑art accuracy of 58.6% on 3D VQA, significantly outperforming prior methods. These results highlight PSFM as a scalable path toward advancing spatial intelligence. Code is available at https://github.com/hustvl/Spa3R.
Authors:Sepehr Salem Ghahfarokhi, M. Moein Esfahani, Raj Sunderraman, Vince Calhoun, Mohammed Alser
Abstract:
Deep learning has significantly advanced automated brain tumor diagnosis, yet clinical adoption remains limited by interpretability and computational constraints. Conventional models often act as opaque ''black boxes'' and fail to quantify the complex, irregular tumor boundaries that characterize malignant growth. To address these challenges, we present XMorph, an explainable and computationally efficient framework for fine‑grained classification of three prominent brain tumor types: glioma, meningioma, and pituitary tumors. We propose an Information‑Weighted Boundary Normalization (IWBN) mechanism that emphasizes diagnostically relevant boundary regions alongside nonlinear chaotic and clinically validated features, enabling a richer morphological representation of tumor growth. A dual‑channel explainable AI module combines GradCAM++ visual cues with LLM‑generated textual rationales, translating model reasoning into clinically interpretable insights. The proposed framework achieves a classification accuracy of 96.0%, demonstrating that explainability and high performance can co‑exist in AI‑based medical imaging systems. The source code and materials for XMorph are all publicly available at: https://github.com/ALSER‑Lab/XMorph.
Authors:Jianglin Lu, Simon Jenni, Kushal Kafle, Jing Shi, Handong Zhao, Yun Fu
Abstract:
Text‑to‑image retrieval is a fundamental task in vision‑language learning, yet in real‑world scenarios it is often challenged by short and underspecified user queries. Such queries are typically only one or two words long, rendering them semantically ambiguous, prone to collisions across diverse visual interpretations, and lacking explicit control over the quality of retrieved images. To address these issues, we propose a new paradigm of quality‑controllable retrieval, which enriches short queries with contextual details while incorporating explicit notions of image quality. Our key idea is to leverage a generative language model as a query completion function, extending underspecified queries into descriptive forms that capture fine‑grained visual attributes such as pose, scene, and aesthetics. We introduce a general framework that conditions query completion on discretized quality levels, derived from relevance and aesthetic scoring models, so that query enrichment is not only semantically meaningful but also quality‑aware. The resulting system provides three key advantages: 1) flexibility, it is compatible with any pretrained vision‑language model (VLMs) without modification; 2) transparency, enriched queries are explicitly interpretable by users; and 3) controllability, enabling retrieval results to be steered toward user‑preferred quality levels. Extensive experiments demonstrate that our proposed approach significantly improves retrieval results and provides effective quality control, bridging the gap between the expressive capacity of modern VLMs and the underspecified nature of short user queries. Our code is available at https://github.com/Jianglin954/QCQC.
Authors:Bastien Gimbert
Abstract:
We present SPRITETOMESH, a fully automatic pipeline for converting 2D game sprite images into triangle meshes compatible with skeletal animation frameworks such as Spine2D. Creating animation‑ready meshes is traditionally a tedious manual process requiring artists to carefully place vertices along visual boundaries, a task that typically takes 15‑60 minutes per sprite. Our method addresses this through a hybrid learned‑algorithmic approach. A segmentation network (EfficientNet‑B0 encoder with U‑Net decoder) trained on over 100,000 sprite‑mask pairs from 172 games achieves an IoU of 0.87, providing accurate binary masks from arbitrary input images. From these masks, we extract exterior contour vertices using Douglas‑Peucker simplification with adaptive arc subdivision, and interior vertices along visual boundaries detected via bilateral‑filtered multi‑channel Canny edge detection with contour‑following placement. Delaunay triangulation with mask‑based centroid filtering produces the final mesh. Through controlled experiments, we demonstrate that direct vertex position prediction via neural network heatmap regression is fundamentally not viable for this task: the heatmap decoder consistently fails to converge (loss plateau at 0.061) while the segmentation decoder trains normally under identical conditions. We attribute this to the inherently artistic nature of vertex placement ‑ the same sprite can be meshed validly in many different ways. This negative result validates our hybrid design: learned segmentation where ground truth is unambiguous, algorithmic placement where domain heuristics are appropriate. The complete pipeline processes a sprite in under 3 seconds, representing a speedup of 300x‑1200x over manual creation. We release our trained model to the game development community.
Authors:Joseph Raj Vishal, Nagasiri Poluri, Katha Naik, Rutuja Patil, Kashyap Hegde Kota, Krishna Vinod, Prithvi Jai Ramesh, Mohammad Farhadi, Yezhou Yang, Bharatesh Chakravarthi
Abstract:
Understanding the complex, multi‑agent dynamics of urban traffic remains a fundamental challenge for video language models. This paper introduces Urban Dynamics VideoQA, a benchmark dataset that captures the unscripted real‑world behavior of dynamic urban scenes. UDVideoQA is curated from 16 hours of traffic footage recorded at multiple city intersections under diverse traffic, weather, and lighting conditions. It employs an event‑driven dynamic blur technique to ensure privacy preservation without compromising scene fidelity. Using a unified annotation pipeline, the dataset contains 28K question‑answer pairs generated across 8 hours of densely annotated video, averaging one question per second. Its taxonomy follows a hierarchical reasoning level, spanning basic understanding and attribution to event reasoning, reverse reasoning, and counterfactual inference, enabling systematic evaluation of both visual grounding and causal reasoning. Comprehensive experiments benchmark 10 SOTA VideoLMs on UDVideoQA and 8 models on a complementary video question generation benchmark. Results reveal a persistent perception‑reasoning gap, showing models that excel in abstract inference often fail with fundamental visual grounding. While models like Gemini Pro achieve the highest zero‑shot accuracy, fine‑tuning the smaller Qwen2.5‑VL 7B model on UDVideoQA bridges this gap, achieving performance comparable to proprietary systems. In VideoQGen, Gemini 2.5 Pro, and Qwen3 Max generate the most relevant and complex questions, though all models exhibit limited linguistic diversity, underscoring the need for human‑centric evaluation. The UDVideoQA suite, including the dataset, annotation tools, and benchmarks for both VideoQA and VideoQGen, provides a foundation for advancing robust, privacy‑aware, and real‑world multimodal reasoning. UDVideoQA is available at https://ud‑videoqa.github.io/UD‑VideoQA/UD‑VideoQA/.
Authors:Noé Artru, Rukhshanda Hussain, Emeline Got, Alexandre Messier, David B. Lindell, Abdallah Dib
Abstract:
Reconstructing high‑fidelity 3D head geometry from images is critical for a wide range of applications, yet existing methods face fundamental limitations. Traditional photogrammetry achieves exceptional detail but requires extensive camera arrays (25‑200+ views), substantial computation, and manual cleanup in challenging areas like facial hair. Recent alternatives present a fundamental trade‑off: foundation models enable efficient single‑image reconstruction but lack fine geometric detail, while optimization‑based methods achieve higher fidelity but require dense views and expensive computation. We bridge this gap with a hybrid approach that combines the strengths of both paradigms. Our method introduces a multi‑view surface normal prediction model that extends monocular foundation models with cross‑view attention to produce geometrically consistent normals in a feed‑forward pass. We then leverage these predictions as strong geometric priors within an inverse rendering optimization framework to recover high‑frequency surface details. Our approach outperforms state‑of‑the‑art single‑image and multi‑view methods, achieving high‑fidelity reconstruction on par with dense‑view photogrammetry while reducing camera requirements and computational cost.
Authors:Duowen Chen, Yan Wang
Abstract:
Federated Semi‑Supervised Learning (FSSL) aims to collaboratively train a global model across clients by leveraging partially‑annotated local data in a privacy‑preserving manner. In FSSL, data heterogeneity is a challenging issue, which exists both across clients and within clients. External heterogeneity refers to the data distribution discrepancy across different clients, while internal heterogeneity represents the mismatch between labeled and unlabeled data within clients. Most FSSL methods typically design fixed or dynamic parameter aggregation strategies to collect client knowledge on the server (external) and / or filter out low‑confidence unlabeled samples to reduce mistakes in local client (internal). But, the former is hard to precisely fit the ideal global distribution via direct weights, and the latter results in fewer data participation into FL training. To this end, we propose a proxy‑guided framework called ProxyFL that focuses on simultaneously mitigating external and internal heterogeneity via a unified proxy. I.e., we consider the learnable weights of classifier as proxy to simulate the category distribution both locally and globally. For external, we explicitly optimize global proxy against outliers instead of direct weights; for internal, we re‑include the discarded samples into training by a positive‑negative proxy pool to mitigate the impact of potentially‑incorrect pseudo‑labels. Insight experiments & theoretical analysis show our significant performance and convergence in FSSL.
Authors:Shimin Wen, Zeyu Zhang, Xingdou Bian, Hongjie Zhu, Lulu He, Layi Shama, Daji Ergu, Ying Cai
Abstract:
Large Vision‑Language Models (VLMs) have demonstrated significant potential on complex visual understanding tasks through iterative optimization methods.However, these models generally lack effective self‑correction mechanisms, making it difficult for them to independently rectify cognitive biases. Consequently, during multi‑turn revisions, they often fall into repetitive and ineffective attempts, failing to achieve stable improvements in answer quality.To address this issue, we propose a novel iterative self‑correction framework that endows models with two key capabilities: Capability Reflection and Memory Reflection. This framework guides the model to first diagnose errors and generate a correction plan via Capability Reflection, then leverage Memory Reflection to review past attempts to avoid repetition and explore new solutions, and finally, optimize the answer through rigorous re‑reasoning. Experiments on the challenging OCRBench v2 benchmark show that OCR‑Agent outperforms the current open‑source SOTA model InternVL3‑8B by +2.0 on English and +1.2 on Chinese subsets, while achieving state‑of‑the‑art results in Visual Understanding (79.9) and Reasoning (66.5) ‑ surpassing even larger fine‑tuned models. Our method demonstrates that structured, self‑aware reflection can significantly enhance VLMs' reasoning robustness without additional training. Code: https://github.com/AIGeeksGroup/OCR‑Agent.
Authors:Bonan Liu, Zeyu Zhang, Bingbing Meng, Han Wang, Hanshuo Zhang, Chengping Wang, Daji Ergu, Ying Cai
Abstract:
Optical character recognition (OCR) has advanced rapidly with deep learning and multimodal models, yet most methods focus on well‑resourced scripts such as Latin and Chinese. Ethnic minority languages remain underexplored due to complex writing systems, scarce annotations, and diverse historical and modern forms, making generalization in low‑resource or zero‑shot settings challenging. To address these challenges, we present OmniOCR, a universal framework for ethnic minority scripts. OmniOCR introduces Dynamic Low‑Rank Adaptation (Dynamic LoRA) to allocate model capacity across layers and scripts, enabling effective adaptation while preserving knowledge.A sparsity regularization prunes redundant updates, ensuring compact and efficient adaptation without extra inference cost. Evaluations on TibetanMNIST, Shui, ancient Yi, and Dongba show that OmniOCR outperforms zero‑shot foundation models and standard post training, achieving state‑of‑the‑art accuracy with superior parameter efficiency, and compared with the state‑of‑the‑art baseline models, it improves accuracy by 39%‑66% on these four datasets. Code: https://github.com/AIGeeksGroup/OmniOCR.
Authors:Tianhao Fu, Yucheng Chen
Abstract:
Medical image processing demands specialized software that handles high‑dimensional volumetric data, heterogeneous file formats, and domain‑specific training procedures. Existing frameworks either provide low‑level components that require substantial integration effort or impose rigid, monolithic pipelines that resist modification. We present MIP Candy (MIPCandy), a freely available, PyTorch‑based framework designed specifically for medical image processing. MIPCandy provides a complete, modular pipeline spanning data loading, training, inference, and evaluation, allowing researchers to obtain a fully functional process workflow by implementing a single method, \textttbuild_network, while retaining fine‑grained control over every component. Central to the design is \textttLayerT, a deferred configuration mechanism that enables runtime substitution of convolution, normalization, and activation modules without subclassing. The framework further offers built‑in k‑fold cross‑validation, dataset inspection with automatic region‑of‑interest detection, deep supervision, exponential moving average, multi‑frontend experiment tracking (Weights & Biases, Notion, MLflow), training state recovery, and validation score prediction via quotient regression. An extensible bundle ecosystem provides pre‑built model implementations that follow a consistent trainer‑‑predictor pattern and integrate with the core framework without modification. MIPCandy is open‑source under the Apache‑2.0 license and requires Python~3.12 or later. Source code and documentation are available at https://github.com/ProjectNeura/MIPCandy.
Authors:Yuhao Wu, Maojia Song, Yihuai Lan, Lei Wang, Zhiqiang Hu, Yao Xiao, Heng Zhou, Weihua Zheng, Dylan Raharja, Soujanya Poria, Roy Ka-Wei Lee
Abstract:
Understanding the physical structure is essential for real‑world applications such as embodied agents, interactive design, and long‑horizon manipulation. Yet, prevailing Vision‑Language Model (VLM) evaluations still center on structure‑agnostic, single‑turn setups (e.g., VQA), which fail to assess agents' ability to reason about how geometry, contact, and support relations jointly constrain what actions are possible in a dynamic environment. To address this gap, we introduce the Causal Hierarchy of Actions and Interactions (CHAIN) benchmark, an interactive 3D, physics‑driven testbed designed to evaluate whether models can understand, plan, and execute structured action sequences grounded in physical constraints. CHAIN shifts evaluation from passive perception to active problem solving, spanning tasks such as interlocking mechanical puzzles and 3D stacking and packing. We conduct a comprehensive study of state‑of‑the‑art VLMs and diffusion‑based models under unified interactive settings. Our results show that top‑performing models still struggle to internalize physical structure and causal constraints, often failing to produce reliable long‑horizon plans and cannot robustly translate perceived structure into effective actions. The project is available at https://social‑ai‑studio.github.io/CHAIN/.
Authors:Bowen Zheng, Yongli Xiang, Ziming Hong, Zerong Lin, Chaojian Yu, Tongliang Liu, Xinge You
Abstract:
Image‑to‑Video (I2V) generation models, which condition video generation on reference images, have shown emerging visual instruction‑following capability, allowing certain visual cues in reference images to act as implicit control signals for video generation. However, this capability also introduces a previously overlooked risk: adversaries may exploit visual instructions to inject malicious intent through the image modality. In this work, we uncover this risk by proposing Visual Instruction Injection (VII), a training‑free and transferable jailbreaking framework that intentionally disguises the malicious intent of unsafe text prompts as benign visual instructions in the safe reference image. Specifically, VII coordinates a Malicious Intent Reprogramming module to distill malicious intent from unsafe text prompts while minimizing their static harmfulness, and a Visual Instruction Grounding module to ground the distilled intent onto a safe input image by rendering visual instructions that preserve semantic consistency with the original unsafe text prompt, thereby inducing harmful content during I2V generation. Empirically, our extensive experiments on four state‑of‑the‑art commercial I2V models (Kling‑v2.5‑turbo, Gemini Veo‑3.1, Seedance‑1.5‑pro, and PixVerse‑V5) demonstrate that VII achieves Attack Success Rates of up to 83.5% while reducing Refusal Rates to near zero, significantly outperforming existing baselines.
Authors:Jihao Qiu, Lingxi Xie, Xinyue Huo, Qi Tian, Qixiang Ye
Abstract:
This paper addresses the critical and underexplored challenge of long video understanding with low computational budgets. We propose LongVideo‑R1, an active, reasoning‑equipped multimodal large language model (MLLM) agent designed for efficient video context navigation, avoiding the redundancy of exhaustive search. At the core of LongVideo‑R1 lies a reasoning module that leverages high‑level visual cues to infer the most informative video clip for subsequent processing. During inference, the agent initiates traversal from top‑level visual summaries and iteratively refines its focus, immediately halting the exploration process upon acquiring sufficient knowledge to answer the query. To facilitate training, we first extract hierarchical video captions from CGBench, a video corpus with grounding annotations, and guide GPT‑5 to generate 33K high‑quality chain‑of‑thought‑with‑tool trajectories. The LongVideo‑R1 agent is fine‑tuned upon the Qwen‑3‑8B model through a two‑stage paradigm: supervised fine‑tuning (SFT) followed by reinforcement learning (RL), where RL employs a specifically designed reward function to maximize selective and efficient clip navigation. Experiments on multiple long video benchmarks validate the effectiveness of name, which enjoys superior tradeoff between QA accuracy and efficiency. All curated data and source code are provided in the supplementary material and will be made publicly available. Code and data are available at: https://github.com/qiujihao19/LongVideo‑R1
Authors:Hanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang, Hao Feng, Dingkang Yang, Chao Feng, Can Huang, Jingqun Tang, Xiang Bai
Abstract:
Visual Text Rendering (VTR) remains a critical challenge in text‑to‑image generation, where even advanced models frequently produce text with structural anomalies such as distortion, blurriness, and misalignment. However, we find that leading MLLMs and specialist OCR models largely fail to perceive these structural anomalies, creating a critical bottleneck for both VTR evaluation and RL‑based optimization. As a result, even state‑of‑the‑art generators (e.g., Seedream4.0, Qwen‑Image) still struggle to render structurally faithful text. To address this, we propose TextPecker, a plug‑and‑play structural anomaly perceptive RL strategy that mitigates noisy reward signals and works with any textto‑image generator. To enable this capability, we construct a recognition dataset with character‑level structural‑anomaly annotations and develop a stroke‑editing synthesis engine to expand structural‑error coverage. Experiments show that TextPecker consistently improves diverse text‑to‑image models; even on the well‑optimized Qwen‑Image, it significantly yields average gains of 4% in structural fidelity and 8.7% in semantic alignment for Chinese text rendering, establishing a new state‑of‑the‑art in high‑fidelity VTR. Our work fills a gap in VTR optimization, providing a foundational step towards reliable and structural faithful visual text generation.
Authors:Yuechen Xie, Xiaoyan Zhang, Yicheng Shan, Hao Zhu, Rui Tang, Rong Wei, Mingli Song, Yuanyu Wan, Jie Song
Abstract:
Vision‑Language Models (VLMs) have been increasingly applied in real‑world scenarios due to their outstanding understanding and reasoning capabilities. Although VLMs have already demonstrated impressive capabilities in common visual question answering and logical reasoning, they still lack the ability to make reasonable decisions in complex real‑world environments. We define this ability as spatial logical reasoning, which not only requires understanding the spatial relationships among objects in complex scenes, but also the logical dependencies between steps in multi‑step tasks. To bridge this gap, we introduce Spatial Logical Question Answering (SpatiaLQA), a benchmark designed to evaluate the spatial logical reasoning capabilities of VLMs. SpatiaLQA consists of 9,605 question answer pairs derived from 241 real‑world indoor scenes. We conduct extensive experiments on 41 mainstream VLMs, and the results show that even the most advanced models still struggle with spatial logical reasoning. To address this issue, we propose a method called recursive scene graph assisted reasoning, which leverages visual foundation models to progressively decompose complex scenes into task‑relevant scene graphs, thereby enhancing the spatial logical reasoning ability of VLMs, outperforming all previous methods. Code and dataset are available at https://github.com/xieyc99/SpatiaLQA.
Authors:Yongli Xiang, Ziming Hong, Zhaoqing Wang, Xiangyu Zhao, Bo Han, Tongliang Liu
Abstract:
Text‑to‑Image (T2I) diffusion models have demonstrated significant advancements in generating high‑quality images, while raising potential safety concerns regarding harmful content generation. Safety‑guidance‑based methods have been proposed to mitigate harmful outputs by steering generation away from harmful zones, where the zones are averaged across multiple harmful categories based on predefined keywords. However, these approaches fail to capture the complex interplay among different harm categories, leading to "harmful conflicts" where mitigating one type of harm may inadvertently amplify another, thus increasing overall harmful rate. To address this issue, we propose Conflict‑aware Adaptive Safety Guidance (CASG), a training‑free framework that dynamically identifies and applies the category‑aligned safety direction during generation. CASG is composed of two components: (i) Conflict‑aware Category Identification (CaCI), which identifies the harmful category most aligned with the model's evolving generative state, and (ii) Conflict‑resolving Guidance Application (CrGA), which applies safety steering solely along the identified category to avoid multi‑category interference. CASG can be applied to both latent‑space and text‑space safeguards. Experiments on T2I safety benchmarks demonstrate CASG's state‑of‑the‑art performance, reducing the harmful rate by up to 15.4% compared to existing methods.
Authors:Jiahao Xu, Sheng Huang, Xin Zhang, Zhixiong Nan, Jiajun Dong, Nankun Mu
Abstract:
In computational pathology, few‑shot whole slide image classification is primarily driven by the extreme scarcity of expert‑labeled slides. Recent vision‑language methods incorporate textual semantics generated by large language models, but treat these descriptions as static class‑level priors that are shared across all samples and lack sample‑wise refinement. This limits both the diversity and precision of visual‑semantic alignment, hindering generalization under limited supervision. To overcome this, we propose the stochastic MUlti‑view Semantic Enhancement (MUSE), a framework that first refines semantic precision via sample‑wise adaptation and then enhances semantic richness through retrieval‑augmented multi‑view generation. Specifically, MUSE introduces Sample‑wise Fine‑grained Semantic Enhancement (SFSE), which yields a fine‑grained semantic prior for each sample through MoE‑based adaptive visual‑semantic interaction. Guided by this prior, Stochastic Multi‑view Model Optimization (SMMO) constructs an LLM‑generated knowledge base of diverse pathological descriptions per class, then retrieves and stochastically integrates multiple matched textual views during training. These dynamically selected texts serve as enriched semantic supervisions to stochastically optimize the vision‑language model, promoting robustness and mitigating overfitting. Experiments on three benchmark WSI datasets show that MUSE consistently outperforms existing vision‑language baselines in few‑shot settings, demonstrating that effective few‑shot pathology learning requires not only richer semantic sources but also their active and sample‑aware semantic optimization. Our code is available at: https://github.com/JiahaoXu‑god/CVPR2026_MUSE.
Authors:Ran Zhang, Xuanhua He, Liu Liu
Abstract:
Image fusion seeks to integrate complementary information from multiple sources into a single, superior image. While traditional methods are fast, they lack adaptability and performance. Conversely, deep learning approaches achieve state‑of‑the‑art (SOTA) results but suffer from critical inefficiencies: their reliance on slow, resource‑intensive, patch‑based training introduces a significant gap with full‑resolution inference. We propose a novel hybrid framework that resolves this trade‑off. Our method utilizes a learnable U‑Net to generate a dynamic guidance map that directs a classic, fixed Laplacian pyramid fusion kernel. This decoupling of policy learning from pixel synthesis enables remarkably efficient full‑resolution training, eliminating the train‑inference gap. Consequently, our model achieves SOTA‑comparable performance in about one minute on a RTX 4090 or two minutes on a consumer laptop GPU from scratch without any external model and demonstrates powerful zero‑shot generalization across diverse tasks, from infrared‑visible to medical imaging. By design, the fused output is linearly constructed solely from source information, ensuring high faithfulness for critical applications. The codes are available at https://github.com/Zirconium233/HybridFusion
Authors:Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva, Alexandre Alahi, Paolo Favaro
Abstract:
Discrete image tokenizers have emerged as a key component of modern vision and multimodal systems, providing a sequential interface for transformer‑based architectures. However, most existing approaches remain primarily optimized for reconstruction and compression, often yielding tokens that capture local texture rather than object‑level semantic structure. Inspired by the incremental and compositional nature of human communication, we introduce COMmunication inspired Tokenization (COMiT), a framework for learning structured discrete visual token sequences. COMiT constructs a latent message within a fixed token budget by iteratively observing localized image crops and recurrently updating its discrete representation. At each step, the model integrates new visual information while refining and reorganizing the existing token sequence. After several encoding iterations, the final message conditions a flow‑matching decoder that reconstructs the full image. Both encoding and decoding are implemented within a single transformer model and trained end‑to‑end using a combination of flow‑matching reconstruction and semantic representation alignment losses. Our experiments demonstrate that while semantic alignment provides grounding, attentive sequential tokenization is critical for inducing interpretable, object‑centric token structure and substantially improving compositional generalization and relational reasoning over prior methods.
Authors:Yichen Xie, Chensheng Peng, Mazen Abdelfattah, Yihan Hu, Jiezhi Yang, Eric Higgins, Ryan Brigden, Masayoshi Tomizuka, Wei Zhan
Abstract:
World foundation models aim to simulate the evolution of the real world with physically plausible behavior. Unlike prior methods that handle spatial and temporal correlations separately, we propose RAYNOVA, a geometry‑agonistic multiview world model for driving scenarios that employs a dual‑causal autoregressive framework. It follows both scale‑wise and temporal topological orders in the autoregressive process, and leverages global attention for unified 4D spatio‑temporal reasoning. Different from existing works that impose strong 3D geometric priors, RAYNOVA constructs an isotropic spatio‑temporal representation across views, frames, and scales based on relative Plücker‑ray positional encoding, enabling robust generalization to diverse camera setups and ego motions. We further introduce a recurrent training paradigm to alleviate distribution drift in long‑horizon video generation. RAYNOVA achieves state‑of‑the‑art multi‑view video generation results on nuScenes, while offering higher throughput and strong controllability under diverse input conditions, generalizing to novel views and camera configurations without explicit 3D scene representation. Our code will be released at https://raynova‑ai.github.io/.
Authors:Xiaokai Bai, Jiahao Cheng, Songkai Wang, Yixuan Luo, Lianqing Zheng, Xiaohan Zhang, Si-Yuan Cao, Hui-Liang Shen
Abstract:
4D radar measurements offer an affordable and weather‑robust solution for 3D perception. However, the inherent sparsity and noise of radar point clouds present significant challenges for accurate 3D object detection, underscoring the need for effective and robust point clouds densification. Despite recent progress, existing densification methods often fail to address the extreme sparsity of 4D radar point clouds and exhibit limited robustness when processing scenes with a small number of points. In this paper, we propose SD4R, a novel framework that transforms sparse radar point clouds into dense representations. SD4R begins by utilizing a foreground point generator (FPG) to mitigate noise propagation and produce densified point clouds. Subsequently, a logit‑query encoder (LQE) enhances conventional pillarization, resulting in robust feature representations. Through these innovations, our SD4R demonstrates strong capability in both noise reduction and foreground point densification. Extensive experiments conducted on the publicly available View‑of‑Delft dataset demonstrate that SD4R achieves state‑of‑the‑art performance. Source code is available at https://github.com/lancelot0805/SD4R.
Authors:Yepeng Liu, Hao Li, Liwen Yang, Fangzhen Li, Xudi Ge, Yuliang Gu, kuang Gao, Bing Wang, Guang Chen, Hangjun Ye, Yongchao Xu
Abstract:
Keypoint‑based matching is a fundamental component of modern 3D vision systems, such as Structure‑from‑Motion (SfM) and SLAM. Most existing learning‑based methods are trained on image pairs, a paradigm that fails to explicitly optimize for the long‑term trackability of keypoints across sequences under challenging viewpoint and illumination changes. In this paper, we reframe keypoint detection as a sequential decision‑making problem. We introduce TraqPoint, a novel, end‑to‑end Reinforcement Learning (RL) framework designed to optimize the Track‑quality (Traq) of keypoints directly on image sequences. Our core innovation is a track‑aware reward mechanism that jointly encourages the consistency and distinctiveness of keypoints across multiple views, guided by a policy gradient method. Extensive evaluations on sparse matching benchmarks, including relative pose estimation and 3D reconstruction, demonstrate that TraqPoint significantly outperforms some state‑of‑the‑art (SOTA) keypoint detection and description methods.The code will be available at https://github.com/xiaomi‑research/traqpoint.
Authors:Yuejiao Su, Yi Wang, Lei Yao, Yawen Cui, Lap-Pui Chau
Abstract:
A fine‑grained understanding of egocentric human‑environment interactions is crucial for developing next‑generation embodied agents. One fundamental challenge in this area involves accurately parsing hands and active objects. While transformer‑based architectures have demonstrated considerable potential for such tasks, several key limitations remain unaddressed: 1) existing query initialization mechanisms rely primarily on semantic cues or learnable parameters, demonstrating limited adaptability to changing active objects across varying input scenes; 2) previous transformer‑based methods utilize pixel‑level semantic features to iteratively refine queries during mask generation, which may introduce interaction‑irrelevant content into the final embeddings; and 3) prevailing models are susceptible to "interaction illusion", producing physically inconsistent predictions. To address these issues, we propose an end‑to‑end Interaction‑aware Transformer (InterFormer), which integrates three key components, i.e., a Dynamic Query Generator (DQG), a Dual‑context Feature Selector (DFS), and the Conditional Co‑occurrence (CoCo) loss. The DQG explicitly grounds query initialization in the spatial dynamics of hand‑object contact, enabling targeted generation of interaction‑aware queries for hands and various active objects. The DFS fuses coarse interactive cues with semantic features, thereby suppressing interaction‑irrelevant noise and emphasizing the learning of interactive relationships. The CoCo loss incorporates hand‑object relationship constraints to enhance physical consistency in prediction. Our model achieves state‑of‑the‑art performance on both the EgoHOS and the challenging out‑of‑distribution mini‑HOI4D datasets, demonstrating its effectiveness and strong generalization ability. Code and models are publicly available at https://github.com/yuggiehk/InterFormer.
Authors:Hanhui Li, Xuan Huang, Wanquan Liu, Yuhao Cheng, Long Chen, Yiqiang Yan, Xiaodan Liang, Chenqiang Gao
Abstract:
Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well‑controlled environments and therefore degrade in real‑world settings with severe perturbations, such as hand‑object interactions, extreme poses, illumination changes, and motion blur. To tackle these issues, we introduce WildGHand, an optimization‑based framework that enables self‑adaptive 3D Gaussian splatting on in‑the‑wild videos and produces high‑fidelity hand avatars. WildGHand incorporates two key components: (i) a dynamic perturbation disentanglement module that explicitly represents perturbations as time‑varying biases on 3D Gaussian attributes during optimization, and (ii) a perturbation‑aware optimization strategy that generates per‑frame anisotropic weighted masks to guide optimization. Together, these components allow the framework to identify and suppress perturbations across both spatial and temporal dimensions. We further curate a dataset of monocular hand videos captured under diverse perturbations to benchmark in‑the‑wild hand avatar reconstruction. Extensive experiments on this dataset and two public datasets demonstrate that WildGHand achieves state‑of‑the‑art performance and substantially improves over its base model across multiple metrics (e.g., up to a 15.8% relative gain in PSNR and a 23.1% relative reduction in LPIPS). Our implementation and dataset are available at https://github.com/XuanHuang0/WildGHand.
Authors:Xinyong Cai, Changbin Sun, Yong Wang, Hongyu Yang, Yuankai Wu
Abstract:
Spatiotemporal predictive learning (STPL) aims to forecast future frames from past observations and is essential across a wide range of applications. Compared with recurrent or hybrid architectures, pure convolutional models offer superior efficiency and full parallelism, yet their fixed receptive fields limit their ability to adaptively capture spatially varying motion patterns. Inspired by biological center‑surround organization and frequency‑selective signal processing, we propose PFGNet, a fully convolutional framework that dynamically modulates receptive fields through pixel‑wise frequency‑guided gating. The core Peripheral Frequency Gating (PFG) block extracts localized spectral cues and adaptively fuses multi‑scale large‑kernel peripheral responses with learnable center suppression, effectively forming spatially adaptive band‑pass filters. To maintain efficiency, all large kernels are decomposed into separable 1D convolutions (1 × k followed by k × 1), reducing per‑channel computational cost from O(k^2) to O(2k). PFGNet enables structure‑aware spatiotemporal modeling without recurrence or attention. Experiments on Moving MNIST, TaxiBJ, Human3.6M, and KTH show that PFGNet delivers SOTA or near‑SOTA forecasting performance with substantially fewer parameters and FLOPs. Our code is available at https://github.com/fhjdqaq/PFGNet.
Authors:Limai Jiang, Ruitao Xie, Bokai Yang, Huazhen Huang, Juan He, Yufu Huo, Zikai Wang, Yang Wei, Yunpeng Cai
Abstract:
Medical image segmentation plays a vital role in clinical decision‑making, enabling precise localization of lesions and guiding interventions. Despite significant advances in segmentation accuracy, the black‑box nature of most deep models has raised growing concerns about their trustworthiness in high‑stakes medical scenarios. Current explanation techniques have primarily focused on classification tasks, leaving the segmentation domain relatively underexplored. We introduced an explanation model for segmentation task which employs the causal inference framework and backpropagates the average treatment effect (ATE) into a quantification metric to determine the influence of input regions, as well as network components, on target segmentation areas. Through comparison with recent segmentation explainability techniques on two representative medical imaging datasets, we demonstrated that our approach provides more faithful explanations than existing approaches. Furthermore, we carried out a systematic causal analysis of multiple foundational segmentation models using our method, which reveals significant heterogeneity in perceptual strategies across different models, and even between different inputs for the same model. Suggesting the potential of our method to provide notable insights for optimizing segmentation models. Our code can be found at https://github.com/lcmmai/PdCR.
Authors:Peiliang Cai, Jiacheng Liu, Haowen Xu, Xinyu Wang, Chang Zou, Linfeng Zhang
Abstract:
Diffusion models have achieved remarkable success in image and video generation tasks. However, the high computational demands of Diffusion Transformers (DiTs) pose a significant challenge to their practical deployment. While feature caching is a promising acceleration strategy, existing methods based on simple reusing or training‑free forecasting struggle to adapt to the complex, stage‑dependent dynamics of the diffusion process, often resulting in quality degradation and failing to maintain consistency with the standard denoising process. To address this, we propose a LEarnable Stage‑Aware (LESA) predictor framework based on two‑stage training. Our approach leverages a Kolmogorov‑Arnold Network (KAN) to accurately learn temporal feature mappings from data. We further introduce a multi‑stage, multi‑expert architecture that assigns specialized predictors to different noise‑level stages, enabling more precise and robust feature forecasting. Extensive experiments show our method achieves significant acceleration while maintaining high‑fidelity generation. Experiments demonstrate 5.00x acceleration on FLUX.1‑dev with minimal quality degradation (1.0% drop), 6.25x speedup on Qwen‑Image with a 20.2% quality improvement over the previous SOTA (TaylorSeer), and 5.00x acceleration on HunyuanVideo with a 24.7% PSNR improvement over TaylorSeer. State‑of‑the‑art performance on both text‑to‑image and text‑to‑video synthesis validates the effectiveness and generalization capability of our training‑based framework across different models. Our code is available at https://github.com/caipeiliang2004/LESA.
Authors:Jintu Zheng, Qizhe Liu, HuangXin Xu, Zhuojie Chen
Abstract:
While iterative stereo matching achieves high accuracy, its dependence on Recurrent Neural Networks (RNN) hinders edge deployment, a challenge underexplored in existing researches. We analyze iterative refinement and reveal that disparity updates are spatially sparse and temporally redundant. First, we introduce a progressive iteration pruning strategy that suppresses redundant update steps, effectively collapsing the recursive computation into a near‑single‑pass inference. Second, we propose a collaborative monocular prior transfer framework that implicitly embeds depth priors without requiring a dedicated monocular encoder, thereby eliminating its associated computational burden. Third, we develop FlashGRU, a hardware‑aware RNN operator leveraging structured sparsity and I/O‑conscious design, achieving a 7.28× speedup, 76.6% memory peak reduction and 80.9% global memory requests reduction over natvie ConvGRUs under 2K resolution. Our PipStereo enables real‑time, high‑fidelity stereo matching on edge hardware: it processes 320×640 frames in just 75ms on an NVIDIA Jetson Orin NX (FP16) and 19ms on RTX 4090, matching the accuracy of large iterative based models, and our generalization ability and accuracy far exceeds that of existing real‑time methods. Our embedded AI projects will be updated at: https://github.com/XPENG‑Aridge‑AI.
Authors:Taha Koleilat, Hojat Asgariandehkordi, Omid Nejati Manzari, Berardino Barile, Yiming Xiao, Hassan Rivaz
Abstract:
Medical image segmentation remains challenging due to limited annotations for training, ambiguous anatomical features, and domain shifts. While vision‑language models such as CLIP offer strong cross‑modal representations, their potential for dense, text‑guided medical image segmentation remains underexplored. We present MedCLIPSeg, a novel framework that adapts CLIP for robust, data‑efficient, and uncertainty‑aware medical image segmentation. Our approach leverages patch‑level CLIP embeddings through probabilistic cross‑modal attention, enabling bidirectional interaction between image and text tokens and explicit modeling of predictive uncertainty. Together with a soft patch‑level contrastive loss that encourages more nuanced semantic learning across diverse textual prompts, MedCLIPSeg effectively improves data efficiency and domain generalizability. Extensive experiments across 16 datasets spanning five imaging modalities and six organs demonstrate that MedCLIPSeg outperforms prior methods in accuracy, efficiency, and robustness, while providing interpretable uncertainty maps that highlight local reliability of segmentation results. This work demonstrates the potential of probabilistic vision‑language modeling for text‑driven medical image segmentation.
Authors:Aryan Garg, Sizhuo Ma, Mohit Gupta
Abstract:
Capturing high‑quality images from only a few detected photons is a fundamental challenge in computational imaging. Single‑photon avalanche diode (SPAD) sensors promise high‑quality imaging in regimes where conventional cameras fail, but raw \emphquanta frames contain only sparse, noisy, binary photon detections. Recovering a coherent image from a burst of such frames requires handling alignment, denoising, and demosaicing (for color) under noise statistics far outside those assumed by standard restoration pipelines or modern generative models. We present an approach that adapts large text‑to‑image latent diffusion models to the photon‑limited domain of quanta burst imaging. Our method leverages the structural and semantic priors of internet‑scale diffusion models while introducing mechanisms to handle Bernoulli photon statistics. By integrating latent‑space restoration with burst‑level spatio‑temporal reasoning, our approach produces reconstructions that are both photometrically faithful and perceptually pleasing, even under high‑speed motion. We evaluate the method on synthetic benchmarks and new real‑world datasets, including the first color SPAD burst dataset and a challenging Deforming (XD) video benchmark. Across all settings, the approach substantially improves perceptual quality over classical and modern learning‑based baselines, demonstrating the promise of adapting large generative priors to extreme photon‑limited sensing. Code at \hrefhttps://github.com/Aryan‑Garg/gQIRhttps://github.com/Aryan‑Garg/gQIR.
Authors:Mainak Singha, Sarthak Mehrotra, Paolo Casari, Subhasis Chaudhuri, Elisa Ricci, Biplab Banerjee
Abstract:
Recent vision‑language models (VLMs) such as CLIP demonstrate impressive cross‑modal reasoning, extending beyond images to 3D perception. Yet, these models remain fragile under domain shifts, especially when adapting from synthetic to real‑world point clouds. Conventional 3D domain adaptation approaches rely on heavy trainable encoders, yielding strong accuracy but at the cost of efficiency. We introduce CLIPoint3D, the first framework for few‑shot unsupervised 3D point cloud domain adaptation built upon CLIP. Our approach projects 3D samples into multiple depth maps and exploits the frozen CLIP backbone, refined through a knowledge‑driven prompt tuning scheme that integrates high‑level language priors with geometric cues from a lightweight 3D encoder. To adapt task‑specific features effectively, we apply parameter‑efficient fine‑tuning to CLIP's encoders and design an entropy‑guided view sampling strategy for selecting confident projections. Furthermore, an optimal transport‑based alignment loss and an uncertainty‑aware prototype alignment loss collaboratively bridge source‑target distribution gaps while maintaining class separability. Extensive experiments on PointDA‑10 and GraspNetPC‑10 benchmarks show that CLIPoint3D achieves consistent 3‑16% accuracy gains over both CLIP‑based and conventional encoder‑based baselines. Project page: https://sarthakm320.github.io/CLIPoint3D.
Authors:Bhavik Chandna, Kelsey R. Allen
Abstract:
AI video generation is evolving rapidly. For video generators to be useful for applications ranging from robotics to film‑making, they must consistently produce realistic videos. However, evaluating the realism of generated videos remains a largely manual process ‑‑ requiring human annotation or bespoke evaluation datasets which have restricted scope. Here we develop an automated evaluation framework for video realism which captures both semantics and coherent 3D structure and which does not require access to a reference video. Our method, 3DSPA, is a 3D spatiotemporal point autoencoder which integrates 3D point trajectories, depth cues, and DINO semantic features into a unified representation for video evaluation. 3DSPA models how objects move and what is happening in the scene, enabling robust assessments of realism, temporal consistency, and physical plausibility. Experiments show that 3DSPA reliably identifies videos which violate physical laws, is more sensitive to motion artifacts, and aligns more closely with human judgments of video quality and realism across multiple datasets. Our results demonstrate that enriching trajectory‑based representations with 3D semantics offers a stronger foundation for benchmarking generative video models, and implicitly captures physical rule violations. The code and pretrained model weights will be available at https://github.com/TheProParadox/3dspa_code.
Authors:C. J. Díaz Baso, I. J. Soler Poquet, C. Kuckein, M. van Noort, N. Poirier
Abstract:
The Sun is observed in unprecedented detail, enabling studies of its activity on very small spatiotemporal scales. However, the large volume of data collected by our telescopes cannot be fully analyzed with conventional methods. Popular machine learning methods identify general trends from observations, but tend to overlook unusual events due to their low frequency of occurrence. We study the applicability of unsupervised probabilistic methods to efficiently identify rare events in multidimensional solar observations and optimize our computational resources to the study of these extreme phenomena. We introduce Inspectorch, an open‑source framework that utilizes flow‑based models: flexible density estimators capable of learning the multidimensional distribution of solar observations. Once optimized, it assigns a probability to each sample, allowing us to identify unusual events. We apply this approach by applying it to observations from the Hinode Spectro‑Polarimeter, the Interface Region Imaging Spectrograph, the Microlensed Hyperspectral Imager at Swedish 1‑m Solar Telescope, the Atmospheric Imaging Assembly on board the Solar Dynamics Observatory and the Extreme Ultraviolet Imager on board Solar Orbiter. We find that the algorithm assigns consistently lower probabilities to spectra that exhibit unusual features. For example, it identifies profiles with very strong Doppler shifts, uncommon broadening, and temporal dynamics associated with small‑scale reconnection events, among others. As a result, Inspectorch demonstrates that density estimation using flow‑based models offers a powerful approach to identifying rare events in large solar datasets. The resulting probabilistic anomaly scores allow computational resources to be focused on the most informative and physically relevant events. We make our Python package publicly available at https://github.com/cdiazbas/inspectorch.
Authors:Guodong Chen, Huanshuo Dong, Mallesham Dasari
Abstract:
We present N4MC, the first 4D neural compression framework to efficiently compress time‑varying mesh sequences by exploiting their temporal redundancy. Unlike prior neural mesh compression methods that treat each mesh frame independently, N4MC takes inspiration from inter‑frame compression in 2D video codecs, and learns motion compensation in long mesh sequences. Specifically, N4MC converts consecutive irregular mesh frames into regular 4D tensors to provide a uniform and compact representation. These tensors are then condensed using an auto‑decoder, which captures both spatial and temporal correlations for redundancy removal. To enhance temporal coherence, we introduce a transformer‑based interpolation model that predicts intermediate mesh frames conditioned on latent embeddings derived from tracked volume centers, eliminating motion ambiguities. Extensive evaluations show that N4MC outperforms state‑of‑the‑art in rate‑distortion performance, while enabling real‑time decoding of 4D mesh sequences. The implementation of our method is available at: https://github.com/frozzzen3/N4MC.
Authors:Manish Kumar Govind, Dominick Reilly, Pu Wang, Srijan Das
Abstract:
Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision‑language‑action (VLA) models without explicit robot action supervision. However, latent actions derived solely from RGB observations primarily encode appearance‑driven dynamics and lack explicit 3D geometric structure, which is essential for precise and contact‑rich manipulation. To address this limitation, we introduce UniLACT, a transformer‑based VLA model that incorporates geometric structure through depth‑aware latent pretraining, enabling downstream policies to inherit stronger spatial priors. To facilitate this process, we propose UniLARN, a unified latent action learning framework based on inverse and forward dynamics objectives that learns a shared embedding space for RGB and depth while explicitly modeling their cross‑modal interactions. This formulation produces modality‑specific and unified latent action representations that serve as pseudo‑labels for the depth‑aware pretraining of UniLACT. Extensive experiments in both simulation and real‑world settings demonstrate the effectiveness of depth‑aware unified latent action representations. UniLACT consistently outperforms RGB‑based latent action baselines under in‑domain and out‑of‑domain pretraining regimes, as well as on both seen and unseen manipulation tasks.The project page is at https://manishgovind.github.io/unilact‑vla/
Authors:Xiwen Chen, Wenhui Zhu, Gen Li, Xuanzhao Dong, Yujian Xiong, Hao Wang, Peijie Qiu, Qingquan Song, Zhipeng Wang, Shao Tang, Yalin Wang, Abolfazl Razi
Abstract:
Multi‑modal large language models (MLLMs) achieve strong visual‑language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual token pruning to accelerate inference, while existing pruning methods overlook the underlying distributional structure of visual representations. We propose OTPrune, a training‑free framework that formulates pruning as distribution alignment via optimal transport (OT). By minimizing the 2‑Wasserstein distance between the full and pruned token distributions, OTPrune preserves both local diversity and global representativeness while reducing inference cost. Moreover, we derive a tractable submodular objective that enables efficient optimization, and theoretically prove its monotonicity and submodularity, providing a principled foundation for stable and efficient pruning. We further provide a comprehensive analysis that explains how distributional alignment contributes to stable and semantically faithful pruning. Comprehensive experiments on wider benchmarks demonstrate that OTPrune achieves superior performance‑efficiency tradeoffs compared to state‑of‑the‑art methods. The code is available at https://github.com/xiwenc1/OTPrune.
Authors:Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad, Ritesh Thawkar, Omkar Thawakar, Senmao Li, Hisham Cholakkal, Ian Reid, Eric P. Xing, Salman Khan, Fahad Shahbaz Khan
Abstract:
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data‑hungry and too heavy for deployment on edge devices. We present Mobile‑O, a compact vision‑language‑diffusion model that brings unified multimodal intelligence to a mobile device. Its core module, the Mobile Conditioning Projector (MCP), fuses vision‑language features with a diffusion generator using depthwise‑separable convolutions and layerwise alignment. This design enables efficient cross‑modal conditioning with minimal computational cost. Trained on only a few million samples and post‑trained in a novel quadruplet format (generation prompt, image, question, answer), Mobile‑O jointly enhances both visual understanding and generation capabilities. Despite its efficiency, Mobile‑O attains competitive or superior performance compared to other unified models, achieving 74% on GenEval and outperforming Show‑O and JanusFlow by 5% and 11%, while running 6x and 11x faster, respectively. For visual understanding, Mobile‑O surpasses them by 15.3% and 5.1% averaged across seven benchmarks. Running in only ~3s per 512x512 image on an iPhone, Mobile‑O establishes the first practical framework for real‑time unified multimodal understanding and generation on edge devices. We hope Mobile‑O will ease future research in real‑time unified multimodal intelligence running entirely on‑device with no cloud dependency. Our code, models, datasets, and mobile application are publicly available at https://amshaker.github.io/Mobile‑O/
Authors:Chen Wang, Hao Tan, Wang Yifan, Zhiqin Chen, Yuheng Liu, Kalyan Sunkavalli, Sai Bi, Lingjie Liu, Yiwei Hu
Abstract:
We propose tttLRM, a novel large 3D reconstruction model that leverages a Test‑Time Training (TTT) layer to enable long‑context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability. Our framework efficiently compresses multiple image observations into the fast weights of the TTT layer, forming an implicit 3D representation in the latent space that can be decoded into various explicit formats, such as Gaussian Splats (GS) for downstream applications. The online learning variant of our model supports progressive 3D reconstruction and refinement from streaming observations. We demonstrate that pretraining on novel view synthesis tasks effectively transfers to explicit 3D modeling, resulting in improved reconstruction quality and faster convergence. Extensive experiments show that our method achieves superior performance in feedforward 3D Gaussian reconstruction compared to state‑of‑the‑art approaches on both objects and scenes.
Authors:Wei-Cheng Huang, Jiaheng Han, Xiaohan Ye, Zherong Pan, Kris Hauser
Abstract:
Estimating simulation‑ready scenes from real‑world observations is crucial for downstream planning and policy learning tasks. Regretfully, existing methods struggle in cluttered environments, often exhibiting prohibitive computational cost, poor robustness, and restricted generality when scaling to multiple interacting objects. We propose a unified optimization‑based formulation for real‑to‑sim scene estimation that jointly recovers the shapes and poses of multiple rigid objects under physical constraints. Our method is built on two key technical innovations. First, we leverage the recently introduced shape‑differentiable contact model, whose global differentiability permits joint optimization over object geometry and pose while modeling inter‑object contacts. Second, we exploit the structured sparsity of the augmented Lagrangian Hessian to derive an efficient linear system solver whose computational cost scales favorably with scene complexity. Building on this formulation, we develop an end‑to‑end Simulation‑ready Physics‑Aware Reconstruction for Cluttered Scenes (SPARCS) pipeline, which integrates learning‑based object initialization, physics‑constrained joint shape‑pose optimization, and differentiable texture refinement. Experiments on cluttered scenes with up to 5 objects and 22 convex hulls demonstrate that our approach robustly reconstructs physically valid, simulation‑ready object shapes and poses. Project webpage: https://rory‑weicheng.github.io/SPARCS/.
Authors:Zanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang, Marco Cristani
Abstract:
Edge‑based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision‑language alignment, showing that isolating and aligning structural cues across modalities can greatly benefit fine‑tuning on long, detail‑rich captions, with a specific focus on improving cross‑modal retrieval. We introduce StructXLIP, a fine‑tuning alignment paradigm that extracts edge maps (e.g., Canny), treating them as proxies for the visual structure of an image, and filters the corresponding captions to emphasize structural cues, making them "structure‑centric". Fine‑tuning augments the standard alignment loss with three structure‑centric losses: (i) aligning edge maps with structural text, (ii) matching local edge regions to textual chunks, and (iii) connecting edge maps to color images to prevent representation drift. From a theoretical standpoint, while standard CLIP maximizes the mutual information between visual and textual embeddings, StructXLIP additionally maximizes the mutual information between multimodal structural representations. This auxiliary optimization is intrinsically harder, guiding the model toward more robust and semantically stable minima, enhancing vision‑language alignment. Beyond outperforming current competitors on cross‑modal retrieval in both general and specialized domains, our method serves as a general boosting recipe that can be integrated into future approaches in a plug‑and‑play manner. Code and pretrained models are publicly available at: https://github.com/intelligolabs/StructXLIP.
Authors:Harry Anthony, Ziyun Liang, Hermione Warr, Konstantinos Kamnitsas
Abstract:
Deep Neural Networks achieve high performance in vision tasks by learning features from regions of interest (ROI) within images, but their performance degrades when deployed on out‑of‑distribution (OOD) data that differs from training data. This challenge has led to OOD detection methods that aim to identify and reject unreliable predictions. Although prior work shows that OOD detection performance varies by artefact type, the underlying causes remain underexplored. To this end, we identify a previously unreported bias in OOD detection: for hard‑to‑detect artefacts (near‑OOD), detection performance typically improves when the artefact shares visual similarity (e.g. colour) with the model's ROI and drops when it does not ‑ a phenomenon we term the Invisible Gorilla Effect. For example, in a skin lesion classifier with red lesion ROI, we show the method Mahalanobis Score achieves a 31.5% higher AUROC when detecting OOD red ink (similar to ROI) compared to black ink (dissimilar) annotations. We annotated artefacts by colour in 11,355 images from three public datasets (e.g. ISIC) and generated colour‑swapped counterfactuals to rule out dataset bias. We then evaluated 40 OOD methods across 7 benchmarks and found significant performance drops for most methods when artefacts differed from the ROI. Our findings highlight an overlooked failure mode in OOD detection and provide guidance for more robust detectors. Code and annotations are available at: https://github.com/HarryAnthony/Invisible_Gorilla_Effect.
Authors:Junli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing, Pengfei Li, Guang Li, Kun Ma, Guang Chen, Hangjun Ye, Zhongpu Xia, Long Chen, Qichao Zhang
Abstract:
Generative models have shown great potential in trajectory planning. Recent studies demonstrate that anchor‑guided generative models are effective in modeling the uncertainty of driving behaviors and improving overall performance. However, these methods rely on discrete anchor vocabularies that must sufficiently cover the trajectory distribution during testing to ensure robustness, inducing an inherent trade‑off between vocabulary size and model performance. To overcome this limitation, we propose MeanFuser, an end‑to‑end autonomous driving method that enhances both efficiency and robustness through three key designs. (1) We introduce Gaussian Mixture Noise (GMN) to guide generative sampling, enabling a continuous representation of the trajectory space and eliminating the dependency on discrete anchor vocabularies. (2) We adapt ``MeanFlow Identity" to end‑to‑end planning, which models the mean velocity field between GMN and trajectory distribution instead of the instantaneous velocity field used in vanilla flow matching methods, effectively eliminating numerical errors from ODE solvers and significantly accelerating inference. (3) We design a lightweight Adaptive Reconstruction Module (ARM) that enables the model to implicitly select from all sampled proposals or reconstruct a new trajectory when none is satisfactory via attention weights.Experiments on the NAVSIM closed‑loop benchmark demonstrate that MeanFuser achieves outstanding performance without the supervision of the PDM Score and exceptional inference efficiency, offering a robust and efficient solution for end‑to‑end autonomous driving. Our code and model are available at https://github.com/wjl2244/MeanFuser.
Authors:Christof Leitgeb, Thomas Puchleitner, Max Peter Ronecker, Daniel Watzenig
Abstract:
Automotive perception systems are obligated to meet high requirements. While optical sensors such as Camera and Lidar struggle in adverse weather conditions, Radar provides a more robust perception performance, effectively penetrating fog, rain, and snow. Since full Radar tensors have large data sizes and very few datasets provide them, most Radar‑based approaches work with sparse point clouds or 2D projections, which can result in information loss. Additionally, deep learning methods show potential to extract richer and more dense features from low level Radar data and therefore significantly increase the perception performance. Therefore, we propose a 3D projection method for fast‑Fourier‑transformed 4D Range‑Azimuth‑Doppler‑Elevation (RADE) tensors. Our method preserves rich Doppler and Elevation features while reducing the required data size for a single frame by 91.9% compared to a full tensor, thus achieving higher training and inference speed as well as lower model complexity. We introduce RADE‑Net, a lightweight model tailored to 3D projections of the RADE tensor. The backbone enables exploitation of low‑level and high‑level cues of Radar tensors with spatial and channel‑attention. The decoupled detection heads predict object center‑points directly in the Range‑Azimuth domain and regress rotated 3D bounding boxes from rich feature maps in the cartesian scene. We evaluate the model on scenes with multiple different road users and under various weather conditions on the large‑scale K‑Radar dataset and achieve a 16.7% improvement compared to their baseline, as well as 6.5% improvement over current Radar‑only models. Additionally, we outperform several Lidar approaches in scenarios with adverse weather conditions. The code is available under https://github.com/chr‑is‑tof/RADE‑Net.
Authors:Yixin Yang, Bojian Wu, Yang Zhou, Hui Huang
Abstract:
Due to the real‑time rendering performance, 3D Gaussian Splatting (3DGS) has emerged as the leading method for radiance field reconstruction. However, its reliance on spherical harmonics for color encoding inherently limits its ability to separate diffuse and specular components, making it challenging to accurately represent complex reflections. To address this, we propose a novel enhanced Gaussian kernel that explicitly models specular effects through view‑dependent opacity. Meanwhile, we introduce an error‑driven compensation strategy to improve rendering quality in existing 3DGS scenes. Our method begins with 2D Gaussian initialization and then adaptively inserts and optimizes enhanced Gaussian kernels, ultimately producing an augmented radiance field. Experiments demonstrate that our method not only surpasses state‑of‑the‑art NeRF methods in rendering performance but also achieves greater parameter efficiency. Project page at: https://xiaoxinyyx.github.io/augs.
Authors:Junyi Wang, Yudong Guo, Boyang Guo, Shengming Yang, Juyong Zhang
Abstract:
While diffusion models have shown great potential in portrait generation, generating expressive, coherent, and controllable cinematic portrait videos remains a significant challenge. Existing intermediate signals for portrait generation, such as 2D landmarks and parametric models, have limited disentanglement capabilities and cannot express personalized details due to their sparse or low‑rank representation. Therefore, existing methods based on these models struggle to accurately preserve subject identity and expressions, hindering the generation of highly expressive portrait videos. To overcome these limitations, we propose a high‑fidelity personalized head representation that more effectively disentangles expression and identity. This representation captures both static, subject‑specific global geometry and dynamic, expression‑related details. Furthermore, we introduce an expression transfer module to achieve personalized transfer of head pose and expression details between different identities. We use this sophisticated and highly expressive head model as a conditional signal to train a diffusion transformer (DiT)‑based generator to synthesize richly detailed portrait videos. Extensive experiments on self‑ and cross‑reenactment tasks demonstrate that our method outperforms previous models in terms of identity preservation, expression accuracy, and temporal stability, particularly in capturing fine‑grained details of complex motion.
Authors:Blaž Rolih, Matic Fučka, Filip Wolf, Luka Čehovin Zajc
Abstract:
Unsupervised change detection (UCD) in remote sensing aims to localise semantic changes between two images of the same region without relying on labelled data during training. Most recent approaches rely either on frozen foundation models in a training‑free manner or on training with synthetic changes generated in pixel space. Both strategies inherently rely on predefined assumptions about change types, typically introduced through handcrafted rules, external datasets, or auxiliary generative models. Due to these assumptions, such methods fail to generalise beyond a few change types, limiting their real‑world usage, especially in rare or complex scenarios. To address this, we propose MaSoN (Make Some Noise), an end‑to‑end UCD framework that synthesises diverse changes directly in the latent feature space during training. It generates changes that are dynamically estimated using feature statistics of target data, enabling diverse yet data‑driven variation aligned with the target domain. It also easily extends to new modalities, such as SAR. MaSoN generalises strongly across diverse change types and achieves state‑of‑the‑art performance on five benchmarks, improving the average F1 score by 14.1 percentage points. Project page: https://blaz‑r.github.io/mason_ucd
Authors:Lucas Martini, Alexander Lappe, Anna Bognár, Rufin Vogels, Martin A. Giese
Abstract:
The recognition of dynamic and social behavior in animals is fundamental for advancing ethology, ecology, medicine and neuroscience. Recent progress in deep learning has enabled automated behavior recognition from video, yet an accurate reconstruction of the three‑dimensional (3D) pose and shape has not been integrated into this process. Especially for non‑human primates, mesh‑based tracking efforts lag behind those for other species, leaving pose descriptions restricted to sparse keypoints that are unable to fully capture the richness of action dynamics. To address this gap, we introduce the Big MacaQue 3D Motion and Animation Dataset (\textttBigMaQ), a large‑scale dataset comprising more than 750 scenes of interacting rhesus macaques with detailed 3D pose descriptions. Extending previous surface‑based animal tracking methods, we construct subject‑specific textured avatars by adapting a high‑quality macaque template mesh to individual monkeys. This allows us to provide pose descriptions that are more accurate than previous state‑of‑the‑art surface‑based animal tracking methods. From the original dataset, we derive BigMaQ500, an action recognition benchmark that links surface‑based pose vectors to single frames across multiple individual monkeys. By pairing features extracted from established image and video encoders with and without our pose descriptors, we demonstrate substantial improvements in mean average precision (mAP) when pose information is included. With these contributions, \textttBigMaQ establishes the first dataset that both integrates dynamic 3D pose‑shape representations into the learning task of animal action recognition and provides a rich resource to advance the study of visual appearance, posture, and social interaction in non‑human primates. The code and data are publicly available at https://martinivis.github.io/BigMaQ/ .
Authors:Qiankun Ma, Ziyao Zhang, Haofei Wang, Jie Chen, Zhen Song, Hairong Zheng
Abstract:
Recent Vision‑Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead and degrade inference efficiency. Prior studies typically relies on [CLS] attention or text‑vision cross‑attention to identify and discard redundant visual tokens. Despite promising results, such solutions are prone to introduce positional bias and, more critically, are incompatible with efficient attention kernels such as FlashAttention, limiting their practical deployment for VLM acceleration. In this paper, we step away from attention dependencies and revisit visual token compression from an information‑theoretic perspective, aiming to maximally preserve visual information without any attention involvement. We present ApET, an Approximation‑Error guided Token compression framework. ApET first reconstructs the original visual tokens with a small set of basis tokens via linear approximation, then leverages the approximation error to identify and drop the least informative tokens. Extensive experiments across multiple VLMs and benchmarks demonstrate that ApET retains 95.2% of the original performance on image‑understanding tasks and even attains 100.4% on video‑understanding tasks, while compressing the token budgets by 88.9% and 87.5%, respectively. Thanks to its attention‑free design, ApET seamlessly integrates with FlashAttention, enabling further inference acceleration and making VLM deployment more practical. Code is available at https://github.com/MaQianKun0/ApET.
Authors:Filip Wolf, Blaž Rolih, Luka Čehovin Zajc
Abstract:
Foundation models are transforming Earth Observation (EO), yet the diversity of EO sensors and modalities makes a single universal model unrealistic. Multiple specialized EO foundation models (EOFMs) will likely coexist, making efficient knowledge transfer across modalities essential. Most existing EO pretraining relies on masked image modeling, which emphasizes local reconstruction but provides limited control over global semantic structure. To address this, we propose a dual‑teacher contrastive distillation framework for multispectral imagery that aligns the student's pretraining objective with the contrastive self‑distillation paradigm of modern optical vision foundation models (VFMs). Our approach combines a multispectral teacher with an optical VFM teacher, enabling coherent cross‑modal representation learning. Experiments across diverse optical and multispectral benchmarks show that our model adapts to multispectral data without compromising performance on optical‑only inputs, achieving state‑of‑the‑art results in both settings, with an average improvement of 3.64 percentage points in semantic segmentation, 1.2 in change detection, and 1.31 in classification tasks. This demonstrates that contrastive distillation provides a principled and efficient approach to scalable representation learning across heterogeneous EO data sources. Project page: \textcolormagentahttps://wolfilip.github.io/DEO/.
Authors:Penghui Niu, Taotao Cai, Suqi Zhang, Junhua Gu, Ping Zhang, Qiqi Liu, Jianxin Li
Abstract:
The inherent intermittency and high‑frequency variability of solar irradiance, particularly during rapid cloud advection, present significant stability challenges to high‑penetration photovoltaic grids. Although multimodal forecasting has emerged as a viable mitigation strategy, existing architectures predominantly rely on shallow feature concatenation and binary cloud segmentation, thereby failing to capture the fine‑grained optical features of clouds and the complex spatiotemporal coupling between visual and meteorological modalities. To bridge this gap, this paper proposes M3S‑Net, a novel multimodal feature fusion network based on multi‑scale data for ultra‑short‑term PV power forecasting. First, a multi‑scale partial channel selection network leverages partial convolutions to explicitly isolate the boundary features of optically thin clouds, effectively transcending the precision limitations of coarse‑grained binary masking. Second, a multi‑scale sequence to image analysis network employs Fast Fourier Transform (FFT)‑based time‑frequency representation to disentangle the complex periodicity of meteorological data across varying time horizons. Crucially, the model incorporates a cross‑modal Mamba interaction module featuring a novel dynamic C‑matrix swapping mechanism. By exchanging state‑space parameters between visual and temporal streams, this design conditions the state evolution of one modality on the context of the other, enabling deep structural coupling with linear computational complexity, thus overcoming the limitations of shallow concatenation. Experimental validation on the newly constructed fine‑grained PV power dataset demonstrates that M3S‑Net achieves a mean absolute error reduction of 6.2% in 10‑minute forecasts compared to state‑of‑the‑art baselines. The dataset and source code will be available at https://github.com/she1110/FGPD.
Authors:Kaifa Yang, Qi Yang, Yiling Xu, Zhu Li
Abstract:
3D Gaussian Splatting (3DGS) has emerged as a leading technology for high‑quality 3D scene reconstruction. However, the iterative refinement and densification process leads to the generation of a large number of primitives, each contributing to the reconstruction to a substantially different extent. Estimating primitive importance is thus crucial, both for removing redundancy during reconstruction and for enabling efficient compression and transmission. Existing methods typically rely on rendering‑based analyses, where each primitive is evaluated through its contribution across multiple camera viewpoints. However, such methods are sensitive to the number and selection of views, rely on specialized differentiable rasterizers, and have long calculation times that grow linearly with view count, making them difficult to integrate as plug‑and‑play modules and limiting scalability and generalization. To address these issues, we propose RAP, a fast feedforward rendering‑free attribute‑guided method for efficient importance score prediction in 3DGS. RAP infers primitive significance directly from intrinsic Gaussian attributes and local neighborhood statistics, avoiding rendering‑based or visibility‑dependent computations. A compact MLP predicts per‑primitive importance scores using rendering loss, pruning‑aware loss, and significance distribution regularization. After training on a small set of scenes, RAP generalizes effectively to unseen data and can be seamlessly integrated into reconstruction, compression, and transmission pipelines. Our code is publicly available at https://github.com/yyyykf/RAP.
Authors:Amir Hamza, Davide Boscaini, Weihang Li, Benjamin Busam, Fabio Poiesi
Abstract:
Existing methods for instance‑level 6D pose estimation typically rely on neural networks that either directly regress the pose in \mathrmSE(3) or estimate it indirectly via local feature matching. The former struggle with object symmetries, while the latter fail in the absence of distinctive local features. To overcome these limitations, we propose a novel formulation of 6D pose estimation as a conditional flow matching problem in \mathbbR^3. We introduce Flose, a generative method that infers object poses via a denoising process conditioned on local features. While prior approaches based on conditional flow matching perform denoising solely based on geometric guidance, Flose integrates appearance‑based semantic features to mitigate ambiguities caused by object symmetries. We further incorporate RANSAC‑based registration to handle outliers. We validate Flose on five datasets from the established BOP benchmark. Flose outperforms prior methods with an average improvement of +4.5 Average Recall. Project Website : https://tev‑fbk.github.io/Flose/
Authors:Kartik Kuckreja, Parul Gupta, Muhammad Haris Khan, Abhinav Dhall
Abstract:
Deepfake detection models often generate natural‑language explanations, yet their reasoning is frequently ungrounded in visual evidence, limiting reliability. Existing evaluations measure classification accuracy but overlook reasoning fidelity. We propose DeepfakeJudge, a framework for scalable reasoning supervision and evaluation, that integrates an out‑of‑distribution benchmark containing recent generative and editing forgeries, a human‑annotated subset with visual reasoning labels, and a suite of evaluation models, that specialize in evaluating reasoning rationales without the need for explicit ground truth reasoning rationales. The Judge is optimized through a bootstrapped generator‑evaluator process that scales human feedback into structured reasoning supervision and supports both pointwise and pairwise evaluation. On the proposed meta‑evaluation benchmark, our reasoning‑bootstrapped model achieves an accuracy of 96.2%, outperforming \texttt30x larger baselines. The reasoning judge attains very high correlation with human ratings and 98.9% percent pairwise agreement on the human‑annotated meta‑evaluation subset. These results establish reasoning fidelity as a quantifiable dimension of deepfake detection and demonstrate scalable supervision for interpretable deepfake reasoning. Our user study shows that participants preferred the reasonings generated by our framework 70% of the time, in terms of faithfulness, groundedness, and usefulness, compared to those produced by other models and datasets. All of our datasets, models, and codebase are \hrefhttps://github.com/KjAeRsTuIsK/DeepfakeJudgeopen‑sourced.
Authors:Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu
Abstract:
Existing Vision‑Language‑Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high‑level perception with sparse, embodiment‑specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA), they excel at semantic identification but often overlook subtle 3D state variations that dictate distinct action patterns. To resolve these misalignments, we propose Pose‑VLA, a decoupled paradigm that separates VLA training into a pre‑training phase for extracting universal 3D spatial priors in a unified camera‑centric space, and a post‑training phase for efficient embodiment alignment within robot‑specific action space. By introducing discrete pose tokens as a universal representation, Pose‑VLA seamlessly integrates spatial grounding from diverse 3D datasets with geometry‑level trajectories from robotic demonstrations. Our framework follows a two‑stage pre‑training pipeline, establishing fundamental spatial grounding via poses followed by motion alignment through trajectory supervision. Extensive evaluations demonstrate that Pose‑VLA achieves state‑of‑the‑art results on RoboTwin 2.0 with a 79.5% average success rate and competitive performance on LIBERO at 96.0%. Real‑world experiments further showcase robust generalization across diverse objects using only 100 demonstrations per task, validating the efficiency of our pre‑training paradigm.
Authors:Yo-Tin Lin, Su-Kai Chen, Hou-Ning Hu, Yen-Yu Lin, Yu-Lun Liu
Abstract:
Single LDR to HDR reconstruction remains challenging for over‑exposed regions where traditional methods often fail due to complete information loss. We present a training‑free approach that enhances existing indirect and direct HDR reconstruction methods through diffusion‑based inpainting. Our method combines text‑guided diffusion models with SDEdit refinement to generate plausible content in over‑exposed areas while maintaining consistency across multi‑exposure LDR images. Unlike previous approaches requiring extensive training, our method seamlessly integrates with existing HDR reconstruction techniques through an iterative compensation mechanism that ensures luminance coherence across multiple exposures. We demonstrate significant improvements in both perceptual quality and quantitative metrics on standard HDR datasets and in‑the‑wild captures. Results show that our method effectively recovers natural details in challenging scenarios while preserving the advantages of existing HDR reconstruction pipelines. Project page: https://github.com/EusdenLin/HDR‑Reconstruction‑Boosting
Authors:Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta
Abstract:
Key part of robotics, augmented reality, and digital inspection is dense 3D reconstruction from depth observations. Traditional volumetric fusion techniques, including truncated signed distance functions (TSDF), enable efficient and deterministic geometry reconstruction; however, they depend on heuristic weighting and fail to transparently convey uncertainty in a systematic way. Recent neural implicit methods, on the other hand, get very high fidelity but usually need a lot of GPU power for optimization and aren't very easy to understand for making decisions later on. This work presents BayesFusion‑SDF, a CPU‑centric probabilistic signed distance fusion framework that conceptualizes geometry as a sparse Gaussian random field with a defined posterior distribution over voxel distances. First, a rough TSDF reconstruction is used to create an adaptive narrow‑band domain. Then, depth observations are combined using a heteroscedastic Bayesian formulation that is solved using sparse linear algebra and preconditioned conjugate gradients. Randomized diagonal estimators are a quick way to get an idea of posterior uncertainty. This makes it possible to extract surfaces and plan the next best view while taking into account uncertainty. Tests on a controlled ablation scene and a CO3D object sequence show that the new method is more accurate geometrically than TSDF baselines and gives useful estimates of uncertainty for active sensing. The proposed formulation provides a clear and easy‑to‑use alternative to GPU‑heavy neural reconstruction methods while still being able to be understood in a probabilistic way and acting in a predictable way. GitHub: https://mazumdarsoumya.github.io/BayesFusionSDF
Authors:Jonas Serych, Jiri Matas
Abstract:
We present SAM‑H and WOFTSAM, novel planar trackers that combine robust long‑term segmentation tracking provided by SAM 2 with 8 degrees‑of‑freedom homography pose estimation. SAM‑H estimates homographies from segmentation mask contours and is thus highly robust to target appearance changes. WOFTSAM significantly improves the current state‑of‑the‑art planar tracker WOFT by exploiting lost target re‑detection provided by SAM‑H. The proposed methods are evaluated on POT‑210 and PlanarTrack tracking benchmarks, setting the new state‑of‑the‑art performance on both. On the latter, they outperform the second best by a large margin, +12.4 and +15.2pp on the p@15 metric. We also present improved ground‑truth annotations of initial PlanarTrack poses, enabling more accurate benchmarking in the high‑precision p@5 metric. The code and the re‑annotations are available at https://github.com/serycjon/WOFTSAM
Authors:Mingxiu Cai, Zhe Zhang, Gaochang Wu, Tianyou Chai, Xiatian Zhu
Abstract:
Unsupervised Anomaly Detection (UAD) aims to identify abnormal regions by establishing correspondences between test images and normal templates. Existing methods primarily rely on image reconstruction or template retrieval but face a fundamental challenge: matching between test images and normal templates inevitably introduces noise due to intra‑class variations, imperfect correspondences, and limited templates. Observing that Retrieval‑Augmented Generation (RAG) leverages retrieved samples directly in the generation process, we reinterpret UAD through this lens and introduce RAID, a retrieval‑augmented UAD framework designed for noise‑resilient anomaly detection and localization. Unlike standard RAG that enriches context or knowledge, we focus on using retrieved normal samples to guide noise suppression in anomaly map generation. RAID retrieves class‑, semantic‑, and instance‑level representations from a hierarchical vector database, forming a coarse‑to‑fine pipeline. A matching cost volume correlates the input with retrieved exemplars, followed by a guided Mixture‑of‑Experts (MoE) network that leverages the retrieved samples to adaptively suppress matching noise and produce fine‑grained anomaly maps. RAID achieves state‑of‑the‑art performance across full‑shot, few‑shot, and multi‑dataset settings on MVTec, VisA, MPDD, and BTAD benchmarks. \hrefhttps://github.com/Mingxiu‑Cai/RAIDhttps://github.com/Mingxiu‑Cai/RAID.
Authors:Girmaw Abebe Tadesse, Titien Bartette, Andrew Hassanali, Allen Kim, Jonathan Chemla, Andrew Zolli, Yves Ubelmann, Caleb Robinson, Inbal Becker-Reshef, Juan Lavista Ferres
Abstract:
Looting at archaeological sites poses a severe risk to cultural heritage, yet monitoring thousands of remote locations remains operationally difficult. We present a scalable and satellite‑based pipeline to detect looted archaeological sites, using PlanetScope monthly mosaics (4.7m/pixel) and a curated dataset of 1,943 archaeological sites in Afghanistan (898 looted, 1,045 preserved) with multi‑year imagery (2016‑‑2023) and site‑footprint masks. We compare (i) end‑to‑end CNN classifiers trained on raw RGB patches and (ii) traditional machine learning (ML) trained on handcrafted spectral/texture features and embeddings from recent remote‑sensing foundation models. Results indicate that ImageNet‑pretrained CNNs combined with spatial masking reach an F1 score of 0.926, clearly surpassing the strongest traditional ML setup, which attains an F1 score of 0.710 using SatCLIP‑V+RF+Mean, i.e., location and vision embeddings fed into a Random Forest with mean‑based temporal aggregation. Ablation studies demonstrate that ImageNet pretraining (even in the presence of domain shift) and spatial masking enhance performance. In contrast, geospatial foundation model embeddings perform competitively with handcrafted features, suggesting that looting signatures are extremely localized. The repository is available at https://github.com/microsoft/looted_site_detection.
Authors:Yihang Tao, Senkang Hu, Haonan An, Zhengru Fang, Hangcheng Cao, Yuguang Fang
Abstract:
Collaborative perception (CP) enables data sharing among connected and autonomous vehicles (CAVs) to enhance driving safety. However, CP systems are vulnerable to adversarial attacks where malicious agents forge false objects via feature‑level perturbations. Current defensive systems use threshold‑based consensus verification by comparing collaborative and ego detection results. Yet, these defenses remain vulnerable to more sophisticated attack strategies that could exploit two critical weaknesses: (i) lack of robustness against attacks with systematic timing and target region optimization, and (ii) inadvertent disclosure of vulnerability knowledge through implicit confidence information in shared collaboration data. In this paper, we propose MVIG attack, a novel adaptive adversarial CP framework learning to capture vulnerability knowledge disclosed by different defensive CP systems from a unified mutual view information graph (MVIG) representation. Our approach combines MVIG representation with temporal graph learning to generate evolving fabrication risk maps and employs entropy‑aware vulnerability search to optimize attack location, timing and persistence, enabling adaptive attacks with generalizability across various defensive configurations. Extensive evaluations on OPV2V and Adv‑OPV2V datasets demonstrate that MVIG attack reduces defense success rates by up to 62% against state‑of‑the‑art defenses while achieving 47% lower detection for persistent attacks at 29.9 FPS, exposing critical security gaps in CP systems. Code will be released at https://github.com/yihangtao/MVIG.git
Authors:Minseo Kim, Minchan Kwon, Dongyeun Lee, Yunho Jeon, Junmo Kim
Abstract:
Personalized text‑to‑image (T2I) generation has emerged as a key application for creating user‑specific concepts from a few reference images. The core challenge is concept disentanglement: separating the target concept from irrelevant residual information. Lacking such disentanglement, capturing high‑fidelity features often incorporates undesired attributes that conflict with user prompts, compromising the trade‑off between concept fidelity and text alignment. While existing methods rely on manual guidance, they often fail to represent intricate visual details and lack scalability. We introduce ConceptPrism, a framework that extracts shared features exclusively through cross‑image comparison without external information. We jointly optimize a target token and image‑wise residual tokens via reconstruction and exclusion losses. By suppressing shared information in residual tokens, the exclusion loss creates an information vacuum that forces the target token to capture the common concept. Extensive evaluations demonstrate that ConceptPrism achieves accurate concept disentanglement and significantly improves overall performance across diverse and complex visual concepts. The code is available at https://github.com/Minseo‑Kimm/ConceptPrism.
Authors:Jingyuan Wang, Li Niu
Abstract:
Generative image composition aims to regenerate the given foreground object in the background image to produce a realistic composite image. Some high‑authenticity methods can adjust foreground pose/view to be compatible with background, while some high‑fidelity methods can preserve the foreground details accurately. However, existing methods can hardly achieve both goals at the same time. In this work, we propose a two‑stage strategy to achieve both goals. In the first stage, we use high‑authenticity method to generate reasonable foreground shape, serving as the condition of high‑fidelity method in the second stage. The experiments on MureCOM dataset verify the effectiveness of our two‑stage strategy. The code and model have been released at https://github.com/bcmi/OSInsert‑Image‑Composition.
Authors:Chongyang Gao, Diji Yang, Shuyan Zhou, Xichen Yan, Luchuan Song, Shuo Li, Kezhen Chen
Abstract:
We introduce CFE‑Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE‑Bench is curated from repeatedly used, authentic university homework and exam problems, paired with reference solutions provided by course instructors. CFE‑Bench remains challenging for frontier models: the newly released Gemini‑3.1‑pro‑preview achieves 59.69% overall accuracy, while the second‑best model, Gemini‑3‑flash‑preview, reaches 55.46%, leaving substantial room for improvement. Beyond aggregate scores, we conduct a diagnostic analysis by decomposing instructor reference solutions into structured reasoning flows. We find that while frontier models often answer intermediate sub‑questions correctly, they struggle to reliably derive and maintain correct intermediate states throughout multi‑step solutions. We further observe that model‑generated solutions typically contain more reasoning steps than instructor solutions, indicating lower step efficiency and a higher risk of error accumulation. Data and code are available at https://github.com/Analogy‑AI/CFE_Bench.
Authors:Pengxi Liu, Zeyu Michael Li, Xiang Cheng
Abstract:
We introduce a variational framework for diffusion models with anisotropic noise schedules parameterized by a matrix‑valued path M_t(θ) that allocates noise across subspaces. Central to our framework is a trajectory‑level objective that jointly trains the score network and learns M_t(θ), which encompasses general parameterization classes of matrix‑valued noise schedules. We further derive an estimator for the derivative with respect to θ of the score that enables efficient optimization of the M_t(θ) schedule. For inference, we develop an efficiently‑implementable reverse‑ODE solver that is an anisotropic generalization of the second‑order Heun discretization algorithm. Across CIFAR‑10, AFHQv2, FFHQ, and ImageNet‑64, our method consistently improves upon the baseline EDM model in all NFE regimes. Code is available at https://github.com/lizeyu090312/anisotropic‑diffusion‑paper.
Authors:Mingrui Wu, Hao Chen, Jiayi Ji, Xiaoshuai Sun, Zhiyuan Liu, Liujuan Cao, Ming-Ming Cheng, Rongrong Ji
Abstract:
We propose ControlMLLM++, a novel test‑time adaptation framework that injects learnable visual prompts into frozen multimodal large language models (MLLMs) to enable fine‑grained region‑based visual reasoning without any model retraining or fine‑tuning. Leveraging the insight that cross‑modal attention maps intrinsically encode semantic correspondences between textual tokens and visual regions, ControlMLLM++ optimizes a latent visual token modifier during inference via a task‑specific energy function to steer model attention towards user‑specified areas. To enhance optimization stability and mitigate language prompt biases, ControlMLLM++ incorporates an improved optimization strategy (Optim++) and a prompt debiasing mechanism (PromptDebias). Supporting diverse visual prompt types including bounding boxes, masks, scribbles, and points, our method demonstrates strong out‑of‑domain generalization and interpretability. The code is available at https://github.com/mrwu‑mac/ControlMLLM.
Authors:Mingrui Wu, Hang Liu, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji
Abstract:
Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini‑2.5‑Flash‑Image show emerging abilities to reason over multiple related images, existing benchmarks rarely address the challenges of multi‑image context generation, focusing mainly on text‑to‑image or single‑image editing tasks. In this work, we introduce MICON‑Bench, a comprehensive benchmark covering six tasks that evaluate cross‑image composition, contextual reasoning, and identity preservation. We further propose an MLLM‑driven Evaluation‑by‑Checkpoint framework for automatic verification of semantic and visual consistency, where multimodal large language model (MLLM) serves as a verifier. Additionally, we present Dynamic Attention Rebalancing (DAR), a training‑free, plug‑and‑play mechanism that dynamically adjusts attention during inference to enhance coherence and reduce hallucinations. Extensive experiments on various state‑of‑the‑art open‑source models demonstrate both the rigor of MICON‑Bench in exposing multi‑image reasoning challenges and the efficacy of DAR in improving generation quality and cross‑image coherence. Github: https://github.com/Angusliuuu/MICON‑Bench.
Authors:Wei Feng, Haiyong Zheng
Abstract:
We propose a template‑driven triangulation framework that embeds raster‑ or segmentation‑derived boundaries into a regular triangular grid for stable PDE discretization on image‑derived domains. Unlike constrained Delaunay triangulation (CDT), which may trigger global connectivity updates, our method retriangulates only triangles intersected by the boundary, preserves the base mesh, and supports synchronization‑free parallel execution. To ensure determinism and scalability, we classify all local boundary‑intersection configurations up to discrete equivalence and triangle symmetries, yielding a finite symbolic lookup table that maps each case to a conflict‑free retriangulation template. We prove that the resulting mesh is closed, has bounded angles, and is compatible with cotangent‑based discretizations and standard finite element methods. Experiments on elliptic and parabolic PDEs, signal interpolation, and structural metrics show fewer sliver elements, more regular triangles, and improved geometric fidelity near complex boundaries. The framework is well suited for real‑time geometric analysis and physically based simulation over image‑derived domains.
Authors:Yifeng Huang, Gia Khanh Nguyen, Minh Hoai
Abstract:
This paper presents CountEx, a discriminative visual counting framework designed to address a key limitation of existing prompt‑based methods: the inability to explicitly exclude visually similar distractors. While current approaches allow users to specify what to count via inclusion prompts, they often struggle in cluttered scenes with confusable object categories, leading to ambiguity and overcounting. CountEx enables users to express both inclusion and exclusion intent, specifying what to count and what to ignore, through multimodal prompts including natural language descriptions and optional visual exemplars. At the core of CountEx is a novel Discriminative Query Refinement module, which jointly reasons over inclusion and exclusion cues by first identifying shared visual features, then isolating exclusion‑specific patterns, and finally applying selective suppression to refine the counting query. To support systematic evaluation of fine‑grained counting methods, we introduce CoCount, a benchmark comprising 1,780 videos and 10,086 annotated frames across 97 category pairs. Experiments show that CountEx achieves substantial improvements over state‑of‑the‑art methods for counting objects from both known and novel categories. The data and code are available at https://github.com/bbvisual/CountEx.
Authors:Yuxuan Yang, Zhonghao Yan, Yi Zhang, Bo Yun, Muxi Diao, Guowei Zhao, Kongming Liang, Wenbin Li, Zhanyu Ma
Abstract:
Hepatocellular Carcinoma diagnosis relies heavily on the interpretation of gigapixel Whole Slide Images. However, current computational approaches are constrained by fixed‑resolution processing mechanisms and inefficient feature aggregation, which inevitably lead to either severe information loss or high feature redundancy. To address these challenges, we propose Hepato‑LLaVA, a specialized Multi‑modal Large Language Model designed for fine‑grained hepatocellular pathology analysis. We introduce a novel Sparse Topo‑Pack Attention mechanism that explicitly models 2D tissue topology. This mechanism effectively aggregates local diagnostic evidence into semantic summary tokens while preserving global context. Furthermore, to overcome the lack of multi‑scale data, we present HepatoPathoVQA, a clinically grounded dataset comprising 33K hierarchically structured question‑answer pairs validated by expert pathologists. Our experiments demonstrate that Hepato‑LLaVA achieves state‑of‑the‑art performance on HCC diagnosis and captioning tasks, significantly outperforming existing methods. Our code and implementation details are available at https://pris‑cv.github.io/Hepto‑LLaVA/.
Authors:Hefei Mei, Zirui Wang, Chang Xu, Jianyuan Guo, Minjing Dong
Abstract:
Large Vision‑Language Models (LVLMs) are foundational to modern multimodal applications, yet their susceptibility to adversarial attacks remains a critical concern. Prior white‑box attacks rarely generalize across tasks, and black‑box methods depend on expensive transfer, which limits efficiency. The vision encoder, standardized and often shared across LVLMs, provides a stable gray‑box pivot with strong cross‑model transfer. Building on this premise, we introduce PA‑Attack (Prototype‑Anchored Attentive Attack). PA‑Attack begins with a prototype‑anchored guidance that provides a stable attack direction towards a general and dissimilar prototype, tackling the attribute‑restricted issue and limited task generalization of vanilla attacks. Building on this, we propose a two‑stage attention enhancement mechanism: (i) leverage token‑level attention scores to concentrate perturbations on critical visual tokens, and (ii) adaptively recalibrate attention weights to track the evolving attention during the adversarial process. Extensive experiments across diverse downstream tasks and LVLM architectures show that PA‑Attack achieves an average 75.1% score reduction rate (SRR), demonstrating strong attack effectiveness, efficiency, and task generalization in LVLMs. Code is available at https://github.com/hefeimei06/PA‑Attack.
Authors:Huayu Wang, Bahaa Alattar, Cheng-Yen Yang, Hsiang-Wei Huang, Jung Heon Kim, Linda Shapiro, Nathan White, Jenq-Neng Hwang
Abstract:
Temporal stability in glottic opening localization remains challenging due to the complementary weaknesses of single‑frame detectors and foundation‑model trackers: the former lacks temporal context, while the latter suffers from memory drift. Specifically, in video laryngoscopy, rapid tissue deformation, occlusions, and visual ambiguities in emergency settings require a robust, temporally aware solution that can prevent progressive tracking errors. We propose Closed‑Loop Memory Correction (CL‑MC), a detector‑in‑the‑loop framework that supervises Segment Anything Model 2(SAM2) through confidence‑aligned state decisions and active memory rectification. High‑confidence detections trigger semantic resets that overwrite corrupted tracker memory, effectively mitigating drift accumulation with a training‑free foundation tracker in complex endoscopic scenes. On emergency intubation videos, CL‑MC achieves state‑of‑the‑art performance, significantly reducing drift and missing rate compared with the SAM2 variants and open loop based methods. Our results establish memory correction as a crucial component for reliable clinical video tracking. Our code will be available in https://github.com/huayuww/CL‑MR.
Authors:Alexandros Haliassos, Rodrigo Mira, Stavros Petridis
Abstract:
Unified Speech Recognition (USR) has emerged as a semi‑supervised framework for training a single model for audio, visual, and audiovisual speech recognition, achieving state‑of‑the‑art results on in‑distribution benchmarks. However, its reliance on autoregressive pseudo‑labelling makes training expensive, while its decoupled supervision of CTC and attention branches increases susceptibility to self‑reinforcing errors, particularly under distribution shifts involving longer sequences, noise, or unseen domains. We propose CTC‑driven teacher forcing, where greedily decoded CTC pseudo‑labels are fed into the decoder to generate attention targets in a single forward pass. Although these can be globally incoherent, in the pseudo‑labelling setting they enable efficient and effective knowledge transfer. Because CTC and CTC‑driven attention pseudo‑labels have the same length, the decoder can predict both simultaneously, benefiting from the robustness of CTC and the expressiveness of attention without costly beam search. We further propose mixed sampling to mitigate the exposure bias of the decoder relying solely on CTC inputs. The resulting method, USR 2.0, halves training time, improves robustness to out‑of‑distribution inputs, and achieves state‑of‑the‑art results on LRS3, LRS2, and WildVSR, surpassing USR and modality‑specific self‑supervised baselines.
Authors:Guoliang Gong, Man Yu
Abstract:
The image purification strategy constructs an intermediate distribution with aligned anatomical structures, which effectively corrects the spatial misalignment between real‑world ultra‑low‑dose CT and normal‑dose CT images and significantly enhances the structural preservation ability of denoising models. However, this strategy exhibits two inherent limitations. First, it suppresses noise only in the chest wall and bone regions while leaving the image background untreated. Second, it lacks a dedicated mechanism for denoising the lung parenchyma. To address these issues, we systematically redesign the original image purification strategy and propose an improved version termed IPv2. The proposed strategy introduces three core modules, namely Remove Background, Add noise, and Remove noise. These modules endow the model with denoising capability in both background and lung tissue regions during training data construction and provide a more reasonable evaluation protocol through refined label construction at the testing stage. Extensive experiments on our previously established real‑world patient lung CT dataset acquired at 2% radiation dose demonstrate that IPv2 consistently improves background suppression and lung parenchyma restoration across multiple mainstream denoising models. The code is publicly available at https://github.com/MonkeyDadLufy/Image‑Purification‑Strategy‑v2.
Authors:Hardik Shah, Erica Tevere, Deegan Atha, Marcel Kaufmann, Shehryar Khattak, Manthan Patel, Marco Hutter, Jonas Frey, Patrick Spieler
Abstract:
Autonomous navigation in complex, unstructured outdoor environments requires robots to operate over long ranges without prior maps and limited depth sensing. In such settings, relying solely on geometric frontiers for exploration is often insufficient. In such settings, the ability to reason semantically about where to go and what is safe to traverse is crucial for robust, efficient exploration. This work presents WildOS, a unified system for long‑range, open‑vocabulary object search that combines safe geometric exploration with semantic visual reasoning. WildOS builds a sparse navigation graph to maintain spatial memory, while utilizing a foundation‑model‑based vision module, ExploRFM, to score frontier nodes of the graph. ExploRFM simultaneously predicts traversability, visual frontiers, and object similarity in image space, enabling real‑time, onboard semantic navigation tasks. The resulting vision‑scored graph enables the robot to explore semantically meaningful directions while ensuring geometric safety. Furthermore, we introduce a particle‑filter‑based method for coarse localization of the open‑vocabulary target query, that estimates candidate goal positions beyond the robot's immediate depth horizon, enabling effective planning toward distant goals. Extensive closed‑loop field experiments across diverse off‑road and urban terrains demonstrate that WildOS enables robust navigation, significantly outperforming purely geometric and purely vision‑based baselines in both efficiency and autonomy. Our results highlight the potential of vision foundation models to drive open‑world robotic behaviors that are both semantically informed and geometrically grounded. Project Page: https://leggedrobotics.github.io/wildos/
Authors:Jindi Kong, Yuting He, Cong Xia, Rongjun Ge, Shuo Li
Abstract:
Clinical MRI contrast acquisition suffers from inefficient information yield, which presents as a mismatch between the risky and costly acquisition protocol and the fixed and sparse acquisition sequence. Applying world models to simulate the contrast enhancement kinetics in the human body enables continuous contrast‑free dynamics. However, the low temporal resolution in MRI acquisition restricts the training of world models, leading to a sparsely sampled dataset. Directly training a generative model to capture the kinetics leads to two limitations: (a) Due to the absence of data on missing time, the model tends to overfit to irrelevant features, leading to content distortion. (b) Due to the lack of continuous temporal supervision, the model fails to learn the continuous kinetics law over time, causing temporal discontinuities. For the first time, we propose MRI Contrast Enhancement Kinetics World model (MRI CEKWorld) with SpatioTemporal Consistency Learning (STCL). For (a), guided by the spatial law that patient‑level structures remain consistent during enhancement, we propose Latent Alignment Learning (LAL) that constructs a patient‑specific template to constrain contents to align with this template. For (b), guided by the temporal law that the kinetics follow a consistent smooth trend, we propose Latent Difference Learning (LDL) which extends the unobserved intervals by interpolation and constrains smooth variations in the latent space among interpolated sequences. Extensive experiments on two datasets show our MRI CEKWorld achieves better realistic contents and kinetics. Codes will be available at https://github.com/DD0922/MRI‑Contrast‑Enhancement‑Kinetics‑World‑Model.
Authors:Zunkai Dai, Ke Li, Jiajia Liu, Jie Yang, Yuanyuan Qiao
Abstract:
The collection and detection of video anomaly data has long been a challenging problem due to its rare occurrence and spatio‑temporal scarcity. Existing video anomaly detection (VAD) methods under perform in open‑world scenarios. Key contributing factors include limited dataset diversity, and inadequate understanding of context‑dependent anomalous semantics. To address these issues, i) we propose LAVIDA, an end‑to‑end zero‑shot video anomaly detection framework. ii) LAVIDA employs an Anomaly Exposure Sampler that transforms segmented objects into pseudo‑anomalies to enhance model adaptability to unseen anomaly categories. It further integrates a Multimodal Large Language Model (MLLM) to bolster semantic comprehension capabilities. Additionally, iii) we design a token compression approach based on reverse attention to handle the spatio‑temporal scarcity of anomalous patterns and decrease computational cost. The training process is conducted solely on pseudo anomalies without any VAD data. Evaluations across four benchmark VAD datasets demonstrate that LAVIDA achieves SOTA performance in both frame‑level and pixel‑level anomaly detection under the zero‑shot setting. Our code is available in https://github.com/VitaminCreed/LAVIDA.
Authors:Zehao Deng, An Liu, Yan Wang
Abstract:
Zero‑shot 3D Anomaly Detection is an emerging task that aims to detect anomalies in a target dataset without any target training data, which is particularly important in scenarios constrained by sample scarcity and data privacy concerns. While current methods adapt CLIP by projecting 3D point clouds into 2D representations, they face challenges. The projection inherently loses some geometric details, and the reliance on a single 2D modality provides an incomplete visual understanding, limiting their ability to detect diverse anomaly types. To address these limitations, we propose the Geometry‑Aware Prompt and Synergistic View Representation Learning (GS‑CLIP) framework, which enables the model to identify geometric anomalies through a two‑stage learning process. In stage 1, we dynamically generate text prompts embedded with 3D geometric priors. These prompts contain global shape context and local defect information distilled by our Geometric Defect Distillation Module (GDDM). In stage 2, we introduce Synergistic View Representation Learning architecture that processes rendered and depth images in parallel. A Synergistic Refinement Module (SRM) subsequently fuses the features of both streams, capitalizing on their complementary strengths. Comprehensive experimental results on four large‑scale public datasets show that GS‑CLIP achieves superior performance in detection. Code can be available at https://github.com/zhushengxinyue/GS‑CLIP.
Authors:Gang Xu, Zhiyu Zhu, Junhui Hou
Abstract:
Event cameras excel at high‑speed, low‑power, and high‑dynamic‑range scene perception. However, as they fundamentally record only relative intensity changes rather than absolute intensity, the resulting data streams suffer from a significant loss of spatial information and static texture details. In this paper, we address this limitation by leveraging the generative prior of a pre‑trained video diffusion model to reconstruct high‑fidelity video frames from sparse event data. Specifically, we first establish a baseline model by directly applying event data as a condition to synthesize videos. Then, based on the physical correlation between the event stream and video frames, we further introduce the event‑based inter‑frame residual guidance to enhance the accuracy of video frame reconstruction. Furthermore, we extend our method to video frame interpolation and prediction in a zero‑shot manner by modulating the reverse diffusion sampling process, thereby creating a unified event‑to‑frame reconstruction framework. Experimental results on real‑world and synthetic datasets demonstrate that our method significantly outperforms previous approaches both quantitatively and qualitatively. We also refer the reviewers to the video demo contained in the supplementary material for video results. The code will be publicly available at https://github.com/CS‑GangXu/UniE2F.
Authors:Kai Liu, Yanhao Zheng, Kai Wang, Shengqiong Wu, Rongjunchen Zhang, Jiebo Luo, Dimitrios Hatzinakos, Ziwei Liu, Hao Fei, Tat-Seng Chua
Abstract:
AIGC has rapidly expanded from text‑to‑image generation toward high‑quality multimodal synthesis across video and audio. Within this context, joint audio‑video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision from textual descriptions. However, compared with advanced commercial models such as Veo3, existing open‑source methods still suffer from limitations in generation quality, temporal synchrony, and alignment with human preferences. To bridge the gap, this paper presents JavisDiT++, a concise yet powerful framework for unified modeling and optimization of JAVG. First, we introduce a modality‑specific mixture‑of‑experts (MS‑MoE) design that enables cross‑modal interaction efficacy while enhancing single‑modal generation quality. Then, we propose a temporal‑aligned RoPE (TA‑RoPE) strategy to achieve explicit, frame‑level synchronization between audio and video tokens. Besides, we develop an audio‑video direct preference optimization (AV‑DPO) method to align model outputs with human preference across quality, consistency, and synchrony dimensions. Built upon Wan2.1‑1.3B‑T2V, our model achieves state‑of‑the‑art performance merely with around 1M public training entries, significantly outperforming prior approaches in both qualitative and quantitative evaluations. Comprehensive ablation studies have been conducted to validate the effectiveness of our proposed modules. All the code, model, and dataset are released at https://JavisVerse.github.io/JavisDiT2‑page.
Authors:Lunjie Zhu, Yushi Huang, Xingtong Ge, Yufei Xue, Zhening Liu, Yumeng Zhang, Zehong Lin, Jun Zhang
Abstract:
Latent diffusion models have enabled high‑quality video synthesis, yet their inference remains costly and time‑consuming. As diffusion transformers become increasingly efficient, the latency bottleneck inevitably shifts to VAE decoders. To reduce their latency while maintaining quality, we propose a universal acceleration framework for VAE decoders that preserves full alignment with the original latent distribution. Specifically, we propose (1) an independence‑aware channel pruning method to effectively mitigate severe channel redundancy, and (2) a stage‑wise dominant operator optimization strategy to address the high inference cost of the widely used causal 3D convolutions in VAE decoders. Based on these innovations, we construct a Flash‑VAED family. Moreover, we design a three‑phase dynamic distillation framework that efficiently transfers the capabilities of the original VAE decoder to Flash‑VAED. Extensive experiments on Wan and LTX‑Video VAE decoders demonstrate that our method outperforms baselines in both quality and speed, achieving approximately a 6× speedup while maintaining the reconstruction performance up to 96.9%. Notably, Flash‑VAED accelerates the end‑to‑end generation pipeline by up to 36% with negligible quality drops on VBench‑2.0.
Authors:Qi Sun, Can Wang, Jiaxiang Shang, Yingchun Liu, Jing Liao
Abstract:
Current 3D human animation methods struggle to achieve photorealism: kinematics‑based approaches lack non‑rigid dynamics (e.g., clothing dynamics), while methods that leverage video diffusion priors can synthesize non‑rigid motion but suffer from quality artifacts and identity loss. To overcome these limitations, we present Ani3DHuman, a framework that marries kinematics‑based animation with video diffusion priors. We first introduce a layered motion representation that disentangles rigid motion from residual non‑rigid motion. Rigid motion is generated by a kinematic method, which then produces a coarse rendering to guide the video diffusion model in generating video sequences that restore the residual non‑rigid motion. However, this restoration task, based on diffusion sampling, is highly challenging, as the initial renderings are out‑of‑distribution, causing standard deterministic ODE samplers to fail. Therefore, we propose a novel self‑guided stochastic sampling method, which effectively addresses the out‑of‑distribution problem by combining stochastic sampling (for photorealistic quality) with self‑guidance (for identity fidelity). These restored videos provide high‑quality supervision, enabling the optimization of the residual non‑rigid motion field. Extensive experiments demonstrate that \MethodName can generate photorealistic 3D human animation, outperforming existing methods. Code is available in https://github.com/qiisun/ani3dhuman.
Authors:Rui-Yang Ju, Kohei Yamashita, Hirotaka Kameko, Shinsuke Mori
Abstract:
Kuzushiji was one of the most popular writing styles in pre‑modern Japan and was widely used in both personal letters and official documents. However, due to its highly cursive forms and extensive glyph variations, most modern Japanese readers cannot directly interpret Kuzushiji characters. Therefore, recent research has focused on developing automated Kuzushiji character recognition methods, which have achieved satisfactory performance on relatively clean Kuzushiji document images. However, existing methods struggle to maintain recognition accuracy under seal interference (e.g., when seals overlap characters), despite the frequent occurrence of seals in pre‑modern Japanese documents. To address this challenge, we propose a three‑stage restoration‑guided Kuzushiji character recognition (RG‑KCR) framework specifically designed to mitigate seal interference. We construct datasets for evaluating Kuzushiji character detection (Stage 1) and classification (Stage 3). Experimental results show that the YOLOv12‑medium model achieves a precision of 98.0% and a recall of 93.3% on the constructed test set. We quantitatively evaluate the restoration performance of Stage 2 using PSNR and SSIM. In addition, we conduct an ablation study to demonstrate that Stage 2 improves the Top‑1 accuracy of Metom, a Vision Transformer (ViT)‑based Kuzushiji classifier employed in Stage 3, from 93.45% to 95.33%. The implementation code of this work is available at https://ruiyangju.github.io/RG‑KCR.
Authors:Qingwen Zhang, Chenhan Jiang, Xiaomeng Zhu, Yunqi Miao, Yushan Zhang, Olov Andersson, Patric Jensfelt
Abstract:
Self‑supervised feed‑forward methods for scene flow estimation offer real‑time efficiency, but their supervision from two‑frame point correspondences is unreliable and often breaks down under occlusions. Multi‑frame supervision has the potential to provide more stable guidance by incorporating motion cues from past frames, yet naive extensions of two‑frame objectives are ineffective because point correspondences vary abruptly across frames, producing inconsistent signals. In the paper, we present TeFlow, enabling multi‑frame supervision for feed‑forward models by mining temporally consistent supervision. TeFlow introduces a temporal ensembling strategy that forms reliable supervisory signals by aggregating the most temporally consistent motion cues from a candidate pool built across multiple frames. Extensive evaluations demonstrate that TeFlow establishes a new state‑of‑the‑art for self‑supervised feed‑forward methods, achieving performance gains of up to 33% on the challenging Argoverse 2 and nuScenes datasets. Our method performs on par with leading optimization‑based methods, yet speeds up 150 times. The code is open‑sourced at https://github.com/Kin‑Zhang/TeFlow along with trained model weights.
Authors:Zheng Miao, Tien-Chieh Hung
Abstract:
Accurate sex identification in fish is vital for optimizing breeding and management strategies in aquaculture, particularly for species at the risk of extinction. However, most existing methods are invasive or stressful and may cause additional mortality, posing severe risks to threatened or endangered fish populations. To address these challenges, we propose FishProtoNet, a robust, non‑invasive computer vision‑based framework for sex identification of delta smelt (Hypomesus transpacificus), an endangered fish species native to California, across its full life cycle. Unlike the traditional deep learning methods, FishProtoNet provides interpretability through learned prototype representations while improving robustness by leveraging foundation models to reduce the influence of background noise. Specifically, the FishProtoNet framework consists of three key components: fish regions of interest (ROIs) extraction using visual foundation model, feature extraction from fish ROIs and fish sex identification based on an interpretable prototype network. FishProtoNet demonstrates strong performance in delta smelt sex identification during early spawning and post‑spawning stages, achieving the accuracies of 74.40% and 81.16% and corresponding F1 scores of 74.27% and 79.43% respectively. In contrast, delta smelt sex identification at the subadult stage remains challenging for current computer vision methods, likely due to less pronounced morphological differences in immature fish. The source code of FishProtoNet is publicly available at: https://github.com/zhengmiao1/Fish_sex_identification
Authors:Duc Duy Nguyen, Tat-Jun Chin, Minh Hoai
Abstract:
We aim to learn a joint representation between inertial measurement unit (IMU) signals and 2D pose sequences extracted from video, enabling accurate cross‑modal retrieval, temporal synchronization, subject and body‑part localization, and action recognition. To this end, we introduce MoBind, a hierarchical contrastive learning framework designed to address three challenges: (1) filtering out irrelevant visual background, (2) modeling structured multi‑sensor IMU configurations, and (3) achieving fine‑grained, sub‑second temporal alignment. To isolate motion‑relevant cues, MoBind aligns IMU signals with skeletal motion sequences rather than raw pixels. We further decompose full‑body motion into local body‑part trajectories, pairing each with its corresponding IMU to enable semantically grounded multi‑sensor alignment. To capture detailed temporal correspondence, MoBind employs a hierarchical contrastive strategy that first aligns token‑level temporal segments, then fuses local (body‑part) alignment with global (body‑wide) motion aggregation. Evaluated on mRi, TotalCapture, and EgoHumans, MoBind consistently outperforms strong baselines across all four tasks, demonstrating robust fine‑grained temporal alignment while preserving coarse semantic consistency across modalities. Code is available at https://github.com/bbvisual/ MoBind.
Authors:Shannan Yan, Leqi Zheng, Keyu Lv, Jingchen Ni, Hongyang Wei, Jiajun Zhang, Guangting Wang, Jing Lyu, Chun Yuan, Fengyun Rao
Abstract:
We study the task of establishing object‑level visual correspondence across different viewpoints in videos, focusing on the challenging egocentric‑to‑exocentric and exocentric‑to‑egocentric scenarios. We propose a simple yet effective framework based on conditional binary segmentation, where an object query mask is encoded into a latent representation to guide the localization of the corresponding object in a target video. To encourage robust, view‑invariant representations, we introduce a cycle‑consistency training objective: the predicted mask in the target view is projected back to the source view to reconstruct the original query mask. This bidirectional constraint provides a strong self‑supervisory signal without requiring ground‑truth annotations and enables test‑time training (TTT) at inference. Experiments on the Ego‑Exo4D and HANDAL‑X benchmarks demonstrate the effectiveness of our optimization objective and TTT strategy, achieving state‑of‑the‑art performance. The code is available at https://github.com/shannany0606/CCMP.
Authors:Jiwoo Chung, Sangeek Hyun, MinKyu Lee, Byeongju Han, Geonho Cha, Dongyoon Wee, Youngjun Hong, Jae-Pil Heo
Abstract:
Diffusion models are a strong backbone for visual generation, but their inherently sequential denoising process leads to slow inference. Previous methods accelerate sampling by caching and reusing intermediate outputs based on feature distances between adjacent timesteps. However, existing caching strategies typically rely on raw feature differences that entangle content and noise. This design overlooks spectral evolution, where low‑frequency structure appears early and high‑frequency detail is refined later. We introduce Spectral‑Evolution‑Aware Cache (SeaCache), a training‑free cache schedule that bases reuse decisions on a spectrally aligned representation. Through theoretical and empirical analysis, we derive a Spectral‑Evolution‑Aware (SEA) filter that preserves content‑relevant components while suppressing noise. Employing SEA‑filtered input features to estimate redundancy leads to dynamic schedules that adapt to content while respecting the spectral priors underlying the diffusion model. Extensive experiments on diverse visual generative models and the baselines show that SeaCache achieves state‑of‑the‑art latency‑quality trade‑offs.
Authors:Thinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina Roitberg
Abstract:
Adapting image‑pretrained backbones to video typically relies on time‑domain adapters tuned to a single temporal scale. Our experiments show that these modules pick up static image cues and very fast flicker changes, while overlooking medium‑speed motion. Capturing dynamics across multiple time‑scales is, however, crucial for fine‑grained temporal analysis (i.e., opening vs. closing bottle).
To address this, we introduce Frame2Freq ‑‑ a family of frequency‑aware adapters that perform spectral encoding during image‑to‑video adaptation of pretrained Vision Foundation Models (VFMs), improving fine‑grained action recognition. Frame2Freq uses Fast Fourier Transform (FFT) along time and learns frequency‑band specific embeddings that adaptively highlight the most discriminative frequency ranges. Across five fine‑grained activity recognition datasets, Frame2Freq outperforms prior PEFT methods and even surpasses fully fine‑tuned models on four of them. These results provide encouraging evidence that frequency analysis methods are a powerful tool for modeling temporal dynamics in image‑to‑video transfer. Code is available at https://github.com/th‑nesh/Frame2Freq.
Authors:Kaiming Jin, Yuefan Wu, Shengqiong Wu, Bobo Li, Shuicheng Yan, Tat-Seng Chua
Abstract:
Vision‑and‑Language Scene navigation is a fundamental capability for embodied human‑AI collaboration, requiring agents to follow natural language instructions to execute coherent action sequences in complex environments. Existing approaches either rely on multiple agents, incurring high coordination and resource costs, or adopt a single‑agent paradigm, which overloads the agent with both global planning and local perception, often leading to degraded reasoning and instruction drift in long‑horizon settings. To address these issues, we introduce DACo, a planning‑grounding decoupled architecture that disentangles global deliberation from local grounding. Concretely, it employs a Global Commander for high‑level strategic planning and a Local Operative for egocentric observing and fine‑grained execution. By disentangling global reasoning from local action, DACo alleviates cognitive overload and improves long‑horizon stability. The framework further integrates dynamic subgoal planning and adaptive replanning to enable structured and resilient navigation. Extensive evaluations on R2R, REVERIE, and R4R demonstrate that DACo achieves 4.9%, 6.5%, 5.4% absolute improvements over the best‑performing baselines in zero‑shot settings, and generalizes effectively across both closed‑source (e.g., GPT‑4o) and open‑source (e.g., Qwen‑VL Series) backbones. DACo provides a principled and extensible paradigm for robust long‑horizon navigation. Project page: https://github.com/ChocoWu/DACo
Authors:Hao Lu, Onur C. Koyun, Yongxin Guo, Zhengjie Zhu, Abbas Alili, Metin Nafi Gurcan
Abstract:
Vector Quantization (VQ) underpins many modern generative frameworks such as VQ‑VAE, VQ‑GAN, and latent diffusion models. Yet, it suffers from the persistent problem of codebook collapse, where a large fraction of code vectors remains unused during training. This work provides a new theoretical explanation by identifying the nonstationary nature of encoder updates as the fundamental cause of this phenomenon. We show that as the encoder drifts, unselected code vectors fail to receive updates and gradually become inactive. To address this, we propose two new methods: Non‑Stationary Vector Quantization (NSVQ), which propagates encoder drift to non‑selected codes through a kernel‑based rule, and Transformer‑based Vector Quantization (TransVQ), which employs a lightweight mapping to adaptively transform the entire codebook while preserving convergence to the k‑means solution. Experiments on the CelebA‑HQ dataset demonstrate that both methods achieve near‑complete codebook utilization and superior reconstruction quality compared to baseline VQ variants, providing a principled and scalable foundation for future VQ‑based generative models. The code is available at: https://github.com/CAIR‑ LAB‑ WFUSM/NSVQ‑TransVQ.git
Authors:Miaowei Wang, Qingxuan Yan, Zhi Cao, Yayuan Li, Oisin Mac Aodha, Jason J. Corso, Amir Vaxman
Abstract:
Text‑guided dynamic 3D character generation has advanced rapidly, yet producing high‑quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub‑actions or incoherent motion due to fixed‑length temporal inputs and discrete frame‑wise representations that fail to capture rich motion semantics. We address these limitations by representing motion with continuous differentiable B‑spline curves, enabling more effective motion generation without modifying the capabilities of the underlying generative model. Specifically, our closed‑form, Laplacian‑regularized B‑spline solver efficiently compresses variable‑length motion sequences into compact representations with a fixed number of control points. Further, we introduce a normal‑fusion strategy for input shape adherence along with correspondence‑aware and local‑rigidity losses for motion‑restoration quality. To train our model, we collate BIMO, a new dataset containing diverse variable‑length 3D motion sequences with rich, high‑quality text annotations. Extensive evaluations show that our feed‑forward framework BiMotion generates more expressive, higher‑quality, and better prompt‑aligned motions than existing state‑of‑the‑art methods, while also achieving faster generation. Our project page is at: https://wangmiaowei.github.io/BiMotion.github.io/.
Authors:Ziheng Chen, Bernhard Schölkopf, Nicu Sebe
Abstract:
Hyperbolic spaces provide a natural geometry for representing hierarchical and tree‑structured data due to their exponential volume growth. To leverage these benefits, neural networks require intrinsic and efficient components that operate directly in hyperbolic space. In this work, we lift two core components of neural networks, Multinomial Logistic Regression (MLR) and Fully Connected (FC) layers, into hyperbolic space via Busemann functions, resulting in Busemann MLR (BMLR) and Busemann FC (BFC) layers with a unified mathematical interpretation. BMLR provides compact parameters, a point‑to‑horosphere distance interpretation, batch‑efficient computation, and a Euclidean limit, while BFC generalizes FC and activation layers with comparable complexity. Experiments on image classification, genome sequence learning, node classification, and link prediction demonstrate improvements in effectiveness and efficiency over prior hyperbolic layers. The code is available at https://github.com/GitZH‑Chen/HBNN.
Authors:Aditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu, Emad Barsoum
Abstract:
Vision‑language models (VLMs) have achieved remarkable multimodal understanding and reasoning capabilities, yet remain computationally expensive due to dense visual tokenization. Existing efficiency approaches either merge redundant visual tokens or drop them progressively in language backbone, often trading accuracy for speed. In this work, we propose DUET‑VLM, a versatile plug‑and‑play dual compression framework that consists of (a) vision‑only redundancy aware compression of vision encoder's output into information‑preserving tokens, followed by (b) layer‑wise, salient text‑guided dropping of visual tokens within the language backbone to progressively prune less informative tokens. This coordinated token management enables aggressive compression while retaining critical semantics. On LLaVA‑1.5‑7B, our approach maintains over 99% of baseline accuracy with 67% fewer tokens, and still retains >97% even at 89% reduction. With this dual‑stage compression during training, it achieves 99.7% accuracy at 67% and 97.6% at 89%, surpassing prior SoTA visual token reduction methods across multiple benchmarks. When integrated into Video‑LLaVA‑7B, it even surpasses the baseline ‑‑ achieving >100% accuracy with a substantial 53.1% token reduction and retaining 97.6% accuracy under an extreme 93.4% setting. These results highlight end‑to‑end training with DUET‑VLM, enabling robust adaptation to reduced visual (image/video) input without sacrificing accuracy, producing compact yet semantically rich representations within the same computational budget. Our code is available at https://github.com/AMD‑AGI/DUET‑VLM.
Authors:Chengwei Xia, Fan Ma, Ruijie Quan, Yunqiu Xu, Kun Zhan, Yi Yang
Abstract:
With the rapid deployment of multimodal large language models (MLLMs), disputes regarding model ownership have become increasingly frequent, raising significant concerns about intellectual property protection. In this paper, we propose a framework for generating copyright triggers for MLLMs, enabling model publishers to embed verifiable ownership information into the model. The goal is to construct trigger images that elicit ownership‑related textual responses exclusively in fine‑tuned derivatives, while remaining inert in other non‑derivative models. Our method constructs a tracking trigger image by treating the image as a learnable tensor, performing adversarial optimization with dual‑injection of ownership‑relevant semantic information. The first injection is achieved by enforcing textual consistency between the output of an auxiliary MLLM and a predefined ownership‑relevant target text; the consistency loss is backpropagated to inject this ownership‑related information into the image. The second injection is performed at the semantic‑level by minimizing the distance between the CLIP features of the image and those of the target text. Furthermore, we introduce an additional adversarial training stage involving the auxiliary model. It is specifically trained to resist generating ownership‑relevant target text, thereby enhancing robustness in heavily fine‑tuned derivative models. Extensive experiments demonstrate the effectiveness of our dual‑injection approach in tracking model lineage under various fine‑tuning and domain‑shift scenarios. Code is at https://github.com/kunzhan/AGDI
Authors:Chongyang Xu, Shen Cheng, Haipeng Li, Haoqiang Fan, Ziliang Feng, Shuaicheng Liu
Abstract:
Imitation learning for robotic manipulation has progressed from 2D image policies to 3D representations that explicitly encode geometry. Yet purely geometric policies often lack explicit part‑level semantics, which are critical for pose‑aware manipulation (e.g., distinguishing a shoe's toe from heel). In this paper, we present HeRO, a diffusion‑based policy that couples geometry and semantics via hierarchical semantic fields. HeRO employs dense semantics lifting to fuse discriminative, geometry‑sensitive features from DINOv2 with the smooth, globally coherent correspondences from Stable Diffusion, yielding dense features that are both fine‑grained and spatially consistent. These features are processed and partitioned to construct a global field and a set of local fields. A hierarchical conditioning module conditions the generative denoiser on global and local fields using permutation‑invariant network architecture, thereby avoiding order‑sensitive bias and producing a coherent control policy for pose‑aware manipulation. In various tests, HeRO establishes a new state‑of‑the‑art, improving success on Place Dual Shoes by 12.3% and averaging 6.5% gains across six challenging pose‑aware tasks. Code is available at https://github.com/Chongyang‑99/HeRO.
Authors:Haobo Lin, Tianyi Bai, Jiajun Zhang, Xuanhao Chang, Sheng Lu, Fangming Gu, Zengjie Hu, Wentao Zhang
Abstract:
Facial Expression Recognition (FER) is a fine‑grained visual understanding task where reliable predictions require reasoning over localized and meaningful facial cues. Recent vision‑‑language models (VLMs) enable natural language explanations for FER, but their reasoning is often ungrounded, producing fluent yet unverifiable rationales that are weakly tied to visual evidence and prone to hallucination, leading to poor robustness across different datasets. We propose TAG (Thinking with Action Unit Grounding), a vision‑‑language framework that explicitly constrains multimodal reasoning to be supported by facial Action Units (AUs). TAG requires intermediate reasoning steps to be grounded in AU‑related facial regions, yielding predictions accompanied by verifiable visual evidence. The model is trained via supervised fine‑tuning on AU‑grounded reasoning traces followed by reinforcement learning with an AU‑aware reward that aligns predicted regions with external AU detectors. Evaluated on RAF‑DB, FERPlus, and AffectNet, TAG consistently outperforms strong open‑source and closed‑source VLM baselines while simultaneously improving visual faithfulness. Ablation and preference studies further show that AU‑grounded rewards stabilize reasoning and mitigate hallucination, demonstrating the importance of structured grounded intermediate representations for trustworthy multimodal reasoning in FER. The code will be available at https://github.com/would1920/FER_TAG .
Authors:Yuran Dong, Hang Dai, Mang Ye
Abstract:
Multimodal editing large models have demonstrated powerful editing capabilities across diverse tasks. However, a persistent and long‑standing limitation is the decline in facial identity (ID) consistency during realistic portrait editing. Due to the human eye's high sensitivity to facial features, such inconsistency significantly hinders the practical deployment of these models. Current facial ID preservation methods struggle to achieve consistent restoration of both facial identity and edited element IP due to Cross‑source Distribution Bias and Cross‑source Feature Contamination. To address these issues, we propose EditedID, an Alignment‑Disentanglement‑Entanglement framework for robust identity‑specific facial restoration. By systematically analyzing diffusion trajectories, sampler behaviors, and attention properties, we introduce three key components: 1) Adaptive mixing strategy that aligns cross‑source latent representations throughout the diffusion process. 2) Hybrid solver that disentangles source‑specific identity attributes and details. 3) Attentional gating mechanism that selectively entangles visual elements. Extensive experiments show that EditedID achieves state‑of‑the‑art performance in preserving original facial ID and edited element IP consistency. As a training‑free and plug‑and‑play solution, it establishes a new benchmark for practical and reliable single/multi‑person facial identity restoration in open‑world settings, paving the way for the deployment of multimodal editing large models in real‑person editing scenarios. The code is available at https://github.com/NDYBSNDY/EditedID.
Authors:Haobo Lin, Tianyi Bai, Chen Chen, Jiajun Zhang, Bohan Zeng, Wentao Zhang, Binhang Yuan
Abstract:
Multimodal geometry reasoning requires models to jointly understand visual diagrams and perform structured symbolic inference, yet current vision‑‑language models struggle with complex geometric constructions due to limited training data and weak visual‑‑symbolic alignment. We propose a pipeline for synthesizing complex multimodal geometry problems from scratch and construct a dataset named GeoCode, which decouples problem generation into symbolic seed construction, grounded instantiation with verification, and code‑based diagram rendering, ensuring consistency across structure, text, reasoning, and images. Leveraging the plotting code provided in GeoCode, we further introduce code prediction as an explicit alignment objective, transforming visual understanding into a supervised structured prediction task. GeoCode exhibits substantially higher structural complexity and reasoning difficulty than existing benchmarks, while maintaining mathematical correctness through multi‑stage validation. Extensive experiments show that models trained on GeoCode achieve consistent improvements on multiple geometry benchmarks, demonstrating both the effectiveness of the dataset and the proposed alignment strategy. The code will be available at https://github.com/would1920/GeoCode.
Authors:Seungku Kim, Suhyeok Jang, Byungjun Yoon, Dongyoung Kim, John Won, Jinwoo Shin
Abstract:
Synthetic data generated by video generative models has shown promise for robot learning as a scalable pipeline, but it often suffers from inconsistent action quality due to imperfectly generated videos. Recently, vision‑language models (VLMs) have been leveraged to validate video quality, but they have limitations in distinguishing physically accurate videos and, even then, cannot directly evaluate the generated actions themselves. To tackle this issue, we introduce RoboCurate, a novel synthetic robot data generation framework that evaluates and filters the quality of annotated actions by comparing them with simulation replay. Specifically, RoboCurate replays the predicted actions in a simulator and assesses action quality by measuring the consistency of motion between the simulator rollout and the generated video. In addition, we unlock observation diversity beyond the available dataset via image‑to‑image editing and apply action‑preserving video‑to‑video transfer to further augment appearance. We observe RoboCurate's generated data yield substantial relative improvements in success rates compared to using real data only, achieving +70.1% on GR‑1 Tabletop (300 demos), +16.1% on DexMimicGen in the pre‑training setup, and +179.9% in the challenging real‑world ALLEX humanoid dexterous manipulation setting.
Authors:Weilong Yan, Haipeng Li, Hao Xu, Nianjin Ye, Yihao Ai, Shuaicheng Liu, Jingyu Hu
Abstract:
This paper introduces LaS‑Comp, a zero‑shot and category‑agnostic approach that leverages the rich geometric priors of 3D foundation models to enable 3D shape completion across diverse types of partial observations. Our contributions are threefold: First, \ourname harnesses these powerful generative priors for completion through a complementary two‑stage design: (i) an explicit replacement stage that preserves the partial observation geometry to ensure faithful completion; and (ii) an implicit refinement stage ensures seamless boundaries between the observed and synthesized regions. Second, our framework is training‑free and compatible with different 3D foundation models. Third, we introduce Omni‑Comp, a comprehensive benchmark combining real‑world and synthetic data with diverse and challenging partial patterns, enabling a more thorough and realistic evaluation. Both quantitative and qualitative experiments demonstrate that our approach outperforms previous state‑of‑the‑art approaches. Our code and data will be available at \hrefhttps://github.com/DavidYan2001/LaS‑CompLaS‑Comp.
Authors:Yufan Wang, Sokratis Makrogiannis, Chandra Kambhamettu
Abstract:
State Space Models (SSMs) have recently gained traction in remote sensing change detection (CD) for their favorable scaling properties. In this paper, we explore the potential of modern convolutional and attention‑based architectures as a competitive alternative. We propose NeXt2Former‑CD, an end‑to‑end framework that integrates a Siamese ConvNeXt encoder initialized with DINOv3 weights, a deformable attention‑based temporal fusion module, and a Mask2Former decoder. This design is intended to better tolerate residual co‑registration noise and small object‑level spatial shifts, as well as semantic ambiguity in bi‑temporal imagery. Experiments on LEVIR‑CD, WHU‑CD, and CDD datasets show that our method achieves the best results among the evaluated methods, improving over recent Mamba‑based baselines in both F1 score and IoU. Furthermore, despite a larger parameter count, our model maintains inference latency comparable to SSM‑based approaches, suggesting it is practical for high‑resolution change detection tasks.
Authors:Massoud Dehghan, Ramona Woitek, Amirreza Mahbod
Abstract:
Vision Transformers (ViTs) and their variants have become state‑of‑the‑art in many computer vision tasks and are widely used as backbones in large‑scale vision and vision‑language foundation models. While substantial research has focused on architectural improvements, the impact of patch size, a crucial initial design choice in ViTs, remains underexplored, particularly in medical domains where both two‑dimensional (2D) and three‑dimensional (3D) imaging modalities exist.
In this study, using 12 medical imaging datasets from various imaging modalities (including seven 2D and five 3D datasets), we conduct a thorough evaluation of how different patch sizes affect ViT classification performance. Using a single graphical processing unit (GPU) and a range of patch sizes (1, 2, 4, 7, 14, 28), we fine‑tune ViT models and observe consistent improvements in classification performance with smaller patch sizes (1, 2, and 4), which achieve the best results across nearly all datasets. More specifically, our results indicate improvements in balanced accuracy of up to 12.78% for 2D datasets (patch size 2 vs. 28) and up to 23.78% for 3D datasets (patch size 1 vs. 14), at the cost of increased computational expense. Moreover, by applying a straightforward ensemble strategy that fuses the predictions of the models trained with patch sizes 1, 2, and 4, we demonstrate a further boost in performance in most cases, especially for the 2D datasets. Our implementation is publicly available on GitHub: https://github.com/HealMaDe/MedViT
Authors:Xiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang, Runze Yang, Yihang Luo, Zhonghua Wu, Wei-Shi Zheng, Chen Change Loy
Abstract:
Following the rise of large foundation models, Vision‑Language‑Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision‑Language Models for general‑purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving space, we reexamine the VLA design space under a unified framework and evaluation setup. Starting from a simple VLA baseline similar to RT‑2, which is the origin of VLA, we systematically dissect design choices along three dimensions: foundational components, perception essentials, and action modelling perspectives. From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. The outcome of this exploration is a simple yet effective model, VLANeXt. It outperforms the state‑of‑the‑art methods on the LIBERO and LIBERO‑plus benchmarks and demonstrates strong performance in real‑world experiments. We release a unified and easy‑to‑use codebase to reproduce our findings, explore the design space, and develop new VLA variants on top of a shared foundation. The codebase is available at https://github.com/DravenALG/VLANeXt.
Authors:Zhan Liu, Changli Tang, Yuxin Wang, Zhiyuan Zhu, Youjun Chen, Yiwen Shao, Tianzi Wang, Lei Ke, Zengrui Jin, Chao Zhang
Abstract:
Current audio‑visual large language models (AV‑LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environments. We address this limitation by presenting JAEGER, a framework that extends AV‑LLMs to 3D space, to enable joint spatial grounding and reasoning through the integration of RGB‑D observations and multi‑channel first‑order ambisonics. A core contribution of our work is the neural intensity vector (Neural IV), a learned spatial audio representation that encodes robust directional cues to enhance direction‑of‑arrival estimation, even in adverse acoustic scenarios with overlapping sources. To facilitate large‑scale training and systematic evaluation, we propose SpatialSceneQA, a benchmark of 61k instruction‑tuning samples curated from simulated physical environments. Extensive experiments demonstrate that our approach consistently surpasses 2D‑centric baselines across diverse spatial perception and reasoning tasks, underscoring the necessity of explicit 3D modelling for advancing AI in physical environments. Our source code, pre‑trained model checkpoints, and datasets are available at https://github.com/liuzhan22/JAEGER.
Authors:Sarah Müller, Philipp Berens
Abstract:
Although deep learning models in medical imaging often achieve excellent classification performance, they can rely on shortcut learning, exploiting spurious correlations or confounding factors that are not causally related to the target task. This poses risks in clinical settings, where models must generalize across institutions, populations, and acquisition conditions. Feature disentanglement is a promising approach to mitigate shortcut learning by separating task‑relevant information from confounder‑related features in latent representations. In this study, we systematically evaluated feature disentanglement methods for mitigating shortcuts in medical imaging, including adversarial learning and latent space splitting based on dependence minimization. We assessed classification performance and disentanglement quality using latent space analyses across one artificial and two medical datasets with natural and synthetic confounders. We also examined robustness under varying levels of confounding and compared computational efficiency across methods. We found that shortcut mitigation methods improved classification performance under strong spurious correlations during training. Latent space analyses revealed differences in representation quality not captured by classification metrics, highlighting the strengths and limitations of each method. Model reliance on shortcuts depended on the degree of confounding in the training data. The best‑performing models combine data‑centric rebalancing with model‑centric disentanglement, achieving stronger and more robust shortcut mitigation than rebalancing alone while maintaining similar computational efficiency. The project code is publicly available at https://github.com/berenslab/medical‑shortcut‑mitigation.
Authors:Vatsal Agarwal, Saksham Suri, Matthew Gwilliam, Pulkit Kumar, Abhinav Shrivastava
Abstract:
Streaming video understanding requires models to robustly encode, store, and retrieve information from a continuous video stream to support accurate video question answering (VQA). Existing state‑of‑the‑art approaches rely on key‑value caching to accumulate frame‑level information over time, but use a limited number of tokens per frame, leading to the loss of fine‑grained visual details. In this work, we propose scaling the token budget to enable more granular spatiotemporal understanding and reasoning. First, we find that current methods are ill‑equipped to handle dense streams: their feature encoding causes query‑frame similarity scores to increase over time, biasing retrieval toward later frames. To address this, we introduce an adaptive selection strategy that reduces token redundancy while preserving local spatiotemporal information. We further propose a training‑free retrieval mixture‑of‑experts that leverages external models to better identify relevant frames. Our method, MemStream, achieves +8.0% on CG‑Bench, +8.5% on LVBench, and +2.4% on VideoMME (Long) over ReKV with Qwen2.5‑VL‑7B.
Authors:Evonne Ng, Siwei Zhang, Zhang Chen, Michael Zollhoefer, Alexander Richard
Abstract:
As embodied agents become central to VR, telepresence, and digital human applications, their motion must go beyond speech‑aligned gestures: agents should turn toward users, respond to their movement, and maintain natural gaze. Current methods lack this spatial awareness. We close this gap with the first real‑time, fully causal method for spatially‑aware conversational motion, deployable on a streaming VR headset. Given a user's position and dyadic audio, our approach produces full‑body motion that aligns gestures with speech while orienting the agent according to the user. Our architecture combines a causal transformer‑based VAE with interleaved latent tokens for streaming inference and a flow matching model conditioned on user trajectory and audio. To support varying gaze preferences, we introduce a gaze scoring mechanism with classifier‑free guidance to decouple learning from control: the model captures natural spatial alignment from data, while users can adjust eye contact intensity at inference time. On the Embody 3D dataset, our method achieves state‑of‑the‑art motion quality at over 300 FPS ‑‑ 3x faster than non‑causal baselines ‑‑ while capturing the subtle spatial dynamics of natural conversation. We validate our approach on a live VR system, bringing spatially‑aware conversational agents to real‑time deployment. Please see https://evonneng.github.io/sarah/ for details.
Authors:Xia Su, Ruiqi Chen, Benlin Liu, Jingwei Ma, Zonglin Di, Ranjay Krishna, Jon Froehlich
Abstract:
Vision‑Language Models (VLMs) have shown remarkable progress in Vision‑Language Navigation (VLN), offering new possibilities for navigation decision‑making that could benefit both robotic platforms and human users. However, real‑world navigation is inherently conditioned by the agent's mobility constraints. For example, a sweeping robot cannot traverse stairs, while a quadruped can. We introduce Capability‑Conditioned Navigation (CapNav), a benchmark designed to evaluate how well VLMs can navigate complex indoor spaces given an agent's specific physical and operational capabilities. CapNav defines five representative human and robot agents, each described with physical dimensions, mobility capabilities, and environmental interaction abilities. CapNav provides 45 real‑world indoor scenes, 473 navigation tasks, and 2365 QA pairs to test if VLMs can traverse indoor environments based on agent capabilities. We evaluate 13 modern VLMs and find that current VLM's navigation performance drops sharply as mobility constraints tighten, and that even state‑of‑the‑art models struggle with obstacle types that require reasoning on spatial dimensions. We conclude by discussing the implications for capability‑aware navigation and the opportunities for advancing embodied spatial reasoning in future VLMs. The benchmark is available at https://github.com/makeabilitylab/CapNav
Authors:Linxi Xie, Lisong C. Sun, Ashley Neall, Tong Wu, Shengqu Cai, Gordon Wetzstein
Abstract:
Extended reality (XR) demands generative models that respond to users' tracked real‑world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied interaction. We introduce a human‑centric video world model that is conditioned on both tracked head pose and joint‑level hand poses. For this purpose, we evaluate existing diffusion transformer conditioning strategies and propose an effective mechanism for 3D head and hand control, enabling dexterous hand‑‑object interactions. We train a bidirectional video diffusion model teacher using this strategy and distill it into a causal, interactive system that generates egocentric virtual environments. We evaluate this generated reality system with human subjects and demonstrate improved task performance as well as a significantly higher level of perceived amount of control over the performed actions compared with relevant baselines.
Authors:Minh Dinh, Stéphane Deny
Abstract:
Despite the successes of deep learning in computer vision, difficulties persist in recognizing objects that have undergone group‑symmetric transformations rarely seen during training\unicodex2013for example objects seen in unusual poses, scales, positions, or combinations thereof. Equivariant neural networks are a solution to the problem of generalizing across symmetric transformations, but require knowledge of transformations a priori. An alternative family of architectures proposes to learn equivariant operators in a latent space, from examples of symmetric transformations. Here, using simple datasets of rotated and translated noisy MNIST, we illustrate how such architectures can successfully be harnessed for out‑of‑distribution classification, thus overcoming the limitations of both traditional and equivariant networks. While conceptually enticing, we discuss challenges ahead on the path of scaling these architectures to more complex datasets. Our code is available at https://github.com/BRAIN‑Aalto/equivariant_operator.
Authors:Ziyue Liu, Davide Talon, Federico Girella, Zanxi Ruan, Mattia Mondo, Loris Bazzani, Yiming Wang, Marco Cristani
Abstract:
Sketches offer designers a concise yet expressive medium for early‑stage fashion ideation by specifying structure, silhouette, and spatial relationships, while textual descriptions complement sketches to convey material, color, and stylistic details. Effectively combining textual and visual modalities requires adherence to the sketch visual structure when leveraging the guidance of localized attributes from text. We present LOcalized Text and Sketch with multi‑level guidance (LOTS), a framework that enhances fashion image generation by combining global sketch guidance with multiple localized sketch‑text pairs. LOTS employs a Multi‑level Conditioning Stage to independently encode local features within a shared latent space while maintaining global structural coordination. Then, the Diffusion Pair Guidance stage integrates both local and global conditioning via attention‑based guidance within the diffusion model's multi‑step denoising process. To validate our method, we develop Sketchy, the first fashion dataset where multiple text‑sketch pairs are provided per image. Sketchy provides high‑quality, clean sketches with a professional look and consistent structure. To assess robustness beyond this setting, we also include an "in the wild" split with non‑expert sketches, featuring higher variability and imperfections. Experiments demonstrate that our method strengthens global structural adherence while leveraging richer localized semantic guidance, achieving improvement over state‑of‑the‑art. The dataset, platform, and code are publicly available.
Authors:Gwangtak Bae, Jaeho Shin, Seunggu Kang, Junho Kim, Ayoung Kim, Young Min Kim
Abstract:
Event cameras in motion tend to detect object boundaries or texture edges, which produce lines of brightness changes, especially in man‑made environments. While lines can constitute a robust intermediate representation that is consistently observed, the sparse nature of lines may lead to drastic deterioration with minor estimation errors. Only a few previous works, often accompanied by additional sensors, utilize lines to compensate for the severe domain discrepancies of event sensors along with unpredictable noise characteristics. We propose a method that can stably extract tracks of varying appearances of lines using a clever algorithmic process that observes multiple representations from various time slices of events, compensating for potential adversaries within the event data. We then propose geometric cost functions that can refine the 3D line maps and camera poses, eliminating projective distortions and depth ambiguities. The 3D line maps are highly compact and can be equipped with our proposed cost function, which can be adapted for any observations that can detect and extract line structures or projections of them, including 3D point cloud maps or image observations. We demonstrate that our formulation is powerful enough to exhibit a significant performance boost in event‑based mapping and pose refinement across diverse datasets, and can be flexibly applied to multimodal scenarios. Our results confirm that the proposed line‑based formulation is a robust and effective approach for the practical deployment of event‑based perceptual modules. Project page: https://gwangtak.github.io/roel/
Authors:Ziyue Wang, Linghan Cai, Chang Han Low, Haofeng Liu, Junde Wu, Jingyu Wang, Rui Wang, Lei Song, Jiang Bian, Jingjing Fu, Yueming Jin
Abstract:
3D CT analysis spans a continuum from low‑level perception to high‑level clinical understanding. Existing 3D‑oriented analysis methods adopt either isolated task‑specific modeling or task‑agnostic end‑to‑end paradigms to produce one‑hop outputs, impeding the systematic accumulation of perceptual evidence for downstream reasoning. In parallel, recent multimodal large language models (MLLMs) exhibit improved visual perception and can integrate visual and textual information effectively, yet their predominantly 2D‑oriented designs fundamentally limit their ability to perceive and analyze volumetric medical data. To bridge this gap, we propose 3DMedAgent, a unified agent that enables 2D MLLMs to perform general 3D CT analysis without 3D‑specific fine‑tuning. 3DMedAgent coordinates heterogeneous visual and textual tools through a flexible MLLM agent, progressively decomposing complex 3D analysis into tractable subtasks that transition from global to regional views, from 3D volumes to informative 2D slices, and from visual evidence to structured textual representations. Central to this design, 3DMedAgent maintains a long‑term structured memory that aggregates intermediate tool outputs and supports query‑adaptive, evidence‑driven multi‑step reasoning. We further introduce the DeepChestVQA benchmark for evaluating unified perception‑to‑understanding capabilities in 3D thoracic imaging. Experiments across over 40 tasks demonstrate that 3DMedAgent consistently outperforms general, medical, and 3D‑specific MLLMs, highlighting a scalable path toward general‑purpose 3D clinical assistants.Code and data are available at \hrefhttps://github.com/jinlab‑imvr/3DMedAgenthttps://github.com/jinlab‑imvr/3DMedAgent.
Authors:Hongsong Wang, Wenjing Yan, Qiuxia Lai, Xin Geng
Abstract:
Text‑to‑Motion (T2M) generation aims to synthesize realistic human motion sequences from natural language descriptions. While two‑stage frameworks leveraging discrete motion representations have advanced T2M research, they often neglect cross‑sequence temporal consistency, i.e., the shared temporal structures present across different instances of the same action. This leads to semantic misalignments and physically implausible motions. To address this limitation, we propose TCA‑T2M, a framework for temporal consistency‑aware T2M generation. Our approach introduces a temporal consistency‑aware spatial VQ‑VAE (TCaS‑VQ‑VAE) for cross‑sequence temporal alignment, coupled with a masked motion transformer for text‑conditioned motion generation. Additionally, a kinematic constraint block mitigates discretization artifacts to ensure physical plausibility. Experiments on HumanML3D and KIT‑ML benchmarks demonstrate that TCA‑T2M achieves state‑of‑the‑art performance, highlighting the importance of temporal consistency in robust and coherent T2M generation.
Authors:Zengtian Deng, Yimeng He, Yu Shi, Lixia Wang, Touseef Ahmad Qureshi, Xiuzhen Huang, Debiao Li
Abstract:
Radiomics and deep learning both offer powerful tools for quantitative medical imaging, but most existing fusion approaches only leverage global radiomic features and overlook the complementary value of spatially resolved radiomic parametric maps. We propose a unified framework that first selects discriminative radiomic features and then injects them into a radiomics‑enhanced nnUNet at both the global and voxel levels for pancreatic ductal adenocarcinoma (PDAC) detection. On the PANORAMA dataset, our method achieved AUC = 0.96 and AP = 0.84 in cross‑validation. On an external in‑house cohort, it achieved AUC = 0.95 and AP = 0.78, outperforming the baseline nnUNet; it also ranked second in the PANORAMA Grand Challenge. This demonstrates that handcrafted radiomics, when injected at both global and voxel levels, provide complementary signals to deep learning models for PDAC detection. Our code can be found at https://github.com/briandzt/dl‑pdac‑radiomics‑global‑n‑paramaps
Authors:Guoheng Sun, Tingting Du, Kaixi Feng, Chenxiang Luo, Xingguo Ding, Zheyu Shen, Ziyao Wang, Yexiao He, Ang Li
Abstract:
Vision‑Language‑Action (VLA) models enable instruction‑following robotic manipulation, but they are typically pretrained on 2D data and lack 3D spatial understanding. An effective approach is representation alignment, where a strong vision foundation model is used to guide a 2D VLA model. However, existing methods usually apply supervision at only a single layer, failing to fully exploit the rich information distributed across depth; meanwhile, naïve multi‑layer alignment can cause gradient interference. We introduce ROCKET, a residual‑oriented multi‑layer representation alignment framework that formulates multi‑layer alignment as aligning one residual stream to another. Concretely, ROCKET employs a shared projector to align multiple layers of the VLA backbone with multiple layers of a powerful 3D vision foundation model via a layer‑invariant mapping, which reduces gradient conflicts. We provide both theoretical justification and empirical analyses showing that a shared projector is sufficient and outperforms prior designs, and further propose a Matryoshka‑style sparse activation scheme for the shared projector to balance multiple alignment losses. Our experiments show that, combined with a training‑free layer selection strategy, ROCKET requires only about 4% of the compute budget while achieving 98.5% state‑of‑the‑art success rate on LIBERO. We further demonstrate the superior performance of ROCKET across LIBERO‑Plus and RoboTwin, as well as multiple VLA models. The code and model weights can be found at https://github.com/CASE‑Lab‑UMD/ROCKET‑VLA.
Authors:Athanasios Angelakis
Abstract:
Vision Transformers rely on positional embeddings and class tokens encoding fixed spatial priors. While effective for natural images, these priors may be suboptimal when spatial layout is weakly informative, a frequent condition in medical imaging. We introduce ZACH‑ViT (Zero‑token Adaptive Compact Hierarchical Vision Transformer), a compact Vision Transformer that removes positional embeddings and the [CLS] token, achieving permutation‑invariant patch processing via global average pooling. Zero‑token denotes removal of the dedicated aggregation token and positional encodings. Patch tokens remain unchanged. Adaptive residual projections preserve training stability under strict parameter constraints. We evaluate ZACH‑ViT across seven MedMNIST datasets under a strict few‑shot protocol (50 samples/class, fixed hyperparameters, five seeds). Results reveal regime‑dependent behavior: ZACH‑ViT (0.25M parameters, trained from scratch) achieves strongest advantage on BloodMNIST and remains competitive on PathMNIST, while relative advantage decreases on datasets with stronger anatomical priors (OCTMNIST, OrganAMNIST), consistent with our hypothesis. Component and pooling ablations show positional support becomes mildly beneficial as spatial structure increases, whereas reintroducing a [CLS] token is consistently unfavorable. These findings support that architectural alignment with data structure can outweigh universal benchmark dominance. Despite minimal size and no pretraining, ZACH‑ViT achieves competitive performance under data‑scarce conditions, relevant for compact medical imaging and low‑resource settings. Code: https://github.com/Bluesman79/ZACH‑ViT
Authors:Amirhosein Javadi, Chi-Shiang Gau, Konstantinos D. Polyzos, Tara Javidi
Abstract:
Diffusion‑based approaches have recently demonstrated strong performance for single‑image novel view synthesis by conditioning generative models on geometry inferred from monocular depth estimation. However, in practice, the quality and consistency of the synthesized views are fundamentally limited by the reliability of the underlying depth estimates, which are often fragile under low‑texture, adverse weather, and occlusion‑heavy real‑world conditions. In this work, we show that incorporating sparse multimodal range measurements provides a simple yet effective way to overcome these limitations. We introduce a multimodal depth reconstruction framework that leverages extremely sparse range sensing data, such as automotive radar or LiDAR, to produce dense depth maps that serve as robust geometric conditioning for diffusion‑based novel view synthesis. Our approach models depth in an angular domain using a localized Gaussian Process formulation, enabling computationally efficient inference while explicitly quantifying uncertainty in regions with limited observations. The reconstructed depth and uncertainty are used as a drop‑in replacement for monocular depth estimators in existing diffusion‑based rendering pipelines, without modifying the generative model itself. Experiments on real‑world multimodal driving scenes demonstrate that replacing vision‑only depth with our sparse range‑based reconstruction substantially improves both geometric consistency and visual quality in single‑image novel‑view video generation. These results highlight the importance of reliable geometric priors for diffusion‑based view synthesis and demonstrate the practical benefits of multimodal sensing even at extreme levels of sparsity. Code is publicly available at: https://github.com/importAmir/MultiModalNVS
Authors:Junkai Liu, Ling Shao, Le Zhang
Abstract:
Self‑supervised learning (SSL) and diffusion models have advanced representation learning and image synthesis, but in 3D medical imaging they are still largely used separately for analysis and synthesis, respectively. Unifying them is appealing but difficult, because multi‑source data exhibit pronounced style shifts while downstream tasks rely primarily on anatomy, causing anatomical content and acquisition style to become entangled. In this paper, we propose MeDUET, a 3D Medical image Disentangled UnifiEd PreTraining framework in the variational autoencoder latent space. Our central idea is to treat unified pretraining under heterogeneous multi‑center data as a factor identifiability problem, where content should consistently capture anatomy and style should consistently capture appearance. MeDUET addresses this problem through three components. Token demixing provides controllable supervision for factor separation, Mixed Factor Token Distillation reduces factor leakage under mixed regions, and Swap‑invariance Quadruplet Contrast promotes factor‑wise invariance and discriminability. With these learned factors, MeDUET transfers effectively to both synthesis and analysis, yielding higher fidelity, faster convergence, and better controllability for synthesis, while achieving competitive or superior domain generalization and label efficiency on diverse medical benchmarks. Overall, MeDUET shows that multi‑source heterogeneity can serve as useful supervision, with disentanglement providing an effective interface for unifying 3D medical image synthesis and analysis. Our code is available at https://github.com/JK‑Liu7/MeDUET.
Authors:Adrian Catalin Lutu, Eduard Poesina, Radu Tudor Ionescu
Abstract:
Query performance prediction (QPP) is an important and actively studied information retrieval task, having various applications, such as query reformulation, query expansion, and retrieval system selection, among many others. The task has been primarily studied in the context of text and image retrieval, whereas QPP for content‑based video retrieval (CBVR) remains largely underexplored. To this end, we propose the first benchmark for video query performance prediction (VQPP), comprising two text‑to‑video retrieval datasets and two CBVR systems, respectively. VQPP contains a total of 56K text queries and 51K videos, and comes with official training, validation and test splits, fostering direct comparisons and reproducible results. We explore multiple pre‑retrieval and post‑retrieval performance predictors, creating a representative benchmark for future exploration of QPP in the video domain. Our results show that pre‑retrieval predictors obtain competitive performance, enabling applications before performing the retrieval step. We also demonstrate the applicability of VQPP by employing the best performing pre‑retrieval predictor as reward model for training a large language model (LLM) on the query reformulation task via direct preference optimization (DPO). We release our benchmark and code at https://github.com/AdrianLutu/VQPP.
Authors:Jose Sosa, Danila Rukhovich, Anis Kacem, Djamila Aouada
Abstract:
Recent advances in Vision Language Models (VLMs) and Vision Foundation Models (VFMs) have opened new opportunities for zero‑shot text‑guided segmentation of remote sensing imagery. However, most existing approaches still rely on additional trainable components, limiting their generalisation and practical applicability. In this work, we investigate to what extent text‑based remote sensing segmentation can be achieved without additional training, by relying solely on existing foundation models. We propose a simple yet effective approach that integrates contrastive and generative VLMs with the Segment Anything Model (SAM), enabling a fully training‑free or lightweight LoRA‑tuned pipeline. Our contrastive approach employs CLIP as mask selector for SAM's grid‑based proposals, achieving state‑of‑the‑art open‑vocabulary semantic segmentation (OVSS) in a completely zero‑shot setting. In parallel, our generative approach enables reasoning and referring segmentation by generating click prompts for SAM using GPT‑5 in a zero‑shot setting and a LoRA‑tuned Qwen‑VL model, with the latter yielding the best results. Extensive experiments across 19 remote sensing benchmarks, including open‑vocabulary, referring, and reasoning‑based tasks, demonstrate the strong capabilities of our approach. Code will be released at https://github.com/josesosajs/trainfree‑rs‑segmentation.
Authors:Balamurugan Thambiraja, Omid Taheri, Radek Danecek, Giorgio Becherini, Gerard Pons-Moll, Justus Thies
Abstract:
Hands play a central role in daily life, yet modeling natural hand motions remains underexplored. Existing methods that tackle text‑to‑hand‑motion generation or hand animation captioning rely on studio‑captured datasets with limited actions and contexts, making them costly to scale to "in‑the‑wild" settings. Further, contemporary models and their training schemes struggle to capture animation fidelity with text‑motion alignment. To address this, we (1) introduce '3D Hands in the Wild' (3D‑HIW), a dataset of 32K 3D hand‑motion sequences and aligned text, and (2) propose CLUTCH, an LLM‑based hand animation system with two critical innovations: (a) SHIFT, a novel VQ‑VAE architecture to tokenize hand motion, and (b) a geometric refinement stage to finetune the LLM. To build 3D‑HIW, we propose a data annotation pipeline that combines vision‑language models (VLMs) and state‑of‑the‑art 3D hand trackers, and apply it to a large corpus of egocentric action videos covering a wide range of scenarios. To fully capture motion in‑the‑wild, CLUTCH employs SHIFT, a part‑modality decomposed VQ‑VAE, which improves generalization and reconstruction fidelity. Finally, to improve animation quality, we introduce a geometric refinement stage, where CLUTCH is co‑supervised with a reconstruction loss applied directly to decoded hand motion parameters. Experiments demonstrate state‑of‑the‑art performance on text‑to‑motion and motion‑to‑text tasks, establishing the first benchmark for scalable in‑the‑wild hand motion modelling. Code, data and models will be released.
Authors:Ziyuan Liu, Shizhao Sun, Danqing Huang, Yingdong Shi, Meisheng Zhang, Ji Li, Jingsong Yu, Jiang Bian
Abstract:
Graphic design generation demands a delicate balance between high visual fidelity and fine‑grained structural editability. However, existing approaches typically bifurcate into either non‑editable raster image synthesis or abstract layout generation devoid of visual content. Recent combinations of these two approaches attempt to bridge this gap but often suffer from rigid composition schemas and unresolvable visual dissonances (e.g., text‑background conflicts) due to their inexpressive representation and open‑loop nature. To address these challenges, we propose DesignAsCode, a novel framework that reimagines graphic design as a programmatic synthesis task using HTML/CSS. Specifically, we introduce a Plan‑Implement‑Reflect pipeline, incorporating a Semantic Planner to construct dynamic, variable‑depth element hierarchies and a Visual‑Aware Reflection mechanism that iteratively optimizes the code to rectify rendering artifacts. Extensive experiments demonstrate that DesignAsCode significantly outperforms state‑of‑the‑art baselines in both structural validity and aesthetic quality. Furthermore, our code‑native representation unlocks advanced capabilities, including automatic layout retargeting, complex document generation (e.g., resumes), and CSS‑based animation. Our project page is available at https://liuziyuan1109.github.io/design‑as‑code/.
Authors:Irene Iele, Giulia Romoli, Daniele Molino, Elena Mulero Ayllón, Filippo Ruffini, Paolo Soda, Matteo Tortora
Abstract:
Short‑term forecasting of vegetation dynamics is a key enabler for data‑driven decision support in precision agriculture. Normalized Difference Vegetation Index (NDVI) forecasting from satellite observations, however, remains challenging due to sparse and irregular sampling caused by cloud masking, as well as the heterogeneous climatic conditions under which crops evolve. In this work, we propose a probabilistic forecasting framework for field‑level NDVI prediction under sparse, irregular clear‑sky acquisitions. The architecture separates the encoding of historical NDVI and meteorological observations from future exogenous covariates, fusing both representations for multi‑step quantile prediction. To address irregular revisit patterns and horizon‑dependent uncertainty, we introduce a temporal‑distance weighted quantile loss that aligns the training objective with the effective forecasting horizon. In addition, we incorporate cumulative and extreme‑weather feature engineering to capture delayed meteorological effects relevant to vegetation response. Experiments on European satellite data show that the proposed approach outperforms statistical, deep learning, and time‑series baselines on both pointwise and probabilistic evaluation metrics. Ablation studies confirm that target history is the primary driver of performance, with meteorological covariates providing additional gains in the full multimodal setting. The code is available at https://github.com/arco‑group/ndvi‑forecasting.
Authors:Tyler Bonnen, Jitendra Malik, Angjoo Kanazawa
Abstract:
Humans can infer the three‑dimensional structure of objects from two‑dimensional visual inputs. Modeling this ability has been a longstanding goal for the science and engineering of visual intelligence, yet decades of computational methods have fallen short of human performance. Here we develop a modeling framework that predicts human 3D shape inferences for arbitrary objects, directly from experimental stimuli. We achieve this with a novel class of neural networks trained using a visual‑spatial objective over naturalistic sensory data; given a set of images taken from different locations within a natural scene, these models learn to predict spatial information related to these images, such as camera location and visual depth, without relying on any object‑related inductive biases. Notably, these visual‑spatial signals are analogous to sensory cues readily available to humans. We design a zero‑shot evaluation approach to determine the performance of these 'multi‑view' models on a well established 3D perception task, then compare model and human behavior. Our modeling framework is the first to match human accuracy on 3D shape inferences, even without task‑specific training or fine‑tuning. Remarkably, independent readouts of model responses predict fine‑grained measures of human behavior, including error patterns and reaction times, revealing a natural correspondence between model dynamics and human perception. Taken together, our findings indicate that human‑level 3D perception can emerge from a simple, scalable learning objective over naturalistic visual‑spatial data. Code, images, and human data needed to reproduce all analyses can be found at https://tzler.github.io/human_multiview/
Authors:Xiaohan Zhao, Zhaoyi Li, Yaxin Luo, Jiacheng Cui, Zhiqiang Shen
Abstract:
Black‑box adversarial attacks on Large Vision‑Language Models (LVLMs) are challenging due to missing gradients and complex multimodal boundaries. While prior state‑of‑the‑art transfer‑based approaches like M‑Attack perform well using local crop‑level matching between source and target images, we find this induces high‑variance, nearly orthogonal gradients across iterations, violating coherent local alignment and destabilizing optimization. We attribute this to (i) ViT translation sensitivity that yields spike‑like gradients and (ii) structural asymmetry between source and target crops. We reformulate local matching as an asymmetric expectation over source transformations and target semantics, and build a gradient‑denoising upgrade to M‑Attack. On the source side, Multi‑Crop Alignment (MCA) averages gradients from multiple independently sampled local views per iteration to reduce variance. On the target side, Auxiliary Target Alignment (ATA) replaces aggressive target augmentation with a small auxiliary set from a semantically correlated distribution, producing a smoother, lower‑variance target manifold. We further reinterpret momentum as Patch Momentum, replaying historical crop gradients; combined with a refined patch‑size ensemble (PE+), this strengthens transferable directions. Together these modules form M‑Attack‑V2, a simple, modular enhancement over M‑Attack that substantially improves transfer‑based black‑box attacks on frontier LVLMs: boosting success rates on Claude‑4.0 from 8% to 30%, Gemini‑2.5‑Pro from 83% to 97%, and GPT‑5 from 98% to 100%, outperforming prior black‑box LVLM attacks. Code and data are publicly available at: https://github.com/vila‑lab/M‑Attack‑V2.
Authors:Yichen Lu, Siwei Nie, Minlong Lu, Xudong Yang, Xiaobo Zhang, Peng Zhang
Abstract:
Image Copy Detection (ICD) aims to identify manipulated content between image pairs through robust feature representation learning. While self‑supervised learning (SSL) has advanced ICD systems, existing view‑level contrastive methods struggle with sophisticated edits due to insufficient fine‑grained correspondence learning. We address this limitation by exploiting the inherent geometric traceability in edited content through two key innovations. First, we propose PixTrace ‑ a pixel coordinate tracking module that maintains explicit spatial mappings across editing transformations. Second, we introduce CopyNCE, a geometrically‑guided contrastive loss that regularizes patch affinity using overlap ratios derived from PixTrace's verified mappings. Our method bridges pixel‑level traceability with patch‑level similarity learning, suppressing supervision noise in SSL training. Extensive experiments demonstrate not only state‑of‑the‑art performance (88.7% uAP / 83.9% RP90 for matcher, 72.6% uAP / 68.4% RP90 for descriptor on DISC21 dataset) but also better interpretability over existing methods.
Authors:Xuan-Bac Nguyen, Hoang-Quan Nguyen, Sankalp Pandey, Tim Faltermeier, Nicholas Borys, Hugh Churchill, Khoa Luu
Abstract:
Characterizing two‑dimensional quantum materials from optical microscopy images is challenging due to the subtle layer‑dependent contrast, limited labeled data, and significant variation across laboratories and imaging setups. Existing vision models struggle in this domain since they lack physical priors and cannot generalize to new materials or hardware conditions. This work presents a new physics‑aware multimodal framework that addresses these limitations from both the data and model perspectives. We first present Synthia, a physics‑based synthetic data generator that simulates realistic optical responses of quantum material flakes under thin‑film interference. Synthia produces diverse and high‑quality samples, helping reduce the dependence on expert manual annotation. We introduce QMat‑Instruct, the first large‑scale instruction dataset for quantum materials, comprising multimodal, physics‑informed question‑answer pairs designed to teach Multimodal Large Language Models (MLLMs) to understand the appearance and thickness of flakes. Then, we propose Physics‑Aware Instruction Tuning (QuPAINT), a multimodal architecture that incorporates a Physics‑Informed Attention module to fuse visual embeddings with optical priors, enabling more robust and discriminative flake representations. Finally, we establish QF‑Bench, a comprehensive benchmark spanning multiple materials, substrates, and imaging settings, offering standardized protocols for fair and reproducible evaluation.
Authors:Jiwei Shan, Zeyu Cai, Cheng-Tai Hsieh, Yirui Li, Hao Liu, Lijun Han, Hesheng Wang, Shing Shin Cheng
Abstract:
Reconstructing deformable surgical scenes from endoscopic videos is challenging and clinically important. Recent state‑of‑the‑art methods based on implicit neural representations or 3D Gaussian splatting have made notable progress. However, most are designed for deformable scenes with fixed endoscope viewpoints and rely on stereo depth priors or accurate structure‑from‑motion for initialization and optimization, limiting their ability to handle monocular sequences with large camera motion in real clinical settings. To address this, we propose Local‑EndoGS, a high‑quality 4D reconstruction framework for monocular endoscopic sequences with arbitrary camera motion. Local‑EndoGS introduces a progressive, window‑based global representation that allocates local deformable scene models to each observed window, enabling scalability to long sequences with substantial motion. To overcome unreliable initialization without stereo depth or accurate structure‑from‑motion, we design a coarse‑to‑fine strategy integrating multi‑view geometry, cross‑window information, and monocular depth priors, providing a robust foundation for optimization. We further incorporate long‑range 2D pixel trajectory constraints and physical motion priors to improve deformation plausibility. Experiments on three public endoscopic datasets with deformable scenes and varying camera motions show that Local‑EndoGS consistently outperforms state‑of‑the‑art methods in appearance quality and geometry. Ablation studies validate the effectiveness of our key designs. Code will be released upon acceptance at: https://github.com/IRMVLab/Local‑EndoGS.
Authors:Lorenzo Caselli, Marco Mistretta, Simone Magistri, Andrew D. Bagdanov
Abstract:
Generalized Category Discovery (GCD) aims to identify novel categories in unlabeled data while leveraging a small labeled subset of known classes. Training a parametric classifier solely on image features often leads to overfitting to old classes, and recent multimodal approaches improve performance by incorporating textual information. However, they treat modalities independently and incur high computational cost. We propose SpectralGCD, an efficient and effective multimodal approach to GCD that uses CLIP cross‑modal image‑concept similarities as a unified cross‑modal representation. Each image is expressed as a mixture over semantic concepts from a large task‑agnostic dictionary, which anchors learning to explicit semantics and reduces reliance on spurious visual cues. To maintain the semantic quality of representations learned by an efficient student, we introduce Spectral Filtering which exploits a cross‑modal covariance matrix over the softmaxed similarities measured by a strong teacher model to automatically retain only relevant concepts from the dictionary. Forward and reverse knowledge distillation from the same teacher ensures that the cross‑modal representations of the student remain both semantically sufficient and well‑aligned. Across six benchmarks, SpectralGCD delivers accuracy comparable to or significantly superior to state‑of‑the‑art methods at a fraction of the computational cost. The code is publicly available at: https://github.com/miccunifi/SpectralGCD.
Authors:Antoine Legouhy, Cosimo Campo, Ross Callaghan, Hojjat Azadbakht, Hui Zhang
Abstract:
In this work we present Polaffini, a robust and versatile framework for anatomically grounded registration. Medical image registration is dominated by intensity‑based registration methods that rely on surrogate measures of alignment quality. In contrast, feature‑based approaches that operate by identifying explicit anatomical correspondences, while more desirable in theory, have largely fallen out of favor due to the challenges of reliably extracting features. However, such challenges are now significantly overcome thanks to recent advances in deep learning, which provide pre‑trained segmentation models capable of instantly delivering reliable, fine‑grained anatomical delineations. We aim to demonstrate that these advances can be leveraged to create new anatomically‑grounded image registration algorithms. To this end, we propose Polaffini, which obtains, from these segmented regions, anatomically grounded feature points with 1‑to‑1 correspondence in a particularly simple way: extracting their centroids. These enable efficient global and local affine matching via closed‑form solutions. Those are used to produce an overall transformation ranging from affine to polyaffine with tunable smoothness. Polyaffine transformations can have many more degrees of freedom than affine ones allowing for finer alignment, and their embedding in the log‑Euclidean framework ensures diffeomorphic properties. Polaffini has applications both for standalone registration and as pre‑alignment for subsequent non‑linear registration, and we evaluate it against popular intensity‑based registration techniques. Results demonstrate that Polaffini outperforms competing methods in terms of structural alignment and provides improved initialisation for downstream non‑linear registration. Polaffini is fast, robust, and accurate, making it particularly well‑suited for integration into medical image processing pipelines.
Authors:Yiming Xu, Yi Yang, Hao Cheng, Monika Sester
Abstract:
Accurate motion forecasting is critical for autonomous driving, yet most predictors rely on multi‑object tracking (MOT) with identity association, assuming that objects are correctly and continuously tracked. When tracking fails due to, e.g., occlusion, identity switches, or missed detections, prediction quality degrades and safety risks increase. We present HiMAP, a tracking‑free, trajectory prediction framework that remains reliable under MOT failures. HiMAP converts past detections into spatiotemporally invariant historical occupancy maps and introduces a historical query module that conditions on the current agent state to iteratively retrieve agent‑specific history from unlabeled occupancy representations. The retrieved history is summarized by a temporal map embedding and, together with the final query and map context, drives a DETR‑style decoder to produce multi‑modal future trajectories. This design lifts identity reliance, supports streaming inference via reusable encodings, and serves as a robust fallback when tracking is unavailable. On Argoverse~2, HiMAP achieves performance comparable to tracking‑based methods while operating without IDs, and it substantially outperforms strong baselines in the no‑tracking setting, yielding relative gains of 11% in FDE, 12% in ADE, and a 4% reduction in MR over a fine‑tuned QCNet. Beyond aggregate metrics, HiMAP delivers stable forecasts for all agents simultaneously without waiting for tracking to recover, highlighting its practical value for safety‑critical autonomy. The code is available under: https://github.com/XuYiMing83/HiMAP.
Authors:Ye Zhu, Kaleb S. Newman, Johannes F. Lutzeyer, Adriana Romero-Soriano, Michal Drozdzal, Olga Russakovsky
Abstract:
Despite high semantic alignment, modern text‑to‑image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I diversity through a geometric lens. Unlike most existing methods that rely primarily on entropy‑based guidance to increase sample dissimilarity, we introduce Geometry‑Aware Spherical Sampling (GASS) to enhance diversity by explicitly controlling both prompt‑dependent and prompt‑independent sources of variation. Specifically, we decompose the diversity measure in CLIP embeddings using two orthogonal directions: the text embedding, which captures semantic variation related to the prompt, and an identified orthogonal direction that captures prompt‑independent variation (e.g., backgrounds). Based on this decomposition, GASS increases the geometric projection spread of generated image embeddings along both axes and guides the T2I sampling process via expanded predictions along the generation trajectory. Our experiments on different frozen T2I backbones (U‑Net and DiT, diffusion and flow) and benchmarks demonstrate the effectiveness of disentangled diversity enhancement with minimal impact on image fidelity and semantic alignment.
Authors:Yahong Wang, Juncheng Wu, Zhangkai Ni, Chengmei Yang, Yihang Liu, Longzhen Yang, Yuyin Zhou, Ying Wen, Lianghua He
Abstract:
Multimodal large language models (MLLMs) incur substantial inference cost due to the processing of hundreds of visual tokens per image. Although token pruning has proven effective for accelerating inference, determining when and where to prune remains largely heuristic. Existing approaches typically rely on static, empirically selected layers, which limit interpretability and transferability across models. In this work, we introduce a matrix‑entropy perspective and identify an "Entropy Collapse Layer" (ECL), where the information content of visual representations exhibits a sharp and consistent drop, which provides a principled criterion for selecting the pruning stage. Building on this observation, we propose EntropyPrune, a novel matrix‑entropy‑guided token pruning framework that quantifies the information value of individual visual tokens and prunes redundant ones without relying on attention maps. Moreover, to enable efficient computation, we exploit the spectral equivalence of dual Gram matrices, reducing the complexity of entropy computation and yielding up to a 64x theoretical speedup. Extensive experiments on diverse multimodal benchmarks demonstrate that EntropyPrune consistently outperforms state‑of‑the‑art pruning methods in both accuracy and efficiency. On LLaVA‑1.5‑7B, our method achieves a 68.2% reduction in FLOPs while preserving 96.0% of the original performance. Furthermore, EntropyPrune generalizes effectively to high‑resolution and video‑based models, highlighting the strong robustness and scalability in practical MLLM acceleration. The code will be publicly available at https://github.com/YahongWang1/EntropyPrune.
Authors:Hiromichi Kamata, Samuel Arthur Munro, Fuminori Homma
Abstract:
Interactive 3D Gaussian Splatting (3DGS) segmentation is essential for real‑time editing of pre‑reconstructed assets in film and game production. However, existing methods rely on predefined camera viewpoints, ground‑truth labels, or costly retraining, making them impractical for low‑latency use. We propose B^3‑Seg (Beta‑Bernoulli Bayesian Segmentation for 3DGS), a fast and theoretically grounded method for open‑vocabulary 3DGS segmentation under camera‑free and training‑free conditions. Our approach reformulates segmentation as sequential Beta‑Bernoulli Bayesian updates and actively selects the next view via analytic Expected Information Gain (EIG). This Bayesian formulation guarantees the adaptive monotonicity and submodularity of EIG, which produces a greedy (1‑1/e) approximation to the optimal view sampling policy. Experiments on multiple datasets show that B^3‑Seg achieves competitive results to high‑cost supervised methods while operating end‑to‑end segmentation within a few seconds. The results demonstrate that B^3‑Seg enables practical, interactive 3DGS segmentation with provable information efficiency.
Authors:Dayeon Lee, Donghyeong Kim, Chaewon Park, Sungmin Woo, Sangyoun Lee
Abstract:
Weakly supervised video anomaly detection aims to detect anomalies and identify abnormal categories with only video‑level labels. We propose CPL‑VAD, a dual‑branch framework with cross pseudo labeling. The binary anomaly detection branch focuses on snippet‑level anomaly localization, while the category classification branch leverages vision‑language alignment to recognize abnormal event categories. By exchanging pseudo labels, the two branches transfer complementary strengths, combining temporal precision with semantic discrimination. Experiments on XD‑Violence and UCF‑Crime demonstrate that CPL‑VAD achieves state‑of‑the‑art performance in both anomaly detection and abnormal category classification.
Authors:Peize Li, Zeyu Zhang, Hao Tang
Abstract:
Single‑image 3D generation with part‑level structure remains challenging: learned priors struggle to cover the long tail of part geometries and maintain multi‑view consistency, and existing systems provide limited support for precise, localized edits. We present PartRAG, a retrieval‑augmented framework that integrates an external part database with a diffusion transformer to couple generation with an editable representation. To overcome the first challenge, we introduce a Hierarchical Contrastive Retrieval module that aligns dense image patches with 3D part latents at both part and object granularity, retrieving from a curated bank of 1,236 part‑annotated assets to inject diverse, physically plausible exemplars into denoising. To overcome the second challenge, we add a masked, part‑level editor that operates in a shared canonical space, enabling swaps, attribute refinements, and compositional updates without regenerating the whole object while preserving non‑target parts and multi‑view consistency. PartRAG achieves competitive results on Objaverse, ShapeNet, and ABO‑reducing Chamfer Distance from 0.1726 to 0.1528 and raising F‑Score from 0.7472 to 0.844 on Objaverse‑with inference of 38s and interactive edits in 5‑8s. Qualitatively, PartRAG produces sharper part boundaries, better thin‑structure fidelity, and robust behavior on articulated objects. Code: https://github.com/AIGeeksGroup/PartRAG. Website: https://aigeeksgroup.github.io/PartRAG.
Authors:Zeyu Ren, Xiang Li, Yiran Wang, Zeyu Zhang, Hao Tang
Abstract:
Stereo depth estimation is fundamental to underwater robotic perception, yet suffers from severe domain shifts caused by wavelength‑dependent light attenuation, scattering, and refraction. Recent approaches leverage monocular foundation models with GRU‑based iterative refinement for underwater adaptation; however, the sequential gating and local convolutional kernels in GRUs necessitate multiple iterations for long‑range disparity propagation, limiting performance in large‑disparity and textureless underwater regions. In this paper, we propose StereoAdapter‑2, which replaces the conventional ConvGRU updater with a novel ConvSS2D operator based on selective state space models. The proposed operator employs a four‑directional scanning strategy that naturally aligns with epipolar geometry while capturing vertical structural consistency, enabling efficient long‑range spatial propagation within a single update step at linear computational complexity. Furthermore, we construct UW‑StereoDepth‑80K, a large‑scale synthetic underwater stereo dataset featuring diverse baselines, attenuation coefficients, and scattering parameters through a two‑stage generative pipeline combining semantic‑aware style transfer and geometry‑consistent novel view synthesis. Combined with dynamic LoRA adaptation inherited from StereoAdapter, our framework achieves state‑of‑the‑art zero‑shot performance on underwater benchmarks with 17% improvement on TartanAir‑UW and 7.2% improvment on SQUID, with real‑world validation on the BlueROV2 platform demonstrates the robustness of our approach. Code: https://github.com/AIGeeksGroup/StereoAdapter‑2. Website: https://aigeeksgroup.github.io/StereoAdapter‑2.
Authors:Mehrshad Taji, Arad Mahdinezhad Kashani, Iman Ahmadi, AmirHossein Jadidi, Saina Kashani, Babak Khalaj
Abstract:
Task planning for robotic manipulation with large language models (LLMs) is an emerging area. Prior approaches rely on specialized models, fine tuning, or prompt tuning, and often operate in an open loop manner without robust environmental feedback, making them fragile in dynamic settings. MALLVI presents a Multi Agent Large Language and Vision framework that enables closed‑loop feedback driven robotic manipulation. Given a natural language instruction and an image of the environment, MALLVI generates executable atomic actions for a robot manipulator. After action execution, a Vision Language Model (VLM) evaluates environmental feedback and decides whether to repeat the process or proceed to the next step. Rather than using a single model, MALLVI coordinates specialized agents, Decomposer, Localizer, Thinker, and Reflector, to manage perception, localization, reasoning, and high level planning. An optional Descriptor agent provides visual memory of the initial state. The Reflector supports targeted error detection and recovery by reactivating only relevant agents, avoiding full replanning. Experiments in simulation and real‑world settings show that iterative closed loop multi agent coordination improves generalization and increases success rates in zero shot manipulation tasks. Code available at https://github.com/iman1234ahmadi/MALLVI .
Authors:Namitha Padmanabhan, Matthew Gwilliam, Abhinav Shrivastava
Abstract:
Implicit Neural Representations (INRs) have recently demonstrated impressive performance for video compression. However, since a separate INR must be overfit for each video, scaling to high‑resolution videos while maintaining encoding efficiency remains a significant challenge. Hypernetwork‑based approaches predict INR weights (hyponetworks) for unseen videos at high speeds, but with low quality, large compressed size, and prohibitive memory needs at higher resolutions. We address these fundamental limitations through three key contributions: (1) an approach that decomposes the weight prediction task spatially and temporally, by breaking short video segments into patch tubelets, to reduce the pretraining memory overhead by 20×; (2) a residual‑based storage scheme that captures only differences between consecutive segment representations, significantly reducing bitstream size; and (3) a temporal coherence regularization framework that encourages changes in the weight space to be correlated with video content. Our proposed method, TeCoNeRV, achieves substantial improvements of 2.47dB and 5.35dB PSNR over the baseline at 480p and 720p on UVG, with 36% lower bitrates and 1.5‑3× faster encoding speeds. With our low memory usage, we are the first hypernetwork approach to demonstrate results at 480p, 720p and 1080p on UVG, HEVC and MCL‑JCV. Our project page is available at https://namithap10.github.io/teconerv/ .
Authors:Yingyuan Yang, Tian Lan, Yifei Gao, Yimeng Lu, Wenjun He, Meng Wang, Chenghao Liu, Chen Zhang
Abstract:
Time‑series anomaly detection (TSAD) requires identifying both immediate Point Anomalies and long‑range Context Anomalies. However, existing foundation models face a fundamental trade‑off: 1D temporal models provide fine‑grained pointwise localization but lack a global contextual perspective, while 2D vision‑based models capture global patterns but suffer from information bottlenecks due to a lack of temporal alignment and coarse‑grained pointwise detection. To resolve this dilemma, we propose VETime, the first TSAD framework that unifies temporal and visual modalities through fine‑grained visual‑temporal alignment and dynamic fusion. VETime introduces a Reversible Image Conversion and a Patch‑Level Temporal Alignment module to establish a shared visual‑temporal timeline, preserving discriminative details while maintaining temporal sensitivity. Furthermore, we design an Anomaly Window Contrastive Learning mechanism and a Task‑Adaptive Multi‑Modal Fusion to adaptively integrate the complementary perceptual strengths of both modalities. Extensive experiments demonstrate that VETime significantly outperforms state‑of‑the‑art models in zero‑shot scenarios, achieving superior localization precision with lower computational overhead than current vision‑based approaches. Code available at: https://github.com/yyyangcoder/VETime.
Authors:Qi You, Yitai Cheng, Zichao Zeng, James Haworth
Abstract:
Street‑view image attribute classification is a vital downstream task of image classification, enabling applications such as autonomous driving, urban analytics, and high‑definition map construction. It remains computationally demanding whether training from scratch, initialising from pre‑trained weights, or fine‑tuning large models. Although pre‑trained vision‑language models such as CLIP offer rich image representations, existing adaptation or fine‑tuning methods often rely on their global image embeddings, limiting their ability to capture fine‑grained, localised attributes essential in complex, cluttered street scenes. To address this, we propose CLIP‑MHAdapter, a variant of the current lightweight CLIP adaptation paradigm that appends a bottleneck MLP equipped with multi‑head self‑attention operating on patch tokens to model inter‑patch dependencies. With approximately 1.4 million trainable parameters, CLIP‑MHAdapter achieves superior or competitive accuracy across eight attribute classification tasks on the Global StreetScapes dataset, attaining new state‑of‑the‑art results while maintaining low computational cost. The code is available at https://github.com/SpaceTimeLab/CLIP‑MHAdapter.
Authors:Kaiting Liu, Hazel Doughty
Abstract:
Video recognition models are typically trained on fixed taxonomies which are often too coarse, collapsing distinctions in object, manner or outcome under a single label. As tasks and definitions evolve, such models cannot accommodate emerging distinctions and collecting new annotations and retraining to accommodate such changes is costly. To address these challenges, we introduce category splitting, a new task where an existing classifier is edited to refine a coarse category into finer subcategories, while preserving accuracy elsewhere. We propose a zero‑shot editing method that leverages the latent compositional structure of video classifiers to expose fine‑grained distinctions without additional data. We further show that low‑shot fine‑tuning, while simple, is highly effective and benefits from our zero‑shot initialization. Experiments on our new video benchmarks for category splitting demonstrate that our method substantially outperforms vision‑language baselines, improving accuracy on the newly split categories without sacrificing performance on the rest. Project page: https://kaitingliu.github.io/Category‑Splitting/.
Authors:Yihao Lu, Wanru Cheng, Zeyu Zhang, Hao Tang
Abstract:
Long‑horizon multimodal agents depend on external memory; however, similarity‑based retrieval often surfaces stale, low‑credibility, or conflicting items, which can trigger overconfident errors. We propose Multimodal Memory Agent (MMA), which assigns each retrieved memory item a dynamic reliability score by combining source credibility, temporal decay, and conflict‑aware network consensus, and uses this signal to reweight evidence and abstain when support is insufficient. We also introduce MMA‑Bench, a programmatically generated benchmark for belief dynamics with controlled speaker reliability and structured text‑vision contradictions. Using this framework, we uncover the "Visual Placebo Effect", revealing how RAG‑based agents inherit latent visual biases from foundation models. On FEVER, MMA matches baseline accuracy while reducing variance by 35.2% and improving selective utility; on LoCoMo, a safety‑oriented configuration improves actionable accuracy and reduces wrong answers; on MMA‑Bench, MMA reaches 41.18% Type‑B accuracy in Vision mode, while the baseline collapses to 0.0% under the same protocol. Code: https://github.com/AIGeeksGroup/MMA.
Authors:Tiou Wang, Zhuoqian Yang, Markus Flierl, Mathieu Salzmann, Sabine Süsstrunk
Abstract:
We propose the Subtractive Modulative Network (SMN), a novel, parameter‑efficient Implicit Neural Representation (INR) architecture inspired by classical subtractive synthesis. The SMN is designed as a principled signal processing pipeline, featuring a learnable periodic activation layer (Oscillator) that generates a multi‑frequency basis, and a series of modulative mask modules (Filters) that actively generate high‑order harmonics. We provide both theoretical analysis and empirical validation for our design. Our SMN achieves a PSNR of 40+ dB on two image datasets, comparing favorably against state‑of‑the‑art methods in terms of both reconstruction accuracy and parameter efficiency. Furthermore, consistent advantage is observed on the challenging 3D NeRF novel view synthesis task. Supplementary materials are available at https://inrainbws.github.io/smn/.
Authors:David Smerkous, Zian Wang, Behzad Najafian
Abstract:
Self‑supervised pretraining has transformed computer vision by enabling data‑efficient fine‑tuning, yet high‑resolution training typically requires server‑scale infrastructure, limiting in‑domain foundation model development for many research laboratories. Masked Autoencoders (MAE) reduce computation by encoding only visible tokens, but combining MAE with hierarchical downsampling architectures remains structurally challenging due to dense grid priors and mask‑aware design compromises. We introduce AFFMAE, a masking‑friendly hierarchical pretraining framework built on adaptive, off‑grid token merging. By discarding masked tokens and performing dynamic merging exclusively over visible tokens, AFFMAE removes dense‑grid assumptions while preserving hierarchical scalability. We developed numerically stable mixed‑precision Flash‑style cluster attention kernels, and mitigate sparse‑stage representation collapse via deep supervision. On high‑resolution electron microscopy segmentation, AFFMAE matches ViT‑MAE performance at equal parameter count while reducing FLOPs by up to 7x, halving memory usage, and achieving faster training on a single RTX 5090. Code available at https://github.com/najafian‑lab/affmae.
Authors:J. Dhar, M. K. Pandey, D. Chakladar, M. Haghighat, A. Alavi, S. Mistry, N. Zaidi
Abstract:
Multimodal fusion frameworks, which integrate diverse medical imaging modalities (e.g., MRI, CT), have shown great potential in applications such as skin cancer detection, dementia diagnosis, and brain tumor prediction. However, existing multimodal fusion methods face significant challenges. First, they often rely on computationally expensive models, limiting their applicability in low‑resource environments. Second, they often employ cascaded attention modules, which potentially increase risk of information loss during inter‑module transitions and hinder their capacity to effectively capture robust shared representations across modalities. This restricts their generalization in multi‑disease analysis tasks. To address these limitations, we propose a Hybrid Parallel‑Fusion Cascaded Attention Network (HyPCA‑Net), composed of two core novel blocks: (a) a computationally efficient residual adaptive learning attention block for capturing refined modality‑specific representations, and (b) a dual‑view cascaded attention block aimed at learning robust shared representations across diverse modalities. Extensive experiments on ten publicly available datasets exhibit that HyPCA‑Net significantly outperforms existing leading methods, with improvements of up to 5.2% in performance and reductions of up to 73.1% in computational cost. Code: https://github.com/misti1203/HyPCA‑Net.
Authors:Tianwei Lin, Zhongwei Qiu, Wenqiao Zhang, Jiang Liu, Yihan Xie, Mingjian Gao, Zhenxuan Fan, Zhaocheng Li, Sijing Li, Zhongle Xie, Peng LU, Yueting Zhuang, Ling Zhang, Beng Chin Ooi, Yingda Xia
Abstract:
Computed Tomography (CT) is one of the most widely used and diagnostically information‑dense imaging modalities, covering critical organs such as the heart, lungs, liver, and colon. Clinical interpretation relies on both slice‑driven local features (e.g., sub‑centimeter nodules, lesion boundaries) and volume‑driven spatial representations (e.g., tumor infiltration, inter‑organ anatomical relations). However, existing Large Vision‑Language Models (LVLMs) remain fragmented in CT slice versus volumetric understanding: slice‑driven LVLMs show strong generalization but lack cross‑slice spatial consistency, while volume‑driven LVLMs explicitly capture volumetric semantics but suffer from coarse granularity and poor compatibility with slice inputs. The absence of a unified modeling paradigm constitutes a major bottleneck for the clinical translation of medical LVLMs. We present OmniCT, a powerful unified slice‑volume LVLM for CT scenarios, which makes three contributions: (i) Spatial Consistency Enhancement (SCE): volumetric slice composition combined with tri‑axial positional embedding that introduces volumetric consistency, and an MoE hybrid projection enables efficient slice‑volume adaptation; (ii) Organ‑level Semantic Enhancement (OSE): segmentation and ROI localization explicitly align anatomical regions, emphasizing lesion‑ and organ‑level semantics; (iii) MedEval‑CT: the largest slice‑volume CT dataset and hybrid benchmark integrates comprehensive metrics for unified evaluation. OmniCT consistently outperforms existing methods with a substantial margin across diverse clinical tasks and satisfies both micro‑level detail sensitivity and macro‑level spatial reasoning. More importantly, it establishes a new paradigm for cross‑modal medical imaging understanding. Our project is available at https://github.com/ZJU4HealthCare/OmniCT.
Authors:Idil Bilge Altun, Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban
Abstract:
Discrete image tokenization is a key bottleneck for scalable visual generation: a tokenizer must remain compact for efficient latent‑space priors while preserving semantic structure and using discrete capacity effectively. Existing quantizers face a trade‑off: vector‑quantized tokenizers learn flexible geometries but often suffer from biased straight‑through optimization, codebook under‑utilization, and representation collapse at large vocabularies. Structured scalar or implicit tokenizers ensure stable, near‑complete utilization by design, yet rely on fixed discretization geometries that may allocate capacity inefficiently under heterogeneous latent statistics.
We introduce Learnable Geometric Quantization (LGQ), a discrete image tokenizer that learns discretization geometry end‑to‑end. LGQ replaces hard nearest‑neighbor lookup with temperature‑controlled soft assignments, enabling fully differentiable training while recovering hard assignments at inference. The assignments correspond to posterior responsibilities of an isotropic Gaussian mixture and minimize a variational free‑energy objective, provably converging to nearest‑neighbor quantization in the low‑temperature limit. LGQ combines a token‑level peakedness regularizer with a global usage regularizer to encourage confident yet balanced code utilization without imposing rigid grids.
Under a controlled VQGAN‑style backbone on ImageNet across multiple vocabulary sizes, LGQ achieves stable optimization and balanced utilization. At 16K codebook size, LGQ improves rFID by 11.88% over FSQ while using 49.96% fewer active codes, and improves rFID by 6.06% over SimVQ with 49.45% lower effective representation rate, achieving comparable fidelity with substantially fewer active entries. Our GitHub repository is available at: https://github.com/KurbanIntelligenceLab/LGQ
Authors:Juampablo E. Heras Rivera, Dickson T. Chen, Tianyi Ren, Daniel K. Low, Asma Ben Abacha, Alberto Santamaria-Pang, Mehmet Kurt
Abstract:
Recent advances in radiology report generation (RRG) have been driven by large paired image‑text datasets; however, progress in neuro‑oncology has been limited due to a lack of open paired image‑report datasets. Here, we introduce BTReport, an open‑source framework for brain tumor RRG that constructs natural language radiology reports using deterministically extracted imaging features. Unlike existing approaches that rely on large general‑purpose or fine‑tuned vision‑language models for both image interpretation and report composition, BTReport performs deterministic feature extraction for image analysis and uses large language models only for syntactic structuring and narrative formatting. By separating RRG into a deterministic feature extraction step and a report generation step, the generated reports are completely interpretable and less prone to hallucinations. We show that the features used for report generation are predictive of key clinical outcomes, including survival and IDH mutation status, and reports generated by BTReport are more closely aligned with reference clinical reports than existing baselines for RRG. Finally, we introduce BTReport‑BraTS, a companion dataset that augments BraTS imaging with synthetically generated radiology reports produced with BTReport. Code for this project can be found at https://github.com/KurtLabUW/BTReport.
Authors:Xitong Yang, Devansh Kukreja, Don Pinkus, Anushka Sagar, Taosha Fan, Jinhyung Park, Soyong Shin, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli, Jitendra Malik, Piotr Dollar, Kris Kitani
Abstract:
We introduce SAM 3D Body (3DB), a promptable model for single‑image full‑body 3D human mesh recovery (HMR) that demonstrates state‑of‑the‑art performance, with strong generalization and consistent accuracy in diverse in‑the‑wild conditions. 3DB estimates the human pose of the body, feet, and hands. It is the first model to use a new parametric mesh representation, Momentum Human Rig (MHR), which decouples skeletal structure and surface shape. 3DB employs an encoder‑decoder architecture and supports auxiliary prompts, including 2D keypoints and masks, enabling user‑guided inference similar to the SAM family of models. We derive high‑quality annotations from a multi‑stage annotation pipeline that uses various combinations of manual keypoint annotation, differentiable optimization, multi‑view geometry, and dense keypoint detection. Our data engine efficiently selects and processes data to ensure data diversity, collecting unusual poses and rare imaging conditions. We present a new evaluation dataset organized by pose and appearance categories, enabling nuanced analysis of model behavior. Our experiments demonstrate superior generalization and substantial improvements over prior methods in both qualitative user preference studies and traditional quantitative analysis. Both 3DB and MHR are open‑source.
Authors:Gregory Cohen, Alexandre Marcireau
Abstract:
Neuromorphic engineering has a data problem. Despite the meteoric rise in the number of neuromorphic datasets published over the past ten years, the conclusion of a significant portion of neuromorphic research papers still states that there is a need for yet more data and even larger datasets. Whilst this need is driven in part by the sheer volume of data required by modern deep learning approaches, it is also fuelled by the current state of the available neuromorphic datasets and the difficulties in finding them, understanding their purpose, and determining the nature of their underlying task. This is further compounded by practical difficulties in downloading and using these datasets. This review starts by capturing a snapshot of the existing neuromorphic datasets, covering over 423 datasets, and then explores the nature of their tasks and the underlying structure of the presented data. Analysing these datasets shows the difficulties arising from their size, the lack of standardisation, and difficulties in accessing the actual data. This paper also highlights the growth in the size of individual datasets and the complexities involved in working with the data. However, a more important concern is the rise of synthetic datasets, created by either simulation or video‑to‑events methods. This review explores the benefits of simulated data for testing existing algorithms and applications, highlighting the potential pitfalls for exploring new applications of neuromorphic technologies. This review also introduces the concepts of meta‑datasets, created from existing datasets, as a way of both reducing the need for more data, and to remove potential bias arising from defining both the dataset and the task.
Authors:Yiwen Wang, Jiahao Qin
Abstract:
We address the problem of cross‑domain image registration, where paired images exhibit coupled geometric misalignment and domain‑specific appearance shift. We formalize this as a factorization problem: decomposing each image into a domain‑invariant scene representation and a global appearance statistic, such that registration reduces to recombining the scene structure of the moving image with the appearance of the fixed image via Adaptive Instance Normalization (AdaIN). This factorization eliminates the need for explicit deformation field estimation. To exploit temporal coherence in sequential acquisitions, we introduce a position‑encoded cross‑frame attention mechanism that fuses learnable and sinusoidal position embeddings with multi‑head attention over a sliding window of neighboring frames, enriching the scene representation with inter‑frame context. We instantiate this framework as GPEReg‑Net and evaluate on two benchmarks: FIRE‑Reg‑256 (retinal fundus, semi‑rigid) and HPatches‑Reg‑256 (synthetic textured patches, affine). GPEReg‑Net achieves state‑of‑the‑art performance on both benchmarks (FIRE: SSIM = 0.928, PSNR = 33.47 dB; HPatches: SSIM = 0.450, PSNR = 21.01 dB), surpassing all baselines, including deformation‑based methods, while running 1.87x faster than SAS‑Net. Code: https://github.com/JiahaoQin/GPEReg‑Net.
Authors:Christian Schlarmann, Matthias Hein
Abstract:
Generative large vision‑language models (LVLMs) have recently achieved impressive performance gains, and their user base is growing rapidly. However, the security of LVLMs, in particular in a long‑context multi‑turn setting, is largely underexplored. In this paper, we consider the realistic scenario in which an attacker uploads a manipulated image to the web/social media. A benign user downloads this image and uses it as input to the LVLM. Our novel stealthy Visual Memory Injection (VMI) attack is designed such that on normal prompts the LVLM exhibits nominal behavior, but once the user gives a triggering prompt, the LVLM outputs a specific prescribed target message to manipulate the user, e.g. for adversarial marketing or political persuasion. Compared to previous work that focused on single‑turn attacks, VMI is effective even after a long multi‑turn conversation with the user. We demonstrate our attack on several recent open‑weight LVLMs. This article thereby shows that large‑scale manipulation of users is feasible with perturbed images in multi‑turn conversation settings, calling for better robustness of LVLMs against these attacks. We release the source code at https://github.com/chs20/visual‑memory‑injection
Authors:Sen Ye, Mengde Xu, Shuyang Gu, Di He, Liwei Wang, Han Hu
Abstract:
Current research in multimodal models faces a key challenge where enhancing generative capabilities often comes at the expense of understanding, and vice versa. We analyzed this trade‑off and identify the primary cause might be the potential conflict between generation and understanding, which creates a competitive dynamic within the model. To address this, we propose the Reason‑Reflect‑Refine (R3) framework. This innovative algorithm re‑frames the single‑step generation task into a multi‑step process of "generate‑understand‑regenerate". By explicitly leveraging the model's understanding capability during generation, we successfully mitigate the optimization dilemma, achieved stronger generation results and improved understanding ability which are related to the generation process. This offers valuable insights for designing next‑generation unified multimodal models. Code is available at https://github.com/sen‑ye/R3.
Authors:Abhiram Shenoi, Philipp Lindenberger, Paul-Edouard Sarlin, Marc Pollefeys
Abstract:
This paper introduces RaCo, a lightweight neural network designed to learn robust and versatile keypoints suitable for a variety of 3D computer vision tasks. The model integrates three key components: the repeatable keypoint detector, a differentiable ranker to maximize matches with a limited number of keypoints, and a covariance estimator to quantify spatial uncertainty in metric scale. Trained on perspective image crops only, RaCo operates without the need for covisible image pairs. It achieves strong rotational robustness through extensive data augmentation, even without the use of computationally expensive equivariant network architectures. The method is evaluated on several challenging datasets, where it demonstrates state‑of‑the‑art performance in keypoint repeatability and two‑view matching, particularly under large in‑plane rotations. Ultimately, RaCo provides an effective and simple strategy to independently estimate keypoint ranking and metric covariance without additional labels, detecting interpretable and repeatable interest points. The code is available at https://github.com/cvg/RaCo.
Authors:Qiangong Zhou, Nagasaka Tomohiro
Abstract:
This work considers merging two independent models, TTS and A2F, into a unified model to enable internal feature transfer, thereby improving the consistency between audio and facial expressions generated from text. We also discuss the extension of the emotion control mechanism from TTS to the joint model. This work does not aim to showcase generation quality; instead, from a system design perspective, it validates the feasibility of reusing intermediate representations from TTS for joint modeling of speech and facial expressions, and provides engineering practice references for subsequent speech expression co‑design. The project code has been open source at: https://github.com/GoldenFishes/UniTAF
Authors:Aman Verma, Seshan Srirangarajan, Sumantra Dutta Roy
Abstract:
Quantifying biometric characteristics within hand gestures involve derivation of fitness scores from a gesture and identity aware feature space. However, evaluating the quality of these scores remains an open question. Existing biometric capacity estimation literature relies upon error rates. But these rates do not indicate goodness of scores. Thus, in this manuscript we present an exhaustive set of evaluation measures. We firstly identify ranking order and relevance of output scores as the primary basis for evaluation. In particular, we consider both rank deviation as well as rewards for: (i) higher scores of high ranked gestures and (ii) lower scores of low ranked gestures. We also compensate for correspondence between trends of output and ground truth scores. Finally, we account for disentanglement between identity features of gestures as a discounting factor. Integrating these elements with adequate weighting, we formulate advanced acceptance score as a holistic evaluation measure. To assess effectivity of the proposed we perform in‑depth experimentation over three datasets with five state‑of‑the‑art (SOTA) models. Results show that the optimal score selected with our measure is more appropriate than existing other measures. Also, our proposed measure depicts correlation with existing measures. This further validates its reliability. We have made our \hrefhttps://github.com/AmanVerma2307/MeasureSuitecode public.
Authors:Joy Dhar, Nayyar Zaidi, Maryam Haghighat
Abstract:
Multimodal Fusion Learning (MFL), leveraging disparate data from various imaging modalities (e.g., MRI, CT, SPECT), has shown great potential for addressing medical problems such as skin cancer and brain tumor prediction. However, existing MFL methods face three key limitations: a) they often specialize in specific modalities, and overlook effective shared complementary information across diverse modalities, hence limiting their generalizability for multi‑disease analysis; b) they rely on computationally expensive models, restricting their applicability in resource‑limited settings; and c) they lack robustness against adversarial attacks, compromising reliability in medical AI applications. To address these limitations, we propose a novel Multi‑Attention Integration Learning (MAIL) network, incorporating two key components: a) an efficient residual learning attention block for capturing refined modality‑specific multi‑scale patterns and b) an efficient multimodal cross‑attention module for learning enriched complementary shared representations across diverse modalities. Furthermore, to ensure adversarial robustness, we extend MAIL network to design Robust‑MAIL by incorporating random projection filters and modulated attention noise. Extensive evaluations on 20 public datasets show that both MAIL and Robust‑MAIL outperform existing methods, achieving performance gains of up to 9.34% while reducing computational costs by up to 78.3%. These results highlight the superiority of our approaches, ensuring more reliable predictions than top competitors. Code: https://github.com/misti1203/MAIL‑Robust‑MAIL.
Authors:Siwei Wen, Zhangcheng Wang, Xingjian Zhang, Lei Huang, Wenjun Wu
Abstract:
Online video understanding requires models to perform continuous perception and long‑range reasoning within potentially infinite visual streams. Its fundamental challenge lies in the conflict between the unbounded nature of streaming media input and the limited context window of Multimodal Large Language Models (MLLMs). Current methods primarily rely on passive processing, which often face a trade‑off between maintaining long‑range context and capturing the fine‑grained details necessary for complex tasks. To address this, we introduce EventMemAgent, an active online video agent framework based on a hierarchical memory module. Our framework employs a dual‑layer strategy for online videos: short‑term memory detects event boundaries and utilizes event‑granular reservoir sampling to process streaming video frames within a fixed‑length buffer dynamically; long‑term memory structuredly archives past observations on an event‑by‑event basis. Furthermore, we integrate a multi‑granular perception toolkit for active, iterative evidence capture and employ Agentic Reinforcement Learning (Agentic RL) to end‑to‑end internalize reasoning and tool‑use strategies into the agent's intrinsic capabilities. Experiments show that EventMemAgent achieves competitive results on online video benchmarks. The code will be released here: https://github.com/lingcco/EventMemAgent.
Authors:Muhammad J. Alahmadi, Peng Gao, Feiyi Wang, Dongkuan Xu
Abstract:
Dataset distillation compresses the original data into compact synthetic datasets, reducing training time and storage while retaining model performance, enabling deployment under limited resources. Although recent decoupling‑based distillation methods enable dataset distillation at large scale, they continue to face an efficiency gap: optimization‑based decoupling methods achieve higher accuracy but demand intensive computation, whereas optimization‑free decoupling methods are efficient but sacrifice accuracy. To overcome this trade‑off, we propose Exploration‑‑Exploitation Distillation (E^2D), a simple, practical method that minimizes redundant computation through an efficient pipeline that begins with full‑image initialization to preserve semantic integrity and feature diversity. It then uses a two‑phase optimization strategy: an exploration phase that performs uniform updates and identifies high‑loss regions, and an exploitation phase that focuses updates on these regions to accelerate convergence. We evaluate E^2D on large‑scale benchmarks, surpassing the state‑of‑the‑art on ImageNet‑1K while being 18× faster, and on ImageNet‑21K, our method substantially improves accuracy while remaining 4.3× faster. These results demonstrate that targeted, redundancy‑reducing updates, rather than brute‑force optimization, bridge the gap between accuracy and efficiency in large‑scale dataset distillation. Code is available at https://github.com/ncsu‑dk‑lab/E2D.
Authors:Tianyu Xiong, Skylar Wurster, Han-Wei Shen
Abstract:
Implicit Neural Representations (INRs) have emerged as promising surrogates for large 3D scientific simulations due to their ability to continuously model spatial and conditional fields, yet they face a critical fidelity‑speed dilemma: deep MLPs suffer from high inference cost, while efficient embedding‑based models lack sufficient expressiveness. To resolve this, we propose the Decoupled Representation Refinement (DRR) architectural paradigm. DRR leverages a deep refiner network, alongside non‑parametric transformations, in a one‑time offline process to encode rich representations into a compact and efficient embedding structure. This approach decouples slow neural networks with high representational capacity from the fast inference path. We introduce DRR‑Net, a simple network that validates this paradigm, and a novel data augmentation strategy, Variational Pairs (VP) for improving INRs under complex tasks like high‑dimensional surrogate modeling. Experiments on several ensemble simulation datasets demonstrate that our approach achieves state‑of‑the‑art fidelity, while being up to 27× faster at inference than high‑fidelity baselines and remaining competitive with the fastest models. The DRR paradigm offers an effective strategy for building powerful and practical neural field surrogates and INRs in broader applications, with a minimal compromise between speed and quality.
Authors:Shiyu Xuan, Dongkai Wang, Zechao Li, Jinhui Tang
Abstract:
Zero‑shot Human‑object interaction (HOI) detection aims to locate humans and objects in images and recognize their interactions. While advances in open‑vocabulary object detection provide promising solutions for object localization, interaction recognition (IR) remains challenging due to the combinatorial diversity of interactions. Existing methods, including two‑stage methods, tightly couple IR with a specific detector and rely on coarse‑grained vision‑language model (VLM) features, which limit generalization to unseen interactions. In this work, we propose a decoupled framework that separates object detection from IR and leverages multi‑modal large language models (MLLMs) for zero‑shot IR. We introduce a deterministic generation method that formulates IR as a visual question answering task and enforces deterministic outputs, enabling training‑free zero‑shot IR. To further enhance performance and efficiency by fine‑tuning the model, we design a spatial‑aware pooling module that integrates appearance and pairwise spatial cues, and a one‑pass deterministic matching method that predicts all candidate interactions in a single forward pass. Extensive experiments on HICO‑DET and V‑COCO demonstrate that our method achieves superior zero‑shot performance, strong cross‑dataset generalization, and the flexibility to integrate with any object detectors without retraining. The codes are publicly available at https://github.com/SY‑Xuan/DA‑HOI.
Authors:Abdul Joseph Fofanah, Lian Wen, Alpha Alimamy Kamara, Zhongyi Zhang, David Chen, Albert Patrick Sankoh
Abstract:
Accurate polyp segmentation in colonoscopy is essential for cancer prevention but remains challenging due to: (1) high morphological variability (from flat to protruding lesions), (2) strong visual similarity to normal structures such as folds and vessels, and (3) the need for robust multi‑scale detection. Existing deep learning approaches suffer from unidirectional processing, weak multi‑scale fusion, and the absence of anatomical constraints, often leading to false positives (over‑segmentation of normal structures) and false negatives (missed subtle flat lesions). We propose GRAFNet, a biologically inspired architecture that emulates the hierarchical organisation of the human visual system. GRAFNet integrates three key modules: (1) a Guided Asymmetric Attention Module (GAAM) that mimics orientation‑tuned cortical neurones to emphasise polyp boundaries, (2) a MultiScale Retinal Module (MSRM) that replicates retinal ganglion cell pathways for parallel multi‑feature analysis, and (3) a Guided Cortical Attention Feedback Module (GCAFM) that applies predictive coding for iterative refinement. These are unified in a Polyp Encoder‑Decoder Module (PEDM) that enforces spatial‑semantic consistency via resolution‑adaptive feedback. Extensive experiments on five public benchmarks (Kvasir‑SEG, CVC‑300, CVC‑ColonDB, CVC‑Clinic, and PolypGen) demonstrate consistent state‑of‑the‑art performance, with 3‑8% Dice improvements and 10‑20% higher generalisation over leading methods, while offering interpretable decision pathways. This work establishes a paradigm in which neural computation principles bridge the gap between AI accuracy and clinically trustworthy reasoning. Code is available at https://github.com/afofanah/GRAFNet.
Authors:Yehonathan Litman, Shikun Liu, Dario Seyb, Nicholas Milef, Yang Zhou, Carl Marshall, Shubham Tulsiani, Caleb Leak
Abstract:
High‑fidelity generative video editing has seen significant quality improvements by leveraging pre‑trained video foundation models. However, their computational cost is a major bottleneck, as they are often designed to inefficiently process the full video context regardless of the inpainting mask's size, even for sparse, localized edits. In this paper, we introduce EditCtrl, an efficient video inpainting control framework that focuses computation only where it is needed. Our approach features a novel local video context module that operates solely on masked tokens, yielding a computational cost proportional to the edit size. This local‑first generation is then guided by a lightweight temporal global context embedder that ensures video‑wide context consistency with minimal overhead. Not only is EditCtrl 10 times more compute efficient than state‑of‑the‑art generative editing methods, it even improves editing quality compared to methods designed with full‑attention. Finally, we showcase how EditCtrl unlocks new capabilities, including multi‑region editing with text prompts and autoregressive content propagation.
Authors:Zun Wang, Han Lin, Jaehong Yoon, Jaemin Cho, Yue Zhang, Mohit Bansal
Abstract:
Maintaining spatial world consistency over long horizons remains a central challenge for camera‑controllable video generation. Existing memory‑based approaches often condition generation on globally reconstructed 3D scenes by rendering anchor videos from the reconstructed geometry in the history. However, reconstructing a global 3D scene from multiple views inevitably introduces cross‑view misalignment, as pose and depth estimation errors cause the same surfaces to be reconstructed at slightly different 3D locations across views. When fused, these inconsistencies accumulate into noisy geometry that contaminates the conditioning signals and degrades generation quality. We introduce AnchorWeave, a memory‑augmented video generation framework that replaces a single misaligned global memory with multiple clean local geometric memories and learns to reconcile their cross‑view inconsistencies. To this end, AnchorWeave performs coverage‑driven local memory retrieval aligned with the target trajectory and integrates the selected local memories through a multi‑anchor weaving controller during generation. Extensive experiments demonstrate that AnchorWeave significantly improves long‑term scene consistency while maintaining strong visual quality, with ablation and analysis studies further validating the effectiveness of local geometric conditioning, multi‑anchor control, and coverage‑driven retrieval.
Authors:Zhenjun Zhao, Heng Yang, Bangyan Liao, Yingping Zeng, Shaocheng Yan, Yingdong Gu, Peidong Liu, Yi Zhou, Haoang Li, Javier Civera
Abstract:
Global solvers have emerged as a powerful paradigm for 3D vision, offering certifiable solutions to nonconvex geometric optimization problems traditionally addressed by local or heuristic methods. This survey presents the first systematic review of global solvers in geometric vision, unifying the field through a comprehensive taxonomy of three core paradigms: Branch‑and‑Bound (BnB), Convex Relaxation (CR), and Graduated Non‑Convexity (GNC). We present their theoretical foundations, algorithmic designs, and practical enhancements for robustness and scalability, examining how each addresses the fundamental nonconvexity of geometric estimation problems. Our analysis spans ten core vision tasks, from Wahba problem to bundle adjustment, revealing the optimality‑robustness‑scalability trade‑offs that govern solver selection. We identify critical future directions: scaling algorithms while maintaining guarantees, integrating data‑driven priors with certifiable optimization, establishing standardized benchmarks, and addressing societal implications for safety‑critical deployment. By consolidating theoretical foundations, practical advances, and broader impacts, this survey provides a unified perspective and roadmap toward certifiable, trustworthy perception for real‑world applications. A continuously‑updated literature summary and companion code tutorials are available at https://github.com/ericzzj1989/Awesome‑Global‑Solvers‑for‑3D‑Vision.
Authors:Joanna Wojciechowicz, Maria Łubniewska, Jakub Antczak, Justyna Baczyńska, Wojciech Gromski, Wojciech Kozłowski, Maciej Zięba
Abstract:
We introduce VIGIL (Visual Inconsistency & Generative In‑context Lucidity), the first benchmark dataset and framework providing a fine‑grained categorization of hallucinations in the multimodal image recontextualization task for large multimodal models (LMMs). While existing research often treats hallucinations as a uniform issue, our work addresses a significant gap in multimodal evaluation by decomposing these errors into five categories: pasted object hallucinations, background hallucinations, object omission, positional & logical inconsistencies, and physical law violations. To address these complexities, we propose a multi‑stage detection pipeline. Our architecture processes recontextualized images through a series of specialized steps targeting object‑level fidelity, background consistency, and omission detection, leveraging a coordinated ensemble of open‑source models, whose effectiveness is demonstrated through extensive experimental evaluations. Our approach enables a deeper understanding of where the models fail with an explanation; thus, we fill a gap in the field, as no prior methods offer such categorization and decomposition for this task. To promote transparency and further exploration, we openly release VIGIL, along with the detection pipeline and benchmark code, through our GitHub repository: https://github.com/mlubneuskaya/vigil and Data repository: https://huggingface.co/datasets/joannaww/VIGIL.
Authors:Aswathi Varma, Suprosanna Shit, Chinmay Prabhakar, Daniel Scholz, Hongwei Bran Li, Bjoern Menze, Daniel Rueckert, Benedikt Wiestler
Abstract:
Vision Transformers (ViTs) have emerged as the state‑of‑the‑art architecture in representation learning, leveraging self‑attention mechanisms to excel in various tasks. ViTs split images into fixed‑size patches, constraining them to a predefined size and necessitating pre‑processing steps like resizing, padding, or cropping. This poses challenges in medical imaging, particularly with irregularly shaped structures like tumors. A fixed bounding box crop size produces input images with highly variable foreground‑to‑background ratios. Resizing medical images can degrade information and introduce artefacts, impacting diagnosis. Hence, tailoring variable‑sized crops to regions of interest can enhance feature representation capabilities. Moreover, large images are computationally expensive, and smaller sizes risk information loss, presenting a computation‑accuracy tradeoff. We propose VariViT, an improved ViT model crafted to handle variable image sizes while maintaining a consistent patch size. VariViT employs a novel positional embedding resizing scheme for a variable number of patches. We also implement a new batching strategy within VariViT to reduce computational complexity, resulting in faster training and inference times. In our evaluations on two 3D brain MRI datasets, VariViT surpasses vanilla ViTs and ResNet in glioma genotype prediction and brain tumor classification. It achieves F1‑scores of 75.5% and 76.3%, respectively, learning more discriminative features. Our proposed batching strategy reduces computation time by up to 30% compared to conventional architectures. These findings underscore the efficacy of VariViT in image representation learning. Our code can be found here: https://github.com/Aswathi‑Varma/varivit
Authors:Chenxu Dang, Sining Ang, Yongkang Li, Haochen Tian, Jie Wang, Guang Li, Hangjun Ye, Jie Ma, Long Chen, Yan Wang
Abstract:
Vision‑Language‑Action (VLA) models for autonomous driving increasingly adopt generative planners trained with imitation learning followed by reinforcement learning. Diffusion‑based planners suffer from modality alignment difficulties, low training efficiency, and limited generalization. Token‑based planners are plagued by cumulative causal errors and irreversible decoding. In summary, the two dominant paradigms exhibit complementary strengths and weaknesses. In this paper, we propose DriveFine, a masked diffusion VLA model that combines flexible decoding with self‑correction capabilities. In particular, we design a novel plug‑and‑play block‑MoE, which seamlessly injects a refinement expert on top of the generation expert. By enabling explicit expert selection during inference and gradient blocking during training, the two experts are fully decoupled, preserving the foundational capabilities and generic patterns of the pretrained weights, which highlights the flexibility and extensibility of the block‑MoE design. Furthermore, we design a hybrid reinforcement learning strategy that encourages effective exploration of refinement expert while maintaining training stability. Extensive experiments on NAVSIM v1, v2, and Navhard benchmarks demonstrate that DriveFine exhibits strong efficacy and robustness. The code will be released at https://github.com/MSunDYY/DriveFine.
Authors:Zhaotong Yang, Yong Du, Shengfeng He, Yuhui Li, Xinzhe Li, Yangyang Xu, Junyu Dong, Jian Yang
Abstract:
Image‑based Virtual Try‑On (VTON) concerns the synthesis of realistic person imagery through garment re‑rendering under human pose and body constraints. In practice, however, existing approaches are typically optimized for specific data conditions, making their deployment reliant on retraining and limiting their generalization as a unified solution. We present OmniVTON++, a training‑free VTON framework designed for universal applicability. It addresses the intertwined challenges of garment alignment, human structural coherence, and boundary continuity by coordinating Structured Garment Morphing for correspondence‑driven garment adaptation, Principal Pose Guidance for step‑wise structural regulation during diffusion sampling, and Continuous Boundary Stitching for boundary‑aware refinement, forming a cohesive pipeline without task‑specific retraining. Experimental results demonstrate that OmniVTON++ achieves state‑of‑the‑art performance across diverse generalization settings, including cross‑dataset and cross‑garment‑type evaluations, while reliably operating across scenarios and diffusion backbones within a single formulation. In addition to single‑garment, single‑human cases, the framework supports multi‑garment, multi‑human, and anime character virtual try‑on, expanding the scope of virtual try‑on applications. The code is available at https://github.com/Jerome‑Young/OmniVTON‑PlusPlus.
Authors:Hongpeng Wang, Zeyu Zhang, Wenhao Li, Hao Tang
Abstract:
Human motion understanding and generation are crucial for vision and robotics but remain limited in reasoning capability and test‑time planning. We propose MoRL, a unified multimodal motion model trained with supervised fine‑tuning and reinforcement learning with verifiable rewards. Our task‑specific reward design combines semantic alignment and reasoning coherence for understanding with physical plausibility and text‑motion consistency for generation, improving both logical reasoning and perceptual realism. To further enhance inference, we introduce Chain‑of‑Motion (CoM), a test‑time reasoning method that enables step‑by‑step planning and reflection. We also construct two large‑scale CoT datasets, MoUnd‑CoT‑140K and MoGen‑CoT‑140K, to align motion sequences with reasoning traces and action descriptions. Experiments on HumanML3D and KIT‑ML show that MoRL achieves significant gains over state‑of‑the‑art baselines. Code: https://github.com/AIGeeksGroup/MoRL. Website: https://aigeeksgroup.github.io/MoRL.
Authors:Jindong Zhao, Yuan Gao, Yang Xia, Sheng Nie, Jun Yue, Weiwei Sun, Shaobo Xia
Abstract:
Domain‑generalized LiDAR semantic segmentation (LSS) seeks to train models on source‑domain point clouds that generalize reliably to multiple unseen target domains, which is essential for real‑world LiDAR applications. However, existing approaches assume similar acquisition views (e.g., vehicle‑mounted) and struggle in cross‑view scenarios, where observations differ substantially due to viewpoint‑dependent structural incompleteness and non‑uniform point density. Accordingly, we formulate cross‑view domain generalization for LiDAR semantic segmentation and propose a novel framework, termed CVGC (Cross‑View Geometric Consistency). Specifically, we introduce a cross‑view geometric augmentation module that models viewpoint‑induced variations in visibility and sampling density, generating multiple cross‑view observations of the same scene. Subsequently, a geometric consistency module enforces consistent semantic and occupancy predictions across geometrically augmented point clouds of the same scene. Extensive experiments on six public LiDAR datasets establish the first systematic evaluation of cross‑view domain generalization for LiDAR semantic segmentation, demonstrating that CVGC consistently outperforms state‑of‑the‑art methods when generalizing from a single source domain to multiple target domains with heterogeneous acquisition viewpoints. The source code will be publicly available at https://github.com/KintomZi/CVGC‑DG
Authors:Aryan Das, Koushik Biswas, Swalpa Kumar Roy, Badri Narayana Patro, Vinay Kumar Verma
Abstract:
We introduce the Nexus Adapters, novel text‑guided efficient adapters to the diffusion‑based framework for the Structure Preserving Conditional Generation (SPCG). Recently, structure‑preserving methods have achieved promising results in conditional image generation by using a base model for prompt conditioning and an adapter for structure input, such as sketches or depth maps. These approaches are highly inefficient and sometimes require equal parameters in the adapter compared to the base architecture. It is not always possible to train the model since the diffusion model is itself costly, and doubling the parameter is highly inefficient. In these approaches, the adapter is not aware of the input prompt; therefore, it is optimal only for the structural input but not for the input prompt. To overcome the above challenges, we proposed two efficient adapters, Nexus Prime and Slim, which are guided by prompts and structural inputs. Each Nexus Block incorporates cross‑attention mechanisms to enable rich multimodal conditioning. Therefore, the proposed adapter has a better understanding of the input prompt while preserving the structure. We conducted extensive experiments on the proposed models and demonstrated that the Nexus Prime adapter significantly enhances performance, requiring only 8M additional parameters compared to the baseline, T2I‑Adapter. Furthermore, we also introduced a lightweight Nexus Slim adapter with 18M fewer parameters than the T2I‑Adapter, which still achieved state‑of‑the‑art results. Code: https://github.com/arya‑domain/Nexus‑Adapters
Authors:Mingrui Ma, Chentao Li, Pan Huang, Jing Qin
Abstract:
Whole slide images (WSIs) are the gold standard for pathological diagnosis and sub‑typing. Current main‑stream two‑step frameworks employ offline feature encoders trained without domain‑specific knowledge. Among them, attention‑based multiple instance learning (MIL) methods are outcome‑oriented and offer limited interpretability. Clustering‑based approaches can provide explainable decision‑making process but suffer from high dimension features and semantically ambiguous centroids. To this end, we propose an end‑to‑end MIL framework that integrates Grassmann re‑embedding and manifold adaptive clustering, where the manifold geometric structure facilitates robust clustering results. Furthermore, we design a prior knowledge guiding proxy instance labeling and aggregation strategy to approximate patch labels and focus on pathologically relevant tumor regions. Experiments on multicentre WSI datasets demonstrate that: 1) our cluster‑incorporated model achieves superior performance in both grading accuracy and interpretability; 2) end‑to‑end learning refines better feature representations and it requires acceptable computation resources.
Authors:Chentao Li, Pan Huang
Abstract:
The tumor region plays a key role in pathological diagnosis. Tumor tissues are highly similar to precancerous lesions and non tumor instances often greatly exceed tumor instances in whole slide images (WSIs). These issues cause instance‑semantic entanglement in multi‑instance learning frameworks, degrading both model representation capability and interpretability. To address this, we propose an end‑to‑end prototype instance semantic disentanglement framework with low‑rank regularized subspace clustering, PID‑LRSC, in two aspects. First, we use secondary instance subspace learning to construct low‑rank regularized subspace clustering (LRSC), addressing instance entanglement caused by an excessive proportion of non tumor instances. Second, we employ enhanced contrastive learning to design prototype instance semantic disentanglement (PID), resolving semantic entanglement caused by the high similarity between tumor and precancerous tissues. We conduct extensive experiments on multicentre pathology datasets, implying that PID‑LRSC outperforms other SOTA methods. Overall, PID‑LRSC provides clearer instance semantics during decision‑making and significantly enhances the reliability of auxiliary diagnostic outcomes.
Authors:Aryan Das, Tanishq Rachamalla, Koushik Biswas, Swalpa Kumar Roy, Vinay Kumar Verma
Abstract:
We introduce a novel uncertainty‑aware multimodal segmentation framework that leverages both radiological images and associated clinical text for precise medical diagnosis. We propose a Modality Decoding Attention Block (MoDAB) with a lightweight State Space Mixer (SSMix) to enable efficient cross‑modal fusion and long‑range dependency modelling. To guide learning under ambiguity, we propose the Spectral‑Entropic Uncertainty (SEU) Loss, which jointly captures spatial overlap, spectral consistency, and predictive uncertainty in a unified objective. In complex clinical circumstances with poor image quality, this formulation improves model reliability. Extensive experiments on various publicly available medical datasets, QATA‑COVID19, MosMed++, and Kvasir‑SEG, demonstrate that our method achieves superior segmentation performance while being significantly more computationally efficient than existing State‑of‑the‑Art (SoTA) approaches. Our results highlight the importance of incorporating uncertainty modelling and structured modality alignment in vision‑language medical segmentation tasks. Code: https://github.com/arya‑domain/UA‑VLS
Authors:Xinpeng Liu, Fumio Okura
Abstract:
3D Gaussian Splatting (3DGS) has enabled high‑fidelity virtualization with fast rendering and optimization for novel view synthesis. On the other hand, triangle mesh models still remain a popular choice for surface reconstruction but suffer from slow or heavy optimization in traditional mesh‑based differentiable renderers. To address this problem, we propose a new lightweight differentiable mesh renderer leveraging the efficient rasterization process of 3DGS, named Gaussian Mesh Renderer (GMR), which tightly integrates the Gaussian and mesh representations. Each Gaussian primitive is analytically derived from the corresponding mesh triangle, preserving structural fidelity and enabling the gradient flow. Compared to the traditional mesh renderers, our method achieves smoother gradients, which especially contributes to better optimization using smaller batch sizes with limited memory. Our implementation is available in the public GitHub repository at https://github.com/huntorochi/Gaussian‑Mesh‑Renderer.
Authors:Lanqing Guo, Xi Liu, Yufei Wang, Zhihao Li, Siyu Huang
Abstract:
Recent advances in image generation have achieved remarkable visual quality, while a fundamental challenge remains: Can image generation be controlled at the element level, enabling intuitive modifications such as adjusting shapes, altering colors, or adding and removing objects? In this work, we address this challenge by introducing layer‑wise controllable generation through simplified vector graphics (VGs). Our approach first efficiently parses images into hierarchical VG representations that are semantic‑aligned and structurally coherent. Building on this representation, we design a novel image synthesis framework guided by VGs, allowing users to freely modify elements and seamlessly translate these edits into photorealistic outputs. By leveraging the structural and semantic features of VGs in conjunction with noise prediction, our method provides precise control over geometry, color, and object semantics. Extensive experiments demonstrate the effectiveness of our approach in diverse applications, including image editing, object‑level manipulation, and fine‑grained content creation, establishing a new paradigm for controllable image generation. Project page: https://guolanqing.github.io/Vec2Pix/
Authors:In Chong Choi, Jiacheng Zhang, Feng Liu, Yiliao Song
Abstract:
Multi‑turn jailbreak attacks have proven effective against text‑only large language models (LLMs), where malicious content is gradually introduced to bypass safety alignment. However, effectively extending such attacks to large vision‑language models (LVLMs) remains underexplored. In this paper, we find that naively incorporating visual inputs can make multi‑turn jailbreaks easier to defend against; for example, overly malicious visual content will easily trigger the defense mechanism in safety‑aligned LVLMs, resulting in more conservative responses. Based on this finding, we propose multi‑turn adaptive prompting attack (MAPA) that 1) at each turn, alternates text‑vision attack actions to elicit the most malicious response; and 2) across turns, adjusts the attack trajectory through iterative back‑and‑forth refinement to gradually amplify response maliciousness. This two‑level design enables MAPA to consistently outperform state‑of‑the‑art methods, improving attack success rates by 15‑30% on recent benchmarks against LLaVA‑v1.6‑Mistral‑7B, Qwen2.5‑VL‑7B‑Instruct, Llama‑3.2‑Vision‑11B‑Instruct and GPT‑4o‑mini. Our code is available at: https://github.com/thomaschoi143/MAPA.
Authors:Ryan Fosdick
Abstract:
We describe an adaptation of VACE (Video All‑in‑one Creation and Editing) for real‑time autoregressive video generation. VACE provides unified video control (reference guidance, structural conditioning, inpainting, and temporal extension) but assumes bidirectional attention over full sequences, making it incompatible with streaming pipelines that require fixed chunk sizes and causal attention. The key modification moves reference frames from the diffusion latent space into a parallel conditioning pathway, preserving the fixed chunk sizes and KV caching that autoregressive models require. This adaptation reuses existing pretrained VACE weights without additional training. Across 1.3B and 14B model scales, VACE adds 20‑30% latency overhead for structural control and inpainting, with negligible VRAM cost relative to the base model. Reference‑to‑video fidelity is severely degraded compared to batch VACE due to causal attention constraints. A reference implementation is available at https://github.com/daydreamlive/scope.
Authors:A. Said Gurbuz, Sunghwan Hong, Ahmed Nassar, Marc Pollefeys, Peter Staar
Abstract:
Modern computer‑use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain, before they can reliably ground instructions and act. Yet, most available grounding datasets provide sparse supervision, with insufficient and low‑diversity labels that annotate only a small subset of task‑relevant elements per screen, which limits both coverage and generalization; moreover, practical deployment requires efficiency to enable low‑latency, on‑device use. We introduce ScreenParse, a large‑scale dataset for complete screen parsing, with dense annotations of all visible UI elements (boxes, 55‑class types, and text) across 771K web screenshots (21M elements). ScreenParse is generated by Webshot, an automated, scalable pipeline that renders diverse urls, extracts annotations and applies VLM‑based relabeling and quality filtering. Using ScreenParse, we train ScreenVLM, a compact, 316M‑parameter vision language model (VLM) that decodes a compact ScreenTag markup representation with a structure‑aware loss that upweights structure‑critical tokens. ScreenVLM substantially outperforms much larger foundation VLMs on dense parsing (e.g., 0.592 vs. 0.294 PageIoU on ScreenParse) and shows strong transfer to public benchmarks. Moreover, finetuning foundation VLMs on ScreenParse consistently improves their grounding performance, suggesting that dense screen supervision provides transferable structural priors for UI understanding. Project page: https://saidgurbuz.github.io/screenparse/.
Authors:Kaixuan Fang, Yuzhen Lu, Xinyang Mu
Abstract:
Traditional mechanized chestnut harvesting is too costly for small producers, non‑selective, and prone to damaging nuts. Accurate, reliable detection of chestnuts on the orchard floor is crucial for developing low‑cost, vision‑guided automated harvesting technology. However, developing a reliable chestnut detection system faces challenges in complex environments with shading, varying natural light conditions, and interference from weeds, fallen leaves, stones, and other foreign on‑ground objects, which have remained unaddressed. This study collected 319 images of chestnuts on the orchard floor, containing 6524 annotated chestnuts. A comprehensive set of 29 state‑of‑the‑art real‑time object detectors, including 14 in the YOLO (v11‑13) and 15 in the RT‑DETR (v1‑v4) families at varied model scales, was systematically evaluated through replicated modeling experiments for chestnut detection. Experimental results show that the YOLOv12m model achieves the best mAP@0.5 of 95.1% among all the evaluated models, while the RT‑DETRv2‑R101 was the most accurate variant among RT‑DETR models, with mAP@0.5 of 91.1%. In terms of mAP@[0.5:0.95], the YOLOv11x model achieved the best accuracy of 80.1%. All models demonstrate significant potential for real‑time chestnut detection, and YOLO models outperformed RT‑DETR models in terms of both detection accuracy and inference, making them better suited for on‑board deployment. Both the dataset and software programs in this study have been made publicly available at https://github.com/AgFood‑Sensing‑and‑Intelligence‑Lab/ChestnutDetection.
Authors:Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun, Tianwen Qian, Yuzheng Wu, Danda Pani Paudel, Xiangyang Xue, Yanwei Fu
Abstract:
Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision‑language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among these modalities, sound provides indispensable cues about spatial layout, off‑screen events, and causal interactions, particularly in egocentric settings where auditory and visual signals are tightly coupled. To this end, we introduce EgoSound, the first benchmark designed to systematically evaluate egocentric sound understanding in MLLMs. EgoSound unifies data from Ego4D and EgoBlind, encompassing both sighted and sound‑dependent experiences. It defines a seven‑task taxonomy spanning intrinsic sound perception, spatial localization, causal inference, and cross‑modal reasoning. Constructed through a multi‑stage auto‑generative pipeline, EgoSound contains 7315 validated QA pairs across 900 videos. Comprehensive experiments on nine state‑of‑the‑art MLLMs reveal that current models exhibit emerging auditory reasoning abilities but remain limited in fine‑grained spatial and causal understanding. EgoSound establishes a challenging foundation for advancing multisensory egocentric intelligence, bridging the gap between seeing and truly hearing the world. Project page: https://groolegend.github.io/EgoSound/ .
Authors:Youqi Wang, Shen Chen, Haowei Wang, Rongxuan Peng, Taiping Yao, Shunquan Tan, Changsheng Chen, Bin Li, Shouhong Ding
Abstract:
Existing Multimodal Large Language Models (MLLMs) for image forgery detection and localization predominantly operate under a text‑centric Chain‑of‑Thought (CoT) paradigm. However, forcing these models to textually characterize imperceptible low‑level tampering traces inevitably leads to hallucinations, as linguistic modalities are insufficient to capture such fine‑grained pixel‑level inconsistencies. To overcome this, we propose ForgeryVCR, a framework that incorporates a forensic toolbox to materialize imperceptible traces into explicit visual intermediates via Visual‑Centric Reasoning. To enable efficient tool utilization, we introduce a Strategic Tool Learning post‑training paradigm, encompassing gain‑driven trajectory construction for Supervised Fine‑Tuning (SFT) and subsequent Reinforcement Learning (RL) optimization guided by a tool utility reward. This paradigm empowers the MLLM to act as a proactive decision‑maker, learning to spontaneously invoke multi‑view reasoning paths including local zoom‑in for fine‑grained inspection and the analysis of invisible inconsistencies in compression history, noise residuals, and frequency domains. Extensive experiments reveal that ForgeryVCR achieves state‑of‑the‑art (SOTA) performance in both detection and localization tasks, demonstrating superior generalization and robustness with minimal tool redundancy. The project page is available at https://youqiwong.github.io/projects/ForgeryVCR/.
Authors:Kai Guan, Rongyuan Wu, Shuai Li, Wentao Zhu, Wenjun Zeng, Lei Zhang
Abstract:
In real‑world scenarios, the performance of semantic segmentation often deteriorates when processing low‑quality (LQ) images, which may lack clear semantic structures and high‑frequency details. Although image restoration techniques offer a promising direction for enhancing degraded visual content, conventional real‑world image restoration (Real‑IR) models primarily focus on pixel‑level fidelity and often fail to recover task‑relevant semantic cues, limiting their effectiveness when directly applied to downstream vision tasks. Conversely, existing segmentation models trained on high‑quality data lack robustness under real‑world degradations. In this paper, we propose Restoration Adaptation for Semantic Segmentation (RASS), which effectively integrates semantic image restoration into the segmentation process, enabling high‑quality semantic segmentation on the LQ images directly. Specifically, we first propose a Semantic‑Constrained Restoration (SCR) model, which injects segmentation priors into the restoration model by aligning its cross‑attention maps with segmentation masks, encouraging semantically faithful image reconstruction. Then, RASS transfers semantic restoration knowledge into segmentation through LoRA‑based module merging and task‑specific fine‑tuning, thereby enhancing the model's robustness to LQ images. To validate the effectiveness of our framework, we construct a real‑world LQ image segmentation dataset with high‑quality annotations, and conduct extensive experiments on both synthetic and real‑world LQ benchmarks. The results show that SCR and RASS significantly outperform state‑of‑the‑art methods in segmentation and restoration tasks. Code, models, and datasets will be available at https://github.com/Ka1Guan/RASS.git.
Authors:Yuang Ai, Jiaming Han, Shaobin Zhuang, Weijia Mao, Xuefeng Hu, Ziyan Yang, Zhenheng Yang, Yali Wang, Huaibo Huang, Xiangyu Yue, Hao Chen
Abstract:
We present BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices. With high‑entropy binary latents, BitDance lets each token represent up to 2^256 states, yielding a compact yet highly expressive discrete representation. Sampling from such a huge token space is difficult with standard classification. To resolve this, BitDance uses a binary diffusion head: instead of predicting an index with softmax, it employs continuous‑space diffusion to generate the binary tokens. Furthermore, we propose next‑patch diffusion, a new decoding method that predicts multiple tokens in parallel with high accuracy, greatly speeding up inference. On ImageNet 256x256, BitDance achieves an FID of 1.24, the best among AR models. With next‑patch diffusion, BitDance beats state‑of‑the‑art parallel AR models that use 1.4B parameters, while using 5.4x fewer parameters (260M) and achieving 8.7x speedup. For text‑to‑image generation, BitDance trains on large‑scale multimodal tokens and generates high‑resolution, photorealistic images efficiently, showing strong performance and favorable scaling. When generating 1024x1024 images, BitDance achieves a speedup of over 30x compared to prior AR models. We release code and models to facilitate further research on AR foundation models. Code and models are available at: https://github.com/shallowdream204/BitDance.
Authors:Jia Li, Xiaomeng Fu, Xurui Peng, Weifeng Chen, Youwei Zheng, Tianyu Zhao, Jiexi Wang, Fangmin Chen, Xing Wang, Hayden Kwok-Hay So
Abstract:
Autoregressive video diffusion models have emerged as a scalable paradigm for long video generation. However, they often suffer from severe extrapolation failure, where rapid error accumulation leads to significant temporal degradation when extending beyond training horizons. We identify that this failure primarily stems from the spectral bias of 3D positional embeddings and the lack of dynamic priors in noise sampling. To address these issues, we propose FLEX (Frequency‑aware Length EXtension), a training‑free inference‑time framework that bridges the gap between short‑term training and long‑term inference. FLEX introduces Frequency‑aware RoPE Modulation to adaptively interpolate under‑trained low‑frequency components while extrapolating high‑frequency ones to preserve multi‑scale temporal discriminability. This is integrated with Antiphase Noise Sampling (ANS) to inject high‑frequency dynamic priors and Inference‑only Attention Sink to anchor global structure. Extensive evaluations on VBench demonstrate that FLEX significantly outperforms state‑of‑the‑art models at 6x extrapolation (30s duration) and matches the performance of long‑video fine‑tuned baselines at 12x scale (60s duration). As a plug‑and‑play augmentation, FLEX seamlessly integrates into existing inference pipelines for horizon extension. It effectively pushes the generation limits of models such as LongLive, supporting consistent and dynamic video synthesis at a 4‑minute scale. Project page is available at https://ga‑lee.github.io/FLEX_demo.
Authors:Shenhan Qian, Ganlin Zhang, Shangzhe Wu, Daniel Cremers
Abstract:
Reconstructing and tracking dynamic 3D scenes remains a fundamental challenge in computer vision. Existing approaches often decouple geometry from motion: multi‑view reconstruction methods assume static scenes, while dynamic tracking frameworks rely on explicit camera pose estimation or separate motion models. We propose Flow4R, a unified framework that treats camera‑space scene flow as the central representation linking 3D structure, object motion, and camera motion. Flow4R predicts a minimal per‑pixel property set‑3D point position, scene flow, pose weight, and confidence‑from two‑view inputs using a Vision Transformer. This flow‑centric formulation allows local geometry and bidirectional motion to be inferred symmetrically with a shared decoder in a single forward pass, without requiring explicit pose regressors or bundle adjustment. Trained jointly on static and dynamic datasets, Flow4R achieves state‑of‑the‑art performance on 4D reconstruction and tracking tasks, demonstrating the effectiveness of the flow‑central representation for spatiotemporal scene understanding.
Authors:Jiangshan Wang, Zeqiang Lai, Jiarui Chen, Jiayi Guo, Hang Guo, Xiu Li, Xiangyu Yue, Chunchao Guo
Abstract:
Diffusion Transformers (DiT) have demonstrated remarkable generative capabilities but remain highly computationally expensive. Previous acceleration methods, such as pruning and distillation, typically rely on a fixed computational capacity, leading to insufficient acceleration and degraded generation quality. To address this limitation, we propose Elastic Diffusion Transformer (E‑DiT), an adaptive acceleration framework for DiT that effectively improves efficiency while maintaining generation quality. Specifically, we observe that the generative process of DiT exhibits substantial sparsity (i.e., some computations can be skipped with minimal impact on quality), and this sparsity varies significantly across samples. Motivated by this observation, E‑DiT equips each DiT block with a lightweight router that dynamically identifies sample‑dependent sparsity from the input latent. Each router adaptively determines whether the corresponding block can be skipped. If the block is not skipped, the router then predicts the optimal MLP width reduction ratio within the block. During inference, we further introduce a block‑level feature caching mechanism that leverages router predictions to eliminate redundant computations in a training‑free manner. Extensive experiments across 2D image (Qwen‑Image and FLUX) and 3D asset (Hunyuan3D‑3.0) demonstrate the effectiveness of E‑DiT, achieving up to ~2× speedup with negligible loss in generation quality. Code will be available at https://github.com/wangjiangshan0725/Elastic‑DiT.
Authors:Shuoyuan Wang, Yiran Wang, Hongxin Wei
Abstract:
Data‑driven approaches like deep learning are rapidly advancing planetary science, particularly in Mars exploration. Despite recent progress, most existing benchmarks remain confined to closed‑set supervised visual tasks and do not support text‑guided retrieval for geospatial discovery. We introduce MarsRetrieval, a retrieval benchmark for evaluating vision‑language models for Martian geospatial discovery. MarsRetrieval includes three tasks: (1) paired image‑text retrieval, (2) landform retrieval, and (3) global geo‑localization, covering multiple spatial scales and diverse geomorphic origins. We propose a unified retrieval‑centric protocol to benchmark multimodal embedding architectures, including contrastive dual‑tower encoders and generative vision‑language models. Our evaluation shows MarsRetrieval is challenging: even strong foundation models often fail to capture domain‑specific geomorphic distinctions. We further show that domain‑specific fine‑tuning is critical for generalizable geospatial discovery in planetary settings. Our code is available at https://github.com/ml‑stat‑Sustech/MarsRetrieval
Authors:Minghao Han, Dingkang Yang, Linhao Qu, Zizhi Chen, Gang Li, Han Wang, Jiacong Wang, Lihua Zhang
Abstract:
Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited pathological supervision, leading to representational bottlenecks. In this paper, we propose STAMP, a Spatial Transcriptomics‑Augmented Multimodal Pathology representation learning framework that integrates spatially‑resolved gene expression profiles to enable molecule‑guided joint embedding of pathology images and transcriptomic data. Our study shows that self‑supervised, gene‑guided training provides a robust and task‑agnostic signal for learning pathology image representations. Incorporating spatial context and multi‑scale information further enhances model performance and generalizability. To support this, we constructed SpaVis‑6M, the largest Visium‑based spatial transcriptomics dataset to date, and trained a spatially‑aware gene encoder on this resource. Leveraging hierarchical multi‑scale contrastive alignment and cross‑scale patch localization mechanisms, STAMP effectively aligns spatial transcriptomics with pathology images, capturing spatial structure and molecular variation. We validate STAMP across six datasets and four downstream tasks, where it consistently achieves strong performance. These results highlight the value and necessity of integrating spatially resolved molecular supervision for advancing multimodal learning in computational pathology. The code is included in the supplementary materials. The pretrained weights and SpaVis‑6M are available at: https://github.com/Hanminghao/STAMP.
Authors:Adson Duarte, Davide Vitturini, Emanuele Milillo, Andrea Bragagnolo, Carlo Alberto Barbano, Riccardo Renzulli, Michele Cannito, Federico Giacobbe, Francesco Bruno, Ovidio de Filippo, Fabrizio D'Ascenzo, Marco Grangetto
Abstract:
Cardiac Output (CO) is a key parameter in the diagnosis and management of cardiovascular diseases. However, its accurate measurement requires right‑heart catheterization, an invasive and time‑consuming procedure, motivating the development of reliable non‑invasive alternatives using echocardiography. In this work, we propose a self‑supervised learning (SSL) pretraining strategy based on SimCLR to improve CO prediction from apical four‑chamber echocardiographic videos. The pretraining is performed using the same limited dataset available for the downstream task, demonstrating the potential of SSL even under data scarcity. Our results show that SSL mitigates overfitting and improves representation learning, achieving an average Pearson correlation of 0.41 on the test set and outperforming PanEcho, a model trained on over one million echocardiographic exams. Source code is available at https://github.com/EIDOSLAB/cardiac‑output.
Authors:Giorgio Chiesa, Rossella Borra, Vittorio Lauro, Sabrina De Cillis, Daniele Amparore, Cristian Fiori, Riccardo Renzulli, Marco Grangetto
Abstract:
This paper presents a comprehensive workflow for generating and validating a synthetic dataset designed for robotic surgery instrument segmentation. A 3D reconstruction of the Da Vinci robotic arms was refined and animated in Autodesk Maya through a fully automated Python‑based pipeline capable of producing photorealistic, labeled video sequences. Each scene integrates randomized motion patterns, lighting variations, and synthetic blood textures to mimic intraoperative variability while preserving pixel‑accurate ground truth masks. To validate the realism and effectiveness of the generated data, several segmentation models were trained under controlled ratios of real and synthetic data. Results demonstrate that a balanced composition of real and synthetic samples significantly improves model generalization compared to training on real data only, while excessive reliance on synthetic data introduces a measurable domain shift. The proposed framework provides a reproducible and scalable tool for surgical computer vision, supporting future research in data augmentation, domain adaptation, and simulation‑based pretraining for robotic‑assisted surgery. Data and code are available at https://github.com/EIDOSLAB/Sintetic‑dataset‑DaVinci.
Authors:Michele Cannito, Riccardo Renzulli, Adson Duarte, Farzad Nikfam, Carlo Alberto Barbano, Enrico Chiesa, Francesco Bruno, Federico Giacobbe, Wojciech Wanha, Arturo Giordano, Marco Grangetto, Fabrizio D'Ascenzo
Abstract:
Severe aortic stenosis is a common and life‑threatening condition in elderly patients, often treated with Transcatheter Aortic Valve Implantation (TAVI). Despite procedural advances, paravalvular aortic regurgitation (PVR) remains one of the most frequent post‑TAVI complications, with a proven impact on long‑term prognosis.
In this work, we investigate the potential of deep learning to predict the occurrence of PVR from preoperative cardiac CT. To this end, a dataset of preoperative TAVI patients was collected, and 3D convolutional neural networks were trained on isotropic CT volumes. The results achieved suggest that volumetric deep learning can capture subtle anatomical features from pre‑TAVI imaging, opening new perspectives for personalized risk assessment and procedural optimization. Source code is available at https://github.com/EIDOSLAB/tavi.
Authors:Hengtong Shen, Li Yan, Hong Xie, Yaxuan Wei, Xinhao Li, Wenfei Shen, Peixian Lv, Fei Tan
Abstract:
Remote sensing (RS) change detection is essential for interpreting surface dynamics. Semantic change detection (SCD) further enables pixel‑level understanding of multi‑class transitions, yet remains sensitive to pseudo‑changes induced by imaging conditions. Recent RS foundation models extract semantically consistent features across temporal and environmental variations, which is critical for mitigating pseudo‑changes. However, existing SCD methods are often rigid and backbone‑specific, lacking the flexibility to integrate diverse multi‑scale features from emerging foundation models. To this end, we introduce a modular Cascaded Gated Decoder (CG‑Decoder) that bridges various backbones and SCD tasks, processing multi‑scale features in a coarse‑to‑fine manner while enabling adaptive change extraction. Building upon the RS foundation model PerA, we present PerASCD, a unified SCD framework. We further propose a Soft Semantic Consistency Loss (SSCLoss) to mitigate numerical instability in mixed‑precision training. Extensive experiments on SECOND and LandsatSCD show that PerASCD achieves new state‑of‑the‑art Sek scores (26.11% and 65.21%), surpassing the previous best by 0.61% and 4.95%, respectively. It also demonstrates exceptional data efficiency (outperforming the full‑data baseline with 50% data), seamless cross‑backbone generalization, and enhanced interpretability. Our approach maintains robust semantic consistency under radiometric variations, providing a reliable SCD solution. Code: https://github.com/SathShen/PerASCD.git.
Authors:Jidong Jia, Youjian Zhang, Huan Fu, Dacheng Tao
Abstract:
Despite advances in dance generation, most methods are trained in the skeletal domain and ignore mesh‑level physical constraints. As a result, motions that look plausible as joint trajectories often exhibit body self‑penetration and Foot‑Ground Contact (FGC) anomalies when visualized with a human body mesh, reducing the aesthetic appeal of generated dances and limiting their real‑world applications. We address this skeleton‑to‑mesh gap by deriving physics‑based rewards from the body mesh and applying Reinforcement Learning Fine‑Tuning (RLFT) to steer the diffusion model toward physically plausible motion synthesis under mesh visualization. Our reward design combines (i) an imitation reward that measures a motion's general plausibility by its imitability in a physical simulator (penalizing penetration and foot skating), and (ii) a Foot‑Ground Deviation (FGD) reward with test‑time FGD guidance to better capture the dynamic foot‑ground interaction in dance. However, we find that the physics‑based rewards tend to push the model to generate freezing motions for fewer physical anomalies and better imitability. To mitigate it, we propose an anti‑freezing reward to preserve motion dynamics while maintaining physical plausibility. Experiments on multiple dance datasets consistently demonstrate that our method can significantly improve the physical plausibility of generated motions, yielding more realistic and aesthetically pleasing dances. The project page is available at: https://jjd1123.github.io/Skeleton2Stage/
Authors:Khang Nguyen Quoc, Phuong D. Dao, Luyl-Da Quach
Abstract:
Foundation models and vision‑language pre‑training have significantly advanced Vision‑Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their application in domain‑specific agricultural tasks, such as plant pathology, remains limited due to the lack of large‑scale, comprehensive multimodal image‑‑text datasets and benchmarks. To address this gap, we introduce LeafNet, a comprehensive multimodal dataset, and LeafBench, a visual question‑answering benchmark developed to systematically evaluate the capabilities of VLMs in understanding plant diseases. The dataset comprises 186,000 leaf digital images spanning 97 disease classes, paired with metadata, generating 13,950 question‑answer pairs spanning six critical agricultural tasks. The questions assess various aspects of plant pathology understanding, including visual symptom recognition, taxonomic relationships, and diagnostic reasoning. Benchmarking 12 state‑of‑the‑art VLMs on our LeafBench dataset, we reveal substantial disparity in their disease understanding capabilities. Our study shows performance varies markedly across tasks: binary healthy‑‑diseased classification exceeds 90% accuracy, while fine‑grained pathogen and species identification remains below 65%. Direct comparison between vision‑only models and VLMs demonstrates the critical advantage of multimodal architectures: fine‑tuned VLMs outperform traditional vision models, confirming that integrating linguistic representations significantly enhances diagnostic precision. These findings highlight critical gaps in current VLMs for plant pathology applications and underscore the need for LeafBench as a rigorous framework for methodological advancement and progress evaluation toward reliable AI‑assisted plant disease diagnosis. Code is available at https://github.com/EnalisUs/LeafBench.
Authors:Yang Zhou, Derui Ding, Ran Sun, Ying Sun, Haohua Zhang
Abstract:
Visual object tracking (VOT) plays a pivotal role in unmanned aerial vehicle (UAV) applications. Addressing the trade‑off between accuracy and efficiency, especially under challenging conditions like unpredictable occlusion, remains a significant challenge. This paper introduces LGTrack, a unified UAV tracking framework that integrates dynamic layer selection, efficient feature enhancement, and robust representation learning for occlusions. By employing a novel lightweight Global‑Grouped Coordinate Attention (GGCA) module, LGTrack captures long‑range dependencies and global contexts, enhancing feature discriminability with minimal computational overhead. Additionally, a lightweight Similarity‑Guided Layer Adaptation (SGLA) module replaces knowledge distillation, achieving an optimal balance between tracking precision and inference efficiency. Experiments on three datasets demonstrate LGTrack's state‑of‑the‑art real‑time speed (258.7 FPS on UAVDT) while maintaining competitive tracking accuracy (82.8% precision). Code is available at https://github.com/XiaoMoc/LGTrack
Authors:Jiacheng Zhang, Feng Liu, Chao Du, Tianyu Pang
Abstract:
A line of recent training‑free methods for mitigating hallucinations in large vision‑language models (LVLMs) operates by amplifying attention to visual tokens during autoregressive generation within a single forward pass. We refer to this paradigm as visual attention amplification (VAA). In this paper, we identify a dual failure pattern in existing VAA methods caused by their use of a fixed amplification factor across generation steps: it can be too weak at some steps, leaving hallucinations unresolved, while too strong at others, introducing new hallucinations. Motivated by this finding, we propose Step‑wise Adaptive Visual Attention Amplification (SAVAA), a new VAA framework that estimates hallucination risk for each generated token and uses the estimated risk to adaptively amplify visual attention at the next generation step. Specifically, we introduce Visual Grounding Entropy (VGE), a lightweight hallucination‑risk estimator that augments predictive entropy with visual grounding, assigning higher risk to tokens that are uncertain, weakly grounded in the image, or both. Guided by VGE, SAVAA uses the estimated risk to calibrate the VAA factor for the next generation step, applying stronger amplification to higher‑risk steps and weaker amplification to lower‑risk steps. Across LLaVA‑NeXT‑7B, Qwen3‑VL‑8B, and InternVL3.5‑8B, SAVAA significantly outperforms baseline methods on generative hallucination benchmarks such as CHAIR, SHR and AMBER. Code is available at: https://github.com/JiachengZ01/SAVVA.
Authors:Feng Gao, Zheng Gong, Wenli Liu, Yanhai Gan, Zhuoran Zheng, Junyu Dong, Qian Du
Abstract:
While Mamba models offer efficient sequence modeling, vanilla versions struggle with temporal correlations and boundary details in Arctic sea ice concentration (SIC) prediction. To address these limitations, we propose Frequency‑enhanced Hilbert scanning Mamba Framework (FH‑Mamba) for short‑term Arctic SIC prediction. Specifically, we introduce a 3D Hilbert scan mechanism that traverses the 3D spatiotemporal grid along a locality‑preserving path, ensuring that adjacent indices in the flattened sequence correspond to neighboring voxels in both spatial and temporal dimensions. Additionally, we incorporate wavelet transform to amplify high‑frequency details and we also design a Hybrid Shuffle Attention module to adaptively aggregate sequence and frequency features. Experiments conducted on the OSI‑450a1 and AMSR2 datasets demonstrate that our FH‑Mamba achieves superior prediction performance compared with state‑of‑the‑art baselines. The results confirm the effectiveness of Hilbert scanning and frequency‑aware attention in improving both temporal consistency and edge reconstruction for Arctic SIC forecasting. Our codes are publicly available at https://github.com/oucailab/FH‑Mamba.
Authors:Huajian Zeng, Lingyun Chen, Jiaqi Yang, Yuantai Zhang, Fan Shi, Peidong Liu, Xingxing Zuo
Abstract:
Recent vision‑language‑action (VLA) models can generate plausible end‑effector motions, yet they often fail in long‑horizon, contact‑rich tasks because the underlying hand‑object interaction (HOI) structure is not explicitly represented. An embodiment‑agnostic interaction representation that captures this structure would make manipulation behaviors easier to validate and transfer across robots. We propose FlowHOI, a two‑stage flow‑matching framework that generates semantically grounded, temporally coherent HOI sequences, comprising hand poses, object poses, and hand‑object contact states, conditioned on an egocentric observation, a language instruction, and a 3D Gaussian splatting (3DGS) scene reconstruction. We decouple geometry‑centric grasping from semantics‑centric manipulation, conditioning the latter on compact 3D scene tokens and employing a motion‑text alignment loss to semantically ground the generated interactions in both the physical scene layout and the language instruction. To address the scarcity of high‑fidelity HOI supervision, we introduce a reconstruction pipeline that recovers aligned hand‑object trajectories and meshes from large‑scale egocentric videos, yielding an HOI prior for robust generation. Across the GRAB and HOT3D benchmarks, FlowHOI achieves the highest action recognition accuracy and a 1.7× higher physics simulation success rate than the strongest diffusion‑based baseline, while delivering a 40× inference speedup. We further demonstrate real‑robot execution on four dexterous manipulation tasks, illustrating the feasibility of retargeting generated HOI representations to real‑robot execution pipelines.
Authors:Sebastian-Ion Nae, Mihai-Eugen Barbu, Sebastian Mocanu, Marius Leordeanu
Abstract:
Autonomous agents such as indoor drones must learn new object classes in real‑time while limiting catastrophic forgetting, motivating Class‑Incremental Learning (CIL). However, most unmanned aerial vehicle (UAV) datasets focus on outdoor scenes and offer limited temporally coherent indoor videos. We introduce an indoor dataset of 14,400 frames capturing inter‑drone and ground vehicle footage, annotated via a semi‑automatic workflow with a 98.6% first‑pass labeling agreement before final manual verification. Using this dataset, we benchmark 3 replay‑based CIL strategies: Experience Replay (ER), Maximally Interfered Retrieval (MIR), and Forgetting‑Aware Replay (FAR), using YOLOv11‑nano as a resource‑efficient detector for deployment‑constrained UAV platforms. Under tight memory budgets (5‑10% replay), FAR performs better than the rest, achieving an average accuracy (ACC, mAP_50‑95 across increments) of 82.96% with 5% replay. Gradient‑weighted class activation mapping (Grad‑CAM) analysis shows attention shifts across classes in mixed scenes, which is associated with reduced localization quality for drones. The experiments further demonstrate that replay‑based continual learning can be effectively applied to edge aerial systems. Overall, this work contributes an indoor UAV video dataset with preserved temporal coherence and an evaluation of replay‑based CIL under limited replay budgets. Project page: https://spacetime‑vision‑robotics‑laboratory.github.io/learning‑on‑the‑fly‑cl
Authors:Zhen Wang, Yiming Gao, Jieyuan Liu, Enze Ma, Jefferson Chen, Mark Antkowiak, Mengzhou Hu, JungHo Kong, Dexter Pratt, Zhiting Hu, Wei Wang, Trey Ideker, Eric P. Xing
Abstract:
Single‑cell RNA‑seq (scRNA‑seq) enables atlas‑scale profiling of complex tissues, revealing rare lineages and transient states. Yet, assigning biologically valid cell identities remains a bottleneck because markers are tissue‑ and state‑dependent, and novel states lack references. We present CellMaster, an AI agent that mimics expert practice for zero‑shot cell‑type annotation. Unlike existing automated tools, CellMaster leverages LLM‑encoded knowledge (e.g., GPT‑4o) to perform on‑the‑fly annotation with interpretable rationales, without pre‑training or fixed marker databases. Across 9 datasets spanning 8 tissues, CellMaster improved accuracy by 7.1% over best‑performing baselines (including CellTypist and scTab) in automatic mode. With human‑in‑the‑loop refinement, this advantage increased to 18.6%, with a 22.1% gain on subtype populations. The system demonstrates particular strength in rare and novel cell states where baselines often fail. Source code and the web application are available at \hrefhttps://github.com/AnonymousGym/CellMasterhttps://github.com/AnonymousGym/CellMaster.
Authors:Jiamiao Lu, Wei Wu, Ke Gao, Ping Mao, Weichuan Zhang, Tuo Wang, Lingkun Ma, Jiapan Guo, Zanyi Wu, Yuqing Hu, Changming Sun
Abstract:
The biological behavior and treatment response of meningiomas depend on their grade, making an accurate diagnosis essential for treatment planning and prognosis assessment. We observed that the weighted fusion of spatial‑frequency domain features significantly influences meningioma classification performance. Notably, the contribution of specific frequency bands obtained by discrete wavelet transform varies considerably across different images. A feature fusion architecture with adaptive weights of different frequency band information and spatial domain information is proposed for few‑shot meningioma learning. To verify the effectiveness of the proposed method, a new MRI dataset of meningiomas is introduced. The experimental results demonstrate the superiority of the proposed method compared with existing state‑of‑the‑art methods in three datasets. The code will be available at: https://github.com/ICL‑SUST/AMSF‑Net
Authors:Daesik Jang, Morgan Lindsay Heisler, Linzi Xing, Yifei Li, Edward Wang, Ying Xiong, Yong Zhang, Zhenan Fan
Abstract:
Automatically generating and iteratively editing academic slide decks requires more than document summarization. It demands faithful content selection, coherent slide organization, layout‑aware rendering, and robust multi‑turn instruction following. However, existing benchmarks and evaluation protocols do not adequately measure these challenges. To address this gap, we introduce the Deck Edits and Compliance Kit Benchmark (DECKBench), an evaluation framework for multi‑agent slide generation and editing. DECKBench is built on a curated dataset of paper to slide pairs augmented with realistic, simulated editing instructions. Our evaluation protocol systematically assesses slide‑level and deck‑level fidelity, coherence, layout quality, and multi‑turn instruction following. We further implement a modular multi‑agent baseline system that decomposes the slide generation and editing task into paper parsing and summarization, slide planning, HTML creation, and iterative editing. Experimental results demonstrate that the proposed benchmark highlights strengths, exposes failure modes, and provides actionable insights for improving multi‑agent slide generation and editing systems. Overall, this work establishes a standardized foundation for reproducible and comparable evaluation of academic presentation generation and editing. Code and data are publicly available at https://github.com/morgan‑heisler/DeckBench .
Authors:Yifan Tan, Yifu Sun, Shirui Huang, Hong Liu, Guanghua Yu, Jianchen Zhu, Yangdong Deng
Abstract:
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities, yet they encounter significant computational bottlenecks due to the massive volume of visual tokens. Consequently, visual token pruning, which substantially reduces the token count, has emerged as a critical technique for accelerating MLLM inference. Existing approaches focus on token importance, diversity, or an intuitive combination of both, without a principled framework for their optimal integration. To address this issue, we first conduct a systematic analysis to characterize the trade‑off between token importance and semantic diversity. Guided by this analysis, we propose the Importance and Diversity Pruner (IDPruner), which leverages the Maximal Marginal Relevance (MMR) algorithm to achieve a Pareto‑optimal balance between these two objectives. Crucially, our method operates without requiring attention maps, ensuring full compatibility with FlashAttention and efficient deployment via one‑shot pruning. We conduct extensive experiments across various model architectures and multimodal benchmarks, demonstrating that IDPruner achieves state‑of‑the‑art performance and superior generalization across diverse architectures and tasks. Notably, on Qwen2.5‑VL‑7B‑Instruct, IDPruner retains 95.18% of baseline performance when pruning 75% of the tokens, and still maintains 86.40% even under an extreme 90% pruning ratio. Our code is available at https://github.com/Tencent/AngelSlim.
Authors:Aydin Ayanzadeh, Prakhar Dixit, Sadia Kamal, Milton Halem
Abstract:
Wildfires are a growing threat to ecosystems, human lives, and infrastructure, with their frequency and intensity rising due to climate change and human activities. Early detection is critical, yet satellite‑based monitoring remains challenging due to faint smoke signals, dynamic weather conditions, and the need for real‑time analysis over large areas. We introduce WildfireVLM, an AI framework that combines satellite imagery wildfire detection with language‑driven risk assessment. We construct a labeled wildfire and smoke dataset using imagery from Landsat‑8/9, GOES‑16, and other publicly available Earth observation sources, including harmonized products with aligned spectral bands. WildfireVLM employs YOLOv12 to detect fire zones and smoke plumes, leveraging its ability to detect small, complex patterns in satellite imagery. We integrate Multimodal Large Language Models (MLLMs) that convert detection outputs into contextualized risk assessments and prioritized response recommendations for disaster management. We validate the quality of risk reasoning using an LLM‑as‑judge evaluation with a shared rubric. The system is deployed using a service‑oriented architecture that supports real‑time processing, visual risk dashboards, and long‑term wildfire tracking, demonstrating the value of combining computer vision with language‑based reasoning for scalable wildfire monitoring. The code and dataset are publicly available on GitHub at https://github.com/Ayanzadeh93/_WildfireVLM_.
Authors:Jiahao Qin
Abstract:
Deformable image registration across heterogeneous domains remains challenging because coupled appearance variation and geometric misalignment violate the brightness constancy assumption underlying conventional methods. We propose PCReg‑Net, a progressive contrast‑guided registration framework that performs coarse‑to‑fine alignment through four lightweight modules: (1)~a registration U‑Net for initial coarse alignment, (2)~a reference feature extractor capturing multi‑scale structural cues from the fixed image, (3)~a multi‑scale contrast module that identifies residual misalignment by comparing coarse‑registered and reference features, and (4)~a refinement U‑Net with feature injection that produces the final high‑fidelity output. We evaluate on the FIRE‑Reg‑256 retinal fundus benchmark, demonstrating improvements over both traditional and deep learning baselines. Additional experiments on two microscopy benchmarks further confirm cross‑domain applicability. With only 2.56M parameters, PCReg‑Net achieves real‑time inference at 141 FPS. Code is available at https://github.com/JiahaoQin/PCReg‑Net.
Authors:Jiarong Liang, Max Ku, Ka-Hei Hui, Ping Nie, Wenhu Chen
Abstract:
Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks rely on recognition‑style protocols such as Visual Question Answering (VQA) and Violation of Expectation (VoE), which can often be answered without committing to an explicit, testable physical hypothesis. We propose VisPhyWorld, an execution‑based framework that evaluates physical reasoning by requiring models to generate executable simulator code from visual observations. By producing runnable code, the inferred world representation is directly inspectable, editable, and falsifiable. This separates physical reasoning from rendering. Building on this framework, we introduce VisPhyBench, comprising 209 evaluation scenes derived from 108 physical templates and a systematic protocol that evaluates how well models reconstruct appearance and reproduce physically plausible motion. Our pipeline produces valid reconstructed videos in 97.7% of benchmark runs before fallback. Experiments show that while state‑of‑the‑art MLLMs achieve strong semantic scene understanding, they struggle to accurately infer physical parameters and to simulate consistent physical dynamics. Our code is available https://github.com/TIGER‑AI‑Lab/VisPhyWorld
Authors:Xiaoxu Peng, Dong Zhou, Jianwen Zhang, Guanghui Sun, Anh Tu Ngo, Anupam Chattopadhyay
Abstract:
Vision Language Models (VLMs) have advanced perception in autonomous driving (AD), but they remain vulnerable to adversarial threats. These risks range from localized physical patches to imperceptible global perturbations. Existing defense methods for VLMs remain limited and often fail to reconcile robustness with clean‑sample performance. To bridge these gaps, we propose NutVLM, a comprehensive self‑adaptive defense framework designed to secure the entire perception‑decision lifecycle. Specifically, we first employ NutNet++ as a sentinel, which is a unified detection‑purification mechanism. It identifies benign samples, local patches, and global perturbations through three‑way classification. Subsequently, localized threats are purified via efficient grayscale masking, while global perturbations trigger Expert‑guided Adversarial Prompt Tuning (EAPT). Instead of the costly parameter updates of full‑model fine‑tuning, EAPT generates "corrective driving prompts" via gradient‑based latent optimization and discrete projection. These prompts refocus the VLM's attention without requiring exhaustive full‑model retraining. Evaluated on the Dolphins benchmark, our NutVLM yields a 4.89% improvement in overall metrics (e.g., Accuracy, Language Score, and GPT Score). These results validate NutVLM as a scalable security solution for intelligent transportation. Our code is available at https://github.com/PXX/NutVLM.
Authors:Yuqi Xiong, Chunyi Peng, Zhipeng Xu, Zhenghao Liu, Zulong Chen, Yukun Yan, Shuo Wang, Yu Gu, Ge Yu
Abstract:
Visual Retrieval‑Augmented Generation (VRAG) enhances Vision‑Language Models (VLMs) by incorporating external visual documents to address a given query. Existing VRAG frameworks usually depend on rigid, pre‑defined external tools to extend the perceptual capabilities of VLMs, typically by explicitly separating visual perception from subsequent reasoning processes. However, this decoupled design can lead to unnecessary loss of visual information, particularly when image‑based operations such as cropping are applied. In this paper, we propose Lang2Act, which enables fine‑grained visual perception and reasoning through self‑emergent linguistic toolchains. Rather than invoking fixed external engines, Lang2Act collects self‑emergent actions as linguistic tools and leverages them to enhance the visual perception capabilities of VLMs. To support this mechanism, we design a two‑stage Reinforcement Learning (RL)‑based training framework. Specifically, the first stage optimizes VLMs to self‑explore high‑quality actions for constructing a reusable linguistic toolbox, and the second stage further optimizes VLMs to exploit these linguistic tools for downstream reasoning effectively. Experimental results demonstrate the effectiveness of Lang2Act in substantially enhancing the visual perception capabilities of VLMs, achieving performance improvements of over 4%. All code and data are available at https://github.com/NEUIR/Lang2Act.
Authors:Aadarsh Sahoo, Georgia Gkioxari
Abstract:
Conversational image segmentation grounds abstract, intent‑driven concepts into pixel‑accurate masks. Prior work on referring image grounding focuses on categorical and spatial queries (e.g., "left‑most apple") and overlooks functional and physical reasoning (e.g., "where can I safely store the knife?"). We address this gap and introduce Conversational Image Segmentation (CIS) and ConverSeg, a benchmark spanning entities, spatial relations, intent, affordances, functions, safety, and physical reasoning. We also present ConverSeg‑Net, which fuses strong segmentation priors with language understanding, and an AI‑powered data engine that generates prompt‑mask pairs without human supervision. We show that current language‑guided segmentation models are inadequate for CIS, while ConverSeg‑Net trained on our data engine achieves significant gains on ConverSeg and maintains strong performance on existing language‑guided segmentation benchmarks. Project webpage: https://glab‑caltech.github.io/converseg/
Authors:Sayan Deb Sarkar, Rémi Pautrat, Ondrej Miksik, Marc Pollefeys, Iro Armeni, Mahdi Rad, Mihai Dusmanu
Abstract:
Video Language Models (VideoLMs) enable AI systems to understand temporal dynamics in videos. To fit within the maximum context window constraint, current methods use keyframe sampling which often misses both macro‑level events and micro‑level details due to the sparse temporal coverage. Furthermore, processing full images and their tokens for each frame incurs substantial computational overhead. We address these limitations by leveraging video codec primitives (specifically motion vectors and residuals) which natively encode video redundancy and sparsity without requiring expensive full‑image encoding for most frames. To this end, we introduce lightweight transformer‑based encoders that aggregate codec primitives and align their representations with image encoder embeddings through a pre‑training strategy that accelerates convergence during end‑to‑end fine‑tuning. Our approach, CoPE‑VideoLM, reduces the time‑to‑first‑token by up to 86% and token usage by up to 93% compared to standard VideoLMs. Moreover, by varying the keyframe and codec primitive densities we maintain or exceed performance on 14 diverse video understanding benchmarks spanning general question answering, temporal and motion reasoning, long‑form understanding, and spatial scene understanding.
Authors:Mingzhi Sheng, Zekai Gu, Peng Li, Cheng Lin, Hao-Xiang Guo, Ying-Cong Chen, Yuan Liu
Abstract:
Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task‑specific signals, we argue that a fundamental disentanglement of "appearance" and "motion" provides a more robust and scalable pathway. We propose FlexAM, a unified framework built upon a novel 3D control signal. This signal represents video dynamics as a point cloud, introducing three key enhancements: multi‑frequency positional encoding to distinguish fine‑grained motion, depth‑aware positional encoding, and a flexible control signal for balancing precision and generative quality. This representation allows FlexAM to effectively disentangle appearance and motion, enabling a wide range of tasks including I2V/V2V editing, camera control, and spatial object editing. Extensive experiments demonstrate that FlexAM achieves superior performance across all evaluated tasks.
Authors:Chong Cheng, Xianda Chen, Tao Xie, Wei Yin, Weiqiang Ren, Qian Zhang, Xiaoyang Guo, Hao Wang
Abstract:
Long‑sequence streaming 3D reconstruction remains a significant open challenge. Existing autoregressive models often fail when processing long sequences because they anchor poses to the first frame, leading to attention decay, scale drift, and extrapolation errors. We introduce LongStream, a novel gauge‑decoupled streaming visual geometry model for metric‑scale scene reconstruction across thousands of frames under a strictly online, future‑invisible setting. Our approach is threefold. First, we discard the first‑frame anchor and predict keyframe‑relative poses. This reformulates long‑range extrapolation into a constant‑difficulty local task. Second, we introduce orthogonal scale learning. This method fully disentangles geometry from scale estimation to suppress drift. Finally, we identify attention bias issues in Transformers, including attention‑sink reliance and long‑term KV‑cache saturation. We propose cache‑consistent training combined with periodic cache refresh. This approach suppresses attention biases and contamination over ultra‑long sequences and reduces the gap between training and inference. Experiments show that LongStream achieves state‑of‑the‑art performance, enabling stable, metric‑scale reconstruction over kilometer‑scale sequences at 18 FPS. Project Page: https://3dagentworld.github.io/longstream/
Authors:Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe, Mubarak Shah
Abstract:
Direct Preference Optimization (DPO) has been proposed as an effective and efficient alternative to reinforcement learning from human feedback (RLHF). However, neither RLHF nor DPO take into account the fact that learning certain preferences is more difficult than learning other preferences, rendering the optimization process suboptimal. To address this gap in text‑to‑image generation, we recently proposed Curriculum‑DPO, a method that organizes image pairs by difficulty. In this paper, we introduce Curriculum‑DPO++, an enhanced method that combines the original data‑level curriculum with a novel model‑level curriculum. More precisely, we propose to dynamically increase the learning capacity of the denoising network as training advances. We implement this capacity increase via two mechanisms. First, we initialize the model with only a subset of the trainable layers used in the original Curriculum‑DPO. As training progresses, we sequentially unfreeze layers until the configuration matches the full baseline architecture. Second, as the fine‑tuning is based on Low‑Rank Adaptation (LoRA), we implement a progressive schedule for the dimension of the low‑rank matrices. Instead of maintaining a fixed capacity, we initialize the low‑rank matrices with a dimension significantly smaller than that of the baseline. As training proceeds, we incrementally increase their rank, allowing the capacity to grow until it converges to the same rank value as in Curriculum‑DPO. Furthermore, we propose an alternative ranking strategy to the one employed by Curriculum‑DPO. Finally, we compare Curriculum‑DPO++ against Curriculum‑DPO and other state‑of‑the‑art preference optimization approaches on nine benchmarks, outperforming the competing methods in terms of text alignment, aesthetics and human preference. Our code is available at https://github.com/CroitoruAlin/Curriculum‑DPO.
Authors:Alejandro Dopico-Castro, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Iván Pérez Digón
Abstract:
Federated Learning (FL) enables collaborative training without centralizing data, essential for privacy compliance in real‑world scenarios involving sensitive visual information. Most FL approaches rely on expensive, iterative deep network optimization, which still risks privacy via shared gradients. In this work, we propose FedHENet, extending the FedHEONN framework to image classification. By using a fixed, pre‑trained feature extractor and learning only a single output layer, we avoid costly local fine‑tuning. This layer is learned by analytically aggregating client knowledge in a single round of communication using homomorphic encryption (HE). Experiments show that FedHENet achieves competitive accuracy compared to iterative FL baselines while demonstrating superior stability performance and up to 70% better energy efficiency. Crucially, our method is hyperparameter‑free, removing the carbon footprint associated with hyperparameter tuning in standard FL. Code available in https://github.com/AlejandroDopico2/FedHENet/
Authors:Boujemaa Guermazi, Riadh Ksantini, Naimul Khan
Abstract:
Unsupervised image segmentation is a critical task in computer vision. It enables dense scene understanding without human annotations, which is especially valuable in domains where labelled data is scarce. However, existing methods often struggle to reconcile global semantic structure with fine‑grained boundary accuracy. This paper introduces DynaGuide, an adaptive segmentation framework that addresses these challenges through a novel dual‑guidance strategy and dynamic loss optimization. Building on our previous work, DynaSeg, DynaGuide combines global pseudo‑labels from zero‑shot models such as DiffSeg or SegFormer with local boundary refinement using a lightweight CNN trained from scratch. This synergy allows the model to correct coarse or noisy global predictions and produce high‑precision segmentations. At the heart of DynaGuide is a multi‑component loss that dynamically balances feature similarity, Huber‑smoothed spatial continuity, including diagonal relationships, and semantic alignment with the global pseudo‑labels. Unlike prior approaches, DynaGuide trains entirely without ground‑truth labels in the target domain and supports plug‑and‑play integration of diverse guidance sources. Extensive experiments on BSD500, PASCAL VOC2012, and COCO demonstrate that DynaGuide achieves state‑of‑the‑art performance, improving mIoU by 17.5% on BSD500, 3.1% on PASCAL VOC2012, and 11.66% on COCO. With its modular design, strong generalization, and minimal computational footprint, DynaGuide offers a scalable and practical solution for unsupervised segmentation in real‑world settings. Code available at: https://github.com/RyersonMultimediaLab/DynaGuide
Authors:Feng Yu, Xiangyu Wu, Yang Yang, Jianfeng Lu
Abstract:
Multimodal learning integrates data from diverse sensors to effectively harness information from different modalities. However, recent studies reveal that joint learning often overfits certain modalities while neglecting others, leading to performance inferior to that of unimodal learning. Although previous efforts have sought to balance modal contributions or combine joint and unimodal learning, thereby mitigating the degradation of weaker modalities with promising outcomes, few have examined the relationship between joint and unimodal learning from an information‑theoretic perspective. In this paper, we theoretically analyze modality competition and propose a method for multimodal classification by maximizing the total correlation between multimodal features and labels. By maximizing this objective, our approach alleviates modality competition while capturing inter‑modal interactions via feature alignment. Building on Mutual Information Neural Estimation (MINE), we introduce Total Correlation Neural Estimation (TCNE) to derive a lower bound for total correlation. Subsequently, we present TCMax, a hyperparameter‑free loss function that maximizes total correlation through variational bound optimization. Extensive experiments demonstrate that TCMax outperforms state‑of‑the‑art joint and unimodal learning approaches. Our code is available at https://github.com/hubaak/TCMax.
Authors:Mohammed Amine Bencheikh Lehocine, Julian Schmidt, Frank Moosmann, Dikshant Gupta, Fabian Flohr
Abstract:
Classical autonomous driving systems connect perception and prediction modules via hand‑crafted bounding‑box interfaces, limiting information flow and propagating errors to downstream tasks. Recent research aims to develop end‑to‑end models that jointly address perception and prediction; however, they often fail to fully exploit the synergy between appearance and motion cues, relying mainly on short‑term visual features. We follow the idea of "looking backward to look forward", and propose MASAR, a novel fully differentiable framework for joint 3D detection and trajectory forecasting compatible with any transformer‑based 3D detector. MASAR employs an object‑centric spatio‑temporal mechanism that jointly encodes appearance and motion features. By predicting past trajectories and refining them using guidance from appearance cues, MASAR captures long‑term temporal dependencies that enhance future trajectory forecasting. Experiments conducted on the nuScenes dataset demonstrate MASAR's effectiveness, showing improvements of over 20% in minADE and minFDE while maintaining robust detection performance. Code and models are available at https://github.com/aminmed/MASAR.
Authors:Weicheng Gao
Abstract:
Non‑line‑of‑sight sensing of human activities in complex environments is enabled by multiple‑input multiple‑output through‑the‑wall radar (TWR). However, the distinctiveness of micro‑Doppler signature between similar indoor human activities such as gun carrying and normal walking is minimal, while the large scale of input images required for effective identification utilizing time‑frequency spectrograms creates challenges for model training and inference efficiency. To address this issue, the Chebyshev‑time map is proposed in this paper, which is a method characterizing micro‑Doppler signature using polynomial orders. The parametric kinematic models for human motion and the TWR echo model are first established. Then, a time‑frequency feature representation method based on orthogonal Chebyshev polynomial decomposition is proposed. The kinematic envelopes of the torso and limbs are extracted, and the time‑frequency spectrum slices are mapped into a robust Chebyshev‑time coefficient space, preserving the multi‑order morphological detail information of time‑frequency spectrum. Numerical simulations and experiments are conducted to verify the effectiveness of the proposed method, which demonstrates the capability to characterize armed and unarmed indoor human activities while effectively compressing the scale of the time‑frequency spectrum to achieve a balance between recognition accuracy and input data dimensions. The open‑source code of this paper can be found in: https://github.com/JoeyBGOfficial/Represent‑Micro‑Doppler‑Signature‑in‑Orders.
Authors:Filippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello, Simone Calderara
Abstract:
Adapting large pre‑trained models to downstream tasks often produces task‑specific parameter updates that are expensive to relearn for every model variant. While recent work has shown that such updates can be transferred between models with identical architectures, transferring them across models of different widths remains unexplored. In this work, we introduce Theseus, a training‑free method for transporting task updates across heterogeneous‑width models. Rather than matching parameters, we characterize a task update by the functional effect it induces on intermediate representations. We formalize task‑vector transport as a functional matching problem on observed activations and show that, after aligning representation spaces via orthogonal Procrustes analysis, it admits a stable closed‑form solution that preserves the geometry of the update. We evaluate Theseus on vision and language models across different widths, showing consistent improvements over baselines without additional training or backpropagation. Our results show that task updates can be meaningfully transferred across architectures when task identity is defined functionally rather than parametrically. Code is available at https://github.com/apanariello4/merge‑and‑rebase.
Authors:Xiao Wang, Xingxing Xiong, Jinfeng Gao, Xufeng Lou, Bo Jiang, Si-bao Chen, Yaowei Wang, Yonghong Tian
Abstract:
Event stream‑based Visual Place Recognition (VPR) is an emerging research direction that offers a compelling solution to the instability of conventional visible‑light cameras under challenging conditions such as low illumination, overexposure, and high‑speed motion. Recognizing the current scarcity of dedicated datasets in this domain, we introduce EPRBench, a high‑quality benchmark specifically designed for event stream‑based VPR. EPRBench comprises 10K event sequences and 65K event frames, collected using both handheld and vehicle‑mounted setups to comprehensively capture real‑world challenges across diverse viewpoints, weather conditions, and lighting scenarios. To support semantic‑aware and language‑integrated VPR research, we provide LLM‑generated scene descriptions, subsequently refined through human annotation, establishing a solid foundation for integrating LLMs into event‑based perception pipelines. To facilitate systematic evaluation, we implement and benchmark 15 state‑of‑the‑art VPR algorithms on EPRBench, offering a strong baseline for future algorithmic comparisons. Furthermore, we propose a novel multi‑modal fusion paradigm for VPR: leveraging LLMs to generate textual scene descriptions from raw event streams, which then guide spatially attentive token selection, cross‑modal feature fusion, and multi‑scale representation learning. This framework not only achieves highly accurate place recognition but also produces interpretable reasoning processes alongside its predictions, significantly enhancing model transparency and explainability. The dataset and source code will be released on https://github.com/Event‑AHU/Neuromorphic_ReID
Authors:Yunshuang Nie, Bingqian Lin, Minzhe Niu, Kun Xiang, Jianhua Han, Guowei Huang, Xingyue Quan, Hang Xu, Bokui Chen, Xiaodan Liang
Abstract:
Pre‑trained Multi‑modal Large Language Models (MLLMs) provide a knowledge‑rich foundation for post‑training by leveraging their inherent perception and reasoning capabilities to solve complex tasks. However, the lack of an efficient evaluation framework impedes the diagnosis of their performance bottlenecks. Current evaluation primarily relies on testing after supervised fine‑tuning, which introduces laborious additional training and autoregressive decoding costs. Meanwhile, common pre‑training metrics cannot quantify a model's perception and reasoning abilities in a disentangled manner. Furthermore, existing evaluation benchmarks are typically limited in scale or misaligned with pre‑training objectives. Thus, we propose RADAR, an efficient ability‑centric evaluation framework for Revealing Asymmetric Development of Abilities in MLLM pRe‑training. RADAR involves two key components: (1) Soft Discrimination Score, a novel metric for robustly tracking ability development without fine‑tuning, based on quantifying nuanced gradations of the model preference for the correct answer over distractors; and (2) Multi‑Modal Mixture Benchmark, a new 15K+ sample benchmark for comprehensively evaluating pre‑trained MLLMs' perception and reasoning abilities in a 0‑shot manner, where we unify authoritative benchmark datasets and carefully collect new datasets, extending the evaluation scope and addressing the critical gaps in current benchmarks. With RADAR, we comprehensively reveal the asymmetric development of perceptual and reasoning capabilities in pretrained MLLMs across diverse factors, including data volume, model size, and pretraining strategy. Our RADAR underscores the need for a decomposed perspective on pre‑training ability bottlenecks, informing targeted interventions to advance MLLMs efficiently. Our code is publicly available at https://github.com/Nieysh/RADAR.
Authors:Mehran Advand, Zahra Dehghanian, Navid Faraji, Reza Barati, Seyed Amir Ahmad Safavi-Naini, Hamid R. Rabiee
Abstract:
Existing medical imaging datasets for abdominal CT often lack three‑dimensional annotations, multi‑organ coverage, or precise lesion‑to‑organ associations, hindering robust representation learning and clinical applications. To address this gap, we introduce 3DLAND, a large‑scale benchmark dataset comprising over 6,000 contrast‑enhanced CT volumes with over 20,000 high‑fidelity 3D lesion annotations linked to seven abdominal organs: liver, kidneys, pancreas, spleen, stomach, and gallbladder. Our streamlined three‑phase pipeline integrates automated spatial reasoning, prompt‑optimized 2D segmentation, and memory‑guided 3D propagation, validated by expert radiologists with surface dice scores exceeding 0.75. By providing diverse lesion types and patient demographics, 3DLAND enables scalable evaluation of anomaly detection, localization, and cross‑organ transfer learning for medical AI. Our dataset establishes a new benchmark for evaluating organ‑aware 3D segmentation models, paving the way for advancements in healthcare‑oriented AI. To facilitate reproducibility and further research, the 3DLAND dataset and implementation code are publicly available at https://mehrn79.github.io/3DLAND.
Authors:Xiao Ren, Yu Liu, Ning An, Jian Cheng, Xin Qiao, He Kong
Abstract:
Recently, 3D Gaussian Splatting has emerged as a prominent research direction owing to its ultrarapid training speed and high‑fidelity rendering capabilities. However, the unstructured and irregular nature of Gaussian point clouds poses challenges to reconstruction accuracy. This limitation frequently causes high‑frequency detail loss in complex surface microstructures when relying solely on routine strategies. To address this limitation, we propose GSM‑GS: a synergistic optimization framework integrating single‑view adaptive sub‑region weighting constraints and multi‑view spatial structure refinement. For single‑view optimization, we leverage image gradient features to partition scenes into texture‑rich and texture‑less sub‑regions. The reconstruction quality is enhanced through adaptive filtering mechanisms guided by depth discrepancy features. This preserves high‑weight regions while implementing a dual‑branch constraint strategy tailored to regional texture variations, thereby improving geometric detail characterization. For multi‑view optimization, we introduce a geometry‑guided cross‑view point cloud association method combined with a dynamic weight sampling strategy. This constructs 3D structural normal constraints across adjacent point cloud frames, effectively reinforcing multi‑view consistency and reconstruction fidelity. Extensive experiments on public datasets demonstrate that our method achieves both competitive rendering quality and geometric reconstruction. See our interactive project page
Authors:Xiaowen Zhang, Zijie Yue, Yong Luo, Cairong Zhao, Qijun Chen, Miaojing Shi
Abstract:
Object counting is a fundamental task in computer vision, with broad applicability in many real‑world scenarios. Fully‑supervised counting methods require costly point‑level annotations per object. Few weakly‑supervised methods leverage only image‑level object counts as supervision and achieve fairly promising results. They are, however, often limited to counting a single category, e.g. person. In this paper, we propose WS‑COC, the first MLLM‑driven weakly‑supervised framework for class‑agnostic object counting. Instead of directly fine‑tuning MLLMs to predict object counts, which can be challenging due to the modality gap, we incorporate three simple yet effective strategies to bootstrap the counting paradigm in both training and testing: First, a divide‑and‑discern dialogue tuning strategy is proposed to guide the MLLM to determine whether the object count falls within a specific range and progressively break down the range through multi‑round dialogue. Second, a compare‑and‑rank count optimization strategy is introduced to train the MLLM to optimize the relative ranking of multiple images according to their object counts. Third, a global‑and‑local counting enhancement strategy aggregates and fuses local and global count predictions to improve counting performance in dense scenes. Extensive experiments on FSC‑147, CARPK, PUCPR+, and ShanghaiTech show that WS‑COC matches or even surpasses many state‑of‑art fully‑supervised methods while significantly reducing annotation costs. Code is available at https://github.com/viscom‑tongji/WS‑COC.
Authors:Ruipeng Wang, Langkun Zhong, Miaowei Wang
Abstract:
State‑of‑the‑art rigging methods typically assume a predefined canonical rest pose. However, this assumption does not hold for dynamic mesh sequences such as DyMesh or DT4D, where no canonical T‑pose is available. When applied independently frame‑by‑frame, existing methods lack pose invariance and often yield temporally inconsistent topologies. To address this limitation, we propose SPRig, a general fine‑tuning framework that enforces cross‑frame consistency across a sequence to learn pose‑invariant rigs on top of existing models, covering both skeleton and skinning generation. For skeleton generation, we introduce novel consistency regularization in both token space and geometry space. For skinning, we improve temporal stability through an articulation‑invariant consistency loss combined with consistency distillation and structural regularization. Extensive experiments show that SPRig achieves superior temporal coherence and significantly reduces artifacts in prior methods, without sacrificing and often even enhancing per‑frame static generation quality. The code is available in the supplemental material and will be made publicly available upon publication.
Authors:Qiuchen Wang, Shihang Wang, Yu Zeng, Qiang Zhang, Fanrui Zhang, Zhuoning Guo, Bosi Zhang, Wenxuan Huang, Lin Chen, Zehui Chen, Pengjun Xie, Ruixue Ding
Abstract:
Effectively retrieving, reasoning, and understanding multimodal information remains a critical challenge for agentic systems. Traditional Retrieval‑augmented Generation (RAG) methods rely on linear interaction histories, which struggle to handle long‑context tasks, especially those involving information‑sparse yet token‑heavy visual data in iterative reasoning scenarios. To bridge this gap, we introduce VimRAG, a framework tailored for multimodal Retrieval‑augmented Reasoning across text, images, and videos. Inspired by our systematic study, we model the reasoning process as a dynamic directed acyclic graph that structures the agent states and retrieved multimodal evidence. Building upon this structured memory, we introduce a Graph‑Modulated Visual Memory Encoding mechanism, with which the significance of memory nodes is evaluated via their topological position, allowing the model to dynamically allocate high‑resolution tokens to pivotal evidence while compressing or discarding trivial clues. To implement this paradigm, we propose a Graph‑Guided Policy Optimization strategy. This strategy disentangles step‑wise validity from trajectory‑level rewards by pruning memory nodes associated with redundant actions, thereby facilitating fine‑grained credit assignment. Extensive experiments demonstrate that VimRAG consistently achieves state‑of‑the‑art performance on diverse multimodal RAG benchmarks. The code is available at https://github.com/Alibaba‑NLP/VRAG.
Authors:Umar Marikkar, Syed Sameed Husain, Muhammad Awais, Sara Atito
Abstract:
Training and evaluating vision encoders on Multi‑Channel Imaging (MCI) data remains challenging as channel configurations vary across datasets, preventing fixed‑channel training and limiting reuse of pre‑trained encoders on new channel settings. Prior work trains MCI encoders but typically evaluates them via full fine‑tuning, leaving probing with frozen pre‑trained encoders comparatively underexplored. Existing studies that perform probing largely focus on improving representations, rather than how to best leverage fixed representations for downstream tasks. Although the latter problem has been studied in other domains, directly transferring those strategies to MCI yields weak results, even worse than training from scratch. We therefore propose Channel‑Aware Probing (CAP), which exploits the intrinsic inter‑channel diversity in MCI datasets by controlling feature flow at both the encoder and probe levels. CAP uses Independent Feature Encoding (IFE) to encode each channel separately, and Decoupled Pooling (DCP) to pool within channels before aggregating across channels. Across three MCI benchmarks, CAP consistently improves probing performance over the default probing protocol, matches fine‑tuning from scratch, and largely reduces the gap to full fine‑tuning from the same MCI pre‑trained checkpoints. Code can be found in https://github.com/umarikkar/CAP.
Authors:Wooseok Jeon, Seunghyun Shin, Dongmin Shin, Hae-Gon Jeon
Abstract:
Recent progress in image‑to‑video (I2V) diffusion models has significantly advanced the field of generative inbetweening, which aims to generate semantically plausible frames between two keyframes. In particular, inference‑time sampling strategies, which leverage the generative priors of large‑scale pre‑trained I2V models without additional training, have become increasingly popular. However, existing inference‑time sampling, either fusing forward and backward paths in parallel or alternating them sequentially, often suffers from temporal discontinuities and undesirable visual artifacts due to the misalignment between the two generated paths. This is because each path follows the motion prior induced by its own conditioning frame. In this work, we propose Motion Prior Distillation (MPD), a simple yet effective inference‑time distillation technique that suppresses bidirectional mismatch by distilling the motion residual of the forward path into the backward path. Our method can deliberately avoid denoising the end‑conditioned path which causes the ambiguity of the path, and yield more temporally coherent inbetweening results with the forward motion prior. We not only perform quantitative evaluations on standard benchmarks, but also conduct extensive user studies to demonstrate the effectiveness of our approach in practical scenarios.
Authors:Marco Stricker, Masakazu Iwamura, Koichi Kise
Abstract:
Clouds are a common phenomenon that distorts optical satellite imagery, which poses a challenge for remote sensing. However, in the literature cloudless analysis is often performed where cloudy images are excluded from machine learning datasets and methods. Such an approach cannot be applied to time sensitive applications, e.g., during natural disasters. A possible solution is to apply cloud removal as a preprocessing step to ensure that cloudfree solutions are not failing under such conditions. But cloud removal methods are still actively researched and suffer from drawbacks, such as generated visual artifacts. Therefore, it is desirable to develop cloud robust methods that are less affected by cloudy weather. Cloud robust methods can be achieved by combining optical data with radar, a modality unaffected by clouds. While many datasets for machine learning combine optical and radar data, most researchers exclude cloudy images. We identify this exclusion from machine learning training and evaluation as a limitation that reduces applicability to cloudy scenarios. To investigate this, we assembled a dataset, named CloudyBigEarthNet (CBEN), of paired optical and radar images with cloud occlusion for training and evaluation. Using average precision (AP) as the evaluation metric, we show that state‑of‑the‑art methods trained on combined clear‑sky optical and radar imagery suffer performance drops of 23‑33 percentage points when evaluated on cloudy images. We then adapt these methods to cloudy optical data during training, achieving relative improvement of 17.2‑28.7 percentage points on cloudy test cases compared with the original approaches. Code and dataset are publicly available at: https://github.com/mstricker13/CBEN
Authors:Sangwoo Jo, Sungjoon Choi
Abstract:
Diffusion‑based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by high sampling costs. While prior work focuses on training objectives or individual solvers, the holistic design of sampling, specifically solver selection and scheduling, remains dominated by static heuristics. In this work, we revisit this challenge through a geometric lens, proposing SDM, a principled framework that aligns the numerical solver with the intrinsic properties of the diffusion trajectory. By analyzing the ODE dynamics, we show that efficient low‑order solvers suffice in early high‑noise stages while higher‑order solvers can be progressively deployed to handle the increasing non‑linearity of later stages. Furthermore, we formalize the scheduling by introducing a Wasserstein‑bounded optimization framework. This method systematically derives adaptive timesteps that explicitly bound the local discretization error, ensuring the sampling process remains faithful to the underlying continuous dynamics. Without requiring additional training or architectural modifications, SDM achieves state‑of‑the‑art performance across standard benchmarks, including an FID of 1.93 on CIFAR‑10, 2.41 on FFHQ, and 1.98 on AFHQv2, with a reduced number of function evaluations compared to existing samplers. Our code is available at https://github.com/aiimaginglab/sdm.
Authors:Ke Xu, Yixin Wang, Zhongcheng Li, Hao Cui, Jinshui Hu, Xingyi Zhang
Abstract:
Elastic precision quantization enables multi‑bit deployment via a single optimization pass, fitting diverse quantization scenarios.Yet, the high storage and optimization costs associated with the Transformer architecture, research on elastic quantization remains limited, particularly for large language models.This paper proposes QuEPT, an efficient post‑training scheme that reconstructs block‑wise multi‑bit errors with one‑shot calibration on a small data slice. It can dynamically adapt to various predefined bit‑widths by cascading different low‑rank adapters, and supports real‑time switching between uniform quantization and mixed precision quantization without repeated optimization. To enhance accuracy and robustness, we introduce Multi‑Bit Token Merging (MB‑ToMe) to dynamically fuse token features across different bit‑widths, improving robustness during bit‑width switching. Additionally, we propose Multi‑Bit Cascaded Low‑Rank adapters (MB‑CLoRA) to strengthen correlations between bit‑width groups, further improve the overall performance of QuEPT. Extensive experiments demonstrate that QuEPT achieves comparable or better performance to existing state‑of‑the‑art post‑training quantization methods.Our code is available at https://github.com/xuke225/QuEPT
Authors:Jinze Chen, Wei Zhai, Han Han, Tiankai Ma, Yang Cao, Bin Li, Zheng-Jun Zha
Abstract:
Event‑based vision encodes dynamic scenes as asynchronous spatio‑temporal spikes called events. To leverage conventional image processing pipelines, events are typically binned into frames. However, binning functions are discontinuous, which truncates gradients at the frame level and forces most event‑based algorithms to rely solely on frame‑based features. Attempts to directly learn from raw events avoid this restriction but instead suffer from biased gradient estimation due to the discontinuities of the binning operation, ultimately limiting their learning efficiency. To address this challenge, we propose a novel framework for unbiased gradient estimation of arbitrary binning functions by synthesizing weak derivatives during backpropagation while keeping the forward output unchanged. The key idea is to exploit integration by parts: lifting the target functions to functionals yields an integral form of the derivative of the binning function during backpropagation, where the cotangent function naturally arises. By reconstructing this cotangent function from the sampled cotangent vector, we compute weak derivatives that provably match long‑range finite differences of both smooth and non‑smooth targets. Experimentally, our method improves simple optimization‑based egomotion estimation with 3.2% lower RMS error and 1.57× faster convergence. On complex downstream tasks, we achieve 9.4% lower EPE in self‑supervised optical flow, and 5.1% lower RMS error in SLAM, demonstrating broad benefits for event‑based visual perception. Source code can be found at https://github.com/chjz1024/EventFBP.
Authors:Bowen Ping, Chengyou Jia, Minnan Luo, Hangwei Qian, Ivor Tsang
Abstract:
Reinforcement learning has emerged as a promising paradigm for aligning diffusion and flow‑matching models with human preferences, yet practitioners face fragmented codebases, model‑specific implementations, and engineering complexity. We introduce Flow‑Factory, a unified framework that decouples algorithms, models, and rewards through through a modular, registry‑based architecture. This design enables seamless integration of new algorithms and architectures, as demonstrated by our support for GRPO, DiffusionNFT, and AWM across Flux, Qwen‑Image, and WAN video models. By minimizing implementation overhead, Flow‑Factory empowers researchers to rapidly prototype and scale future innovations with ease. Flow‑Factory provides production‑ready memory optimization, flexible multi‑reward training, and seamless distributed training support. The codebase is available at https://github.com/X‑GenGroup/Flow‑Factory.
Authors:Ara Yeroyan
Abstract:
Multi‑vector visual retrievers (e.g., ColPali‑style late interaction models) deliver strong accuracy, but scale poorly because each page yields thousands of vectors, making indexing and search increasingly expensive. We present Visual RAG Toolkit, a practical system for scaling visual multi‑vector retrieval with training‑free, model‑aware pooling and multi‑stage retrieval. Motivated by Matryoshka Embeddings, our method performs static spatial pooling ‑ including a lightweight sliding‑window averaging variant ‑ over patch embeddings to produce compact tile‑level and global representations for fast candidate generation, followed by exact MaxSim reranking using full multi‑vector embeddings.
Our design yields a quadratic reduction in vector‑to‑vector comparisons by reducing stored vectors per page from thousands to dozens, notably without requiring post‑training, adapters, or distillation. Across experiments with interaction‑style models such as ColPali and ColSmol‑500M, we observe that over the limited ViDoRe v2 benchmark corpus 2‑stage retrieval typically preserves NDCG and Recall @ 5/10 with minimal degradation, while substantially improving throughput (approximately 4x QPS); with sensitivity mainly at very large k. The toolkit additionally provides robust preprocessing ‑ high resolution PDF to image conversion, optional margin/empty‑region cropping and token hygiene (indexing only visual tokens) ‑ and a reproducible evaluation pipeline, enabling rapid exploration of two‑, three‑, and cascaded retrieval variants. By emphasizing efficiency at common cutoffs (e.g., k <= 10), the toolkit lowers hardware barriers and makes state‑of‑the‑art visual retrieval more accessible in practice.
Authors:Ali Abbasi, Mehdi Taghipour, Rahmatollah Beheshti
Abstract:
Negation is a fundamental linguistic operation in clinical reporting, yet vision‑language models (VLMs) frequently fail to distinguish affirmative from negated medical statements. To systematically characterize this limitation, we introduce a radiology‑specific diagnostic benchmark that evaluates polarity sensitivity under controlled clinical conditions, revealing that common medical VLMs consistently confuse negated and non‑negated findings. To enable learning beyond simple condition absence, we further construct a contextual clinical negation dataset that encodes structured claims and supports attribute‑level negations involving location and severity. Building on these resources, we propose Negation‑Aware Selective Training (NAST), an interpretability‑guided adaptation method that uses causal tracing effects (CTEs) to modulate layer‑wise gradient updates during fine‑tuning. Rather than applying uniform learning rates, NAST scales each layer's update according to its causal contribution to negation processing, transforming mechanistic interpretability signals into a principled optimization rule. Experiments demonstrate improved discrimination of affirmative and negated clinical statements without degrading general vision‑language alignment, highlighting the value of causal interpretability for targeted model adaptation in safety‑critical medical settings. Code and resources are available at https://github.com/healthylaife/NAST.
Authors:Jiacheng Zhang, Jinhao Li, Hanxun Huang, Sarah M. Erfani, Benjamin I. P. Rubinstein, Feng Liu
Abstract:
Recent studies have shown that CLIP model's adversarial robustness in zero‑shot classification tasks can be enhanced by adversarially fine‑tuning its image encoder with adversarial examples (AEs), which are generated by minimizing the cosine similarity between images and a hand‑crafted template (e.g., ''A photo of a label''). However, it has been shown that the cosine similarity between a single image and a single hand‑crafted template is insufficient to measure the similarity for image‑text pairs. Building on this, in this paper, we find that the AEs generated using cosine similarity may fail to fool CLIP when the similarity metric is replaced with semantically enriched alternatives, making the image encoder fine‑tuned with these AEs less robust. To overcome this issue, we first propose a semantic‑ensemble attack to generate semantic‑aware AEs by minimizing the average similarity between the original image and an ensemble of refined textual descriptions. These descriptions are initially generated by a foundation model to capture core semantic features beyond hand‑crafted templates and are then refined to reduce hallucinations. To this end, we propose Semantic‑aware Adversarial Fine‑Tuning (SAFT), which fine‑tunes CLIP's image encoder with semantic‑aware AEs. Extensive experiments show that SAFT outperforms current methods, achieving substantial improvements in zero‑shot adversarial robustness across 16 datasets. Our code is available at: https://github.com/tmlr‑group/SAFT.
Authors:Ali Nasiri-Sarvi, Anh Tien Nguyen, Hassan Rivaz, Dimitris Samaras, Mahdi S. Hosseini
Abstract:
Sparse autoencoders (SAEs) decompose polysemantic neural representations, where neurons respond to multiple unrelated concepts, into monosemantic features that capture single, interpretable concepts. However, standard training objectives only weakly encourage this decomposition, and existing monosemanticity metrics require pairwise comparisons across all dataset samples, making them inefficient during training and evaluation. We study a recent MonoScore metric and derive a single‑pass algorithm that computes exactly the same quantity, but with a cost that grows linearly, rather than quadratically, with the number of dataset images. On OpenImagesV7, we achieve up to a 1200x speedup wall‑clock speedup in evaluation and 159x during training, while adding only ~4% per‑epoch overhead. This allows us to treat MonoScore as a training signal: we introduce the Monosemanticity Loss (MonoLoss), a plug‑in objective that directly rewards semantically consistent activations for learning interpretable monosemantic representations. Across SAEs trained on CLIP, SigLIP2, and pretrained ViT features, using BatchTopK, TopK, and JumpReLU SAEs, MonoLoss increases MonoScore for most latents. MonoLoss also consistently improves class purity (the fraction of a latent's activating images belonging to its dominant class) across all encoder and SAE combinations, with the largest gain raising baseline purity from 0.152 to 0.723. Used as an auxiliary regularizer during ResNet‑50 and CLIP‑ViT‑B/32 finetuning, MonoLoss yields up to 0.6% accuracy gains on ImageNet‑1K and monosemantic activating patterns on standard benchmark datasets. The code is publicly available at https://github.com/AtlasAnalyticsLab/MonoLoss.
Authors:Ali Subhan, Ashir Raza
Abstract:
DragDiffusion is a diffusion‑based method for interactive point‑based image editing that enables users to manipulate images by directly dragging selected points. The method claims that accurate spatial control can be achieved by optimizing a single diffusion latent at an intermediate timestep, together with identity‑preserving fine‑tuning and spatial regularization. This work presents a reproducibility study of DragDiffusion using the authors' released implementation and the DragBench benchmark. We reproduce the main ablation studies on diffusion timestep selection, LoRA‑based fine‑tuning, mask regularization strength, and UNet feature supervision, and observe close agreement with the qualitative and quantitative trends reported in the original work. At the same time, our experiments show that performance is sensitive to a small number of hyperparameter assumptions, particularly the optimized timestep and the feature level used for motion supervision, while other components admit broader operating ranges. We further evaluate a multi‑timestep latent optimization variant and find that it does not improve spatial accuracy while substantially increasing computational cost. Overall, our findings support the central claims of DragDiffusion while clarifying the conditions under which they are reliably reproducible. Code is available at https://github.com/AliSubhan5341/DragDiffusion‑TMLR‑Reproducibility‑Challenge.
Authors:Zekun Li, Sizhe An, Chengcheng Tang, Chuan Guo, Ivan Shugurov, Linguang Zhang, Amy Zhao, Srinath Sridhar, Lingling Tao, Abhay Mittal
Abstract:
Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion‑language generation and understanding remains largely underexplored. Existing approaches often fine‑tune large language models (LLMs) on paired motion‑text data, which can result in catastrophic forgetting of linguistic capabilities due to the limited scale of available text‑motion pairs. Furthermore, prior methods typically convert motion into discrete representations via quantization to integrate with language models, introducing substantial jitter artifacts from discrete tokenization. To address these challenges, we propose LLaMo, a unified framework that extends pretrained LLMs through a modality‑specific Mixture‑of‑Transformers (MoT) architecture. This design inherently preserves the language understanding of the base model while enabling scalable multimodal adaptation. We encode human motion into a causal continuous latent space and maintain the next‑token prediction paradigm in the decoder‑only backbone through a lightweight flow‑matching head, allowing for streaming motion generation in real‑time (>30 FPS). Leveraging the comprehensive language understanding of pretrained LLMs and large‑scale motion‑text pretraining, our experiments demonstrate that LLaMo achieves high‑fidelity text‑to‑motion generation and motion‑to‑text captioning in general settings, especially zero‑shot motion generation, marking a significant step towards a general unified motion‑language large model.
Authors:Junwoon Lee, Yulun Tian
Abstract:
We present LatentAM, an online 3D Gaussian Splatting (3DGS) mapping framework that builds scalable latent feature maps from streaming RGB‑D observations for open‑vocabulary robotic perception. Instead of distilling high‑dimensional Vision‑Language Model (VLM) embeddings using model‑specific decoders, LatentAM proposes an online dictionary learning approach that is both model‑agnostic and pretraining‑free, enabling plug‑and‑play integration with different VLMs at test time. Specifically, our approach associates each Gaussian primitive with a compact query vector that can be converted into approximate VLM embeddings using an attention mechanism with a learnable dictionary. The dictionary is initialized efficiently from streaming observations and optimized online to adapt to evolving scene semantics under trust‑region regularization. To scale to long trajectories and large environments, we further propose an efficient map management strategy based on voxel hashing, where optimization is restricted to an active local map on the GPU, while the global map is stored and indexed on the CPU to maintain bounded GPU memory usage. Experiments on public benchmarks and a large‑scale custom dataset demonstrate that LatentAM attains significantly better feature reconstruction fidelity compared to state‑of‑the‑art methods, while achieving near‑real‑time speed (12‑35 FPS) on the evaluated datasets. Our project page is at: https://junwoonlee.github.io/projects/LatentAM
Authors:Neemias da Silva, Júlio C. W. Scholz, John Harrison, Marina Borges, Paulo Ávila, Frances A Santos, Myriam Delgado, Rodrigo Minetto, Thiago H Silva
Abstract:
Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs‑Teoria‑e‑Pratica. Finally, the chapter discusses the challenges and highlights promising trends.
Authors:Miaosen Zhang, Yishan Liu, Shuxia Lin, Xu Yang, Qi Dai, Chong Luo, Weihao Jiang, Peng Hou, Anxiang Zeng, Xin Geng, Baining Guo
Abstract:
Supervised fine‑tuning (SFT) is computationally efficient but often yields inferior generalization compared to reinforcement learning (RL). This gap is primarily driven by RL's use of on‑policy data. We propose a framework to bridge this chasm by enabling On‑Policy SFT. We first present Distribution Discriminant Theory (DDT), which explains and quantifies the alignment between data and the model‑induced distribution. Leveraging DDT, we introduce two complementary techniques: (i) In‑Distribution Finetuning (IDFT), a loss‑level method to enhance generalization ability of SFT, and (ii) Hinted Decoding, a data‑level technique that can re‑align the training corpus to the model's distribution. Extensive experiments demonstrate that our framework achieves generalization performance surpassing prominent offline RL algorithms, including DPO and SimPO, while maintaining the efficiency of an SFT pipeline. The proposed framework thus offers a practical alternative in domains where RL is infeasible. We open‑source the code here: https://github.com/zhangmiaosen2000/Towards‑On‑Policy‑SFT
Authors:Xu Guo, Fulong Ye, Qichao Sun, Liyang Chen, Bingchuan Li, Pengze Zhang, Jiawei Liu, Songtao Zhao, Qian He, Xiangwang Hou
Abstract:
Recent advancements in foundation models have revolutionized joint audio‑video generation. However, existing approaches typically treat human‑centric tasks including reference‑based audio‑video generation (R2AV), video editing (RV2AV) and audio‑driven video animation (RA2V) as isolated objectives. Furthermore, achieving precise, disentangled control over multiple character identities and voice timbres within a single framework remains an open challenge. In this paper, we propose DreamID‑Omni, a unified framework for controllable human‑centric audio‑video generation. Specifically, we design a Symmetric Conditional Diffusion Transformer that integrates heterogeneous conditioning signals via a symmetric conditional injection scheme. To resolve the pervasive identity‑timbre binding failures and speaker confusion in multi‑person scenarios, we introduce a Dual‑Level Disentanglement strategy: Synchronized RoPE at the signal level to ensure rigid attention‑space binding, and Structured Captions at the semantic level to establish explicit attribute‑subject mappings. Furthermore, we devise a Multi‑Task Progressive Training scheme that leverages weakly‑constrained generative priors to regularize strongly‑constrained tasks, preventing overfitting and harmonizing disparate objectives. Extensive experiments demonstrate that DreamID‑Omni achieves comprehensive state‑of‑the‑art performance across video, audio, and audio‑visual consistency, even outperforming leading proprietary commercial models. We will release our code to bridge the gap between academic research and commercial‑grade applications.
Authors:Ziteng Lu, Yushuang Wu, Chongjie Ye, Yuda Qiu, Jing Shao, Xiaoyang Guo, Jiaqing Zhou, Tianlei Hu, Kun Zhou, Xiaoguang Han
Abstract:
High‑quality 3D texture generation remains a fundamental challenge due to the view‑inconsistency inherent in current mainstream multi‑view diffusion pipelines. Existing representations either rely on UV maps, which suffer from distortion during unwrapping, or point‑based methods, which tightly couple texture fidelity to geometric density that limits high‑resolution texture generation. To address these limitations, we introduce TexSpot, a diffusion‑based texture enhancement framework. At its core is Texlet, a novel 3D texture representation that merges the geometric expressiveness of point‑based 3D textures with the compactness of UV‑based representation. Each Texlet latent vector encodes a local texture patch via a 2D encoder and is further aggregated using a 3D encoder to incorporate global shape context. A cascaded 3D‑to‑2D decoder reconstructs high‑quality texture patches, enabling the Texlet space learning. Leveraging this representation, we train a diffusion transformer conditioned on Texlets to refine and enhance textures produced by multi‑view diffusion methods. Extensive experiments demonstrate that TexSpot significantly improves visual fidelity, geometric consistency, and robustness over existing state‑of‑the‑art 3D texture generation and enhancement approaches. Project page: https://texlet‑arch.github.io/TexSpot‑page.
Authors:Yeyao Ma, Chen Li, Xiaosong Zhang, Han Hu, Weidi Xie
Abstract:
Post‑training of flow matching models‑aligning the output distribution with a high‑quality target‑is mathematically equivalent to imitation learning. While Supervised Fine‑Tuning mimics expert demonstrations effectively, it cannot correct policy drift in unseen states. Preference optimization methods address this but require costly preference pairs or reward modeling. We propose Flow Matching Adversarial Imitation Learning (FAIL), which minimizes policy‑expert divergence through adversarial training without explicit rewards or pairwise comparisons. We derive two algorithms: FAIL‑PD exploits differentiable ODE solvers for low‑variance pathwise gradients, while FAIL‑PG provides a black‑box alternative for discrete or computationally constrained settings. Fine‑tuning FLUX with only 13,000 demonstrations from Nano Banana pro, FAIL achieves competitive performance on prompt following and aesthetic benchmarks. Furthermore, the framework generalizes effectively to discrete image and video generation, and functions as a robust regularizer to mitigate reward hacking in reward‑based optimization. Code and data are available at https://github.com/HansPolo113/FAIL.
Authors:Lingting Zhu, Shengju Qian, Haidi Fan, Jiayu Dong, Zhenchao Jin, Siwei Zhou, Gen Dong, Xin Wang, Lequan Yu
Abstract:
The digital industry demands high‑quality, diverse modular 3D assets, especially for user‑generated content~(UGC). In this work, we introduce AssetFormer, an autoregressive Transformer‑based model designed to generate modular 3D assets from textual descriptions. Our pilot study leverages real‑world modular assets collected from online platforms. AssetFormer tackles the challenge of creating assets composed of primitives that adhere to constrained design parameters for various applications. By innovatively adapting module sequencing and decoding techniques inspired by language models, our approach enhances asset generation quality through autoregressive modeling. Initial results indicate the effectiveness of AssetFormer in streamlining asset creation for professional development and UGC scenarios. This work presents a flexible framework extendable to various types of modular 3D assets, contributing to the broader field of 3D content generation. The code is available at https://github.com/Advocate99/AssetFormer.
Authors:GigaBrain Team, Boyuan Wang, Bohan Li, Chaojun Ni, Guan Huang, Guosheng Zhao, Hao Li, Jie Li, Jindi Lv, Jingyu Liu, Lv Feng, Mingming Yu, Peng Li, Qiuping Deng, Tianze Liu, Xinyu Zhou, Xinze Chen, Xiaofeng Wang, Yang Wang, Yifan Li, Yifei Nie, Yilong Li, Yukun Zhou, Yun Ye, Zhichao Liu, Zheng Zhu
Abstract:
Vision‑language‑action (VLA) models that directly predict multi‑step action chunks from current observations face inherent limitations due to constrained scene understanding and weak future anticipation capabilities. In contrast, video world models pre‑trained on web‑scale video corpora exhibit robust spatiotemporal reasoning and accurate future prediction, making them a natural foundation for enhancing VLA learning. Therefore, we propose GigaBrain‑0.5M, a VLA model trained via world model‑based reinforcement learning. Built upon GigaBrain‑0.5, which is pre‑trained on over 10,000 hours of robotic manipulation data, whose intermediate version currently ranks first on the international RoboChallenge benchmark. GigaBrain‑0.5M further integrates world model‑based reinforcement learning via RAMP (Reinforcement leArning via world Model‑conditioned Policy) to enable robust cross‑task adaptation. Empirical results demonstrate that RAMP achieves substantial performance gains over the RECAP baseline, yielding improvements of approximately 30% on challenging tasks including \textttLaundry Folding, \textttBox Packing, and \textttEspresso Preparation. Critically, GigaBrain‑0.5M^ exhibits reliable long‑horizon execution, consistently accomplishing complex manipulation tasks without failure as validated by real‑world deployment videos on our \hrefhttps://gigabrain05m.github.ioproject page.
Authors:Bingxu Xie, Fang Zhou, Jincan Wu, Yonghui Liu, Weiqing Li, Zhiyong Su
Abstract:
While no‑reference point cloud quality assessment (NR‑PCQA) approaches have achieved significant progress over the past decade, their performance often degrades substantially when a distribution gap exists between the training (source domain) and testing (target domain) data. However, to date, limited attention has been paid to transferring NR‑PCQA models across domains. To address this challenge, we propose the first unsupervised progressive domain adaptation (UPDA) framework for NR‑PCQA, which introduces a two‑stage coarse‑to‑fine alignment paradigm to address domain shifts. At the coarse‑grained stage, a discrepancy‑aware coarse‑grained alignment method is designed to capture relative quality relationships between cross‑domain samples through a novel quality‑discrepancy‑aware hybrid loss, circumventing the challenges of direct absolute feature alignment. At the fine‑grained stage, a perception fusion fine‑grained alignment approach with symmetric feature fusion is developed to identify domain‑invariant features, while a conditional discriminator selectively enhances the transfer of quality‑relevant features. Extensive experiments demonstrate that the proposed UPDA effectively enhances the performance of NR‑PCQA methods in cross‑domain scenarios, validating its practical applicability. The code is available at https://github.com/yokeno1/UPDA‑main.
Authors:Suraj Ranganath, Anish Patnaik, Vaishak Menon
Abstract:
Efficient spatial reasoning requires world models that remain reliable under tight precision budgets. We study whether low‑bit planning behavior is determined mostly by total bitwidth or by where bits are allocated across modules. Using DINO‑WM on the Wall planning task, we run a paired‑goal mixed‑bit evaluation across uniform, mixed, asymmetric, and layerwise variants under two planner budgets. We observe a consistent three‑regime pattern: 8‑bit and 6‑bit settings remain close to FP16, 3‑bit settings collapse, and 4‑bit settings are allocation‑sensitive. In that transition region, preserving encoder precision improves planning relative to uniform quantization, and near‑size asymmetric variants show the same encoder‑side direction. In a later strict 22‑cell replication with smaller per‑cell episode count, the mixed‑versus‑uniform INT4 sign becomes budget‑conditioned, which further highlights the sensitivity of this transition regime. These findings motivate module‑aware, budget‑aware quantization policies as a broader research direction for efficient spatial reasoning. Code and run artifacts are available at https://github.com/suraj‑ranganath/DINO‑MBQuant.
Authors:Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong, Yutong Cai, Siyuan Li, Huijia Zhu, Weiqiang Wang, Linghe Kong, Yue Wang, Zhuosheng Zhang, Weiran Huang
Abstract:
Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine‑grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking‑with‑Images" methods alleviate this by iteratively zooming in and out regions of interest during inference, but incur high latency due to repeated tool calls and visual re‑encoding. To address this, we propose Region‑to‑Image Distillation, which transforms zooming from an inference‑time tool into a training‑time primitive, thereby internalizing the benefits of agentic zooming into a single forward pass of an MLLM. In particular, we first zoom in to micro‑cropped regions to let strong teacher models generate high‑quality VQA data, and then distill this region‑grounded supervision back to the full image. After training on such data, the smaller student model improves "single‑glance" fine‑grained perception without tool use. To rigorously evaluate this capability, we further present ZoomBench, a hybrid‑annotated benchmark of 845 VQA data spanning six fine‑grained perceptual dimensions, together with a dual‑view protocol that quantifies the global‑‑regional "zooming gap". Experiments show that our models achieve leading performance across multiple fine‑grained perception benchmarks, and also improve general multimodal cognition on benchmarks such as visual reasoning and GUI agents. We further discuss when "Thinking‑with‑Images" is necessary versus when its gains can be distilled into a single forward pass. Our code is available at https://github.com/inclusionAI/Zooming‑without‑Zooming.
Authors:Qisen Wang, Yifan Zhao, Jia Li
Abstract:
Dynamic reconstruction has achieved remarkable progress, but there remain challenges in monocular input for more practical applications. The prevailing works attempt to construct efficient motion representations, but lack a unified spatiotemporal decomposition framework, suffering from either holistic temporal optimization or coupled hierarchical spatial composition. To this end, we propose WorldTree, a unified framework comprising Temporal Partition Tree (TPT) that enables coarse‑to‑fine optimization based on the inheritance‑based partition tree structure for hierarchical temporal decomposition, and Spatial Ancestral Chains (SAC) that recursively query ancestral hierarchical structure to provide complementary spatial dynamics while specializing motion representations across ancestral nodes. Experimental results on different datasets indicate that our proposed method achieves 8.26% improvement of LPIPS on NVIDIA‑LS and 9.09% improvement of mLPIPS on DyCheck compared to the second‑best method. Code: https://github.com/iCVTEAM/WorldTree.
Authors:Zhenghuang Wu, Kang Chen, Zeyu Zhang, Hao Tang
Abstract:
Recent advances in diffusion‑based generative models have established a new paradigm for image and video relighting. However, extending these capabilities to 4D relighting remains challenging, due primarily to the scarcity of paired 4D relighting training data and the difficulty of maintaining temporal consistency across extreme viewpoints. In this work, we propose Light4D, a novel training‑free framework designed to synthesize consistent 4D videos under target illumination, even under extreme viewpoint changes. First, we introduce Disentangled Flow Guidance, a time‑aware strategy that effectively injects lighting control into the latent space while preserving geometric integrity. Second, to reinforce temporal consistency, we develop Temporal Consistent Attention within the IC‑Light architecture and further incorporate deterministic regularization to eliminate appearance flickering. Extensive experiments demonstrate that our method achieves competitive performance in temporal consistency and lighting fidelity, robustly handling camera rotations from ‑90 to 90. Code: https://github.com/AIGeeksGroup/Light4D. Website: https://aigeeksgroup.github.io/Light4D.
Authors:Yi Zhang, Yunshuang Wang, Zeyu Zhang, Hao Tang
Abstract:
Achieving spatial intelligence requires moving beyond visual plausibility to build world simulators grounded in physical laws. While coding LLMs have advanced static 3D scene generation, extending this paradigm to 4D dynamics remains a critical frontier. This task presents two fundamental challenges: multi‑scale context entanglement, where monolithic generation fails to balance local object structures with global environmental layouts; and a semantic‑physical execution gap, where open‑loop code generation leads to physical hallucinations lacking dynamic fidelity. We introduce Code2Worlds, a framework that formulates 4D generation as language‑to‑simulation code generation. First, we propose a dual‑stream architecture that disentangles retrieval‑augmented object generation from hierarchical environmental orchestration. Second, to ensure dynamic fidelity, we establish a physics‑aware closed‑loop mechanism in which a PostProcess Agent scripts dynamics, coupled with a VLM‑Motion Critic that performs self‑reflection to iteratively refine simulation code. Evaluations on the Code4D benchmark show Code2Worlds outperforms baselines with a 41% SGS gain and 49% higher Richness, while uniquely generating physics‑aware dynamics absent in prior static methods. Code: https://github.com/AIGeeksGroup/Code2Worlds. Website: https://aigeeksgroup.github.io/Code2Worlds.
Authors:Xiangyu Wu, Dongming Jiang, Feng Yu, Yueying Tian, Jiaqi Tang, Qing-Guo Chen, Yang Yang, Jianfeng Lu
Abstract:
Mainstream Test‑Time Adaptation (TTA) methods for adapting vision‑language models, e.g., CLIP, typically rely on Shannon Entropy (SE) at test time to measure prediction uncertainty and inconsistency. However, since CLIP has a built‑in bias from pretraining on highly imbalanced web‑crawled data, SE inevitably results in producing biased estimates of uncertainty entropy. To address this issue, we notably find and demonstrate that Tsallis Entropy (TE), a generalized form of SE, is naturally suited for characterizing biased distributions by introducing a non‑extensive parameter q, with the performance of SE serving as a lower bound for TE. Building upon this, we generalize TE into Adaptive Debiasing Tsallis Entropy (ADTE) for TTA, customizing a class‑specific parameter q^l derived by normalizing the estimated label bias from continuously incoming test instances, for each category. This adaptive approach allows ADTE to accurately select high‑confidence views and seamlessly integrate with a label adjustment strategy to enhance adaptation, without introducing distribution‑specific hyperparameter tuning. Besides, our investigation reveals that both TE and ADTE can serve as direct, advanced alternatives to SE in TTA, without any other modifications. Experimental results show that ADTE outperforms state‑of‑the‑art methods on ImageNet and its five variants, and achieves the highest average performance on 10 cross‑domain benchmarks, regardless of the model architecture or text prompts used. Our code is available at https://github.com/Jinx630/ADTE.
Authors:Zehao Xia, Yiqun Wang, Zhengda Lu, Kai Liu, Jun Xiao, Peter Wonka
Abstract:
Creating high‑fidelity, animatable 3D avatars from a single image remains a formidable challenge. We identified three desirable attributes of avatar generation: 1) the method should be feed‑forward, 2) model a 360° full‑head, and 3) should be animation‑ready. However, current work addresses only two of the three points simultaneously. To address these limitations, we propose OMEGA‑Avatar, the first feed‑forward framework that simultaneously generates a generalizable, 360°‑complete, and animatable 3D Gaussian head from a single image. Starting from a feed‑forward and animatable framework, we address the 360° full‑head avatar generation problem with two novel components. First, to overcome poor hair modeling in full‑head avatar generation, we introduce a semantic‑aware mesh deformation module that integrates multi‑view normals to optimize a FLAME head with hair while preserving its topology structure. Second, to enable effective feed‑forward decoding of full‑head features, we propose a multi‑view feature splatting module that constructs a shared canonical UV representation from features across multiple views through differentiable bilinear splatting, hierarchical UV mapping, and visibility‑aware fusion. This approach preserves both global structural coherence and local high‑frequency details across all viewpoints, ensuring 360° consistency without per‑instance optimization. Extensive experiments demonstrate that OMEGA‑Avatar achieves state‑of‑the‑art performance, significantly outperforming existing baselines in 360° full‑head completeness while robustly preserving identity across different viewpoints.
Authors:Chengwei Ma, Zhen Tian, Zhou Zhou, Zhixian Xu, Xiaowei Zhu, Xia Hua, Si Shi, F. Richard Yu
Abstract:
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state‑of‑the‑art models fail to capture topology and symbolic logic in engineering schematics, as their pixel‑driven paradigm discards the explicit vector‑defined relations needed for reasoning. To overcome this, we propose a Vector‑to‑Graph (V2G) pipeline that converts CAD diagrams into property graphs where nodes represent components and edges encode connectivity, making structural dependencies explicit and machine‑auditable. On a diagnostic benchmark of electrical compliance checks, V2G yields large accuracy gains across all error categories, while leading MLLMs remain near chance level. These results highlight the systemic inadequacy of pixel‑based methods and demonstrate that structure‑aware representations provide a reliable path toward practical deployment of multimodal AI in engineering domains. To facilitate further research, we release our benchmark and implementation at https://github.com/gm‑embodied/V2G‑Audit.
Authors:Khanh Nguyen, Dasith de Silva Edirimuni, Ghulam Mubashar Hassan, Ajmal Mian
Abstract:
3D assets have rapidly expanded in quantity and diversity due to the growing popularity of virtual reality and gaming. As a result, text‑to‑shape retrieval has become essential in facilitating intuitive search within large repositories. However, existing methods require canonical poses and support few object categories, limiting their real‑world applicability where objects can belong to diverse classes and appear in random orientations. To address this challenge, we propose RI‑Mamba, the first rotation‑invariant state‑space model for point clouds. RI‑Mamba defines global and local reference frames to disentangle pose from geometry and uses Hilbert sorting to construct token sequences with meaningful geometric structure while maintaining rotation invariance. We further introduce a novel strategy to compute orientational embeddings and reintegrate them via feature‑wise linear modulation, effectively recovering spatial context and enhancing model expressiveness. Our strategy is inherently compatible with state‑space models and operates in linear time. To scale up retrieval, we adopt cross‑modal contrastive learning with automated triplet generation, allowing training on diverse datasets without manual annotation. Extensive experiments demonstrate RI‑Mamba's superior representational capacity and robustness, achieving state‑of‑the‑art performance on the OmniObject3D benchmark across more than 200 object categories under arbitrary orientations. Our code will be made available at https://github.com/ndkhanh360/RI‑Mamba.git.
Authors:Jeongho Noh, Tai Hyoung Rhee, Eunho Lee, Jeongyun Kim, Sunwoo Lee, Ayoung Kim
Abstract:
Reliable 3D instance segmentation is fundamental to language‑grounded robotic manipulation. Its critical application lies in cluttered environments, where occlusions, limited viewpoints, and noisy masks degrade perception. To address these challenges, we present Clutt3R‑Seg, a zero‑shot pipeline for robust 3D instance segmentation for language‑grounded grasping in cluttered scenes. Our key idea is to introduce a hierarchical instance tree of semantic cues. Unlike prior approaches that attempt to refine noisy masks, our method leverages them as informative cues: through cross‑view grouping and conditional substitution, the tree suppresses over‑ and under‑segmentation, yielding view‑consistent masks and robust 3D instances. Each instance is enriched with open‑vocabulary semantic embeddings, enabling accurate target selection from natural language instructions. To handle scene changes during multi‑stage tasks, we further introduce a consistency‑aware update that preserves instance correspondences from only a single post‑interaction image, allowing efficient adaptation without rescanning. Clutt3R‑Seg is evaluated on both synthetic and real‑world datasets, and validated on a real robot. Across all settings, it consistently outperforms state‑of‑the‑art baselines in cluttered and sparse‑view scenarios. Even on the most challenging heavy‑clutter sequences, Clutt3R‑Seg achieves an AP@25 of 61.66, over 2.2x higher than baselines, and with only four input views it surpasses MaskClustering with eight views by more than 2x. The code is available at: https://github.com/jeonghonoh/clutt3r‑seg.
Authors:Yufeng Tian, Shuiqi Cheng, Tianming Wei, Tianxing Zhou, Yuanhang Zhang, Zixian Liu, Qianwei Han, Zhecheng Yuan, Huazhe Xu
Abstract:
Tactile information plays a crucial role in human manipulation tasks and has recently garnered increasing attention in robotic manipulation. However, existing approaches mostly focus on the alignment of visual and tactile features and the integration mechanism tends to be direct concatenation. Consequently, they struggle to effectively cope with occluded scenarios due to neglecting the inherent complementary nature of both modalities and the alignment may not be exploited enough, limiting the potential of their real‑world deployment. In this paper, we present ViTaS, a simple yet effective framework that incorporates both visual and tactile information to guide the behavior of an agent. We introduce Soft Fusion Contrastive Learning, an advanced version of conventional contrastive learning method and a CVAE module to utilize the alignment and complementarity within visuo‑tactile representations. We demonstrate the effectiveness of our method in 12 simulated and 3 real‑world environments, and our experiments show that ViTaS significantly outperforms existing baselines. Project page: https://skyrainwind.github.io/ViTaS/index.html.
Authors:Changti Wu, Jiahuai Mao, Yuzhuo Miao, Shijie Lian, Bin Yu, Xiaopeng Lin, Cong Huang, Lei Zhang, Kai Chen
Abstract:
Large‑scale Visual Instruction Tuning (VIT) has become a key paradigm for advancing the performance of vision‑language models (VLMs) across various multimodal tasks. However, training on the large‑scale datasets is computationally expensive and inefficient due to redundancy in the data, which motivates the need for multimodal data selection to improve training efficiency. Existing data selection methods for VIT either require costly training or gradient computation. Training‑free alternatives often depend on proxy models or datasets, instruction‑agnostic representations, and pairwise similarity with quadratic complexity, limiting scalability and representation fidelity. In this work, we propose ScalSelect, a scalable training‑free multimodal data selection method with linear‑time complexity with respect to the number of samples, eliminating the need for external models or auxiliary datasets. ScalSelect first constructs sample representations by extracting visual features most attended by instruction tokens in the target VLM, capturing instruction‑relevant information. It then identifies samples whose representations best approximate the dominant subspace of the full dataset representations, enabling scalable importance scoring without pairwise comparisons. Extensive experiments across multiple VLMs, datasets, and selection budgets demonstrate that ScalSelect achieves over 97.5% of the performance of training on the full dataset using only 16% of the data, and even outperforms full‑data training in some settings. The code is available at \hrefhttps://github.com/ChangtiWu/ScalSelectScalSelect.
Authors:Zedong Chu, Shichao Xie, Xiaolong Wu, Yanfen Shen, Minghua Luo, Zhengbo Wang, Fei Liu, Xiaoxu Leng, Junjun Hu, Mingyang Yin, Jia Lu, Yingnan Guo, Kai Yang, Jiawei Han, Xu Chen, Yanqing Zhu, Yuxiang Zhao, Xin Liu, Yirong Yang, Ye He, Jiahang Wang, Yang Cai, Tianlin Zhang, Li Gao, Liu Liu, Mingchao Sun, Fan Jiang, Chiyu Wang, Zhicheng Liu, Hongyu Pan, Honglin Han, Zhining Gu, Kuan Yang, Jianfang Zhang, Di Jing, Zihao Guan, Wei Guo, Guoqing Liu, Di Yang, Xiangpo Yang, Menglin Yang, Hongguang Xing, Weiguo Li, Mu Xu
Abstract:
Embodied navigation has long been fragmented by task‑specific architectures. We introduce ABot‑N0, a unified Vision‑Language‑Action (VLA) foundation model that achieves a ``Grand Unification'' across 5 core tasks: Point‑Goal, Object‑Goal, Instruction‑Following, POI‑Goal, and Person‑Following. ABot‑N0 utilizes a hierarchical ``Brain‑Action'' architecture, pairing an LLM‑based Cognitive Brain for semantic reasoning with a Flow Matching‑based Action Expert for precise, continuous trajectory generation.
To support large‑scale learning, we developed the ABot‑N0 Data Engine, curating 16.9M expert trajectories and 5.0M reasoning samples across 7,802 high‑fidelity 3D scenes (10.7 \textkm^2). ABot‑N0 achieves new SOTA performance across 7 benchmarks, significantly outperforming specialized models. Furthermore, our Agentic Navigation System integrates a planner with hierarchical topological memory, enabling robust, long‑horizon missions in dynamic real‑world environments.
Authors:Seungyeon Yoo, Youngseok Jang, Dabin Kim, Youngsoo Han, Seungwoo Jung, H. Jin Kim
Abstract:
Visual navigation models often struggle in real‑world dynamic environments due to limited robustness to the sim‑to‑real gap and the difficulty of training policies tailored to target deployment environments (e.g., households, restaurants, and factories). Although real‑to‑sim navigation simulation using 3D Gaussian Splatting (GS) can mitigate these challenges, prior GS‑based works have considered only static scenes or non‑photorealistic human obstacles built from simulator assets, despite the importance of safe navigation in dynamic environments. To address these issues, we propose ReaDy‑Go, a novel real‑to‑sim simulation pipeline that synthesizes photorealistic dynamic scenarios in target environments by augmenting a reconstructed static GS scene with dynamic human GS obstacles, and trains navigation policies using the generated datasets. The pipeline provides three key contributions: (1) a dynamic GS simulator that integrates static scene GS with a human animation module, enabling the insertion of animatable human GS avatars and the synthesis of plausible human motions from 2D trajectories, (2) a navigation dataset generation framework that leverages the simulator along with a robot expert planner designed for dynamic GS representations and a human planner, and (3) robust navigation policies to both the sim‑to‑real gap and moving obstacles. The proposed simulator generates thousands of photorealistic navigation scenarios with animatable human GS avatars from arbitrary viewpoints. ReaDy‑Go outperforms baselines across target environments in both simulation and real‑world experiments, demonstrating improved navigation performance even after sim‑to‑real transfer and in the presence of moving obstacles. Moreover, zero‑shot sim‑to‑real deployment in an unseen environment indicates its generalization potential. Project page: https://syeon‑yoo.github.io/ready‑go‑site/.
Authors:Chen Zhao, Jiawei Chen, Hongyu Li, Zhuoliang Kang, Shilin Lu, Xiaoming Wei, Kai Zhang, Jian Yang, Ying Tai
Abstract:
Recent advances in video diffusion models have significantly improved visual quality, yet ultra‑high‑resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning, and detail synthesis. To address these limitations, we propose LUVE, a Latent‑cascaded UHR Video generation framework built upon dual frequency Experts. LUVE employs a three‑stage architecture comprising low‑resolution motion generation for motion‑consistent latent synthesis, video latent upsampling that performs resolution upsampling directly in the latent space to mitigate memory and computational overhead, and high‑resolution content refinement that integrates low‑frequency and high‑frequency experts to jointly enhance semantic coherence and fine‑grained detail generation. Extensive experiments demonstrate that our LUVE achieves superior photorealism and content fidelity in UHR video generation, and comprehensive ablation studies further validate the effectiveness of each component. The project is available at \hrefhttps://unicornanrocinu.github.io/LUVE_web/https://github.io/LUVE/.
Authors:De-Xing Huang, Chaohui Yu, Xiao-Hu Zhou, Tian-Yu Xiang, Qin-Yi Zhang, Mei-Jiang Gui, Rui-Ze Ma, Chen-Yu Wang, Nu-Fang Xiao, Fan Wang, Zeng-Guang Hou
Abstract:
X‑ray angiography is the gold standard imaging modality for cardiovascular diseases. However, current deep learning approaches for X‑ray angiogram analysis are severely constrained by the scarcity of annotated data. While large‑scale self‑supervised learning (SSL) has emerged as a promising solution, its potential in this domain remains largely unexplored, primarily due to the lack of effective SSL frameworks and large‑scale datasets. To bridge this gap, we introduce a vascular anatomy‑aware masked image modeling (VasoMIM) framework that explicitly integrates domain‑specific anatomical knowledge. Specifically, VasoMIM comprises two key designs: an anatomy‑guided masking strategy and an anatomical consistency loss. The former strategically masks vessel‑containing patches to compel the model to learn robust vascular semantics, while the latter preserves structural consistency of vessels between original and reconstructed images, enhancing the discriminability of the learned representations. In conjunction with VasoMIM, we curate XA‑170K, the largest X‑ray angiogram pre‑training dataset to date. We validate VasoMIM on four downstream tasks across six datasets, where it demonstrates superior transferability and achieves state‑of‑the‑art performance compared to existing methods. These findings highlight the significant potential of VasoMIM as a foundation model for advancing a wide range of X‑ray angiogram analysis tasks. VasoMIM and XA‑170K will be available at https://github.com/Dxhuang‑CASIA/XA‑SSL.
Authors:David Wan, Han Wang, Ziyang Wang, Elias Stengel-Eskin, Hyunji Lee, Mohit Bansal
Abstract:
Multimodal large language models (MLLMs) are increasingly used for real‑world tasks involving multi‑step reasoning and long‑form generation, where reliability requires grounding model outputs in heterogeneous input sources and verifying individual factual claims. However, existing multimodal grounding benchmarks and evaluation methods focus on simplified, observation‑based scenarios or limited modalities and fail to assess attribution in complex multimodal reasoning. We introduce MuRGAt (Multimodal Reasoning with Grounded Attribution), a benchmark for evaluating fact‑level multimodal attribution in settings that require reasoning beyond direct observation. Given inputs spanning video, audio, and other modalities, MuRGAt requires models to generate answers with explicit reasoning and precise citations, where each citation specifies both modality and temporal segments. To enable reliable assessment, we introduce an automatic evaluation framework that strongly correlates with human judgments. Benchmarking with human and automated scores reveals that even strong MLLMs frequently hallucinate citations despite correct reasoning. Moreover, we observe a key trade‑off: increasing reasoning depth or enforcing structured grounding often degrades accuracy, highlighting a significant gap between internal reasoning and verifiable attribution.
Authors:Mark D. Olchanyi, Annabel Sorby-Adams, John Kirsch, Brian L. Edlow, Ava Farnan, Renfei Liu, Matthew S. Rosen, Emery N. Brown, W. Taylor Kimberly, Juan Eugenio Iglesias
Abstract:
Portable, ultra‑low‑field (ULF) magnetic resonance imaging has the potential to expand access to neuroimaging but currently suffers from coarse spatial and angular resolutions and low signal‑to‑noise ratios. Diffusion tensor imaging (DTI), a sequence tailored to detect and reconstruct white matter tracts within the brain, is particularly prone to such imaging degradation due to inherent sequence design coupled with prolonged scan times. In addition, ULF DTI scans exhibit artifacting that spans both the space and angular domains, requiring a custom modelling algorithm for subsequent correction. We introduce a nine‑direction, single‑shell ULF DTI sequence, as well as a companion Bayesian bias field correction algorithm that possesses angular dependence and convolutional neural network‑based superresolution algorithm that is generalizable across DTI datasets and does not require re‑training (''DiffSR''). We show through a synthetic downsampling experiment and white matter assessment in real, matched ULF and high‑field DTI scans that these algorithms can recover microstructural and volumetric white matter information at ULF. We also show that DiffSR can be directly applied to white matter‑based Alzheimers disease classification in synthetically degraded scans, with notable improvements in agreement between DTI metrics, as compared to un‑degraded scans. We freely disseminate the Bayesian bias correction algorithm and DiffSR with the goal of furthering progress on both ULF reconstruction methods and general DTI sequence harmonization. We release all code related to DiffSR for \hrefhttps://github.com/markolchanyi/DiffSRpublic \space use.
Authors:Evgeney Bogatyrev, Khaled Abud, Ivan Molodetskikh, Nikita Alutis, Dmitriy Vatolin
Abstract:
Recent advancements in real‑time super‑resolution have enabled higher‑quality video streaming, yet existing methods struggle with the unique challenges of compressed video content. Commonly used datasets do not accurately reflect the characteristics of streaming media, limiting the relevance of current benchmarks. To address this gap, we introduce a comprehensive dataset ‑ StreamSR ‑ sourced from YouTube, covering a wide range of video genres and resolutions representative of real‑world streaming scenarios. We benchmark 11 state‑of‑the‑art real‑time super‑resolution models to evaluate their performance for the streaming use‑case.
Furthermore, we propose EfRLFN, an efficient real‑time model that integrates Efficient Channel Attention and a hyperbolic tangent activation function ‑ a novel design choice in the context of real‑time super‑resolution. We extensively optimized the architecture to maximize efficiency and designed a composite loss function that improves training convergence. EfRLFN combines the strengths of existing architectures while improving both visual quality and runtime performance.
Finally, we show that fine‑tuning other models on our dataset results in significant performance gains that generalize well across various standard benchmarks. We made the dataset, the code, and the benchmark available at https://github.com/EvgeneyBogatyrev/EfRLFN.
Authors:Yandan Yang, Shuang Zeng, Tong Lin, Xinyuan Chang, Dekang Qi, Junjin Xiao, Haoyun Liu, Ronghan Chen, Yuzhi Chen, Dongjie Huo, Feng Xiong, Xing Wei, Zhiheng Ma, Mu Xu
Abstract:
Building general‑purpose embodied agents across diverse hardware remains a central challenge in robotics, often framed as the ''one‑brain, many‑forms'' paradigm. Progress is hindered by fragmented data, inconsistent representations, and misaligned training objectives. We present ABot‑M0, a framework that builds a systematic data curation pipeline while jointly optimizing model architecture and training strategies, enabling end‑to‑end transformation of heterogeneous raw data into unified, efficient representations. From six public datasets, we clean, standardize, and balance samples to construct UniACT‑dataset, a large‑scale dataset with over 6 million trajectories and 9,500 hours of data, covering diverse robot morphologies and task scenarios. Unified pre‑training improves knowledge transfer and generalization across platforms and tasks, supporting general‑purpose embodied intelligence. To improve action prediction efficiency and stability, we propose the Action Manifold Hypothesis: effective robot actions lie not in the full high‑dimensional space but on a low‑dimensional, smooth manifold governed by physical laws and task constraints. Based on this, we introduce Action Manifold Learning (AML), which uses a DiT backbone to predict clean, continuous action sequences directly. This shifts learning from denoising to projection onto feasible manifolds, improving decoding speed and policy stability. ABot‑M0 supports modular perception via a dual‑stream mechanism that integrates VLM semantics with geometric priors and multi‑view inputs from plug‑and‑play 3D modules such as VGGT and Qwen‑Image‑Edit, enhancing spatial understanding without modifying the backbone and mitigating standard VLM limitations in 3D reasoning. Experiments show components operate independently with additive benefits. We will release all code and pipelines for reproducibility and future research.
Authors:Manuel Hetzel, Kerim Turacan, Hannes Reichert, Konrad Doll, Bernhard Sick
Abstract:
Human Trajectory Forecasting (HTF) predicts future human movements from past trajectories and environmental context, with applications in Autonomous Driving, Smart Surveillance, and Human‑Robot Interaction. While prior work has focused on accuracy, social interaction modeling, and diversity, little attention has been paid to uncertainty modeling, calibration, and forecasts from short observation periods, which are crucial for downstream tasks such as path planning and collision avoidance. We propose DD‑MDN, an end‑to‑end probabilistic HTF model that combines high positional accuracy, calibrated uncertainty, and robustness to short observations. Using a few‑shot denoising diffusion backbone and a dual mixture density network, our method learns self‑calibrated residence areas and probability‑ranked anchor paths, from which diverse trajectory hypotheses are derived, without predefined anchors or endpoints. Experiments on the ETH/UCY, SDD, inD, and IMPTC datasets demonstrate state‑of‑the‑art accuracy, robustness at short observation intervals, and reliable uncertainty modeling. The code is available at: https://github.com/kav‑institute/ddmdn.
Authors:Gongye Liu, Bo Yang, Yida Zhi, Zhizhou Zhong, Lei Ke, Didan Deng, Han Gao, Yongxiang Huang, Kaihao Zhang, Hongbo Fu, Wenhan Luo
Abstract:
Preference optimization for diffusion and flow‑matching models relies on reward functions that are both discriminatively robust and computationally efficient. Vision‑Language Models (VLMs) have emerged as the primary reward provider, leveraging their rich multimodal priors to guide alignment. However, their computation and memory cost can be substantial, and optimizing a latent diffusion generator through a pixel‑space reward introduces a domain mismatch that complicates alignment. In this paper, we propose DiNa‑LRM, a diffusion‑native latent reward model that formulates preference learning directly on noisy diffusion states. Our method introduces a noise‑calibrated Thurstone likelihood with diffusion‑noise‑dependent uncertainty. DiNa‑LRM leverages a pretrained latent diffusion backbone with a timestep‑conditioned reward head, and supports inference‑time noise ensembling, providing a diffusion‑native mechanism for test‑time scaling and robust rewarding. Across image alignment benchmarks, DiNa‑LRM substantially outperforms existing diffusion‑based reward baselines and achieves performance competitive with state‑of‑the‑art VLMs at a fraction of the computational cost. In preference optimization, we demonstrate that DiNa‑LRM improves preference optimization dynamics, enabling faster and more resource‑efficient model alignment.
Authors:Ruichuan An, Sihan Yang, Ziyu Guo, Wei Dai, Zijun Shen, Haodong Li, Renrui Zhang, Xinyu Wei, Guopeng Li, Wenshan Wu, Wentao Zhang
Abstract:
Unified Multimodal Models (UMMs) have shown remarkable progress in visual generation. Yet, existing benchmarks predominantly assess Crystallized Intelligence, which relies on recalling accumulated knowledge and learned schemas. This focus overlooks Generative Fluid Intelligence (GFI): the capacity to induce patterns, reason through constraints, and adapt to novel scenarios on the fly. To rigorously assess this capability, we introduce GENIUS (GEN Fluid Intelligence EvalUation Suite). We formalize GFI as a synthesis of three primitives. These include Inducing Implicit Patterns (e.g., inferring personalized visual preferences), Executing Ad‑hoc Constraints (e.g., visualizing abstract metaphors), and Adapting to Contextual Knowledge (e.g., simulating counter‑intuitive physics). Collectively, these primitives challenge models to solve problems grounded entirely in the immediate context. Our systematic evaluation of 12 representative models reveals significant performance deficits in these tasks. Crucially, our diagnostic analysis disentangles these failure modes. It demonstrates that deficits stem from limited context comprehension rather than insufficient intrinsic generative capability. To bridge this gap, we propose a training‑free attention intervention strategy. Ultimately, GENIUS establishes a rigorous standard for GFI, guiding the field beyond knowledge utilization toward dynamic, general‑purpose reasoning. Our dataset and code will be released at: \hrefhttps://github.com/arctanxarc/GENIUShttps://github.com/arctanxarc/GENIUS.
Authors:Di Chang, Ji Hou, Aljaz Bozic, Assaf Neuberger, Felix Juefei-Xu, Olivier Maury, Gene Wei-Chin Lin, Tuur Stuyck, Doug Roble, Mohammad Soleymani, Stephane Grabli
Abstract:
We present HairWeaver, a diffusion‑based pipeline that animates a single human image with realistic and expressive hair dynamics. While existing methods successfully control body pose, they lack specific control over hair, and as a result, fail to capture the intricate hair motions, resulting in stiff and unrealistic animations. HairWeaver overcomes this limitation using two specialized modules: a Motion‑Context‑LoRA to integrate motion conditions and a Sim2Real‑Domain‑LoRA to preserve the subject's photoreal appearance across different data domains. These lightweight components are designed to guide a video diffusion backbone while maintaining its core generative capabilities. By training on a specialized dataset of dynamic human motion generated from a CG simulator, HairWeaver affords fine control over hair motion and ultimately learns to produce highly realistic hair that responds naturally to movement. Comprehensive evaluations demonstrate that our approach sets a new state of the art, producing lifelike human hair animations with dynamic details.
Authors:Divya Jyoti Bajpai, Dhruv Bhardwaj, Soumya Roy, Tejas Duseja, Harsh Agarwal, Aashay Sandansing, Manjesh Kumar Hanawal
Abstract:
Flow‑matching models deliver state‑of‑the‑art fidelity in image and video generation, but the inherent sequential denoising process renders them slower. Existing acceleration methods like distillation, trajectory truncation, and consistency approaches are static, require retraining, and often fail to generalize across tasks. We propose FastFlow, a plug‑and‑play adaptive inference framework that accelerates generation in flow matching models. FastFlow identifies denoising steps that produce only minor adjustments to the denoising path and approximates them without using the full neural network models used for velocity predictions. The approximation utilizes finite‑difference velocity estimates from prior predictions to efficiently extrapolate future states, enabling faster advancements along the denoising path at zero compute cost. This enables skipping computation at intermediary steps. We model the decision of how many steps to safely skip before requiring a full model computation as a multi‑armed bandit problem. The bandit learns the optimal skips to balance speed with performance. FastFlow integrates seamlessly with existing pipelines and generalizes across image generation, video generation, and editing tasks. Experiments demonstrate a speedup of over 2.6x while maintaining high‑quality outputs. The source code for this work can be found at https://github.com/Div290/FastFlow.
Authors:Yujie Chen, Li Zhang, Xiaomeng Chu, Tian Zhang
Abstract:
We propose PuriLight, a lightweight and efficient framework for self‑supervised monocular depth estimation, to address the dual challenges of computational efficiency and detail preservation. While recent advances in self‑supervised depth estimation have reduced reliance on ground truth supervision, existing approaches remain constrained by either bulky architectures compromising practicality or lightweight models sacrificing structural precision. These dual limitations underscore the critical need to develop lightweight yet structurally precise architectures. Our framework addresses these limitations through a three‑stage architecture incorporating three novel modules: the Shuffle‑Dilation Convolution (SDC) module for local feature extraction, the Rotation‑Adaptive Kernel Attention (RAKA) module for hierarchical feature enhancement, and the Deep Frequency Signal Purification (DFSP) module for global feature purification. Through effective collaboration, these modules enable PuriLight to achieve both lightweight and accurate feature extraction and processing. Extensive experiments demonstrate that PuriLight achieves state‑of‑the‑art performance with minimal training parameters while maintaining exceptional computational efficiency. Codes will be available at https://github.com/ishrouder/PuriLight.
Authors:Lei Yao, Yi Wang, Yawen Cui, Moyun Liu, Lap-Pui Chau
Abstract:
Query‑based 3D scene instance segmentation from point clouds has attained notable performance. However, existing methods suffer from the query initialization dilemma due to the sparse nature of point clouds and rely on computationally intensive attention mechanisms in query decoders. We accordingly introduce LaSSM, prioritizing simplicity and efficiency while maintaining competitive performance. Specifically, we propose a hierarchical semantic‑spatial query initializer to derive the query set from superpoints by considering both semantic cues and spatial distribution, achieving comprehensive scene coverage and accelerated convergence. We further present a coordinate‑guided state space model (SSM) decoder that progressively refines queries. The novel decoder features a local aggregation scheme that restricts the model to focus on geometrically coherent regions and a spatial dual‑path SSM block to capture underlying dependencies within the query set by integrating associated coordinates information. Our design enables efficient instance prediction, avoiding the incorporation of noisy information and reducing redundant computation. LaSSM ranks first place on the latest ScanNet++ V2 leaderboard, outperforming the previous best method by 2.5% mAP with only 1/3 FLOPs, demonstrating its superiority in challenging large‑scale scene instance segmentation. LaSSM also achieves competitive performance on ScanNet, ScanNet200, S3DIS and ScanNet++ V1 benchmarks with less computational cost. Extensive ablation studies and qualitative results validate the effectiveness of our design. The code and weights are available at https://github.com/RayYoh/LaSSM.
Authors:Nuno Gonçalves, Diogo Nunes, Carla Guerra, João Marcos
Abstract:
Ensuring compliance with ISO/IEC and ICAO standards for facial images in machine‑readable travel documents (MRTDs) is essential for reliable identity verification, but current manual inspection methods are inefficient in high‑demand environments. This paper introduces the DFIC dataset, a novel comprehensive facial image dataset comprising around 58,000 annotated images and 2706 videos of more than 1000 subjects, that cover a broad range of non‑compliant conditions, in addition to compliant portraits. Our dataset provides a more balanced demographic distribution than the existing public datasets, with one partition that is nearly uniformly distributed, facilitating the development of automated ICAO compliance verification methods.
Using DFIC, we fine‑tuned a novel method that heavily relies on spatial attention mechanisms for the automatic validation of ICAO compliance requirements, and we have compared it with the state‑of‑the‑art aimed at ICAO compliance verification, demonstrating improved results. DFIC dataset is now made public (https://github.com/visteam‑isr‑uc/DFIC) for the training and validation of new models, offering an unprecedented diversity of faces, that will improve both robustness and adaptability to the intrinsically diverse combinations of faces and props that can be presented to the validation system. These results emphasize the potential of DFIC to enhance automated ICAO compliance methods but it can also be used in many other applications that aim to improve the security, privacy, and fairness of facial recognition systems.
Authors:Jinqing Zhang, Zehua Fu, Zelin Xu, Wenying Dai, Qingjie Liu, Yunhong Wang
Abstract:
The comprehensive understanding capabilities of world models for driving scenarios have significantly improved the planning accuracy of end‑to‑end autonomous driving frameworks. However, the redundant modeling of static regions and the lack of deep interaction with trajectories hinder world models from exerting their full effectiveness. In this paper, we propose Temporal Residual World Model (TR‑World), which focuses on dynamic object modeling. By calculating the temporal residuals of scene representations, the information of dynamic objects can be extracted without relying on detection and tracking. TR‑World takes only temporal residuals as input, thus predicting the future spatial distribution of dynamic objects more precisely. By combining the prediction with the static object information contained in the current BEV features, accurate future BEV features can be obtained. Furthermore, we propose Future‑Guided Trajectory Refinement (FGTR) module, which conducts interaction between prior trajectories (predicted from the current scene representation) and the future BEV features. This module can not only utilize future road conditions to refine trajectories, but also provides sparse spatial‑temporal supervision on future BEV features to prevent world model collapse. Comprehensive experiments conducted on the nuScenes and NAVSIM datasets demonstrate that our method, namely ResWorld, achieves state‑of‑the‑art planning performance. The code is available at https://github.com/mengtan00/ResWorld.git.
Authors:Minggui He, Mingchen Dai, Jian Zhang, Yilun Liu, Shimin Tao, Pufan Zeng, Osamu Yoshie, Yuya Ieiri
Abstract:
Vision‑Language Models (VLMs) have shown promise in generating plotting code from chart images, yet achieving structural fidelity remains challenging. Existing approaches largely rely on supervised fine‑tuning, encouraging surface‑level token imitation rather than faithful modeling of underlying chart structure, which often leads to hallucinated or semantically inconsistent outputs. We propose Chart Specification, a structured intermediate representation that shifts training from text imitation to semantically grounded supervision. Chart Specification filters syntactic noise to construct a structurally balanced training set and supports a Spec‑Align Reward that provides fine‑grained, verifiable feedback on structural correctness, enabling reinforcement learning to enforce consistent plotting logic. Experiments on three public benchmarks show that our method consistently outperforms prior approaches. With only 3K training samples, we achieve strong data efficiency, surpassing leading baselines by up to 61.7% on complex benchmarks, and scaling to 4K samples establishes new state‑of‑the‑art results across all evaluated metrics. Overall, our results demonstrate that precise structural supervision offers an efficient pathway to high‑fidelity chart‑to‑code generation. Code and dataset are available at: https://github.com/Mighten/chart‑specification‑paper
Authors:Darakshan Rashid, Raza Imam, Dwarikanath Mahapatra, Brejesh Lall
Abstract:
Deep neural networks for chest X‑ray classification achieve strong average performance, yet often underperform for specific demographic subgroups, raising critical concerns about clinical safety and equity. Existing debiasing methods frequently yield inconsistent improvements across datasets or attain fairness by degrading overall diagnostic utility, treating fairness as a post hoc constraint rather than a property of the learned representation. In this work, we propose Stride‑Net (Sensitive Attribute Resilient Learning via Disentanglement and Learnable Masking with Embedding Alignment), a fairness‑aware framework that learns disease‑discriminative yet demographically invariant representations for chest X‑ray analysis. Stride‑Net operates at the patch level, using a learnable stride‑based mask to select label‑aligned image regions while suppressing sensitive attribute information through adversarial confusion loss. To anchor representations in clinical semantics and discourage shortcut learning, we further enforce semantic alignment between image features and BioBERT‑based disease label embeddings via Group Optimal Transport. We evaluate Stride‑Net on the MIMIC‑CXR and CheXpert benchmarks across race and intersectional race‑gender subgroups. Across architectures including ResNet and Vision Transformers, Stride‑Net consistently improves fairness metrics while matching or exceeding baseline accuracy, achieving a more favorable accuracy‑fairness trade‑off than prior debiasing approaches. Our code is available at https://github.com/Daraksh/Fairness_StrideNet.
Authors:Yuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu, Feng Ling, Xiawu Zheng, Huafeng Kuang, Huixia Li, Xing Wang, Xuefeng Xiao, Fei Chao, Rongrong Ji
Abstract:
Autoregressive models, often built on Transformer architectures, represent a powerful paradigm for generating ultra‑long videos by synthesizing content in sequential chunks. However, this sequential generation process is notoriously slow. While caching strategies have proven effective for accelerating traditional video diffusion models, existing methods assume uniform denoising across all frames‑an assumption that breaks down in autoregressive models where different video chunks exhibit varying similarity patterns at identical timesteps. In this paper, we present FlowCache, the first caching framework specifically designed for autoregressive video generation. Our key insight is that each video chunk should maintain independent caching policies, allowing fine‑grained control over which chunks require recomputation at each timestep. We introduce a chunkwise caching strategy that dynamically adapts to the unique denoising characteristics of each chunk, complemented by a joint importance‑redundancy optimized KV cache compression mechanism that maintains fixed memory bounds while preserving generation quality. Our method achieves remarkable speedups of 2.38 times on MAGI‑1 and 6.7 times on SkyReels‑V2, with negligible quality degradation (VBench: 0.87 increase and 0.79 decrease respectively). These results demonstrate that FlowCache successfully unlocks the potential of autoregressive models for real‑time, ultra‑long video generation‑establishing a new benchmark for efficient video synthesis at scale. The code is available at https://github.com/mikeallen39/FlowCache.
Authors:Aojun Lu, Tao Feng, Hangjie Yuan, Wei Li, Yanan Sun
Abstract:
The adaptation of large‑scale Vision‑Language Models (VLMs) through post‑training reveals a pronounced generalization gap: models fine‑tuned with Reinforcement Learning (RL) consistently achieve superior out‑of‑distribution (OOD) performance compared to those trained with Supervised Fine‑Tuning (SFT). This paper posits a data‑centric explanation for this phenomenon, contending that RL's generalization advantage arises from an implicit data filtering mechanism that inherently prioritizes medium‑difficulty training samples. To test this hypothesis, we systematically evaluate the OOD generalization of SFT models across training datasets of varying difficulty levels. Our results confirm that data difficulty is a critical factor, revealing that training on hard samples significantly degrades OOD performance. Motivated by this finding, we introduce Difficulty‑Curated SFT (DC‑SFT), a straightforward method that explicitly filters the training set based on sample difficulty. Experiments show that DC‑SFT not only substantially enhances OOD generalization over standard SFT, but also surpasses the performance of RL‑based training, all while providing greater stability and computational efficiency. This work offers a data‑centric account of the OOD generalization gap in VLMs and establishes a more efficient pathway to achieving robust generalization. Code is available at https://github.com/byyx666/DC‑SFT.
Authors:Minglei Li, Mengfan He, Chunyu Li, Chao Chen, Xingyu Shao, Ziyang Meng
Abstract:
Cross‑view geo‑localization (CVGL) is pivotal for GNSS‑denied UAV navigation but remains brittle under the drastic geometric misalignment between oblique aerial views and orthographic satellite references. Existing methods predominantly operate within a 2D manifold, neglecting the underlying 3D geometry where view‑dependent vertical facades (macro‑structure) and scale variations (micro‑scale) severely corrupt feature alignment. To bridge this gap, we propose (MGS)^2, a geometry‑grounded framework. The core of our innovation is the Macro‑Geometric Structure Filtering (MGSF) module. Unlike pixel‑wise matching sensitive to noise, MGSF leverages dilated geometric gradients to physically filter out high‑frequency facade artifacts while enhancing the view‑invariant horizontal plane, directly addressing the domain shift. To guarantee robust input for this structural filtering, we explicitly incorporate a Micro‑Geometric Scale Adaptation (MGSA) module. MGSA utilizes depth priors to dynamically rectify scale discrepancies via multi‑branch feature fusion. Furthermore, a Geometric‑Appearance Contrastive Distillation (GACD) loss is designed to strictly discriminate against oblique occlusions. Extensive experiments demonstrate that (MGS)^2 achieves state‑of‑the‑art performance, recording a Recall@1 of 97.5% on University‑1652 and 97.02% on SUES‑200. Furthermore, the framework exhibits superior cross‑dataset generalization against geometric ambiguity. The code is available at: \hrefhttps://github.com/GabrielLi1473/MGS‑Nethttps://github.com/GabrielLi1473/MGS‑Net.
Authors:Jinjie Shen, Jing Wu, Yaxiong Wang, Lechao Cheng, Shengeng Tang, Tianrui Hui, Nan Pu, Zhun Zhong
Abstract:
Existing forgery detection methods are often limited to uni‑modal or bi‑modal settings, failing to handle the interleaved text, images, and videos prevalent in real‑world misinformation. To bridge this gap, this paper targets to develop a unified framework for omnibus vision‑language forgery detection and grounding. In this unified setting, the interplay between diverse modalities and the dual requirements of simultaneous detection and localization pose a critical ``difficulty bias`` problem: the simpler veracity classification task tends to dominate the gradients, leading to suboptimal performance in fine‑grained grounding during multi‑task optimization. To address this challenge, we propose OmniVL‑Guard, a balanced reinforcement learning framework for omnibus vision‑language forgery detection and grounding. Particularly, OmniVL‑Guard comprises two core designs: Self‑Evolving CoT Generatio and Adaptive Reward Scaling Policy Optimization (ARSPO). Self‑Evolving CoT Generation synthesizes high‑quality reasoning paths, effectively overcoming the cold‑start challenge. Building upon this, Adaptive Reward Scaling Policy Optimization (ARSPO) dynamically modulates reward scales and task weights, ensuring a balanced joint optimization. Extensive experiments demonstrate that OmniVL‑Guard significantly outperforms state‑of‑the‑art methods and exhibits zero‑shot robust generalization across out‑of‑domain scenarios. The dataset and code are publicly available at https://github.com/shen8424/OmniVL‑Guard.
Authors:Junhua Liu, Zhangcheng Wang, Zhike Han, Ningli Wang, Guotao Liang, Kun Kuang
Abstract:
Visual Chain‑of‑Thought (VCoT) has emerged as a promising paradigm for enhancing multimodal reasoning by integrating visual perception into intermediate reasoning steps. However, existing VCoT approaches are largely confined to static scenarios and struggle to capture the temporal dynamics essential for tasks such as instruction, prediction, and camera motion. To bridge this gap, we propose TwiFF‑2.7M, the first large‑scale, temporally grounded VCoT dataset derived from 2.7 million video clips, explicitly designed for dynamic visual question and answer. Accompanying this, we introduce TwiFF‑Bench, a high‑quality evaluation benchmark of 1,078 samples that assesses both the plausibility of reasoning trajectories and the correctness of final answers in open‑ended dynamic settings. Building on these foundations, we propose the TwiFF model, a unified modal that synergistically leverages pre‑trained video generation and image comprehension capabilities to produce temporally coherent visual reasoning cues‑iteratively generating future action frames and textual reasoning. Extensive experiments demonstrate that TwiFF significantly outperforms existing VCoT methods and Textual Chain‑of‑Thought baselines on dynamic reasoning tasks, which fully validates the effectiveness for visual question answering in dynamic scenarios. Our code and data is available at https://github.com/LiuJunhua02/TwiFF.
Authors:Kiarash Ghasemzadeh, Sedigheh Dehghani
Abstract:
Self‑driving cars hold significant potential to reduce traffic accidents, alleviate congestion, and enhance urban mobility. However, developing reliable AI systems for autonomous vehicles remains a substantial challenge. Over the past decade, multi‑task learning has emerged as a powerful approach to address complex problems in driving perception. Multi‑task networks offer several advantages, including increased computational efficiency, real‑time processing capabilities, optimized resource utilization, and improved generalization. In this study, we present AurigaNet, an advanced multi‑task network architecture designed to push the boundaries of autonomous driving perception. AurigaNet integrates three critical tasks: object detection, lane detection, and drivable area instance segmentation. The system is trained and evaluated using the BDD100K dataset, renowned for its diversity in driving conditions. Key innovations of AurigaNet include its end‑to‑end instance segmentation capability, which significantly enhances both accuracy and efficiency in path estimation for autonomous vehicles. Experimental results demonstrate that AurigaNet achieves an 85.2% IoU in drivable area segmentation, outperforming its closest competitor by 0.7%. In lane detection, AurigaNet achieves a remarkable 60.8% IoU, surpassing other models by more than 30%. Furthermore, the network achieves an mAP@0.5:0.95 of 47.6% in traffic object detection, exceeding the next leading model by 2.9%. Additionally, we validate the practical feasibility of AurigaNet by deploying it on embedded devices such as the Jetson Orin NX, where it demonstrates competitive real‑time performance. These results underscore AurigaNet's potential as a robust and efficient solution for autonomous driving perception systems. The code can be found here https://github.com/KiaRational/AurigaNet.
Authors:Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue, Jin Song Dong
Abstract:
Video Large Language Models (VideoLLMs) have recently achieved strong performance in video understanding tasks. However, we identify a previously underexplored generation failure: severe output repetition, where models degenerate into self‑reinforcing loops of repeated phrases or sentences. This failure mode is not captured by existing VideoLLM benchmarks, which focus primarily on task accuracy and factual correctness. We introduce VideoSTF, the first framework for systematically measuring and stress‑testing output repetition in VideoLLMs. VideoSTF formalizes repetition using three complementary n‑gram‑based metrics and provides a standardized testbed of 10,000 diverse videos together with a library of controlled temporal transformations. Using VideoSTF, we conduct pervasive testing, temporal stress testing, and adversarial exploitation across 10 advanced VideoLLMs. We find that output repetition is widespread and, critically, highly sensitive to temporal perturbations of video inputs. Moreover, we show that simple temporal transformations can efficiently induce repetitive degeneration in a black‑box setting, exposing output repetition as an exploitable security vulnerability. Our results reveal output repetition as a fundamental stability issue in modern VideoLLMs and motivate stability‑aware evaluation for video‑language systems. Our evaluation code and scripts are available at: https://github.com/yuxincao22/VideoSTF_benchmark.
Authors:Khanh Linh Tran, Minh Nguyen Dang, Thien Nguyen Trong, Hung Nguyen Quoc, Linh Nguyen Kieu
Abstract:
This paper presents a practical and lightweight solution for enhancing child detection in low‑quality surveillance footage, a critical component in real‑world missing child alert and daycare monitoring systems. Building upon the efficient YOLOv11n architecture, we propose a deployment‑ready pipeline that improves detection under challenging conditions including occlusion, small object size, low resolution, motion blur, and poor lighting commonly found in existing CCTV infrastructures. Our approach introduces a domain‑specific augmentation strategy that synthesizes realistic child placements using spatial perturbations such as partial visibility, truncation, and overlaps, combined with photometric degradations including lighting variation and noise. To improve recall of small and partially occluded instances, we integrate Slicing Aided Hyper Inference (SAHI) at inference time. All components are trained and evaluated on a filtered, child‑only subset of the Roboflow Daycare dataset. Compared to the baseline YOLOv11n, our enhanced system achieves a mean Average Precision at 0.5 IoU (mAP@0.5) of 0.967 and a mean Average Precision averaged over IoU thresholds from 0.5 to 0.95 (mAP@0.5:0.95) of 0.783, yielding absolute improvements of 0.7 percent and 2.3 percent, respectively, without architectural changes. Importantly, the entire pipeline maintains compatibility with low‑power edge devices and supports real‑time performance, making it particularly well suited for low‑cost or resource‑constrained industrial surveillance deployments. The example augmented dataset and the source code used to generate it are available at: https://github.com/html‑ptit/Data‑Augmentation‑YOLOv11n‑child‑detection
Authors:Bosen Lin, Feng Gao, Yanwei Yu, Junyu Dong, Qian Du
Abstract:
Underwater Image Enhancement (UIE) is an ill‑posed problem where natural clean references are not available, and the degradation levels vary significantly across semantic regions. Existing UIE methods treat images with a single global model and ignore the inconsistent degradation of different scene components. This oversight leads to significant color distortions and loss of fine details in heterogeneous underwater scenes, especially where degradation varies significantly across different image regions. Therefore, we propose SUCode (Semantic‑aware Underwater Codebook Network), which achieves adaptive UIE from semantic‑aware discrete codebook representation. Compared with one‑shot codebook‑based methods, SUCode exploits semantic‑aware, pixel‑level codebook representation tailored to heterogeneous underwater degradation. A three‑stage training paradigm is employed to represent raw underwater image features to avoid pseudo ground‑truth contamination. Gated Channel Attention Module (GCAM) and Frequency‑Aware Feature Fusion (FAFF) jointly integrate channel and frequency cues for faithful color restoration and texture recovery. Extensive experiments on multiple benchmarks demonstrate that SUCode achieves state‑of‑the‑art performance, outperforming recent UIE methods on both reference and no‑reference metrics. The code will be made public available at https://github.com/oucailab/SUCode.
Authors:Chenhao Zhang, Yazhe Niu, Hongsheng Li
Abstract:
Metaphorical comprehension in images remains a critical challenge for Nowadays AI systems. While Multimodal Large Language Models (MLLMs) excel at basic Visual Question Answering (VQA), they consistently struggle to grasp the nuanced cultural, emotional, and contextual implications embedded in visual content. This difficulty stems from the task's demand for sophisticated multi‑hop reasoning, cultural context, and Theory of Mind (ToM) capabilities, which current models lack. To fill this gap, we propose MetaphorStar, the first end‑to‑end visual reinforcement learning (RL) framework for image implication tasks. Our framework includes three core components: the fine‑grained dataset TFQ‑Data, the visual RL method TFQ‑GRPO, and the well‑structured benchmark TFQ‑Bench.
Our fully open‑source MetaphorStar family, trained using TFQ‑GRPO on TFQ‑Data, significantly improves performance by an average of 82.6% on the image implication benchmarks. Compared with 20+ mainstream MLLMs, MetaphorStar‑32B achieves state‑of‑the‑art (SOTA) on Multiple‑Choice Question and Open‑Style Question, significantly outperforms the top closed‑source model Gemini‑3.0‑pro on True‑False Question. Crucially, our experiments reveal that learning image implication tasks improves the general understanding ability, especially the complex visual reasoning ability. We further provide a systematic analysis of model parameter scaling, training data scaling, and the impact of different model architectures and training strategies, demonstrating the broad applicability of our method. We open‑sourced all model weights, datasets, and method code at https://metaphorstar.github.io.
Authors:Guanting Ye, Qiyan Zhao, Wenhao Yu, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Ka-Veng Yuen
Abstract:
Recent advances in 3D Large Multimodal Models (LMMs) built on Large Language Models (LLMs) have established the alignment of 3D visual features with LLM representations as the dominant paradigm. However, the inherited Rotary Position Embedding (RoPE) introduces limitations for multimodal processing. Specifically, applying 1D temporal positional indices disrupts the continuity of visual features along the column dimension, resulting in spatial locality loss. Moreover, RoPE follows the prior that temporally closer image tokens are more causally related, leading to long‑term decay in attention allocation and causing the model to progressively neglect earlier visual tokens as the sequence length increases. To address these issues, we propose C^2RoPE, an improved RoPE that explicitly models local spatial Continuity and spatial Causal relationships for visual processing. C^2RoPE introduces a spatio‑temporal continuous positional embedding mechanism for visual tokens. It first integrates 1D temporal positions with Cartesian‑based spatial coordinates to construct a triplet hybrid positional index, and then employs a frequency allocation strategy to encode spatio‑temporal positional information across the three index components. Additionally, we introduce Chebyshev Causal Masking, which determines causal dependencies by computing the Chebyshev distance of image tokens in 2D space. Evaluation results across various benchmarks, including 3D scene reasoning and 3D visual question answering, demonstrate C^2RoPE's effectiveness. The code is be available at https://github.com/ErikZ719/C2RoPE.
Authors:Dongshuo Yin, Xue Yang, Deng-Ping Fan, Shi-Min Hu
Abstract:
Deploying vision foundation models typically relies on efficient adaptation strategies, whereas conventional full fine‑tuning suffers from prohibitive costs and low efficiency. While delta‑tuning has proven effective in boosting the performance and efficiency of LLMs during adaptation, its advantages cannot be directly transferred to the fine‑tuning pipeline of vision foundation models. To push the boundaries of adaptation efficiency for vision tasks, we propose an adapter with Complex Linear Projection Optimization (CoLin). For architecture, we design a novel low‑rank complex adapter that introduces only about 1% parameters to the backbone. For efficiency, we theoretically prove that low‑rank composite matrices suffer from severe convergence issues during training, and address this challenge with a tailored loss. Extensive experiments on object detection, segmentation, image classification, and rotated object detection (remote sensing scenario) demonstrate that CoLin outperforms both full fine‑tuning and classical delta‑tuning approaches with merely 1% parameters for the first time, providing a novel and efficient solution for deployment of vision foundation models. We release the code on https://github.com/DongshuoYin/CoLin.
Authors:Siddhant Katyan, Marc-André Gardner, Jean-François Lalonde
Abstract:
LiDAR sensors are a key modality for 3D perception, yet they are typically designed independently of downstream tasks such as point cloud registration. Conventional registration operates on pre‑acquired datasets with fixed LiDAR configurations, leading to suboptimal data collection and significant computational overhead for sampling, noise filtering, and parameter tuning. In this work, we propose an adaptive LiDAR sensing framework that dynamically adjusts sensor parameters, jointly optimizing LiDAR acquisition and registration hyperparameters. By integrating registration feedback into the sensing loop, our approach optimally balances point density, noise, and sparsity, improving registration accuracy and efficiency. Evaluations in the CARLA simulation demonstrate that our method outperforms fixed‑parameter baselines while retaining generalization abilities, highlighting the potential of adaptive LiDAR for autonomous perception and robotic applications.
Authors:Xi Chen, Arian Maleki, Shirin Jalali
Abstract:
In coherent imaging, speckle is statistically modeled as multiplicative noise, posing a fundamental challenge for image reconstruction. While maximum likelihood estimation (MLE) provides a principled framework for speckle mitigation, its application to coherent imaging system such as digital holography with finite apertures is hindered by the prohibitive cost of high‑dimensional matrix inversion, especially at high resolutions. This computational burden has prevented the use of MLE‑based reconstruction with physically accurate aperture modeling. In this work, we propose a randomized linear algebra approach that enables scalable MLE optimization without explicit matrix inversions in gradient computation. By exploiting the structural properties of sensing matrix and using conjugate gradient for likelihood gradient evaluation, the proposed algorithm supports accurate aperture modeling without the simplifying assumptions commonly imposed for tractability. We term the resulting method projected gradient descent with Monte Carlo estimation (PGD‑MC). The proposed PGD‑MC framework (i) demonstrates robustness to diverse and physically accurate aperture models, (ii) achieves substantial improvements in reconstruction quality and computational efficiency, and (iii) scales effectively to high‑resolution digital holography. Extensive experiments incorporating three representative denoisers as regularization show that PGD‑MC provides a flexible and effective MLE‑based reconstruction framework for digital holography with finite apertures, consistently outperforming prior Plug‑and‑Play model‑based iterative reconstruction methods in both accuracy and speed. Our code is available at: https://github.com/Computational‑Imaging‑RU/MC_Maximum_Likelihood_Digital_Holography_Speckle.
Authors:Qingwu Liu, Nicolas Saunier, Guillaume-Alexandre Bilodeau
Abstract:
This study introduces a new object detection dataset of pedestrians using mobility aids, named PMMA. The dataset was collected in an outdoor environment, where volunteers used wheelchairs, canes, and walkers, resulting in nine categories of pedestrians: pedestrians, cane users, two types of walker users, whether walking or resting, five types of wheelchair users, including wheelchair users, people pushing empty wheelchairs, and three types of users pushing occupied wheelchairs, including the entire pushing group, the pusher and the person seated on the wheelchair. To establish a benchmark, seven object detection models (Faster R‑CNN, CenterNet, YOLOX, DETR, Deformable DETR, DINO, and RT‑DETR) and three tracking algorithms (ByteTrack, BOT‑SORT, and OC‑SORT) were implemented under the MMDetection framework. Experimental results show that YOLOX, Deformable DETR, and Faster R‑CNN achieve the best detection performance, while the differences among the three trackers are relatively small. The PMMA dataset is publicly available at https://doi.org/10.5683/SP3/XJPQUG, and the video processing and model training code is available at https://github.com/DatasetPMMA/PMMA.
Authors:Dominik Galus, Julia Farganus, Tymoteusz Zapala, Mikołaj Czachorowski, Piotr Borycki, Przemysław Spurek, Piotr Syga
Abstract:
3D Gaussian Splatting (3DGS) has rapidly become a standard for high‑fidelity 3D reconstruction, yet its adoption in multiple critical domains is hindered by the lack of interpretability of the generation models as well as classification of the Splats. While explainability methods exist for other 3D representations, like point clouds, they typically rely on ambiguous saliency maps that fail to capture the volumetric coherence of Gaussian primitives. We introduce XSPLAIN, the first ante‑hoc, prototype‑based interpretability framework designed specifically for 3DGS classification. Our approach leverages a voxel‑aggregated PointNet backbone and a novel, invertible orthogonal transformation that disentangles feature channels for interpretability while strictly preserving the original decision boundaries. Explanations are grounded in representative training examples, enabling intuitive ``this looks like that'' reasoning without any degradation in classification performance. A rigorous user study (N=51) demonstrates a decisive preference for our approach: participants selected XSPLAIN explanations 48.4% of the time as the best, significantly outperforming baselines (p<0.001), showing that XSPLAIN provides transparency and user trust. The source code for this work is available at: https://github.com/Solvro/ml‑splat‑xai
Authors:Jiacheng Hou, Yining Sun, Ruochong Jin, Haochen Han, Fangming Liu, Wai Kin Victor Chan, Alex Jinpeng Wang
Abstract:
Recent advances in large image editing models have shifted the paradigm from text‑driven instructions to vision‑prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual‑text prompts. While this paradigm greatly expands usability, it also introduces a critical and underexplored safety risk: the attack surface itself becomes visual. In this work, we propose Vision‑Centric Jailbreak Attack (VJA), the first visual‑to‑visual jailbreak attack that conveys malicious instructions purely through visual inputs. To systematically study this emerging threat, we introduce IESBench, a safety‑oriented benchmark for image editing models. Extensive experiments on IESBench demonstrate that VJA effectively compromises state‑of‑the‑art commercial models, achieving attack success rates of up to 80.9% on Nano Banana Pro and 70.1% on GPT‑Image‑1.5. To mitigate this vulnerability, we propose a training‑free defense based on introspective multimodal reasoning, which substantially improves the safety of poorly aligned models to a level comparable with commercial systems, without auxiliary guard models and with negligible computational overhead. Our findings expose new vulnerabilities, provide both a benchmark and practical defense to advance safe and trustworthy modern image editing systems. Warning: This paper contains offensive images created by large image editing models.
Authors:Mingyang Wu, Ashirbad Mishra, Soumik Dey, Shuo Xing, Naveen Ravipati, Hansi Wu, Binbin Li, Zhengzhong Tu
Abstract:
Image‑to‑Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine‑grained object identity under changing viewpoints remains a persistent challenge. Unlike text‑to‑video models, existing I2V pipelines often suffer from appearance drift and geometric distortion, artifacts we attribute to the sparsity of single‑view 2D observations and weak cross‑modal alignment. Here we address this problem from both data and model perspectives. First, we curate ConsIDVid, a large‑scale object‑centric dataset built with a scalable pipeline for high‑quality, temporally aligned videos, and establish ConsIDVid‑Bench, where we present a novel benchmarking and evaluation framework for multi‑view consistency using metrics sensitive to subtle geometric and appearance deviations. We further propose ConsID‑Gen, a view‑assisted I2V generation framework that augments the first frame with unposed auxiliary views and fuses semantic and structural cues via a dual‑stream visual‑geometric encoder as well as a text‑visual connector, yielding unified conditioning for a Diffusion Transformer backbone. Experiments across ConsIDVid‑Bench demonstrate that ConsID‑Gen consistently outperforms in multiple metrics, with the best overall performance surpassing leading video generation models like Wan2.1 and HunyuanVideo, delivering superior identity fidelity and temporal coherence under challenging real‑world scenarios. We will release our model and dataset at https://myangwu.github.io/ConsID‑Gen.
Authors:Yuxin Jiang, Yuchao Gu, Ivor W. Tsang, Mike Zheng Shou
Abstract:
Scaling action‑controllable world models is limited by the scarcity of action labels. While latent action learning promises to extract control interfaces from unlabeled video, learned latents often fail to transfer across contexts: they entangle scene‑specific cues and lack a shared coordinate system. This occurs because standard objectives operate only within each clip, providing no mechanism to align action semantics across contexts. Our key insight is that although actions are unobserved, their semantic effects are observable and can serve as a shared reference. We introduce SeqΔ‑REPA, a sequence‑level control‑effect alignment objective that anchors integrated latent action to temporal feature differences from a frozen, self‑supervised video encoder. Building on this, we present Olaf‑World, a pipeline that pretrains action‑conditioned video world models from large‑scale passive video. Extensive experiments demonstrate that our method learns a more structured latent action space, leading to stronger zero‑shot action transfer and more data‑efficient adaptation to new control interfaces than state‑of‑the‑art baselines.
Authors:Zhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo, Yao Zhao, Bingyi Kang, Jiashi Feng, Xiaojie Jin
Abstract:
Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the first investigation into learning transferable knowledge directly from raw real‑world videos. At its core, VideoWorld 2 introduces a dynamic‑enhanced Latent Dynamics Model (dLDM) that decouples action dynamics from visual appearance: a pretrained video diffusion model handles visual appearance modeling, enabling the dLDM to learn latent codes that focus on compact and meaningful task‑related dynamics. These latent codes are then modeled autoregressively to learn task policies and support long‑horizon reasoning. We evaluate VideoWorld 2 on challenging real‑world handcraft making tasks, where prior video generation and latent‑dynamics models struggle to operate reliably. Remarkably, VideoWorld 2 achieves up to 70% improvement in task success rate and produces coherent long execution videos. In robotics, we show that VideoWorld 2 can acquire effective manipulation knowledge from the Open‑X dataset, which substantially improves task performance on CALVIN. This study reveals the potential of learning transferable world knowledge directly from raw videos, with all code, data, and models to be open‑sourced for further research.
Authors:Amandeep Kumar, Vishal M. Patel
Abstract:
Leveraging representation encoders for generative modeling offers a path for efficient, high‑fidelity synthesis. However, standard diffusion transformers fail to converge on these representations directly. While recent work attributes this to a capacity bottleneck proposing computationally expensive width scaling of diffusion transformers we demonstrate that the failure is fundamentally geometric. We identify Geometric Interference as the root cause: standard Euclidean flow matching forces probability paths through the low‑density interior of the hyperspherical feature space of representation encoders, rather than following the manifold surface. To resolve this, we propose Riemannian Flow Matching with Jacobi Regularization (RJF). By constraining the generative process to the manifold geodesics and correcting for curvature‑induced error propagation, RJF enables standard Diffusion Transformer architectures to converge without width scaling. Our method RJF enables the standard DiT‑B architecture (131M parameters) to converge effectively, achieving an FID of 3.37 where prior methods fail to converge. Code: https://github.com/amandpkr/RJF
Authors:Florian Hahlbohm, Linus Franke, Martin Eisemann, Marcus Magnor
Abstract:
Recent advances in 3D Gaussian Splatting (3DGS) have focused on accelerating optimization while preserving reconstruction quality. However, many proposed methods entangle implementation‑level improvements with fundamental algorithmic modifications or trade performance for fidelity, leading to a fragmented research landscape that complicates fair comparison. In this work, we consolidate and evaluate the most effective and broadly applicable strategies from prior 3DGS research and augment them with several novel optimizations. We further investigate underexplored aspects of the framework, including numerical stability, Gaussian truncation, and gradient approximation. The resulting system, Faster‑GS, provides a rigorously optimized algorithm that we evaluate across a comprehensive suite of benchmarks. Our experiments demonstrate that Faster‑GS achieves up to 5× faster training while maintaining visual quality, establishing a new cost‑effective and resource efficient baseline for 3DGS optimization. Furthermore, we demonstrate that optimizations can be applied to 4D Gaussian reconstruction, leading to efficient non‑rigid scene optimization.
Authors:Zongrui Li, Xinhua Ma, Minghui Hu, Yunqing Zhao, Yingchen Yu, Qian Zheng, Chang Liu, Xudong Jiang, Song Bai
Abstract:
Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict normal maps. However, they often suffer from 3D misalignment: while the estimated normal maps may appear to have a correct appearance, the reconstructed surfaces often fail to align with the geometric details. We argue that this misalignment stems from the current paradigm: the model struggles to distinguish and reconstruct varying geometry represented in normal maps, as the differences in underlying geometry are reflected only through relatively subtle color variations. To address this issue, we propose a new paradigm that reformulates normal estimation as shading sequence estimation, where shading sequences are more sensitive to various geometric information. Building on this paradigm, we present RoSE, a method that leverages image‑to‑video generative models to predict shading sequences. The predicted shading sequences are then converted into normal maps by solving a simple ordinary least‑squares problem. To enhance robustness and better handle complex objects, RoSE is trained on a synthetic dataset, MultiShade, with diverse shapes, materials, and light conditions. Experiments demonstrate that RoSE achieves state‑of‑the‑art performance on real‑world benchmark datasets for object‑based monocular normal estimation.
Authors:Shaoqiu Zhang, Zizhong Ding, Kaicheng Yang, Junyi Wu, Xianglong Yan, Xi Li, Bingnan Duan, Jianping Fang, Yulun Zhang
Abstract:
Diffusion Transformers (DiTs) have emerged as the state‑of‑the‑art backbone for high‑fidelity image and video generation. However, their massive computational cost and memory footprint hinder deployment on edge devices. While post‑training quantization (PTQ) has proven effective for large language models (LLMs), directly applying existing methods to DiTs yields suboptimal results due to the neglect of the unique temporal dynamics inherent in diffusion processes. In this paper, we propose AdaTSQ, a novel PTQ framework that pushes the Pareto frontier of efficiency and quality by exploiting the temporal sensitivity of DiTs. First, we propose a Pareto‑aware timestep‑dynamic bit‑width allocation strategy. We model the quantization policy search as a constrained pathfinding problem. We utilize a beam search algorithm guided by end‑to‑end reconstruction error to dynamically assign layer‑wise bit‑widths across different timesteps. Second, we propose a Fisher‑guided temporal calibration mechanism. It leverages temporal Fisher information to prioritize calibration data from highly sensitive timesteps, seamlessly integrating with Hessian‑based weight optimization. Extensive experiments on four advanced DiTs (e.g., Flux‑Dev, Flux‑Schnell, Z‑Image, and Wan2.1) demonstrate that AdaTSQ significantly outperforms state‑of‑the‑art methods like SVDQuant and ViDiT‑Q. Our code will be released at https://github.com/Qiushao‑E/AdaTSQ.
Authors:Yuhao Zheng, Li'an Zhong, Yi Wang, Rui Dai, Kaikui Liu, Xiangxiang Chu, Linyuan Lv, Philip Torr, Kevin Qinghong Lin
Abstract:
Autonomous GUI agents interact with environments by perceiving interfaces and executing actions. As a virtual sandbox, the GUI World model empowers agents with human‑like foresight by enabling action‑conditioned prediction. However, existing text‑ and pixel‑based approaches struggle to simultaneously achieve high visual fidelity and fine‑grained structural controllability. To this end, we propose Code2World, a vision‑language coder that simulates the next visual state via renderable code generation. Specifically, to address the data scarcity problem, we construct AndroidCode by translating GUI trajectories into high‑fidelity HTML and refining synthesized code through a visual‑feedback revision mechanism, yielding a corpus of over 80K high‑quality screen‑action pairs. To adapt existing VLMs into code prediction, we first perform SFT as a cold start for format layout following, then further apply Render‑Aware Reinforcement Learning which uses rendered outcome as the reward signal by enforcing visual semantic fidelity and action consistency. Extensive experiments demonstrate that Code2World‑8B achieves the top‑performing next UI prediction, rivaling the competitive GPT‑5 and Gemini‑3‑Pro‑Image. Notably, Code2World significantly enhances downstream navigation success rates in a flexible manner, boosting Gemini‑2.5‑Flash by +9.5% on AndroidWorld navigation. The code is available at https://github.com/AMAP‑ML/Code2World.
Authors:Peng Chen, Chao Huang, Yunkang Cao, Chengliang Liu, Wei Wang, Wenqiang Wang, Mingbo Yang, Li Shen, Wenqi Ren, Xiaochun Cao
Abstract:
Industrial anomaly detection demands precise reasoning over fine‑grained defect patterns. However, existing multimodal large language models (MLLMs), pretrained on general‑domain data, often struggle to capture category‑specific anomalies, thereby limiting both detection accuracy and interpretability. To address these limitations, we propose Reason‑IAD, a knowledge‑guided dynamic latent reasoning framework for explainable industrial anomaly detection. Reason‑IAD comprises two core components. First, a retrieval‑augmented knowledge module incorporates category‑specific textual descriptions into the model input, enabling context‑aware reasoning over domain‑specific defects. Second, an entropy‑driven latent reasoning mechanism conducts iterative exploration within a compact latent space using optimizable latent think tokens, guided by an entropy‑based reward that encourages confident and stable predictions. Furthermore, a dynamic visual injection strategy selectively incorporates the most informative image patches into the latent sequence, directing the reasoning process toward regions critical for anomaly detection. Extensive experimental results demonstrate that Reason‑IAD consistently outperforms state‑of‑the‑art methods across multiple tasks. The code will be publicly available at https://github.com/chenpeng052/Reason‑IAD.
Authors:Boya Wang, Ruizhe Li, Chao Chen, Xin Chen
Abstract:
Liver fibrosis poses a substantial challenge in clinical practice, emphasizing the necessity for precise liver segmentation and accurate disease staging. Based on the CARE Liver 2025 Track 4 Challenge, this study introduces a multi‑task deep learning framework developed for liver segmentation (LiSeg) and liver fibrosis staging (LiFS) using multiparametric MRI. The LiSeg phase addresses the challenge of limited annotated images and the complexities of multi‑parametric MRI data by employing a semi‑supervised learning model that integrates image segmentation and registration. By leveraging both labeled and unlabeled data, the model overcomes the difficulties introduced by domain shifts and variations across modalities. In the LiFS phase, we employed a patchbased method which allows the visualization of liver fibrosis stages based on the classification outputs. Our approach effectively handles multimodality imaging data, limited labels, and domain shifts. The proposed method has been tested by the challenge organizer on an independent test set that includes in‑distribution (ID) and out‑of‑distribution (OOD) cases using three‑channel MRIs (T1, T2, DWI) and seven‑channel MRIs (T1, T2, DWI, GED1‑GED4). The code is freely available. Github link: https://github.com/mileywang3061/Care‑Liver
Authors:Deyang Jiang, Jing Huang, Xuanle Zhao, Lei Chen, Liming Zheng, Fanfan Liu, Haibo Qiu, Peng Shi, Zhixiong Zeng
Abstract:
Effectively scaling GUI automation is essential for computer‑use agents (CUAs); however, existing work primarily focuses on scaling GUI grounding rather than the more crucial GUI planning, which requires more sophisticated data collection. In reality, the exploration process of a CUA across apps/desktops/web pages typically follows a tree structure, with earlier functional entry points often being explored more frequently. Thus, organizing large‑scale trajectories into tree structures can reduce data cost and streamline the data scaling of GUI planning. In this work, we propose TreeCUA to efficiently scale GUI automation with tree‑structured verifiable evolution. We propose a multi‑agent collaborative framework to explore the environment, verify actions, summarize trajectories, and evaluate quality to generate high‑quality and scalable GUI trajectories. To improve efficiency, we devise a novel tree‑based topology to store and replay duplicate exploration nodes, and design an adaptive exploration algorithm to balance the depth (\emphi.e., trajectory difficulty) and breadth (\emphi.e., trajectory diversity). Moreover, we develop world knowledge guidance and global memory backtracking to avoid low‑quality generation. Finally, we naturally extend and propose the TreeCUA‑DPO method from abundant tree node information, improving GUI planning capability by referring to the branch information of adjacent trajectories. Experimental results show that TreeCUA and TreeCUA‑DPO offer significant improvements, and out‑of‑domain (OOD) studies further demonstrate strong generalization. All trajectory node information and code will be available at https://github.com/UITron‑hub/TreeCUA.
Authors:Jiayi Lyu, Leigang Qu, Wenjing Zhang, Hanyu Jiang, Kai Liu, Zhenglin Zhou, Xiaobo Xia, Jian Xue, Tat-Seng Chua
Abstract:
Realistic talking‑head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nuanced emotional expressions due to the lack of fine‑grained emotion control. To address this issue, we introduce a novel two‑stage method (AUHead) to disentangle fine‑grained emotion control, i.e. , Action Units (AUs), from audio and achieve controllable generation. In the first stage, we explore the AU generation abilities of large audio‑language models (ALMs), by spatial‑temporal AU tokenization and an "emotion‑then‑AU" chain‑of‑thought mechanism. It aims to disentangle AUs from raw speech, effectively capturing subtle emotional cues. In the second stage, we propose an AU‑driven controllable diffusion model that synthesizes realistic talking‑head videos conditioned on AU sequences. Specifically, we first map the AU sequences into the structured 2D facial representation to enhance spatial fidelity, and then model the AU‑vision interaction within cross‑attention modules. To achieve flexible AU‑quality trade‑off control, we introduce an AU disentanglement guidance strategy during inference, further refining the emotional expressiveness and identity consistency of the generated videos. Results on benchmark datasets demonstrate that our approach achieves competitive performance in emotional realism, accurate lip synchronization, and visual coherence, significantly surpassing existing techniques. Our implementation is available at https://github.com/laura990501/AUHead_ICLR
Authors:Hung-Shuo Chang, Yue-Cheng Yang, Yu-Hsi Chen, Wei-Hsin Chen, Chien-Yao Wang, James C. Liao, Chien-Chang Chen, Hen-Hsen Huang, Hong-Yuan Mark Liao
Abstract:
Analyzing animal and human behavior has long been a challenging task in computer vision. Early approaches from the 1970s to the 1990s relied on hand‑crafted edge detection, segmentation, and low‑level features such as color, shape, and texture to locate objects and infer their identities‑an inherently ill‑posed problem. Behavior analysis in this era typically proceeded by tracking identified objects over time and modeling their trajectories using sparse feature points, which further limited robustness and generalization. A major shift occurred with the introduction of ImageNet by Deng and Li in 2010, which enabled large‑scale visual recognition through deep neural networks and effectively served as a comprehensive visual dictionary. This development allowed object recognition to move beyond complex low‑level processing toward learned high‑level representations. In this work, we follow this paradigm to build a large‑scale Universal Action Space (UAS) using existing labeled human‑action datasets. We then use this UAS as the foundation for analyzing and categorizing mammalian and chimpanzee behavior datasets. The source code is released on GitHub at https://github.com/franktpmvu/Universal‑Action‑Space.
Authors:Lin Chen, Xiaoke Zhao, Kun Ding, Weiwei Feng, Changtao Miao, Zili Wang, Wenxuan Guo, Ying Wang, Kaiyuan Zheng, Bo Zhang, Zhe Li, Shiming Xiang
Abstract:
Multimodal Large Language Models (MLLMs) demonstrate impressive cross‑modal capabilities, yet their substantial size poses significant deployment challenges. Knowledge distillation (KD) is a promising solution for compressing these models, but existing methods primarily rely on static next‑token alignment, neglecting the dynamic token interactions, which embed essential capabilities for multimodal understanding and generation. To this end, we introduce Align‑TI, a novel KD framework designed from the perspective of Token Interactions. Our approach is motivated by the insight that MLLMs rely on two primary interactions: vision‑instruction token interactions to extract relevant visual information, and intra‑response token interactions for coherent generation. Accordingly, Align‑TI introduces two components: IVA enables the student model to imitate the teacher's instruction‑relevant visual information extract capability by aligning on salient visual regions. TPA captures the teacher's dynamic generative logic by aligning the sequential token‑to‑token transition probabilities. Extensive experiments demonstrate Align‑TI's superiority. Notably, our approach achieves 2.6% relative improvement over Vanilla KD, and our distilled Align‑TI‑2B even outperforms LLaVA‑1.5‑7B (a much larger MLLM) by 7.0%, establishing a new state‑of‑the‑art distillation framework for training parameter‑efficient MLLMs. Code is available at https://github.com/lchen1019/Align‑TI.
Authors:Chuanhai Zang, Jiabao Hu, XW Song
Abstract:
Synthetic data provide low‑cost, accurately annotated samples for geometry‑sensitive vision tasks, but appearance and imaging differences between synthetic and real domains cause severe domain shift and degrade downstream performance. Unpaired synthetic‑to‑real translation can reduce this gap without paired supervision, yet existing methods often face a trade‑off between photorealism and structural stability: unconstrained generation may introduce deformation or spurious textures, while overly rigid constraints limit adaptation to real‑domain statistics. We propose FD‑DB, a frequency‑decoupled dual‑branch model that separates appearance transfer into low‑frequency interpretable editing and high‑frequency residual compensation. The interpretable branch predicts physically meaningful editing parameters (white balance, exposure, contrast, saturation, blur, and grain) to build a stable low‑frequency appearance base with strong content preservation. The free branch complements fine details through residual generation, and a gated fusion mechanism combines the two branches under explicit frequency constraints to limit low‑frequency drift. We further adopt a two‑stage training schedule that first stabilizes the editing branch and then releases the residual branch to improve optimization stability. Experiments on the YCB‑V dataset show that FD‑DB improves real‑domain appearance consistency and significantly boosts downstream semantic segmentation performance while preserving geometric and semantic structures.
Authors:James Burgess, Rameen Abdal, Dan Stoddart, Sergey Tulyakov, Serena Yeung-Levy, Kuan-Chieh Jackson Wang
Abstract:
Modern image generators produce strikingly realistic images, where only artifacts like distorted hands or warped objects reveal their synthetic origin. Detecting these artifacts is essential: without detection, we cannot benchmark generators or train reward models to improve them. Current detectors fine‑tune VLMs on tens of thousands of labeled images, but this is expensive to repeat whenever generators evolve or new artifact types emerge. We show that pretrained VLMs already encode the knowledge needed to detect artifacts ‑ with the right scaffolding, this capability can be unlocked using only a few hundred labeled examples per artifact category. Our system, ArtifactLens, achieves state‑of‑the‑art on five human artifact benchmarks (the first evaluation across multiple datasets) while requiring orders of magnitude less labeled data. The scaffolding consists of a multi‑component architecture with in‑context learning and text instruction optimization, with novel improvements to each. Our methods generalize to other artifact types ‑ object morphology, animal anatomy, and entity interactions ‑ and to the distinct task of AIGC detection.
Authors:Zhikai Li, Jiatong Li, Xuewen Liu, Wangbo Zhao, Pan Du, Kaicheng Zhou, Qingyi Gu, Yang You, Zhen Dong, Kurt Keutzer
Abstract:
The rapid development of visual generative models raises the need for more scalable and human‑aligned evaluation methods. While the crowdsourced Arena platforms offer human preference assessments by collecting human votes, they are costly and time‑consuming, inherently limiting their scalability. Leveraging vision‑language model (VLMs) as substitutes for manual judgments presents a promising solution. However, the inherent hallucinations and biases of VLMs hinder alignment with human preferences, thus compromising evaluation reliability. Additionally, the static evaluation approach lead to low efficiency. In this paper, we propose K‑Sort Eval, a reliable and efficient VLM‑based evaluation framework that integrates posterior correction and dynamic matching. Specifically, we curate a high‑quality dataset from thousands of human votes in K‑Sort Arena, with each instance containing the outputs and rankings of K models. When evaluating a new model, it undergoes (K+1)‑wise free‑for‑all comparisons with existing models, and the VLM provide the rankings. To enhance alignment and reliability, we propose a posterior correction method, which adaptively corrects the posterior probability in Bayesian updating based on the consistency between the VLM prediction and human supervision. Moreover, we propose a dynamic matching strategy, which balances uncertainty and diversity to maximize the expected benefit of each comparison, thus ensuring more efficient evaluation. Extensive experiments show that K‑Sort Eval delivers evaluation results consistent with K‑Sort Arena, typically requiring fewer than 90 model runs, demonstrating both its efficiency and reliability.
Authors:Venus Team, Changlong Gao, Zhangxuan Gu, Yulin Liu, Xinyu Qiu, Shuheng Shen, Yue Wen, Tianyu Xia, Zhenyu Xu, Zhengwen Zeng, Beitong Zhou, Xingran Zhou, Weizhi Chen, Sunhao Dai, Jingya Dou, Yichen Gong, Yuan Guo, Zhenlin Guo, Feng Li, Qian Li, Jinzhen Lin, Yuqi Zhou, Linchao Zhu, Liang Chen, Zhenyu Guo, Changhua Meng, Weiqiang Wang
Abstract:
GUI agents have emerged as a powerful paradigm for automating interactions in digital environments, yet achieving both broad generality and consistently strong task performance remains challenging. In this report, we present UI‑Venus‑1.5, a unified, end‑to‑end GUI Agent designed for robust real‑world applications. The proposed model family comprises two dense variants (2B and 8B) and one mixture‑of‑experts variant (30B‑A3B) to meet various downstream application scenarios. Compared to our previous version, UI‑Venus‑1.5 introduces three key technical advances: (1) a comprehensive Mid‑Training stage leveraging 10 billion tokens across 30+ datasets to establish foundational GUI semantics; (2) Online Reinforcement Learning with full‑trajectory rollouts, aligning training objectives with long‑horizon, dynamic navigation in large‑scale environments; and (3) a single unified GUI Agent constructed via Model Merging, which synthesizes domain‑specific models (grounding, web, and mobile) into one cohesive checkpoint. Extensive evaluations demonstrate that UI‑Venus‑1.5 establishes new state‑of‑the‑art performance on benchmarks such as ScreenSpot‑Pro (69.6%), VenusBench‑GD (75.0%), and AndroidWorld (77.6%), significantly outperforming previous strong baselines. In addition, UI‑Venus‑1.5 demonstrates robust navigation capabilities across a variety of Chinese mobile apps, effectively executing user instructions in real‑world scenarios. Code: https://github.com/inclusionAI/UI‑Venus; Model: https://huggingface.co/collections/inclusionAI/ui‑venus
Authors:Jiahao Qin
Abstract:
Cross‑domain image registration requires aligning images acquired under heterogeneous imaging physics, where the classical brightness constancy assumption is fundamentally violated. We formulate this problem through an image formation model I = R(s, a) + epsilon, where each observation is generated by a rendering function R acting on domain‑invariant scene structure s and domain‑specific appearance statistics a. Registration then reduces to an inverse rendering problem: given observations from two domains, recover the shared structure and re‑render it under the target appearance to obtain the registered output. We instantiate this framework as SAS‑Net (Scene‑Appearance Separation Network), where instance normalization implements the structure‑appearance decomposition and Adaptive Instance Normalization (AdaIN) realizes the differentiable forward renderer. A scene consistency loss enforces geometric correspondence in the factorized latent space. Experiments on EuroSAT‑Reg‑256 (satellite remote sensing) and FIRE‑Reg‑256 (retinal fundus) demonstrate state‑of‑the‑art performance across heterogeneous imaging domains. SAS‑Net (3.35M parameters) achieves 89 FPS on an RTX 5090 GPU. Code:
https://github.com/D‑ST‑Sword/SAS‑Net.
Authors:Hao Phung, Hadar Averbuch-Elor
Abstract:
Reconstructing a structured vector‑graphics representation from a rasterized floorplan image is typically an important prerequisite for computational tasks involving floorplans such as automated understanding or CAD workflows. However, existing techniques struggle in faithfully generating the structure and semantics conveyed by complex floorplans that depict large indoor spaces with many rooms and a varying numbers of polygon corners. To this end, we propose Raster2Seq, framing floorplan reconstruction as a sequence‑to‑sequence task in which floorplan elements‑‑such as rooms, windows, and doors‑‑are represented as labeled polygon sequences that jointly encode geometry and semantics. Our approach introduces an autoregressive decoder that learns to predict the next corner conditioned on image features and previously generated corners using guidance from learnable anchors. These anchors represent spatial coordinates in image space, hence allowing for effectively directing the attention mechanism to focus on informative image regions. By embracing the autoregressive mechanism, our method offers flexibility in the output format, enabling for efficiently handling complex floorplans with numerous rooms and diverse polygon structures. Our method achieves state‑of‑the‑art performance on standard benchmarks such as Structure3D, CubiCasa5K, and Raster2Graph, while also demonstrating strong generalization to more challenging datasets like WAFFLE, which contain diverse room structures and complex geometric variations.
Authors:Haodong Li, Jingwei Wu, Quan Sun, Guopeng Li, Juanxi Tian, Huanyu Zhang, Yanlin Lai, Ruichuan An, Hongbo Peng, Yuhong Dai, Chenxi Li, Chunmei Qing, Jia Wang, Ziyang Meng, Zheng Ge, Xiangyu Zhang, Daxin Jiang
Abstract:
Recent advancements in image generation models have enabled the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving the evaluation of state transitions and temporal coherence in GUI‑specific contexts underexplored. To address this gap, we introduce GEBench, a comprehensive benchmark for evaluating dynamic interaction and temporal coherence in GUI generation. GEBench comprises 700 carefully curated samples spanning five task categories, covering both single‑step interactions and multi‑step trajectories across real‑world and fictional scenarios, as well as grounding point localization. To support systematic evaluation, we propose GE‑Score, a novel five‑dimensional metric that assesses Goal Achievement, Interaction Logic, Content Consistency, UI Plausibility, and Visual Quality. Extensive evaluations on current models indicate that while they perform well on single‑step transitions, they struggle significantly with maintaining temporal coherence and spatial grounding over longer interaction sequences. Our findings identify icon interpretation, text rendering, and localization precision as critical bottlenecks. This work provides a foundation for systematic assessment and suggests promising directions for future research toward building high‑fidelity generative GUI environments. The code is available at: https://github.com/stepfun‑ai/GEBench.
Authors:Guangxun Zhu, Xuan Liu, Nicolas Pugeault, Chongfeng Wei, Edmond S. L. Ho
Abstract:
Accurately predicting pedestrian motion is crucial for safe and reliable autonomous driving in complex urban environments. In this work, we present a 3D vehicle‑conditioned pedestrian pose forecasting framework that explicitly incorporates surrounding vehicle information. To support this, we enhance the Waymo‑3DSkelMo dataset with aligned 3D vehicle bounding boxes, enabling realistic modeling of multi‑agent pedestrian‑vehicle interactions. We introduce a sampling scheme to categorize scenes by pedestrian and vehicle count, facilitating training across varying interaction complexities. Our proposed network adapts the TBIFormer architecture with a dedicated vehicle encoder and pedestrian‑vehicle interaction cross‑attention module to fuse pedestrian and vehicle features, allowing predictions to be conditioned on both historical pedestrian motion and surrounding vehicles. Extensive experiments demonstrate substantial improvements in forecasting accuracy and validate different approaches for modeling pedestrian‑vehicle interactions, highlighting the importance of vehicle‑aware 3D pose prediction for autonomous driving. Code is available at: https://github.com/GuangxunZhu/VehCondPose3D
Authors:Ruijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han, Jianfei Cai, Ying Shan, Chuanxia Zheng
Abstract:
We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint representation of dense 3D point maps and 3D scene flows in a shared coordinate system, together with a 4D VAE tailored to learn this representation effectively. Unlike prior work that strictly aligns 3D values and latents with RGB VAE latents‑despite their fundamentally different distributions‑we show that such alignment is unnecessary and can hurt performance. Instead, we propose a new data normalization and VAE training strategy that better transfers diffusion priors and greatly improves reconstruction quality. Extensive experiments on multiple datasets show that MotionCrafter achieves state‑of‑the‑art performance in both geometry reconstruction and dense scene flow estimation, delivering 38.64% and 25.0% improvements in geometry and motion reconstruction, respectively, all without any post‑optimization. Project page: https://ruijiezhu94.github.io/MotionCrafter_Page
Authors:Hao Tan, Jun Lan, Senyuan Shi, Zichang Tan, Zijian Yu, Huijia Zhu, Weiqiang Wang, Jun Wan, Zhen Lei
Abstract:
The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce VideoVeritas, a framework that integrates fine‑grained perception and fact‑based reasoning. We observe that while current multi‑modal large language models (MLLMs) exhibit strong reasoning capacity, their granular perception ability remains limited. To mitigate this, we introduce Joint Preference Alignment and Perception Pretext Reinforcement Learning (PPRL). Specifically, rather than directly optimizing for detection task, we adopt general spatiotemporal grounding and self‑supervised object counting in the RL stage, enhancing detection performance with simple perception pretext tasks. To facilitate robust evaluation, we further introduce MintVid, a light yet high‑quality dataset containing 3K videos from 9 state‑of‑the‑art generators, along with a real‑world collected subset that has factual errors in content. Experimental results demonstrate that existing methods tend to bias towards either superficial reasoning or mechanical analysis, while VideoVeritas achieves more balanced performance across diverse benchmarks.
Authors:Hao Yang, Zhiyu Tan, Jia Gong, Luozheng Qin, Hesen Chen, Xiaomeng Yang, Yuqing Sun, Yuetan Lin, Mengping Yang, Hao Li
Abstract:
We present Omni‑Video 2, a scalable and computationally efficient model that connects pretrained multimodal large‑language models (MLLMs) with video diffusion models for unified video generation and editing. Our key idea is to exploit the understanding and reasoning capabilities of MLLMs to produce explicit target captions to interpret user instructions. In this way, the rich contextual representations from the understanding model are directly used to guide the generative process, thereby improving performance on complex and compositional editing. Moreover, a lightweight adapter is developed to inject multimodal conditional tokens into pretrained text‑to‑video diffusion models, allowing maximum reuse of their powerful generative priors in a parameter‑efficient manner. Benefiting from these designs, we scale up Omni‑Video 2 to a 14B video diffusion model on meticulously curated training data with quality, supporting high quality text‑to‑video generation and various video editing tasks such as object removal, addition, background change, complex motion editing, \emphetc. We evaluate the performance of Omni‑Video 2 on the FiVE benchmark for fine‑grained video editing and the VBench benchmark for text‑to‑video generation. The results demonstrate its superior ability to follow complex compositional instructions in video editing, while also achieving competitive or superior quality in video generation tasks.
Authors:OpenMOSS Team, Donghua Yu, Mingshu Chen, Qi Chen, Qi Luo, Qianyi Wu, Qinyuan Cheng, Ruixiao Li, Tianyi Liang, Wenbo Zhang, Wenming Tu, Xiangyu Peng, Yang Gao, Yanru Huo, Ying Zhu, Yinze Luo, Yiyang Zhang, Yuerong Song, Zhe Xu, Zhiyu Zhang, Chenchen Yang, Cheng Chang, Chushu Zhou, Hanfu Chen, Hongnan Ma, Jiaxi Li, Jingqi Tong, Junxi Liu, Ke Chen, Shimin Li, Shiqi Jiang, Songlin Wang, Wei Jiang, Zhaoye Fei, Zhiyuan Ning, Chunguo Li, Chenhui Li, Ziwei He, Zengfeng Huang, Xie Chen, Xipeng Qiu
Abstract:
Audio is indispensable for real‑world video, yet generation models have largely overlooked audio components. Current approaches to producing audio‑visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and degrade overall quality. While systems such as Veo 3 and Sora 2 emphasize the value of simultaneous generation, joint multimodal modeling introduces unique challenges in architecture, data, and training. Moreover, the closed‑source nature of existing systems limits progress in the field. In this work, we introduce MOVA (MOSS Video and Audio), an open‑source model capable of generating high‑quality, synchronized audio‑visual content, including realistic lip‑synced speech, environment‑aware sound effects, and content‑aligned music. MOVA employs a Mixture‑of‑Experts (MoE) architecture, with a total of 32B parameters, of which 18B are active during inference. It supports IT2VA (Image‑Text to Video‑Audio) generation task. By releasing the model weights and code, we aim to advance research and foster a vibrant community of creators. The released codebase features comprehensive support for efficient inference, LoRA fine‑tuning, and prompt enhancement.
Authors:Vineet Kumar Rakesh, Ahana Bhattacharjee, Soumya Mazumdar, Tapas Samanta, Hemendra Kumar Pandey, Amitabha Das, Sarbajit Pal
Abstract:
Talking‑head avatars are increasingly adopted in educational technology to deliver content with social presence and improved engagement. However, many recent talking‑head generation (THG) methods rely on GPU‑centric neural rendering, large training sets, or high‑capacity diffusion models, which limits deployment in offline or resource‑constrained learning environments. A deterministic and CPU‑oriented THG framework is described, termed Symbolic Vedic Computation, that converts speech to a time‑aligned phoneme stream, maps phonemes to a compact viseme inventory, and produces smooth viseme trajectories through symbolic coarticulation inspired by Vedic sutra Urdhva Tiryakbhyam. A lightweight 2D renderer performs region‑of‑interest (ROI) warping and mouth compositing with stabilization to support real‑time synthesis on commodity CPUs. Experiments report synchronization accuracy, temporal stability, and identity consistency under CPU‑only execution, alongside benchmarking against representative CPU‑feasible baselines. Results indicate that acceptable lip‑sync quality can be achieved while substantially reducing computational load and latency, supporting practical educational avatars on low‑end hardware. GitHub: https://vineetkumarrakesh.github.io/vedicthg
Authors:Shanshan Wang, Ziying Feng, Xiaozheng Shen, Xun Yang, Pichao Wang, Zhenwei He, Xingyi Zhang
Abstract:
Source‑Free Domain Adaptation (SFDA) tackles the problem of adapting a pre‑trained source model to an unlabeled target domain without accessing any source data, which is quite suitable for the field of data security. Although recent advances have shown that pseudo‑labeling strategies can be effective, they often fail in fine‑grained scenarios due to subtle inter‑class similarities. A critical but underexplored issue is the presence of asymmetric and dynamic class confusion, where visually similar classes are unequally and inconsistently misclassified by the source model. Existing methods typically ignore such confusion patterns, leading to noisy pseudo‑labels and poor target discrimination. To address this, we propose CLIP‑Guided Alignment(CGA), a novel framework that explicitly models and mitigates class confusion in SFDA. Generally, our method consists of three parts: (1) MCA: detects first directional confusion pairs by analyzing the predictions of the source model in the target domain; (2) MCC: leverages CLIP to construct confusion‑aware textual prompts (e.g. a truck that looks like a bus), enabling more context‑sensitive pseudo‑labeling; and (3) FAM: builds confusion‑guided feature banks for both CLIP and the source model and aligns them using contrastive learning to reduce ambiguity in the representation space. Extensive experiments on various datasets demonstrate that CGA consistently outperforms state‑of‑the‑art SFDA methods, with especially notable gains in confusion‑prone and fine‑grained scenarios. Our results highlight the importance of explicitly modeling inter‑class confusion for effective source‑free adaptation. Our code can be find at https://github.com/soloiro/CGA
Authors:Johannes Thalhammer, Tina Dorosti, Sebastian Peterhansl, Daniela Pfeiffer, Franz Pfeiffer, Florian Schaff
Abstract:
Undersampled CT volumes minimize acquisition time and radiation exposure but introduce artifacts degrading image quality and diagnostic utility. Reducing these artifacts is critical for high‑quality imaging. We propose a computationally efficient hybrid deep‑learning framework that combines the strengths of 2D and 3D models. First, a 2D U‑Net operates on individual slices of undersampled CT volumes to extract feature maps. These slice‑wise feature maps are then stacked across the volume and used as input to a 3D decoder, which utilizes contextual information across slices to predict an artifact‑free 3D CT volume. The proposed two‑stage approach balances the computational efficiency of 2D processing with the volumetric consistency provided by 3D modeling. The results show substantial improvements in inter‑slice consistency in coronal and sagittal direction with low computational overhead. This hybrid framework presents a robust and efficient solution for high‑quality 3D CT image post‑processing. The code of this project can be found on github: https://github.com/J‑3TO/2D‑3DCNN_sparseview/.
Authors:Yongwen Lai, Chaoqun Wang, Shaobo Min
Abstract:
Text‑guided image editing aims to modify specific regions according to the target prompt while preserving the identity of the source image. Recent methods exploit explicit binary masks to constrain editing, but hard mask boundaries introduce artifacts and reduce editability. To address these issues, we propose FusionEdit, a training‑free image editing framework that achieves precise and controllable edits. First, editing and preserved regions are automatically identified by measuring semantic discrepancies between source and target prompts. To mitigate boundary artifacts, FusionEdit performs distance‑aware latent fusion along region boundaries to yield the soft and accurate mask, and employs a total variation loss to enforce smooth transitions, obtaining natural editing results. Second, FusionEdit leverages AdaIN‑based modulation within DiT attention layers to perform a statistical attention fusion in the editing region, enhancing editability while preserving global consistency with the source image. Extensive experiments demonstrate that our FusionEdit significantly outperforms state‑of‑the‑art methods. Code is available at \hrefhttps://github.com/Yvan1001/FusionEdithttps://github.com/Yvan1001/FusionEdit.
Authors:Linli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li, Xinlong Chen, Feifan Song, Ziyue Wang, Kun Ouyang, Yuanxin Liu, Lingpeng Kong, Qi Liu, Pengfei Wan, Kun Gai, Yuanxing Zhang, Xu Sun
Abstract:
This paper proposes Omni Dense Captioning, a novel task designed to generate continuous, fine‑grained, and structured audio‑visual narratives with explicit timestamps. To ensure dense semantic coverage, we introduce a six‑dimensional structural schema to create "script‑like" captions, enabling readers to vividly imagine the video content scene by scene, akin to a cinematographic screenplay. To facilitate research, we construct OmniDCBench, a high‑quality, human‑annotated benchmark, and propose SodaM, a unified metric that evaluates time‑aware detailed descriptions while mitigating scene boundary ambiguity. Furthermore, we construct a training dataset, TimeChatCap‑42K, and present TimeChat‑Captioner‑7B, a strong baseline trained via SFT and GRPO with task‑specific rewards. Extensive experiments demonstrate that TimeChat‑Captioner‑7B achieves state‑of‑the‑art performance, surpassing Gemini‑2.5‑Pro, while its generated dense descriptions significantly boost downstream capabilities in audio‑visual reasoning (DailyOmni and WorldSense) and temporal grounding (Charades‑STA). All datasets, models, and code will be made publicly available at https://github.com/yaolinli/TimeChat‑Captioner.
Authors:Ying Guo, Qijun Gan, Yifu Zhang, Jinlai Liu, Yifei Hu, Pan Xie, Dongjun Qian, Yu Zhang, Ruiqi Li, Yuqi Zhang, Ruibiao Lu, Xiaofeng Mei, Bo Han, Xiang Yin, Bingyue Peng, Zehuan Yuan
Abstract:
Video generation is rapidly evolving towards unified audio‑video generation. In this paper, we present ALIVE, a generation model that adapts a pretrained Text‑to‑Video (T2V) model to Sora‑style audio‑video generation and animation. In particular, the model unlocks the Text‑to‑Video&Audio (T2VA) and Reference‑to‑Video&Audio (animation) capabilities compared to the T2V foundation models. To support the audio‑visual synchronization and reference animation, we augment the popular MMDiT architecture with a joint audio‑video branch which includes TA‑CrossAttn for temporally‑aligned cross‑modal fusion and UniTemp‑RoPE for precise audio‑visual alignment. Meanwhile, a comprehensive data pipeline consisting of audio‑video captioning, quality control, etc., is carefully designed to collect high‑quality finetuning data. Additionally, we introduce a new benchmark to perform a comprehensive model test and comparison. After continue pretraining and finetuning on million‑level high‑quality data, ALIVE demonstrates outstanding performance, consistently outperforming open‑source models and matching or surpassing state‑of‑the‑art commercial solutions. With detailed recipes and benchmarks, we hope ALIVE helps the community develop audio‑video generation models more efficiently. Official page: https://github.com/FoundationVision/Alive.
Authors:Yi Dao, Lankai Zhang, Hao Liu, Haiwei Zhang, Wenbo Wang
Abstract:
Human pose estimation is fundamental to intelligent perception in the Internet of Things (IoT), enabling applications ranging from smart healthcare to human‑computer interaction. While WiFi‑based methods have gained traction, they often struggle with continuous motion and high computational overhead. This work presents WiFlow, a novel framework for continuous human pose estimation using WiFi signals. Unlike vision‑based approaches such as two‑dimensional deep residual networks that treat Channel State Information (CSI) as images, WiFlow employs an encoder‑decoder architecture. The encoder captures spatio‑temporal features of CSI using temporal and asymmetric convolutions, preserving the original sequential structure of signals. It then refines keypoint features of human bodies to be tracked and capture their structural dependencies via axial attention. The decoder subsequently maps the encoded high‑dimensional features into keypoint coordinates. Trained on a self‑collected dataset of 360,000 synchronized CSI‑pose samples from 5 subjects performing continuous sequences of 8 daily activities, WiFlow achieves a Percentage of Correct Keypoints (PCK) of 97.25% at a threshold of 20% (PCK@20) and 99.48% at PCK@50, with a mean per‑joint position error of 0.007 m. With only 2.23M parameters, WiFlow significantly reduces model complexity and computational cost, establishing a new performance baseline for practical WiFi‑based human pose estimation. Our code and datasets are available at https://github.com/DY2434/WiFlow‑WiFi‑Pose‑Estimation‑with‑Spatio‑Temporal‑Decoupling.git.
Authors:Siyu Liu, Chujie Qin, Hubery Yin, Qixin Yan, Zheng-Peng Duan, Chen Li, Jing Lyu, Chun-Le Guo, Chongyi Li
Abstract:
Recent work leverages Vision Foundation Models as image encoders to boost the generative performance of latent diffusion models (LDMs), as their semantic feature distributions are easy to learn. However, such semantic features often lack low‑level information (\eg, color and texture), leading to degraded reconstruction fidelity, which has emerged as a primary bottleneck in further scaling LDMs. To address this limitation, we propose LV‑RAE, a representation autoencoder that augments semantic features with missing low‑level information, enabling high‑fidelity reconstruction while remaining highly aligned with the semantic distribution. We further observe that the resulting high‑dimensional, information‑rich latent make decoders sensitive to latent perturbations, causing severe artifacts when decoding generated latent and consequently degrading generation quality. Our analysis suggests that this sensitivity primarily stems from excessive decoder responses along directions off the data manifold. Building on these insights, we propose fine‑tuning the decoder to increase its robustness and smoothing the generated latent via controlled noise injection, thereby enhancing generation quality. Experiments demonstrate that LV‑RAE significantly improves reconstruction fidelity while preserving the semantic abstraction and achieving strong generative quality. Our code is available at https://github.com/modyu‑liu/LVRAE.
Authors:Kfir Goldberg, Elad Richardson, Yael Vinker
Abstract:
While generative models have become powerful tools for image synthesis, they are typically optimized for executing carefully crafted textual prompts, offering limited support for the open‑ended visual exploration that often precedes idea formation. In contrast, designers frequently draw inspiration from loosely connected visual references, seeking emergent connections that spark new ideas. We propose Inspiration Seeds, a generative framework that shifts image generation from final execution to exploratory ideation. Given two input images, our model produces diverse, visually coherent compositions that reveal latent relationships between inputs, without relying on user‑specified text prompts. Our approach is feed‑forward, trained on synthetic triplets of decomposed visual aspects derived entirely through visual means: we use CLIP Sparse Autoencoders to extract editing directions in CLIP latent space and isolate concept pairs. By removing the reliance on language and enabling fast, intuitive recombination, our method supports visual ideation at the early and ambiguous stages of creative work.
Authors:Melany Yang, Yuhang Yu, Diwang Weng, Jinwei Chen, Wei Dong
Abstract:
Photorealistic color retouching plays a vital role in visual content creation, yet manual retouching remains inaccessible to non‑experts due to its reliance on specialized expertise. Reference‑based methods offer a promising alternative by transferring the preset color of a reference image to a source image. However, these approaches often operate as novice learners, performing global color mappings derived from pixel‑level statistics, without a true understanding of semantic context or human aesthetics. To address this issue, we propose SemiNFT, a Diffusion Transformer (DiT)‑based retouching framework that mirrors the trajectory of human artistic training: beginning with rigid imitation and evolving into intuitive creation. Specifically, SemiNFT is first taught with paired triplets to acquire basic structural preservation and color mapping skills, and then advanced to reinforcement learning (RL) on unpaired data to cultivate nuanced aesthetic perception. Crucially, during the RL stage, to prevent catastrophic forgetting of old skills, we design a hybrid online‑offline reward mechanism that anchors aesthetic exploration with structural review. % experiments Extensive experiments show that SemiNFT not only outperforms state‑of‑the‑art methods on standard preset transfer benchmarks but also demonstrates remarkable intelligence in zero‑shot tasks, such as black‑and‑white photo colorization and cross‑domain (anime‑to‑photo) preset transfer. These results confirm that SemiNFT transcends simple statistical matching and achieves a sophisticated level of aesthetic comprehension. Our project can be found at https://melanyyang.github.io/SemiNFT/.
Authors:Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin
Abstract:
Human perception for effective object tracking in 2D video streams arises from the implicit use of prior 3D knowledge and semantic reasoning. In contrast, most generic object tracking (GOT) methods primarily rely on 2D features of the target and its surroundings, while neglecting 3D geometric cues, making them susceptible to partial occlusion, distractors, and variations in geometry and appearance. To address this limitation, we introduce GOT‑Edit, an online cross‑modality model editing approach that integrates geometry‑aware cues into a generic object tracker from a 2D video stream. Our approach leverages features from a pre‑trained Visual Geometry Grounded Transformer to infer geometric cues from only a few 2D images. To address the challenge of seamlessly combining geometry and semantics, GOT‑Edit performs online model editing. By leveraging null‑space constraints during model updates, it incorporates geometric information while preserving semantic discrimination, yielding consistently better performance across diverse scenarios. Extensive experiments on multiple GOT benchmarks demonstrate that GOT‑Edit achieves superior robustness and accuracy, particularly under occlusion and clutter, establishing a new paradigm for combining 2D semantics with 3D geometric reasoning for generic object tracking. The project page is available at https://chenshihfang.github.io/GOT‑EDIT.
Authors:Linger Deng, Yuliang Liu, Wenwen Yu, Zujia Zhang, Jianzhong Ju, Zhenbo Luo, Xiang Bai
Abstract:
Geometry problem‑solving remains a significant challenge for Large Multimodal Models (LMMs), requiring not only global shape recognition but also attention to intricate local relationships related to geometric theory. To address this, we propose GeoFocus, a novel framework comprising two core modules. 1) Critical Local Perceptor, which automatically identifies and emphasizes critical local structure (e.g., angles, parallel lines, comparative distances) through thirteen theory‑based perception templates, boosting critical local feature coverage by 61% compared to previous methods. 2) VertexLang, a compact topology formal language, encodes global figures through vertex coordinates and connectivity relations. By replacing bulky code‑based encodings, VertexLang reduces global perception training time by 20% while improving topology recognition accuracy. When evaluated in Geo3K, GeoQA, and FormalGeo7K, GeoFocus achieves a 4.7% accuracy improvement over leading specialized models and demonstrates superior robustness in MATHVERSE under diverse visual conditions. Project Page ‑‑ https://github.com/dle666/GeoFocus
Authors:Yiyang Cao, Yunze Deng, Ziyu Lin, Bin Feng, Xinggang Wang, Wenyu Liu, Dandan Zheng, Jingdong Chen
Abstract:
Text‑to‑motion generation, a rapidly evolving field in computer vision, aims to produce realistic and text‑aligned motion sequences. Current methods primarily focus on spatial‑temporal modeling or independent frequency domain analysis, lacking a unified framework for joint optimization across spatial, temporal, and frequency domains. This limitation hinders the model's ability to leverage information from all domains simultaneously, leading to suboptimal generation quality. Additionally, in motion generation frameworks, motion‑irrelevant cues caused by noise are often entangled with features that contribute positively to generation, thereby leading to motion distortion. To address these issues, we propose Tri‑Domain Causal Text‑to‑Motion Generation (TriC‑Motion), a novel diffusion‑based framework integrating spatial‑temporal‑frequency‑domain modeling with causal intervention. TriC‑Motion includes three core modeling modules for domain‑specific modeling, namely Temporal Motion Encoding, Spatial Topology Modeling, and Hybrid Frequency Analysis. After comprehensive modeling, a Score‑guided Tri‑domain Fusion module integrates valuable information from the triple domains, simultaneously ensuring temporal consistency, spatial topology, motion trends, and dynamics. Moreover, the Causality‑based Counterfactual Motion Disentangler is meticulously designed to expose motion‑irrelevant cues to eliminate noise, disentangling the real modeling contributions of each domain for superior generation. Extensive experimental results validate that TriC‑Motion achieves superior performance compared to state‑of‑the‑art methods, attaining an outstanding R@1 of 0.612 on the HumanML3D dataset. These results demonstrate its capability to generate high‑fidelity, coherent, diverse, and text‑aligned motion sequences. Code is available at: https://caoyiyang1105.github.io/TriC‑Motion/.
Authors:Guoqi Yu, Xiaowei Hu, Angelica I. Aviles-Rivero, Anqi Qiu, Shujun Wang
Abstract:
Functional magnetic resonance imaging (fMRI) enables non‑invasive brain disorder classification by capturing blood‑oxygen‑level‑dependent (BOLD) signals. However, most existing methods rely on functional connectivity (FC) via Pearson correlation, which reduces 4D BOLD signals to static 2D matrices, discarding temporal dynamics and capturing only linear inter‑regional relationships. In this work, we benchmark state‑of‑the‑art temporal models (e.g., time‑series models such as PatchTST, TimesNet, and TimeMixer) on raw BOLD signals across five public datasets. Results show these models consistently outperform traditional FC‑based approaches, highlighting the value of directly modeling temporal information such as cycle‑like oscillatory fluctuations and drift‑like slow baseline trends. Building on this insight, we propose DeCI, a simple yet effective framework that integrates two key principles: (i) Cycle and Drift Decomposition to disentangle cycle and drift within each ROI (Region of Interest); and (ii) Channel‑Independence to model each ROI separately, improving robustness and reducing overfitting. Extensive experiments demonstrate that DeCI achieves superior classification accuracy and generalization compared to both FC‑based and temporal baselines. Our findings advocate for a shift toward end‑to‑end temporal modeling in fMRI analysis to better capture complex brain dynamics. The code is available at https://github.com/Levi‑Ackman/DeCI.
Authors:Jing Zhang, Zhikai Li, Xuewen Liu, Qingyi Gu
Abstract:
Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real‑time video processing. Although there have been efforts to improve the efficiency of SAM2, most of them focus on retraining a lightweight backbone, with little exploration into post‑training acceleration. In this paper, we observe that SAM2 exhibits sparse perception pattern as biological vision, which provides opportunities for eliminating redundant computation and acceleration: i) In mask decoder, the attention primarily focuses on the foreground objects, whereas the image encoder in the earlier stage exhibits a broad attention span, which results in unnecessary computation to background regions. ii) In memory bank, only a small subset of tokens in each frame contribute significantly to memory attention, and the salient regions exhibit temporal consistency, making full‑token computation redundant. With these insights, we propose Efficient‑SAM2, which promotes SAM2 to adaptively focus on object regions while eliminating task‑irrelevant computations, thereby significantly improving inference efficiency. Specifically, for image encoder, we propose object‑aware Sparse Window Routing (SWR), a window‑level computation allocation mechanism that leverages the consistency and saliency cues from the previous‑frame decoder to route background regions into a lightweight shortcut branch. Moreover, for memory attention, we propose object‑aware Sparse Memory Retrieval (SMR), which allows only the salient memory tokens in each frame to participate in computation, with the saliency pattern reused from their first recollection. With negligible additional parameters and minimal training overhead, Efficient‑SAM2 delivers 1.68x speedup on SAM2.1‑L model with only 1.0% accuracy drop on SA‑V test set.
Authors:Seoyeon Jang, Alex Junho Lee, I Made Aswin Nahrendra, Hyun Myung
Abstract:
Online change detection is crucial for mobile robots to efficiently navigate through dynamic environments. Detecting changes in transient settings, such as active construction sites or frequently reconfigured indoor spaces, is particularly challenging due to frequent occlusions and spatiotemporal variations. Existing approaches often struggle to detect changes and fail to update the map across different observations. To address these limitations, we propose a dual‑head network designed for online change detection and long‑term map maintenance. A key difficulty in this task is the collection and alignment of real‑world data, as manually registering structural differences over time is both labor‑intensive and often impractical. To overcome this, we develop a data augmentation strategy that synthesizes structural changes by importing elements from different scenes, enabling effective model training without the need for extensive ground‑truth annotations. Experiments conducted at real‑world construction sites and in indoor office environments demonstrate that our approach generalizes well across diverse scenarios, achieving efficient and accurate map updates.\resubmitOur source code and additional material are available at: https://chamelion‑pages.github.io/.
Authors:Sidike Paheding, Abel Reyes-Angulo, Leo Thomas Ramos, Angel D. Sappa, Rajaneesh A., Hiral P. B., Sajin Kumar K. S., Thomas Oommen
Abstract:
We present MMLSv2, a dataset for landslide segmentation on Martian surfaces. MMLSv2 consists of multimodal imagery with seven bands: RGB, digital elevation model, slope, thermal inertia, and grayscale channels. MMLSv2 comprises 664 images distributed across training, validation, and test splits. In addition, an isolated test set of 276 images from a geographically disjoint region from the base dataset is released to evaluate spatial generalization. Experiments conducted with multiple segmentation models show that the dataset supports stable training and achieves competitive performance, while still posing challenges in fragmented, elongated, and small‑scale landslide regions. Evaluation on the isolated test set leads to a noticeable performance drop, indicating increased difficulty and highlighting its value for assessing model robustness and generalization beyond standard in‑distribution settings. Dataset will be available at: https://github.com/MAIN‑Lab/MMLS_v2
Authors:Issar Tzachor, Dvir Samuel, Rami Ben-Ari
Abstract:
Recent studies have adapted generative Multimodal Large Language Models (MLLMs) into embedding extractors for vision tasks, typically through fine‑tuning to produce universal representations. However, their performance on video remains inferior to Video Foundation Models (VFMs). In this paper, we focus on leveraging MLLMs for video‑text embedding and retrieval. We first conduct a systematic layer‑wise analysis, showing that intermediate (pre‑trained) MLLM layers already encode substantial task‑relevant information. Leveraging this insight, we demonstrate that combining intermediate‑layer embeddings with a calibrated MLLM head yields strong zero‑shot retrieval performance without any training. Building on these findings, we introduce a lightweight text‑based alignment strategy which maps dense video captions to short summaries and enables task‑related video‑text embedding learning without visual supervision. Remarkably, without any fine‑tuning beyond text, our method outperforms current methods, often by a substantial margin, achieving state‑of‑the‑art results across common video retrieval benchmarks.
Authors:Feng Wang, Sucheng Ren, Tiezheng Zhang, Predrag Neskovic, Anand Bhattad, Cihang Xie, Alan Yuille
Abstract:
This work presents a systematic investigation into modernizing Vision Transformer backbones by leveraging architectural advancements from the past five years. While preserving the canonical Attention‑FFN structure, we conduct a component‑wise refinement involving normalization, activation functions, positional encoding, gating mechanisms, and learnable tokens. These updates form a new generation of Vision Transformers, which we call ViT‑5. Extensive experiments demonstrate that ViT‑5 consistently outperforms state‑of‑the‑art plain Vision Transformers across both understanding and generation benchmarks. On ImageNet‑1k classification, ViT‑5‑Base reaches 84.2% top‑1 accuracy under comparable compute, exceeding DeiT‑III‑Base at 83.8%. ViT‑5 also serves as a stronger backbone for generative modeling: when plugged into an SiT diffusion framework, it achieves 1.84 FID versus 2.06 with a vanilla ViT backbone. Beyond headline metrics, ViT‑5 exhibits improved representation learning and favorable spatial reasoning behavior, and transfers reliably across tasks. With a design aligned with contemporary foundation‑model practices, ViT‑5 offers a simple drop‑in upgrade over vanilla ViT for mid‑2020s vision backbones.
Authors:Chunyang Li, Yuanbo Yang, Jiahao Shao, Hongyu Zhou, Katja Schwarz, Yiyi Liao
Abstract:
Video generation with controllable camera viewpoints is essential for applications such as interactive content creation, gaming, and simulation. Existing methods typically adapt pre‑trained video models using camera poses relative to a fixed reference, e.g., the first frame. However, these encodings lack shift‑invariance, often leading to poor generalization and accumulated drift. While relative camera pose embeddings defined between arbitrary view pairs offer a more robust alternative, integrating them into pre‑trained video diffusion models without prohibitive training costs or architectural changes remains challenging. We introduce ReRoPE, a plug‑and‑play framework that incorporates relative camera information into pre‑trained video diffusion models without compromising their generation capability. Our approach is based on the insight that Rotary Positional Embeddings (RoPE) in existing models underutilize their full spectral bandwidth, particularly in the low‑frequency components. By seamlessly injecting relative camera pose information into these underutilized bands, ReRoPE achieves precise control while preserving strong pre‑trained generative priors. We evaluate our method on both image‑to‑video (I2V) and video‑to‑video (V2V) tasks in terms of camera control accuracy and visual fidelity. Our results demonstrate that ReRoPE offers a training‑efficient path toward controllable, high‑fidelity video generation. See project page for more results: https://sisyphe‑lee.github.io/ReRoPE/
Authors:Yixuan Ye, Xuanyu Lu, Yuxin Jiang, Yuchao Gu, Rui Zhao, Qiwei Liang, Jiachun Pan, Fengda Zhang, Weijia Wu, Alex Jinpeng Wang
Abstract:
World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open‑domain closed‑loop revisited benchmark for evaluating Memory consIstency and action coNtrol in worlD models. MIND contains 250 high‑quality videos at 1080p and 24 FPS, including 100 (first‑person) + 100 (third‑person) video clips under a shared action space and 25 + 25 clips across varied action spaces covering eight diverse scenes. We design an efficient evaluation framework to measure two core abilities: memory consistency and action control, capturing temporal stability and contextual coherence across viewpoints. Furthermore, we design various action spaces, including different character movement speeds and camera rotation angles, to evaluate the action generalization capability across different action spaces under shared scenes. To facilitate future performance benchmarking on MIND, we introduce MIND‑World, a novel interactive Video‑to‑World baseline. Extensive experiments demonstrate the completeness of MIND and reveal key challenges in current world models, including the difficulty of maintaining long‑term memory consistency and generalizing across action spaces. Code: https://github.com/CSU‑JPG/MIND.
Authors:Ziyang Fan, Keyu Chen, Ruilong Xing, Yulin Li, Li Jiang, Zhuotao Tian
Abstract:
Although Video Large Language Models (VLLMs) have shown remarkable capabilities in video understanding, they are required to process high volumes of visual tokens, causing significant computational inefficiency. Existing VLLMs acceleration frameworks usually compress spatial and temporal redundancy independently, which overlooks the spatiotemporal relationships, thereby leading to suboptimal spatiotemporal compression. The highly correlated visual features are likely to change in spatial position, scale, orientation, and other attributes over time due to the dynamic nature of video. Building on this insight, we introduce FlashVID, a training‑free inference acceleration framework for VLLMs. Specifically, FlashVID utilizes Attention and Diversity‑based Token Selection (ADTS) to select the most representative tokens for basic video representation, then applies Tree‑based Spatiotemporal Token Merging (TSTM) for fine‑grained spatiotemporal redundancy elimination. Extensive experiments conducted on three representative VLLMs across five video understanding benchmarks demonstrate the effectiveness and generalization of our method. Notably, by retaining only 10% of visual tokens, FlashVID preserves 99.1% of the performance of LLaVA‑OneVision. Consequently, FlashVID can serve as a training‑free and plug‑and‑play module for extending long video frames, which enables a 10x increase in video frame input to Qwen2.5‑VL, resulting in a relative improvement of 8.6% within the same computational budget. Code is available at https://github.com/Fanziyang‑v/FlashVID.
Authors:Xiaofeng Tan, Wanjiang Weng, Haodong Lei, Hongsong Wang
Abstract:
In recent years, motion generative models have undergone significant advancement, yet pose challenges in aligning with downstream objectives. Recent studies have shown that using differentiable rewards to directly align the preference of diffusion models yields promising results. However, these methods suffer from (1) inefficient and coarse‑grained optimization with (2) high memory consumption. In this work, we first theoretically and empirically identify the key reason of these limitations: the recursive dependence between different steps in the denoising trajectory. Inspired by this insight, we propose EasyTune, which fine‑tunes diffusion at each denoising step rather than over the entire trajectory. This decouples the recursive dependence, allowing us to perform (1) a dense and fine‑grained, and (2) memory‑efficient optimization. Furthermore, the scarcity of preference motion pairs restricts the availability of motion reward model training. To this end, we further introduce a Self‑refinement Preference Learning (SPL) mechanism that dynamically identifies preference pairs and conducts preference learning. Extensive experiments demonstrate that EasyTune outperforms DRaFT‑50 by 8.2% in alignment (MM‑Dist) improvement while requiring only 31.16% of its additional memory overhead and achieving a 7.3x training speedup. The project page is available at this link https://xiaofeng‑tan.github.io/projects/EasyTune/index.html.
Authors:Changli Tang, Tianyi Wang, Fengyun Rao, Jing Lyu, Chao Zhang
Abstract:
Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D‑ORCA, a dialogue‑centric omni‑modal large language model optimized for robust audio‑visual captioning. We further curate DVD, a large‑scale, high‑quality bilingual dataset comprising nearly 40,000 multi‑party dialogue videos for training and 2000 videos for evaluation in English and Mandarin, addressing a critical gap in the open‑source ecosystem. To ensure fine‑grained captioning accuracy, we adopt group relative policy optimization with three novel reward functions that assess speaker attribution accuracy, global speech content accuracy, and sentence‑level temporal boundary alignment. These rewards are derived from evaluation metrics widely used in speech processing and, to our knowledge, are applied for the first time as reinforcement learning objectives for audio‑visual captioning. Extensive experiments demonstrate that D‑ORCA substantially outperforms existing open‑source models in speaker identification, speech recognition, and temporal grounding. Notably, despite having only 8 billion parameters, D‑ORCA achieves performance competitive with Qwen3‑Omni across several general‑purpose audio‑visual understanding benchmarks. Demos are available at \hrefhttps://d‑orca‑llm.github.io/https://d‑orca‑llm.github.io/. Our code, data, and checkpoints will be available at \hrefhttps://github.com/WeChatCV/D‑ORCA/https://github.com/WeChatCV/D‑ORCA/.
Authors:Mert Sonmezer, Serge Vasylechko, Duygu Atasoy, Seyda Ertekin, Sila Kurugol
Abstract:
Retrieving wrist radiographs with analogous fracture patterns is challenging because clinically important cues are subtle, highly localized and often obscured by overlapping anatomy or variable imaging views. Progress is further limited by the scarcity of large, well‑annotated datasets for case‑based medical image retrieval. We introduce WristMIR, a region‑aware pediatric wrist radiograph retrieval framework that leverages dense radiology reports and bone‑specific localization to learn fine‑grained, clinically meaningful image representations without any manual image‑level annotations. Using MedGemma‑based structured report mining to generate both global and region‑level captions, together with pre‑processed wrist images and bone‑specific crops of the distal radius, distal ulna, and ulnar styloid, WristMIR jointly trains global and local contrastive encoders and performs a two‑stage retrieval process: (1) coarse global matching to identify candidate exams, followed by (2) region‑conditioned reranking aligned to a predefined anatomical bone region. WristMIR improves retrieval performance over strong vision‑language baselines, raising image‑to‑text Recall@5 from 0.82% to 9.35%. Its embeddings also yield stronger fracture classification (AUROC 0.949, AUPRC 0.953). In region‑aware evaluation, the two‑stage design markedly improves retrieval‑based fracture diagnosis, increasing mean F_1 from 0.568 to 0.753, and radiologists rate its retrieved cases as more clinically relevant, with mean scores rising from 3.36 to 4.35. These findings highlight the potential of anatomically guided retrieval to enhance diagnostic reasoning and support clinical decision‑making in pediatric musculoskeletal imaging. The source code is publicly available at https://github.com/quin‑med‑harvard‑edu/WristMIR.
Authors:Fei Yu, Shudan Guo, Shiqing Xin, Beibei Wang, Haisen Zhao, Wenzheng Chen
Abstract:
We consider the problem of 3D shape recovery from ultra‑fast motion‑blurred images. While 3D reconstruction from static images has been extensively studied, recovering geometry from extreme motion‑blurred images remains challenging. Such scenarios frequently occur in both natural and industrial settings, such as fast‑moving objects in sports (e.g., balls) or rotating machinery, where rapid motion distorts object appearance and makes traditional 3D reconstruction techniques like Multi‑View Stereo (MVS) ineffective.
In this paper, we propose a novel inverse rendering approach for shape recovery from ultra‑fast motion‑blurred images. While conventional rendering techniques typically synthesize blur by averaging across multiple frames, we identify a major computational bottleneck in the repeated computation of barycentric weights. To address this, we propose a fast barycentric coordinate solver, which significantly reduces computational overhead and achieves a speedup of up to 4.57x, enabling efficient and photorealistic simulation of high‑speed motion. Crucially, our method is fully differentiable, allowing gradients to propagate from rendered images to the underlying 3D shape, thereby facilitating shape recovery through inverse rendering.
We validate our approach on two representative motion types: rapid translation and rotation. Experimental results demonstrate that our method enables efficient and realistic modeling of ultra‑fast moving objects in the forward simulation. Moreover, it successfully recovers 3D shapes from 2D imagery of objects undergoing extreme translational and rotational motion, advancing the boundaries of vision‑based 3D reconstruction. Project page: https://maxmilite.github.io/rec‑from‑ultrafast‑blur/
Authors:Sanoojan Baliah, Yohan Abeysinghe, Rusiru Thushara, Khan Muhammad, Abhinav Dhall, Karthik Nandakumar, Muhammad Haris Khan
Abstract:
We present a training‑free, plug‑and‑play method, namely VFace, for high‑quality face swapping in videos. It can be seamlessly integrated with image‑based face swapping approaches built on diffusion models. First, we introduce a Frequency Spectrum Attention Interpolation technique to facilitate generation and intact key identity characteristics. Second, we achieve Target Structure Guidance via plug‑and‑play attention injection to better align the structural features from the target frame to the generation. Third, we present a Flow‑Guided Attention Temporal Smoothening mechanism that enforces spatiotemporal coherence without modifying the underlying diffusion model to reduce temporal inconsistencies typically encountered in frame‑wise generation. Our method requires no additional training or video‑specific fine‑tuning. Extensive experiments show that our method significantly enhances temporal consistency and visual fidelity, offering a practical and modular solution for video‑based face swapping. Our code is available at https://github.com/Sanoojan/VFace.
Authors:Weijiang Lv, Yaoxuan Feng, Xiaobo Xia, Jiayu Wang, Yan Jing, Wenchao Chen, Bo Chen
Abstract:
Chain‑of‑Thought reasoning is widely used to improve the interpretability of multimodal large language models (MLLMs), yet the faithfulness of the generated reasoning traces remains unclear. Prior work has mainly focused on perceptual hallucinations, leaving reasoning level unfaithfulness underexplored. To isolate faithfulness from linguistic priors, we introduce SPD‑Faith Bench, a diagnostic benchmark based on fine‑grained image difference reasoning that enforces explicit visual comparison. Evaluations on state‑of‑the‑art MLLMs reveal two systematic failure modes, perceptual blindness and perception‑reasoning dissociation. We trace these failures to decaying visual attention and representation shifts in the residual stream. Guided by this analysis, we propose SAGE, a train‑free visual evidence‑calibrated framework that improves visual routing and aligns reasoning with perception. Our results highlight the importance of explicitly evaluating faithfulness beyond response correctness. Our benchmark and codes are available at https://github.com/Johanson‑colab/SPD‑Faith‑Bench.
Authors:Wenping Jin, Yuyang Tang, Li Zhu
Abstract:
Reliable foreign‑object anomaly detection and pixel‑level localization in conveyor‑belt coal scenes are essential for safe and intelligent mining operations. This task is particularly challenging due to the highly unstructured environment: coal and gangue are randomly piled, backgrounds are complex and variable, and foreign objects often exhibit low contrast, deformation, occlusion, resulting in coupling with their surroundings. These characteristics weaken the stability and regularity assumptions that many anomaly detection methods rely on in structured industrial settings, leading to notable performance degradation. To support evaluation and comparison in this setting, we construct CoalAD, a benchmark for unsupervised foreign‑object anomaly detection with pixel‑level localization in coal‑stream scenes. We further propose a complementary‑cue collaborative perception framework that extracts and fuses complementary anomaly evidence from three perspectives: object‑level semantic composition modeling, semantic‑attribution‑based global deviation analysis, and fine‑grained texture matching. The fused outputs provide robust image‑level anomaly scoring and accurate pixel‑level localization. Experiments on CoalAD demonstrate that our method outperforms widely used baselines across the evaluated image‑level and pixel‑level metrics, and ablation studies validate the contribution of each component. The code is available at https://github.com/xjpp2016/USAD.
Authors:Binxiao Xu, Junyu Feng, Xiaopeng Lin, Haodong Li, Zhiyuan Feng, Bohan Zeng, Shaolin Lu, Ming Lu, Qi She, Wentao Zhang
Abstract:
Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, despite excelling at general search, existing agents often struggle to bridge the cognitive gap between pixel‑level perception and high‑level marketing logic. To address this challenge, we introduce AD‑MIR, a framework designed to decode advertising intent via a two‑stage architecture. First, in the Structure‑Aware Memory Construction phase, the system converts raw video into a structured database by integrating semantic retrieval with exact keyword matching. This approach prioritizes fine‑grained brand details (e.g., logos, on‑screen text) while dynamically filtering out irrelevant background noise to isolate key protagonists. Second, the Structured Reasoning Agent mimics a marketing expert through an iterative inquiry loop, decomposing the narrative to deduce implicit persuasion tactics. Crucially, it employs an evidence‑based self‑correction mechanism that rigorously validates these insights against specific video frames, automatically backtracking when visual support is lacking. Evaluation on the AdsQA benchmark demonstrates that AD‑MIR achieves state‑of‑the‑art performance, surpassing the strongest general‑purpose agent, DVD, by 1.8% in strict and 9.5% in relaxed accuracy. These results underscore that effective advertising understanding demands explicitly grounding abstract marketing strategies in pixel‑level evidence. The code is available at https://github.com/Little‑Fridge/AD‑MIR.
Authors:Hulingxiao He, Zijun Geng, Yuxin Peng
Abstract:
Any entity in the visual world can be hierarchically grouped based on shared characteristics and mapped to fine‑grained sub‑categories. While Multi‑modal Large Language Models (MLLMs) achieve strong performance on coarse‑grained visual tasks, they often struggle with Fine‑Grained Visual Recognition (FGVR). Adapting general‑purpose MLLMs to FGVR typically requires large amounts of annotated data, which is costly to obtain, leaving a substantial performance gap compared to contrastive CLIP models dedicated for discriminative tasks. Moreover, MLLMs tend to overfit to seen sub‑categories and generalize poorly to unseen ones. To address these challenges, we propose Fine‑R1, an MLLM tailored for FGVR through an R1‑style training framework: (1) Chain‑of‑Thought Supervised Fine‑tuning, where we construct a high‑quality FGVR CoT dataset with rationales of "visual analysis, candidate sub‑categories, comparison, and prediction", transition the model into a strong open‑world classifier; and (2) Triplet Augmented Policy Optimization, where Intra‑class Augmentation mixes trajectories from anchor and positive images within the same category to improve robustness to intra‑class variance, while Inter‑class Augmentation maximizes the response distinction conditioned on images across sub‑categories to enhance discriminative ability. With only 4‑shot training, Fine‑R1 outperforms existing general MLLMs, reasoning MLLMs, and even contrastive CLIP models in identifying both seen and unseen sub‑categories, showing promise in working in knowledge‑intensive domains where gathering expert annotations for all sub‑categories is arduous. Code is available at https://github.com/PKU‑ICST‑MIPL/FineR1_ICLR2026.
Authors:Wenjie Liu, Hao Wu, Xin Qiu, Xudong Wang, Yingqi Fan, Yihan Zhang, Anhao Zhao, Yunpu Ma, Xiaoyu Shen
Abstract:
Modern multimodal large language models (MLLMs) adopt a unified self‑attention design that processes visual and textual tokens at every Transformer layer, incurring substantial computational overhead. In this work, we revisit the necessity of such dense visual processing and show that projected visual embeddings are already well‑aligned with the language space, while effective vision‑language interaction occurs in only a small subset of layers. Based on these insights, we propose ViCA (Vision‑only Cross‑Attention), a minimal MLLM architecture in which visual tokens bypass all self‑attention and feed‑forward layers, interacting with text solely through sparse cross‑attention at selected layers. Extensive evaluations across three MLLM backbones, nine multimodal benchmarks, and 26 pruning‑based baselines show that ViCA preserves 98% of baseline accuracy while reducing visual‑side computation to 4%, consistently achieving superior performance‑efficiency trade‑offs. Moreover, ViCA provides a regular, hardware‑friendly inference pipeline that yields over 3.5x speedup in single‑batch inference and over 10x speedup in multi‑batch inference, reducing visual grounding to near‑zero overhead compared with text‑only LLMs. It is also orthogonal to token pruning methods and can be seamlessly combined for further efficiency gains. Our code is available at https://github.com/EIT‑NLP/ViCA.
Authors:Hussni Mohd Zakir, Eric Tatt Wei Ho
Abstract:
Recent self‑supervised Vision Transformers (ViTs), such as DINOv3, provide rich feature representations for dense vision tasks. This study investigates the intrinsic few‑shot semantic segmentation (FSS) capabilities of frozen DINOv3 features through a training‑free baseline, FSSDINO, utilizing class‑specific prototypes and Gram‑matrix refinement. Our results across binary, multi‑class, and cross‑domain (CDFSS) benchmarks demonstrate that this minimal approach, applied to the final backbone layer, is highly competitive with specialized methods involving complex decoders or test‑time adaptation. Crucially, we conduct an Oracle‑guided layer analysis, identifying a significant performance gap between the standard last‑layer features and globally optimal intermediate representations. We reveal a "Safest vs. Optimal" dilemma: while the Oracle proves higher performance is attainable, matching the results of compute‑intensive adaptation methods, current unsupervised and support‑guided selection metrics consistently yield lower performance than the last‑layer baseline. This characterizes a "Semantic Selection Gap" in Foundation Models, a disconnect where traditional heuristics fail to reliably identify high‑fidelity features. Our work establishes the "Last‑Layer" as a deceptively strong baseline and provides a rigorous diagnostic of the latent semantic potentials in DINOv3.The code is publicly available at https://github.com/hussni0997/fssdino.
Authors:Sebastian Bock, Leonie Schüßler, Krishnakant Singh, Simone Schaub-Meyer, Stefan Roth
Abstract:
Unsupervised object‑centric learning (OCL) decomposes visual scenes into distinct entities. Slot attention is a popular approach that represents individual objects as latent vectors, called slots. Current methods obtain these slot representations solely from the last layer of a pre‑trained vision transformer (ViT), ignoring valuable, semantically rich information encoded across the other layers. To better utilize this latent semantic information, we introduce MUFASA, a lightweight plug‑and‑play framework for slot attention‑based approaches to unsupervised object segmentation. Our model computes slot attention across multiple feature layers of the ViT encoder, fully leveraging their semantic richness. We propose a fusion strategy to aggregate slots obtained on multiple layers into a unified object‑centric representation. Integrating MUFASA into existing OCL methods improves their segmentation results across multiple datasets, setting a new state of the art while simultaneously improving training convergence with only minor inference overhead.
Authors:Krishnakant Singh, Simone Schaub-Meyer, Stefan Roth
Abstract:
Object‑centric learning (OCL) aims to learn structured scene representations that support compositional generalization and robustness to out‑of‑distribution (OOD) data. However, OCL models are often not evaluated regarding these goals. Instead, most prior work focuses on evaluating OCL models solely through object discovery and simple reasoning tasks, such as probing the representation via image classification. We identify two limitations in existing benchmarks: (1) They provide limited insights on the representation usefulness of OCL models, and (2) localization and representation usefulness are assessed using disjoint metrics. To address (1), we use instruction‑tuned VLMs as evaluators, enabling scalable benchmarking across diverse VQA datasets to measure how well VLMs leverage OCL representations for complex reasoning tasks. To address (2), we introduce a unified evaluation task and metric that jointly assess localization (where) and representation usefulness (what), thereby eliminating inconsistencies introduced by disjoint evaluation. Finally, we include a simple multi‑feature reconstruction baseline as a reference point.
Authors:Tao Wang, Chenyu Lin, Chenwei Tang, Jizhe Zhou, Deng Xiong, Jianan Li, Jian Zhao, Jiancheng Lv
Abstract:
Detecting objects from UAV‑captured images is challenging due to the small object size. In this work, a simple and efficient adaptive zoom‑in framework is explored for object detection on UAV images. The main motivation is that the foreground objects are generally smaller and sparser than those in common scene images, which hinders the optimization of effective object detectors. We thus aim to zoom in adaptively on the objects to better capture object features for the detection task. To achieve the goal, two core designs are required: \textcolorblacki) How to conduct non‑uniform zooming on each image efficiently? ii) How to enable object detection training and inference with the zoomed image space? Correspondingly, a lightweight offset prediction scheme coupled with a novel box‑based zooming objective is introduced to learn non‑uniform zooming on the input image. Based on the learned zooming transformation, a corner‑aligned bounding box transformation method is proposed. The method warps the ground‑truth bounding boxes to the zoomed space to learn object detection, and warps the predicted bounding boxes back to the original space during inference. We conduct extensive experiments on three representative UAV object detection datasets, including VisDrone, UAVDT, and SeaDronesSee. The proposed ZoomDet is architecture‑independent and can be applied to an arbitrary object detection architecture. Remarkably, on the SeaDronesSee dataset, ZoomDet offers more than 8.4 absolute gain of mAP with a Faster R‑CNN model, with only about 3 ms additional latency. The code is available at https://github.com/twangnh/zoomdet_code.
Authors:Naqcho Ali Mehdi, Aamir Ali Drigh
Abstract:
Electrocardiogram (ECG) digitization‑converting paper‑based or scanned ECG images back into time‑series signals‑is critical for leveraging decades of legacy clinical data in modern deep learning applications. However, progress has been hindered by the lack of large‑scale datasets providing both ECG images and their corresponding ground truth signals with comprehensive annotations. We introduce PTB‑XL‑Image‑17K, a complete synthetic ECG image dataset comprising 17,271 high‑quality 12‑lead ECG images generated from the PTB‑XL signal database. Our dataset uniquely provides five complementary data types per sample: (1) realistic ECG images with authentic grid patterns and annotations (50% with visible grid, 50% without), (2) pixel‑level segmentation masks, (3) ground truth time‑series signals, (4) bounding box annotations in YOLO format for both lead regions and lead name labels, and (5) comprehensive metadata including visual parameters and patient information. We present an open‑source Python framework enabling customizable dataset generation with controllable parameters including paper speed (25/50 mm/s), voltage scale (5/10 mm/mV), sampling rate (500 Hz), grid appearance (4 colors), and waveform characteristics. The dataset achieves 100% generation success rate with an average processing time of 1.35 seconds per sample. PTB‑XL‑Image‑17K addresses critical gaps in ECG digitization research by providing the first large‑scale resource supporting the complete pipeline: lead detection, waveform segmentation, and signal extraction with full ground truth for rigorous evaluation. The dataset, generation framework, and documentation are publicly available at https://github.com/naqchoalimehdi/PTB‑XL‑Image‑17K and https://doi.org/10.5281/zenodo.18197519.
Authors:Changhua Xu, En Yu, Junyu Xuan, Jie Lu
Abstract:
Vision‑‑Language‑‑Action (VLA) models bridge multimodal reasoning with physical control, but adapting them to new tasks with scarce demonstrations remains unreliable. While fine‑tuned VLA policies often produce semantically plausible trajectories, failures often arise from unresolved geometric ambiguities, where near‑miss actions lead to divergent execution outcomes under limited supervision. We study few‑shot VLA adaptation from a \emphgeneration‑‑selection perspective and propose a novel framework VGAS (Value‑Guided Action‑chunk Selection). It performs inference‑time best‑of‑N selection to identify action chunks that are both semantically faithful and geometrically precise. Specifically, VGAS employs a finetuned VLA as a high‑recall proposal generator and introduces the \textrmQ‑Chunk‑Former, a geometrically grounded Transformer critic to resolve fine‑grained geometric ambiguities. In addition, we propose Explicit Geometric Regularization (\textttEGR), which shapes a discriminative value landscape to preserve action ranking resolution among near‑miss candidates while mitigating value instability under scarce supervision. Experiments and theoretical analysis demonstrate that VGAS consistently improves success rates and robustness under limited demonstrations and distribution shifts. Our code is available at https://github.com/Jyugo‑15/VGAS.
Authors:Yang Zhang, Zhangkai Ni, Wenhan Yang, Hanli Wang
Abstract:
High Dynamic Range (HDR) video reconstruction aims to recover fine brightness, color, and details from Low Dynamic Range (LDR) videos. However, existing methods often suffer from color inaccuracies and temporal inconsistencies. To address these challenges, we propose WMNet, a novel HDR video reconstruction network that leverages Wavelet domain Masked Image Modeling (W‑MIM). WMNet adopts a two‑phase training strategy: In Phase I, W‑MIM performs self‑reconstruction pre‑training by selectively masking color and detail information in the wavelet domain, enabling the network to develop robust color restoration capabilities. A curriculum learning scheme further refines the reconstruction process. Phase II fine‑tunes the model using the pre‑trained weights to improve the final reconstruction quality. To improve temporal consistency, we introduce the Temporal Mixture of Experts (T‑MoE) module and the Dynamic Memory Module (DMM). T‑MoE adaptively fuses adjacent frames to reduce flickering artifacts, while DMM captures long‑range dependencies, ensuring smooth motion and preservation of fine details. Additionally, since existing HDR video datasets lack scene‑based segmentation, we reorganize HDRTV4K into HDRTV4K‑Scene, establishing a new benchmark for HDR video reconstruction. Extensive experiments demonstrate that WMNet achieves state‑of‑the‑art performance across multiple evaluation metrics, significantly improving color fidelity, temporal coherence, and perceptual quality. The code is available at: https://github.com/eezkni/WMNet
Authors:Junbo Jacob Lian, Feng Xiong, Yujun Sun, Kaichen Ouyang, Zong Ke, Mingyang Yu, Shengwei Fu, Zhong Rui, Zhang Yujun, Huiling Chen
Abstract:
Second‑order feature statistics are central to texture recognition, yet existing mechanisms exhibit a structural tension: bilinear pooling and Gram matrices capture global channel correlations but discard spatial structure, whereas self‑attention models capture cross‑position relations through weighted sums rather than explicit pairwise products. We propose TwistNet‑2D, a lightweight module that computes local pairwise channel products under directional spatial displacement, jointly encoding where features co‑occur and how they interact. The core component, Spiral‑Twisted Channel Interaction (STCI), shifts one feature map along a prescribed direction before L2‑normalized channel multiplication, capturing cross‑position co‑occurrence patterns that characterize structured and periodic textures. Four directional heads are aggregated through content‑adaptive channel reweighting, and the result is injected via a sigmoid‑gated residual path with near‑zero initialization. TwistNet‑2D adds only approximately 3.5% parameters and approximately 2% FLOPs over ResNet‑18. To isolate the contribution of architectural inductive bias from that of transfer learning, all models in this study are trained from scratch without ImageNet pretraining. Under this protocol, TwistNet‑2D consistently surpasses parameter‑matched baselines and substantially larger ConvNeXt and Swin Transformer backbones across four texture and fine‑grained recognition benchmarks, while the multi‑head structure produces interpretable, orientation‑selective representations that align with classical texture analysis.
Authors:Heyuan Li, Huimin Zhang, Yuda Qiu, Zhengwentai Sun, Keru Zheng, Lingteng Qiu, Peihao Li, Qi Zuo, Ce Chen, Yujian Zheng, Yuming Gu, Zilong Dong, Xiaoguang Han
Abstract:
Conditioning is crucial for stable training of full‑head 3D GANs. Without any conditioning signal, the model suffers from severe mode collapse, making it impractical to training. However, a series of previous full‑head 3D GANs conventionally choose the view angle as the conditioning input, which leads to a bias in the learned 3D full‑head space along the conditional view direction. This is evident in the significant differences in generation quality and diversity between the conditional view and non‑conditional views of the generated 3D heads, resulting in global incoherence across different head regions. In this work, we propose to use view‑invariant semantic feature as the conditioning input, thereby decoupling the generative capability of 3D heads from the viewing direction. To construct a view‑invariant semantic condition for each training image, we create a novel synthesized head image dataset. We leverage FLUX.1 Kontext to extend existing high‑quality frontal face datasets to a wide range of view angles. The image clip feature extracted from the frontal view is then used as a shared semantic condition across all views in the extended images, ensuring semantic alignment while eliminating directional bias. This also allows supervision from different views of the same subject to be consolidated under a shared semantic condition, which accelerates training and enhances the global coherence of the generated 3D heads. Moreover, as GANs often experience slower improvements in diversity once the generator learns a few modes that successfully fool the discriminator, our semantic conditioning encourages the generator to follow the true semantic distribution, thereby promoting continuous learning and diverse generation. Extensive experiments on full‑head synthesis and single‑view GAN inversion demonstrate that our method achieves significantly higher fidelity, diversity, and generalizability.
Authors:Yongheng Sun, Jun Shu, Jianhua Ma, Fan Wang
Abstract:
Accurate segmentation of brain tissues from MRI scans is critical for neuroscience and clinical applications, but achieving consistent performance across the human lifespan remains challenging due to dynamic, age‑related changes in brain appearance and morphology. While prior work has sought to mitigate these shifts by using self‑supervised regularization with paired longitudinal data, such data are often unavailable in practice. To address this, we propose \emphDuMeta++, a dual meta‑learning framework that operates without paired longitudinal data. Our approach integrates: (1) meta‑feature learning to extract age‑agnostic semantic representations of spatiotemporally evolving brain structures, and (2) meta‑initialization learning to enable data‑efficient adaptation of the segmentation model. Furthermore, we propose a memory‑bank‑based class‑aware regularization strategy to enforce longitudinal consistency without explicit longitudinal supervision. We theoretically prove the convergence of our DuMeta++, ensuring stability. Experiments on diverse datasets (iSeg‑2019, IBIS, OASIS, ADNI) under few‑shot settings demonstrate that DuMeta++ outperforms existing methods in cross‑age generalization. Code will be available at https://github.com/ladderlab‑xjtu/DuMeta++.
Authors:Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han, Soochahn Lee, Sukanta Ganguly, Yong Jae Lee
Abstract:
Universal Multimodal Retrieval (UMR) seeks any‑to‑any search across text and vision, yet modern embedding models remain brittle when queries require latent reasoning (e.g., resolving underspecified references or matching compositional constraints). We argue this brittleness is often data‑induced: when images carry "silent" evidence and queries leave key semantics implicit, a single embedding pass must both reason and compress, encouraging spurious feature matching. We propose a data‑centric framework that decouples these roles by externalizing reasoning before retrieval. Using a strong Vision‑‑Language Model, we make implicit semantics explicit by densely captioning visual evidence in corpus entries, resolving ambiguous multimodal references in queries, and rewriting verbose instructions into concise retrieval constraints. Inference‑time enhancement alone is insufficient; the retriever must be trained on these semantically dense representations to avoid distribution shift and fully exploit the added signal. Across M‑BEIR, our reasoning‑augmented training method yields consistent gains over strong baselines, with ablations showing that corpus enhancement chiefly benefits knowledge‑intensive queries while query enhancement is critical for compositional modification requests. We publicly release our code at https://github.com/AugmentedRetrieval/ReasoningAugmentedRetrieval.
Authors:Biao Xiong, Zhen Peng, Ping Wang, Qiegen Liu, Xian Zhong
Abstract:
Automated floorplan generation aims to improve design quality, architectural efficiency, and sustainability by jointly modeling global spatial organization and precise geometric detail. However, existing approaches operate in raster space and rely on post hoc vectorization, which introduces structural inconsistencies and hinders end‑to‑end learning. Motivated by compositional spatial reasoning, we propose TLC‑Plan, a hierarchical generative model that directly synthesizes vector floorplans from input boundaries, aligning with human architectural workflows based on modular and reusable patterns. TLC‑Plan employs a two‑level VQ‑VAE to encode global layouts as semantically labeled room bounding boxes and to refine local geometries using polygon‑level codes. This hierarchy is unified in a CodeTree representation, while an autoregressive transformer samples codes conditioned on the boundary to generate diverse and topologically valid designs, without requiring explicit room topology or dimensional priors. Extensive experiments show state‑of‑the‑art performance on RPLAN dataset (FID = 1.84, MSE = 2.06) and leading results on LIFULL dataset. The proposed framework advances constraint‑aware and scalable vector floorplan generation for real‑world architectural applications. Source code and trained models are released at https://github.com/rosolose/TLC‑PLAN.
Authors:Zihao Fan, Xin Lu, Yidi Liu, Jie Huang, Dong Li, Xueyang Fu, Baocai Yin
Abstract:
Powered by multimodal text‑to‑image priors, diffusion‑based super‑resolution excels at synthesizing intricate details; however, models trained on synthetic low‑resolution (LR) and high‑resolution (HR) image pairs often degrade when applied to real‑world LR images due to significant distribution shifts. We propose Bird‑SR, a bidirectional reward‑guided diffusion framework that formulates super‑resolution as trajectory‑level preference optimization via reward feedback learning (ReFL), jointly leveraging synthetic LR‑HR pairs and real‑world LR images. For structural fidelity easily affected in ReFL, the model is directly optimized on synthetic pairs at early diffusion steps, which also facilitates structure preservation for real‑world inputs under smaller distribution gap in structure levels. For perceptual enhancement, quality‑guided rewards are applied to both synthetic and real LR images at the later trajectory phase. To mitigate reward hacking, the rewards for synthetic results are formulated in a relative advantage space bounded by their ground‑truth counterparts, while real‑world optimization is regularized via a semantic alignment constraint. Furthermore, to balance structural and perceptual learning, we introduce a dynamic fidelity‑perception weighting strategy that emphasizes structure preservation at early stages and progressively shifts focus toward perceptual optimization at later diffusion steps. Extensive experiments on real‑world SR benchmarks demonstrate that Bird‑SR consistently outperforms state‑of‑the‑art methods in perceptual quality while preserving structural consistency, validating its effectiveness for real‑world super‑resolution. Our code can be obtained at https://github.com/fanzh03/Bird‑SR.
Authors:Zhiming Luo, Di Wang, Haonan Guo, Jing Zhang, Bo Du
Abstract:
Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward perception tasks, such as object recognition and scene classification. This limitation hinders the development of MLLMs for cognitively demanding RS applications. To address this, we propose a Vision Language ReaSoning Benchmark (VLRS‑Bench), which is the first benchmark exclusively dedicated to complex RS reasoning. Structured across the three core dimensions of Cognition, Decision, and Prediction, VLRS‑Bench comprises 2,000 question‑answer pairs with an average question length of 130.19 words, spanning 14 tasks and up to eight temporal phases. VLRS‑Bench is constructed via a specialized pipeline that integrates RS‑specific priors and expert knowledge to ensure geospatial realism and reasoning complexity. Experimental results reveal significant bottlenecks in existing state‑of‑the‑art MLLMs, providing critical insights for advancing multimodal reasoning within the remote sensing community. The project repository is available at https://github.com/MiliLab/VLRS‑Bench.
Authors:Yifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen, Qian Zhang, Zhibo Yang, Junyang Lin, Yu Gu, Ge Yu, Maosong Sun
Abstract:
Key Information Extraction (KIE) from real‑world documents remains challenging due to substantial variations in layout structures, visual quality, and task‑specific information requirements. Recent Large Multimodal Models (LMMs) have shown promising potential for performing end‑to‑end KIE directly from document images. To enable a comprehensive and systematic evaluation across realistic and diverse application scenarios, we introduce UNIKIE‑BENCH, a unified benchmark designed to rigorously evaluate the KIE capabilities of LMMs. UNIKIE‑BENCH consists of two complementary tracks: a constrained‑category KIE track with scenario‑predefined schemas that reflect practical application needs, and an open‑category KIE track that extracts any key information that is explicitly present in the document. Experiments on 15 state‑of‑the‑art LMMs reveal substantial performance degradation under diverse schema definitions, long‑tail key fields, and complex layouts, along with pronounced performance disparities across different document types and scenarios. These findings underscore persistent challenges in grounding accuracy and layout‑aware reasoning for LMM‑based KIE. All codes and datasets are available at https://github.com/NEUIR/UNIKIE‑BENCH.
Authors:Ankan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal, Salman Khan, Imran Razzak
Abstract:
Multimodal large language models have advanced rapidly, but their adoption in medicine is constrained by limited domain coverage, imperfect modality alignment, and insufficient grounded reasoning. We introduce MedMO, a medical multimodal foundation model built on a general MLLM architecture and trained exclusively on large‑scale domain‑specific data. MedMO uses a multi‑stage training recipe that includes cross‑modal pretraining to align heterogeneous visual encoders with a medical language backbone, instruction tuning with multi‑task supervision spanning captioning, VQA, report generation, retrieval, and bounding‑box disease localization, and reinforcement learning with verifiable rewards that combine factuality checks with a box‑level GIoU signal to improve spatial grounding and step‑by‑step reasoning in challenging clinical settings. Across modalities and tasks, MedMO surpasses strong open‑source medical baselines. MedMO‑8B‑Next achieves consistent gains on VQA benchmarks, improving by 6.6% on average over Fleming‑VL‑8B, including gains of 6.0% on MMMU‑Med, 9.8% on PMC‑VQA, and 21.3% on MedXpertQA. On text‑based QA, it improves by 14.4% over Fleming‑VL‑8B, driven by gains of 8.4% on MMLU‑Med and 30.1% on MedQA. For medical report generation, it improves by 6.7% on MIMIC‑CXR. MedMO‑8B‑Next also demonstrates strong grounding performance, reaching 56.1 IoU on Bacteria, which is a 47.8 IoU gain over Fleming‑VL‑8B. At smaller scale, MedMO‑4B‑Next remains competitive and exceeds Fleming‑VL‑8B across VQA, QA, and report generation. Evaluations spanning radiology, ophthalmology, and pathology microscopy further confirm broad cross‑modality generalization. Project is available at https://genmilab.github.io/MedMO‑Page
Authors:Kaiyi Huang, Yukun Huang, Yu Li, Jianhong Bai, Xintao Wang, Zinan Lin, Xuefei Ning, Jiwen Yu, Pengfei Wan, Yu Wang, Xihui Liu
Abstract:
Cinematic video production requires control over scene‑subject composition and camera movement, but live‑action shooting remains costly due to the need for constructing physical sets. To address this, we introduce the task of cinematic video generation with decoupled scene context: given multiple images of a static environment, the goal is to synthesize high‑quality videos featuring dynamic subject while preserving the underlying scene consistency and following a user‑specified camera trajectory. We present CineScene, a framework that leverages implicit 3D‑aware scene representation for cinematic video generation. Our key innovation is a novel context conditioning mechanism that injects 3D‑aware features in an implicit way: By encoding scene images into visual representations through VGGT, CineScene injects spatial priors into a pretrained text‑to‑video generation model by additional context concatenation, enabling camera‑controlled video synthesis with consistent scenes and dynamic subjects. To further enhance the model's robustness, we introduce a simple yet effective random‑shuffling strategy for the input scene images during training. To address the lack of training data, we construct a scene‑decoupled dataset with Unreal Engine 5, containing paired videos of scenes with and without dynamic subjects, panoramic images representing the underlying static scene, along with their camera trajectories. Experiments show that CineScene achieves state‑of‑the‑art performance in scene‑consistent cinematic video generation, handling large camera movements and demonstrating generalization across diverse environments.
Authors:Mohammadreza Salehi, Mehdi Noroozi, Luca Morreale, Ruchika Chavhan, Malcolm Chadwick, Alberto Gil Ramos, Abhinav Mehrotra
Abstract:
Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural‑language control. Despite rapid progress, most methods still require fixed‑length inputs and substantial compute. Meanwhile, autoregressive video generation enables efficient variable‑length synthesis, yet remains under‑explored for video editing. We introduce a causal, efficient video editing model that edits variable‑length videos frame by frame. For efficiency, we start from a 2D image‑to‑image (I2I) diffusion model and adapt it to video‑to‑video (V2V) editing by conditioning the edit at time step t on the model's prediction at t‑1. To leverage videos' temporal redundancy, we propose a new I2I diffusion forward process formulation that encourages the model to predict the residual between the target output and the previous prediction. We call this Residual Flow Diffusion Model (RFDM), which focuses the denoising process on changes between consecutive frames. Moreover, we propose a new benchmark that better ranks state‑of‑the‑art methods for editing tasks. Trained on paired video data for global/local style transfer and object removal, RFDM surpasses I2I‑based methods and competes with fully spatiotemporal (3D) V2V models, while matching the compute of image models and scaling independently of input video length. More content can be found in: https://smsd75.github.io/RFDM_page/
Authors:Meng Lou, Stanley Yu, Yizhou Yu
Abstract:
Adapting pre‑trained vision models using parameter‑efficient fine‑tuning (PEFT) remains challenging, as it aims to achieve performance comparable to full fine‑tuning using a minimal number of trainable parameters. When applied to complex dense prediction tasks, existing methods exhibit limitations, including input‑agnostic modeling and redundant cross‑layer representations. To this end, we propose ParaX, a new adapter‑style method featuring a simple mixture‑of‑experts (MoE) architecture. Specifically, we introduce shared expert centers, where each expert is a trainable parameter matrix. During a feedforward pass, each ParaX module in the network dynamically generates weight matrices tailored for the current module via a simple dynamic parameter routing mechanism, which selectively aggregates parameter matrices in the corresponding expert center. Dynamic weight matrices in ParaX modules facilitate low‑rank adaptation in an input‑dependent manner, thus generating more customized and powerful feature representations. Moreover, since ParaX modules across multiple network layers share the same expert center, they improve feature diversity by promoting implicit cross‑layer feature interaction. Extensive experimental results demonstrate the superiority of ParaX across diverse visual recognition tasks. Code is publicly released at: https://github.com/LMMMEng/ParaX.
Authors:Joao Baptista Cardia Neto, Claudio Ferrari, Stefano Berretti
Abstract:
Facial emotion recognition has been typically cast as a single‑label classification problem of one out of six prototypical emotions. However, that is an oversimplification that is unsuitable for representing the multifaceted spectrum of spontaneous emotional states, which are most often the result of a combination of multiple emotions contributing at different intensities. Building on this, a promising direction that was explored recently is to cast emotion recognition as a distribution learning problem. Still, such approaches are limited in that research datasets are typically annotated with a single emotion class. In this paper, we contribute a novel approach to describe complex emotional states as probability distributions over a set of emotion classes. To do so, we propose a solution to automatically re‑label existing datasets by exploiting the result of a study in which a large set of both basic and compound emotions is mapped to probability distributions in the Valence‑Arousal‑Dominance (VAD) space. In this way, given a face image annotated with VAD values, we can estimate the likelihood of it belonging to each of the distributions, so that emotional states can be described as a mixture of emotions, enriching their description, while also accounting for the ambiguous nature of their perception. In a preliminary set of experiments, we illustrate the advantages of this solution and a new possible direction of investigation. Data annotations are available at https://github.com/jbcnrlz/affectnet‑b‑annotation.
Authors:Bo Du, Xiaochen Ma, Xuekang Zhu, Zhe Yang, Chaogun Niu, Chenfan Qu, Mingqi Fang, Zhenming Wang, Jingjing Liu, Jian Liu, Ji-Zhe Zhou
Abstract:
Fake Image Detection (FID), aiming at unified detection across four image forensic subdomains, is critical in real‑world forensic scenarios. Compared with ensemble approaches, monolithic FID models are theoretically more promising, but to date, consistently yield inferior performance in practice. In this work, we identify the intrinsic distinctness of artifacts across subdomains, a critical barrier we term the ``Ji‑Zhe phenomenon". Driven by this phenomenon, we diagnose the cause of this underperformance for the first time: the collapse of the artifact feature space. The core challenge for developing a practical monolithic FID model thus boils down to the ``unified‑yet‑discriminative" reconstruction of the artifact feature space. To address this paradoxical challenge, we hypothesize that high‑level semantics can serve as a structural prior for the reconstruction, and further propose Semantic‑Induced Constrained Adaptation (SICA), the first monolithic FID paradigm. Extensive experiments on our OpenMMSec dataset demonstrate that SICA outperforms 15 state‑of‑the‑art methods and reconstructs the target unified‑yet‑discriminative artifact feature space in a near‑orthogonal manner, thus firmly validating our hypothesis. The code and dataset are available at: https://github.com/venus‑guangjian/SICA_OpenMMSec.
Authors:Junxian Li, Kai Liu, Leyang Chen, Weida Wang, Zhixin Wang, Jiaqi Xu, Fan Li, Renjing Pei, Linghe Kong, Yulun Zhang
Abstract:
Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting computer‑use planning tasks, which are closely related to our lives, remain underexplored. Image generation and editing in computer‑use tasks require capabilities like spatial reasoning and procedural understanding, and it is still unknown whether UMMs have these capabilities to finish these tasks or not. Therefore, we propose PlanViz, a new benchmark designed to evaluate image generation and editing for computer‑use tasks. To achieve the goal of our evaluation, we focus on sub‑tasks which frequently involve in daily life and require planning. Specifically, three representative sub‑tasks are designed: route planning, work diagramming, and web&UI displaying. We address challenges in data quality ensuring by curating human‑annotated questions and reference images, and a quality control process. For detailed and exact evaluation, a task‑adaptive score, PlanScore, is proposed. The score helps understanding the correctness, visual quality and efficiency of generated images. Through experiments, we highlight key limitations and opportunities for future research on this topic.
Authors:Mingxi Xu, Qi Wang, Zhengyu Wen, Phong Dao Thien, Zhengyu Li, Ning Zhang, Xiaoyu He, Wei Zhao, Kehong Gong, Mingyuan Zhang
Abstract:
Motion tokenization is a key component of generalizable motion models, yet most existing approaches are restricted to species‑specific skeletons, limiting their applicability across diverse morphologies. We propose NECromancer (NEC), a universal motion tokenizer that operates directly on arbitrary BVH skeletons. NEC consists of three components: (1) an Ontology‑aware Skeletal Graph Encoder (OwO) that encodes structural priors from BVH files, including joint semantics, rest‑pose offsets, and skeletal topology, into skeletal embeddings; (2) a Topology‑Agnostic Tokenizer (TAT) that compresses motion sequences into a universal, topology‑invariant discrete representation; and (3) the Unified BVH Universe (UvU), a large‑scale dataset aggregating BVH motions across heterogeneous skeletons. Experiments show that NEC achieves high‑fidelity reconstruction under substantial compression and effectively disentangles motion from skeletal structure. The resulting token space supports cross‑species motion transfer, composition, denoising, generation with token‑based models, and text‑motion retrieval, establishing a unified framework for motion analysis and synthesis across diverse morphologies. Demo page: https://animotionlab.github.io/NECromancer/
Authors:Mingyu Dou, Shi Qiu, Ming Hu, Yifan Chen, Huping Ye, Xiaohan Liao, Zhe Sun
Abstract:
Remote sensing change detection plays a pivotal role in domains such as environmental monitoring, urban planning, and disaster assessment. However, existing methods typically rely on predefined categories and large‑scale pixel‑level annotations, which limit their generalization and applicability in open‑world scenarios. To address these limitations, this paper proposes AdaptOVCD, a training‑free Open‑Vocabulary Change Detection (OVCD) architecture based on dual‑dimensional multi‑level information fusion. The framework integrates multi‑level information fusion across data, feature, and decision levels vertically while incorporating targeted adaptive designs horizontally, achieving deep synergy among heterogeneous pre‑trained models to effectively mitigate error propagation. Specifically, (1) at the data level, Adaptive Radiometric Alignment (ARA) fuses radiometric statistics with original texture features and synergizes with SAM‑HQ to achieve radiometrically consistent segmentation; (2) at the feature level, Adaptive Change Thresholding (ACT) combines global difference distributions with edge structure priors and leverages DINOv3 to achieve robust change detection; (3) at the decision level, Adaptive Confidence Filtering (ACF) integrates semantic confidence with spatial constraints and collaborates with DGTRS‑CLIP to achieve high‑confidence semantic identification. Comprehensive evaluations across nine scenarios demonstrate that AdaptOVCD detects arbitrary category changes in a zero‑shot manner, significantly outperforming existing training‑free methods. Meanwhile, it achieves 84.89% of the fully‑supervised performance upper bound in cross‑dataset evaluations and exhibits superior generalization capabilities. The code is available at https://github.com/Dmygithub/AdaptOVCD.
Authors:Feiyang jia, Lin Liu, Ziying Song, Caiyan Jia, Hangjun Ye, Xiaoshuai Hao, Long Chen
Abstract:
End‑to‑end (E2E) autonomous driving has recently attracted increasing interest in unifying Vision‑Language‑Action (VLA) with World Models to enhance decision‑making and forward‑looking imagination. However, existing methods fail to effectively unify future scene evolution and action planning within a single architecture due to inadequate sharing of latent states, limiting the impact of visual imagination on action decisions. To address this limitation, we propose DriveWorld‑VLA, a novel framework that unifies world modeling and planning within a latent space by tightly integrating VLA and world models at the representation level, which enables the VLA planner to benefit directly from holistic scene‑evolution modeling and reducing reliance on dense annotated supervision. Additionally, DriveWorld‑VLA incorporates the latent states of the world model as core decision‑making states for the VLA planner, facilitating the planner to assess how candidate actions impact future scene evolution. By conducting world modeling entirely in the latent space, DriveWorld‑VLA supports controllable, action‑conditioned imagination at the feature level, avoiding expensive pixel‑level rollouts. Extensive open‑loop and closed‑loop evaluations demonstrate the effectiveness of DriveWorld‑VLA, which achieves state‑of‑the‑art performance with 91.3 PDMS on NAVSIMv1, 86.8 EPDMS on NAVSIMv2, and 0.16 3‑second average collision rate on nuScenes. Code and models will be released in https://github.com/liulin815/DriveWorld‑VLA.git.
Authors:Xingsong Ye, Yongkun Du, JiaXin Zhang, Chen Li, Jing Lyu, Zhineng Chen
Abstract:
Large‑scale and categorical‑balanced text data is essential for training effective Scene Text Recognition (STR) models, which is hard to achieve when collecting real data. Synthetic data offers a cost‑effective and perfectly labeled alternative. However, its performance often lags behind, revealing a significant domain gap between real and current synthetic data. In this work, we systematically analyze mainstream rendering‑based synthetic datasets and identify their key limitations: insufficient diversity in corpus, font, and layout, which restricts their realism in complex scenarios. To address these issues, we introduce UnionST, a strong data engine synthesizes text covering a union of challenging samples and better aligns with the complexity observed in the wild. We then construct UnionST‑S, a large‑scale synthetic dataset with improved simulations in challenging scenarios. Furthermore, we develop a self‑evolution learning (SEL) framework for effective real data annotation. Experiments show that models trained on UnionST‑S achieve significant improvements over existing synthetic datasets. They even surpass real‑data performance in certain scenarios. Moreover, when using SEL, the trained models achieve competitive performance by only seeing 9% of real data labels. Code is available at https://github.com/YesianRohn/UnionST.
Authors:Yunze Tong, Mushui Liu, Canyu Zhao, Wanggui He, Shiyi Zhang, Hongwei Zhang, Peng Zhang, Jinlong Liu, Ju Huang, Jiamang Wang, Hao Jiang, Pipei Huang
Abstract:
Deploying GRPO on Flow Matching models has proven effective for text‑to‑image generation. However, existing paradigms typically propagate an outcome‑based reward to all preceding denoising steps without distinguishing the local effect of each step. Moreover, current group‑wise ranking mainly compares trajectories at matched timesteps and ignores within‑trajectory dependencies, where certain early denoising actions can affect later states via delayed, implicit interactions. We propose TurningPoint‑GRPO (TP‑GRPO), a GRPO framework that alleviates step‑wise reward sparsity and explicitly models long‑term effects within the denoising trajectory. TP‑GRPO makes two key innovations: (i) it replaces outcome‑based rewards with step‑level incremental rewards, providing a dense, step‑aware learning signal that better isolates each denoising action's "pure" effect, and (ii) it identifies turning points‑steps that flip the local reward trend and make subsequent reward evolution consistent with the overall trajectory trend‑and assigns these actions an aggregated long‑term reward to capture their delayed impact. Turning points are detected solely via sign changes in incremental rewards, making TP‑GRPO efficient and hyperparameter‑free. Extensive experiments also demonstrate that TP‑GRPO exploits reward signals more effectively and consistently improves generation. Demo code is available at https://github.com/YunzeTong/TurningPoint‑GRPO.
Authors:Zhenxing Ming, Yaoqi Huang, Julie Stephany Berrio, Mao Shan, Stewart Worrall
Abstract:
The prediction of 3D semantic occupancy enables autonomous vehicles (AVs) to perceive the fine‑grained geometric and semantic scene structure for safe navigation and decision‑making. Existing methods mainly rely on either voxel‑based representations, which incur redundant computation over empty regions, or on object‑centric Gaussian primitives, which are limited in modeling complex, non‑convex, and asymmetric structures. In this paper, we present TFusionOcc, a T‑primitive‑based object‑centric multi‑sensor fusion framework for 3D semantic occupancy prediction. Specifically, we introduce a family of Students t‑distribution‑based T‑primitives, including the plain T‑primitive, T‑Superquadric, and deformable T‑Superquadric with inverse warping, where the deformable T‑Superquadric serves as the key geometry‑enhancing primitive. We further develop a unified probabilistic formulation based on the Students t‑distribution and the T‑mixture model (TMM) to jointly model occupancy and semantics, and design a tightly coupled multi‑stage fusion architecture to effectively integrate camera and LiDAR cues. Extensive experiments on nuScenes show state‑of‑the‑art performance, while additional evaluations on nuScenes‑C demonstrate strong robustness under most corruption scenarios. The code will be available at: https://github.com/DanielMing123/TFusionOcc
Authors:Fuxi Zhang, Yifan Wang, Hengrun Zhao, Zhuohan Sun, Changxing Xia, Lijun Wang, Huchuan Lu, Yangrui Shao, Chen Yang, Long Teng
Abstract:
Salient object detection is inherently a subjective problem, as observers with different priors may perceive different objects as salient. However, existing methods predominantly formulate it as an objective prediction task with a single groundtruth segmentation map for each image, which renders the problem under‑determined and fundamentally ill‑posed. To address this issue, we propose Observer‑Centric Salient Object Detection (OC‑SOD), where salient regions are predicted by considering not only the visual cues but also the observer‑specific factors such as their preferences or intents. As a result, this formulation captures the intrinsic ambiguity and diversity of human perception, enabling personalized and context‑aware saliency prediction. By leveraging multi‑modal large language models, we develop an efficient data annotation pipeline and construct the first OC‑SOD dataset named OC‑SODBench, comprising 33k training, validation and test images with 152k textual prompts and object pairs. Built upon this new dataset, we further design OC‑SODAgent, an agentic baseline which performs OC‑SOD via a human‑like "Perceive‑Reflect‑Adjust" process. Extensive experiments on our proposed OC‑SODBench have justified the effectiveness of our contribution. Through this observer‑centric perspective, we aim to bridge the gap between human perception and computational modeling, offering a more realistic and flexible understanding of what makes an object truly "salient." Code and dataset are publicly available at: https://github.com/Dustzx/OC_SOD
Authors:Gensheng Pei, Xiruo Jiang, Yazhou Yao, Xiangbo Shu, Fumin Shen, Byeungwoo Jeon
Abstract:
The recent introduction of \textttSAM3 has revolutionized Open‑Vocabulary Segmentation (OVS) through promptable concept segmentation, which grounds pixel predictions in flexible concept prompts. However, this reliance on pre‑defined concepts makes the model vulnerable: when visual distributions shift (data drift) or conditional label distributions evolve (concept drift) in the target domain, the alignment between visual evidence and prompts breaks down. In this work, we present \textscConceptBank, a parameter‑free calibration framework to restore this alignment on the fly. Instead of adhering to static prompts, we construct a dataset‑specific concept bank from the target statistics. Our approach (i) anchors target‑domain evidence via class‑wise visual prototypes, (ii) mines representative supports to suppress outliers under data drift, and (iii) fuses candidate concepts to rectify concept drift. We demonstrate that \textscConceptBank effectively adapts \textttSAM3 to distribution drifts, including challenging natural‑scene and remote‑sensing scenarios, establishing a new baseline for robustness and efficiency in OVS. Code and model are available at https://github.com/pgsmall/ConceptBank.
Authors:Jiazheng Wang, Zeyu Liu, Min Liu, Xiang Chen, Xinyao Yu, Yaonan Wang, Hang Zhang
Abstract:
Surgical navigation based on multimodal image registration has played a significant role in providing intraoperative guidance to surgeons by showing the relative position of the target area to critical anatomical structures during surgery. However, due to the differences between multimodal images and intraoperative image deformation caused by tissue displacement and removal during the surgery, effective registration of preoperative and intraoperative multimodal images faces significant challenges. To address the multimodal image registration challenges in Learn2Reg 2025, an unsupervised multimodal medical image registration method based on Multilevel Correlation Pyramidal Optimization (MCPO) is designed to solve these problems. First, the features of each modality are extracted based on the modality independent neighborhood descriptor, and the multimodal images is mapped to the feature space. Second, a multilevel pyramidal fusion optimization mechanism is designed to achieve global optimization and local detail complementation of the displacement field through dense correlation analysis and weight‑balanced coupled convex optimization for input features at different scales. Our method focuses on the ReMIND2Reg task in Learn2Reg 2025. Based on the results, our method achieved the first place in the validation phase and test phase of ReMIND2Reg. The MCPO is also validated on the Resect dataset, achieving an average TRE of 1.798 mm. This demonstrates the broad applicability of our method in preoperative‑to‑intraoperative image registration. The code is available at https://github.com/wjiazheng/MCPO.
Authors:Yuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang, Zhaojie Fang, Chenghong Li, Xiaoguang Han
Abstract:
The ubiquity of monocular videos capturing daily hand‑object interactions presents a valuable resource for embodied intelligence. While 3D hand reconstruction from in‑the‑wild videos has seen significant progress, reconstructing the involved objects remains challenging due to severe occlusions and the complex, coupled motion of the camera, hands, and object. In this paper, we introduce ForeHOI, a novel feed‑forward model that directly reconstructs 3D object geometry from monocular hand‑object interaction videos within one minute of inference time, eliminating the need for any pre‑processing steps. Our key insight is that, the joint prediction of 2D mask inpainting and 3D shape completion in a feed‑forward framework can effectively address the problem of severe occlusion in monocular hand‑held object videos, thereby achieving results that outperform the performance of optimization‑based methods. The information exchanges between the 2D and 3D shape completion boosts the overall reconstruction quality, enabling the framework to effectively handle severe hand‑object occlusion. Furthermore, to support the training of our model, we contribute the first large‑scale, high‑fidelity synthetic dataset of hand‑object interactions with comprehensive annotations. Extensive experiments demonstrate that ForeHOI achieves state‑of‑the‑art performance in object reconstruction, significantly outperforming previous methods with around a 100x speedup. Code and data are available at: https://github.com/Tao‑11‑chen/ForeHOI.
Authors:Anika Knupfer, Johanna P. Müller, Jordina A. Verdera, Martin Fenske, Claudius S. Mathy, Smiti Tripathy, Sebastian Arndt, Matthias May, Michael Uder, Matthias W. Beckmann, Stefanie Burghaus, Jana Hutter
Abstract:
Pelvic diseases in women of reproductive age represent a major global health burden, with diagnosis frequently delayed due to high anatomical variability, complicating MRI interpretation. Existing AI approaches are largely disease‑specific and lack real‑time compatibility, limiting generalizability and clinical integration. To address these challenges, we establish a benchmark framework for disease‑ and parameter‑agnostic, real‑time‑compatible unsupervised anomaly detection in pelvic MRI. The method uses a residual variational autoencoder trained exclusively on healthy sagittal T2‑weighted scans acquired across diverse imaging protocols to model normal pelvic anatomy. During inference, reconstruction error heatmaps indicate deviations from learned healthy structure, enabling detection of pathological regions without labeled abnormal data. The model is trained on 294 healthy scans and augmented with diffusion‑generated synthetic data to improve robustness. Quantitative evaluation on the publicly available Uterine Myoma MRI Dataset yields an average area‑under‑the‑curve (AUC) value of 0.736, with 0.828 sensitivity and 0.692 specificity. Additional inter‑observer clinical evaluation extends analysis to endometrial cancer, endometriosis, and adenomyosis, revealing the influence of anatomical heterogeneity and inter‑observer variability on performance interpretation. With a reconstruction time of approximately 92.6 frames per second, the proposed framework establishes a baseline for unsupervised anomaly detection in the female pelvis and supports future integration into real‑time MRI. Code is available upon request (https://github.com/AniKnu/UADPelvis), prospective data sets are available for academic collaboration.
Authors:Xuyang Chen, Conglang Zhang, Chuanheng Fu, Zihao Yang, Kaixuan Zhou, Yizhi Zhang, Jianan He, Yanfeng Zhang, Mingwei Sun, Zengmao Wang, Zhen Dong, Xiaoxiao Long, Liqiu Meng
Abstract:
Driven by the emergence of Controllable Video Diffusion, existing Sim2Real methods for autonomous driving video generation typically rely on explicit intermediate representations to bridge the domain gap. However, these modalities face a fundamental Consistency‑Realism Dilemma. Low‑level signals (e.g., edges, blurred images) ensure precise control but compromise realism by "baking in" synthetic artifacts, whereas high‑level priors (e.g., depth, semantics, HDMaps) facilitate photorealism but lack the structural detail required for consistent guidance. In this work, we present Driving with DINO (DwD), a novel framework that leverages Vision Foundation Module (VFM) features as a unified bridge between the simulation and real‑world domains. We first identify that these features encode a spectrum of information, from high‑level semantics to fine‑grained structure. To effectively utilize this, we employ Principal Subspace Projection to discard the high‑frequency elements responsible for "texture baking," while concurrently introducing Random Channel Tail Drop to mitigate the structural loss inherent in rigid dimensionality reduction, thereby reconciling realism with control consistency. Furthermore, to fully leverage DINOv3's high‑resolution capabilities for enhancing control precision, we introduce a learnable Spatial Alignment Module that adapts these high‑resolution features to the diffusion backbone. Finally, we propose a Causal Temporal Aggregator employing causal convolutions to explicitly preserve historical motion context when integrating frame‑wise DINO features, which effectively mitigates motion blur and guarantees temporal stability. Project page: https://albertchen98.github.io/DwD‑project/
Authors:Ding-Jiun Huang, Yuanhao Wang, Shao-Ji Yuan, Albert Mosella-Montoro, Francisco Vicente Carrasco, Cheng Zhang, Fernando De la Torre
Abstract:
Creating high‑fidelity, animatable 3D talking heads is crucial for immersive applications, yet often hindered by the prevalence of low‑quality image or video sources, which yield poor 3D reconstructions. In this paper, we introduce SuperHead, a novel framework for enhancing low‑resolution, animatable 3D head avatars. The core challenge lies in synthesizing high‑quality geometry and textures, while ensuring both 3D and temporal consistency during animation and preserving subject identity. Despite recent progress in image, video and 3D‑based super‑resolution (SR), existing SR techniques are ill‑equipped to handle dynamic 3D inputs. To address this, SuperHead leverages the rich priors from pre‑trained 3D generative models via a novel dynamics‑aware 3D inversion scheme. This process optimizes the latent representation of the generative model to produce a super‑resolved 3D Gaussian Splatting (3DGS) head model, which is subsequently rigged to an underlying parametric head model (e.g., FLAME) for animation. The inversion is jointly supervised using a sparse collection of upscaled 2D face renderings and corresponding depth maps, captured from diverse facial expressions and camera viewpoints, to ensure realism under dynamic facial motions. Experiments demonstrate that SuperHead generates avatars with fine‑grained facial details under dynamic motions, significantly outperforming baseline methods in visual quality.
Authors:Jongha Kim, Byungoh Ko, Jeehye Na, Jinsung Yoon, Hyunwoo J. Kim
Abstract:
Despite the remarkable capabilities of Large Vision Language Models (LVLMs), they still lack detailed knowledge about specific entities. Retrieval‑augmented Generation (RAG) is a widely adopted solution that enhances LVLMs by providing additional contexts from an external Knowledge Base. However, we observe that previous decoding methods for RAG are sub‑optimal as they fail to sufficiently leverage multiple relevant contexts and suppress the negative effects of irrelevant contexts. To this end, we propose Relevance‑aware Multi‑context Contrastive Decoding (RMCD), a novel decoding method for RAG. RMCD outputs a final prediction by combining outputs predicted with each context, where each output is weighted based on its relevance to the question. By doing so, RMCD effectively aggregates useful information from multiple relevant contexts while also counteracting the negative effects of irrelevant ones. Experiments show that RMCD consistently outperforms other decoding methods across multiple LVLMs, achieving the best performance on three knowledge‑intensive visual question‑answering benchmarks. Also, RMCD can be simply applied by replacing the decoding method of LVLMs without additional training. Analyses also show that RMCD is robust to the retrieval results, consistently performing the best across the weakest to the strongest retrieval results. Code is available at https://github.com/mlvlab/RMCD.
Authors:Jintao Tong, Shilin Yan, Hongwei Xue, Xiaojun Tang, Kunyu Shi, Guannan Zhang, Ruixuan Li, Yixiong Zou
Abstract:
Multimodal Large Language Models (MLLMs) have made remarkable progress in multimodal perception and reasoning by bridging vision and language. However, most existing MLLMs perform reasoning primarily with textual CoT, which limits their effectiveness on vision‑intensive tasks. Recent approaches inject a fixed number of continuous hidden states as "visual thoughts" into the reasoning process and improve visual performance, but often at the cost of degraded text‑based logical reasoning. We argue that the core limitation lies in a rigid, pre‑defined reasoning pattern that cannot adaptively choose the most suitable thinking modality for different user queries. We introduce SwimBird, a reasoning‑switchable MLLM that dynamically switches among three reasoning modes conditioned on the input: (1) text‑only reasoning, (2) vision‑only reasoning (continuous hidden states as visual thoughts), and (3) interleaved vision‑text reasoning. To enable this capability, we adopt a hybrid autoregressive formulation that unifies next‑token prediction for textual thoughts with next‑embedding prediction for visual thoughts, and design a systematic reasoning‑mode curation strategy to construct SwimBird‑SFT‑92K, a diverse supervised fine‑tuning dataset covering all three reasoning patterns. By enabling flexible, query‑adaptive mode selection, SwimBird preserves strong textual logic while substantially improving performance on vision‑dense tasks. Experiments across diverse benchmarks covering textual reasoning and challenging visual understanding demonstrate that SwimBird achieves state‑of‑the‑art results and robust gains over prior fixed‑pattern multimodal reasoning methods.
Authors:Haoyuan Li, Qihang Cao, Tao Tang, Kun Xiang, Zihan Guo, Jianhua Han, JiaWang Bian, Hang Xu, Xiaodan Liang
Abstract:
Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is exposed as a global stream and fused in an indiscriminate manner, which often induces semantic‑geometry misalignment and redundant signals. We propose GeoThinker, a framework that shifts the paradigm from passive fusion to active perception. Instead of feature mixing, GeoThinker enables the model to selectively retrieve geometric evidence conditioned on its internal reasoning demands. GeoThinker achieves this through Spatial‑Grounded Fusion applied at carefully selected VLM layers, where semantic visual priors selectively query and integrate task‑relevant geometry via frame‑strict cross‑attention, further calibrated by Importance Gating that biases per‑frame attention toward task‑relevant structures. Comprehensive evaluation results show that GeoThinker sets a new state‑of‑the‑art in spatial intelligence, achieving a peak score of 72.6 on the VSI‑Bench. Furthermore, GeoThinker demonstrates robust generalization and significantly improved spatial perception across complex downstream scenarios, including embodied referring and autonomous driving. Our results indicate that the ability to actively integrate spatial structures is essential for next‑generation spatial intelligence. Code can be found at https://github.com/Li‑Hao‑yuan/GeoThinker.
Authors:Sirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He, Xiaohan Fei, Yu-Xiong Wang, Liangyan Gui
Abstract:
Humans rarely plan whole‑body interactions with objects at the level of explicit whole‑body movements. High‑level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally from underlying physical and motor priors. Scaling such priors is key to enabling humanoids to compose and generalize loco‑manipulation skills across diverse contexts while maintaining physically coherent whole‑body coordination. To this end, we introduce InterPrior, a scalable framework that learns a unified generative controller through large‑scale imitation pretraining and post‑training by reinforcement learning. InterPrior first distills a full‑reference imitation expert into a versatile, goal‑conditioned variational policy that reconstructs motion from multimodal observations and high‑level intent. While the distilled policy reconstructs training behaviors, it does not generalize reliably due to the vast configuration space of large‑scale human‑object interactions. To address this, we apply data augmentation with physical perturbations, and then perform reinforcement learning finetuning to improve competence on unseen goals and initializations. Together, these steps consolidate the reconstructed latent skills into a valid manifold, yielding a motion prior that generalizes beyond the training data, e.g., it can incorporate new behaviors such as interactions with unseen objects. We further demonstrate its effectiveness for user‑interactive control and its potential for real robot deployment.
Authors:Dongyang Chen, Chaoyang Wang, Dezhao Su, Xi Xiao, Zeyu Zhang, Jing Xiong, Qing Li, Yuzhang Shang, Shichao Kan
Abstract:
Multimodal Large Language Models (MLLMs) have recently been applied to universal multimodal retrieval, where Chain‑of‑Thought (CoT) reasoning improves candidate reranking. However, existing approaches remain largely language‑driven, relying on static visual encodings and lacking the ability to actively verify fine‑grained visual evidence, which often leads to speculative reasoning in visually ambiguous cases. We propose V‑Retrver, an evidence‑driven retrieval framework that reformulates multimodal retrieval as an agentic reasoning process grounded in visual inspection. V‑Retrver enables an MLLM to selectively acquire visual evidence during reasoning via external visual tools, performing a multimodal interleaved reasoning process that alternates between hypothesis generation and targeted visual verification.To train such an evidence‑gathering retrieval agent, we adopt a curriculum‑based learning strategy combining supervised reasoning activation, rejection‑based refinement, and reinforcement learning with an evidence‑aligned objective. Experiments across multiple multimodal retrieval benchmarks demonstrate consistent improvements in retrieval accuracy (with 23.0% improvements on average), perception‑driven reasoning reliability, and generalization.
Authors:David Shavin, Sagie Benaim
Abstract:
Vision Foundation Models (VFMs) have achieved remarkable success when applied to various downstream 2D tasks. Despite their effectiveness, they often exhibit a critical lack of 3D awareness. To this end, we introduce Splat and Distill, a framework that instills robust 3D awareness into 2D VFMs by augmenting the teacher model with a fast, feed‑forward 3D reconstruction pipeline. Given 2D features produced by a teacher model, our method first lifts these features into an explicit 3D Gaussian representation, in a feedforward manner. These 3D features are then ``splatted" onto novel viewpoints, producing a set of novel 2D feature maps used to supervise the student model, ``distilling" geometrically grounded knowledge. By replacing slow per‑scene optimization of prior work with our feed‑forward lifting approach, our framework avoids feature‑averaging artifacts, creating a dynamic learning process where the teacher's consistency improves alongside that of the student. We conduct a comprehensive evaluation on a suite of downstream tasks, including monocular depth estimation, surface normal estimation, multi‑view correspondence, and semantic segmentation. Our method significantly outperforms prior works, not only achieving substantial gains in 3D awareness but also enhancing the underlying semantic richness of 2D features. Project page is available at https://davidshavin4.github.io/Splat‑and‑Distill/
Authors:Ruihang Li, Leigang Qu, Jingxu Zhang, Dongnan Gui, Mengde Xu, Xiaosong Zhang, Han Hu, Wenjie Wang, Jiaqi Wang
Abstract:
The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision‑Language Models as surrogate judges. In this work, we systematically investigate the reliability of the prevailing absolute pointwise scoring standard, across a wide spectrum of visual generation tasks. Our analysis reveals that this paradigm is limited due to stochastic inconsistency and poor alignment with human perception. To resolve these limitations, we introduce GenArena, a unified evaluation framework that leverages a pairwise comparison paradigm to ensure stable and human‑aligned evaluation. Crucially, our experiments uncover a transformative finding that simply adopting this pairwise protocol enables off‑the‑shelf open‑source models to outperform top‑tier proprietary models. Notably, our method boosts evaluation accuracy by over 20% and achieves a Spearman correlation of 0.86 with the authoritative LMArena leaderboard, drastically surpassing the 0.36 correlation of pointwise methods. Based on GenArena, we benchmark state‑of‑the‑art visual generation models across diverse tasks, providing the community with a rigorous and automated evaluation standard for visual generation.
Authors:Mingxin Liu, Shuran Ma, Shibei Meng, Xiangyu Zhao, Zicheng Zhang, Shaofeng Zhang, Zhihang Zhong, Peixian Chen, Haoyu Cao, Xing Sun, Haodong Duan, Xue Yang
Abstract:
While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under‑explored frontier. To bridge this gap, we present RISE‑Video, a pioneering reasoning‑oriented benchmark for Text‑Image‑to‑Video (TI2V) synthesis that shifts the evaluative focus from surface‑level aesthetics to deep cognitive reasoning. RISE‑Video comprises 467 meticulously human‑annotated samples spanning eight rigorous categories, providing a structured testbed for probing model intelligence across diverse dimensions, ranging from commonsense and spatial dynamics to specialized subject domains. Our framework introduces a multi‑dimensional evaluation protocol consisting of four metrics: Reasoning Alignment, Temporal Consistency, Physical Rationality, and Visual Quality. To further support scalable evaluation, we propose an automated pipeline leveraging Large Multimodal Models (LMMs) to emulate human‑centric assessment. Extensive experiments on 11 state‑of‑the‑art TI2V models reveal pervasive deficiencies in simulating complex scenarios under implicit constraints, offering critical insights for the advancement of future world‑simulating generative models.
Authors:Mirlan Karimov, Teodora Spasojevic, Markus Braun, Julian Wiederer, Vasileios Belagiannis, Marc Pollefeys
Abstract:
Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control signals at inference time to guide the generative model towards temporally consistent generation of dynamic objects, limiting their utility as scalable and generalizable data engines. In this work, we propose Localized Semantic Alignment (LSA), a simple yet effective framework for fine‑tuning pre‑trained video generation models. LSA enhances temporal consistency by aligning semantic features between ground‑truth and generated video clips. Specifically, we compare the output of an off‑the‑shelf feature extraction model between the ground‑truth and generated video clips localized around dynamic objects inducing a semantic feature consistency loss. We fine‑tune the base model by combining this loss with the standard diffusion loss. The model fine‑tuned for a single epoch with our novel loss outperforms the baselines in common video generation evaluation metrics. To further test the temporal consistency in generated videos we adapt two additional metrics from object detection task, namely mAP and mIoU. Extensive experiments on nuScenes and KITTI datasets show the effectiveness of our approach in enhancing temporal consistency in video generation without the need for external control signals during inference and any computational overheads.
Authors:Enwei Tong, Yuanchao Bai, Yao Zhu, Junjun Jiang, Xianming Liu
Abstract:
Vision‑language models (VLMs) often generate massive visual tokens that greatly increase inference latency and memory footprint; while training‑free token pruning offers a practical remedy, existing methods still struggle to balance local evidence and global context under aggressive compression. We propose Focus‑Scan‑Refine (FSR), a human‑inspired, plug‑and‑play pruning framework that mimics how humans answer visual questions: focus on key evidence, then scan globally if needed, and refine the scanned context by aggregating relevant details. FSR first focuses on key evidence by combining visual importance with instruction relevance, avoiding the bias toward visually salient but query‑irrelevant regions. It then scans for complementary context conditioned on the focused set, selecting tokens that are most different from the focused evidence. Finally, FSR refines the scanned context by aggregating nearby informative tokens into the scan anchors via similarity‑based assignment and score‑weighted merging, without increasing the token budget. Extensive experiments across multiple VLM backbones and vision‑language benchmarks show that FSR consistently improves the accuracy‑efficiency trade‑off over existing state‑of‑the‑art pruning methods. The source codes can be found at https://github.com/ILOT‑code/FSR.
Authors:Ti Wang, Xiaohang Yu, Mackenzie Weygandt Mathis
Abstract:
Monocular 3D pose estimation is fundamentally ill‑posed due to depth ambiguity and occlusions, thereby motivating probabilistic methods that generate multiple plausible 3D pose hypotheses. In particular, diffusion‑based models have recently demonstrated strong performance, but their iterative denoising process typically requires many timesteps for each prediction, making inference computationally expensive. In contrast, we leverage Flow Matching (FM) to learn a velocity field defined by an Ordinary Differential Equation (ODE), enabling efficient generation of 3D pose samples with only a few integration steps. We propose a novel generative pose estimation framework, FMPose3D, that formulates 3D pose estimation as a conditional distribution transport problem. It continuously transports samples from a standard Gaussian prior to the distribution of plausible 3D poses conditioned only on 2D inputs. Although ODE trajectories are deterministic, FMPose3D naturally generates various pose hypotheses by sampling different noise seeds. To obtain a single accurate prediction from those hypotheses, we further introduce a Reprojection‑based Posterior Expectation Aggregation (RPEA) module, which approximates the Bayesian posterior expectation over 3D hypotheses. FMPose3D surpasses existing methods on the widely used human pose estimation benchmarks Human3.6M and MPI‑INF‑3DHP, and further achieves state‑of‑the‑art performance on the 3D animal pose datasets Animal3D and CtrlAni3D, demonstrating strong performance across both 3D pose domains. The code is available at https://github.com/AdaptiveMotorControlLab/FMPose3D.
Authors:Moussa Kassem Sbeyti, Nadja Klein
Abstract:
Detecting small and distant objects remains challenging for object detectors due to scale variation, low resolution, and background clutter. Safety‑critical applications require reliable detection of these objects for safe planning. Depth information can improve detection, but existing approaches require complex, model‑specific architectural modifications. We provide a theoretical analysis followed by an empirical investigation of the depth‑detection relationship. Together, they explain how depth causes systematic performance degradation and why depth‑informed supervision mitigates it. We introduce DepthPrior, a framework that uses depth as prior knowledge rather than as a fused feature, providing comparable benefits without modifying detector architectures. DepthPrior consists of Depth‑Based Loss Weighting (DLW) and Depth‑Based Loss Stratification (DLS) during training, and Depth‑Aware Confidence Thresholding (DCT) during inference. The only overhead is the initial cost of depth estimation. Experiments across four benchmarks (KITTI, MS COCO, VisDrone, SUN RGB‑D) and two detectors (YOLOv11, EfficientDet) demonstrate the effectiveness of DepthPrior, achieving up to +9% mAP_S and +7% mAR_S for small objects, with inference recovery rates as high as 95:1 (true vs. false detections). DepthPrior offers these benefits without additional sensors, architectural changes, or performance costs. Code is available at https://github.com/mos‑ks/DepthPrior.
Authors:Inbar Gat, Dana Cohen-Bar, Guy Levy, Elad Richardson, Daniel Cohen-Or
Abstract:
Recent advancements in 3D foundation models have enabled the generation of high‑fidelity assets, yet precise 3D manipulation remains a significant challenge. Existing 3D editing frameworks often face a difficult trade‑off between visual controllability, geometric consistency, and scalability. Specifically, optimization‑based methods are prohibitively slow, multi‑view 2D propagation techniques suffer from visual drift, and training‑free latent manipulation methods are inherently bound by frozen priors and cannot directly benefit from scaling. In this work, we present ShapeUP, a scalable, image‑conditioned 3D editing framework that formulates editing as a supervised latent‑to‑latent translation within a native 3D representation. This formulation allows ShapeUP to build on a pretrained 3D foundation model, leveraging its strong generative prior while adapting it to editing through supervised training. In practice, ShapeUP is trained on triplets consisting of a source 3D shape, an edited 2D image, and the corresponding edited 3D shape, and learns a direct mapping using a 3D Diffusion Transformer (DiT). This image‑as‑prompt approach enables fine‑grained visual control over both local and global edits and achieves implicit, mask‑free localization, while maintaining strict structural consistency with the original asset. Our extensive evaluations demonstrate that ShapeUP consistently outperforms current trained and training‑free baselines in both identity preservation and edit fidelity, offering a robust and scalable paradigm for native 3D content creation.
Authors:Nikolay Patakin, Arsenii Shirokov, Anton Konushin, Dmitry Senushkin
Abstract:
In this work, we introduce XSIM, a sensor simulation framework for autonomous driving. XSIM extends 3DGUT splatting with a generalized rolling‑shutter modeling tailored for autonomous driving applications. Our framework provides a unified and flexible formulation for appearance and geometric sensor modeling, enabling rendering of complex sensor distortions in dynamic environments. We identify spherical cameras, such as LiDARs, as a critical edge case for existing 3DGUT splatting due to cyclic projection and time discontinuities at azimuth boundaries leading to incorrect particle projection. To address this issue, we propose a phase modeling mechanism that explicitly accounts temporal and shape discontinuities of Gaussians projected by the Unscented Transform at azimuth borders. In addition, we introduce an extended 3D Gaussian representation that incorporates two distinct opacity parameters to resolve mismatches between geometry and color distributions. As a result, our framework provides enhanced scene representations with improved geometric consistency and photorealistic appearance. We evaluate our framework extensively on multiple autonomous driving datasets, including Waymo Open Dataset, Argoverse 2, and PandaSet. Our framework consistently outperforms strong recent baselines and achieves state‑of‑the‑art performance across all datasets. The source code is publicly available at \hrefhttps://github.com/whesense/XSIMhttps://github.com/whesense/XSIM.
Authors:Zongliang Zhang, Shuxiang Li, Xingwang Huang, Zongyue Wang
Abstract:
Most existing robust fitting methods are designed for classical models, such as lines, circles, and planes. In contrast, fewer methods have been developed to robustly handle non‑classical models, such as spiral curves, procedural character models, and free‑form surfaces. Furthermore, existing methods primarily focus on reconstructing a single instance of a non‑classical model. This paper aims to reconstruct multiple instances of non‑classical models from noisy data. We formulate this multi‑instance fitting task as an optimization problem, which comprises an estimator and an optimizer. Specifically, we propose a novel estimator based on the model‑to‑data error, capable of handling outliers without a predefined error threshold. Since the proposed estimator is non‑differentiable with respect to the model parameters, we employ a meta‑heuristic algorithm as the optimizer to seek the global optimum. The effectiveness of our method are demonstrated through experimental results on various non‑classical models. The code is available at https://github.com/zhangzongliang/fitting.
Authors:Arsenii Shirokov, Mikhail Kuznetsov, Danila Stepochkin, Egor Evdokimov, Daniil Glazkov, Nikolay Patakin, Anton Konushin, Dmitry Senushkin
Abstract:
We introduce the Visual Implicit Geometry Transformer (ViGT), an autonomous driving geometric model that estimates continuous 3D occupancy fields from surround‑view camera rigs. ViGT represents a step towards foundational geometric models for autonomous driving, prioritizing scalability, architectural simplicity, and generalization across diverse sensor configurations. Our approach achieves this through a calibration‑free architecture, enabling a single model to adapt to different sensor setups. Unlike general‑purpose geometric foundational models that focus on pixel‑aligned predictions, ViGT estimates a continuous 3D occupancy field in a birds‑eye‑view (BEV) addressing domain‑specific requirements. ViGT naturally infers geometry from multiple camera views into a single metric coordinate frame, providing a common representation for multiple geometric tasks. Unlike most existing occupancy models, we adopt a self‑supervised training procedure that leverages synchronized image‑LiDAR pairs, eliminating the need for costly manual annotations. We validate the scalability and generalizability of our approach by training our model on a mixture of five large‑scale autonomous driving datasets (NuScenes, Waymo, NuPlan, ONCE, and Argoverse) and achieving state‑of‑the‑art performance on the pointmap estimation task, with the best average rank across all evaluated baselines. We further evaluate ViGT on the Occ3D‑nuScenes benchmark, where ViGT achieves comparable performance with supervised methods. The source code is publicly available at \hrefhttps://github.com/whesense/ViGThttps://github.com/whesense/ViGT.
Authors:Michael Schwingshackl, Fabio F. Oberweger, Mario Niedermeyer, Huemer Johannes, Markus Murschitz
Abstract:
We present PIRATR, an end‑to‑end 3D object detection framework for robotic use cases in point clouds. Extending PI3DETR, our method streamlines parametric 3D object detection by jointly estimating multi‑class 6‑DoF poses and class‑specific parametric attributes directly from occlusion‑affected point cloud data. This formulation enables not only geometric localization but also the estimation of task‑relevant properties for parametric objects, such as a gripper's opening, where the 3D model is adjusted according to simple, predefined rules. The architecture employs modular, class‑specific heads, making it straightforward to extend to novel object types without re‑designing the pipeline. We validate PIRATR on an automated forklift platform, focusing on three structurally and functionally diverse categories: crane grippers, loading platforms, and pallets. Trained entirely in a synthetic environment, PIRATR generalizes effectively to real outdoor LiDAR scans, achieving a detection mAP of 0.919 without additional fine‑tuning. PIRATR establishes a new paradigm of pose‑aware, parameterized perception. This bridges the gap between low‑level geometric reasoning and actionable world models, paving the way for scalable, simulation‑trained perception systems that can be deployed in dynamic robotic environments. Code available at https://github.com/swingaxe/piratr.
Authors:Panagiotis Sapoutzoglou, Orestis Vaggelis, Athina Zacharia, Evangelos Sartinas, Maria Pateraki
Abstract:
We introduce IndustryShapes, a new RGB‑D benchmark dataset of industrial tools and components, designed for both instance‑level and novel object 6D pose estimation approaches. The dataset provides a realistic and application‑relevant testbed for benchmarking these methods in the context of industrial robotics bridging the gap between lab‑based research and deployment in real‑world manufacturing scenarios. Unlike many previous datasets that focus on household or consumer products or use synthetic, clean tabletop datasets, or objects captured solely in controlled lab environments, IndustryShapes introduces five new object types with challenging properties, also captured in realistic industrial assembly settings. The dataset has diverse complexity, from simple to more challenging scenes, with single and multiple objects, including scenes with multiple instances of the same object and it is organized in two parts: the classic set and the extended set. The classic set includes a total of 4,6k images and 6k annotated poses. The extended set introduces additional data modalities to support the evaluation of model‑free and sequence‑based approaches. To the best of our knowledge, IndustryShapes is the first dataset to offer RGB‑D static onboarding sequences. We further evaluate the dataset on a representative set of state‑of‑the art methods for instance‑based and novel object 6D pose estimation, including also object detection, segmentation, showing that there is room for improvement in this domain. The dataset page can be found in https://pose‑lab.github.io/IndustryShapes.
Authors:Yue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng, Hongyu Liu, Jiayi Guo, Kunyu Feng, Yuxuan Xue, Zixiang Zhao, Konrad Schindler, Qifeng Chen, Linfeng Zhang
Abstract:
Video motion transfer aims to synthesize videos by generating visual content according to a text prompt while transferring the motion pattern observed in a reference video. Recent methods predominantly use the Diffusion Transformer (DiT) architecture. To achieve satisfactory runtime, several methods attempt to accelerate the computations in the DiT, but fail to address structural sources of inefficiency. In this work, we identify and remove two types of computational redundancy in earlier work: motion redundancy arises because the generic DiT architecture does not reflect the fact that frame‑to‑frame motion is small and smooth; gradient redundancy occurs if one ignores that gradients change slowly along the diffusion trajectory. To mitigate motion redundancy, we mask the corresponding attention layers to a local neighborhood such that interaction weights are not computed unnecessarily distant image regions. To exploit gradient redundancy, we design an optimization scheme that reuses gradients from previous diffusion steps and skips unwarranted gradient computations. On average, FastVMT achieves a 3.43x speedup without degrading the visual fidelity or the temporal consistency of the generated videos.
Authors:Yayuan Li, Ze Peng, Jian Zhang, Jintao Guo, Yue Duan, Yinghuan Shi
Abstract:
Model merging combines multiple fine‑tuned models into a single model by adding their weight updates, providing a lightweight alternative to retraining. Existing methods primarily target resolving conflicts between task updates, leaving the failure mode of over‑counting shared knowledge unaddressed. We show that when tasks share aligned spectral directions (i.e., overlapping singular vectors), a simple linear combination repeatedly accumulates these directions, inflating the singular values and biasing the merged model toward shared subspaces. To mitigate this issue, we propose Singular Value Calibration (SVC), a training‑free and data‑free post‑processing method that quantifies subspace overlap and rescales inflated singular values to restore a balanced spectrum. Across vision and language benchmarks, SVC consistently improves strong merging baselines and achieves state‑of‑the‑art performance. Furthermore, by modifying only the singular values, SVC improves the performance of Task Arithmetic by 13.0%. Code is available at https://github.com/lyymuwu/SVC.
Authors:Youngwoo Shin, Jiwan Hur, Junmo Kim
Abstract:
Visual autoregressive (VAR) models generate images through next‑scale prediction, naturally achieving coarse‑to‑fine, fast, high‑fidelity synthesis mirroring human perception. In practice, this hierarchy can drift at inference time, as limited capacity and accumulated error cause the model to deviate from its coarse‑to‑fine nature. We revisit this limitation from an information‑theoretic perspective and deduce that ensuring each scale contributes high‑frequency content not explained by earlier scales mitigates the train‑inference discrepancy. With this insight, we propose Scaled Spatial Guidance (SSG), training‑free, inference‑time guidance that steers generation toward the intended hierarchy while maintaining global coherence. SSG emphasizes target high‑frequency signals, defined as the semantic residual, isolated from a coarser prior. To obtain this prior, we leverage a principled frequency‑domain procedure, Discrete Spatial Enhancement (DSE), which is devised to sharpen and better isolate the semantic residual through frequency‑aware construction. SSG applies broadly across VAR models leveraging discrete visual tokens, regardless of tokenization design or conditioning modality. Experiments demonstrate SSG yields consistent gains in fidelity and diversity while preserving low latency, revealing untapped efficiency in coarse‑to‑fine image generation. Code is available at https://github.com/Youngwoo‑git/SSG.
Authors:Peihao Wu, Yongxiang Yao, Yi Wan, Wenfei Zhang, Ruipeng Zhao, Jiayuan Li, Yongjun Zhang
Abstract:
Synthetic Aperture Radar (SAR) and optical imagery provide complementary strengths that constitute the critical foundation for transcending single‑modality constraints and facilitating cross‑modal collaborative processing and intelligent interpretation. However, existing benchmark datasets often suffer from limitations such as single spatial resolution, insufficient data scale, and low alignment accuracy, making them inadequate for supporting the training and generalization of multi‑scale foundation models. To address these challenges, we introduce SOMA‑1M (SAR‑Optical Multi‑resolution Alignment), a pixel‑level precisely aligned dataset containing over 1.3 million pairs of georeferenced images with a specification of 512 x 512 pixels. This dataset integrates imagery from Sentinel‑1, PIESAT‑1, Capella Space, and Google Earth, achieving global multi‑scale coverage from 0.5 m to 10 m. It encompasses 12 typical land cover categories, effectively ensuring scene diversity and complexity. To address multimodal projection deformation and massive data registration, we designed a rigorous coarse‑to‑fine image matching framework ensuring pixel‑level alignment. Based on this dataset, we established comprehensive evaluation benchmarks for four hierarchical vision tasks, including image matching, image fusion, SAR‑assisted cloud removal, and cross‑modal translation, involving over 30 mainstream algorithms. Experimental results demonstrate that supervised training on SOMA‑1M significantly enhances performance across all tasks. Notably, multimodal remote sensing image (MRSI) matching performance achieves current state‑of‑the‑art (SOTA) levels. SOMA‑1M serves as a foundational resource for robust multimodal algorithms and remote sensing foundation models. The dataset will be released publicly at: https://github.com/PeihaoWu/SOMA‑1M.
Authors:Dekang Qi, Shuang Zeng, Xinyuan Chang, Feng Xiong, Shichao Xie, Xiaolong Wu, Mu Xu
Abstract:
Visual Language Navigation (VLN) is one of the fundamental capabilities for embodied intelligence and a critical challenge that urgently needs to be addressed. However, existing methods are still unsatisfactory in terms of both success rate (SR) and generalization: Supervised Fine‑Tuning (SFT) approaches typically achieve higher SR, while Training‑Free (TF) approaches often generalize better, but it is difficult to obtain both simultaneously. To this end, we propose a Memory‑Execute‑Review framework. It consists of three parts: a hierarchical memory module for providing information support, an execute module for routine decision‑making and actions, and a review module for handling abnormal situations and correcting behavior. We validated the effectiveness of this framework on the Object Goal Navigation task. Across 4 datasets, our average SR achieved absolute improvements of 7% and 5% compared to all baseline methods under TF and Zero‑Shot (ZS) settings, respectively. On the most commonly used HM3D_v0.1 and the more challenging open vocabulary dataset HM3D_OVON, the SR improved by 8% and 6%, under ZS settings. Furthermore, on the MP3D and HM3D_OVON datasets, our method not only outperformed all TF methods but also surpassed all SFT methods, achieving comprehensive leadership in both SR (5% and 2%) and generalization. Additionally, we deployed the MerNav model on the humanoid robot and conducted experiments in the real world. The project address is: https://qidekang.github.io/MerNav.github.io/
Authors:Budhaditya Mukhopadhyay, Chirag Mandal, Pavan Tummala, Naghmeh Mahmoodian, Andreas Nürnberger, Soumick Chatterjee
Abstract:
Liver tumour ablation presents a significant clinical challenge: whilst tumours are clearly visible on pre‑operative MRI, they are often effectively invisible on intra‑operative CT due to minimal contrast between pathological and healthy tissue. This work investigates the feasibility of cross‑modality weak supervision for scenarios where pathology is visible in one modality (MRI) but absent in another (CT). We present a hybrid registration‑segmentation framework that combines MSCGUNet for inter‑modal image registration with a UNet‑based segmentation module, enabling registration‑assisted pseudo‑label generation for CT images. Our evaluation on the CHAOS dataset demonstrates that the pipeline can successfully register and segment healthy liver anatomy, achieving a Dice score of 0.72. However, when applied to clinical data containing tumours, performance degrades substantially (Dice score of 0.16), revealing the fundamental limitations of current registration methods when the target pathology lacks corresponding visual features in the target modality. We analyse the "domain gap" and "feature absence" problems, demonstrating that whilst spatial propagation of labels via registration is feasible for visible structures, segmenting truly invisible pathology remains an open challenge. Our findings highlight that registration‑based label transfer cannot compensate for the absence of discriminative features in the target modality, providing important insights for future research in cross‑modality medical image analysis. Code an weights are available at: https://github.com/BudhaTronix/Weakly‑Supervised‑Tumour‑Detection
Authors:Chang Zou, Changlin Li, Yang Li, Patrol Li, Jianbing Wu, Xiao He, Songtao Liu, Zhao Zhong, Kailin Huang, Linfeng Zhang
Abstract:
While diffusion models have achieved great success in the field of video generation, this progress is accompanied by a rapidly escalating computational burden. Among the existing acceleration methods, Feature Caching is popular due to its training‑free property and considerable speedup performance, but it inevitably faces semantic and detail drop with further compression. Another widely adopted method, training‑aware step‑distillation, though successful in image generation, also faces drastic degradation in video generation with a few steps. Furthermore, the quality loss becomes more severe when simply applying training‑free feature caching to the step‑distilled models, due to the sparser sampling steps. This paper novelly introduces a distillation‑compatible learnable feature caching mechanism for the first time. We employ a lightweight learnable neural predictor instead of traditional training‑free heuristics for diffusion models, enabling a more accurate capture of the high‑dimensional feature evolution process. Furthermore, we explore the challenges of highly compressed distillation on large‑scale video models and propose a conservative Restricted MeanFlow approach to achieve more stable and lossless distillation. By undertaking these initiatives, we further push the acceleration boundaries to 11.8× while preserving generation quality. Extensive experiments demonstrate the effectiveness of our method. Code has been made publicly available: https://github.com/Tencent‑Hunyuan/DisCa
Authors:Ngoc Doan-Minh Huynh, Duong Nguyen-Ngoc Tran, Long Hoang Pham, Tai Huu-Phuong Tran, Hyung-Joon Jeon, Huy-Hung Nguyen, Duong Khac Vu, Hyung-Min Jeon, Son Hong Phan, Quoc Pham-Nam Ho, Chi Dai Tran, Trinh Le Ba Khanh, Jae Wook Jeon
Abstract:
Global warming has intensified the frequency and severity of extreme weather events, which degrade CCTV signal and video quality while disrupting traffic flow, thereby increasing traffic accident rates. Existing datasets, often limited to light haze, rain, and snow, fail to capture extreme weather conditions. To address this gap, this study introduces the Traffic Surveillance Benchmark for Occluded vehicles under various Weather conditions (TSBOW), a comprehensive dataset designed to enhance occluded vehicle detection across diverse annual weather scenarios. Comprising over 32 hours of real‑world traffic data from densely populated urban areas, TSBOW includes more than 48,000 manually annotated and 3.2 million semi‑labeled frames; bounding boxes spanning eight traffic participant classes from large vehicles to micromobility devices and pedestrians. We establish an object detection benchmark for TSBOW, highlighting challenges posed by occlusions and adverse weather. With its varied road types, scales, and viewpoints, TSBOW serves as a critical resource for advancing Intelligent Transportation Systems. Our findings underscore the potential of CCTV‑based traffic monitoring, pave the way for new research and applications. The TSBOW dataset is publicly available at: https://github.com/SKKUAutoLab/TSBOW.
Authors:Zolnamar Dorjsembe, Hung-Yi Chen, Furen Xiao, Hsing-Kuo Pao
Abstract:
MRI provides superior soft tissue contrast without ionizing radiation; however, the absence of electron density information limits its direct use for dose calculation. As a result, current radiotherapy workflows rely on combined MRI and CT acquisitions, increasing registration uncertainty and procedural complexity. Synthetic CT generation enables MRI only planning but remains challenging due to nonlinear MRI‑CT relationships and anatomical variability. We propose Parallel Swin Transformer‑Enhanced Med2Transformer, a 3D architecture that integrates convolutional encoding with dual Swin Transformer branches to model both local anatomical detail and long‑range contextual dependencies. Multi‑scale shifted window attention with hierarchical feature aggregation improves anatomical fidelity. Experiments on public and clinical datasets demonstrate higher image similarity and improved geometric accuracy compared with baseline methods. Dosimetric evaluation shows clinically acceptable performance, with a mean target dose error of 1.69%. Code is available at: https://github.com/mobaidoctor/med2transformer.
Authors:Zhuokun Chen, Jianfei Cai, Bohan Zhuang
Abstract:
Generating long‑form content, such as minute‑long videos and extended texts, is increasingly important for modern generative models. Block diffusion improves inference efficiency via KV caching and block‑wise causal inference and has been widely adopted in diffusion language models and video generation. However, in long‑context settings, block diffusion still incurs substantial overhead from repeatedly computing attention over a growing KV cache. We identify an underexplored property of block diffusion: cross‑step redundancy of attention within a block. Our analysis shows that attention outputs from tokens outside the current block remain largely stable across diffusion steps, while block‑internal attention varies significantly. Based on this observation, we propose FlashBlock, a cached block‑external attention mechanism that reuses stable attention output, reducing attention computation and KV cache access without modifying the diffusion process. Moreover, FlashBlock is orthogonal to sparse attention and can be combined as a complementary residual reuse strategy, substantially improving model accuracy under aggressive sparsification. Experiments on diffusion language models and video generation demonstrate up to 1.44× higher token throughput and up to 1.6× reduction in attention time, with negligible impact on generation quality. Project page: https://caesarhhh.github.io/FlashBlock/.
Authors:Changhoon Song, Teng Yuan Chang, Youngjoon Hong
Abstract:
Accurate forecasting of extreme weather events such as heavy rainfall or storms is critical for risk management and disaster mitigation. Although high‑resolution radar observations have spurred extensive research on nowcasting models, precipitation nowcasting remains particularly challenging due to pronounced spatial locality, intricate fine‑scale rainfall structures, and variability in forecasting horizons. While recent diffusion‑based generative ensembles show promising results, they are computationally expensive and unsuitable for real‑time applications. In contrast, deterministic models are computationally efficient but remain biased toward normal rainfall. Furthermore, the benchmark datasets commonly used in prior studies are themselves skewed‑‑either dominated by ordinary rainfall events or restricted to extreme rainfall episodes‑‑thereby hindering general applicability in real‑world settings. In this paper, we propose exPreCast, an efficient deterministic framework for generating finely detailed radar forecasts, and introduce a newly constructed balanced radar dataset from the Korea Meteorological Administration (KMA), which encompasses both ordinary precipitation and extreme events. Our model integrates local spatiotemporal attention, a texture‑preserving cubic dual upsampling decoder, and a temporal extractor to flexibly adjust forecasting horizons. Experiments on established benchmarks (SEVIR and MeteoNet) as well as on the balanced KMA dataset demonstrate that our approach achieves state‑of‑the‑art performance, delivering accurate and reliable nowcasts across both normal and extreme rainfall regimes.
Authors:Grzegorz Wilczyński, Rafał Tobiasz, Paweł Gora, Marcin Mazur, Przemysław Spurek
Abstract:
Recent advances in neural rendering, particularly 3D Gaussian Splatting (3DGS), have enabled real‑time rendering of complex scenes. However, standard 3DGS relies on spherical harmonics, which often struggle to accurately capture high‑frequency view‑dependent effects such as sharp reflections and transparency. While hybrid approaches like Viewing Direction Gaussian Splatting (VDGS) mitigate this limitation using classical Multi‑Layer Perceptrons (MLPs), they remain limited by the expressivity of classical networks in low‑parameter regimes. In this paper, we introduce QuantumGS, a novel hybrid framework that integrates Variational Quantum Circuits (VQC) into the Gaussian Splatting pipeline. We propose a unique encoding strategy that maps the viewing direction directly onto the Bloch sphere, leveraging the natural geometry of qubits to represent 3D directional data. By replacing classical color‑modulating networks with quantum circuits generated via a hypernetwork or conditioning mechanism, we achieve higher expressivity and better generalization. Source code is available in the supplementary material. Code is available at https://github.com/gwilczynski95/QuantumGS
Authors:Jiahao Zhan, Zizhang Li, Hong-Xing Yu, Jiajun Wu
Abstract:
We introduce PerpetualWonder, a hybrid generative simulator that enables long‑horizon, action‑conditioned 4D scene generation from a single image. Current works fail at this task because their physical state is decoupled from their visual representation, which prevents generative refinements to update the underlying physics for subsequent interactions. PerpetualWonder solves this by introducing the first true closed‑loop system. It features a novel unified representation that creates a bidirectional link between the physical state and visual primitives, allowing generative refinements to correct both the dynamics and appearance. It also introduces a robust update mechanism that gathers supervision from multiple viewpoints to resolve optimization ambiguity. Experiments demonstrate that from a single image, PerpetualWonder can successfully simulate complex, multi‑step interactions from long‑horizon actions, maintaining physical plausibility and visual consistency.
Authors:Yi Gu, Yukang Gao, Yangchen Zhou, Xingyu Chen, Yixiao Feng, Mingle Zhao, Yunyang Mo, Zhaorui Wang, Lixin Xu, Renjing Xu
Abstract:
Pose and motion priors play a crucial role in humanoid robotics. Although such priors have been widely studied in human motion recovery (HMR) domain with a range of models, their adoption for humanoid robots remains limited, largely due to the scarcity of high‑quality humanoid motion data. In this work, we introduce Pose Distance Fields for Humanoid Robots (PDF‑HR), a lightweight prior that represents the robot pose distribution as a continuous and differentiable manifold. Given an arbitrary pose, PDF‑HR predicts its distance to a large corpus of retargeted robot poses, yielding a smooth measure of pose plausibility that is well suited for optimization and control. PDF‑HR can be integrated as a reward shaping term, a regularizer, or a standalone plausibility scorer across diverse pipelines. We evaluate PDF‑HR on various humanoid tasks, including single‑trajectory motion tracking, general motion tracking, style‑based motion mimicry, and general motion retargeting. Experiments show that this plug‑and‑play prior consistently and substantially strengthens strong baselines. Code and models will be released.
Authors:Ronghuan Wu, Wanchao Su, Kede Ma, Jing Liao, Rafał K. Mantiuk
Abstract:
High‑dynamic‑range (HDR) formats and displays are becoming increasingly prevalent, yet state‑of‑the‑art image generators (e.g., Stable Diffusion and FLUX) typically remain limited to low‑dynamic‑range (LDR) output due to the lack of large‑scale HDR training data. In this work, we show that existing pretrained diffusion models can be easily adapted to HDR generation without retraining from scratch. A key challenge is that HDR images are natively represented in linear RGB, whose intensity and color statistics differ substantially from those of sRGB‑encoded LDR images. This gap, however, can be effectively bridged by converting HDR inputs into perceptually uniform encodings (e.g., using PU21 or PQ). Empirically, we find that LDR‑pretrained variational autoencoders (VAEs) reconstruct PU21‑encoded HDR inputs with fidelity comparable to LDR data, whereas linear RGB inputs cause severe degradations. Motivated by this finding, we describe an efficient adaptation strategy that freezes the VAE and finetunes only the denoiser via low‑rank adaptation in a perceptually uniform space. This results in a unified computational method that supports both text‑to‑HDR synthesis and single‑image RAW‑to‑HDR reconstruction. Experiments demonstrate that our perceptually encoded adaptation consistently improves perceptual fidelity, text‑image alignment, and effective dynamic range, relative to previous techniques.
Authors:Chengtao Lv, Yumeng Shi, Yushi Huang, Ruihao Gong, Shen Ren, Wenya Wang
Abstract:
Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention remains a primary bottleneck for efficient deployment. While existing sparse attention solutions have shown promise on bidirectional models, we identify that applying these solutions to AR models leads to considerable performance degradation for two reasons: isolated consideration of chunk generation and insufficient utilization of past informative context. Motivated by these observations, we propose \textscLight Forcing, the first sparse attention solution tailored for AR video generation models. It incorporates a Chunk‑Aware Growth mechanism to quantitatively estimate the contribution of each chunk, which determines their sparsity allocation. This progressive sparsity increase strategy enables the current chunk to inherit prior knowledge in earlier chunks during generation. Additionally, we introduce a Hierarchical Sparse Attention to capture informative historical and local context in a coarse‑to‑fine manner. Such two‑level mask selection strategy (\ie, frame and block level) can adaptively handle diverse attention patterns. Extensive experiments demonstrate that our method outperforms existing sparse attention in quality (\eg, 84.5 on VBench) and efficiency (\eg, 1.2~1.3× end‑to‑end speedup). Combined with FP8 quantization and LightVAE, \textscLight Forcing further achieves a 2.3× speedup and 19.7\,FPS on an RTX~5090 GPU. Code will be released at \hrefhttps://github.com/chengtao‑lv/LightForcinghttps://github.com/chengtao‑lv/LightForcing.
Authors:Mingyang Deng, He Li, Tianhong Li, Yilun Du, Kaiming He
Abstract:
Generative modeling can be formulated as learning a mapping f such that its pushforward distribution matches the data distribution. The pushforward behavior can be carried out iteratively at inference time, for example in diffusion and flow‑based models. In this paper, we propose a new paradigm called Drifting Models, which evolve the pushforward distribution during training and naturally admit one‑step inference. We introduce a drifting field that governs the sample movement and achieves equilibrium when the distributions match. This leads to a training objective that allows the neural network optimizer to evolve the distribution. In experiments, our one‑step generator achieves state‑of‑the‑art results on ImageNet at 256 x 256 resolution, with an FID of 1.54 in latent space and 1.61 in pixel space. We hope that our work opens up new opportunities for high‑quality one‑step generation.
Authors:Buddhi Wijenayake, Nichula Wasalathilake, Roshan Godaliyadda, Vijitha Herath, Parakrama Ekanayake, Vishal M. Patel
Abstract:
Long‑tailed class imbalance remains a fundamental obstacle in semantic segmentation of high‑resolution remote‑sensing imagery, where dominant classes shape learned representations and rare classes are systematically under‑segmented. This challenge becomes more acute in cross‑domain settings such as LoveDA, which exhibits an explicit Urban/Rural split with substantial appearance differences and inconsistent class‑frequency statistics across domains. We propose a prompt‑controlled diffusion augmentation framework that generates paired label‑image samples with explicit control over semantic composition and domain, enabling targeted enrichment of underrepresented classes rather than indiscriminate dataset expansion. A domain‑aware, masked, ratio‑conditioned discrete diffusion model first synthesizes layouts that satisfy class‑ratio targets while preserving realistic spatial co‑occurrence, and a ControlNet‑guided diffusion model then renders photorealistic, domain‑consistent images from these layouts. When mixed with real data, the resulting synthetic pairs improve multiple segmentation backbones, especially on minority classes and under domain shift, showing that better downstream segmentation comes from adding the right samples in the right proportions.
Source codes, pretrained models, and synthetic datasets are available at \hrefhttps://buddhi19.github.io/SyntheticGen\textttbuddhi19.github.io/SyntheticGen.
Authors:Samet Hicsonmez, Jose Sosa, Dan Pineau, Inder Pal Singh, Arunkumar Rathinam, Abd El Rahman Shabayek, Djamila Aouada
Abstract:
Vision Language Models (VLMs) have demonstrated remarkable performance in open‑world zero‑shot visual recognition. However, their potential in space‑related applications remains largely unexplored. In the space domain, accurate manual annotation is particularly challenging due to factors such as low visibility, illumination variations, and object blending with planetary backgrounds. Developing methods that can detect and segment spacecraft and orbital targets without requiring extensive manual labeling is therefore of critical importance. In this work, we propose an annotation‑free detection and segmentation pipeline for space targets using VLMs. Our approach begins by automatically generating pseudo‑labels for a small subset of unlabeled real data with a pre‑trained VLM. These pseudo‑labels are then leveraged in a teacher‑student label distillation framework to train lightweight models. Despite the inherent noise in the pseudo‑labels, the distillation process leads to substantial performance gains over direct zero‑shot VLM inference. Experimental evaluations on the SPARK‑2024, SPEED+, and TANGO datasets on segmentation tasks demonstrate consistent improvements in average precision (AP) by up to 10 points. Code and models are available at https://github.com/giddyyupp/annotation‑free‑spacecraft‑segmentation.
Authors:Sijia Chen, Lijuan Ma, Yanqiu Yu, En Yu, Liman Liu, Wenbing Tao
Abstract:
Referring Multi‑Object Tracking (RMOT) aims to track specific targets based on language descriptions and is vital for interactive AI systems such as robotics and autonomous driving. However, existing RMOT models rely solely on 2D RGB data, making it challenging to accurately detect and associate targets characterized by complex spatial semantics (e.g., ``the person closest to the camera'') and to maintain reliable identities under severe occlusion, due to the absence of explicit 3D spatial information. In this work, we propose a novel task, RGBD Referring Multi‑Object Tracking (DRMOT), which explicitly requires models to fuse RGB, Depth (D), and Language (L) modalities to achieve 3D‑aware tracking. To advance research on the DRMOT task, we construct a tailored RGBD referring multi‑object tracking dataset, named DRSet, designed to evaluate models' spatial‑semantic grounding and tracking capabilities. Specifically, DRSet contains RGB images and depth maps from 187 scenes, along with 240 language descriptions, among which 56 descriptions incorporate depth‑related information. Furthermore, we propose DRTrack, a MLLM‑guided depth‑referring tracking framework. DRTrack performs depth‑aware target grounding from joint RGB‑D‑L inputs and enforces robust trajectory association by incorporating depth cues. Extensive experiments on the DRSet dataset demonstrate the effectiveness of our framework.
Authors:Haokui Zhang, Congyang Ou, Dawei Yan, Peng Wang, Qingsen Yan, Yu Zhang, Ying Li, Rong Xiao
Abstract:
Recently, reducing redundant visual tokens in vision‑language models (VLMs) to accelerate VLM inference has emerged as a hot topic. However, most existing methods rely on heuristics constructed based on inter‑visual‑token similarity or cross‑modal visual‑text similarity, which gives rise to certain limitations in compression performance and practical deployment. In contrast, we propose TRIO from the perspective of inference objectives, which transforms visual token compression into preserving output result invariance and selects tokens primarily by their importance to this goal. Specifically, vision tokens are reordered with the guidance of token‑level gradient saliency generated by our designed layer‑local proxy loss, a coarse constraint from the current layer to the final result. Then the most valuable vision tokens are selected following the non‑maximum suppression (NMS) principle.The proposed TRIO is training‑free and compatible with FlashAttention, friendly to practical application and deployment. It can be deployed independently as an encoder‑free method, or combined with encoder compression approaches like VisionZip for use as an encoder‑involved method. On LLaVA‑Next‑7B, TRIO retains just 11.1% of visual tokens but maintains 97.2% of the original performance, with a 2.75× prefill speedup, 2.14× inference speedup, 6.22× lower FLOPs, and 6.05× reduced KV Cache overhead.Our code is available at https://github.com/ocy1/TRIO.
Authors:Cem Eteke, Enzo Tartaglione
Abstract:
3D Gaussian Splatting (3DGS) revolutionized novel view rendering. Instead of inferring from dense spatial points, as implicit representations do, 3DGS uses sparse Gaussians. This enables real‑time performance but increases space requirements, hindering rate‑constrained applications. 3DGS compression emerged as a field aimed at alleviating this issue. While impressive progress has been made, at low rates, compression introduces artifacts that degrade visual quality significantly. We introduce NiFi, a method for extreme 3DGS compression through restoration via artifact‑aware, diffusion‑based one‑step distillation. We show that our method achieves state‑of‑the‑art perceptual quality at extremely low rates, down to 0.1 MB, and towards 1000x rate improvement over 3DGS at comparable perceptual performance. Code is available at: https://github.com/ceteke/nifi
Authors:Timothy Schaumlöffel, Arthur Aubret, Gemma Roig, Jochen Triesch
Abstract:
Humans acquire semantic object representations from egocentric visual streams with minimal supervision, but the underlying mechanisms remain unclear. Importantly, the visual system only processes the center of its field of view with high resolution and it learns similar representations for visual inputs occurring close in time. This emphasizes slowly changing information around gaze locations. This study investigates the role of central vision and slowness learning in the formation of semantic object representations from human‑like visual experience. We simulate five months of human‑like visual experience using the Ego4D dataset and a state‑of‑the‑art gaze prediction model. We extract image crops around predicted gaze locations to train a time‑contrastive Self‑Supervised Learning model. Our results show that exploiting temporal slowness when learning from central visual field experience improves the encoding of different facets of object semantics. Specifically, focusing on central vision strengthens the extraction of foreground object features, while considering temporal slowness, especially in conjunction with eye movements, allows the model to encode broader semantic information about objects. These findings provide new insights into the mechanisms by which humans may develop semantic object representations from natural visual experience. Our code will be made public upon acceptance. Code is available at https://github.com/t9s9/central‑vision‑ssl.
Authors:Tianming Liang, Qirui Du, Jian-Fang Hu, Haichao Jiang, Zicheng Lin, Wei-Shi Zheng
Abstract:
Segmentation based on language has been a popular topic in computer vision. While recent advances in multimodal large language models (MLLMs) have endowed segmentation systems with reasoning capabilities, these efforts remain confined by the frozen internal knowledge of MLLMs, which limits their potential for real‑world scenarios that involve up‑to‑date information or domain‑specific concepts. In this work, we propose Seg‑ReSearch, a novel segmentation paradigm that overcomes the knowledge bottleneck of existing approaches. By enabling interleaved reasoning and external search, Seg‑ReSearch empowers segmentation systems to handle dynamic, open‑world queries that extend beyond the frozen knowledge of MLLMs. To effectively train this capability, we introduce a hierarchical reward design that harmonizes initial guidance with progressive incentives, mitigating the dilemma between sparse outcome signals and rigid step‑wise supervision. For evaluation, we construct OK‑VOS, a challenging benchmark that explicitly requires outside knowledge for video object segmentation. Experiments on OK‑VOS and two existing reasoning segmentation benchmarks demonstrate that our Seg‑ReSearch improves state‑of‑the‑art approaches by a substantial margin. Code and data will be released at https://github.com/iSEE‑Laboratory/Seg‑ReSearch.
Authors:Aavash Chhetri, Bibek Niroula, Pratik Shrestha, Yash Raj Shrestha, Lesley A Anderson, Prashnna K Gyawali, Loris Bazzani, Binod Bhattarai
Abstract:
Federated learning (FL) enables collaborative model training across decentralized medical institutions while preserving data privacy. However, medical FL benchmarks remain scarce, with existing efforts focusing mainly on unimodal or bimodal modalities and a limited range of medical tasks. This gap underscores the need for standardized evaluation to advance systematic understanding in medical MultiModal FL (MMFL). To this end, we introduce Med‑MMFL, the first comprehensive MMFL benchmark for the medical domain, encompassing diverse modalities, tasks, and federation scenarios. Our benchmark evaluates six representative state‑of‑the‑art FL algorithms, covering different aggregation strategies, loss formulations, and regularization techniques. It spans datasets with 2 to 4 modalities, comprising a total of 10 unique medical modalities, including text, pathology images, ECG, X‑ray, radiology reports, and multiple MRI sequences. Experiments are conducted across naturally federated, synthetic IID, and synthetic non‑IID settings to simulate real‑world heterogeneity. We assess segmentation, classification, modality alignment (retrieval), and VQA tasks. To support reproducibility and fair comparison of future multimodal federated learning (MMFL) methods under realistic medical settings, we release the complete benchmark implementation, including data processing and partitioning pipelines, at https://github.com/bhattarailab/Med‑MMFL‑Benchmark .
Authors:Jue Gong, Zihan Zhou, Jingkai Wang, Shu Li, Libo Liu, Jianliang Lan, Yulun Zhang
Abstract:
Existing methods for restoring degraded human‑centric images often struggle with insufficient fidelity, particularly in human body restoration (HBR). Recent diffusion‑based restoration methods commonly adapt pre‑trained text‑to‑image diffusion models, where the variational autoencoder (VAE) can significantly bottleneck restoration fidelity. We propose LCUDiff, a stable one‑step framework that upgrades a pre‑trained latent diffusion model from the 4‑channel latent space to the 16‑channel latent space. For VAE fine‑tuning, channel splitting distillation (CSD) is used to keep the first four channels aligned with pre‑trained priors while allocating the additional channels to effectively encode high‑frequency details. We further design prior‑preserving adaptation (PPA) to smoothly bridge the mismatch between 4‑channel diffusion backbones and the higher‑dimensional 16‑channel latent. In addition, we propose a decoder router (DeR) for per‑sample decoder routing using restoration‑quality score annotations, which improves visual quality across diverse conditions. Experiments on synthetic and real‑world datasets show competitive results with higher fidelity and fewer artifacts under mild degradations, while preserving one‑step efficiency. The code and model will be at https://github.com/gobunu/LCUDiff.
Authors:Yixin Zhu, Long Lv, Pingping Zhang, Xuehu Liu, Tongdan Tang, Feng Tian, Weibing Sun, Huchuan Lu
Abstract:
Multi‑Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently, some MMIF methods incorporate frequency domain information to enhance spatial features. However, these methods typically rely on simple serial or parallel spatial‑frequency fusion without interaction. In this paper, we propose a novel Interactive Spatial‑Frequency Fusion Mamba (ISFM) framework for MMIF. Specifically, we begin with a Modality‑Specific Extractor (MSE) to extract features from different modalities. It models long‑range dependencies across the image with linear computational complexity. To effectively leverage frequency information, we then propose a Multi‑scale Frequency Fusion (MFF). It adaptively integrates low‑frequency and high‑frequency components across multiple scales, enabling robust representations of frequency features. More importantly, we further propose an Interactive Spatial‑Frequency Fusion (ISF). It incorporates frequency features to guide spatial features across modalities, enhancing complementary representations. Extensive experiments are conducted on six MMIF datasets. The experimental results demonstrate that our ISFM can achieve better performances than other state‑of‑the‑art methods. The source code is available at https://github.com/Namn23/ISFM.
Authors:Zekun Li, Ning Wang, Tongxin Bai, Changwang Mei, Peisong Wang, Shuang Qiu, Jian Cheng
Abstract:
Visual AutoRegressive (VAR) modeling has garnered significant attention for its innovative next‑scale prediction paradigm. However, mainstream VAR paradigms attend to all tokens across historical scales at each autoregressive step. As the next scale resolution grows, the computational complexity of attention increases quartically with resolution, causing substantial latency. Prior accelerations often skip high‑resolution scales, which speeds up inference but discards high‑frequency details and harms image quality. To address these problems, we present SparVAR, a training‑free acceleration framework that exploits three properties of VAR attention: (i) strong attention sinks, (ii) cross‑scale activation similarity, and (iii) pronounced locality. Specifically, we dynamically predict the sparse attention pattern of later high‑resolution scales from a sparse decision scale, and construct scale self‑similar sparse attention via an efficient index‑mapping mechanism, enabling high‑efficiency sparse attention computation at large scales. Furthermore, we propose cross‑scale local sparse attention and implement an efficient block‑wise sparse kernel, which achieves \mathbf> 5× faster forward speed than FlashAttention. Extensive experiments demonstrate that the proposed SparVAR can reduce the generation time of an 8B model producing 1024×1024 high‑resolution images to the 1s, without skipping the last scales. Compared with the VAR baseline accelerated by FlashAttention, our method achieves a \mathbf1.57× speed‑up while preserving almost all high‑frequency details. When combined with existing scale‑skipping strategies, SparVAR attains up to a \mathbf2.28× acceleration, while maintaining competitive visual generation quality. Code is available at \hrefhttps://github.com/CAS‑CLab/SparVARSparVAR.
Authors:Jaehyun Kwak, Nam Cao, Boryeong Cho, Segyu Lee, Sumyeong Ahn, Se-Young Yun
Abstract:
Adversarial attacks against Large Vision‑Language Models (LVLMs) are crucial for exposing safety vulnerabilities in modern multimodal systems. Recent attacks based on input transformations, such as random cropping, suggest that spatially localized perturbations can be more effective than global image manipulation. However, randomly cropping the entire image is inherently stochastic and fails to use the limited per‑pixel perturbation budget efficiently. We make two key observations: (i) regional attention scores are positively correlated with adversarial loss sensitivity, and (ii) attacking high‑attention regions induces a structured redistribution of attention toward subsequent salient regions. Based on these findings, we propose Stage‑wise Attention‑Guided Attack (SAGA), an attention‑guided framework that progressively concentrates perturbations on high‑attention regions. SAGA enables more efficient use of constrained perturbation budgets, producing highly imperceptible adversarial examples while consistently achieving state‑of‑the‑art attack success rates across ten LVLMs. The source code is available at https://github.com/jackwaky/SAGA.
Authors:Teng-Fang Hsiao, Bo-Kai Ruan, Yu-Lun Liu, Hong-Han Shuai
Abstract:
3D editing has emerged as a critical research area to provide users with flexible control over 3D assets. While current editing approaches predominantly focus on 3D Gaussian Splatting or multi‑view images, the direct editing of 3D meshes remains underexplored. Prior attempts, such as VoxHammer, rely on voxel‑based representations that suffer from limited resolution and necessitate labor‑intensive 3D mask. To address these limitations, we propose VecSet‑Edit, the first pipeline that leverages the high‑fidelity VecSet Large Reconstruction Model (LRM) as a backbone for mesh editing. Our approach is grounded on a analysis of the spatial properties in VecSet tokens, revealing that token subsets govern distinct geometric regions. Based on this insight, we introduce Mask‑guided Token Seeding and Attention‑aligned Token Gating strategies to precisely localize target regions using only 2D image conditions. Also, considering the difference between VecSet diffusion process versus voxel we design a Drift‑aware Token Pruning to reject geometric outliers during the denoising process. Finally, our Detail‑preserving Texture Baking module ensures that we not only preserve the geometric details of original mesh but also the textural information. More details can be found in our project page: https://github.com/BlueDyee/VecSet‑Edit/tree/main
Authors:Zihan Lou, Jinlong Fan, Sihan Ma, Yuxiang Yang, Jing Zhang
Abstract:
Reconstructing high‑fidelity animatable 3D human avatars from monocular RGB videos remains challenging, particularly in unconstrained in‑the‑wild scenarios where camera parameters and human poses from off‑the‑shelf methods (e.g., COLMAP, HMR2.0) are often inaccurate. Splatting (3DGS) advances demonstrate impressive rendering quality and real‑time performance, they critically depend on precise camera calibration and pose annotations, limiting their applicability in real‑world settings. We present JOintGS, a unified framework that jointly optimizes camera extrinsics, human poses, and 3D Gaussian representations from coarse initialization through a synergistic refinement mechanism. Our key insight is that explicit foreground‑background disentanglement enables mutual reinforcement: static background Gaussians anchor camera estimation via multi‑view consistency; refined cameras improve human body alignment through accurate temporal correspondence; optimized human poses enhance scene reconstruction by removing dynamic artifacts from static constraints. We further introduce a temporal dynamics module to capture fine‑grained pose‑dependent deformations and a residual color field to model illumination variations. Extensive experiments on NeuMan and EMDB datasets demonstrate that JOintGS achieves superior reconstruction quality, with 2.1~dB PSNR improvement over state‑of‑the‑art methods on NeuMan dataset, while maintaining real‑time rendering. Notably, our method shows significantly enhanced robustness to noisy initialization compared to the baseline.Our source code is available at https://github.com/MiliLab/JOintGS.
Authors:Guoqing Ma, Siheng Wang, Zeyu Zhang, Shan Yu, Hao Tang
Abstract:
Large foundation models have shown strong open‑world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robotics. One fundamental challenge is that the models exhibit limited zero‑shot capability, which hampers their ability to generalize effectively to unseen scenarios. In this work, we propose GeneralVLA (Generalizable Vision‑Language‑Action Models with Knowledge‑Guided Trajectory Planning), a hierarchical vision‑language‑action (VLA) model that can be more effective in utilizing the generalization of foundation models, enabling zero‑shot manipulation and automatically generating data for robotics. In particular, we study a class of hierarchical VLA model where the high‑level ASM (Affordance Segmentation Module) is finetuned to perceive image keypoint affordances of the scene; the mid‑level 3DAgent carries out task understanding, skill knowledge, and trajectory planning to produce a 3D path indicating the desired robot end‑effector trajectory. The intermediate 3D path prediction is then served as guidance to the low‑level, 3D‑aware control policy capable of precise manipulation. Compared to alternative approaches, our method requires no real‑world robotic data collection or human demonstration, making it much more scalable to diverse tasks and viewpoints. Empirically, GeneralVLA successfully generates trajectories for 14 tasks, significantly outperforming state‑of‑the‑art methods such as VoxPoser. The generated demonstrations can train more robust behavior cloning policies than training with human demonstrations or from data generated by VoxPoser, Scaling‑up, and Code‑As‑Policies. We believe GeneralVLA can be the scalable method for both generating data for robotics and solving novel tasks in a zero‑shot setting. Code: https://github.com/AIGeeksGroup/GeneralVLA. Website: https://aigeeksgroup.github.io/GeneralVLA.
Authors:Jue Gong, Zihan Zhou, Jingkai Wang, Xiaohong Liu, Yulun Zhang, Xiaokang Yang
Abstract:
Face fill‑light enhancement (FFE) brightens underexposed faces by adding virtual fill light while keeping the original scene illumination and background unchanged. Most face relighting methods aim to reshape overall lighting, which can suppress the input illumination or modify the entire scene, leading to foreground‑background inconsistency and mismatching practical FFE needs. To support scalable learning, we introduce LightYourFace‑160K (LYF‑160K), a large‑scale paired dataset built with a physically consistent renderer that injects a disk‑shaped area fill light controlled by six disentangled factors, producing 160K before‑and‑after pairs. We first pretrain a physics‑aware lighting prompt (PALP) that embeds the 6D parameters into conditioning tokens, using an auxiliary planar‑light reconstruction objective. Building on a pretrained diffusion backbone, we then train a fill‑light diffusion (FiLitDiff), an efficient one‑step model conditioned on physically grounded lighting codes, enabling controllable and high‑fidelity fill lighting at low computational cost. Experiments on held‑out paired sets demonstrate strong perceptual quality and competitive full‑reference metrics, while better preserving background illumination. The dataset and model will be at https://github.com/gobunu/Light‑Up‑Your‑Face.
Authors:Lifan Wu, Ruijie Zhu, Yubo Ai, Tianzhu Zhang
Abstract:
4D generation has made remarkable progress in synthesizing dynamic 3D objects from input text, images, or videos. However, existing methods often represent motion as an implicit deformation field, which limits direct control and editability. To address this issue, we propose SkeletonGaussian, a novel framework for generating editable dynamic 3D Gaussians from monocular video input. Our approach introduces a hierarchical articulated representation that decomposes motion into sparse rigid motion explicitly driven by a skeleton and fine‑grained non‑rigid motion. Concretely, we extract a robust skeleton and drive rigid motion via linear blend skinning, followed by a hexplane‑based refinement for non‑rigid deformations, enhancing interpretability and editability. Experimental results demonstrate that SkeletonGaussian surpasses existing methods in generation quality while enabling intuitive motion editing, establishing a new paradigm for editable 4D generation. Project page: https://wusar.github.io/projects/skeletongaussian/
Authors:Suzeyu Chen, Leheng Li, Ying-Cong Chen
Abstract:
Achieving highly accurate and real‑time 3D occupancy prediction from cameras is a critical requirement for the safe and practical deployment of autonomous vehicles. While this shift to sparse 3D representations solves the encoding bottleneck, it creates a new challenge for the decoder: how to efficiently aggregate information from a sparse, non‑uniformly distributed set of voxel features without resorting to computationally prohibitive dense attention.
In this paper, we propose a novel Prototype‑based Sparse Transformer Decoder that replaces this costly interaction with an efficient, two‑stage process of guided feature selection and focused aggregation. Our core idea is to make the decoder's attention prototype‑guided. We achieve this through a sparse prototype selection mechanism, where each query adaptively identifies a compact set of the most salient voxel features, termed prototypes, for focused feature aggregation.
To ensure this dynamic selection is stable and effective, we introduce a complementary denoising paradigm. This approach leverages ground‑truth masks to provide explicit guidance, guaranteeing a consistent query‑prototype association across decoder layers. Our model, dubbed SPOT‑Occ, outperforms previous methods with a significant margin in speed while also improving accuracy. Source code is released at https://github.com/chensuzeyu/SpotOcc.
Authors:Ning Zhang, Zhengyu Li, Kwong Weng Loh, Mingxi Xu, Qi Wang, Zhengyu Wen, Xiaoyu He, Wei Zhao, Kehong Gong, Mingyuan Zhang
Abstract:
Prior masked modeling motion generation methods predominantly study text‑to‑motion. We present DiMo, a discrete diffusion‑style framework, which extends masked modeling to bidirectional text‑‑motion understanding and generation. Unlike GPT‑style autoregressive approaches that tokenize motion and decode sequentially, DiMo performs iterative masked token refinement, unifying Text‑to‑Motion (T2M), Motion‑to‑Text (M2T), and text‑free Motion‑to‑Motion (M2M) within a single model. This decoding paradigm naturally enables a quality‑latency trade‑off at inference via the number of refinement steps. We further improve motion token fidelity with residual vector quantization (RVQ) and enhance alignment and controllability with Group Relative Policy Optimization (GRPO). Experiments on HumanML3D and KIT‑ML show strong motion quality and competitive bidirectional understanding under a unified framework. In addition, we demonstrate model ability in text‑free motion completion, text‑guided motion prediction and motion caption correction without architectural change. Additional qualitative results are available on our project page: https://animotionlab.github.io/DiMo/.
Authors:Angel Martinez-Sanchez, Parthib Roy, Ross Greer
Abstract:
Instruction‑grounded driving, where passenger language guides trajectory planning, requires vehicles to understand intent before motion. However, most prior instruction‑following planners rely on simulation or fixed command vocabularies, limiting real‑world generalization. doScenes, the first real‑world dataset linking free‑form instructions (with referentiality) to nuScenes ground‑truth motion, enables instruction‑conditioned planning. In this work, we adapt OpenEMMA, an open‑source MLLM‑based end‑to‑end driving framework that ingests front‑camera views and ego‑state and outputs 10‑step speed‑curvature trajectories, to this setting, presenting a reproducible instruction‑conditioned baseline on doScenes and investigate the effects of human instruction prompts on predicted driving behavior. We integrate doScenes directives as passenger‑style prompts within OpenEMMA's vision‑language interface, enabling linguistic conditioning before trajectory generation. Evaluated on 849 annotated scenes using ADE, we observe that instruction conditioning substantially improves robustness by preventing extreme baseline failures, yielding a 98.7% reduction in mean ADE. When such outliers are removed, instructions still influence trajectory alignment, with well‑phrased prompts improving ADE by up to 5.1%. We use this analysis to discuss what makes a "good" instruction for the OpenEMMA framework. We release the evaluation prompts and scripts to establish a reproducible baseline for instruction‑aware planning. GitHub: https://github.com/Mi3‑Lab/doScenes‑VLM‑Planning
Authors:Chenhe Du, Qing Wu, Xuanyu Tian, Jingyi Yu, Hongjiang Wei, Yuyao Zhang
Abstract:
3D medical imaging is in high demand and essential for clinical diagnosis and scientific research. Currently, diffusion models (DMs) have become an effective tool for medical imaging reconstruction thanks to their ability to learn rich, high‑quality data priors. However, learning the 3D data distribution with DMs in medical imaging is challenging, not only due to the difficulties in data collection but also because of the significant computational burden during model training. A common compromise is to train the DMs on 2D data priors and reconstruct stacked 2D slices to address 3D medical inverse problems. However, the intrinsic randomness of diffusion sampling causes severe inter‑slice discontinuities of reconstructed 3D volumes. Existing methods often enforce continuity regularizations along the z‑axis, which introduces sensitive hyper‑parameters and may lead to over‑smoothing results. In this work, we revisit the origin of stochasticity in diffusion sampling and introduce Inter‑Slice Consistent Stochasticity (ISCS), a simple yet effective strategy that encourages interslice consistency during diffusion sampling. Our key idea is to control the consistency of stochastic noise components during diffusion sampling, thereby aligning their sampling trajectories without adding any new loss terms or optimization steps. Importantly, the proposed ISCS is plug‑and‑play and can be dropped into any 2D trained diffusion based 3D reconstruction pipeline without additional computational cost. Experiments on several medical imaging problems show that our method can effectively improve the performance of medical 3D imaging problems based on 2D diffusion models. Our findings suggest that controlling inter‑slice stochasticity is a principled and practically attractive route toward high‑fidelity 3D medical imaging with 2D diffusion priors. The code is available at: https://github.com/duchenhe/ISCS
Authors:Rio Aguina-Kang, Kevin James Blackburn-Matzen, Thibault Groueix, Vladimir Kim, Matheus Gadelha
Abstract:
We present SeeingThroughClutter, a method for reconstructing structured 3D representations from single images by segmenting and modeling objects individually. Prior approaches rely on intermediate tasks such as semantic segmentation and depth estimation, which often underperform in complex scenes, particularly in the presence of occlusion and clutter. We address this by introducing an iterative object removal and reconstruction pipeline that decomposes complex scenes into a sequence of simpler subtasks. Using VLMs as orchestrators, foreground objects are removed one at a time via detection, segmentation, object removal, and 3D fitting. We show that removing objects allows for cleaner segmentations of subsequent objects, even in highly occluded scenes. Our method requires no task‑specific training and benefits directly from ongoing advances in foundation models. We demonstrate stateof‑the‑art robustness on 3D‑Front and ADE20K datasets. Project Page: https://rioak.github.io/seeingthroughclutter/
Authors:Joanna Kaleta, Bartosz Świrta, Kacper Kania, Przemysław Spurek, Marek Kowalski
Abstract:
The growing demand for rapid and scalable 3D asset creation has driven interest in feed‑forward 3D reconstruction methods, with 3D Gaussian Splatting (3DGS) emerging as an effective scene representation. While recent approaches have demonstrated pose‑free reconstruction from unposed image collections, integrating stylization or appearance control into such pipelines remains underexplored. Existing attempts largely rely on image‑based conditioning, which limits both controllability and flexibility. In this work, we introduce AnyStyle, a feed‑forward 3D reconstruction and stylization framework that enables pose‑free, zero‑shot stylization through multimodal conditioning. Our method supports both textual and visual style inputs, allowing users to control the scene appearance using natural language descriptions or reference images. We propose a modular stylization architecture that requires only minimal architectural modifications and can be integrated into existing feed‑forward 3D reconstruction backbones. Experiments demonstrate that AnyStyle improves style controllability over prior feed‑forward stylization methods while preserving high‑quality geometric reconstruction. A user study further confirms that AnyStyle achieves superior stylization quality compared to an existing state‑of‑the‑art approach. Repository: https://github.com/joaxkal/AnyStyle.
Authors:Ahmed Alagha, Christopher Leclerc, Yousef Kotp, Omar Metwally, Calvin Moras, Peter Rentopoulos, Ghodsiyeh Rostami, Bich Ngoc Nguyen, Jumanah Baig, Abdelhakim Khellaf, Vincent Quoc-Huy Trinh, Rabeb Mizouni, Hadi Otrok, Jamal Bentahar, Mahdi S. Hosseini
Abstract:
Whole‑slide image (WSI) preprocessing, comprising tissue detection followed by patch extraction, is foundational to AI‑driven computational pathology but remains a major bottleneck for scaling to large and heterogeneous cohorts. We present AtlasPatch, a scalable framework that couples foundation‑model tissue detection with high‑throughput patch extraction at minimal computational overhead. Our tissue detector achieves high precision (0.986) and remains robust across varying tissue conditions (e.g., brightness, fragmentation, boundary definition, tissue heterogeneity) and common artifacts (e.g., pen/ink markings, scanner streaks). This robustness is enabled by our annotated, heterogeneous multi‑cohort training set of ~30,000 WSI thumbnails combined with efficient adaptation of the Segment‑Anything (SAM) model. AtlasPatch also reduces end‑to‑end WSI preprocessing time by up to 16× versus widely used deep‑learning pipelines, without degrading downstream task performance. The AtlasPatch tool is open‑source, efficiently parallelized for practical deployment, and supports options to save extracted patches or stream them into common feature‑extraction models for on‑the‑fly embedding, making it adaptable to both pathology departments (tissue detection and quality control) and AI researchers (dataset creation and model training). AtlasPatch software package is available at https://github.com/AtlasAnalyticsLab/AtlasPatch.
Authors:Shuo Liu, Ishneet Sukhvinder Singh, Yiqing Xu, Jiafei Duan, Ranjay Krishna
Abstract:
Why do pretrained diffusion or flow‑matching policies fail when the same task is performed near an obstacle, on a shifted support surface, or amid mild clutter? Such failures rarely reflect missing motor skills; instead, they expose a limitation of imitation learning under train‑test shifts, where action generation is tightly coupled to training‑specific spatial configurations and task specifications. Retraining or fine‑tuning to address these failures is costly and conceptually misaligned, as the required behaviors already exist but cannot be selectively adapted at test time. We propose Vision‑Language Steering (VLS), a training‑free framework for inference‑time adaptation of frozen generative robot policies. VLS treats adaptation as an inference‑time control problem, steering the sampling process of a pretrained diffusion or flow‑matching policy in response to out‑of‑distribution observation‑language inputs without modifying policy parameters. By leveraging vision‑language models to synthesize trajectory‑differentiable reward functions, VLS guides denoising toward action trajectories that satisfy test‑time spatial and task requirements. Across simulation and real‑world evaluations, VLS consistently outperforms prior steering methods, achieving a 31% improvement on CALVIN and a 13% gain on LIBERO‑PRO. Real‑world deployment on a Franka robot further demonstrates robust inference‑time adaptation under test‑time spatial and semantic shifts. Project page: https://vision‑language‑steering.github.io/webpage/
Authors:Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik, Munem Shahriar, Mohsin Mahmud Topu, Sadia Tasnim Meem, Rahatun Nesa Priti, Sabrina Afroz Mitu, Md. Iqramul Hoque, Shahriyar Zaman Ridoy, Mohammed Eunus Ali, Majd Hawasly, Mohammad Raza, Md Rizwan Parvez
Abstract:
Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision‑language models (VLMs). Prior work largely relied on synthetic or LLM‑generated environments with limited task designs and puzzle‑like setups, failing to capture the real‑world complexity, visual noise, and diverse spatial relationships that VLMs encounter. To address this, we introduce SpatiaLab, a comprehensive benchmark for evaluating VLMs' spatial reasoning in realistic, unconstrained contexts. SpatiaLab comprises 1,400 visual question‑answer pairs across six major categories: Relative Positioning, Depth & Occlusion, Orientation, Size & Scale, Spatial Navigation, and 3D Geometry, each with five subcategories, yielding 30 distinct task types. Each subcategory contains at least 25 questions, and each main category includes at least 200 questions, supporting both multiple‑choice and open‑ended evaluation. Experiments across diverse state‑of‑the‑art VLMs, including open‑ and closed‑source models, reasoning‑focused, and specialized spatial reasoning models, reveal a substantial gap in spatial reasoning capabilities compared with humans. In the multiple‑choice setup, InternVL3.5‑72B achieves 54.93% accuracy versus 87.57% for humans. In the open‑ended setting, all models show a performance drop of around 10‑25%, with GPT‑5‑mini scoring highest at 40.93% versus 64.93% for humans. These results highlight key limitations in handling complex spatial relationships, depth perception, navigation, and 3D geometry. By providing a diverse, real‑world evaluation framework, SpatiaLab exposes critical challenges and opportunities for advancing VLMs' spatial reasoning, offering a benchmark to guide future research toward robust, human‑aligned spatial understanding. SpatiaLab is available at: https://spatialab‑reasoning.github.io/.
Authors:Jinxing Zhou, Yanghao Zhou, Yaoting Wang, Zongyan Han, Jiaqi Ma, Henghui Ding, Rao Muhammad Anwer, Hisham Cholakkal
Abstract:
Language‑referred audio‑visual segmentation (Ref‑AVS) aims to segment target objects described by natural language by jointly reasoning over video, audio, and text. Beyond generating segmentation masks, providing rich and interpretable diagnoses of mask quality remains largely underexplored. In this work, we introduce Mask Quality Assessment in the Ref‑AVS context (MQA‑RefAVS), a new task that evaluates the quality of candidate segmentation masks without relying on ground‑truth annotations as references at inference time. Given audio‑visual‑language inputs and each provided segmentation mask, the task requires estimating its IoU with the unobserved ground truth, identifying the corresponding error type, and recommending an actionable quality‑control decision. To support this task, we construct MQ‑RAVSBench, a benchmark featuring diverse and representative mask error modes that span both geometric and semantic issues. We further propose MQ‑Auditor, a multimodal large language model (MLLM)‑based auditor that explicitly reasons over multimodal cues and mask information to produce quantitative and qualitative mask quality assessments. Extensive experiments demonstrate that MQ‑Auditor outperforms strong open‑source and commercial MLLMs and can be integrated with existing Ref‑AVS systems to detect segmentation failures and support downstream segmentation improvement. Data and codes will be released at https://github.com/jasongief/MQA‑RefAVS.
Authors:Weiming Chen, Xitong Ling, Xidong Wang, Zhenyang Cai, Yijia Guo, Mingxi Fu, Ziyi Zeng, Minxi Ouyang, Jiawen Li, Yizhi Wang, Tian Guan, Benyou Wang, Yonghong He
Abstract:
Pathology foundation models (PFMs) have rapidly advanced and are becoming a common backbone for downstream clinical tasks, offering strong transferability across tissues and institutions. However, for dense prediction (e.g., segmentation), practical deployment still lacks a clear, reproducible understanding of how different PFMs behave across datasets and how adaptation choices affect performance and stability. We present PFM‑DenseBench, a large‑scale benchmark for dense pathology prediction, evaluating 17 PFMs across 18 public segmentation datasets. Under a unified protocol, we systematically assess PFMs with multiple adaptation and fine‑tuning strategies, and derive insightful, practice‑oriented findings on when and why different PFMs and tuning choices succeed or fail across heterogeneous datasets. We release containers, configs, and dataset cards to enable reproducible evaluation and informed PFM selection for real‑world dense pathology tasks. Project Website: https://m4a1tastegood.github.io/PFM‑DenseBench
Authors:Longjie Zhao, Ziming Hong, Jiaxin Huang, Runnan Chen, Mingming Gong, Tongliang Liu
Abstract:
3D Gaussian Splatting (3DGS) has become a mainstream representation for real‑time 3D scene synthesis, enabling applications in virtual and augmented reality, robotics, and 3D content creation. Its rising commercial value and explicit parametric structure raise emerging intellectual property (IP) protection concerns, prompting a surge of research on 3DGS IP protection. However, current progress remains fragmented, lacking a unified view of the underlying mechanisms, protection paradigms, and robustness challenges. To address this gap, we present the first systematic survey on 3DGS IP protection and introduce a bottom‑up framework that examines (i) underlying Gaussian‑based perturbation mechanisms, (ii) passive and active protection paradigms, and (iii) robustness threats under emerging generative AI era, revealing gaps in technical foundations and robustness characterization and indicating opportunities for deeper investigation. Finally, we outline six research directions across robustness, efficiency, and protection paradigms, offering a roadmap toward reliable and trustworthy IP protection for 3DGS assets.
Authors:Minjun Zhu, Zhen Lin, Yixuan Weng, Panzhong Lu, Qiujie Xie, Yifan Wei, Sifan Liu, Qiyao Sun, Yue Zhang
Abstract:
High‑quality scientific illustrations are crucial for effectively communicating complex scientific and technical concepts, yet their manual creation remains a well‑recognized bottleneck in both academia and industry. We present FigureBench, the first large‑scale benchmark for generating scientific illustrations from long‑form scientific texts. It contains 3,300 high‑quality scientific text‑figure pairs, covering diverse text‑to‑illustration tasks from scientific papers, surveys, blogs, and textbooks. Moreover, we propose AutoFigure, the first agentic framework that automatically generates high‑quality scientific illustrations based on long‑form scientific text. Specifically, before rendering the final result, AutoFigure engages in extensive thinking, recombination, and validation to produce a layout that is both structurally sound and aesthetically refined, outputting a scientific illustration that achieves both structural completeness and aesthetic appeal. Leveraging the high‑quality data from FigureBench, we conduct extensive experiments to test the performance of AutoFigure against various baseline methods. The results demonstrate that AutoFigure consistently surpasses all baseline methods, producing publication‑ready scientific illustrations. The code, dataset and huggingface space are released in https://github.com/ResearAI/AutoFigure.
Authors:Dingkun Zhang, Shuhan Qi, Yulin Wu, Xinyu Xiao, Xuan Wang, Long Chen
Abstract:
Multimodal Large Language Models (MLLMs) suffer from severe training inefficiency issue, which is associated with their massive model sizes and visual token numbers. Existing efforts in efficient training focus on reducing model sizes or trainable parameters. Inspired by the success of Visual Token Pruning (VTP) in improving inference efficiency, we are exploring another substantial research direction for efficient training by reducing visual tokens. However, applying VTP at the training stage results in a training‑inference mismatch: pruning‑trained models perform poorly when inferring on non‑pruned full visual token sequences. To close this gap, we propose DualSpeed, a fast‑slow framework for efficient training of MLLMs. The fast‑mode is the primary mode, which incorporates existing VTP methods as plugins to reduce visual tokens, along with a mode isolator to isolate the model's behaviors. The slow‑mode is the auxiliary mode, where the model is trained on full visual sequences to retain training‑inference consistency. To boost its training, it further leverages self‑distillation to learn from the sufficiently trained fast‑mode. Together, DualSpeed can achieve both training efficiency and non‑degraded performance. Experiments show DualSpeed accelerates the training of LLaVA‑1.5 by 2.1× and LLaVA‑NeXT by 4.0×, retaining over 99% performance. Code: https://github.com/dingkun‑zhang/DualSpeed
Authors:Zimu Lu, Houxing Ren, Yunqiao Yang, Ke Wang, Zhuofan Zong, Mingjie Zhan, Hongsheng Li
Abstract:
Assisting non‑expert users to develop complex interactive websites has become a popular task for LLM‑powered code agents. However, existing code agents tend to only generate frontend web pages, masking the lack of real full‑stack data processing and storage with fancy visual effects. Notably, constructing production‑level full‑stack web applications is far more challenging than only generating frontend web pages, demanding careful control of data flow, comprehensive understanding of constantly updating packages and dependencies, and accurate localization of obscure bugs in the codebase. To address these difficulties, we introduce FullStack‑Agent, a unified agent system for full‑stack agentic coding that consists of three parts: (1) FullStack‑Dev, a multi‑agent framework with strong planning, code editing, codebase navigation, and bug localization abilities. (2) FullStack‑Learn, an innovative data‑scaling and self‑improving method that back‑translates crawled and synthesized website repositories to improve the backbone LLM of FullStack‑Dev. (3) FullStack‑Bench, a comprehensive benchmark that systematically tests the frontend, backend and database functionalities of the generated website. Our FullStack‑Dev outperforms the previous state‑of‑the‑art method by 8.7%, 38.2%, and 15.9% on the frontend, backend, and database test cases respectively. Additionally, FullStack‑Learn raises the performance of a 30B model by 9.7%, 9.5%, and 2.8% on the three sets of test cases through self‑improvement, demonstrating the effectiveness of our approach. The code is released at https://github.com/mnluzimu/FullStack‑Agent.
Authors:Zhixue Fang, Xu He, Songlin Tang, Haoxian Zhang, Qingfeng Li, Xiaoqiang Liu, Pengfei Wan, Kun Gai
Abstract:
Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding novel‑view synthesis. Explicit 3D models, though structurally informative, suffer from inherent inaccuracies (e.g., depth ambiguity and inaccurate dynamics) which, when used as a strong constraint, override the powerful intrinsic 3D awareness of large‑scale video generators. In this work, we revisit motion control from a 3D‑aware perspective, advocating for an implicit, view‑agnostic motion representation that naturally aligns with the generator's spatial priors rather than depending on externally reconstructed constraints. We introduce 3DiMo, which jointly trains a motion encoder with a pretrained video generator to distill driving frames into compact, view‑agnostic motion tokens, injected semantically via cross‑attention. To foster 3D awareness, we train with view‑rich supervision (i.e., single‑view, multi‑view, and moving‑camera videos), forcing motion consistency across diverse viewpoints. Additionally, we use auxiliary geometric supervision that leverages SMPL only for early initialization and is annealed to zero, enabling the model to transition from external 3D guidance to learning genuine 3D spatial motion understanding from the data and the generator's priors. Experiments confirm that 3DiMo faithfully reproduces driving motions with flexible, text‑driven camera control, significantly surpassing existing methods in both motion fidelity and visual quality.
Authors:Jingjing Peng, Giorgio Fiore, Yang Liu, Ksenia Ellum, Debayan Daspupta, Keyoumars Ashkan, Andrew McEvoy, Anna Miserocchi, Sebastien Ourselin, John Duncan, Alejandro Granados
Abstract:
Introduction: In neurosurgery, image‑guided Neurosurgery Systems (IGNS) highly rely on preoperative brain magnetic resonance images (MRI) to assist surgeons in locating surgical targets and determining surgical paths. However, brain shift invalidates the preoperative MRI after dural opening. Updated intraoperative brain MRI with brain shift compensation is crucial for enhancing the precision of neuronavigation systems and ensuring the optimal outcome of surgical interventions. Methodology: We propose NeuralShift, a U‑Net‑based model that predicts brain shift entirely from pre‑operative MRI for patients undergoing temporal lobe resection. We evaluated our results using Target Registration Errors (TREs) computed on anatomical landmarks located on the resection side and along the midline, and DICE scores comparing predicted intraoperative masks with masks derived from intraoperative MRI. Results: Our experimental results show that our model can predict the global deformation of the brain (DICE of 0.97) with accurate local displacements (achieve landmark TRE as low as 1.12 mm), compensating for large brain shifts during temporal lobe removal neurosurgery. Conclusion: Our proposed model is capable of predicting the global deformation of the brain during temporal lobe resection using only preoperative images, providing potential opportunities to the surgical team to increase safety and efficiency of neurosurgery and better outcomes to patients. Our contributions will be publicly available after acceptance in https://github.com/SurgicalDataScienceKCL/NeuralShift.
Authors:Nicolas Sereyjol-Garros, Ellington Kirby, Victor Letzelter, Victor Besnier, Nermin Samet
Abstract:
While representation alignment with self‑supervised models has been shown to improve diffusion model training, its potential for enhancing inference‑time conditioning remains largely unexplored. We introduce Representation‑Aligned Guidance (REPA‑G), a framework that leverages these aligned representations, with rich semantic properties, to enable test‑time conditioning from features in generation. By optimizing a similarity objective (the potential) at inference, we steer the denoising process toward a conditioned representation extracted from a pre‑trained feature extractor. Our method provides versatile control at multiple scales, ranging from fine‑grained texture matching via single patches to broad semantic guidance using global image feature tokens. We further extend this to multi‑concept composition, allowing for the faithful combination of distinct concepts. REPA‑G operates entirely at inference time, offering a flexible and precise alternative to often ambiguous text prompts or coarse class labels. We theoretically justify how this guidance enables sampling from the potential‑induced tilted distribution. Quantitative results on ImageNet and COCO demonstrate that our approach achieves high‑quality, diverse generations. Code is available at https://github.com/valeoai/REPA‑G.
Authors:Pengfei Yue, Xiaokang Jiang, Yilin Lu, Jianghang Lin, Shengchuan Zhang, Liujuan Cao
Abstract:
Industrial Anomaly Detection (IAD) is vital for manufacturing, yet traditional methods face significant challenges: unsupervised approaches yield rough localizations requiring manual thresholds, while supervised methods overfit due to scarce, imbalanced data. Both suffer from the "One Anomaly Class, One Model" limitation. To address this, we propose Referring Industrial Anomaly Segmentation (RIAS), a paradigm leveraging language to guide detection. RIAS generates precise masks from text descriptions without manual thresholds and uses universal prompts to detect diverse anomalies with a single model. We introduce the MVTec‑Ref dataset to support this, designed with diverse referring expressions and focusing on anomaly patterns, notably with 95% small anomalies. We also propose the Dual Query Token with Mask Group Transformer (DQFormer) benchmark, enhanced by Language‑Gated Multi‑Level Aggregation (LMA) to improve multi‑scale segmentation. Unlike traditional methods using redundant queries, DQFormer employs only "Anomaly" and "Background" tokens for efficient visual‑textual integration. Experiments demonstrate RIAS's effectiveness in advancing IAD toward open‑set capabilities. Code: https://github.com/swagger‑coder/RIAS‑MVTec‑Ref.
Authors:Wei Zhang, Xiang Liu, Ningjing Liu, Mingxin Liu, Wei Liao, Chunyan Xu, Xue Yang
Abstract:
A consistent trend throughout the research of oriented object detection has been the pursuit of maintaining comparable performance with fewer and weaker annotations. This is particularly crucial in the remote sensing domain, where the dense object distribution and a wide variety of categories contribute to prohibitively high costs. Based on the supervision level, existing oriented object detection algorithms can be broadly grouped into fully supervised, semi‑supervised, and weakly supervised methods. Within the scope of this work, we further categorize them to include sparsely supervised and partially weakly‑supervised methods. To address the challenges of large‑scale labeling, we introduce the first Sparse Partial Weakly‑Supervised Oriented Object Detection framework, designed to efficiently leverage only a few sparse weakly‑labeled data and plenty of unlabeled data. Our framework incorporates three key innovations: (1) We design a Sparse‑annotation‑Orientation‑and‑Scale‑aware Student (SOS‑Student) model to separate unlabeled objects from the background in a sparsely‑labeled setting, and learn orientation and scale information from orientation‑agnostic or scale‑agnostic weak annotations. (2) We construct a novel Multi‑level Pseudo‑label Filtering strategy that leverages the distribution of model predictions, which is informed by the model's multi‑layer predictions. (3) We propose a unique sparse partitioning approach, ensuring equal treatment for each category. Extensive experiments on the DOTA and DIOR datasets show that our framework achieves a significant performance gain over traditional oriented object detection methods mentioned above, offering a highly cost‑effective solution. Our code is publicly available at https://github.com/VisionXLab/SPWOOD.
Authors:Estelle Chigot, Thomas Oberlin, Manon Huguenin, Dennis Wilson
Abstract:
Semantic segmentation networks require large amounts of pixel‑level annotated data, which are costly to obtain for real‑world images. Computer graphics engines can generate synthetic images alongside their ground‑truth annotations. However, models trained on such images can perform poorly on real images due to the domain gap between real and synthetic images. Style transfer methods can reduce this difference by applying a realistic style to synthetic images. Choosing effective data transformations and their sequence is difficult due to the large combinatorial search space of style transfer operators. Using multi‑objective genetic algorithms, we optimize pipelines to balance structural coherence and style similarity to target domains. We study the use of paired‑image metrics on individual image samples during evolution to enable rapid pipeline evaluation, as opposed to standard distributional metrics that require the generation of many images. After optimization, we evaluate the resulting Pareto front using distributional metrics and segmentation performance. We apply this approach to standard datasets in synthetic‑to‑real domain adaptation: from the video game GTA5 to real image datasets Cityscapes and ACDC, focusing on adverse conditions. Results demonstrate that evolutionary algorithms can propose diverse augmentation pipelines adapted to different objectives. The contribution of this work is the formulation of style transfer as a sequencing problem suitable for evolutionary optimization and the study of efficient metrics that enable feasible search in this space. The source code is available at: https://github.com/echigot/MOOSS.
Authors:Basile Terver, Randall Balestriero, Megi Dervishi, David Fan, Quentin Garrido, Tushar Nagarajan, Koustuv Sinha, Wancong Zhang, Mike Rabbat, Yann LeCun, Amir Bar
Abstract:
We present EB‑JEPA, an open‑source library for learning representations and world models using Joint‑Embedding Predictive Architectures (JEPAs). JEPAs learn to predict in representation space rather than pixel space, avoiding the pitfalls of generative modeling while capturing semantically meaningful features suitable for downstream tasks. Our library provides modular, self‑contained implementations that illustrate how representation learning techniques developed for image‑level self‑supervised learning can transfer to video, where temporal dynamics add complexity, and ultimately to action‑conditioned world models, where the model must additionally learn to predict the effects of control inputs. Each example is designed for single‑GPU training within a few hours, making energy‑based self‑supervised learning accessible for research and education. We provide ablations of JEA components on CIFAR‑10. Probing these representations yields 91% accuracy, indicating that the model learns useful features. Extending to video, we include a multi‑step prediction example on Moving MNIST that demonstrates how the same principles scale to temporal modeling. Finally, we show how these representations can drive action‑conditioned world models, achieving a 97% planning success rate on the Two Rooms navigation task. Comprehensive ablations reveal the critical importance of each regularization component for preventing representation collapse. Code is available at https://github.com/facebookresearch/eb_jepa.
Authors:Haichao Jiang, Tianming Liang, Wei-Shi Zheng, Jian-Fang Hu
Abstract:
Referring Video Object Segmentation (RVOS) aims to segment objects in videos based on textual queries. Current methods mainly rely on large‑scale supervised fine‑tuning (SFT) of Multi‑modal Large Language Models (MLLMs). However, this paradigm suffers from heavy data dependence and limited scalability against the rapid evolution of MLLMs. Although recent zero‑shot approaches offer a flexible alternative, their performance remains significantly behind SFT‑based methods, due to the straightforward workflow designs. To address these limitations, we propose Refer‑Agent, a collaborative multi‑agent system with alternating reasoning‑reflection mechanisms. This system decomposes RVOS into step‑by‑step reasoning process. During reasoning, we introduce a Coarse‑to‑Fine frame selection strategy to ensure the frame diversity and textual relevance, along with a Dynamic Focus Layout that adaptively adjusts the agent's visual focus. Furthermore, we propose a Chain‑of‑Reflection mechanism, which employs a Questioner‑Responder pair to generate a self‑reflection chain, enabling the system to verify intermediate results and generates feedback for next‑round reasoning refinement. Extensive experiments on five challenging benchmarks demonstrate that Refer‑Agent significantly outperforms state‑of‑the‑art methods, including both SFT‑based models and zero‑shot approaches. Moreover, Refer‑Agent is flexible and enables fast integration of new MLLMs without any additional fine‑tuning costs. Code will be released at https://github.com/iSEE‑Laboratory/Refer‑Agent.
Authors:Wenji Wu, Shuo Ye, Yiyu Liu, Jiguang He, Zhuo Wang, Zitong Yu
Abstract:
Underwater Camouflaged Object Detection (UCOD) is a challenging task due to the extreme visual similarity between targets and backgrounds across varying marine depths. Existing methods often struggle with topological fragmentation of slender creatures in the deep sea and the subtle feature extraction of transparent organisms. In this paper, we propose DeepTopo‑Net, a novel framework that integrates topology‑aware modeling with frequency‑decoupled perception. To address physical degradation, we design the Water‑Conditioned Adaptive Perceptor (WCAP), which employs Riemannian metric tensors to dynamically deform convolutional sampling fields. Furthermore, the Abyssal‑Topology Refinement Module (ATRM) is developed to maintain the structural connectivity of spindly targets through skeletal priors. Specifically, we first introduce GBU‑UCOD, the first high‑resolution (2K) benchmark tailored for marine vertical zonation, filling the data gap for hadal and abyssal zones. Extensive experiments on MAS3K, RMAS, and our proposed GBU‑UCOD datasets demonstrate that DeepTopo‑Net achieves state‑of‑the‑art performance, particularly in preserving the morphological integrity of complex underwater patterns. The datasets and codes will be released at https://github.com/Wuwenji18/GBU‑UCOD.
Authors:Yongwei Chen, Tianyi Wei, Yushi Lan, Zhaoyang Lyu, Shangchen Zhou, Xudong Xu, Xingang Pan
Abstract:
The rapid progress of large multimodal models has inspired efforts toward unified frameworks that couple understanding and generation. While such paradigms have shown remarkable success in 2D, extending them to 3D remains largely underexplored. Existing attempts to unify 3D tasks under a single autoregressive (AR) paradigm lead to significant performance degradation due to forced signal quantization and prohibitive training cost. Our key insight is that the essential challenge lies not in enforcing a unified autoregressive paradigm, but in enabling effective information interaction between generation and understanding while minimally compromising their inherent capabilities and leveraging pretrained models to reduce training cost. Guided by this perspective, we present the first unified framework for 3D understanding and generation that combines autoregression with diffusion. Specifically, we adopt an autoregressive next‑token prediction paradigm for 3D understanding, and a continuous diffusion paradigm for 3D generation. A lightweight transformer bridges the feature space of large language models and the conditional space of 3D diffusion models, enabling effective cross‑modal information exchange while preserving the priors learned by standalone models. Extensive experiments demonstrate that our framework achieves state‑of‑the‑art performance across diverse 3D understanding and generation benchmarks, while also excelling in 3D editing tasks. These results highlight the potential of unified AR+diffusion models as a promising direction for building more general‑purpose 3D intelligence.
Authors:Yingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang, Youcheng Pan, Xiaoqiang Zhou, Min Zhang
Abstract:
Reasoning over table images remains challenging for Large Vision‑Language Models (LVLMs) due to complex layouts and tightly coupled structure‑content information. Existing solutions often depend on expensive supervised training, reinforcement learning, or external tools, limiting efficiency and scalability. This work addresses a key question: how to adapt LVLMs to table reasoning with minimal annotation and no external tools? Specifically, we first introduce DiSCo, a Disentangled Structure‑Content alignment framework that explicitly separates structural abstraction from semantic grounding during multimodal alignment, efficiently adapting LVLMs to tables structures. Building on DiSCo, we further present Table‑GLS, a Global‑to‑Local Structure‑guided reasoning framework that performs table reasoning via structured exploration and evidence‑grounded inference. Extensive experiments across diverse benchmarks demonstrate that our framework efficiently enhances LVLM's table understanding and reasoning capabilities, particularly generalizing to unseen table structures. Our data and code are available at https://github.com/AAAndy‑Zhu/TableVLM.
Authors:Meng Lou, Yunxiang Fu, Yizhou Yu
Abstract:
Continual learning, especially class‑incremental learning (CIL), on the basis of a pre‑trained model (PTM) has garnered substantial research interest in recent years. However, how to effectively learn both discriminative and comprehensive feature representations while maintaining stability and plasticity over very long task sequences remains an open problem. We propose CaRE, a scalable Continual Learner with efficient Bi‑Level Routing Mixture‑of‑Experts (BR‑MoE). The core idea of BR‑MoE is a bi‑level routing mechanism: a router selection stage that dynamically activates relevant task‑specific routers, followed by an expert routing phase that dynamically activates and aggregates experts, aiming to inject discriminative and comprehensive representations into every intermediate network layer. On the other hand, we introduce a challenging dataset, OmniBenchmark‑1K, for CIL performance evaluation on very long task sequences with hundreds of tasks. Extensive experiments show that CaRE demonstrates leading performance across a variety of datasets and task settings, including commonly used CIL datasets with classical CIL settings (e.g., 5‑20 tasks). To the best of our knowledge, CaRE is the first continual learner that scales to very long task sequences (ranging from 100 to over 300 non‑overlapping tasks), while outperforming all baselines by a large margin on such task sequences. We hope that this work will inspire further research into continual learning over extremely long task sequences. Code and dataset are publicly released at https://github.com/LMMMEng/CaRE.
Authors:Yu-Hsiang Chen, Wei-Jer Chang, Christian Kotulla, Thomas Keutgens, Steffen Runde, Tobias Moers, Christoph Klas, Wei Zhan, Masayoshi Tomizuka, Yi-Ting Chen
Abstract:
We present HetroD, a dataset and benchmark for developing autonomous driving systems in heterogeneous environments. HetroD targets the critical challenge of navi‑ gating real‑world heterogeneous traffic dominated by vulner‑ able road users (VRUs), including pedestrians, cyclists, and motorcyclists that interact with vehicles. These mixed agent types exhibit complex behaviors such as hook turns, lane splitting, and informal right‑of‑way negotiation. Such behaviors pose significant challenges for autonomous vehicles but remain underrepresented in existing datasets focused on structured, lane‑disciplined traffic. To bridge the gap, we collect a large‑ scale drone‑based dataset to provide a holistic observation of traffic scenes with centimeter‑accurate annotations, HD maps, and traffic signal states. We further develop a modular toolkit for extracting per‑agent scenarios to support downstream task development. In total, the dataset comprises over 65.4k high‑ fidelity agent trajectories, 70% of which are from VRUs. HetroD supports modeling of VRU behaviors in dense, het‑ erogeneous traffic and provides standardized benchmarks for forecasting, planning, and simulation tasks. Evaluation results reveal that state‑of‑the‑art prediction and planning models struggle with the challenges presented by our dataset: they fail to predict lateral VRU movements, cannot handle unstructured maneuvers, and exhibit limited performance in dense and multi‑agent scenarios, highlighting the need for more robust approaches to heterogeneous traffic. See our project page for more examples: https://hetroddata.github.io/HetroD/
Authors:Xiaofeng Tan, Jun Liu, Yuanting Fan, Bin-Bin Gao, Xi Jiang, Xiaochen Chen, Jinlong Peng, Chengjie Wang, Hongsong Wang, Feng Zheng
Abstract:
Reinforcement Fine‑Tuning (RFT) on flow‑based models is crucial for preference alignment. However, they often introduce visual hallucinations like over‑optimized details and semantic misalignment. This work preliminarily explores why visual hallucinations arise and how to reduce them. We first investigate RFT methods from a unified perspective, and reveal the core problems stemming from two aspects, exploration and exploitation: (1) limited exploration during stochastic differential equation (SDE) rollouts, leading to an over‑emphasis on local details at the expense of global semantics, and (2) trajectory imitation process inherent in policy gradient methods, distorting the model's foundational vector field and its cross‑step consistency. Building on this, we propose ConsistentRFT, a general framework to mitigate these hallucinations. Specifically, we design a Dynamic Granularity Rollout (DGR) mechanism to balance exploration between global semantics and local details by dynamically scheduling different noise sources. We then introduce a Consistent Policy Gradient Optimization (CPGO) that preserves the model's consistency by aligning the current policy with a more stable prior. Extensive experiments demonstrate that ConsistentRFT significantly mitigates visual hallucinations, achieving average reductions of 49% for low‑level and 38% for high‑level perceptual hallucinations. Furthermore, ConsistentRFT outperforms other RFT methods on out‑of‑domain metrics, showing an improvement of 5.1% (v.s. the baseline's decrease of ‑0.4%) over FLUX1.dev. This is \hrefhttps://xiaofeng‑tan.github.io/projects/ConsistentRFTProject Page.
Authors:Hyun Seok Seong, WonJun Moon, Jae-Pil Heo
Abstract:
Unsupervised object‑centric learning models, particularly slot‑based architectures, have shown great promise in decomposing complex scenes. However, their reliance on reconstruction‑based training creates a fundamental conflict between the sharp, high‑frequency attention maps of the encoder and the spatially consistent but blurry reconstruction maps of the decoder. We identify that this discrepancy gives rise to a vicious cycle: the noisy feature map from the encoder forces the decoder to average over possibilities and produce even blurrier outputs, while the gradient computed from blurry reconstruction maps lacks high‑frequency details necessary to supervise encoder features. To break this cycle, we introduce Synergistic Representation Learning (SRL) that establishes a virtuous cycle where the encoder and decoder mutually refine one another. SRL leverages the encoder's sharpness to deblur the semantic boundary within the decoder output, while exploiting the decoder's spatial consistency to denoise the encoder's features. This mutual refinement process is stabilized by a warm‑up phase with a slot regularization objective that initially allocates distinct entities per slot. By bridging the representational gap between the encoder and decoder, SRL achieves state‑of‑the‑art results on video object‑centric learning benchmarks. Codes are available at https://github.com/hynnsk/SRL.
Authors:Constantin Selzer, Fabina B. Flohr
Abstract:
Trajectory prediction and planning are fundamental yet disconnected components in autonomous driving. Prediction models forecast surrounding agent motion under unknown intentions, producing multimodal distributions, while planning assumes known ego objectives and generates deterministic trajectories. This mismatch creates a critical bottleneck: prediction lacks supervision for agent intentions, while planning requires this information. Existing prediction models, despite strong benchmarking performance, often remain disconnected from planning constraints such as collision avoidance and dynamic feasibility. We introduce Plan TRansformer (PTR), a unified Gaussian Mixture Transformer framework integrating goal‑conditioned prediction, dynamic feasibility, interaction awareness, and lane‑level topology reasoning. A teacher‑student training strategy progressively masks surrounding agent commands during training to align with inference conditions where agent intentions are unavailable. PTR achieves 4.3%/3.5% improvement in marginal/joint mAP compared to the baseline Motion Transformer (MTR) and 15.5% planning error reduction at 5s horizon compared to GameFormer. The architecture‑agnostic design enables application to diverse Transformer‑based prediction models. Project Website: https://github.com/SelzerConst/PlanTRansformer
Authors:Mario Pascual-González, Ariadna Jiménez-Partinen, R. M. Luque-Baena, Fátima Nagib-Raya, Ezequiel López-Rubio
Abstract:
Focal cortical dysplasia (FCD) lesions in epilepsy FLAIR MRI are subtle and scarce, making joint image‑‑mask generative modeling prone to instability and memorization. We propose SLIM‑Diff, a compact joint diffusion model whose main contributions are (i) a single shared‑bottleneck U‑Net that enforces tight coupling between anatomy and lesion geometry from a 2‑channel image+mask representation, and (ii) loss‑geometry tuning via a tunable L_p objective. As an internal baseline, we include the canonical DDPM‑style objective (ε‑prediction with L_2 loss) and isolate the effect of prediction parameterization and L_p geometry under a matched setup. Experiments show that x_0‑prediction is consistently the strongest choice for joint synthesis, and that fractional sub‑quadratic penalties (L_1.5) improve image fidelity while L_2 better preserves lesion mask morphology. Our code and model weights are available in https://github.com/MarioPasc/slim‑diff
Authors:Zhiwen Yang, Yuxin Peng
Abstract:
Camera‑based 3D semantic scene completion (SSC) offers a cost‑effective solution for assessing the geometric occupancy and semantic labels of each voxel in the surrounding 3D scene with image inputs, providing a voxel‑level scene perception foundation for the perception‑prediction‑planning autonomous driving systems. Although significant progress has been made in existing methods, their optimization rely solely on the supervision from voxel labels and face the challenge of voxel sparsity as a large portion of voxels in autonomous driving scenarios are empty, which limits both optimization efficiency and model performance. To address this issue, we propose a Multi‑Resolution Alignment (MRA) approach to mitigate voxel sparsity in camera‑based 3D semantic scene completion, which exploits the scene and instance level alignment across multi‑resolution 3D features as auxiliary supervision. Specifically, we first propose the Multi‑resolution View Transformer module, which projects 2D image features into multi‑resolution 3D features and aligns them at the scene level through fusing discriminative seed features. Furthermore, we design the Cubic Semantic Anisotropy module to identify the instance‑level semantic significance of each voxel, accounting for the semantic differences of a specific voxel against its neighboring voxels within a cubic area. Finally, we devise a Critical Distribution Alignment module, which selects critical voxels as instance‑level anchors with the guidance of cubic semantic anisotropy, and applies a circulated loss for auxiliary supervision on the critical feature distribution consistency across different resolutions. The code is available at https://github.com/PKU‑ICST‑MIPL/MRA_TIP.
Authors:Nikita Drozdov, Andrey Lemeshko, Nikita Gavrilov, Anton Konushin, Danila Rukhovich, Maksim Kolodiazhnyi
Abstract:
3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero‑shot 3DVG from multi‑view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi‑view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero‑shot methods causing significant performance degradation and address them with (i) a state‑of‑the‑art zero‑shot 3D instance segmentation method to generate high‑quality 3D bounding box proposals and (ii) advanced reasoning via prompt‑based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state‑of‑the‑art performance among zero‑shot methods. Code is available at https://github.com/col14m/z3d .
Authors:Bryan Sangwoo Kim, Jonghyun Park, Jong Chul Ye
Abstract:
Text‑conditioned diffusion models have advanced image and video super‑resolution by using prompts as semantic priors, and modern super‑resolution pipelines typically rely on latent tiling to scale to high resolutions. In practice, a single global caption is used with the latent tiling, often causing prompt misguidance. Specifically, a coarse global prompt often misses localized details (errors of omission) and provides locally irrelevant guidance (errors of commission) which leads to substandard results at the tile level. To solve this, we propose Tiled Prompts, a unified framework for image and video super‑resolution that generates a tile‑specific prompt for each latent tile and performs super‑resolution under locally text‑conditioned posteriors to resolve prompt misguidance with minimal overhead. Our experiments on high resolution real‑world images and videos show that tiled prompts bring consistent gains in perceptual quality and fidelity, while reducing hallucinations and tile‑level artifacts that can be found in global‑prompt baselines. Project Page: https://bryanswkim.github.io/tiled‑prompts/.
Authors:Haoran Li, Renyang Liu, Hongjia Liu, Chen Wang, Long Yin, Jian Xu
Abstract:
Recent progress in adversarial attacks on 3D point clouds, particularly in achieving spatial imperceptibility and high attack performance, presents significant challenges for defenders. Current defensive approaches remain cumbersome, often requiring invasive model modifications, expensive training procedures or auxiliary data access. To address these threats, in this paper, we propose a plug‑and‑play and non‑invasive defense mechanism in the spectral domain, grounded in a theoretical and empirical analysis of the relationship between imperceptible perturbations and high‑frequency spectral components. Building upon these insights, we introduce a novel purification framework, termed PWAVEP, which begins by computing a spectral graph wavelet domain saliency score and local sparsity score for each point. Guided by these values, PWAVEP adopts a hierarchical strategy, it eliminates the most salient points, which are identified as hardly recoverable adversarial outliers. Simultaneously, it applies a spectral filtering process to a broader set of moderately salient points. This process leverages a graph wavelet transform to attenuate high‑frequency coefficients associated with the targeted points, thereby effectively suppressing adversarial noise. Extensive evaluations demonstrate that the proposed PWAVEP achieves superior accuracy and robustness compared to existing approaches, advancing the state‑of‑the‑art in 3D point cloud purification. Code and datasets are available at https://github.com/a772316182/pwavep
Authors:Shengyuan Liu, Liuxin Bao, Qi Yang, Wanting Geng, Boyun Zheng, Chenxin Li, Wenting Chen, Houwen Peng, Yixuan Yuan
Abstract:
Medical image segmentation is evolving from task‑specific models toward generalizable frameworks. Recent research leverages Multi‑modal Large Language Models (MLLMs) as autonomous agents, employing reinforcement learning with verifiable reward (RLVR) to orchestrate specialized tools like the Segment Anything Model (SAM). However, these approaches often rely on single‑turn, rigid interaction strategies and lack process‑level supervision during training, which hinders their ability to fully exploit the dynamic potential of interactive tools and leads to redundant actions. To bridge this gap, we propose MedSAM‑Agent, a framework that reformulates interactive segmentation as a multi‑step autonomous decision‑making process. First, we introduce a hybrid prompting strategy for expert‑curated trajectory generation, enabling the model to internalize human‑like decision heuristics and adaptive refinement strategies. Furthermore, we develop a two‑stage training pipeline that integrates multi‑turn, end‑to‑end outcome verification with a clinical‑fidelity process reward design to promote interaction parsimony and decision efficiency. Extensive experiments across 6 medical modalities and 21 datasets demonstrate that MedSAM‑Agent achieves state‑of‑the‑art performance, effectively unifying autonomous medical reasoning with robust, iterative optimization. Code is available \hrefhttps://github.com/CUHK‑AIM‑Group/MedSAM‑Agenthere.
Authors:Songming Liu, Bangguo Li, Kai Ma, Lingxuan Wu, Hengkai Tan, Xiao Ouyang, Hang Su, Jun Zhu
Abstract:
Vision‑Language‑Action (VLA) models hold promise for generalist robotics but currently struggle with data scarcity, architectural inefficiencies, and the inability to generalize across different hardware platforms. We introduce RDT2, a robotic foundation model built upon a 7B parameter VLM designed to enable zero‑shot deployment on novel embodiments for open‑vocabulary tasks. To achieve this, we collected one of the largest open‑source robotic datasets‑‑over 10,000 hours of demonstrations in diverse families‑‑using an enhanced, embodiment‑agnostic Universal Manipulation Interface (UMI). Our approach employs a novel three‑stage training recipe that aligns discrete linguistic knowledge with continuous control via Residual Vector Quantization (RVQ), flow‑matching, and distillation for real‑time inference. Consequently, RDT2 becomes one of the first models that simultaneously zero‑shot generalizes to unseen objects, scenes, instructions, and even robotic platforms. Besides, it outperforms state‑of‑the‑art baselines in dexterous, long‑horizon, and dynamic downstream tasks like playing table tennis. See https://rdt‑robotics.github.io/rdt2/ for more information.
Authors:Jianghao Wu, Xiangde Luo, Yubo Zhou, Lianming Wu, Guotai Wang, Shaoting Zhang
Abstract:
Test‑Time Adaptation (TTA) offers a practical solution for deploying image segmentation models under domain shift without accessing source data or retraining. Among existing TTA strategies, pseudo‑label‑based methods have shown promising performance. However, they often rely on perturbation‑ensemble heuristics (e.g., dropout sampling, test‑time augmentation, Gaussian noise), which lack distributional grounding and yield unstable training signals. This can trigger error accumulation and catastrophic forgetting during adaptation. To address this, we propose A3‑TTA, a TTA framework that constructs reliable pseudo‑labels through anchor‑guided supervision. Specifically, we identify well‑predicted target domain images using a class compact density metric, under the assumption that confident predictions imply distributional proximity to the source domain. These anchors serve as stable references to guide pseudo‑label generation, which is further regularized via semantic consistency and boundary‑aware entropy minimization. Additionally, we introduce a self‑adaptive exponential moving average strategy to mitigate label noise and stabilize model update during adaptation. Evaluated on both multi‑domain medical images (heart structure and prostate segmentation) and natural images, A3‑TTA significantly improves average Dice scores by 10.40 to 17.68 percentage points compared to the source model, outperforming several state‑of‑the‑art TTA methods under different segmentation model architectures. A3‑TTA also excels in continual TTA, maintaining high performance across sequential target domains with strong anti‑forgetting ability. The code will be made publicly available at https://github.com/HiLab‑git/A3‑TTA.
Authors:Yi Yu, Qixin Zhang, Shuhan Ye, Xun Lin, Qianshan Wei, Kun Wang, Wenhan Yang, Dacheng Tao, Xudong Jiang
Abstract:
Spiking neural networks (SNNs) compute with discrete spikes and exploit temporal structure, yet most adversarial attacks change intensities or event counts instead of timing. We study a timing‑only adversary that retimes existing spikes while preserving spike counts and amplitudes in event‑driven SNNs, thus remaining rate‑preserving. We formalize a capacity‑1 spike‑retiming threat model with a unified trio of budgets: per‑spike jitter \mathcalB_\infty, total delay \mathcalB_1, and tamper count \mathcalB_0. Feasible adversarial examples must satisfy timeline consistency and non‑overlap, which makes the search space discrete and constrained. To optimize such retimings at scale, we use projected‑in‑the‑loop (PIL) optimization: shift‑probability logits yield a differentiable soft retiming for backpropagation, and a strict projection in the forward pass produces a feasible discrete schedule that satisfies capacity‑1, non‑overlap, and the chosen budget at every step. The objective maximizes task loss on the projected input and adds a capacity regularizer together with budget‑aware penalties, which stabilizes gradients and aligns optimization with evaluation. Across event‑driven benchmarks (CIFAR10‑DVS, DVS‑Gesture, N‑MNIST) and diverse SNN architectures, we evaluate under binary and integer event grids and a range of retiming budgets, and also test models trained with timing‑aware adversarial training designed to counter timing‑only attacks. For example, on DVS‑Gesture the attack attains high success (over 90%) while touching fewer than 2% of spikes under \mathcalB_0. Taken together, our results show that spike retiming is a practical and stealthy attack surface that current defenses struggle to counter, providing a clear reference for temporal robustness in event‑driven SNNs. Code is available at https://github.com/yuyi‑sd/Spike‑Retiming‑Attacks.
Authors:Francesco Di Salvo, Sebastian Doerrich, Jonas Alle, Christian Ledig
Abstract:
Robust generalization beyond training distributions remains a critical challenge for deep neural networks. This is especially pronounced in medical image analysis, where data is often scarce and covariate shifts arise from different hardware devices, imaging protocols, and heterogeneous patient populations. These factors collectively hinder reliable performance and slow down clinical adoption. Despite recent progress, existing learning paradigms primarily rely on the Euclidean manifold, whose flat geometry fails to capture the complex, hierarchical structures present in clinical data. In this work, we exploit the advantages of hyperbolic manifolds to model complex data characteristics. We present the first comprehensive validation of hyperbolic representation learning for medical image analysis and demonstrate statistically significant gains across eleven in‑distribution datasets and three ViT models. We further propose an unsupervised, domain‑invariant hyperbolic cross‑branch consistency constraint. Extensive experiments confirm that our proposed method promotes domain‑invariant features and outperforms state‑of‑the‑art Euclidean methods by an average of +2.1% AUC on three domain generalization benchmarks: Fitzpatrick17k, Camelyon17‑WILDS, and a cross‑dataset setup for retinal imaging. These datasets span different imaging modalities, data sizes, and label granularities, confirming generalization capabilities across substantially different conditions. The code is available at https://github.com/francescodisalvo05/hyperbolic‑cross‑branch‑consistency .
Authors:Ofer Idan, Dan Badur, Yosi Keller, Yoli Shavit
Abstract:
Visual Place Recognition (VPR) often fails under extreme environmental changes and perceptual aliasing. Furthermore, standard systems cannot perform "blind" localization from verbal descriptions alone, a capability needed for applications such as emergency response. To address these challenges, we introduce LaVPR, a large‑scale benchmark that extends existing VPR datasets with over 650,000 rich natural‑language descriptions. Using LaVPR, we investigate two paradigms: Multi‑Modal Fusion for enhanced robustness and Cross‑Modal Retrieval for language‑based localization. Our results show that language descriptions yield consistent gains in visually degraded conditions, with the most significant impact on smaller backbones. Notably, adding language allows compact models to rival the performance of much larger vision‑only architectures. For cross‑modal retrieval, we establish a baseline using Low‑Rank Adaptation (LoRA) and Multi‑Similarity loss, which substantially outperforms standard contrastive methods across vision‑language models. Ultimately, LaVPR enables a new class of localization systems that are both resilient to real‑world stochasticity and practical for resource‑constrained deployment. Our dataset and code are available at https://github.com/oferidan1/LaVPR.
Authors:Zhuoran Yang, Xi Guo, Chenjing Ding, Chiyu Wang, Wei Wu, Yanyong Zhang
Abstract:
Autonomous driving relies on robust models trained on high‑quality, large‑scale multi‑view driving videos. While world models offer a cost‑effective solution for generating realistic driving videos, they struggle to maintain instance‑level temporal consistency and spatial geometric fidelity. To address these challenges, we propose InstaDrive, a novel framework that enhances driving video realism through two key advancements: (1) Instance Flow Guider, which extracts and propagates instance features across frames to enforce temporal consistency, preserving instance identity over time. (2) Spatial Geometric Aligner, which improves spatial reasoning, ensures precise instance positioning, and explicitly models occlusion hierarchies. By incorporating these instance‑aware mechanisms, InstaDrive achieves state‑of‑the‑art video generation quality and enhances downstream autonomous driving tasks on the nuScenes dataset. Additionally, we utilize CARLA's autopilot to procedurally and stochastically simulate rare but safety‑critical driving scenarios across diverse maps and regions, enabling rigorous safety evaluation for autonomous systems. Our project page is https://shanpoyang654.github.io/InstaDrive/page.html.
Authors:Zhuoran Yang, Yanyong Zhang
Abstract:
Autonomous driving relies on robust models trained on large‑scale, high‑quality multi‑view driving videos. Although world models provide a cost‑effective solution for generating realistic driving data, they often suffer from identity drift, where the same object changes its appearance or category across frames due to the absence of instance‑level temporal constraints. We introduce ConsisDrive, an identity‑preserving driving world model designed to enforce temporal consistency at the instance level. Our framework incorporates two key components: (1) Instance‑Masked Attention, which applies instance identity masks and trajectory masks within attention blocks to ensure that visual tokens interact only with their corresponding instance features across spatial and temporal dimensions, thereby preserving object identity consistency; and (2) Instance‑Masked Loss, which adaptively emphasizes foreground regions with probabilistic instance masking, reducing background noise while maintaining overall scene fidelity. By integrating these mechanisms, ConsisDrive achieves state‑of‑the‑art driving video generation quality and demonstrates significant improvements in downstream autonomous driving tasks on the nuScenes dataset. Our project page is https://shanpoyang654.github.io/ConsisDrive/page.html.
Authors:Tianxing Wu, Zheng Chen, Cirou Xu, Bowen Chai, Yong Guo, Yutong Liu, Linghe Kong, Yulun Zhang
Abstract:
One‑Step Diffusion Models have demonstrated promising capability and fast inference in video super‑resolution (VSR) for real‑world. Nevertheless, the substantial model size and high computational cost of Diffusion Transformers (DiTs) limit downstream applications. While low‑bit quantization is a common approach for model compression, the effectiveness of quantized models is challenged by the high dynamic range of input latent and diverse layer behaviors. To deal with these challenges, we introduce LSGQuant, a layer‑sensitivity guided quantizing approach for one‑step diffusion‑based real‑world VSR. Our method incorporates a Dynamic Range Adaptive Quantizer (DRAQ) to fit video token activations. Furthermore, we estimate layer sensitivity and implement a Variance‑Oriented Layer Training Strategy (VOLTS) by analyzing layer‑wise statistics in calibration. We also introduce Quantization‑Aware Optimization (QAO) to jointly refine the quantized branch and a retained high‑precision branch. Extensive experiments demonstrate that our method has nearly performance to origin model with full‑precision and significantly exceeds existing quantization techniques. Code is available at: https://github.com/zhengchen1999/LSGQuant.
Authors:Zheng Chen, Zhi Yang, Xiaoyang Liu, Weihang Zhang, Mengfan Wang, Yifan Fu, Linghe Kong, Yulun Zhang
Abstract:
Image demoiréing aims to remove structured moiré artifacts in recaptured imagery, where degradations are highly frequency‑dependent and vary across scales and directions. While recent deep networks achieve high‑quality restoration, their full‑precision designs remain costly for deployment. Binarization offers an extreme compression regime by quantizing both activations and weights to 1‑bit. Yet, it has been rarely studied for demoiréing and performs poorly when naively applied. In this work, we propose BinaryDemoire, a binarized demoiréing framework that explicitly accommodates the frequency structure of moiré degradations. First, we introduce a moiré‑aware binary gate (MABG) that extracts lightweight frequency descriptors together with activation statistics. It predicts channel‑wise gating coefficients to condition the aggregation of binary convolution responses. Second, we design a shuffle‑grouped residual adapter (SGRA) that performs structured sparse shortcut alignment. It further integrates interleaved mixing to promote information exchange across different channel partitions. Extensive experiments on four benchmarks demonstrate that the proposed BinaryDemoire surpasses current binarization methods. Code: https://github.com/zhengchen1999/BinaryDemoire.
Authors:Chihiro Nakatani, Hiroaki Kawashima, Norimichi Ukita
Abstract:
This paper proposes human‑in‑the‑loop adaptation for Group Activity Feature Learning (GAFL) without group activity annotations. This human‑in‑the‑loop adaptation is employed in a group‑activity video retrieval framework to improve its retrieval performance. Our method initially pre‑trains the GAF space based on the similarity of group activities in a self‑supervised manner, unlike prior work that classifies videos into pre‑defined group activity classes in a supervised learning manner. Our interactive fine‑tuning process updates the GAF space to allow a user to better retrieve videos similar to query videos given by the user. In this fine‑tuning, our proposed data‑efficient video selection process provides several videos, which are selected from a video database, to the user in order to manually label these videos as positive or negative. These labeled videos are used to update (i.e., fine‑tune) the GAF space, so that the positive and negative videos move closer to and farther away from the query videos through contrastive learning. Our comprehensive experimental results on two team sports datasets validate that our method significantly improves the retrieval performance. Ablation studies also demonstrate that several components in our human‑in‑the‑loop adaptation contribute to the improvement of the retrieval performance. Code: https://github.com/chihina/GAFL‑FINE‑CVIU.
Authors:Chen-Bin Feng, Youyang Sha, Longfei Liu, Yongjun Yu, Chi Man Vong, Xuanlong Yu, Xi Shen
Abstract:
In this paper, we present FSOD‑VFM: Few‑Shot Object Detectors with Vision Foundation Models, a framework that leverages vision foundation models to tackle the challenge of few‑shot object detection. FSOD‑VFM integrates three key components: a universal proposal network (UPN) for category‑agnostic bounding box generation, SAM2 for accurate mask extraction, and DINOv2 features for efficient adaptation to new object categories. Despite the strong generalization capabilities of foundation models, the bounding boxes generated by UPN often suffer from overfragmentation, covering only partial object regions and leading to numerous small, false‑positive proposals rather than accurate, complete object detections. To address this issue, we introduce a novel graph‑based confidence reweighting method. In our approach, predicted bounding boxes are modeled as nodes in a directed graph, with graph diffusion operations applied to propagate confidence scores across the network. This reweighting process refines the scores of proposals, assigning higher confidence to whole objects and lower confidence to local, fragmented parts. This strategy improves detection granularity and effectively reduces the occurrence of false‑positive bounding box proposals. Through extensive experiments on Pascal‑5^i, COCO‑20^i, and CD‑FSOD datasets, we demonstrate that our method substantially outperforms existing approaches, achieving superior performance without requiring additional training. Notably, on the challenging CD‑FSOD dataset, which spans multiple datasets and domains, our FSOD‑VFM achieves 31.6 AP in the 10‑shot setting, substantially outperforming previous training‑free methods that reach only 21.4 AP. Code is available at: https://intellindust‑ai‑lab.github.io/projects/FSOD‑VFM.
Authors:Francis Snelgar, Ming Xu, Stephen Gould, Liang Zheng, Akshay Asthana
Abstract:
3D human pose estimation from 2D images is a challenging problem due to depth ambiguity and occlusion. Because of these challenges the task is underdetermined, where there exists multiple ‑‑ possibly infinite ‑‑ poses that are plausible given the image. Despite this, many prior works assume the existence of a deterministic mapping and estimate a single pose given an image. Furthermore, methods based on machine learning require a large amount of paired 2D‑3D data to train and suffer from generalization issues to unseen scenarios. To address both of these issues, we propose a framework for pose estimation using diffusion models, which enables sampling from a probability distribution over plausible poses which are consistent with a 2D image. Our approach falls under the guidance framework for conditional generation, and guides samples from an unconditional diffusion model, trained only on 3D data, using the gradients of the heatmaps from a 2D keypoint detector. We evaluate our method on the Human 3.6M dataset under best‑of‑m multiple hypothesis evaluation, showing state‑of‑the‑art performance among methods which do not require paired 2D‑3D data for training. We additionally evaluate the generalization ability using the MPI‑INF‑3DHP and 3DPW datasets and demonstrate competitive performance. Finally, we demonstrate the flexibility of our framework by using it for novel tasks including pose generation and pose completion, without the need to train bespoke conditional models. We make code available at https://github.com/fsnelgar/diffusion_pose .
Authors:Francis Snelgar, Stephen Gould, Ming Xu, Liang Zheng, Akshay Asthana
Abstract:
Establishing correspondences between image pairs is a long studied problem in computer vision. With recent large‑scale foundation models showing strong zero‑shot performance on downstream tasks including classification and segmentation, there has been interest in using the internal feature maps of these models for the semantic correspondence task. Recent works observe that features from DINOv2 and Stable Diffusion (SD) are complementary, the former producing accurate but sparse correspondences, while the latter produces spatially consistent correspondences. As a result, current state‑of‑the‑art methods for semantic correspondence involve combining features from both models in an ensemble. While the performance of these methods is impressive, they are computationally expensive, requiring evaluating feature maps from large‑scale foundation models. In this work we take a different approach, instead replacing SD features with a superior matching algorithm which is imbued with the desirable spatial consistency property. Specifically, we replace the standard nearest neighbours matching with an optimal transport algorithm that includes a Gromov Wasserstein spatial smoothness prior. We show that we can significantly boost the performance of the DINOv2 baseline, and be competitive and sometimes surpassing state‑of‑the‑art methods using Stable Diffusion features, while being 5‑‑10x more efficient. We make code available at https://github.com/fsnelgar/semantic_matching_gwot .
Authors:Sunoh Kim, Kimin Yun, Daeho Um
Abstract:
Weakly supervised temporal video grounding aims to localize query‑relevant segments in untrimmed videos using only video‑sentence pairs, without requiring ground‑truth segment annotations that specify exact temporal boundaries. Recent approaches tackle this task by utilizing Gaussian‑based temporal proposals to represent query‑relevant segments. However, their inference strategies rely on heuristic mappings from Gaussian parameters to segment boundaries, resulting in suboptimal localization performance. To address this issue, we propose Gaussian Boundary Optimization (GBO), a novel inference framework that predicts segment boundaries by solving a principled optimization problem that balances proposal coverage and segment compactness. We derive a closed‑form solution for this problem and rigorously analyze the optimality conditions under varying penalty regimes. Beyond its theoretical foundations, GBO offers several practical advantages: it is training‑free and compatible with both single‑Gaussian and mixture‑based proposal architectures. Our experiments show that GBO significantly improves localization, achieving state‑of‑the‑art results across standard benchmarks. Extensive experiments demonstrate the efficiency and generalizability of GBO across various proposal schemes. The code is available at https://github.com/sunoh‑kim/gbo.
Authors:Zhichao Sun, Yidong Ma, Gang Liu, Yibo Chen, Xu Tang, Yao Hu, Yongchao Xu
Abstract:
Large Vision‑Language Models (LVLMs) achieve impressive performance across multiple tasks. A significant challenge, however, is their prohibitive inference cost when processing high‑resolution visual inputs. While visual token pruning has emerged as a promising solution, existing methods that primarily focus on semantic relevance often discard tokens that are crucial for spatial reasoning. We address this gap through a novel insight into \emphhow LVLMs process spatial reasoning. Specifically, we reveal that LVLMs implicitly establish visual coordinate systems through Rotary Position Embeddings (RoPE), where specific token positions serve as implicit visual coordinates (IVC tokens) that are essential for spatial reasoning. Based on this insight, we propose IVC‑Prune, a training‑free, prompt‑aware pruning strategy that retains both IVC tokens and semantically relevant foreground tokens. IVC tokens are identified by theoretically analyzing the mathematical properties of RoPE, targeting positions at which its rotation matrices approximate identity matrix or the 90^\circ rotation matrix. Foreground tokens are identified through a robust two‑stage process: semantic seed discovery followed by contextual refinement via value‑vector similarity. Extensive evaluations across four representative LVLMs and twenty diverse benchmarks show that IVC‑Prune reduces visual tokens by approximately 50% while maintaining \geq 99% of the original performance and even achieving improvements on several benchmarks. Source codes are available at https://github.com/FireRedTeam/IVC‑Prune.
Authors:Geonhui Son, Jeong Ryong Lee, Dosik Hwang
Abstract:
Generative Adversarial Networks (GANs) have made significant progress in enhancing the quality of image synthesis. Recent methods frequently leverage pretrained networks to calculate perceptual losses or utilize pretrained feature spaces. In this paper, we extend the capabilities of pretrained networks by incorporating innovative self‑supervised learning techniques and enforcing consistency between discriminators during GAN training. Our proposed method, named HP‑GAN, effectively exploits neural network priors through two primary strategies: FakeTwins and discriminator consistency. FakeTwins leverages pretrained networks as encoders to compute a self‑supervised loss and applies this through the generated images to train the generator, thereby enabling the generation of more diverse and high quality images. Additionally, we introduce a consistency mechanism between discriminators that evaluate feature maps extracted from Convolutional Neural Network (CNN) and Vision Transformer (ViT) feature networks. Discriminator consistency promotes coherent learning among discriminators and enhances training robustness by aligning their assessments of image quality. Our extensive evaluation across seventeen datasets‑including scenarios with large, small, and limited data, and covering a variety of image domains‑demonstrates that HP‑GAN consistently outperforms current state‑of‑the‑art methods in terms of Fréchet Inception Distance (FID), achieving significant improvements in image diversity and quality. Code is available at: https://github.com/higun2/HP‑GAN.
Authors:Haipeng Liu, Yang Wang, Biao Qian, Yong Rui, Meng Wang
Abstract:
Image inpainting has earned substantial progress, owing to the encoder‑and‑decoder pipeline, which is benefited from the Convolutional Neural Networks (CNNs) with convolutional downsampling to inpaint the masked regions semantically from the known regions within the encoder, coupled with an upsampling process from the decoder for final inpainting output. Recent studies intuitively identify the high‑frequency structure and low‑frequency texture to be extracted by CNNs from the encoder, and subsequently for a desirable upsampling recovery. However, the existing arts inevitably overlook the information loss for both structure and texture feature maps during the convolutional downsampling process, hence suffer from a non‑ideal upsampling output. In this paper, we systematically answer whether and how the structure and texture feature map can mutually help to alleviate the information loss during the convolutional downsampling. Given the structure and texture feature maps, we adopt the statistical normalization and denormalization strategy for the reconstruction guidance during the convolutional downsampling process. The extensive experimental results validate its advantages to the state‑of‑the‑arts over the images from low‑to‑high resolutions including 256256 and 512512, especially holds by substituting all the encoders by ours. Our code is available at https://github.com/htyjers/ConvInpaint‑TSGL
Authors:Ruojing Li, Chao Xiao, Qian Yin, Wei An, Nuo Chen, Xinyi Ying, Miao Li, Yingqian Wang
Abstract:
Infrared small targets are typically tiny and locally salient, which belong to high‑frequency components (HFCs) in images. Single‑frame infrared small target (SIRST) detection is challenging, since there are many HFCs along with targets, such as bright corners, broken clouds, and other clutters. Current learning‑based methods rely on the powerful capabilities of deep networks, but neglect explicit modeling and discriminative representation learning of various HFCs, which is important to distinguish targets from other HFCs. To address the aforementioned issues, we propose a dynamic high‑frequency convolution (DHiF) to translate the discriminative modeling process into the generation of a dynamic local filter bank. Especially, DHiF is sensitive to HFCs, owing to the dynamic parameters of its generated filters being symmetrically adjusted within a zero‑centered range according to Fourier transformation properties. Combining with standard convolution operations, DHiF can adaptively and dynamically process different HFC regions and capture their distinctive grayscale variation characteristics for discriminative representation learning. DHiF functions as a drop‑in replacement for standard convolution and can be used in arbitrary SIRST detection networks without significant decrease in computational efficiency. To validate the effectiveness of our DHiF, we conducted extensive experiments across different SIRST detection networks on real‑scene datasets. Compared to other state‑of‑the‑art convolution operations, DHiF exhibits superior detection performance with promising improvement. Codes are available at https://github.com/TinaLRJ/DHiF.
Authors:Keqi Chen, Vinkle Srivastav, Armine Vardazaryan, Cindy Rolland, Didier Mutter, Nicolas Padoy
Abstract:
Privacy preservation is a prerequisite for using video data in Operating Room (OR) research. Effective anonymization relies on the exhaustive localization of every individual; even a single missed detection necessitates extensive manual correction. However, existing approaches face two critical scalability bottlenecks: (1) they usually require manual annotations of each new clinical site for high accuracy; (2) while multi‑camera setups have been widely adopted to address single‑view ambiguity, camera calibration is typically required whenever cameras are repositioned. To address these problems, we propose a novel self‑supervised multi‑view video anonymization framework consisting of whole‑body person detection and whole‑body pose estimation, without annotation or camera calibration. Our core strategy is to enhance the single‑view detector by "retrieving" false negatives using temporal and multi‑view context, and conducting self‑supervised domain adaptation. We first run an off‑the‑shelf whole‑body person detector in each view with a low‑score threshold to gather candidate detections. Then, we retrieve the low‑score false negatives that exhibit consistency with the high‑score detections via tracking and self‑supervised uncalibrated multi‑view association. These recovered detections serve as pseudo labels to iteratively fine‑tune the whole‑body detector. Finally, we apply whole‑body pose estimation on each detected person, and fine‑tune the pose model using its own high‑score predictions. Experiments on the 4D‑OR dataset of simulated surgeries and our dataset of real surgeries show the effectiveness of our approach achieving over 97% recall. Moreover, we train a real‑time whole‑body detector using our pseudo labels, achieving comparable performance and highlighting our method's practical applicability. Code will be available at https://github.com/CAMMA‑public/OR_anonymization.
Authors:Michael Ogezi, Martin Bell, Freda Shi, Ethan Smith
Abstract:
For certain image generation tasks, vector graphics such as Scalable Vector Graphics (SVGs) offer clear benefits such as increased flexibility, size efficiency, and editing ease, but remain less explored than raster‑based approaches. A core challenge is that the numerical, geometric parameters, which make up a large proportion of SVGs, are inefficiently encoded as long sequences of tokens. This slows training, reduces accuracy, and hurts generalization. To address these problems, we propose Continuous Number Modeling (CNM), an approach that directly models numbers as first‑class, continuous values rather than discrete tokens. This formulation restores the mathematical elegance of the representation by aligning the model's inputs with the data's continuous nature, removing discretization artifacts introduced by token‑based encoding. We then train a multimodal transformer on 2 million raster‑to‑SVG samples, followed by fine‑tuning via reinforcement learning using perceptual feedback to further improve visual quality. Our approach improves training speed by over 30% while maintaining higher perceptual fidelity compared to alternative approaches. This work establishes CNM as a practical and efficient approach for high‑quality vector generation, with potential for broader applications. We make our code available http://github.com/mikeogezi/CNM.
Authors:Matteo Bastico, Pierre Onghena, David Ryckelynck, Beatriz Marcotegui, Santiago Velasco-Forero, Laurent Corté, Caroline Robine--Decourcelle, Etienne Decencière
Abstract:
Accurate identification of anatomical landmarks is crucial for various medical applications. Traditional manual landmarking is time‑consuming and prone to inter‑observer variability, while rule‑based methods are often tailored to specific geometries or limited sets of landmarks. In recent years, anatomical surfaces have been effectively represented as point clouds, which are lightweight structures composed of spatial coordinates. Following this strategy and to overcome the limitations of existing landmarking techniques, we propose Landmark Point Transformer (LmPT), a method for automatic anatomical landmark detection on point clouds that can leverage homologous bones from different species for translational research. The LmPT model incorporates a conditioning mechanism that enables adaptability to different input types to conduct cross‑species learning. We focus the evaluation of our approach on femoral landmarking using both human and newly annotated dog femurs, demonstrating its generalization and effectiveness across species. The code and dog femur dataset will be publicly available at: https://github.com/Pierreoo/LandmarkPointTransformer.
Authors:Namhoon Kim, Ashwin Pananjady, Amir Pourmorteza, Sara Fridovich-Keil
Abstract:
Background: Perfusion computed tomography (CT) images the dynamics of a contrast agent through the body over time, and is one of the highest X‑ray dose scans in medical imaging. Recently, a theoretically justified reconstruction algorithm based on a monotone variational inequality (VI) was proposed for single material polychromatic photon‑counting CT, and showed promising early results at low‑dose imaging.
Purpose: We adapt this reconstruction algorithm for perfusion CT, to reconstruct the concentration map of the contrast agent while the static background tissue is assumed known; we call our method VI‑PRISM (VI‑based PeRfusion Imaging and Single Material reconstruction). We evaluate its potential for dose‑reduced perfusion CT, using a digital phantom with water and iodine of varying concentration.
Methods: Simulated iodine concentrations range from 0.05 to 2.5 mg/ml. The simulated X‑ray source emits photons up to 100 keV, with average intensity ranging from 10^5 down to 10^2 photons per detector element. The number of tomographic projections was varied from 984 down to 8 to characterize the tradeoff in photon allocation between views and intensity.
Results: We compare VI‑PRISM against filtered back‑projection (FBP), and find that VI‑PRISM recovers iodine concentration with error below 0.4 mg/ml at all source intensity levels tested. Even with a dose reduction between 10x and 100x compared to FBP, VI‑PRISM exhibits reconstruction quality on par with FBP.
Conclusion: Across all photon budgets and angular sampling densities tested, VI‑PRISM achieved consistently lower RMSE, reduced noise, and higher SNR compared to filtered back‑projection. Even in extremely photon‑limited and sparsely sampled regimes, VI‑PRISM recovered iodine concentrations with errors below 0.4 mg/ml, showing that VI‑PRISM can support accurate and dose‑efficient perfusion imaging in photon‑counting CT.
Authors:Xiaoce Wang, Guibin Zhang, Junzhe Li, Jinzhe Tu, Chun Li, Ming Li
Abstract:
Existing GUI agent models relying on coordinate‑based one‑step visual grounding struggle with generalizing to varying input resolutions and aspect ratios. Alternatives introduce coordinate‑free strategies yet suffer from learning under severe data scarcity. To address the limitations, we propose ToolTok, a novel paradigm of multi‑step pathfinding for GUI agents, where operations are modeled as a sequence of progressive tool usage. Specifically, we devise tools aligned with human interaction habits and represent each tool using learnable token embeddings. To enable efficient embedding learning under limited supervision, ToolTok introduces a semantic anchoring mechanism that grounds each tool with semantically related concepts as natural inductive bias. To further enable a pre‑trained large language model to progressively acquire tool semantics, we construct an easy‑to‑hard curriculum consisting of three tasks: token definition question‑answering, pure text‑guided tool selection, and simplified visual pathfinding. Extensive experiments on multiple benchmarks show that ToolTok achieves superior performance among models of comparable scale (4B) and remains competitive with a substantially larger model (235B). Notably, these results are obtained using less than 1% of the training data required by other post‑training approaches. In addition, ToolTok demonstrates strong generalization across unseen scenarios. Our training & inference code is open‑source at https://github.com/ZephinueCode/ToolTok.
Authors:Tianle Gu, Kexin Huang, Lingyu Li, Ruilin Luo, Shiyang Huang, Zongqi Wang, Yujiu Yang, Yan Teng, Yingchun Wang
Abstract:
Safety moderation is pivotal for identifying harmful content. Despite the success of textual safety moderation, its multimodal counterparts remain hindered by a dual sparsity of data and supervision. Conventional reliance on binary labels lead to shortcut learning, which obscures the intrinsic classification boundaries necessary for effective multimodal discrimination. Hence, we propose a novel learning paradigm (UniMod) that transitions from sparse decision‑making to dense reasoning traces. By constructing structured trajectories encompassing evidence grounding, modality assessment, risk mapping, policy decision, and response generation, we reformulate monolithic decision tasks into a multi‑dimensional boundary learning process. This approach forces the model to ground its decision in explicit safety semantics, preventing the model from converging on superficial shortcuts. To facilitate this paradigm, we develop a multi‑head scalar reward model (UniRM). UniRM provides multi‑dimensional supervision by assigning attribute‑level scores to the response generation stage. Furthermore, we introduce specialized optimization strategies to decouple task‑specific parameters and rebalance training dynamics, effectively resolving interference between diverse objectives in multi‑task learning. Empirical results show UniMod achieves competitive textual moderation performance and sets a new multimodal benchmark using less than 40% of the training data used by leading baselines. Ablations further validate our multi‑attribute trajectory reasoning, offering an effective and efficient framework for multimodal moderation. Supplementary materials are available at \hrefhttps://trustworthylab.github.io/UniMod/project website.
Authors:Yuming Zhao, Peiyi Zhang, Oana Ignat
Abstract:
Memes are a pervasive form of online communication, yet their cultural specificity poses significant challenges for cross‑cultural adaptation. We study cross‑cultural meme transcreation, a multimodal generation task that aims to preserve communicative intent and humor while adapting culture‑specific references. We propose a hybrid transcreation framework based on vision‑language models and introduce a large‑scale bidirectional dataset of Chinese and US memes. Using both human judgments and automated evaluation, we analyze 6,315 meme pairs and assess transcreation quality across cultural directions. Our results show that current vision‑language models can perform cross‑cultural meme transcreation to a limited extent, but exhibit clear directional asymmetries: US‑Chinese transcreation consistently achieves higher quality than Chinese‑US. We further identify which aspects of humor and visual‑textual design transfer across cultures and which remain challenging, and propose an evaluation framework for assessing cross‑cultural multimodal generation. Our code and dataset are publicly available at https://github.com/AIM‑SCU/MemeXGen.
Authors:Zehong Ma, Ruihan Xu, Shiliang Zhang
Abstract:
Pixel diffusion generates images directly in pixel space, avoiding the VAE artifacts and representational bottlenecks of two‑stage latent diffusion. Recent JiT further simplifies pixel diffusion with x‑prediction, where the model predicts clean images rather than velocity. However, the standard pixel‑wise diffusion loss treats all pixels equally, spending model capacity to perceptually insignificant signals and often leading to blurry samples. We propose PixelGen, an end‑to‑end pixel diffusion framework that augments x‑prediction with perceptual supervision. Specifically, PixelGen introduces two complementary perceptual losses on top of x‑prediction: an LPIPS loss for local textures and a P‑DINO loss for global semantics. To preserve sample coverage, PixelGen further proposes a noise‑gating strategy that applies these losses only at lower‑noise timesteps. On ImageNet‑256 without classifier‑free guidance, PixelGen achieves an FID of 5.11 in 80 training epochs, surpassing the latent diffusion baselines. Moreover, PixelGen scales efficiently to text‑to‑image generation, reaching a GenEval score of 0.79 with only 6 days of training on 8xH800 GPUs. These results show that perceptual supervision substantially narrows the gap between pixel and latent diffusion while preserving a simple one‑stage pipeline. Codes are available at https://github.com/Zehong‑Ma/PixelGen.
Authors:Yinjie Wang, Tianbao Xie, Ke Shen, Mengdi Wang, Ling Yang
Abstract:
We propose RLAnything, a reinforcement learning framework that dynamically forges environment, policy, and reward models through closed‑loop optimization, amplifying learning signals and strengthening the overall RL system for any LLM or agentic scenarios. Specifically, the policy is trained with integrated feedback from step‑wise and outcome signals, while the reward model is jointly optimized via consistency feedback, which in turn further improves policy training. Moreover, our theory‑motivated automatic environment adaptation improves training for both the reward and policy models by leveraging critic feedback from each, enabling learning from experience. Empirically, each added component consistently improves the overall system, and RLAnything yields substantial gains across various representative LLM and agentic tasks, boosting Qwen3‑VL‑8B‑Thinking by 9.1% on OSWorld and Qwen2.5‑7B‑Instruct by 18.7% and 11.9% on AlfWorld and LiveBench, respectively. We also that optimized reward‑model signals outperform outcomes that rely on human labels. Code: https://github.com/Gen‑Verse/Open‑AgentRL
Authors:Abid Hassan, Tuan Ngo, Saad Shafiq, Nenad Medvidovic
Abstract:
Out‑of‑distribution (OOD) detection is critical for the safe deployment of deep neural networks. State‑of‑the‑art post‑hoc methods typically derive OOD scores from the output logits or penultimate feature vector obtained via global average pooling (GAP). We contend that this exclusive reliance on the logit or feature vector discards a rich, complementary signal: the raw channel‑wise statistics of the pre‑pooling feature map lost in GAP. In this paper, we introduce Catalyst, a post‑hoc framework that exploits these under‑explored signals. Catalyst computes an input‑dependent scaling factor (γ) on‑the‑fly from these raw statistics (e.g., mean, standard deviation, and maximum activation). This γ is then fused with the existing baseline score, multiplicatively modulating it ‑‑ an elastic scaling ‑‑ to push the ID and OOD distributions further apart. We demonstrate Catalyst is a generalizable framework: it seamlessly integrates with logit‑based methods (e.g., Energy, ReAct, SCALE) and also provides a significant boost to distance‑based detectors like KNN. As a result, Catalyst achieves substantial and consistent performance gains, reducing the average False Positive Rate by 32.87 on CIFAR‑10 (ResNet‑18), 27.94% on CIFAR‑100 (ResNet‑18), and 22.25% on ImageNet (ResNet‑50). Our results highlight the untapped potential of pre‑pooling statistics and demonstrate that Catalyst is complementary to existing OOD detection approaches. Our code is available here: https://github.com/bingabid/Catalyst
Authors:Mu Huang, Hui Wang, Kerui Ren, Linning Xu, Yunsong Zhou, Mulin Yu, Bo Dai, Jiangmiao Pang
Abstract:
Simulating deformable objects under rich interactions remains a fundamental challenge for real‑to‑sim robot manipulation, with dynamics jointly driven by environmental effects and robot actions. Existing simulators rely on predefined physics or data‑driven dynamics without robot‑conditioned control, limiting accuracy, stability, and generalization. This paper presents SoMA, a 3D Gaussian Splat simulator for soft‑body manipulation. SoMA couples deformable dynamics, environmental forces, and robot joint actions in a unified latent neural space for end‑to‑end real‑to‑sim simulation. Modeling interactions over learned Gaussian splats enables controllable, stable long‑horizon manipulation and generalization beyond observed trajectories without predefined physical models. SoMA improves resimulation accuracy and generalization on real‑world robot manipulation by 20%, enabling stable simulation of complex tasks such as long‑horizon cloth folding.
Authors:Ruiqi Wu, Xuanhua He, Meng Cheng, Tianyu Yang, Yong Zhang, Zhuoliang Kang, Xunliang Cai, Xiaoming Wei, Chunle Guo, Chongyi Li, Ming-Ming Cheng
Abstract:
We propose Infinite‑World, a robust interactive world model capable of maintaining coherent visual memory over 1000+ frames in complex real‑world environments. While existing world models can be efficiently optimized on synthetic data with perfect ground‑truth, they lack an effective training paradigm for real‑world videos due to noisy pose estimations and the scarcity of viewpoint revisits. To bridge this gap, we first introduce a Hierarchical Pose‑free Memory Compressor (HPMC) that recursively distills historical latents into a fixed‑budget representation. By jointly optimizing the compressor with the generative backbone, HPMC enables the model to autonomously anchor generations in the distant past with bounded computational cost, eliminating the need for explicit geometric priors. Second, we propose an Uncertainty‑aware Action Labeling module that discretizes continuous motion into a tri‑state logic. This strategy maximizes the utilization of raw video data while shielding the deterministic action space from being corrupted by noisy trajectories, ensuring robust action‑response learning. Furthermore, guided by insights from a pilot toy study, we employ a Revisit‑Dense Finetuning Strategy using a compact, 30‑minute dataset to efficiently activate the model's long‑range loop‑closure capabilities. Extensive experiments, including objective metrics and user studies, demonstrate that Infinite‑World achieves superior performance in visual quality, action controllability, and spatial consistency.
Authors:Yibin Wang, Yuhang Zang, Feng Han, Jiazi Bu, Yujie Zhou, Cheng Jin, Jiaqi Wang
Abstract:
Recent advancements in multimodal reward models (RMs) have significantly propelled the development of visual generation. Existing frameworks typically adopt Bradley‑Terry‑style preference modeling or leverage generative VLMs as judges, and subsequently optimize visual generation models via reinforcement learning. However, current RMs suffer from inherent limitations: they often follow a one‑size‑fits‑all paradigm that assumes a monolithic preference distribution or relies on fixed evaluation rubrics. As a result, they are insensitive to content‑specific visual cues, leading to systematic misalignment with subjective and context‑dependent human preferences. To this end, inspired by human assessment, we propose UnifiedReward‑Flex, a unified personalized reward model for vision generation that couples reward modeling with flexible and context‑adaptive reasoning. Specifically, given a prompt and the generated visual content, it first interprets the semantic intent and grounds on visual evidence, then dynamically constructs a hierarchical assessment by instantiating fine‑grained criteria under both predefined and self‑generated high‑level dimensions. Our training pipeline follows a two‑stage process: (1) we first distill structured, high‑quality reasoning traces from advanced closed‑source VLMs to bootstrap SFT, equipping the model with flexible and context‑adaptive reasoning behaviors; (2) we then perform direct preference optimization (DPO) on carefully curated preference pairs to further strengthen reasoning fidelity and discriminative alignment. To validate the effectiveness, we integrate UnifiedReward‑Flex into the GRPO framework for image and video synthesis, and extensive results demonstrate its superiority.
Authors:Wangduo Xie, Matthew B. Blaschko
Abstract:
Computed Tomography (CT) plays a vital role in inspecting the internal structures of industrial objects. Furthermore, achieving high‑quality CT reconstruction from sparse views is essential for reducing production costs. While classic implicit neural networks have shown promising results for sparse reconstruction, they are unable to leverage shape priors of objects. Motivated by the observation that numerous industrial objects exhibit rectangular structures, we propose a novel Neural Adaptive Binning (NAB) method that effectively integrates rectangular priors into the reconstruction process. Specifically, our approach first maps coordinate space into a binned vector space. This mapping relies on an innovative binning mechanism based on differences between shifted hyperbolic tangent functions, with our extension enabling rotations around the input‑plane normal vector. The resulting representations are then processed by a neural network to predict CT attenuation coefficients. This design enables end‑to‑end optimization of the encoding parameters ‑‑ including position, size, steepness, and rotation ‑‑ via gradient flow from the projection data, thus enhancing reconstruction accuracy. By adjusting the smoothness of the binning function, NAB can generalize to objects with more complex geometries. This research provides a new perspective on integrating shape priors into neural network‑based reconstruction. Extensive experiments demonstrate that NAB achieves superior performance on two industrial datasets. It also maintains robust on medical datasets when the binning function is extended to more general expression. The code is available at https://github.com/Wangduo‑Xie/NAB_CT_reconstruction.
Authors:Ziwen Xu, Chenyan Wu, Hengyu Sun, Haiwen Hong, Mengru Wang, Yunzhi Yao, Longtao Huang, Hui Xue, Shumin Deng, Zhixuan Chu, Huajun Chen, Ningyu Zhang
Abstract:
Methods for controlling large language models (LLMs), including local weight fine‑tuning, LoRA‑based adaptation, and activation‑based interventions, are often studied in isolation, obscuring their connections and making comparison difficult. In this work, we present a unified view that frames these interventions as dynamic weight updates induced by a control signal, placing them within a single conceptual framework. Building on this view, we propose a unified preference‑utility analysis that separates control effects into preference, defined as the tendency toward a target concept, and utility, defined as coherent and task‑valid generation, and measures both on a shared log‑odds scale using polarity‑paired contrastive examples. Across methods, we observe a consistent trade‑off between preference and utility: stronger control increases preference while predictably reducing utility. We further explain this behavior through an activation manifold perspective, in which control shifts representations along target‑concept directions to enhance preference, while utility declines primarily when interventions push representations off the model's valid‑generation manifold. Finally, we introduce a new steering approach SPLIT guided by this analysis that improves preference while better preserving utility. Code is available at https://github.com/zjunlp/EasyEdit/blob/main/examples/SPLIT.md.
Authors:Xiang Li, Yupeng Zheng, Pengfei Li, Yilun Chen, Ya-Qin Zhang, Wenchao Ding
Abstract:
Occupancy prediction provides critical geometric and semantic understanding for robotics but faces efficiency‑accuracy trade‑offs. Current dense methods suffer computational waste on empty voxels, while sparse query‑based approaches lack robustness in diverse and complex indoor scenes. In this paper, we propose DiScene, a novel sparse query‑based framework that leverages multi‑level distillation to achieve efficient and robust occupancy prediction. In particular, our method incorporates two key innovations: (1) a Multi‑level Consistent Knowledge Distillation strategy, which transfers hierarchical representations from large teacher models to lightweight students through coordinated alignment across four levels, including encoder‑level feature alignment, query‑level feature matching, prior‑level spatial guidance, and anchor‑level high‑confidence knowledge transfer and (2) a Teacher‑Guided Initialization policy, employing optimized parameter warm‑up to accelerate model convergence. Validated on the Occ‑Scannet benchmark, DiScene achieves 23.2 FPS without depth priors while outperforming our baseline method, OPUS, by 36.1% and even better than the depth‑enhanced version, OPUS†. With depth integration, DiScene† attains new SOTA performance, surpassing EmbodiedOcc by 3.7% with 1.62× faster inference speed. Furthermore, experiments on the Occ3D‑nuScenes benchmark and in‑the‑wild scenarios demonstrate the versatility of our approach in various environments. Code and models can be accessed at https://github.com/getterupper/DiScene.
Authors:Andrea Matteazzi, Dietmar Tutsch
Abstract:
In autonomous driving scenarios, the collected LiDAR point clouds can be challenged by occlusion and long‑range sparsity, limiting the perception of autonomous driving systems. Scene completion methods can infer the missing parts of incomplete 3D LiDAR scenes. Recent methods adopt local point‑level denoising diffusion probabilistic models, which require predicting Gaussian noise, leading to a mismatch between training and inference initial distributions. This paper introduces the first flow matching framework for 3D LiDAR scene completion, improving upon diffusion‑based methods by ensuring consistent initial distributions between training and inference. The model employs a nearest neighbor flow matching loss and a Chamfer distance loss to enhance both local structure and global coverage in the alignment of point clouds. LiFlow achieves state‑of‑the‑art performance across multiple metrics. Code: https://github.com/matteandre/LiFlow.
Authors:Harold Haodong Chen, Xinxiang Yin, Wen-Jie Shu, Hongfei Zhang, Zixin Zhang, Chenfei Liao, Litao Guo, Qifeng Chen, Ying-Cong Chen
Abstract:
Text‑to‑image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation‑‑a hallmark of human creativity. Current reasoning‑augmented paradigms most rely on explicit thought processes, where intermediate reasoning is decoded into discrete text at fixed steps with frequent image decoding and re‑encoding, leading to inefficiencies, information loss, and cognitive mismatches. To bridge this gap, we introduce LatentMorph, a novel framework that seamlessly integrates implicit latent reasoning into the T2I generation process. At its core, LatentMorph introduces four lightweight components: (i) a condenser for summarizing intermediate generation states into compact visual memory, (ii) a translator for converting latent thoughts into actionable guidance, (iii) a shaper for dynamically steering next image token predictions, and (iv) an RL‑trained invoker for adaptively determining when to invoke reasoning. By performing reasoning entirely in continuous latent spaces, LatentMorph avoids the bottlenecks of explicit reasoning and enables more adaptive self‑refinement. Extensive experiments demonstrate that LatentMorph (I) enhances the base model Janus‑Pro by 16% on GenEval and 25% on T2I‑CompBench; (II) outperforms explicit paradigms (e.g., TwiG) by 15% and 11% on abstract reasoning tasks like WISE and IPV‑Txt, (III) while reducing inference time by 44% and token consumption by 51%; and (IV) exhibits 71% cognitive alignment with human intuition on reasoning invocation.
Authors:Ruiqi Liu, Manni Cui, Ziheng Qin, Zhiyuan Yan, Ruoxin Chen, Yi Han, Zhiheng Li, Junkai Chen, ZhiJin Chen, Kaiqing Lin, Jialiang Shen, Lubin Weng, Jing Dong, Yan Wang, Shu Wu
Abstract:
High‑fidelity generative models have narrowed the perceptual gap between synthetic and real images, posing serious threats to media security. Most existing AI‑generated image (AIGI) detectors rely on artifact‑based classification and struggle to generalize to evolving generative traces. In contrast, human judgment relies on stable real‑world regularities, with deviations from the human cognitive manifold serving as a more generalizable signal of forgery. Motivated by this insight, we reformulate AIGI detection as a Reference‑Comparison problem that verifies consistency with the real‑image manifold rather than fitting specific forgery cues. We propose MIRROR (Manifold Ideal Reference ReconstructOR), a framework that explicitly encodes reality priors using a learnable discrete memory bank. MIRROR projects an input into a manifold‑consistent ideal reference via sparse linear combination, and uses the resulting residuals as robust detection signals. To evaluate whether detectors reach the "superhuman crossover" required to replace human experts, we introduce the Human‑AIGI benchmark, featuring a psychophysically curated human‑imperceptible subset. Across 14 benchmarks, MIRROR consistently outperforms prior methods, achieving gains of 2.1% on six standard benchmarks and 8.1% on seven in‑the‑wild benchmarks. On Human‑AIGI, MIRROR reaches 89.6% accuracy across 27 generators, surpassing both lay users and visual experts, and further approaching the human perceptual limit as pretrained backbones scale. The code is publicly available at: https://github.com/349793927/MIRROR
Authors:Ziqiao Weng, Jiancheng Yang, Kangxian Xie, Bo Zhou, Weidong Cai
Abstract:
Pulmonary trees extracted from CT images frequently exhibit topological incompleteness, such as missing or disconnected branches, which substantially degrades downstream anatomical analysis and limits the applicability of existing pulmonary tree modeling pipelines. Current approaches typically rely on dense volumetric processing or explicit graph reasoning, leading to limited efficiency and reduced robustness under realistic structural corruption. We propose TopoField, a topology‑aware implicit modeling framework that treats topology repair as a first‑class modeling problem and enables unified multi‑task inference for pulmonary tree analysis. TopoField represents pulmonary anatomy using sparse surface and skeleton point clouds and learns a continuous implicit field that supports topology repair without relying on complete or explicit disconnection annotations, by training on synthetically introduced structural disruptions over already incomplete trees. Building upon the repaired implicit representation, anatomical labeling and lung segment reconstruction are jointly inferred through task‑specific implicit functions within a single forward pass.Extensive experiments on the Lung3D+ dataset demonstrate that TopoField consistently improves topological completeness and achieves accurate anatomical labeling and lung segment reconstruction under challenging incomplete scenarios. Owing to its implicit formulation, TopoField attains high computational efficiency, completing all tasks in just over one second per case, highlighting its practicality for large‑scale and time‑sensitive clinical applications. Code and data will be available at https://github.com/HINTLab/TopoField.
Authors:Yu Zeng, Wenxuan Huang, Zhen Fang, Shuang Chen, Yufan Shen, Yishuo Cai, Xiaoman Wang, Zhenfei Yin, Lin Chen, Zehui Chen, Shiting Huang, Yiming Zhao, Xu Tang, Yao Hu, Philip Torr, Wanli Ouyang, Shaosheng Cao
Abstract:
Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision‑DeepResearch systems that use search engines for complex visual‑textual fact‑finding. However, evaluating these visual and textual search abilities is still difficult, and existing benchmarks have two major limitations. First, existing benchmarks are not visual search‑centric: answers that should require visual search are often leaked through cross‑textual cues in the text questions or can be inferred from the prior world knowledge in current MLLMs. Second, overly idealized evaluation scenario: On the image‑search side, the required information can often be obtained via near‑exact matching against the full image, while the text‑search side is overly direct and insufficiently challenging. To address these issues, we construct the Vision‑DeepResearch benchmark (VDR‑Bench) comprising 2,000 VQA instances. All questions are created via a careful, multi‑stage curation pipeline and rigorous expert review, designed to assess the behavior of Vision‑DeepResearch systems under realistic real‑world conditions. Moreover, to address the insufficient visual retrieval capabilities of current MLLMs, we propose a simple multi‑round cropped‑search workflow. This strategy is shown to effectively improve model performance in realistic visual retrieval scenarios. Overall, our results provide practical guidance for the design of future multimodal deep‑research systems. The code will be released in https://github.com/Osilly/Vision‑DeepResearch.
Authors:Wen-Jie Shu, Xuerui Qiu, Rui-Jie Zhu, Harold Haodong Chen, Yexin Liu, Harry Yang
Abstract:
Recent advances in visual reasoning have leveraged vision transformers to tackle the ARC‑AGI benchmark. However, we argue that the feed‑forward architecture, where computational depth is strictly bound to parameter size, falls short of capturing the iterative, algorithmic nature of human induction. In this work, we propose a recursive architecture called Loop‑ViT, which decouples reasoning depth from model capacity through weight‑tied recurrence. Loop‑ViT iterates a weight‑tied Hybrid Block, combining local convolutions and global attention, to form a latent chain of thought. Crucially, we introduce a parameter‑free Dynamic Exit mechanism based on predictive entropy: the model halts inference when its internal state ``crystallizes" into a low‑uncertainty attractor. Empirical results on the ARC‑AGI‑1 benchmark validate this perspective: our 18M model achieves 65.8% accuracy, outperforming massive 73M‑parameter ensembles. These findings demonstrate that adaptive iterative computation offers a far more efficient scaling axis for visual reasoning than simply increasing network width. The code is available at https://github.com/WenjieShu/LoopViT.
Authors:Zhongqian Fu, Tianyi Zhao, Kai Han, Hang Zhou, Xinghao Chen, Yunhe Wang
Abstract:
World models learn an internal representation of environment dynamics, enabling agents to simulate and reason about future states within a compact latent space for tasks such as planning, prediction, and inference. However, running world models rely on hevay computational cost and memory footprint, making model quantization essential for efficient deployment. To date, the effects of post‑training quantization (PTQ) on world models remain largely unexamined. In this work, we present a systematic empirical study of world model quantization using DINO‑WM as a representative case, evaluating diverse PTQ methods under both weight‑only and joint weight‑activation settings. We conduct extensive experiments on different visual planning tasks across a wide range of bit‑widths, quantization granularities, and planning horizons up to 50 iterations. Our results show that quantization effects in world models extend beyond standard accuracy and bit‑width trade‑offs: group‑wise weight quantization can stabilize low‑bit rollouts, activation quantization granularity yields inconsistent benefits, and quantization sensitivity is highly asymmetric between encoder and predictor modules. Moreover, aggressive low‑bit quantization significantly degrades the alignment between the planning objective and task success, leading to failures that cannot be remedied by additional optimization. These findings reveal distinct quantization‑induced failure modes in world model‑based planning and provide practical guidance for deploying quantized world models under strict computational constraints. The code will be available at https://github.com/huawei‑noah/noah‑research/tree/master/QuantWM.
Authors:FSVideo Team, Qingyu Chen, Zhiyuan Fang, Haibin Huang, Xinwei Huang, Tong Jin, Minxuan Lin, Bo Liu, Celong Liu, Chongyang Ma, Xing Mei, Xiaohui Shen, Yaojie Shen, Fuwen Tan, Angtian Wang, Xiao Yang, Yiding Yang, Jiamin Yuan, Lingxi Zhang, Yuxin Zhang
Abstract:
We introduce FSVideo, a fast speed transformer‑based image‑to‑video (I2V) diffusion framework. We build our framework on the following key components: 1.) a new video autoencoder with highly‑compressed latent space (64×64×4 spatial‑temporal downsampling ratio), achieving competitive reconstruction quality; 2.) a diffusion transformer (DIT) architecture with a new layer memory design to enhance inter‑layer information flow and context reuse within DIT, and 3.) a multi‑resolution generation strategy via a few‑step DIT upsampler to increase video fidelity. Our final model, which contains a 14B DIT base model and a 14B DIT upsampler, achieves competitive performance against other popular open‑source models, while being an order of magnitude faster. We discuss our model design as well as training strategies in this report.
Authors:Nikola Cenikj, Özgün Turgut, Alexander Müller, Alexander Steger, Jan Kehrer, Marcus Brugger, Daniel Rueckert, Eimo Martens, Philip Müller
Abstract:
Coronary artery stenosis is a leading cause of cardiovascular disease, diagnosed by analyzing the coronary arteries from multiple angiography views. Although numerous deep‑learning models have been proposed for stenosis detection from a single angiography view, their performance heavily relies on expensive view‑level annotations, which are often not readily available in hospital systems. Moreover, these models fail to capture the temporal dynamics and dependencies among multiple views, which are crucial for clinical diagnosis. To address this, we propose SegmentMIL, a transformer‑based multi‑view multiple‑instance learning framework for patient‑level stenosis classification. Trained on a real‑world clinical dataset, using patient‑level supervision and without any view‑level annotations, SegmentMIL jointly predicts the presence of stenosis and localizes the affected anatomical region, distinguishing between the right and left coronary arteries and their respective segments. SegmentMIL obtains high performance on internal and external evaluations and outperforms both view‑level models and classical MIL baselines, underscoring its potential as a clinically viable and scalable solution for coronary stenosis diagnosis. Our code is available at https://github.com/NikolaCenic/mil‑stenosis.
Authors:Shuo Lu, Haohan Wang, Wei Feng, Weizhen Wang, Shen Zhang, Yaoyu Li, Ao Ma, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Bing Zhan, Yuan Xu, Huizai Yao, Yongcan Yu, Chenyang Si, Jian Liang
Abstract:
Advertising image generation has increasingly focused on online metrics like Click‑Through Rate (CTR), yet existing approaches adopt a ``one‑size‑fits‑all" strategy that optimizes for overall CTR while neglecting preference diversity among user groups. This leads to suboptimal performance for specific groups, limiting targeted marketing effectiveness. To bridge this gap, we present One Size, Many Fits (OSMF), a unified framework that aligns diverse group‑wise click preferences in large‑scale advertising image generation. OSMF begins with product‑aware adaptive grouping, which dynamically organizes users based on their attributes and product characteristics, representing each group with rich collective preference features. Building on these groups, preference‑conditioned image generation employs a Group‑aware Multimodal Large Language Model (G‑MLLM) to generate tailored images for each group. The G‑MLLM is pre‑trained to simultaneously comprehend group features and generate advertising images. Subsequently, we fine‑tune the G‑MLLM using our proposed Group‑DPO for group‑wise preference alignment, which effectively enhances each group's CTR on the generated images. To further advance this field, we introduce the Grouped Advertising Image Preference Dataset (GAIP), the first large‑scale public dataset of group‑wise image preferences, including around 600K groups built from 40M users. Extensive experiments demonstrate that our framework achieves the state‑of‑the‑art performance in both offline and online settings. Our code and datasets will be released at https://github.com/JD‑GenX/OSMF.
Authors:Bing He, Jingnan Gao, Yunuo Chen, Ning Cao, Gang Chen, Zhengxue Cheng, Li Song, Wenjun Zhang
Abstract:
Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches leverage generalizable models to generate 3D scenes using 3D Gaussian Splatting (3DGS) primitive. However, they often fail to produce continuous surfaces and instead yield discrete, color‑biased point clouds that appear plausible at normal resolution but reveal severe artifacts under close‑up views. To address this issue, we present SurfSplat, a feedforward framework based on 2D Gaussian Splatting (2DGS) primitive, which provides stronger anisotropy and higher geometric precision. By incorporating a surface continuity prior and a forced alpha blending strategy, SurfSplat reconstructs coherent geometry together with faithful textures. Furthermore, we introduce High‑Resolution Rendering Consistency (HRRC), a new evaluation metric designed to evaluate high‑resolution reconstruction quality. Extensive experiments on RealEstate10K, DL3DV, and ScanNet demonstrate that SurfSplat consistently outperforms prior methods on both standard metrics and HRRC, establishing a robust solution for high‑fidelity 3D reconstruction from sparse inputs. Project page: https://hebing‑sjtu.github.io/SurfSplat‑website/
Authors:Hongwei Yan, Guanglong Sun, Kanglei Zhou, Qian Li, Liyuan Wang, Yi Zhong
Abstract:
General continual learning (GCL) challenges intelligent systems to learn from single‑pass, non‑stationary data streams without clear task boundaries. While recent advances in continual parameter‑efficient tuning (PET) of pretrained models show promise, they typically rely on multiple training epochs and explicit task cues, limiting their effectiveness in GCL scenarios. Moreover, existing methods often lack targeted design and fail to address two fundamental challenges in continual PET: how to allocate expert parameters to evolving data distributions, and how to improve their representational capacity under limited supervision. Inspired by the fruit fly's hierarchical memory system characterized by sparse expansion and modular ensembles, we propose FlyPrompt, a brain‑inspired framework that decomposes GCL into two subproblems: expert routing and expert competence improvement. FlyPrompt introduces a randomly expanded analytic router for instance‑level expert activation and a temporal ensemble of output heads to dynamically adapt decision boundaries over time. Extensive theoretical and empirical evaluations demonstrate FlyPrompt's superior performance, achieving up to 11.23%, 12.43%, and 7.62% gains over state‑of‑the‑art baselines on CIFAR‑100, ImageNet‑R, and CUB‑200, respectively. Our source code is available at https://github.com/AnAppleCore/FlyGCL.
Authors:Muli Yang, Gabriel James Goenawan, Henan Wang, Huaiyuan Qin, Chenghao Xu, Yanhua Yang, Fen Fang, Ying Sun, Joo-Hwee Lim, Hongyuan Zhu
Abstract:
Despite being trained on balanced datasets, existing AI‑generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training. Specifically, models tend to overfit to superficial artifacts that do not generalize well across different generation methods, leading to a misaligned decision threshold when faced with test‑time distribution shift. To address this, we propose a theoretically grounded post‑hoc calibration framework based on Bayesian decision theory. In particular, we introduce a learnable scalar correction to the model's logits, optimized on a small validation set from the target distribution while keeping the backbone frozen. This parametric adjustment compensates for distributional shift in model output, realigning the decision boundary even without requiring ground‑truth labels. Experiments on challenging benchmarks show that our approach significantly improves robustness without retraining, offering a lightweight and principled solution for reliable and adaptive AI‑generated image detection in the open world. Code is available at https://github.com/muliyangm/AIGI‑Det‑Calib.
Authors:Yuliang Zhan, Jian Li, Wenbing Huang, Wenbing Huang, Yang Liu, Hao Sun
Abstract:
Deep learning has demonstrated remarkable capabilities in simulating complex dynamic systems. However, existing methods require known physical properties as supervision or inputs, limiting their applicability under unknown conditions. To explore this challenge, we introduce Cloth Dynamics Grounding (CDG), a novel scenario for unsupervised learning of cloth dynamics from multi‑view visual observations. We further propose Cloth Dynamics Splatting (CloDS), an unsupervised dynamic learning framework designed for CDG. CloDS adopts a three‑stage pipeline that first performs video‑to‑geometry grounding and then trains a dynamics model on the grounded meshes. To cope with large non‑linear deformations and severe self‑occlusions during grounding, we introduce a dual‑position opacity modulation that supports bidirectional mapping between 2D observations and 3D geometry via mesh‑based Gaussian splatting in video‑to‑geometry grounding stage. It jointly considers the absolute and relative position of Gaussian components. Comprehensive experimental evaluations demonstrate that CloDS effectively learns cloth dynamics from visual data while maintaining strong generalization capabilities for unseen configurations. Our code is available at https://github.com/whynot‑zyl/CloDS. Visualization results are available at https://github.com/whynot‑zyl/CloDS_video.%\footnoteAs in this example.
Authors:Dvir Samuel, Issar Tzachor, Matan Levy, Micahel Green, Gal Chechik, Rami Ben-Ari
Abstract:
Autoregressive video diffusion models enable streaming generation, opening the door to long‑form synthesis, video world models, and interactive neural game engines. However, their core attention layers become a major bottleneck at inference time: as generation progresses, the KV cache grows, causing both increasing latency and escalating GPU memory, which in turn restricts usable temporal context and harms long‑range consistency. In this work, we study redundancy in autoregressive video diffusion and identify three persistent sources: near‑duplicate cached keys across frames, slowly evolving (largely semantic) queries/keys that make many attention computations redundant, and cross‑attention over long prompts where only a small subset of tokens matters per frame. Building on these observations, we propose a unified, training‑free attention framework for autoregressive diffusion: TempCache compresses the KV cache via temporal correspondence to bound cache growth; AnnCA accelerates cross‑attention by selecting frame‑relevant prompt tokens using fast approximate nearest neighbor (ANN) matching; and AnnSA sparsifies self‑attention by restricting each query to semantically matched keys, also using a lightweight ANN. Together, these modules reduce attention, compute, and memory and are compatible with existing autoregressive diffusion backbones and world models. Experiments demonstrate up to x5‑‑x10 end‑to‑end speedups while preserving near‑identical visual quality and, crucially, maintaining stable throughput and nearly constant peak GPU memory usage over long rollouts, where prior methods progressively slow down and suffer from increasing memory usage.
Authors:Shicheng Yin, Kaixuan Yin, Weixing Chen, Yang Liu, Guanbin Li, Liang Lin
Abstract:
World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformerbased models significantly hinders real‑time deployment. To address this efficiency‑performance bottleneck, we introduce DDP‑WM, a novel world model centered on the principle of Disentangled Dynamics Prediction (DDP). We hypothesize that latent state evolution in observed scenes is heterogeneous and can be decomposed into sparse primary dynamics driven by physical interactions and secondary context‑driven background updates. DDP‑WM realizes this decomposition through an architecture that integrates efficient historical processing with dynamic localization to isolate primary dynamics. By employing a crossattention mechanism for background updates, the framework optimizes resource allocation and provides a smooth optimization landscape for planners. Extensive experiments demonstrate that DDP‑WM achieves significant efficiency and performance across diverse tasks, including navigation, precise tabletop manipulation, and complex deformable or multi‑body interactions. Specifically, on the challenging Push‑T task, DDP‑WM achieves an approximately 9 times inference speedup and improves the MPC success rate from 90% to98% compared to state‑of‑the‑art dense models. The results establish a promising path for developing efficient, high‑fidelity world models. Codes is available at https://hcplab‑sysu.github.io/DDP‑WM/.
Authors:Hao Zhang, Yanping Zha, Zizhuo Li, Meiqi Gong, Jiayi Ma
Abstract:
This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi‑modal image fusion under harsh conditions when only visible imaging sensors are available. To achieve this goal, we propose a novel concept of single‑image fusion, which extends conventional data‑level fusion to the knowledge level. Specifically, we develop MagicFuse, a novel single image fusion framework capable of deriving a comprehensive cross‑spectral scene representation from a single low‑quality visible image. MagicFuse first introduces an intra‑spectral knowledge reinforcement branch and a cross‑spectral knowledge generation branch based on the diffusion models. They mine scene information obscured in the visible spectrum and learn thermal radiation distribution patterns transferred to the infrared spectrum, respectively. Building on them, we design a multi‑domain knowledge fusion branch that integrates the probabilistic noise from the diffusion streams of these two branches, from which a cross‑spectral scene representation can be obtained through successive sampling. Then, we impose both visual and semantic constraints to ensure that this scene representation can satisfy human observation while supporting downstream semantic decision‑making. Extensive experiments show that our MagicFuse achieves visual and semantic representation performance comparable to or even better than state‑of‑the‑art fusion methods with multi‑modal inputs, despite relying solely on a single degraded visible image. The code is publicly available at https://github.com/zhayanping/MagicFuse.
Authors:Tushar Anand, Maheswar Bora, Antitza Dantcheva, Abhijit Das
Abstract:
In this work, we propose a novel Mamba block DenVisCoM, as well as a novel hybrid architecture specifically tailored for accurate and real‑time estimation of optical flow and disparity estimation. Given that such multi‑view geometry and motion tasks are fundamentally related, we propose a unified architecture to tackle them jointly. Specifically, the proposed hybrid architecture is based on DenVisCoM and a Transformer‑based attention block that efficiently addresses real‑time inference, memory footprint, and accuracy at the same time for joint estimation of motion and 3D dense perception tasks. We extensively analyze the benchmark trade‑off of accuracy and real‑time processing on a large number of datasets. Our experimental results and related analysis suggest that our proposed model can accurately estimate optical flow and disparity estimation in real time. All models and associated code are available at https://github.com/vimstereo/DenVisCoM.
Authors:Yinchao Ma, Qiang Zhou, Zhibin Wang, Xianing Chen, Hanqing Yang, Jun Song, Bo Zheng
Abstract:
Video large language models have demonstrated remarkable capabilities in video understanding tasks. However, the redundancy of video tokens introduces significant computational overhead during inference, limiting their practical deployment. Many compression algorithms are proposed to prioritize retaining features with the highest attention scores to minimize perturbations in attention computations. However, the correlation between attention scores and their actual contribution to correct answers remains ambiguous. To address the above limitation, we propose a novel Contribution‑aware token Compression algorithm for VIDeo understanding (CaCoVID) that explicitly optimizes the token selection policy based on the contribution of tokens to correct predictions. First, we introduce a reinforcement learning‑based framework that optimizes a policy network to select video token combinations with the greatest contribution to correct predictions. This paradigm shifts the focus from passive token preservation to active discovery of optimal compressed token combinations. Secondly, we propose a combinatorial policy optimization algorithm with online combination space sampling, which dramatically reduces the exploration space for video token combinations and accelerates the convergence speed of policy optimization. Extensive experiments on diverse video understanding benchmarks demonstrate the effectiveness of CaCoVID. Codes are available at https://github.com/LivingFutureLab/CaCoVID.
Authors:Tianyu Yang, Chenwei He, Xiangzhao Hao, Tianyue Wang, Jiarui Guo, Haiyun Guo, Leigang Qu, Jinqiao Wang, Tat-Seng Chua
Abstract:
Composed Image Retrieval (CIR) aims to retrieve target images based on a hybrid query comprising a reference image and a modification text. Early dual‑tower Vision‑Language Models (VLMs) struggle with cross‑modality compositional reasoning required for this task. While adapting generative Multimodal Large Language Models (MLLMs) for retrieval offers a promising direction, we identify that this strategy overlooks a fundamental issue: compressing a generative MLLM into a single‑embedding discriminative retriever triggers a paradigm conflict, which leads to Capability Degradation ‑ the deterioration of native fine‑grained reasoning after retrieval adaptation. To address this challenge, we propose ReCALL, a model‑agnostic framework that follows a diagnose‑generate‑refine pipeline: First, we diagnose cognitive blind spots of the retriever via self‑guided informative instance mining. Next, we generate corrective instructions and triplets by prompting the foundation MLLM and conduct quality control with VQA‑based consistency filtering. Finally, we refine the retriever through continual training on these triplets with a grouped contrastive scheme, thereby internalizing fine‑grained visual‑semantic distinctions and realigning the discriminative embedding space of retriever with intrinsic compositional reasoning within the MLLM. Extensive experiments on CIRR and FashionIQ show that ReCALL consistently recalibrates degraded capabilities and achieves state‑of‑the‑art performance. Code is available at https://github.com/RemRico/Recall.
Authors:Xinyuan Zhao, Yihang Wu, Ahmad Chaddad, Tareef Daqqaq, Reem Kateb
Abstract:
While deep learning models like Vision Transformer (ViT) have achieved significant advances, they typically require large datasets. With data privacy regulations, access to many original datasets is restricted, especially medical images. Federated learning (FL) addresses this challenge by enabling global model aggregation without data exchange. However, the heterogeneity of the data and the class imbalance that exist in local clients pose challenges for the generalization of the model. This study proposes a FL framework leveraging a dynamic adaptive focal loss (DAFL) and a client‑aware aggregation strategy for local training. Specifically, we design a dynamic class imbalance coefficient that adjusts based on each client's sample distribution and class data distribution, ensuring minority classes receive sufficient attention and preventing sparse data from being ignored. To address client heterogeneity, a weighted aggregation strategy is adopted, which adapts to data size and characteristics to better capture inter‑client variations. The classification results on three public datasets (ISIC, Ocular Disease and RSNA‑ICH) show that the proposed framework outperforms DenseNet121, ResNet50, ViT‑S/16, ViT‑L/32, FedCLIP, Swin Transformer, CoAtNet, and MixNet in most cases, with accuracy improvements ranging from 0.98% to 41.69%. Ablation studies on the imbalanced ISIC dataset validate the effectiveness of the proposed loss function and aggregation strategy compared to traditional loss functions and other FL approaches. The codes can be found at: https://github.com/AIPMLab/ViT‑FLDAF.
Authors:Yiwen Jia, Hao Wei, Yanhui Zhou, Chenyang Ge
Abstract:
Diffusion‑based image compression methods have achieved notable progress, delivering high perceptual quality at low bitrates. However, their practical deployment is hindered by significant inference latency and heavy computational overhead, primarily due to the large number of denoising steps required during decoding. To address this problem, we propose a diffusion‑based image compression method that requires only a single‑step diffusion process, significantly improving inference speed. To enhance the perceptual quality of reconstructed images, we introduce a discriminator that operates on compact feature representations instead of raw pixels, leveraging the fact that features better capture high‑level texture and structural details. Experimental results show that our method delivers comparable compression performance while offering a 46× faster inference speed compared to recent diffusion‑based approaches. The source code and models are available at https://github.com/cheesejiang/OSDiff.
Authors:Shuai Liu, Siheng Ren, Xiaoyao Zhu, Quanmin Liang, Zefeng Li, Qiang Li, Xin Hu, Kai Huang
Abstract:
Achieving reliable and efficient planning in complex driving environments requires a model that can reason over the scene's geometry, appearance, and dynamics. We present UniDWM, a unified driving world model that advances autonomous driving through multifaceted representation learning. UniDWM constructs a structure‑ and dynamic‑aware latent world representation that serves as a physically grounded state space, enabling consistent reasoning across perception, prediction, and planning. Specifically, a joint reconstruction pathway learns to recover the scene's structure, including geometry and visual texture, while a collaborative generation framework leverages a conditional diffusion transformer to forecast future world evolution within the latent space. Furthermore, we show that our UniDWM can be deemed as a variation of VAE, which provides theoretical guidance for the multifaceted representation learning. Extensive experiments demonstrate the effectiveness of UniDWM in trajectory planning, 4D reconstruction and generation, highlighting the potential of multifaceted world representations as a foundation for unified driving intelligence. The code will be publicly available at https://github.com/Say2L/UniDWM.
Authors:Minwoo Jung, Nived Chebrolu, Lucas Carvalho de Lima, Haedam Oh, Maurice Fallon, Ayoung Kim
Abstract:
Reliable localization is crucial for navigation in forests, where GPS is often degraded and LiDAR measurements are repetitive, occluded, and structurally complex. These conditions weaken the assumptions of traditional urban‑centric localization methods, which assume that consistent features arise from unique structural patterns, necessitating forest‑centric solutions to achieve robustness in these environments. To address these challenges, we propose TreeLoc, a LiDAR‑based global localization framework for forests that handles place recognition and 6‑DoF pose estimation. We represent scenes using tree stems and their Diameter at Breast Height (DBH), which are aligned to a common reference frame via their axes and summarized using the tree distribution histogram (TDH) for coarse matching, followed by fine matching with a 2D triangle descriptor. Finally, pose estimation is achieved through a two‑step geometric verification. On diverse forest benchmarks, TreeLoc outperforms baselines, achieving precise localization. Ablation studies validate the contribution of each component. We also propose applications for long‑term forest management using descriptors from a compact global tree database. TreeLoc is open‑sourced for the robotics community at https://github.com/minwoo0611/TreeLoc.
Authors:Tal Grutman, Carmel Shinar, Tali Ilovitsh
Abstract:
Ultrasound is the most widely used medical imaging modality, yet the images it produces are fundamentally unique, arising from tissue‑dependent scattering, reflection, and speed‑of‑sound variations that produce a constrained set of characteristic textures that differ markedly from natural‑image statistics. These acoustically driven patterns make ultrasound challenging for algorithms originally designed for natural images. To bridge this gap, the field has increasingly turned to foundation models, hoping to leverage their generalization capabilities. However, these models often falter in ultrasound applications because they are not designed for ultrasound physics, they are merely trained on ultrasound data. Therefore, it is essential to integrate ultrasound‑specific domain knowledge into established learning frameworks. We achieve this by reformulating self‑supervised learning as a texture‑analysis problem, introducing texture ultrasound semantic analysis (TUSA). Using TUSA, models learn to leverage highly scalable contrastive methods to extract true domain‑specific representations directly from simple B‑mode images. We train a TUSA model on a combination of open‑source, simulated, and in vivo data. The latent space is compared to several larger foundation models, demonstrating that our approach gives TUSA models better generalizability for difficult downstream tasks on unique online datasets as well as a clinical eye dataset collected for this study. Our model achieves higher accuracy in detecting COVID (70%), spinal hematoma (100%) and vitreous hemorrhage (97%) and correlates more closely with quantitative parameters like liver steatosis (r = 0.83), ejection fraction (r = 0.63), and oxygen saturation (r = 0.38). We open‑source the model weights and training script: https://github.com/talg2324/tusa
Authors:Soumyaroop Nandi, Prem Natarajan
Abstract:
We propose BioTamperNet, a novel framework for detecting duplicated regions in tampered biomedical images, leveraging affinity‑guided attention inspired by State Space Model (SSM) approximations. Existing forensic models, primarily trained on natural images, often underperform on biomedical data where subtle manipulations can compromise experimental validity. To address this, BioTamperNet introduces an affinity‑guided self‑attention module to capture intra‑image similarities and an affinity‑guided cross‑attention module to model cross‑image correspondences. Our design integrates lightweight SSM‑inspired linear attention mechanisms to enable efficient, fine‑grained localization. Trained end‑to‑end, BioTamperNet simultaneously identifies tampered regions and their source counterparts. Extensive experiments on the benchmark bio‑forensic datasets demonstrate significant improvements over competitive baselines in accurately detecting duplicated regions. Code ‑ https://github.com/SoumyaroopNandi/BioTamperNet
Authors:Christoffer Koo Øhrstrøm, Rafael I. Cabral Muchacho, Yifei Dong, Filippos Moumtzidellis, Ronja Güldenring, Florian T. Pokorny, Lazaros Nalpantidis
Abstract:
We propose Parabolic Position Encoding (PaPE), a parabola‑based position encoding for vision modalities in attention‑based architectures. Given a set of vision tokens‑such as from videos, event camera streams, images, or point clouds‑our objective is to encode their positions while accounting for the characteristics of vision modalities. Prior works have largely extended position encodings from 1D‑sequences in language to nD‑structures in vision, but only with partial account of vision characteristics. We address this gap by designing PaPE from principles distilled from prior work: translation invariance, rotation invariance (PaPE‑RI), distance decay, directionality, and context awareness. Extrapolation experiments on ImageNet‑1K show how PaPE extrapolates remarkably well, improving in absolute terms by up to 10.5% over the next‑best encoding. Generality experiments on 8 datasets across 4 modalities show that PaPE is a general vision position encoding, as PaPE matches the best baseline on 5 datasets and exceeds all on 2 datasets. Code is available at https://github.com/DTU‑PAS/parabolic‑position‑encoding.
Authors:Fu-Yun Wang, Han Zhang, Michael Gharbi, Hongsheng Li, Taesung Park
Abstract:
Flow matching models (FMs) have revolutionized text‑to‑image (T2I) generation, with reinforcement learning (RL) serving as a critical post‑training strategy for alignment with reward objectives. In this research, we show that current RL pipelines for FMs suffer from two underappreciated yet important limitations: sample inefficiency due to insufficient generation diversity, and pronounced prompt overfitting, where models memorize specific training formulations and exhibit dramatic performance collapse when evaluated on semantically equivalent but stylistically varied prompts. We present PromptRL (Prompt Matters in RL for Flow‑Based Image Generation), a framework that incorporates language models (LMs) as trainable prompt refinement agents directly within the flow‑based RL optimization loop. This design yields two complementary benefits: rapid development of sophisticated prompt rewriting capabilities and, critically, a synergistic training regime that reshapes the optimization dynamics. PromptRL achieves state‑of‑the‑art performance across multiple benchmarks, obtaining scores of 0.97 on GenEval, 0.98 on OCR accuracy, and 24.05 on PickScore.
Furthermore, we validate the effectiveness of our RL approach on large‑scale image editing models, improving the EditReward of FLUX.1‑Kontext from 1.19 to 1.43 with only 0.06 million rollouts, surpassing Gemini 2.5 Flash Image (also known as Nano Banana), which scores 1.37, and achieving comparable performance with ReasonNet (1.44), which relied on fine‑grained data annotations along with a complex multi‑stage training. Our extensive experiments empirically demonstrate that PromptRL consistently achieves higher performance ceilings while requiring over 2× fewer rollouts compared to naive flow‑only RL. Our code is available at https://github.com/G‑U‑N/UniRL.
Authors:Yan Ma, Weiyu Zhang, Tianle Li, Linge Du, Xuyang Shen, Pengfei Liu
Abstract:
Vision tool‑use reinforcement learning (RL) can equip vision language models with visual operators such as crop‑and‑zoom and achieves strong performance gains, yet it remains unclear whether these gains are driven by improvements in tool use or evolving intrinsic capabilities. We introduce MED (Measure‑‑Explain‑‑Diagnose), a coarse‑to‑fine framework that disentangles intrinsic capability changes from tool‑induced effects, decomposes the tool‑induced performance difference into gain and harm terms, and probes the mechanisms driving their evolution. Across checkpoint‑level analyses in the crop‑and‑zoom setting on two VLMs with different tool priors and six benchmarks, we find that improvements are dominated by intrinsic learning, while tool‑use RL mainly reduces tool‑induced harm (e.g., fewer call‑induced errors and weaker tool schema interference) and yields limited progress in tool‑based correction of intrinsic failures. Overall, in the crop‑and‑zoom setting studied here, current vision tool‑use RL learns to coexist safely with tools rather than master them.
Authors:Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris
Abstract:
Maintaining visual and semantic consistency across frames is a key challenge in text‑to‑image storytelling. Existing training‑free methods, such as One‑Prompt‑One‑Story, concatenate all prompts into a single sequence, which often induces strong embedding correlation and leads to color leakage, background blending, and identity drift. We propose DeCorStory, a training‑free inference‑time framework that explicitly reduces inter‑frame semantic interference. DeCorStory applies Gram‑Schmidt prompt embedding decorrelation to orthogonalize frame‑level semantics, followed by singular value reweighting to strengthen prompt‑specific information and identity‑preserving cross‑attention to stabilize character identity during diffusion. The method requires no model modification or fine‑tuning and can be seamlessly integrated into existing diffusion pipelines. Experiments demonstrate consistent improvements in prompt‑image alignment, identity consistency, and visual diversity, achieving state‑of‑the‑art performance among training‑free baselines. Code is available at: https://github.com/YuZhenyuLindy/DeCorStory
Authors:Ayushman Sarkar, Zhenyu Yu, Wei Tang, Chu Chen, Kangning Cui, Mohd Yamani Idna Idris
Abstract:
Large multimodal models have enabled one‑click storybook generation, where users provide a short description and receive a multi‑page illustrated story. However, the underlying story state, such as characters, world settings, and page‑level objects, remains implicit, making edits coarse‑grained and often breaking visual consistency. We present StoryState, an agent‑based orchestration layer that introduces an explicit and editable story state on top of training‑free text‑to‑image generation. StoryState represents each story as a structured object composed of a character sheet, global settings, and per‑page scene constraints, and employs a small set of LLM agents to maintain this state and derive 1Prompt1Story‑style prompts for generation and editing. Operating purely through prompts, StoryState is model‑agnostic and compatible with diverse generation backends. System‑level experiments on multi‑page editing tasks show that StoryState enables localized page edits, improves cross‑page consistency, and reduces unintended changes, interaction turns, and editing time compared to 1Prompt1Story, while approaching the one‑shot consistency of Gemini Storybook. Code is available at https://github.com/YuZhenyuLindy/StoryState
Authors:Ayushman Sarkar, Zhenyu Yu, Chu Chen, Wei Tang, Kangning Cui, Mohd Yamani Idna Idris
Abstract:
Generating coherent visual stories requires maintaining subject identity across multiple images while preserving frame‑specific semantics. Recent training‑free methods concatenate identity and frame prompts into a unified representation, but this often introduces inter‑frame semantic interference that weakens identity preservation in complex stories. We propose ReDiStory, a training‑free framework that improves multi‑frame story generation via inference‑time prompt embedding reorganization. ReDiStory explicitly decomposes text embeddings into identity‑related and frame‑specific components, then decorrelates frame embeddings by suppressing shared directions across frames. This reduces cross‑frame interference without modifying diffusion parameters or requiring additional supervision. Under identical diffusion backbones and inference settings, ReDiStory improves identity consistency while maintaining prompt fidelity. Experiments on the ConsiStory+ benchmark show consistent gains over 1Prompt1Story on multiple identity consistency metrics. Code is available at: https://github.com/YuZhenyuLindy/ReDiStory
Authors:Zeran Ke, Bin Tan, Gui-Song Xia, Yujun Shen, Nan Xue
Abstract:
3D line mapping from multi‑view RGB images provides a compact and structured visual representation of scenes. We study the problem from a physical and topological perspective: a 3D line most naturally emerges as the edge of a finite 3D planar patch. We present LiP‑Map, a line‑plane joint optimization framework that explicitly models learnable line and planar primitives. This coupling enables accurate and detailed 3D line mapping while maintaining strong efficiency (typically completing a reconstruction in 3 to 5 minutes per scene). LiP‑Map pioneers the integration of planar topology into 3D line mapping, not by imposing pairwise coplanarity constraints but by explicitly constructing interactions between plane and line primitives, thus offering a principled route toward structured reconstruction in man‑made environments. On more than 100 scenes from ScanNetV2, ScanNet++, Hypersim, 7Scenes, and Tanks\&Temple, LiP‑Map improves both accuracy and completeness over state‑of‑the‑art methods. Beyond line mapping quality, LiP‑Map significantly advances line‑assisted visual localization, establishing strong performance on 7Scenes. Our code is released at https://github.com/calmke/LiPMAP for reproducible research.
Authors:Xianhui Zhang, Chengyu Xie, Linxia Zhu, Yonghui Yang, Weixiang Zhao, Zifeng Cheng, Cong Wang, Fei Shen, Tat-Seng Chua
Abstract:
Multilingual safety remains significantly imbalanced, leaving non‑high‑resource (NHR) languages vulnerable compared to robust high‑resource (HR) ones. Moreover, the neural mechanisms driving safety alignment remain unclear despite observed cross‑lingual representation transfer.
In this paper, we find that LLMs contain a set of cross‑lingual shared safety neurons (SS‑Neurons), a remarkably small yet critical neuronal subset that jointly regulates safety behavior across languages.
We first identify monolingual safety neurons (MS‑Neurons) and validate their causal role in safety refusal behavior through targeted activation and suppression.
Our cross‑lingual analyses then identify SS‑Neurons as the subset of MS‑Neurons shared between HR and NHR languages, serving as a bridge to transfer safety capabilities from HR to NHR domains.
We observe that suppressing these neurons causes concurrent safety drops across NHR languages, whereas reinforcing them improves cross‑lingual defensive consistency.
Building on these insights, we propose a simple neuron‑oriented training strategy that targets SS‑Neurons based on language resource distribution and model architecture. Experiments demonstrate that fine‑tuning this tiny neuronal subset outperforms state‑of‑the‑art methods, significantly enhancing NHR safety while maintaining the model's general capabilities.
The code and dataset will be available athttps://github.com/1518630367/SS‑Neuron‑Expansion.
Authors:Xun Zhang, Kaicheng Yang, Hongliang Lu, Haotong Qin, Yong Guo, Yulun Zhang
Abstract:
Recently, Diffusion Transformers (DiTs) have emerged in Real‑World Image Super‑Resolution (Real‑ISR) to generate high‑quality textures, yet their heavy inference burden hinders real‑world deployment. While Post‑Training Quantization (PTQ) is a promising solution for acceleration, existing methods in super‑resolution mostly focus on U‑Net architectures, whereas generic DiT quantization is typically designed for text‑to‑image tasks. Directly applying these methods to DiT‑based super‑resolution models leads to severe degradation of local textures. Therefore, we propose Q‑DiT4SR, the first PTQ framework specifically tailored for DiT‑based Real‑ISR. We propose H‑SVD, a hierarchical SVD that integrates a global low‑rank branch with a local block‑wise rank‑1 branch under a matched parameter budget. We further propose Variance‑aware Spatio‑Temporal Mixed Precision: VaSMP allocates cross‑layer weight bit‑widths in a data‑free manner based on rate‑distortion theory, while VaTMP schedules intra‑layer activation precision across diffusion timesteps via dynamic programming (DP) with minimal calibration. Experiments on multiple real‑world datasets demonstrate that our Q‑DiT4SR achieves SOTA performance under both W4A6 and W4A4 settings. Notably, the W4A4 quantization configuration reduces model size by 5.8× and computational operations by 6.14×. Our code and models will be available at https://github.com/xunzhang1128/Q‑DiT4SR.
Authors:Qishuai Wen, Zhiyuan Huang, Xianghan Meng, Wei He, Chun-Guang Li
Abstract:
The vanilla self‑attention mechanism in Transformers can be viewed as a two‑layer fast‑weight MLP, whose weights are dynamically induced by inputs and whose hidden dimension is equal to the sequence length N. As the context extends, the expressive capacity of such an N‑width MLP increases, but it becomes unscalable for extremely long sequences. Recently, this fast‑weight perspective has motivated the Mixture‑of‑Experts (MoE) attention mechanism, which partitions the sequence into rigid blocks, treats them as fast‑weight experts, and sparsely routes the tokens to them. In this paper, we elevate this perspective to a unifying framework for efficient attention mechanisms, interpreting them as making fast weights scalable through either routing or compression, and organizing them into a five‑dimensional taxonomy. Then, we propose Mixture‑of‑Top‑k Attention (MiTA), which employs a small set of landmark queries to gather top‑k attended key‑value pairs as query‑aware and deformable routed experts, while compressing the N‑width MLP into a narrower shared expert. Consequently, our MiTA improves the flexibility of prior MoE attention from rigid to deformable fast‑weight experts, as well as the scalability of prior top‑k attention from query‑specific set to reusable top‑k set. We conduct extensive experiments on vision tasks showing the superior effectiveness and efficiency of our MiTA, and also uncovering intriguing properties such as an emergent token‑pruning effect and easy generalization from standard attention. Code is available at https://github.com/QishuaiWen/MiTA.
Authors:Marco Chen, Xianbiao Qi, Yelin He, Jiaquan Ye, Rong Xiao
Abstract:
In this work, we revisit Transformer optimization through the lens of second‑order geometry and establish a direct connection between architectural design, activation scale, the Hessian matrix, and the maximum tolerable learning rate. We introduce a simple normalization strategy, termed SimpleNorm, which stabilizes intermediate activation scales by construction. Then, by analyzing the Hessian of the loss with respect to network activations, we theoretically show that SimpleNorm significantly reduces the spectral norm of the Hessian, thereby permitting larger stable learning rates. We validate our theoretical findings through extensive experiments on large GPT models at parameter scales 1B, 1.4B, 7B and 8B. Empirically, SimpleGPT, our SimpleNorm‑based network, tolerates learning rates 3×‑10× larger than standard convention, consistently demonstrates strong optimization stability, and achieves substantially better performance than well‑established baselines. Specifically, when training 7B‑scale models for 60K steps, SimpleGPT achieves a training loss that is 0.08 lower than that of LLaMA2 with QKNorm, reducing the loss from 2.290 to 2.208. Our source code will be released at https://github.com/Ocram7/SimpleGPT.
Authors:Hao Chen, Tao Han, Jie Zhang, Song Guo, Fenghua Ling, Lei Bai
Abstract:
Long‑term weather forecasting is critical for socioeconomic planning and disaster preparedness. While recent approaches employ finetuning to extend prediction horizons, they remain constrained by the issues of catastrophic forgetting, error accumulation, and high training overhead. To address these limitations, we present a novel pipeline across pretraining, finetuning and forecasting to enhance long‑context modeling while reducing computational overhead. First, we introduce an Efficient Multi‑scale Transformer (EMFormer) to extract multi‑scale features through a single convolution in both training and inference. Based on the new architecture, we further employ an accumulative context finetuning to improve temporal consistency without degrading short‑term accuracy. Additionally, we propose a composite loss that dynamically balances different terms via a sinusoidal weighting, thereby adaptively guiding the optimization trajectory throughout pretraining and finetuning. Experiments show that our approach achieves strong performance in weather forecasting and extreme event prediction, substantially improving long‑term forecast accuracy. Moreover, EMFormer demonstrates strong generalization on vision benchmarks (ImageNet‑1K and ADE20K) while delivering a 5.69x speedup over conventional multi‑scale modules. Code: https://github.com/chenhao‑zju/emformer
Authors:Zhihao Chen, Yiyuan Ge, Ziyang Wang
Abstract:
Diffusion‑based visuomotor policies excel at modeling action distributions but are inference‑inefficient, since recursively denoising from noise to policy requires many steps and heavy UNet backbones, which hinders deployment on resource‑constrained robots. Flow matching alleviates the sampling burden by learning a one‑step vector field, yet prior implementations still inherit large UNet‑style architectures. In this work, we present KAN‑We‑Flow, a flow‑matching policy that draws on recent advances in Receptance Weighted Key Value (RWKV) and Kolmogorov‑Arnold Networks (KAN) from vision to build a lightweight and highly expressive backbone for 3D manipulation. Concretely, we introduce an RWKV‑KAN block: an RWKV first performs efficient time/channel mixing to propagate task context, and a subsequent GroupKAN layer applies learnable spline‑based, groupwise functional mappings to perform feature‑wise nonlinear calibration of the action mapping on RWKV outputs. Moreover, we introduce an Action Consistency Regularization (ACR), a lightweight auxiliary loss that enforces alignment between predicted action trajectories and expert demonstrations via Euler extrapolation, providing additional supervision to stabilize training and improve policy precision. Without resorting to large UNets, our design reduces parameters by 86.8%, maintains fast runtime, and achieves state‑of‑the‑art success rates on Adroit, Meta‑World, and DexArt benchmarks. Our project page can be viewed in \hrefhttps://zhihaochen‑2003.github.io/KAN‑We‑Flow.github.io/\textcolorredlink
Authors:Haopeng Li, Shitong Shao, Wenliang Zhong, Zikai Zhou, Lichen Bai, Hui Xiong, Zeke Xie
Abstract:
Diffusion Transformers are fundamental for video and image generation, but their efficiency is bottlenecked by the quadratic complexity of attention. While block sparse attention accelerates computation by attending only critical key‑value blocks, it suffers from degradation at high sparsity by discarding context. In this work, we discover that attention scores of non‑critical blocks exhibit distributional stability, allowing them to be approximated accurately and efficiently rather than discarded, which is essentially important for sparse attention design. Motivated by this key insight, we propose PISA, a training‑free Piecewise Sparse Attention that covers the full attention span with sub‑quadratic complexity. Unlike the conventional keep‑or‑drop paradigm that directly drop the non‑critical block information, PISA introduces a novel exact‑or‑approximate strategy: it maintains exact computation for critical blocks while efficiently approximating the remainder through block‑wise Taylor expansion. This design allows PISA to serve as a faithful proxy to full attention, effectively bridging the gap between speed and quality. Experimental results demonstrate that PISA achieves 1.91 times and 2.57 times speedups on Wan2.1‑14B and Hunyuan‑Video, respectively, while consistently maintaining the highest quality among sparse attention methods. Notably, even for image generation on FLUX, PISA achieves a 1.2 times acceleration without compromising visual quality. Code is available at: https://github.com/xie‑lab‑ml/piecewise‑sparse‑attention.
Authors:Bo Deng, Yitong Tang, Jiake Li, Yuxin Huang, Li Wang, Yu Zhang, Yufei Zhan, Hua Lu, Xiaoshen Zhang, Jieyun Bai
Abstract:
Ultrasound (US) imaging exhibits substantial heterogeneity across anatomical structures and acquisition protocols, posing significant challenges to the development of generalizable analysis models. Most existing methods are task‑specific, limiting their suitability as clinically deployable foundation models. To address this limitation, the Foundation Model Challenge for Ultrasound Image Analysis (FM\_UIA~2026) introduces a large‑scale multi‑task benchmark comprising 27 subtasks across segmentation, classification, detection, and regression. In this paper, we present the official baseline for FM\_UIA~2026 based on a unified Multi‑Head Multi‑Task Learning (MH‑MTL) framework that supports all tasks within a single shared network. The model employs an ImageNet‑pretrained EfficientNet‑‑B4 backbone for robust feature extraction, combined with a Feature Pyramid Network (FPN) to capture multi‑scale contextual information. A task‑specific routing strategy enables global tasks to leverage high‑level semantic features, while dense prediction tasks exploit spatially detailed FPN representations. Training incorporates a composite loss with task‑adaptive learning rate scaling and a cosine annealing schedule. Validation results demonstrate the feasibility and robustness of this unified design, establishing a strong and extensible baseline for ultrasound foundation model research. The code and dataset are publicly available at \hrefhttps://github.com/lijiake2408/Foundation‑Model‑Challenge‑for‑Ultrasound‑Image‑AnalysisGitHub.
Authors:Guangshuo Qin, Zhiteng Li, Zheng Chen, Weihang Zhang, Linghe Kong, Yulun Zhang
Abstract:
Mixture‑of‑Experts(MoE) Vision‑Language Models (VLMs) offer remarkable performance but incur prohibitive memory and computational costs, making compression essential. Post‑Training Quantization (PTQ) is an effective training‑free technique to address the massive memory and computation overhead. Existing quantization paradigms fall short as they are oblivious to two critical forms of heterogeneity: the inherent discrepancy between vision and language tokens, and the non‑uniform contribution of different experts. To bridge this gap, we propose Visual Expert Quantization (VEQ), a dual‑aware quantization framework designed to simultaneously accommodate cross‑modal differences and heterogeneity between experts. Specifically, VEQ incorporates 1)Modality‑expert‑aware Quantization, which utilizes expert activation frequency to prioritize error minimization for pivotal experts, and 2)Modality‑affinity‑aware Quantization, which constructs an enhanced Hessian matrix by integrating token‑expert affinity with modality information to guide the calibration process. Extensive experiments across diverse benchmarks verify that VEQ consistently outperforms state‑of‑the‑art baselines. Specifically, under the W3A16 configuration, our method achieves significant average accuracy gains of 2.04% on Kimi‑VL and 3.09% on Qwen3‑VL compared to the previous SOTA quantization methods, demonstrating superior robustness across various multimodal tasks. Our code will be available at https://github.com/guangshuoqin/VEQ.
Authors:Kaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma, Sarah Erfani, Christopher Leckie, Hanxun Huang
Abstract:
Vision‑language models (VLMs) extend large language models (LLMs) with vision encoders, enabling text generation conditioned on both images and text. However, this multimodal integration expands the attack surface by exposing the model to image‑based jailbreaks crafted to induce harmful responses. Existing gradient‑based jailbreak methods transfer poorly, as adversarial patterns overfit to a single white‑box surrogate and fail to generalise to black‑box models. In this work, we propose Universal and transferable jailbreak (UltraBreak), a framework that constrains adversarial patterns through transformations and regularisation in the vision space, while relaxing textual targets through semantic‑based objectives. By defining its loss in the textual embedding space of the target LLM, UltraBreak discovers universal adversarial patterns that generalise across diverse jailbreak objectives. This combination of vision‑level regularisation and semantically guided textual supervision mitigates surrogate overfitting and enables strong transferability across both models and attack targets. Extensive experiments show that UltraBreak consistently outperforms prior jailbreak methods. Further analysis reveals why earlier approaches fail to transfer, highlighting that smoothing the loss landscape via semantic objectives is crucial for enabling universal and transferable jailbreaks. The code is publicly available in our \hrefhttps://github.com/kaiyuanCui/UltraBreakGitHub repository.
Authors:Nick DiSanto, Ehsan Khodapanah Aghdam, Han Liu, Jacob Watson, Yuankai K. Tao, Hao Li, Ipek Oguz
Abstract:
Handheld Optical Coherence Tomography Angiography (OCTA) enables noninvasive retinal imaging in uncooperative or pediatric subjects, but is highly susceptible to motion artifacts that severely degrade volumetric image quality. Sudden motion during 3D acquisition can lead to unsampled retinal regions across entire B‑scans (cross‑sectional slices), resulting in blank bands in en face projections. We propose VAMOS‑OCTA, a deep learning framework for inpainting motion‑corrupted B‑scans using vessel‑aware multi‑axis supervision. We employ a 2.5D U‑Net architecture that takes a stack of neighboring B‑scans as input to reconstruct a corrupted center B‑scan, guided by a novel Vessel‑Aware Multi‑Axis Orthogonal Supervision (VAMOS) loss. This loss combines vessel‑weighted intensity reconstruction with axial and lateral projection consistency, encouraging vascular continuity in native B‑scans and across orthogonal planes. Unlike prior work that focuses primarily on restoring the en face MIP, VAMOS‑OCTA jointly enhances both cross‑sectional B‑scan sharpness and volumetric projection accuracy, even under severe motion corruptions. We trained our model on both synthetic and real‑world corrupted volumes and evaluated its performance using both perceptual quality and pixel‑wise accuracy metrics. VAMOS‑OCTA consistently outperforms prior methods, producing reconstructions with sharp capillaries, restored vessel continuity, and clean en face projections. These results demonstrate that multi‑axis supervision offers a powerful constraint for restoring motion‑degraded 3D OCTA data. Our source code is available at https://github.com/MedICL‑VU/VAMOS‑OCTA.
Authors:Meng Luo, Bobo Li, Shanqing Xu, Shize Zhang, Qiuchan Chen, Menglu Han, Wenhao Chen, Yanxiang Huang, Hao Fei, Mong-Li Lee, Wynne Hsu
Abstract:
Despite rapid progress in multimodal large language models (MLLMs), their capability for deep emotional understanding remains limited. We argue that genuine affective intelligence requires explicit modeling of Theory of Mind (ToM), the cognitive substrate from which emotions arise. To this end, we introduce HitEmotion, a ToM‑grounded hierarchical benchmark that diagnoses capability breakpoints across increasing levels of cognitive depth. Second, we propose a ToM‑guided reasoning chain that tracks mental states and calibrates cross‑modal evidence to achieve faithful emotional reasoning. We further introduce TMPO, a reinforcement learning method that uses intermediate mental states as process‑level supervision to guide and strengthen model reasoning. Extensive experiments show that HitEmotion exposes deep emotional reasoning deficits in state‑of‑the‑art models, especially on cognitively demanding tasks. In evaluation, the ToM‑guided reasoning chain and TMPO improve end‑task accuracy and yield more faithful, more coherent rationales. In conclusion, our work provides the research community with a practical toolkit for evaluating and enhancing the cognition‑based emotional understanding capabilities of MLLMs. Our dataset and code are available at: https://HitEmotion.github.io/.
Authors:I-Chun Arthur Liu, Krzysztof Choromanski, Sandy Huang, Connor Schenck
Abstract:
Leveraging pre‑trained 2D image representations in behavior cloning policies has achieved great success and has become a standard approach for robotic manipulation. However, such representations fail to capture the 3D spatial information about objects and scenes that is essential for precise manipulation. In this work, we introduce Contrastive Learning for 3D Multi‑View Action‑Conditioned Robotic Manipulation Pretraining (CLAMP), a novel 3D pre‑training framework that utilizes point clouds and robot actions. From the merged point cloud computed from RGB‑D images and camera extrinsics, we re‑render multi‑view four‑channel image observations with depth and 3D coordinates, including dynamic wrist views, to provide clearer views of target objects for high‑precision manipulation tasks. The pre‑trained encoders learn to associate the 3D geometric and positional information of objects with robot action patterns via contrastive learning on large‑scale simulated robot trajectories. During encoder pre‑training, we pre‑train a Diffusion Policy to initialize the policy weights for fine‑tuning, which is essential for improving fine‑tuning sample efficiency and performance. After pre‑training, we fine‑tune the policy on a limited amount of task demonstrations using the learned image and action representations. We demonstrate that this pre‑training and fine‑tuning design substantially improves learning efficiency and policy performance on unseen tasks. Furthermore, we show that CLAMP outperforms state‑of‑the‑art baselines across six simulated tasks and five real‑world tasks. The project website and videos can be found at https://clamp3d.github.io/CLAMP/.
Authors:Alicja Polowczyk, Agnieszka Polowczyk, Piotr Borycki, Joanna Waczyńska, Jacek Tabor, Przemysław Spurek
Abstract:
Despite impressive results from recent text‑to‑image models like FLUX, visual and anatomical artifacts remain a significant hurdle for practical and professional use. Existing methods for artifact reduction, typically work in a post‑hoc manner, consequently failing to intervene effectively during the core image formation process. Notably, current techniques require problematic and invasive modifications to the model weights, or depend on a computationally expensive and time‑consuming process of regional refinement. To address these limitations, we propose DIAMOND, a training‑free method that applies trajectory correction to mitigate artifacts during inference. By reconstructing an estimate of the clean sample at every step of the generative trajectory, DIAMOND actively steers the generation process away from latent states that lead to artifacts. Furthermore, we extend the proposed method to standard Diffusion Models, demonstrating that DIAMOND provides a robust, zero‑shot path to high‑fidelity, artifact‑free image synthesis without the need for additional training or weight modifications in modern generative architectures. Code is available at https://gmum.github.io/DIAMOND/
Authors:Mingwei Li, Hehe Fan, Yi Yang
Abstract:
Monocular normal estimation for transparent objects is critical for laboratory automation, yet it remains challenging due to complex light refraction and reflection. These optical properties often lead to catastrophic failures in conventional depth and normal sensors, hindering the deployment of embodied AI in scientific environments. We propose TransNormal, a novel framework that adapts pre‑trained diffusion priors for single‑step normal regression. To handle the lack of texture in transparent surfaces, TransNormal integrates dense visual semantics from DINOv3 via a cross‑attention mechanism, providing strong geometric cues. Furthermore, we employ a multi‑task learning objective and wavelet‑based regularization to ensure the preservation of fine‑grained structural details. To support this task, we introduce TransNormal‑Synthetic, a physics‑based dataset with high‑fidelity normal maps for transparent labware. Extensive experiments demonstrate that TransNormal significantly outperforms state‑of‑the‑art methods: on the ClearGrasp benchmark, it reduces mean error by 24.4% and improves 11.25° accuracy by 22.8%; on ClearPose, it achieves a 15.2% reduction in mean error. The code and dataset will be made publicly available at https://longxiang‑ai.github.io/TransNormal.
Authors:Xianzhe Fan, Shengliang Deng, Xiaoyang Wu, Yuxiang Lu, Zhuoling Li, Mi Yan, Yujia Zhang, Zhizheng Zhang, He Wang, Hengshuang Zhao
Abstract:
Existing Vision‑Language‑Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D information to enhance VLA capabilities? We conduct a pilot study across different observation spaces and visual representations. The results show that explicitly lifting visual input into point clouds yields representations that better complement their corresponding 2D representations. To address the challenges of (1) scarce 3D data and (2) the domain gap induced by cross‑environment differences and depth‑scale biases, we propose Any3D‑VLA. It unifies the simulator, sensor, and model‑estimated point clouds within a training pipeline, constructs diverse inputs, and learns domain‑agnostic 3D representations that are fused with the corresponding 2D representations. Simulation and real‑world experiments demonstrate Any3D‑VLA's advantages in improving performance and mitigating the domain gap. Our project homepage is available at https://xianzhefan.github.io/Any3D‑VLA.github.io.
Authors:Ruikui Wang, Jinheng Feng, Lang Tian, Huaishao Luo, Chaochao Li, Liangbo Zhou, Huan Zhang, Youzheng Wu, Xiaodong He
Abstract:
Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions, particularly when the prompts involve complex elements including large full‑body movement, dynamic camera trajectory, background transitions, or human‑object interactions. To break out this limitation, we present JoyAvatar, a framework capable of generating long duration avatar videos, featuring two key technical innovations. Firstly, we introduce a twin‑teacher enhanced training algorithm that enables the model to transfer inherent text‑controllability from the foundation model while simultaneously learning audio‑visual synchronization. Secondly, during training, we dynamically modulate the strength of multi‑modal conditions (e.g., audio and text) based on the distinct denoising timestep, aiming to mitigate conflicts between the heterogeneous conditioning signals. These two key designs serve to substantially expand the avatar model's capacity to generate natural, temporally coherent full‑body motions and dynamic camera movements as well as preserve the basic avatar capabilities, such as accurate lip‑sync and identity consistency. GSB evaluation results demonstrate that our JoyStreamer model outperforms the state‑of‑the‑art models such as Omnihuman‑1.5 and KlingAvatar 2.0. Moreover, our approach enables complex applications including multi‑person dialogues and non‑human subjects role‑playing. Some video samples are provided on https://joystreamer.github.io/.
Authors:Vivek Madhavaram, Vartika Sengar, Arkadipta De, Charu Sharma
Abstract:
Scene understanding and reasoning has been a fundamental problem in 3D computer vision, requiring models to identify objects, their properties, and spatial or comparative relationships among the objects. Existing approaches enable this by creating scene graphs using multiple inputs such as 2D images, depth maps, object labels, and annotated relationships from specific reference view. However, these methods often struggle with generalization and produce inaccurate spatial relationships like "left/right", which become inconsistent across different viewpoints. To address these limitations, we propose Viewpoint‑Invariant Zero‑shot scene graph generation for 3D scene Reasoning (VIZOR). VIZOR is a training‑free, end‑to‑end framework that constructs dense, viewpoint‑invariant 3D scene graphs directly from raw 3D scenes. The generated scene graph is unambiguous, as spatial relationships are defined relative to each object's front‑facing direction, making them consistent regardless of the reference view. Furthermore, it infers open‑vocabulary relationships that describe spatial and proximity relationships among scene objects without requiring annotated training data. We conduct extensive quantitative and qualitative evaluations to assess the effectiveness of VIZOR in scene graph generation and downstream tasks, such as query‑based object grounding. VIZOR outperforms state‑of‑the‑art methods, showing clear improvements in scene graph generation and achieving 22% and 4.81% gains in zero‑shot grounding accuracy on the Replica and Nr3D datasets, respectively.
Authors:Yian Zhao, Rushi Ye, Ruochong Zheng, Zesen Cheng, Chaoran Feng, Jiashu Yang, Pengchong Qiao, Chang Liu, Jie Chen
Abstract:
3D style transfer refers to the artistic stylization of 3D assets based on reference style images. Recently, 3DGS‑based stylization methods have drawn considerable attention, primarily due to their markedly enhanced training and rendering speeds. However, a vital challenge for 3D style transfer is to strike a balance between the content and the patterns and colors of the style. Although the existing methods strive to achieve relatively balanced outcomes, the fixed‑output paradigm struggles to adapt to the diverse content‑style balance requirements from different users. In this work, we introduce a creative intensity‑tunable 3D style transfer paradigm, dubbed Tune‑Your‑Style, which allows users to flexibly adjust the style intensity injected into the scene to match their desired content‑style balance, thus enhancing the customizability of 3D style transfer. To achieve this goal, we first introduce Gaussian neurons to explicitly model the style intensity and parameterize a learnable style tuner to achieve intensity‑tunable style injection. To facilitate the learning of tunable stylization, we further propose the tunable stylization guidance, which obtains multi‑view consistent stylized views from diffusion models through cross‑view style alignment, and then employs a two‑stage optimization strategy to provide stable and efficient guidance by modulating the balance between full‑style guidance from the stylized views and zero‑style guidance from the initial rendering. Extensive experiments demonstrate that our method not only delivers visually appealing results, but also exhibits flexible customizability for 3D style transfer. Project page is available at https://zhao‑yian.github.io/TuneStyle.
Authors:JiaKui Hu, Zhengjian Yao, Lujia Jin, Yanye Lu
Abstract:
Universal image restoration is a critical task in low‑level vision, requiring the model to remove various degradations from low‑quality images to produce clean images with rich detail. The challenges lie in sampling the distribution of high‑quality images and adjusting the outputs on the basis of the degradation. This paper presents a novel approach, Bridging Degradation discrimination and Generation (BDG), which aims to address these challenges concurrently. First, we propose the Multi‑Angle and multi‑Scale Gray Level Co‑occurrence Matrix (MAS‑GLCM) and demonstrate its effectiveness in performing fine‑grained discrimination of degradation types and levels. Subsequently, we divide the diffusion training process into three distinct stages: generation, bridging, and restoration. The objective is to preserve the diffusion model's capability of restoring rich textures while simultaneously integrating the discriminative information from the MAS‑GLCM into the restoration process. This enhances its proficiency in addressing multi‑task and multi‑degraded scenarios. Without changing the architecture, BDG achieves significant performance gains in all‑in‑one restoration and real‑world super‑resolution tasks, primarily evidenced by substantial improvements in fidelity without compromising perceptual quality. The code and pretrained models are provided in https://github.com/MILab‑PKU/BDG.
Authors:Xingyu Luo, Yidong Cai, Jie Liu, Jie Tang, Gangshan Wu, Limin Wang
Abstract:
Vision‑language tracking has gained increasing attention in many scenarios. This task simultaneously deals with visual and linguistic information to localize objects in videos. Despite its growing utility, the development of vision‑language tracking methods remains in its early stage. Current vision‑language trackers usually employ Transformer architectures for interactive integration of template, search, and text features. However, persistent challenges about low‑semantic images including prevalent image blurriness, low resolution and so on, may compromise model performance through degraded cross‑modal understanding. To solve this problem, language assistance is usually used to deal with the obstacles posed by low‑semantic images. However, due to the existing gap between current textual and visual features, direct concatenation and fusion of these features may have limited effectiveness. To address these challenges, we introduce a pioneering Generative Language‑AssisteD tracking model, GLAD, which utilizes diffusion models for the generative multi‑modal fusion of text description and template image to bolster compatibility between language and image and enhance template image semantic information. Our approach demonstrates notable improvements over the existing fusion paradigms. Blurry and semantically ambiguous template images can be restored to improve multi‑modal features in the generative fusion paradigm. Experiments show that our method establishes a new state‑of‑the‑art on multiple benchmarks and achieves an impressive inference speed. The code and models will be released at: https://github.com/Confetti‑lxy/GLAD
Authors:Wenbin Xing, Quanxing Zha, Lizheng Zu, Mengran Li, Ming Li, Junchi Yan
Abstract:
Current research on video hallucination mitigation primarily focuses on isolated error types, leaving compositional hallucinations, arising from incorrect reasoning over multiple interacting spatial and temporal factors largely underexplored. We introduce OmniVCHall, a benchmark designed to systematically evaluate both isolated and compositional hallucinations in video multimodal large language models (VLLMs). OmniVCHall spans diverse video domains, introduces a novel camera‑based hallucination type, and defines a fine‑grained taxonomy, together with adversarial answer options (e.g., "All are correct" and "None of the above") to prevent shortcut reasoning. The evaluations of 39 representative VLLMs reveal that even advanced models (e.g., Qwen3‑VL and GPT‑5) exhibit substantial performance degradation. We propose TriCD, a contrastive decoding framework with a triple‑pathway calibration mechanism. An adaptive perturbation controller dynamically selects distracting operations to construct negative video variants, while a saliency‑guided enhancement module adaptively reinforces grounded token‑wise visual evidences. These components are optimized via reinforcement learning to encourage precise decision‑making under compositional hallucination settings. Experimental results show that TriCD consistently improves performance across two representative backbones, achieving an average accuracy improvement of over 10%. The data and code can be find at https://github.com/BMRETURN/OmniVCHall.
Authors:Daoxuan Zhang, Ping Chen, Xiaobo Xia, Xiu Su, Ruichen Zhen, Jianqiang Xiao, Shuo Yang
Abstract:
Aerial Object Goal Navigation, a challenging frontier in Embodied AI, requires an Unmanned Aerial Vehicle (UAV) agent to autonomously explore, reason, and identify a specific target using only visual perception and language description. However, existing methods struggle with the memorization of complex spatial representations in aerial environments, reliable and interpretable action decision‑making, and inefficient exploration and information gathering. To address these challenges, we introduce APEX (Aerial Parallel Explorer), a novel hierarchical agent designed for efficient exploration and target acquisition in complex aerial settings. APEX is built upon a modular, three‑part architecture: 1) Dynamic Spatio‑Semantic Mapping Memory, which leverages the zero‑shot capability of a Vision‑Language Model (VLM) to dynamically construct high‑resolution 3D Attraction, Exploration, and Obstacle maps, serving as an interpretable memory mechanism. 2) Action Decision Module, trained with reinforcement learning, which translates this rich spatial understanding into a fine‑grained and robust control policy. 3) Target Grounding Module, which employs an open‑vocabulary detector to achieve definitive and generalizable target identification. All these components are integrated into a hierarchical, asynchronous, and parallel framework, effectively bypassing the VLM's inference latency and boosting the agent's proactivity in exploration. Extensive experiments show that APEX outperforms the previous state of the art by +4.2% SR and +2.8% SPL on challenging UAV‑ON benchmarks, demonstrating its superior efficiency and the effectiveness of its hierarchical asynchronous design. Our source code is provided in \hrefhttps://github.com/4amGodvzx/apexGitHub
Authors:Yifan Zhang, Qian Chen, Yi Liu, Wengen Li, Jihong Guan
Abstract:
Cloud contamination severely degrades the usability of remote sensing imagery and poses a fundamental challenge for downstream Earth observation tasks. Recently, diffusion‑based models have emerged as a dominant paradigm for remote sensing cloud removal due to their strong generative capability and stable optimization. However, existing diffusion‑based approaches often suffer from limited sampling efficiency and insufficient exploitation of structural and temporal priors in multi‑temporal remote sensing scenarios. In this work, we propose SADER, a structure‑aware diffusion framework for multi‑temporal remote sensing cloud removal. SADER first develops a scalable Multi‑Temporal Conditional Diffusion Network (MTCDN) to fully capture multi‑temporal and multimodal correlations via temporal fusion and hybrid attention. Then, a cloud‑aware attention loss is introduced to emphasize cloud‑dominated regions by accounting for cloud thickness and brightness discrepancies. In addition, a deterministic resampling strategy is designed for continuous diffusion models to iteratively refine samples under fixed sampling steps by replacing outliers through guided correction. Extensive experiments on multiple multi‑temporal datasets demonstrate that SADER consistently outperforms state‑of‑the‑art cloud removal methods across all evaluation metrics. The code of SADER is publicly available at https://github.com/zyfzs0/SADER.
Authors:Chaoran Xu, Chengkan Lv, Qiyu Chen, Feng Zhang, Zhengtao Zhang
Abstract:
Zero‑shot anomaly detection (ZSAD) often leverages pretrained vision or vision‑language models, but many existing methods use prompt learning or complex modeling to fit the data distribution, resulting in high training or inference cost and limited cross‑domain stability. To address these limitations, we propose Memory‑Retrieval Anomaly Detection method (MRAD), a unified framework that replaces parametric fitting with a direct memory retrieval. The train‑free base model, MRAD‑TF, freezes the CLIP image encoder and constructs a two‑level memory bank (image‑level and pixel‑level) from auxiliary data, where feature‑label pairs are explicitly stored as keys and values. During inference, anomaly scores are obtained directly by similarity retrieval over the memory bank. Based on the MRAD‑TF, we further propose two lightweight variants as enhancements: (i) MRAD‑FT fine‑tunes the retrieval metric with two linear layers to enhance the discriminability between normal and anomaly; (ii) MRAD‑CLIP injects the normal and anomalous region priors from the MRAD‑FT as dynamic biases into CLIP's learnable text prompts, strengthening generalization to unseen categories. Across 16 industrial and medical datasets, the MRAD framework consistently demonstrates superior performance in anomaly classification and segmentation, under both train‑free and training‑based settings. Our work shows that fully leveraging the empirical distribution of raw data, rather than relying only on model fitting, can achieve stronger anomaly detection performance. The code will be publicly released at https://github.com/CROVO1026/MRAD.
Authors:Chia-Ming Lee, Yu-Hao Ho, Yu-Fan Lin, Jen-Wei Lee, Li-Wei Kang, Chih-Chung Hsu
Abstract:
Hyperspectral image (HSI) fusion aims to reconstruct a high‑resolution HSI (HR‑HSI) by combining the rich spectral information of a low‑resolution HSI (LR‑HSI) with the fine spatial details of a high‑resolution multispectral image (HR‑MSI). Although recent deep learning methods have achieved notable progress, they still suffer from limited receptive fields, redundant spectral bands, and the quadratic complexity of self‑attention, which restrict both efficiency and robustness. To overcome these challenges, we propose the Hierarchical Spatial‑Spectral Dense Correlation Network (HSSDCT). The framework introduces two key modules: (i) a Hierarchical Dense‑Residue Transformer Block (HDRTB) that progressively enlarges windows and employs dense‑residue connections for multi‑scale feature aggregation, and (ii) a Spatial‑Spectral Correlation Layer (SSCL) that explicitly factorizes spatial and spectral dependencies, reducing self‑attention to linear complexity while mitigating spectral redundancy. Extensive experiments on benchmark datasets demonstrate that HSSDCT delivers superior reconstruction quality with significantly lower computational costs, achieving new state‑of‑the‑art performance in HSI fusion. Our code is available at https://github.com/jemmyleee/HSSDCT.
Authors:Rong-Lin Jian, Ming-Chi Luo, Chen-Wei Huang, Chia-Ming Lee, Yu-Fan Lin, Chih-Chung Hsu
Abstract:
Multi‑object tracking (MOT) in sports is highly challenging due to irregular player motion, uniform appearances, and frequent occlusions. These difficulties are further exacerbated by the geometric distortion and extreme scale variation introduced by static fisheye cameras. In this work, we present GTATrack, a hierarchical tracking framework that win first place in the SoccerTrack Challenge 2025. GTATrack integrates two core components: Deep Expansion IoU (Deep‑EIoU) for motion‑agnostic online association and Global Tracklet Association (GTA) for trajectory‑level refinement. This two‑stage design enables both robust short‑term matching and long‑term identity consistency. Additionally, a pseudo‑labeling strategy is used to boost detector recall on small and distorted targets. The synergy between local association and global reasoning effectively addresses identity switches, occlusions, and tracking fragmentation. Our method achieved a winning HOTA score of 0.60 and significantly reduced false positives to 982, demonstrating state‑of‑the‑art accuracy in fisheye‑based soccer tracking. Our code is available at https://github.com/ron941/GTATrack‑STC2025.
Authors:Xinlei Yu, Chengming Xu, Zhangquan Chen, Bo Yin, Cheng Yang, Yongbo He, Yihao Hu, Jiangning Zhang, Cheng Tan, Xiaobin Hu, Shuicheng Yan
Abstract:
While Visual Multi‑Agent Systems (VMAS) promise to enhance comprehensive abilities through inter‑agent collaboration, empirical evidence reveals a counter‑intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure to the information bottleneck inherent in text‑centric communication, where converting perceptual and thinking trajectories into discrete natural language inevitably induces semantic loss. To this end, we propose L^2‑VMAS, a novel model‑agnostic framework that enables inter‑agent collaboration with dual latent memories. Furthermore, we decouple the perception and thinking while dynamically synthesizing dual latent memories. Additionally, we introduce an entropy‑driven proactive triggering that replaces passive information transmission with efficient, on‑demand memory access. Extensive experiments among backbones, sizes, and multi‑agent structures demonstrate that our method effectively breaks the "scaling wall" with superb scalability, improving average accuracy by 2.7‑5.4% while reducing token usage by 21.3‑44.8%. Codes: https://github.com/YU‑deep/L2‑VMAS.
Authors:Gabriel Bromonschenkel, Alessandro L. Koerich, Thiago M. Paixão, Hilário Tomaz Alves de Oliveira
Abstract:
Image captioning (IC) refers to the automatic generation of natural language descriptions for images, with applications ranging from social media content generation to assisting individuals with visual impairments. While most research has been focused on English‑based models, low‑resource languages such as Brazilian Portuguese face significant challenges due to the lack of specialized datasets and models. Several studies create datasets by automatically translating existing ones to mitigate resource scarcity. This work addresses this gap by proposing a cross‑native‑translated evaluation of Transformer‑based vision and language models for Brazilian Portuguese IC. We use a version of Flickr30K comprised of captions manually created by native Brazilian Portuguese speakers and compare it to a version with captions automatically translated from English to Portuguese. The experiments include a cross‑context approach, where models trained on one dataset are tested on the other to assess the translation impact. Additionally, we incorporate attention maps for model inference interpretation and use the CLIP‑Score metric to evaluate the image‑description alignment. Our findings show that Swin‑DistilBERTimbau consistently outperforms other models, demonstrating strong generalization across datasets. ViTucano, a Brazilian Portuguese pre‑trained VLM, surpasses larger multilingual models (GPT‑4o, LLaMa 3.2 Vision) in traditional text‑based evaluation metrics, while GPT‑4 models achieve the highest CLIP‑Score, highlighting improved image‑text alignment. Attention analysis reveals systematic biases, including gender misclassification, object enumeration errors, and spatial inconsistencies. The datasets and the models generated and analyzed during the current study are available in: https://github.com/laicsiifes/transformer‑caption‑ptbr.
Authors:Alberto Mario Ceballos-Arroyo, Shrikanth M. Yadav, Chu-Hsuan Lin, Jisoo Kim, Geoffrey S. Young, Lei Qin, Huaizu Jiang
Abstract:
In this study, we develop a novel methodology for annotating the brain vasculature using dynamic 4D‑CTA head scans. By using multiple time points from dynamic CTA acquisitions, we subtract bone and soft tissue to enhance the visualization of arteries and veins, reducing the effort required to obtain manual annotations of brain vessels. We then train deep learning models on our ground truth annotations by using the same segmentation for multiple phases from the dynamic 4D‑CTA collection, effectively enlarging our dataset by 4 to 5 times and inducing robustness to contrast phases. In total, our dataset comprises 110 training images from 25 patients and 165 test images from 14 patients. In comparison with two similarly‑sized datasets for CTA‑based brain vessel segmentation, a nnUNet model trained on our dataset can achieve significantly better segmentations across all vascular regions, with an average mDC of 0.846 for arteries and 0.957 for veins in the TopBrain dataset. Furthermore, metrics such as average directed Hausdorff distance (adHD) and topology sensitivity (tSens) reflected similar trends: using our dataset resulted in low error margins (adHD of 0.304 mm for arteries and 0.078 for veins) and high sensitivity (tSens of 0.877 for arteries and 0.974 for veins), indicating excellent accuracy in capturing vessel morphology. Our code and model weights are available online at https://github.com/alceballosa/robust‑vessel‑segmentation
Authors:Ignacy Kolton, Kacper Marzol, Paweł Batorski, Marcin Mazur, Paul Swoboda, Przemysław Spurek
Abstract:
Machine unlearning is a key defense mechanism for removing unauthorized concepts from text‑to‑image diffusion models, yet recent evidence shows that latent visual information often persists after unlearning. Existing adversarial approaches for exploiting this leakage are constrained by fundamental limitations: optimization‑based methods are computationally expensive due to per‑instance iterative search. At the same time, reasoning‑based and heuristic techniques lack direct feedback from the target model's latent visual representations. To address these challenges, we introduce ReLAPSe, a policy‑based adversarial framework that reformulates concept restoration as a reinforcement learning problem. ReLAPSe trains an agent using Reinforcement Learning with Verifiable Rewards (RLVR), leveraging the diffusion model's noise prediction loss as a model‑intrinsic and verifiable feedback signal. This closed‑loop design directly aligns textual prompt manipulation with latent visual residuals, enabling the agent to learn transferable restoration strategies rather than optimizing isolated prompts. By pioneering the shift from per‑instance optimization to global policy learning, ReLAPSe achieves efficient, near‑real‑time recovery of fine‑grained identities and styles across multiple state‑of‑the‑art unlearning methods, providing a scalable tool for rigorous red‑teaming of unlearned diffusion models. Some experimental evaluations involve sensitive visual concepts, such as nudity. Code is available at https://github.com/gmum/ReLaPSe
Authors:Zhengyi Lu, Ming Lu, Chongyu Qu, Junchao Zhu, Junlin Guo, Marilyn Lionts, Yanfan Zhu, Yuechen Yang, Tianyuan Yao, Jayasai Rajagopal, Bennett Allan Landman, Xiao Wang, Xinqiang Yan, Yuankai Huo
Abstract:
Metal implants in MRI cause severe artifacts that degrade image quality and hinder clinical diagnosis. Traditional approaches address metal artifact reduction (MAR) and accelerated MRI acquisition as separate problems. We propose MASC, a unified reinforcement learning framework that jointly optimizes metal‑aware k‑space sampling and artifact correction for accelerated MRI. To enable supervised training, we construct a paired MRI dataset using physics‑based simulation, generating k‑space data and reconstructions for phantoms with and without metal implants. This paired dataset provides simulated 3D MRI scans with and without metal implants, where each metal‑corrupted sample has an exactly matched clean reference, enabling direct supervision for both artifact reduction and acquisition policy learning. We formulate active MRI acquisition as a sequential decision‑making problem, where an artifact‑aware Proximal Policy Optimization (PPO) agent learns to select k‑space phase‑encoding lines under a limited acquisition budget. The agent operates on undersampled reconstructions processed through a U‑Net‑based MAR network, learning patterns that maximize reconstruction quality. We further propose an end‑to‑end training scheme where the acquisition policy learns to select k‑space lines that best support artifact removal while the MAR network simultaneously adapts to the resulting undersampling patterns. Experiments demonstrate that MASC's learned policies outperform conventional sampling strategies, and end‑to‑end training improves performance compared to using a frozen pre‑trained MAR network, validating the benefit of joint optimization. Cross‑dataset experiments on FastMRI with physics‑based artifact simulation further confirm generalization to realistic clinical MRI data. The code and models of MASC have been made publicly available: https://github.com/hrlblab/masc
Authors:Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius
Abstract:
Fine‑grained spatio‑temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio‑temporal understanding. Inspired by cognitive science, TimeBlind categorizes fine‑grained temporal understanding into three levels: recognizing atomic events, characterizing event properties, and reasoning about event interdependencies. Unlike benchmarks that conflate recognition with temporal reasoning, TimeBlind leverages a minimal‑pairs paradigm: video pairs share identical static visual content but differ solely in temporal structure, utilizing complementary questions to neutralize language priors. Evaluating over 20 state‑of‑the‑art MLLMs (e.g., GPT‑5, Gemini 3 Pro) on 600 curated instances (2400 video‑question pairs), reveals that the Instance Accuracy (correctly distinguishing both videos in a pair) of the best performing MLLM is only 48.2%, far below the human performance (98.2%). These results demonstrate that even frontier models rely heavily on static visual shortcuts rather than genuine temporal logic, positioning TimeBlind as a vital diagnostic tool for next‑generation video understanding. Dataset and code are available at https://baiqi‑li.github.io/timeblind_project/ .
Authors:Beier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun, Yuan Zhou, Xun Yang, Hanwang Zhang
Abstract:
Deep neural networks often exhibit substantial disparities in class‑wise accuracy, even when trained on class‑balanced data, posing concerns for reliable deployment. While prior efforts have explored empirical remedies, a theoretical understanding of such performance disparities in classification remains limited. In this work, we present Margin Regularization for Performance Disparity Reduction (MR^2), a theoretically principled regularization for classification by dynamically adjusting margins in both the logit and representation spaces. Our analysis establishes a margin‑based, class‑sensitive generalization bound that reveals how per‑class feature variability contributes to error, motivating the use of larger margins for hard classes. Guided by this insight, MR^2 optimizes per‑class logit margins proportional to feature spread and penalizes excessive representation margins to enhance intra‑class compactness. Experiments on seven datasets, including ImageNet, and diverse pre‑trained backbones (MAE, MoCov2, CLIP) demonstrate that MR^2 not only improves overall accuracy but also significantly boosts hard class performance without trading off easy classes, thus reducing performance disparity. Code is available at: https://github.com/BeierZhu/MR2
Authors:Shanwen Wang, Xin Sun, Danfeng Hong, Fei Zhou
Abstract:
The semi‑supervised semantic segmentation (S4) can learn rich visual knowledge from low‑cost unlabeled images. However, traditional S4 architectures all face the challenge of low‑quality pseudo‑labels, especially for the teacher‑student framework.We propose a novel SemiEarth model that introduces vision‑language models (VLMs) to address the S4 issues for the remote sensing (RS) domain. Specifically, we invent a VLM pseudo‑label purifying (VLM‑PP) structure to purify the teacher network's pseudo‑labels, achieving substantial improvements. Especially in multi‑class boundary regions of RS images, the VLM‑PP module can significantly improve the quality of pseudo‑labels generated by the teacher, thereby correctly guiding the student model's learning. Moreover, since VLM‑PP equips VLMs with open‑world capabilities and is independent of the S4 architecture, it can correct mispredicted categories in low‑confidence pseudo‑labels whenever a discrepancy arises between its prediction and the pseudo‑label. We conducted extensive experiments on multiple RS datasets, which demonstrate that our SemiEarth achieves SOTA performance. More importantly, unlike previous SOTA RS S4 methods, our model not only achieves excellent performance but also offers good interpretability. The code is released at https://github.com/wangshanwen001/SemiEarth.
Authors:Elif Nebioglu, Emirhan Bilgiç, Adrian Popescu
Abstract:
Modern deep learning‑based inpainting enables realistic local image manipulation, raising critical challenges for reliable detection. However, we observe that current detectors primarily rely on global artifacts that appear as inpainting side effects, rather than on locally synthesized content. We show that this behavior occurs because VAE‑based reconstruction induces a subtle but pervasive spectral shift across the entire image, including unedited regions. To isolate this effect, we introduce Inpainting Exchange (INP‑X), an operation that restores original pixels outside the edited region while preserving all synthesized content. We create a 90K test dataset including real, inpainted, and exchanged images to evaluate this phenomenon. Under this intervention, pretrained state‑of‑the‑art detectors, including commercial ones, exhibit a dramatic drop in accuracy (e.g., from 91% to 55%), frequently approaching chance level. We provide a theoretical analysis linking this behavior to high‑frequency attenuation caused by VAE information bottlenecks. Our findings highlight the need for content‑aware detection. Indeed, training on our dataset yields better generalization and localization than standard inpainting. Our dataset and code are publicly available at https://github.com/emirhanbilgic/INP‑X.
Authors:Yadang Alexis Rouzoumka, Jean Pinsolle, Eugénie Terreaux, Christèle Morisseau, Jean-Philippe Ovarlez, Chengfang Ren
Abstract:
Diffusion models learn a time‑indexed score field \mathbfs_θ(\mathbfx_t,t) that often inherits approximate equivariances (flips, rotations, circular shifts) from in‑distribution (ID) data and convolutional backbones. Most diffusion‑based out‑of‑distribution (OOD) detectors exploit score magnitude or local geometry (energies, curvature, covariance spectra) and largely ignore equivariances. We introduce Group‑Equivariant Posterior Consistency (GEPC), a training‑free probe that measures how consistently the learned score transforms under a finite group \mathcalG, detecting equivariance breaking even when score magnitude remains unchanged. At the population level, we propose the ideal GEPC residual, which averages an equivariance‑residual functional over \mathcalG, and we derive ID upper bounds and OOD lower bounds under mild assumptions. GEPC requires only score evaluations and produces interpretable equivariance‑breaking maps. On OOD image benchmark datasets, we show that GEPC achieves competitive or improved AUROC compared to recent diffusion‑based baselines while remaining computationally lightweight. On high‑resolution synthetic aperture radar imagery where OOD corresponds to targets or anomalies in clutter, GEPC yields strong target‑background separation and visually interpretable equivariance‑breaking maps. Code is available at https://github.com/RouzAY/gepc‑diffusion/.
Authors:Yiyang Wen, Liu Shi, Zekun Zhou, WenZhe Shan, Qiegen Liu
Abstract:
Limited‑angle computed tomography (LACT) offers the advantages of reduced radiation dose and shortened scanning time. Traditional reconstruction algorithms exhibit various inherent limitations in LACT. Currently, most deep learning‑based LACT reconstruction methods focus on multi‑domain fusion or the introduction of generic priors, failing to fully align with the core imaging characteristics of LACT‑such as the directionality of artifacts and directional loss of structural information, which are caused by the absence of projection angles in certain directions. Inspired by the theory of visible and invisible singularities, taking into account the aforementioned core imaging characteristics of LACT, we propose a Visible Singularities Guided Correlation network for LACT reconstruction (VSGC). The design philosophy of VSGC consists of two core steps: First, extract VS edge features from LACT images and focus the model's attention on these VS. Second, establish correlations between the VS edge features and other regions of the image. Additionally, a multi‑scale loss function with anisotropic constraint is employed to constrain the model to converge in multiple aspects. Finally, qualitative and quantitative validations are conducted on both simulated and real datasets to verify the effectiveness and feasibility of the proposed design. Particularly, in comparison with alternative methods, VSGC delivers more prominent performance in small angular ranges, with the PSNR improvement of 2.45 dB and the SSIM enhancement of 1.5%. The code is publicly available at https://github.com/yqx7150/VSGC.
Authors:Jiajun Zhao, Xuan Yang
Abstract:
We propose an intra‑class subdivision pixel contrastive learning (SPCL) framework for cardiac image segmentation to address representation contamination at boundaries. The novel concept ``Unconcerned sample'' is proposed to distinguish pixel representations at the inner and boundary regions within the same class, facilitating a clearer characterization of intra‑class variations. A novel boundary contrastive loss for boundary representations is proposed to enhance representation discrimination across boundaries. The advantages of the unconcerned sample and boundary contrastive loss are analyzed theoretically. Experimental results in public cardiac datasets demonstrate that SPCL significantly improves segmentation performance, outperforming existing methods with respect to segmentation quality and boundary precision. Our code is available at https://github.com/Jrstud203/SPCL.
Authors:Xuan Rao, Mingming Ha, Bo Zhao, Derong Liu, Cesare Alippi
Abstract:
Class‑incremental learning (CIL) with Vision Transformers (ViTs) faces a major computational bottleneck during the classifier reconstruction phase, where most existing methods rely on costly iterative stochastic gradient descent (SGD). We observe that analytic Regularized Gaussian Discriminant Analysis (RGDA) provides a Bayes‑optimal alternative with accuracy comparable to SGD‑based classifiers; however, its quadratic inference complexity limits its use in large‑scale CIL scenarios. To overcome this, we propose Low‑Rank Factorized RGDA (LR‑RGDA), a scalable classifier that combines RGDA's expressivity with the efficiency of linear classifiers. By exploiting the low‑rank structure of the covariance via the Woodbury matrix identity, LR‑RGDA decomposes the discriminant function into a global affine term refined by a low‑rank quadratic perturbation, reducing the inference complexity from \mathcalO(Cd^2) to \mathcalO(d^2 + Crd^2), where C is the class number, d the feature dimension, and r \ll d the subspace rank. To mitigate representation drift caused by backbone updates, we further introduce Hopfield‑based Distribution Compensator (HopDC), a training‑free mechanism that uses modern continuous Hopfield Networks to recalibrate historical class statistics through associative memory dynamics on unlabeled anchors, accompanied by a theoretical bound on the estimation error. Extensive experiments on diverse CIL benchmarks demonstrate that our framework achieves state‑of‑the‑art performance, providing a scalable solution for large‑scale class‑incremental learning with ViTs. Code: https://github.com/raoxuan98‑hash/lr_rgda_hopdc.
Authors:Wing Chan, Richard Allen
Abstract:
Public demos of image editing models are typically best‑case samples; real workflows pay for retries and review time. We introduce HYPE‑EDIT‑1, a 100‑task benchmark of reference‑based marketing/design edits with binary pass/fail judging. For each task we generate 10 independent outputs to estimate per‑attempt pass rate, pass@10, expected attempts under a retry cap, and an effective cost per successful edit that combines model price with human review time. We release 50 public tasks and maintain a 50‑task held‑out private split for server‑side evaluation, plus a standardized JSON schema and tooling for VLM and human‑based judging. Across the evaluated models, per‑attempt pass rates span 34‑83 percent and effective cost per success spans USD 0.66‑1.42. Models that have low per‑image pricing are more expensive when you consider the total effective cost of retries and human reviews.
Authors:Weiyu Sun, Liangliang Chen, Yongnuo Cai, Huiru Xie, Yi Zeng, Ying Zhang
Abstract:
Multimodal Large Language Models (MLLMs) hold significant promise for revolutionizing traditional education and reducing teachers' workload. However, accurately interpreting unconstrained STEM student handwritten solutions with intertwined mathematical formulas, diagrams, and textual reasoning poses a significant challenge due to the lack of authentic and domain‑specific benchmarks. Additionally, current evaluation paradigms predominantly rely on the outcomes of downstream tasks (e.g., auto‑grading), which often probe only a subset of the recognized content, thereby failing to capture the MLLMs' understanding of complex handwritten logic as a whole. To bridge this gap, we release EDU‑CIRCUIT‑HW, a dataset consisting of 1,300+ authentic student handwritten solutions from a university‑level STEM course. Utilizing the expert‑verified verbatim transcriptions and grading reports of student solutions, we simultaneously evaluate various MLLMs' upstream recognition fidelity and downstream auto‑grading performance. Our evaluation uncovers an astonishing scale of latent failures within MLLM‑recognized student handwritten content, highlighting the models' insufficient reliability for auto‑grading and other understanding‑oriented applications in high‑stakes educational settings. As a potential solution, we present a case study demonstrating that leveraging identified error patterns to preemptively detect and correct recognition errors, while requiring only minimal human intervention (e.g., routing 3.3% of assignments to human graders and the remainder to the GPT‑5.1 grader), can effectively enhance the robustness of the deployed AI‑enabled grading system. Code and dataset are available in this GitHub repo: https://gt‑learning‑innovation.github.io/CIRCUIT_EDU_HW_ACL.
Authors:Han Xiao
Abstract:
We present an ε‑bounded compression method for unit‑norm embeddings that achieves 1.5× compression, 25% better than the best prior lossless method. The method exploits that spherical coordinates of high‑dimensional unit vectors concentrate around π/2, causing IEEE 754 exponents to collapse to a single value and high‑order mantissa bits to become predictable, enabling entropy coding of both. Reconstruction error is bounded by float32 machine epsilon (1.19 × 10^‑7), making reconstructed values indistinguishable from originals at float32 precision. Evaluation across 26 configurations spanning text, image, and multi‑vector embeddings confirms consistent compression improvement with zero measurable retrieval degradation on BEIR benchmarks.
Authors:Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zhang, Yan Huang, Liang Wang
Abstract:
In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open‑domain video shot retrieval, which involves richer temporal structure and more complex semantics, still lacks systematic benchmarks and analysis. To fill this gap, we introduce ShotFinder, a benchmark that formalizes editing requirements as keyframe‑oriented shot descriptions and introduces five types of controllable single‑factor constraints: Temporal order, Color, Visual style, Audio, and Resolution. We curate 1,210 high‑quality samples from YouTube across 20 thematic categories, using large models for generation with human verification. Based on the benchmark, we propose ShotFinder, a text‑driven three‑stage retrieval and localization pipeline: (1) query expansion via video imagination, (2) candidate video retrieval with a search engine, and (3) description‑guided temporal localization. Experiments on multiple closed‑source and open‑source models reveal a significant gap to human performance, with clear imbalance across constraints: temporal localization is relatively tractable, while color and visual style remain major challenges. These results reveal that open‑domain video shot retrieval is still a critical capability that multimodal large models have yet to overcome.
Authors:Seungjun Lee, Gim Hee Lee
Abstract:
Scene understanding with free‑form language has been widely explored within diverse modalities such as images, point clouds, and LiDAR. However, related studies on event sensors are scarce or narrowly centered on semantic‑level understanding. We introduce SEAL, the first Semantic‑aware Segment Any Events framework that addresses Open‑Vocabulary Event Instance Segmentation (OV‑EIS). Given the visual prompt, our model presents a unified framework to support both event segmentation and open‑vocabulary mask classification at multiple levels of granularity, including instance‑level and part‑level. To enable thorough evaluation on OV‑EIS, we curate four benchmarks that cover label granularity from coarse to fine class configurations and semantic granularity from instance‑level to part‑level understanding. Extensive experiments show that our SEAL largely outperforms proposed baselines in terms of performance and inference speed with a parameter‑efficient architecture. In the Appendix, we further present a simple variant of our SEAL achieving generic spatiotemporal OV‑EIS that does not require any visual prompts from users in the inference. Check out our project page in https://0nandon.github.io/SEAL
Authors:Ping Chen, Zicheng Huang, Xiangming Wang, Yungeng Liu, Bingyu Liang, Haijin Zeng, Yongyong Chen
Abstract:
We propose VL‑DUN, a principled framework for joint All‑in‑One Medical Image Restoration and Segmentation (AiOMIRS) that bridges the gap between low‑level signal recovery and high‑level semantic understanding. While standard pipelines treat these tasks in isolation, our core insight is that they are fundamentally synergistic: restoration provides clean anatomical structures to improve segmentation, while semantic priors regularize the restoration process. VL‑DUN resolves the sub‑optimality of sequential processing through two primary innovations. (1) We formulate AiOMIRS as a unified optimization problem, deriving an interpretable joint unfolding mechanism where restoration and segmentation are mathematically coupled for mutual refinement. (2) We introduce a frequency‑aware Mamba mechanism to capture long‑range dependencies for global segmentation while preserving the high‑frequency textures necessary for restoration. This allows for efficient global context modeling with linear complexity, effectively mitigating the spectral bias of standard architectures. As a pioneering work in the AiOMIRS task, VL‑DUN establishes a new state‑of‑the‑art across multi‑modal benchmarks, improving PSNR by 0.92 dB and the Dice coefficient by 9.76%. Our results demonstrate that joint collaborative learning offers a superior, more robust solution for complex clinical workflows compared to isolated task processing. The codes are provided in https://github.com/cipi666/VLDUN.
Authors:Yinsong Wang, Thomas Fletcher, Xinzhe Luo, Aine Travers Dineen, Rhodri Cusack, Chen Qin
Abstract:
Reconstructing 3D fetal MR volumes from motion‑corrupted stacks of 2D slices is a crucial and challenging task. Conventional slice‑to‑volume reconstruction (SVR) methods are time‑consuming and require multiple orthogonal stacks for reconstruction. While learning‑based SVR approaches have significantly reduced the time required at the inference stage, they heavily rely on ground truth information for training, which is inaccessible in practice. To address these challenges, we propose GaussianSVR, a self‑supervised framework for slice‑to‑volume reconstruction. GaussianSVR represents the target volume using 3D Gaussian representations to achieve high‑fidelity reconstruction. It leverages a simulated forward slice acquisition model to enable self‑supervised training, alleviating the need for ground‑truth volumes. Furthermore, to enhance both accuracy and efficiency, we introduce a multi‑resolution training strategy that jointly optimizes Gaussian parameters and spatial transformations across different resolution levels. Experiments show that GaussianSVR outperforms the baseline methods on fetal MR volumetric reconstruction. Code is available at https://github.com/Yinsong0510/GaussianSVR‑Self‑Supervised‑Slice‑to‑Volume‑Reconstruction‑with‑Gaussian‑Representations.
Authors:Wulin Xie, Rui Dai, Ruidong Ding, Kaikui Liu, Xiangxiang Chu, Xinwen Hou, Jie Wen
Abstract:
Image Quality Assessment (IQA) predicts perceptual quality scores consistent with human judgments. Recent RL‑based IQA methods built on MLLMs focus on generating visual quality descriptions and scores, ignoring two key reliability limitations: (i) although the model's prediction stability varies significantly across training samples, existing GRPO‑based methods apply uniform advantage weighting, thereby amplifying noisy signals from unstable samples in gradient updates; (ii) most works emphasize text‑grounded reasoning over images while overlooking the model's visual perception ability of image content. In this paper, we propose Q‑Hawkeye, an RL‑based reliable visual policy optimization framework that redesigns the learning signal through unified Uncertainty‑Aware Dynamic Optimization and Perception‑Aware Optimization. Q‑Hawkeye estimates predictive uncertainty using the variance of predicted scores across multiple rollouts and leverages this uncertainty to reweight each sample's update strength, stabilizing policy optimization. To strengthen perceptual reliability, we construct paired inputs of degraded images and their original images and introduce an Implicit Perception Loss that constrains the model to ground its quality judgments in genuine visual evidence. Extensive experiments demonstrate that Q‑Hawkeye outperforms state‑of‑the‑art methods and generalizes better across multiple datasets. Our dataset and code are available at https://github.com/AMAP‑ML/Q‑Hawkeye.
Authors:Siyi Du, Xinzhe Luo, Declan P. O'Regan, Chen Qin
Abstract:
Multimodal deep learning (MDL) has achieved remarkable success across various domains, yet its practical deployment is often hindered by incomplete multimodal data. Existing incomplete MDL methods either discard missing modalities, risking the loss of valuable task‑relevant information, or recover them, potentially introducing irrelevant noise, leading to the discarding‑imputation dilemma. To address this dilemma, in this paper, we propose DyMo, a new inference‑time dynamic modality selection framework that adaptively identifies and fuses reliable recovered modalities, fully exploring task‑relevant information beyond the conventional discard‑or‑impute paradigm. Central to DyMo is a novel selection algorithm that maximizes multimodal task‑relevant information for each test sample. Since direct estimation of such information at test time is intractable due to the unknown data distribution, we theoretically establish a connection between information and the task loss, which we compute at inference time as a tractable proxy. Building on this, a novel principled reward function is proposed to guide modality selection. In addition, we design a flexible multimodal network architecture compatible with arbitrary modality combinations, alongside a tailored training strategy for robust representation learning. Extensive experiments on diverse natural and medical image datasets show that DyMo significantly outperforms state‑of‑the‑art incomplete/dynamic MDL methods across various missing‑data scenarios. Our code is available at https://github.com//siyi‑wind/DyMo.
Authors:Haiyang Wu, Weiliang Mu, Jipeng Zhang, Zhong Dandan, Zhuofei Du, Haifeng Li, Tao Chao
Abstract:
Existing methods for farmland remote sensing image (FRSI) segmentation generally follow a static segmentation paradigm, where analysis relies solely on the limited information contained within a single input patch. Consequently, their reasoning capability is limited when dealing with complex scenes characterized by ambiguity and visual uncertainty. In contrast, human experts, when interpreting remote sensing images in such ambiguous cases, tend to actively query auxiliary images (such as higher‑resolution, larger‑scale, or temporally adjacent data) to conduct cross‑verification and achieve more comprehensive reasoning. Inspired by this, we propose a reasoning‑query‑driven dynamic segmentation framework for FRSIs, named FarmMind. This framework breaks through the limitations of the static segmentation paradigm by introducing a reasoning‑query mechanism, which dynamically and on‑demand queries external auxiliary images to compensate for the insufficient information in a single input image. Unlike direct queries, this mechanism simulates the thinking process of human experts when faced with segmentation ambiguity: it first analyzes the root causes of segmentation ambiguities through reasoning, and then determines what type of auxiliary image needs to be queried based on this analysis. Extensive experiments demonstrate that FarmMind achieves superior segmentation performance and stronger generalization ability compared with existing methods. The source code and dataset used in this work are publicly available at: https://github.com/WithoutOcean/FarmMind.
Authors:Xingwu Zhang, Guanxuan Li, Paul Henderson, Gerardo Aragon-Camarasa, Zijun Long
Abstract:
Current state‑of‑the‑art multi‑class unsupervised anomaly detection (MUAD) methods rely on training encoder‑decoder models to reconstruct anomaly‑free features. We first show these approaches have an inherent fidelity‑stability dilemma in how they detect anomalies via reconstruction residuals. We then abandon the reconstruction paradigm entirely and propose Retrieval‑based Anomaly Detection (RAD). RAD is a training‑free approach that stores anomaly‑free features in a memory and detects anomalies through multi‑level retrieval, matching test patches against the memory. Experiments demonstrate that RAD achieves state‑of‑the‑art performance across four established benchmarks (MVTec‑AD, VisA, Real‑IAD, 3D‑ADAM) under both standard and few‑shot settings. On MVTec‑AD, RAD reaches 96.7% Pixel AUROC with just a single anomaly‑free image compared to 98.5% of RAD's full‑data performance. We further prove that retrieval‑based scores theoretically upper‑bound reconstruction‑residual scores. Collectively, these findings overturn the assumption that MUAD requires task‑specific training, showing that state‑of‑the‑art anomaly detection is feasible with memory‑based retrieval. Our code is available at https://github.com/longkukuhi/RAD.
Authors:Enyi Shi, Pengyang Shao, Yanxin Zhang, Chenhang Cui, Jiayi Lyu, Xiaobo Xia, Fei Shen, Tat-Seng Chua
Abstract:
The robust safety of Vision‑Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current benchmarks typically isolate these dimensions, being either multilingual but text‑only, or multimodal but monolingual. While recent red‑teaming efforts attempt to bridge this gap by rendering harmful prompts as images, their overreliance on typography‑style visuals and lack of semantically grounded image‑text pairs fail to capture realistic cross‑modal interactions under multilingual and multimodal conditions. To address this, we introduce Lingua‑SafetyBench, a comprehensive benchmark of 100,440 harmful image‑text pairs spanning 10 languages. Crucially, Lingua‑SafetyBench explicitly partitions data into image‑dominant and text‑dominant subsets to precisely disentangle sources of risk. Extensive evaluations reveal that current VLLMs retain non‑negligible vulnerabilities under these joint inputs. Linguistically, requests in Non‑High‑Resource Languages (Non‑HRLs) and non‑Latin scripts generally pose greater threats. Furthermore, analyzing modality‑language interactions uncovers a striking asymmetry: in High‑Resource Languages (HRLs), models are most vulnerable to image‑dominant risks, whereas in Non‑HRLs, text‑dominant risks severely degrade safety performance. Finally, a controlled study on the Qwen series demonstrates that while model scaling and iterative upgrades improve overall safety, they disproportionately benefit HRLs. This exacerbates the safety disparity between HRLs and Non‑HRLs under text‑dominant risks, highlighting that achieving robust safety requires dedicated language‑ and modality‑aware alignment strategies beyond mere scaling. The code and dataset will be available at https://github.com/zsxr15/Lingua‑SafetyBench.Warning: this paper contains examples with unsafe content.
Authors:Yanlong Chen, Amirhossein Habibian, Luca Benini, Yawei Li
Abstract:
Vision‑Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post‑training quantization often causes significant accuracy loss. Despite its potential, quantization‑aware training for VLMs remains underexplored. We propose GRACE, a framework unifying knowledge distillation and QAT under the Information Bottleneck principle: quantization constrains information capacity while distillation guides what to preserve within this budget. Treating the teacher as a proxy for task‑relevant information, we introduce confidence‑gated decoupled distillation to filter unreliable supervision, relational centered kernel alignment to transfer visual token structures, and an adaptive controller via Lagrangian relaxation to balance fidelity against capacity constraints. Across extensive benchmarks on LLaVA and Qwen families, our INT4 models consistently outperform FP16 baselines (e.g., LLaVA‑1.5‑7B: 70.1 vs. 66.8 on SQA; Qwen2‑VL‑2B: 76.9 vs. 72.6 on MMBench), nearly matching teacher performance. Using real INT4 kernel, we achieve 3× throughput with 54% memory reduction. This principled framework significantly outperforms existing quantization methods, making GRACE a compelling solution for resource‑constrained deployment. Code and data are available at: https://github.com/ForeverBlue816/GRACE.
Authors:Jiahao Wu, Yunfei Liu, Lijian Lin, Ye Zhu, Lei Zhu, Jingyi Li, Yu Li
Abstract:
Reconstructing detailed 3D human meshes from a single in‑the‑wild image remains a fundamental challenge in computer vision. Existing SMPLX‑based methods often suffer from slow inference, produce only coarse body poses, and exhibit misalignments or unnatural artifacts in fine‑grained regions such as the face and hands. These issues make current approaches difficult to apply to downstream tasks. To address these challenges, we propose PEAR‑a fast and robust framework for pixel‑aligned expressive human mesh recovery. PEAR explicitly tackles three major limitations of existing methods: slow inference, inaccurate localization of fine‑grained human pose details, and insufficient facial expression capture. Specifically, to enable real‑time SMPLX parameter inference, we depart from prior designs that rely on high resolution inputs or multi‑branch architectures. Instead, we adopt a clean and unified ViT‑based model capable of recovering coarse 3D human geometry. To compensate for the loss of fine‑grained details caused by this simplified architecture, we introduce pixel‑level supervision to optimize the geometry, significantly improving the reconstruction accuracy of fine‑grained human details. To make this approach practical, we further propose a modular data annotation strategy that enriches the training data and enhances the robustness of the model. Overall, PEAR is a preprocessing‑free framework that can simultaneously infer EHM‑s (SMPLX and scaled‑FLAME) parameters at over 100 FPS. Extensive experiments on multiple benchmark datasets demonstrate that our method achieves substantial improvements in pose estimation accuracy compared to previous SMPLX‑based approaches. Project page: https://wujh2001.github.io/PEAR
Authors:Binyi Su, Chenghao Huang, Haiyong Chen
Abstract:
Zero‑shot out‑of‑vocabulary detection (ZS‑OOVD) aims to accurately recognize objects of in‑vocabulary (IV) categories provided at zero‑shot inference, while simultaneously rejecting undefined ones (out‑of‑vocabulary, OOV) that lack corresponding category prompts. However, previous methods are prone to overfitting the IV classes, leading to the OOV or undefined classes being misclassified as IV ones with a high confidence score. To address this issue, this paper proposes a zero‑shot OOV detector (OOVDet), a novel framework that effectively detects predefined classes while reliably rejecting undefined ones in zero‑shot scenes. Specifically, due to the model's lack of prior knowledge about the distribution of OOV data, we synthesize region‑level OOV prompts by sampling from the low‑likelihood regions of the class‑conditional Gaussian distributions in the hidden space, motivated by the assumption that unknown semantics are more likely to emerge in low‑density areas of the latent space. For OOV images, we further propose a Dirichlet‑based gradient attribution mechanism to mine pseudo‑OOV image samples, where the attribution gradients are interpreted as Dirichlet evidence to estimate prediction uncertainty, and samples with high uncertainty are selected as pseudo‑OOV images. Building on these synthesized OOV prompts and pseudo‑OOV images, we construct the OOV decision boundary through a low‑density prior constraint, which regularizes the optimization of OOV classes using Gaussian kernel density estimation in accordance with the above assumption.
Experimental results show that our method significantly improves the OOV detection performance in zero‑shot scenes. The code is available at https://github.com/binyisu/OOV‑detector.
Authors:Rameen Abdal, James Burgess, Sergey Tulyakov, Kuan-Chieh Jackson Wang
Abstract:
We introduce the Visual Personalization Turing Test (VPTT), a new paradigm for evaluating contextual visual personalization based on perceptual indistinguishability, rather than identity replication. A model passes the VPTT if its output (image, video, 3D asset, etc.) is indistinguishable to a human or calibrated VLM judge from content a given person might plausibly create or share. To operationalize VPTT, we present the VPTT Framework, integrating a 10k‑persona benchmark (VPTT‑Bench), a visual retrieval‑augmented generator (VPRAG), and the VPTT Score, a text‑only metric calibrated against human and VLM judgments. We show high correlation across human, VLM, and VPTT evaluations, validating the VPTT Score as a reliable perceptual proxy. Experiments demonstrate that VPRAG achieves the best alignment‑originality balance, offering a scalable and privacy‑safe foundation for personalized generative AI.
Authors:Hanxun Yu, Wentong Li, Xuan Qu, Song Wang, Junbo Chen, Jianke Zhu
Abstract:
Multimodal large language models (MLLMs) suffer from high computational costs due to excessive visual tokens, particularly in high‑resolution and video‑based scenarios. Existing token reduction methods typically focus on isolated pipeline components and often neglect textual alignment, leading to performance degradation. In this paper, we propose VisionTrim, a unified framework for training‑free MLLM acceleration, integrating two effective plug‑and‑play modules: 1) the Dominant Vision Token Selection (DVTS) module, which preserves essential visual tokens via a global‑local view, and 2) the Text‑Guided Vision Complement (TGVC) module, which facilitates context‑aware token merging guided by textual cues. Extensive experiments across diverse image and video multimodal benchmarks demonstrate the performance superiority of our VisionTrim, advancing practical MLLM deployment in real‑world applications. The code is available at: https://github.com/hanxunyu/VisionTrim.
Authors:Jiahao Wang, Ting Pan, Haoge Deng, Dongchen Han, Taiqiang Wu, Xinlong Wang, Ping Luo
Abstract:
Autoregressive models with continuous tokens form a promising paradigm for visual generation, especially for text‑to‑image (T2I) synthesis, but they suffer from high computational cost. We study how to design compute‑efficient linear attention within this framework. Specifically, we conduct a systematic empirical analysis of scaling behavior with respect to parameter counts under different design choices, focusing on (1) normalization paradigms in linear attention (division‑based vs. subtraction‑based) and (2) depthwise convolution for locality augmentation.
Our results show that although subtraction‑based normalization is effective for image classification, division‑based normalization scales better for linear generative transformers. In addition, incorporating convolution for locality modeling plays a crucial role in autoregressive generation, consistent with findings in diffusion models.
We further extend gating mechanisms, commonly used in causal linear attention, to the bidirectional setting and propose a KV gate. By introducing data‑independent learnable parameters to the key and value states, the KV gate assigns token‑wise memory weights, enabling flexible memory management similar to forget gates in language models.
Based on these findings, we present LINA, a simple and compute‑efficient T2I model built entirely on linear attention, capable of generating high‑fidelity 1024x1024 images from user instructions. LINA achieves competitive performance on both class‑conditional and T2I benchmarks, obtaining 2.18 FID on ImageNet (about 1.4B parameters) and 0.74 on GenEval (about 1.5B parameters). A single linear attention module reduces FLOPs by about 61 percent compared to softmax attention. Code and models are available at: https://github.com/techmonsterwang/LINA.
Authors:Zhijie Zheng, Xinhao Xiang, Jiawei Zhang
Abstract:
Streaming recurrent models enable efficient 3D reconstruction by maintaining persistent state representations. However, they suffer from catastrophic forgetting over long sequences due to balancing historical information with new observations. Recent methods alleviate this by deriving adaptive signals from attention perspective, but they operate on single dimensions without considering temporal and spatial consistency. To this end, we propose a training‑free framework termed TTSA3R that leverages both temporal state evolution and spatial observation quality for adaptive state updates in 3D reconstruction. In particular, we devise a Temporal Adaptive Update Module that regulates update magnitude by analyzing temporal state evolution patterns. Then, a Spatial Contextual Update Module is introduced to localize spatial regions that require updates through observation‑state alignment and scene dynamics. These complementary signals are finally fused to determine the state updating strategies. Extensive experiments demonstrate the effectiveness of TTSA3R in diverse 3D tasks. Moreover, our method exhibits only 1.33x error increase compared to over 4x degradation in the baseline model on extended sequences of 3D reconstruction, significantly improving long‑term reconstruction stability. Our codes are available at https://github.com/anonus2357/ttsa3r.
Authors:Naeem Paeedeh, Mahardhika Pratama, Ary Shiddiqi, Zehong Cao, Mukesh Prasad, Wisnu Jatmiko
Abstract:
Although cross‑domain few‑shot learning (CDFSL) for hyper‑spectral image (HSI) classification has attracted significant research interest, existing works often rely on an unrealistic data augmentation procedure in the form of external noise to enlarge the sample size, thus greatly simplifying the issue of data scarcity. They involve a large number of parameters for model updates, being prone to the overfitting problem. To the best of our knowledge, none has explored the strength of the foundation model, having strong generalization power to be quickly adapted to downstream tasks. This paper proposes the MIxup FOundation MOdel (MIFOMO) for CDFSL of HSI classifications. MIFOMO is built upon the concept of a remote sensing (RS) foundation model, pre‑trained across a large scale of RS problems, thus featuring generalizable features. The notion of coalescent projection (CP) is introduced to quickly adapt the foundation model to downstream tasks while freezing the backbone network. The concept of mixup domain adaptation (MDM) is proposed to address the extreme domain discrepancy problem. Last but not least, the label smoothing concept is implemented to cope with noisy pseudo‑label problems. Our rigorous experiments demonstrate the advantage of MIFOMO, where it beats prior arts with up to 14% margin. The source code of MIFOMO is open‑sourced at https://github.com/Naeem‑Paeedeh/MIFOMO for reproducibility and convenient further study.
Authors:Hanjiang Zhu, Pedro Martelleto Rezende, Zhang Yang, Tong Ye, Bruce Z. Gao, Feng Luo, Siyu Huang, Jiancheng Yang
Abstract:
This work proposes Bonnet, an ultra‑fast sparse‑volume pipeline for whole‑body bone segmentation from CT scans. Accurate bone segmentation is important for surgical planning and anatomical analysis, but existing 3D voxel‑based models such as nnU‑Net and STU‑Net require heavy computation and often take several minutes per scan, which limits time‑critical use. The proposed Bonnet addresses this by integrating a series of novel framework components including HU‑based bone thresholding, patch‑wise inference with a sparse spconv‑based U‑Net, and multi‑window fusion into a full‑volume prediction. Trained on TotalSegmentator and evaluated without additional tuning on RibSeg, CT‑Pelvic1K, and CT‑Spine1K, Bonnet achieves high Dice across ribs, pelvis, and spine while running in only 2.69 seconds per scan on an RTX A6000. Compared to strong voxel baselines, Bonnet attains a similar accuracy but reduces inference time by roughly 25x on the same hardware and tiling setup. The toolkit and pre‑trained models will be released at https://github.com/HINTLab/Bonnet.
Authors:Xudong Lu, Huankang Guan, Yang Bo, Jinpeng Chen, Xintong Guo, Shuhan Li, Fang Liu, Peiwen Sun, Xueying Li, Wei Zhang, Xue Yang, Rui Liu, Hongsheng Li
Abstract:
Multimodal Large Language Models excel at offline audio‑visual understanding, but their ability to serve as mobile assistants in continuous real‑world streams remains underexplored. In daily phone use, mobile assistants must track streaming audio‑visual inputs and respond at the right time, yet existing benchmarks are often restricted to multiple‑choice questions or use shorter videos. In this paper, we introduce PhoStream, the first mobile‑centric streaming benchmark that unifies on‑screen and off‑screen scenarios to evaluate video, audio, and temporal reasoning. PhoStream contains 5,572 open‑ended QA pairs from 578 videos across 4 scenarios and 10 capabilities. We build it with an Automated Generative Pipeline backed by rigorous human verification, and evaluate models using a realistic Online Inference Pipeline and LLM‑as‑a‑Judge evaluation for open‑ended responses. Experiments reveal a temporal asymmetry in LLM‑judged scores (0‑100): models perform well on Instant and Backward tasks (Gemini 3 Pro exceeds 80), but drop sharply on Forward tasks (16.40), largely due to early responses before the required visual and audio cues appear. This highlights a fundamental limitation: current MLLMs struggle to decide when to speak, not just what to say. Code and datasets used in this work will be made publicly accessible at https://github.com/Lucky‑Lance/PhoStream.
Authors:Aditya Sarkar, Yi Li, Jiacheng Cheng, Shlok Mishra, Nuno Vasconcelos
Abstract:
Selective prediction aims to endow predictors with a reject option, to avoid low confidence predictions. However, existing literature has primarily focused on closed‑set tasks, such as visual question answering with predefined options or fixed‑category classification. This paper considers selective prediction for visual language foundation models, addressing a taxonomy of tasks ranging from closed to open set and from finite to unbounded vocabularies, as in image captioning. We seek training‑free approaches of low‑complexity, applicable to any foundation model and consider methods based on external vision‑language model embeddings, like CLIP. This is denoted as Plug‑and‑Play Selective Prediction (PaPSP). We identify two key challenges: (1) instability of the visual‑language representations, leading to high variance in image‑text embeddings, and (2) poor calibration of similarity scores. To address these issues, we propose a memory augmented PaPSP (MA‑PaPSP) model, which augments PaPSP with a retrieval dataset of image‑text pairs. This is leveraged to reduce embedding variance by averaging retrieved nearest‑neighbor pairs and is complemented by the use of contrastive normalization to improve score calibration. Through extensive experiments on multiple datasets, we show that MA‑PaPSP outperforms PaPSP and other selective prediction baselines for selective captioning, image‑text matching, and fine‑grained classification. Code is publicly available at https://github.com/kingston‑aditya/MA‑PaPSP.
Authors:Zhuoyu Wu, Wenhui Ou, Pei-Sze Tan, Jiayan Yang, Wenqi Fang, Zheng Wang, Raphaël C. -W. Phan
Abstract:
Endoscopic image analysis is vital for colorectal cancer screening, yet real‑world conditions often suffer from lens fogging, motion blur, and specular highlights, which severely compromise automated polyp detection. We propose EndoCaver, a lightweight transformer with a unidirectional‑guided dual‑decoder architecture, enabling joint multi‑task capability for image deblurring and segmentation while significantly reducing computational complexity and model parameters. Specifically, it integrates a Global Attention Module (GAM) for cross‑scale aggregation, a Deblurring‑Segmentation Aligner (DSA) to transfer restoration cues, and a cosine‑based scheduler (LoCoS) for stable multi‑task optimisation. Experiments on the Kvasir‑SEG dataset show that EndoCaver achieves 0.922 Dice on clean data and 0.889 under severe image degradation, surpassing state‑of‑the‑art methods while reducing model parameters by 90%. These results demonstrate its efficiency and robustness, making it well‑suited for on‑device clinical deployment. Code is available at https://github.com/ReaganWu/EndoCaver.
Authors:Gyuwon Han, Young Kyun Jang, Chanho Eom
Abstract:
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing benchmarks consider only visual changes, ignoring videos that differ in audio despite visual similarity. To address this limitation, we introduce Composed retrieval for Video with its Audio CoVA, a new retrieval task that accounts for both visual and auditory variations. To support this, we construct AV‑Comp, a benchmark consisting of video pairs with cross‑modal changes and corresponding textual queries that describe the differences. We also propose AVT Compositional Fusion (AVT), which integrates video, audio, and text features by selectively aligning the query to the most relevant modality. AVT outperforms traditional unimodal fusion and serves as a strong baseline for CoVA. Examples from the proposed dataset, including both visual and auditory information, are available at https://perceptualai‑lab.github.io/CoVA/.
Authors:Shiyu Liu, Xinyi Wen, Zhibin Lan, Ante Wang, Jinsong Su
Abstract:
Despite progress in Large Vision Language Models (LVLMs), object hallucination remains a critical issue in image captioning task, where models generate descriptions of non‑existent objects, compromising their reliability. Previous work attributes this to LVLMs' over‑reliance on language priors and attempts to mitigate it through logits calibration. However, they still lack a thorough analysis of the over‑reliance. To gain a deeper understanding of over‑reliance, we conduct a series of preliminary experiments, indicating that as the generation length increases, LVLMs' over‑reliance on language priors leads to inflated probability of hallucinated object tokens, consequently exacerbating object hallucination. To circumvent this issue, we propose Language‑Prior‑Free Verification to enable LVLMs to faithfully verify the confidence of object existence. Based on this, we propose a novel training‑free Self‑Validation Framework to counter the over‑reliance trap. It first validates objects' existence in sampled candidate captions and further mitigates object hallucination via caption selection or aggregation. Experiment results demonstrate that our framework mitigates object hallucination significantly in image captioning task (e.g., 65.6% improvement on CHAIRI metric with LLaVA‑v1.5‑7B), surpassing the previous SOTA methods. This result highlights a novel path towards mitigating hallucination by unlocking the inherent potential within LVLMs themselves.
Authors:Gonzalo Gomez-Nogales, Yicong Hong, Chongjian Ge, Peiye Zhuang, Marc Comino-Trinidad, Dan Casas, Yi Zhou
Abstract:
Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic scenes. We present C2R (Coarse‑to‑Real), a generative rendering framework that synthesizes real‑style urban crowd videos from coarse 3D simulations. Our approach uses coarse 3D renderings to explicitly control scene layout, camera motion, and human trajectories, while a learned neural renderer generates realistic appearance, lighting, and fine‑scale dynamics guided by text prompts. To overcome the lack of paired training data between coarse simulations and real videos, we adopt a two‑stage synthetic‑real domain‑hedging strategy that first learns a strong generative prior from large‑scale real footage, and then introduces controllability by using a small amount of paired synthetic coarse‑to‑fine data to anchor shared implicit spatio‑temporal features across domains. The resulting system supports coarse‑to‑fine control, generalizes across diverse CG and game inputs, and produces temporally consistent, controllable, and realistic urban scene videos from minimal 3D input. We will release the model and project webpage at https://gonzalognogales.github.io/coarse2real/.
Authors:Shirin Reyhanian, Laurenz Wiskott
Abstract:
Vector‑quantized variational autoencoders (VQ‑VAEs) are central to models that rely on high reconstruction fidelity, from neural compression to generative pipelines. Hierarchical extensions, such as VQ‑VAE2, are often credited with superior reconstruction performance because they split global and local features across multiple levels. However, since higher levels derive all their information from lower levels, they should not carry additional reconstructive content beyond what the lower‑level already encodes. Combined with recent advances in training objectives and quantization mechanisms, this leads us to ask whether a single‑level VQ‑VAE, with matched representational budget and no codebook collapse, can equal the reconstruction fidelity of its hierarchical counterpart. Although the multi‑scale structure of hierarchical models may improve perceptual quality in downstream tasks, the effect of hierarchy on reconstruction accuracy, isolated from codebook utilization and overall representational capacity, remains empirically underexamined. We revisit this question by comparing a two‑level VQ‑VAE and a capacity‑matched single‑level model on high‑resolution ImageNet images. Consistent with prior observations, we confirm that inadequate codebook utilization limits single‑level VQ‑VAEs and that overly high‑dimensional embeddings destabilize quantization and increase codebook collapse. We show that lightweight interventions such as initialization from data, periodic reset of inactive codebook vectors, and systematic tuning of codebook hyperparameters significantly reduce collapse. Our results demonstrate that when representational budgets are matched, and codebook collapse is mitigated, single‑level VQ‑VAEs can match the reconstruction fidelity of hierarchical variants, challenging the assumption that hierarchical quantization is inherently superior for high‑quality reconstructions.
Authors:Jian Shi, Michael Birsak, Wenqing Cui, Zhenyu Li, Peter Wonka
Abstract:
This paper revisits the role of positional embeddings (PEs) within vision transformers (ViTs) from a geometric perspective. We show that PEs are not mere token indices but effectively function as geometric priors that shape the spatial structure of the representation. We introduce token‑level diagnostics that measure how multi‑view geometric consistency in ViT representation depends on consitent PEs. Through extensive experiments on 14 foundation ViT models, we reveal how PEs influence multi‑view geometry and spatial reasoning. Our findings clarify the role of PEs as a causal mechanism that governs spatial structure in ViT representations. Our code is provided in https://github.com/shijianjian/vit‑geometry‑probes
Authors:Yiyang Lu, Susie Lu, Qiao Sun, Hanhong Zhao, Zhicheng Jiang, Xianbang Wang, Tianhong Li, Zhengyang Geng, Kaiming He
Abstract:
Modern diffusion/flow‑based models for image generation typically exhibit two core characteristics: (i) using multi‑step sampling, and (ii) operating in a latent space. Recent advances have made encouraging progress on each aspect individually, paving the way toward one‑step diffusion/flow without latents. In this work, we take a further step towards this goal and propose "pixel MeanFlow" (pMF). Our core guideline is to formulate the network output space and the loss space separately. The network target is designed to be on a presumed low‑dimensional image manifold (i.e., x‑prediction), while the loss is defined via MeanFlow in the velocity space. We introduce a simple transformation between the image manifold and the average velocity field. In experiments, pMF achieves strong results for one‑step latent‑free generation on ImageNet at 256x256 resolution (2.22 FID) and 512x512 resolution (2.48 FID), filling a key missing piece in this regime. We hope that our study will further advance the boundaries of diffusion/flow‑based generative models.
Authors:Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu
Abstract:
Manipulating dynamic objects remains an open challenge for Vision‑Language‑Action (VLA) models, which, despite strong generalization in static manipulation, struggle in dynamic scenarios requiring rapid perception, temporal anticipation, and continuous control. We present DynamicVLA, a framework for dynamic object manipulation that integrates temporal reasoning and closed‑loop adaptation through three key designs: 1) a compact 0.4B VLA using a convolutional vision encoder for spatially efficient, structurally faithful encoding, enabling fast multimodal inference; 2) Continuous Inference, enabling overlapping reasoning and execution for lower latency and timely adaptation to object motion; and 3) Latent‑aware Action Streaming, which bridges the perception‑execution gap by enforcing temporally aligned action execution. To fill the missing foundation of dynamic manipulation data, we introduce the Dynamic Object Manipulation (DOM) benchmark, built from scratch with an auto data collection pipeline that efficiently gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real‑world episodes without teleoperation. Extensive evaluations demonstrate remarkable improvements in response speed, perception, and generalization, positioning DynamicVLA as a unified framework for general dynamic object manipulation across embodiments.
Authors:Hanzhuo Huang, Qingyang Bao, Zekai Gu, Zhongshuo Du, Cheng Lin, Yuan Liu, Sibei Yang
Abstract:
In this paper, we propose a 3D asset‑referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference‑based image generation methods leverage large‑scale pretrained diffusion models and demonstrate strong capability in generating diverse images conditioned on a single reference image. However, these methods are limited to single‑image references and cannot leverage 3D assets, constraining their practical versatility. To address this gap, we present a cross‑domain diffusion model with dual‑branch perception that leverages multi‑view RGB images and point maps of 3D assets to jointly model their colors and canonical‑space coordinates, achieving precise consistency between generated images and the 3D references. Our spatially aligned dual‑branch generation architecture and domain‑decoupled generation mechanism ensure the simultaneous generation of two spatially aligned but content‑disentangled outputs, RGB images and point maps, linking 2D image attributes with 3D asset attributes. Experiments show that our approach effectively uses 3D assets as references to produce images consistent with the given assets, opening new possibilities for combining diffusion models with 3D content creation.
Authors:Wenxuan Huang, Yu Zeng, Qiuchen Wang, Zhen Fang, Shaosheng Cao, Zheng Chu, Qingyu Yin, Shuang Chen, Zhenfei Yin, Lin Chen, Zehui Chen, Xu Tang, Yao Hu, Shaohui Lin, Philip Torr, Feng Zhao, Wanli Ouyang
Abstract:
Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting MLLMs by ``reasoning‑then‑tool‑call'' for visual and textual search engines to obtain substantial gains on tasks requiring extensive factual information. However, these approaches typically define multimodal search in a naive setting, assuming that a single full‑level or entity‑level image query and few text query suffices to retrieve the key evidence needed to answer the question, which is unrealistic in real‑world scenarios with substantial visual noise. Moreover, they are often limited in the reasoning depth and search breadth, making it difficult to solve complex questions that require aggregating evidence from diverse visual and textual sources. Building on this, we propose Vision‑DeepResearch, which proposes one new multimodal deep‑research paradigm, i.e., performs multi‑turn, multi‑entity and multi‑scale visual and textual search to robustly hit real‑world search engines under heavy noise. Our Vision‑DeepResearch supports dozens of reasoning steps and hundreds of engine interactions, while internalizing deep‑research capabilities into the MLLM via cold‑start supervision and RL training, resulting in a strong end‑to‑end multimodal deep‑research MLLM. It substantially outperforming existing multimodal deep‑research MLLMs, and workflows built on strong closed‑source foundation model such as GPT‑5, Gemini‑2.5‑pro and Claude‑4‑Sonnet. The code will be released in https://github.com/Osilly/Vision‑DeepResearch.
Authors:Baorui Ma, Jiahui Yang, Donglin Di, Xuancheng Zhang, Jianxun Cui, Hao Li, Yan Xie, Wei Chen
Abstract:
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camera‑dependent biases, and metric ambiguity in noisy cross‑source 3D data. We introduce Metric Anything, a simple and scalable pretraining framework that learns metric depth from noisy, diverse 3D sources without manually engineered prompts, camera‑specific modeling, or task‑specific architectures. Central to our approach is the Sparse Metric Prompt, created by randomly masking depth maps, which serves as a universal interface that decouples spatial reasoning from sensor and camera biases. Using about 20M image‑depth pairs spanning reconstructed, captured, and rendered 3D data across 10000 camera models, we demonstrate‑for the first time‑a clear scaling trend in the metric depth track. The pretrained model excels at prompt‑driven tasks such as depth completion, super‑resolution and Radar‑camera fusion, while its distilled prompt‑free student achieves state‑of‑the‑art results on monocular depth estimation, camera intrinsics recovery, single/multi‑view metric 3D reconstruction, and VLA planning. We also show that using pretrained ViT of Metric Anything as a visual encoder significantly boosts Multimodal Large Language Model capabilities in spatial intelligence. These results show that metric depth estimation can benefit from the same scaling laws that drive modern foundation models, establishing a new path toward scalable and efficient real‑world metric perception. We open‑source MetricAnything at http://metric‑anything.github.io/metric‑anything‑io/ to support community research.
Authors:Changjian Jiang, Kerui Ren, Xudong Li, Kaiwen Song, Guanghao Li, Linning Xu, Tao Lu, Junting Dong, Yu Zhang, Bo Dai, Mulin Yu
Abstract:
Streaming reconstruction from monocular image sequences remains challenging, as existing methods typically favor either high‑quality rendering or accurate geometry, but rarely both. We present PLANING, an efficient on‑the‑fly reconstruction framework built on a hybrid representation that loosely couples explicit geometric primitives with neural Gaussians, enabling geometry and appearance to be modeled in a decoupled manner. This decoupling supports an online initialization and optimization strategy that separates geometry and appearance updates, yielding stable streaming reconstruction with substantially reduced structural redundancy. PLANING improves dense mesh Chamfer‑L2 by 18.52% over PGSR, surpasses ARTDECO by 1.31 dB PSNR, and reconstructs ScanNetV2 scenes in under 100 seconds, over 5x faster than 2D Gaussian Splatting, while matching the quality of offline per‑scene optimization. Beyond reconstruction quality, the structural clarity and computational efficiency of PLANING make it well suited for a broad range of downstream applications, such as enabling large‑scale scene modeling and simulation‑ready environments for embodied AI. Project page: https://city‑super.github.io/PLANING/ .
Authors:Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, Yinghao Xu
Abstract:
This work highlights that video world modeling, alongside vision‑language pre‑training, establishes a fresh and independent foundation for robot learning. Intuitively, video world models provide the ability to imagine the near future by understanding the causality between actions and visual dynamics. Inspired by this, we introduce LingBot‑VA, an autoregressive diffusion framework that learns frame prediction and policy execution simultaneously. Our model features three carefully crafted designs: (1) a shared latent space, integrating vision and action tokens, driven by a Mixture‑of‑Transformers (MoT) architecture, (2) a closed‑loop rollout mechanism, allowing for ongoing acquisition of environmental feedback with ground‑truth observations, (3) an asynchronous inference pipeline, parallelizing action prediction and motor execution to support efficient control. We evaluate our model on both simulation benchmarks and real‑world scenarios, where it shows significant promise in long‑horizon manipulation, data efficiency in post‑training, and strong generalizability to novel configurations. The code and model are made publicly available to facilitate the community.
Authors:Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, Yue Zhang, Yubo Zhang, Yi Liu, Dianhai Yu, Yanjun Ma
Abstract:
We introduce PaddleOCR‑VL‑1.5, an upgraded model achieving a new state‑of‑the‑art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real‑world physical distortions, including scanning, skew, warping, screen‑photography, and illumination, we propose the Real5‑OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model's capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra‑compact VLM with high efficiency. Code: https://github.com/PaddlePaddle/PaddleOCR
Authors:Yunhao Li, Sijing Wu, Zhilin Gao, Zicheng Zhang, Qi Jia, Huiyu Duan, Xiongkuo Min, Guangtao Zhai
Abstract:
Large multimodal models (LMMs) have demonstrated outstanding capabilities in various visual perception tasks, which has in turn made the evaluation of LMMs significant. However, the capability of video aesthetic quality assessment, which is a fundamental ability for human, remains underexplored for LMMs. To address this, we introduce VideoAesBench, a comprehensive benchmark for evaluating LMMs' understanding of video aesthetic quality. VideoAesBench has several significant characteristics: (1) Diverse content including 1,804 videos from multiple video sources including user‑generated (UGC), AI‑generated (AIGC), compressed, robotic‑generated (RGC), and game videos. (2) Multiple question formats containing traditional single‑choice questions, multi‑choice questions, True or False questions, and a novel open‑ended questions for video aesthetics description. (3) Holistic video aesthetics dimensions including visual form related questions from 5 aspects, visual style related questions from 4 aspects, and visual affectiveness questions from 3 aspects. Based on VideoAesBench, we benchmark 23 open‑source and commercial large multimodal models. Our findings show that current LMMs only contain basic video aesthetics perception ability, their performance remains incomplete and imprecise. We hope our VideoAesBench can be served as a strong testbed and offer insights for explainable video aesthetics assessment. The data will be released on https://github.com/michaelliyunhao/VideoAesBench
Authors:Borja Carrillo-Perez, Felix Sattler, Angel Bueno Rodriguez, Maurice Stephan, Sarah Barnes
Abstract:
Three‑dimensional (3D) reconstruction of ships is an important part of maritime monitoring, allowing improved visualization, inspection, and decision‑making in real‑world monitoring environments. However, most state‑ofthe‑art 3D reconstruction methods require multi‑view supervision, annotated 3D ground truth, or are computationally intensive, making them impractical for real‑time maritime deployment. In this work, we present an efficient pipeline for single‑view 3D reconstruction of real ships by training entirely on synthetic data and requiring only a single view at inference. Our approach uses the Splatter Image network, which represents objects as sparse sets of 3D Gaussians for rapid and accurate reconstruction from single images. The model is first fine‑tuned on synthetic ShapeNet vessels and further refined with a diverse custom dataset of 3D ships, bridging the domain gap between synthetic and real‑world imagery. We integrate a state‑of‑the‑art segmentation module based on YOLOv8 and custom preprocessing to ensure compatibility with the reconstruction network. Postprocessing steps include real‑world scaling, centering, and orientation alignment, followed by georeferenced placement on an interactive web map using AIS metadata and homography‑based mapping. Quantitative evaluation on synthetic validation data demonstrates strong reconstruction fidelity, while qualitative results on real maritime images from the ShipSG dataset confirm the potential for transfer to operational maritime settings. The final system provides interactive 3D inspection of real ships without requiring real‑world 3D annotations. This pipeline provides an efficient, scalable solution for maritime monitoring and highlights a path toward real‑time 3D ship visualization in practical applications. Interactive demo: https://dlr‑mi.github.io/ship3d‑demo/.
Authors:Jiankun Peng, Jianyuan Guo, Ying Xu, Yue Liu, Jiashuang Yan, Xuanwei Ye, Houhua Li, Xiaoming Wang
Abstract:
Vision‑Language Navigation in Continuous Environments (VLN‑CE) presents a core challenge: grounding high‑level linguistic instructions into precise, safe, and long‑horizon spatial actions. Explicit topological maps have proven to be a vital solution for providing robust spatial memory in such tasks. However, existing topological planning methods suffer from a "Granularity Rigidity" problem. Specifically, these methods typically rely on fixed geometric thresholds to sample nodes, which fails to adapt to varying environmental complexities. This rigidity leads to a critical mismatch: the model tends to over‑sample in simple areas, causing computational redundancy, while under‑sampling in high‑uncertainty regions, increasing collision risks and compromising precision. To address this, we propose DGNav, a framework for Dynamic Topological Navigation, introducing a context‑aware mechanism to modulate map density and connectivity on‑the‑fly. Our approach comprises two core innovations: (1) A Scene‑Aware Adaptive Strategy that dynamically modulates graph construction thresholds based on the dispersion of predicted waypoints, enabling "densification on demand" in challenging environments; (2) A Dynamic Graph Transformer that reconstructs graph connectivity by fusing visual, linguistic, and geometric cues into dynamic edge weights, enabling the agent to filter out topological noise and enhancing instruction adherence. Extensive experiments on the R2R‑CE and RxR‑CE benchmarks demonstrate DGNav exhibits superior navigation performance and strong generalization capabilities. Furthermore, ablation studies confirm that our framework achieves an optimal trade‑off between navigation efficiency and safe exploration. The code is available at https://github.com/shannanshouyin/DGNav.
Authors:Baoliang Chen, Danni Huang, Hanwei Zhu, Lingyu Zhu, Wei Zhou, Shiqi Wang, Yuming Fang, Weisi Lin
Abstract:
Evaluation of Image Quality Assessment (IQA) models has long been dominated by global correlation metrics, such as Pearson Linear Correlation Coefficient (PLCC) and Spearman Rank‑Order Correlation Coefficient (SRCC). While widely adopted, these metrics reduce performance to a single scalar, failing to capture how ranking consistency varies across the local quality spectrum. For example, two IQA models may achieve identical SRCC values, yet one ranks high‑quality images (related to high Mean Opinion Score, MOS) more reliably, while the other better discriminates image pairs with small quality/MOS differences (related to |ΔMOS|). Such complementary behaviors are invisible under global metrics. Moreover, SRCC and PLCC are sensitive to test‑sample quality distributions, yielding unstable comparisons across test sets. To address these limitations, we propose Granularity‑Modulated Correlation (GMC), which provides a structured, fine‑grained analysis of IQA performance. GMC includes: (1) a Granularity Modulator that applies Gaussian‑weighted correlations conditioned on absolute MOS values and pairwise MOS differences (|ΔMOS|) to examine local performance variations, and (2) a Distribution Regulator that regularizes correlations to mitigate biases from non‑uniform quality distributions. The resulting correlation surface maps correlation values as a joint function of MOS and |ΔMOS|, providing a 3D representation of IQA performance. Experiments on standard benchmarks show that GMC reveals performance characteristics invisible to scalar metrics, offering a more informative and reliable paradigm for analyzing, comparing, and deploying IQA models. Codes are available at https://github.com/Dniaaa/GMC.
Authors:Mingshuang Luo, Shuang Liang, Zhengkun Rong, Yuxuan Luo, Tianshu Hu, Ruibing Hou, Hong Chang, Yong Li, Yuan Zhang, Mingyuan Gao
Abstract:
Character image animation aims to synthesize high‑fidelity videos by transferring motion from a driving sequence to a static reference image. Despite recent advancements, existing methods suffer from two fundamental challenges: (1) suboptimal motion injection strategies that lead to a trade‑off between identity preservation and motion consistency, manifesting as a "see‑saw", and (2) an over‑reliance on explicit pose priors (e.g., skeletons), which inadequately capture intricate dynamics and hinder generalization to arbitrary, non‑humanoid characters. To address these challenges, we present DreamActor‑M2, a universal animation framework that reimagines motion conditioning as an in‑context learning problem. Our approach follows a two‑stage paradigm. First, we bridge the input modality gap by fusing reference appearance and motion cues into a unified latent space, enabling the model to jointly reason about spatial identity and temporal dynamics by leveraging the generative prior of foundational models. Second, we introduce a self‑bootstrapped data synthesis pipeline that curates pseudo cross‑identity training pairs, facilitating a seamless transition from pose‑dependent control to direct, end‑to‑end RGB‑driven animation. This strategy significantly enhances generalization across diverse characters and motion scenarios. To facilitate comprehensive evaluation, we further introduce AW Bench, a versatile benchmark encompassing a wide spectrum of characters types and motion scenarios. Extensive experiments demonstrate that DreamActor‑M2 achieves state‑of‑the‑art performance, delivering superior visual fidelity and robust cross‑domain generalization. Project Page: https://grisoon.github.io/DreamActor‑M2/
Authors:Shuo Li, Jiajun Sun, Zhekai Wang, Xiaoran Fan, Hui Li, Dingwen Yang, Zhiheng Xi, Yijun Wang, Zifei Shan, Tao Gui, Qi Zhang, Xuanjing Huang
Abstract:
Charts are a fundamental visualization format for structured data analysis. Enabling end‑to‑end chart editing according to user intent is of great practical value, yet remains challenging due to the need for both fine‑grained control and global structural consistency. Most existing approaches adopt pipeline‑based designs, where natural language or code serves as an intermediate representation, limiting their ability to faithfully execute complex edits. We introduce ChartE^3, an End‑to‑End Chart Editing benchmark that directly evaluates models without relying on intermediate natural language programs or code‑level supervision. ChartE^3 focuses on two complementary editing dimensions: local editing, which involves fine‑grained appearance changes such as font or color adjustments, and global editing, which requires holistic, data‑centric transformations including data filtering and trend line addition. ChartE^3 contains over 1,200 high‑quality samples constructed via a well‑designed data pipeline with human curation. Each sample is provided as a triplet of a chart image, its underlying code, and a multimodal editing instruction, enabling evaluation from both objective and subjective perspectives. Extensive benchmarking of state‑of‑the‑art multimodal large language models reveals substantial performance gaps, particularly on global editing tasks, highlighting critical limitations in current end‑to‑end chart editing capabilities.
Authors:Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
Abstract:
Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, while their ability to process sequential audio‑video data remains underexplored. This gap highlights the need for a high‑quality benchmark to systematically evaluate MLLM performance in a real‑world setting. We introduce SONIC‑O1, a comprehensive, fully human‑verified benchmark of 60 hours (231 clips) spanning 13 real‑world conversational domains with 4,958 annotations and demographic metadata. SONIC‑O1 evaluates three capabilities: open‑ended summarization, multiple‑choice question (MCQ) answering, and temporal localization with supporting rationales (reasoning). Across closed‑ and open‑source models, we find that the MCQ accuracy shows the smallest gap between model families, but the best closed‑source model outperforms the best open‑source model by 22.6% on temporal localization. We further observe accuracy gaps of up to 21.4% on temporal localization across demographic groups, indicating persistent disparities in model behaviour. SONIC‑O1 provides an open evaluation suite for temporally grounded and demographically robust multimodal understanding. SONIC‑O1 is publicly available for research: Project page (https://vectorinstitute.github.io/sonic‑o1/), Dataset (https://huggingface.co/datasets/vector‑institute/sonic‑o1), GitHub (https://github.com/vectorinstitute/sonic‑o1), Leaderboard (https://huggingface.co/spaces/vector‑institute/sonic‑o1‑leaderboard).
Authors:Bowen Zhou, Marc-André Fiedler, Ayoub Al-Hamadi
Abstract:
Depression is a prevalent mental health disorder that severely impairs daily functioning and quality of life. While recent deep learning approaches for depression detection have shown promise, most rely on limited feature types, overlook explicit cross‑modal interactions, and employ simple concatenation or static weighting for fusion. To overcome these limitations, we propose CAF‑Mamba, a novel Mamba‑based cross‑modal adaptive attention fusion framework. CAF‑Mamba not only captures cross‑modal interactions explicitly and implicitly, but also dynamically adjusts modality contributions through a modality‑wise attention mechanism, enabling more effective multimodal fusion. Experiments on two in‑the‑wild benchmark datasets, LMVD and D‑Vlog, demonstrate that CAF‑Mamba consistently outperforms existing methods and achieves state‑of‑the‑art performance. Our code is available at https://github.com/zbw‑zhou/CAF‑Mamba.
Authors:Songhan Jiang, Fengchun Liu, Ziyue Wang, Linghan Cai, Yongbing Zhang
Abstract:
Vision‑Language Models (VLMs) are advancing computational pathology with superior visual understanding capabilities. However, current systems often reduce diagnosis to directly output conclusions without verifiable evidence‑linked reasoning, which severely limits clinical trust and hinders expert error rectification. To address these barriers, we construct PathReasoner, the first large‑scale dataset of whole‑slide image (WSI) reasoning. Unlike previous work reliant on unverified distillation, we develop a rigorous knowledge‑guided generation pipeline. By leveraging medical knowledge graphs, we explicitly align structured pathological findings and clinical reasoning with diagnoses, generating over 20K high‑quality instructional samples. Based on the database, we propose PathReasoner‑R1, which synergizes trajectory‑masked supervised fine‑tuning with reasoning‑oriented reinforcement learning to instill structured chain‑of‑thought capabilities. To ensure medical rigor, we engineer a knowledge‑aware multi‑granular reward function incorporating an Entity Reward mechanism strictly aligned with knowledge graphs. This effectively guides the model to optimize for logical consistency rather than mere outcome matching, thereby enhancing robustness. Extensive experiments demonstrate that PathReasoner‑R1 achieves state‑of‑the‑art performance on both PathReasoner and public benchmarks across various image scales, equipping pathology models with transparent, clinically grounded reasoning capabilities. Dataset and code are available at https://github.com/cyclexfy/PathReasoner‑R1.
Authors:Kaito Shiku, Ichika Seo, Tetsuya Matoba, Rissei Hino, Yasuhiro Nakano, Ryoma Bise
Abstract:
In this paper, we present the first attempt to estimate the necessity of debulking coronary artery calcifications from computed tomography (CT) images. We formulate this task as a Multiple‑instance Learning (MIL) problem. The difficulty of this task lies in that physicians adjust their focus and decision criteria for device usage according to tabular data representing each patient's condition. To address this issue, we propose a hypernetwork‑based adaptive aggregation transformer (HyperAdAgFormer), which adaptively modifies the feature aggregation strategy for each patient based on tabular data through a hypernetwork. The experiments using the clinical dataset demonstrated the effectiveness of HyperAdAgFormer. The code is publicly available at https://github.com/Shiku‑Kaito/HyperAdAgFormer.
Authors:Shohei Enomoto, Shin'ya Yamaguchi
Abstract:
In this paper, we address a fundamental gap between pre‑training and fine‑tuning of deep neural networks: while pre‑training has shifted from unimodal to multimodal learning with enhanced visual understanding, fine‑tuning predominantly remains unimodal, limiting the benefits of rich pre‑trained representations. To bridge this gap, we propose a novel approach that transforms unimodal datasets into multimodal ones using Multimodal Large Language Models (MLLMs) to generate synthetic image captions for fine‑tuning models with a multimodal objective. Our method employs carefully designed prompts incorporating class labels and domain context to produce high‑quality captions tailored for classification tasks. Furthermore, we introduce a supervised contrastive loss function that explicitly encourages clustering of same‑class representations during fine‑tuning, along with a new inference technique that leverages class‑averaged text embeddings from multiple synthetic captions per image. Extensive experiments across 13 image classification benchmarks demonstrate that our approach outperforms baseline methods, with particularly significant improvements in few‑shot learning scenarios. Our work establishes a new paradigm for dataset enhancement that effectively bridges the gap between multimodal pre‑training and fine‑tuning. Our code is available at https://github.com/s‑enmt/MMFT.
Authors:Zihan Su, Hongyang Wei, Kangrui Cen, Yong Wang, Guanhua Chen, Chun Yuan, Xiangxiang Chu
Abstract:
Unified Multimodal Models (UMMs) integrate both visual understanding and generation within a single framework. Their ultimate aspiration is to create a cycle where understanding and generation mutually reinforce each other. While recent post‑training methods have successfully leveraged understanding to enhance generation, the reverse direction of utilizing generation to improve understanding remains largely unexplored. In this work, we propose UniMRG (Unified Multi‑Representation Generation), a simple yet effective architecture‑agnostic post‑training method. UniMRG enhances the understanding capabilities of UMMs by incorporating auxiliary generation tasks. Specifically, we train UMMs to generate multiple intrinsic representations of input images, namely pixel (reconstruction), depth (geometry), and segmentation (structure), alongside standard visual understanding objectives. By synthesizing these diverse representations, UMMs capture complementary information regarding appearance, spatial relations, and structural layout. Consequently, UMMs develop a deeper and more comprehensive understanding of visual inputs. Extensive experiments across diverse UMM architectures demonstrate that our method notably enhances fine‑grained perception, reduces hallucinations, and improves spatial understanding, while simultaneously boosting generation capabilities.
Authors:Xuewen Liu, Zhikai Li, Jing Zhang, Mengjuan Chen, Qingyi Gu
Abstract:
AutoRegressive Visual Generation (ARVG) models retain an architecture compatible with language models, while achieving performance comparable to diffusion‑based models. Quantization is commonly employed in neural networks to reduce model size and computational latency. However, applying quantization to ARVG remains largely underexplored, and existing quantization methods fail to generalize effectively to ARVG models. In this paper, we explore this issue and identify three key challenges: (1) severe outliers at channel‑wise level, (2) highly dynamic activations at token‑wise level, and (3) mismatched distribution information at sample‑wise level. To these ends, we propose PTQ4ARVG, a training‑free post‑training quantization (PTQ) framework consisting of: (1) Gain‑Projected Scaling (GPS) mitigates the channel‑wise outliers, which expands the quantization loss via a Taylor series to quantify the gain of scaling for activation‑weight quantization, and derives the optimal scaling factor through differentiation.(2) Static Token‑Wise Quantization (STWQ) leverages the inherent properties of ARVG, fixed token length and position‑invariant distribution across samples, to address token‑wise variance without incurring dynamic calibration overhead.(3) Distribution‑Guided Calibration (DGC) selects samples that contribute most to distributional entropy, eliminating the sample‑wise distribution mismatch. Extensive experiments show that PTQ4ARVG can effectively quantize the ARVG family models to 8‑bit and 6‑bit while maintaining competitive performance. Code is available at http://github.com/BienLuky/PTQ4ARVG .
Authors:Yuji Lin, Qian Zhao, Zongsheng Yue, Junhui Hou, Deyu Meng
Abstract:
This work studies the challenging problem of acquiring high‑quality underwater images via 4‑D light field (LF) imaging. To this end, we propose GeoDiff‑LF, a novel diffusion‑based framework built upon SD‑Turbo to enhance underwater 4‑D LF imaging by leveraging its spatial‑angular structure. GeoDiff‑LF consists of three key adaptations: (1) a modified U‑Net architecture with convolutional and attention adapters to model geometric cues, (2) a geometry‑guided loss function using tensor decomposition and progressive weighting to regularize global structure, and (3) an optimized sampling strategy with noise prediction to improve efficiency. By integrating diffusion priors and LF geometry, GeoDiff‑LF effectively mitigates color distortion in underwater scenes. Extensive experiments demonstrate that our framework outperforms existing methods across both visual fidelity and quantitative performance, advancing the state‑of‑the‑art in enhancing underwater imaging. The code will be publicly available at https://github.com/linlos1234/GeoDiff‑LF.
Authors:Jianzheng Wang, Huan Ni
Abstract:
High‑resolution remote sensing images contain densely distributed objects with pronounced scale variations and complex boundaries, which impose higher demands on both the geometric localization and semantic prediction capabilities of semantic segmentation models. Existing training‑free open‑vocabulary semantic segmentation (OVSS) methods typically fuse Contrastive Language‑Image Pretraining (CLIP) and vision foundation models (VFMs) using one‑way injection and shallow post‑processing strategies, making it difficult to satisfy these requirements. To address this issue, we propose a spatial‑regularization‑aware dual‑branch collaborative inference framework for training‑free OVSS, termed SDCI. First, during feature encoding, SDCI introduces a cross‑model attention fusion (CAF) module, which guides collaborative inference by injecting self‑attention maps into each other. Second, we propose a bidirectional cross‑graph diffusion refinement (BCDR) module that enhances the reliability of dual‑branch segmentation scores through iterative random‑walk diffusion. Finally, we incorporate low‑level superpixel structures and develop a convex‑optimization‑based superpixel collaborative prediction (CSCP) mechanism to further refine object boundaries. Experiments on multiple remote sensing semantic segmentation benchmarks demonstrate that our method achieves better performance than existing approaches. Our code is available at https://github.com/yu‑ni1989/SDCI.
Authors:Jiaqi Li, Guangming Wang, Shuntian Zheng, Minzhe Ni, Xiaoman Lu, Guanghui Ye, Yu Guan
Abstract:
Temporal Action Localization (TAL) requires identifying both the boundaries and categories of actions in untrimmed videos. While vision‑language models (VLMs) offer rich semantics to complement visual evidence, existing approaches tend to overemphasize linguistic priors at the expense of visual performance, leading to a pronounced modality bias. We propose ActionVLM, a vision‑language aggregation framework that systematically mitigates modality bias in TAL. Our key insight is to preserve vision as the dominant signal while adaptively exploiting language only when beneficial. To this end, we introduce (i) a debiasing reweighting module that estimates the language advantage‑the incremental benefit of language over vision‑only predictions‑and dynamically reweights language modality accordingly, and (ii) a residual aggregation strategy that treats language as a complementary refinement rather than the primary driver. This combination alleviates modality bias, reduces overconfidence from linguistic priors, and strengthens temporal reasoning. Experiments on THUMOS14 show that our model outperforms state‑of‑the‑art by up to 3.2% mAP. Our code is available at https://github.com/JiaqiLi404/ActionVLM
Authors:Hongyu Zhou, Zisen Shao, Sheng Miao, Pan Wang, Dongfeng Bai, Bingbing Liu, Yiyi Liao
Abstract:
Neural Radiance Fields and 3D Gaussian Splatting have advanced novel view synthesis, yet still rely on dense inputs and often degrade at extrapolated views. Recent approaches leverage generative models, such as diffusion models, to provide additional supervision, but face a trade‑off between generalization and fidelity: fine‑tuning diffusion models for artifact removal improves fidelity but risks overfitting, while fine‑tuning‑free methods preserve generalization but often yield lower fidelity. We introduce FreeFix, a fine‑tuning‑free approach that pushes the boundary of this trade‑off by enhancing extrapolated rendering with pretrained image diffusion models. We present an interleaved 2D‑3D refinement strategy, showing that image diffusion models can be leveraged for consistent refinement without relying on costly video diffusion models. Furthermore, we take a closer look at the guidance signal for 2D refinement and propose a per‑pixel confidence mask to identify uncertain regions for targeted improvement. Experiments across multiple datasets show that FreeFix improves multi‑frame consistency and achieves performance comparable to or surpassing fine‑tuning‑based methods, while retaining strong generalization ability.
Authors:Hao Sun, Da-Wei Zhou
Abstract:
Traditional machine learning systems are typically designed for static data distributions, which suffer from catastrophic forgetting when learning from evolving data streams. Class‑Incremental Learning (CIL) addresses this challenge by enabling learning systems to continuously learn new classes while preserving prior knowledge. With the rise of pre‑trained models (PTMs) such as CLIP, leveraging their strong generalization and semantic alignment capabilities has become a promising direction in CIL. However, existing CLIP‑based CIL methods are often scattered across disparate codebases, rely on inconsistent configurations, hindering fair comparisons, reproducibility, and practical adoption. Therefore, we propose C3Box (CLIP‑based Class‑inCremental learning toolBOX), a modular and comprehensive Python toolbox. C3Box integrates representative traditional CIL methods, ViT‑based CIL methods, and state‑of‑the‑art CLIP‑based CIL methods into a unified CLIP‑based framework. By inheriting the streamlined design of PyCIL, C3Box provides a JSON‑based configuration and standardized execution pipeline. This design enables reproducible experimentation with low engineering overhead and makes C3Box a reliable benchmark platform for continual learning research. Designed to be user‑friendly, C3Box relies only on widely used open‑source libraries and supports major operating systems. The code is available at https://github.com/LAMDA‑CL/C3Box.
Authors:Ziwei Liu, Borui Kang, Hangjie Yuan, Zixiang Zhao, Wei Li, Yifan Zhu, Tao Feng
Abstract:
As digital environments (data distribution) are in flux, with new GUI data arriving over time‑introducing new domains or resolutions‑agents trained on static environments deteriorate in performance. In this work, we introduce Continual GUI Agents, a new task that requires GUI agents to perform continual learning under shifted domains and resolutions. We find existing methods fail to maintain stable grounding as GUI distributions shift over time, due to the diversity of UI interaction points and regions in fluxing scenarios. To address this, we introduce GUI‑Anchoring in Flux (GUI‑AiF), a new reinforcement fine‑tuning framework that stabilizes continual learning through two novel rewards: Anchoring Point Reward in Flux (APR‑iF) and Anchoring Region Reward in Flux (ARR‑iF). These rewards guide the agents to align with shifting interaction points and regions, mitigating the tendency of existing reward strategies to over‑adapt to static grounding cues (e.g., fixed coordinates or element scales). Extensive experiments show GUI‑AiF surpasses state‑of‑the‑art baselines. Our work establishes the first continual learning framework for GUI agents, revealing the untapped potential of reinforcement fine‑tuning for continual GUI Agents.
Authors:Pankhi Kashyap, Mainak Singha, Biplab Banerjee
Abstract:
Prompt learning (PL) has emerged as an effective strategy to adapt vision‑language models (VLMs), such as CLIP, for downstream tasks under limited supervision. While PL has demonstrated strong generalization on natural image datasets, its transferability to remote sensing (RS) imagery remains underexplored. RS data present unique challenges, including multi‑label scenes, high intra‑class variability, and diverse spatial resolutions, that hinder the direct applicability of existing PL methods. In particular, current prompt‑based approaches often struggle to identify dominant semantic cues and fail to generalize to novel classes in RS scenarios. To address these challenges, we propose BiMoRS, a lightweight bi‑modal prompt learning framework tailored for RS tasks. BiMoRS employs a frozen image captioning model (e.g., BLIP‑2) to extract textual semantic summaries from RS images. These captions are tokenized using a BERT tokenizer and fused with high‑level visual features from the CLIP encoder. A lightweight cross‑attention module then conditions a learnable query prompt on the fused textual‑visual representation, yielding contextualized prompts without altering the CLIP backbone. We evaluate BiMoRS on four RS datasets across three domain generalization (DG) tasks and observe consistent performance gains, outperforming strong baselines by up to 2% on average. Codes are available at https://github.com/ipankhi/BiMoRS.
Authors:Michele Mazzamuto, Daniele Di Mauro, Gianpiero Francesca, Giovanni Maria Farinella, Antonino Furnari
Abstract:
Skill assessment in procedural videos is crucial for the objective evaluation of human performance in settings such as manufacturing and procedural daily tasks. Current research on skill assessment has predominantly focused on sports and lacks large‑scale datasets for complex procedural activities. Existing studies typically involve only a limited number of actions, focus on either pairwise assessments (e.g., A is better than B) or on binary labels (e.g., good execution vs needs improvement). In response to these shortcomings, we introduce ProSkill, the first benchmark dataset for action‑level skill assessment in procedural tasks. ProSkill provides absolute skill assessment annotations, along with pairwise ones. This is enabled by a novel and scalable annotation protocol that allows for the creation of an absolute skill assessment ranking starting from pairwise assessments. This protocol leverages a Swiss Tournament scheme for efficient pairwise comparisons, which are then aggregated into consistent, continuous global scores using an ELO‑based rating system. We use our dataset to benchmark the main state‑of‑the‑art skill assessment algorithms, including both ranking‑based and pairwise paradigms. The suboptimal results achieved by the current state‑of‑the‑art highlight the challenges and thus the value of ProSkill in the context of skill assessment for procedural videos. All data and code are available at https://fpv‑iplab.github.io/ProSkill/
Authors:Jing Wu, Daphne Barretto, Yiye Chen, Nicholas Gydé, Yanan Jian, Yuhang He, Vibhav Vineet
Abstract:
Long‑horizon, repetitive workflows are common in professional settings, such as processing expense reports from receipts and entering student grades from exam papers. These tasks are often tedious for humans since they can extend to extreme lengths proportional to the size of the data to process. However, they are ideal for Computer‑Use Agents (CUAs) due to their structured, recurring sub‑workflows with logic that can be systematically learned. Identifying the absence of an evaluation benchmark as a primary bottleneck, we establish OS‑Marathon, comprising 242 long‑horizon, repetitive tasks across 2 domains to evaluate state‑of‑the‑art (SOTA) agents. We then introduce a cost‑effective method to construct a condensed demonstration using only few‑shot examples to teach agents the underlying workflow logic, enabling them to execute similar workflows effectively on larger, unseen data collections. Extensive experiments demonstrate both the inherent challenges of these tasks and the effectiveness of our proposed method. Project website: https://os‑marathon.github.io/.
Authors:Rohan Asthana, Vasileios Belagiannis
Abstract:
Diffusion‑based image generative models produce high‑fidelity images through iterative denoising but remain vulnerable to memorization, where they unintentionally reproduce exact copies or parts of training images. Recent memorization detection methods are primarily based on the norm of score difference as indicators of memorization. We prove that such norm‑based metrics are mainly effective under the assumption of isotropic log‑probability distributions, which generally holds at high or medium noise levels. In contrast, analyzing the anisotropic regime reveals that memorized samples exhibit strong angular alignment between the guidance vector and unconditional scores in the low‑noise setting. Through these insights, we develop a memorization detection metric by integrating isotropic norm and anisotropic alignment. Our detection metric can be computed directly on pure noise inputs via two conditional and unconditional forward passes, eliminating the need for costly denoising steps. Detection experiments on Stable Diffusion v1.4 and v2 show that our metric outperforms existing denoising‑free detection methods while being at least approximately 5x faster than the previous best approach. Finally, we demonstrate the effectiveness of our approach by utilizing a mitigation strategy that adapts memorized prompts based on our developed metric. The code is available at https://github.com/rohanasthana/memorization‑anisotropy .
Authors:Zhuonan Wang, Wenjie Yan, Wenqiao Zhang, Xiaohui Song, Jian Ma, Ke Yao, Yibo Yu, Beng Chin Ooi
Abstract:
Medical image classification is a core task in computer‑aided diagnosis (CAD), playing a pivotal role in early disease detection, treatment planning, and patient prognosis assessment. In ophthalmic practice, fluorescein fundus angiography (FFA) and indocyanine green angiography (ICGA) provide hemodynamic and lesion‑structural information that conventional fundus photography cannot capture. However, due to the single‑modality nature, subtle lesion patterns, and significant inter‑device variability, existing methods still face limitations in generalization and high‑confidence prediction. To address these challenges, we propose CLEAR‑Mamba, an enhanced framework built upon MedMamba with optimizations in both architecture and training strategy. Architecturally, we introduce HaC, a hypernetwork‑based adaptive conditioning layer that dynamically generates parameters according to input feature distributions, thereby improving cross‑domain adaptability. From a training perspective, we develop RaP, a reliability‑aware prediction scheme built upon evidential uncertainty learning, which encourages the model to emphasize low‑confidence samples and improves overall stability and reliability. We further construct a large‑scale ophthalmic angiography dataset covering both FFA and ICGA modalities, comprising multiple retinal disease categories for model training and evaluation. Experimental results demonstrate that CLEAR‑Mamba consistently outperforms multiple baseline models, including the original MedMamba, across various metrics‑showing particular advantages in multi‑disease classification and reliability‑aware prediction. This study provides an effective solution that balances generalizability and reliability for modality‑specific medical image classification tasks. Our project can be accessed at https://github.com/ZJU4HealthCare/CLEAR‑Mamba.
Authors:Lakshman Balasubramanian
Abstract:
Person Re‑Identification (ReID) remains a challenging problem in computer vision. This work reviews various training paradigm and evaluates the robustness of state‑of‑the‑art ReID models in cross‑domain applications and examines the role of foundation models in improving generalization through richer, more transferable visual representations. We compare three training paradigms, supervised, self‑supervised, and language‑aligned models. Through the study the aim is to answer the following questions: Can supervised models generalize in cross‑domain scenarios? How does foundation models like SigLIP2 perform for the ReID tasks? What are the weaknesses of current supervised and foundational models for ReID? We have conducted the analysis across 11 models and 9 datasets. Our results show a clear split: supervised models dominate their training domain but crumble on cross‑domain data. Language‑aligned models, however, show surprising robustness cross‑domain for ReID tasks, even though they are not explicitly trained to do so. Code and data available at: https://github.com/moiiai‑tech/object‑reid‑benchmark.
Authors:Jia Fu, Litingyu Wang, He Li, Zihao Luo, Huamin Wang, Chenyuan Bian, Zijun Gao, Chunbin Gu, Xin Weng, Jianghao Wu, Yicheng Wu, Jin Ye, Linhao Li, Yiwen Ye, Yong Xia, Elias Tappeiner, Fei He, Abdul qayyum, Moona Mazher, Steven A Niederer, Junqiang Chen, Chuanyi Huang, Lisheng Wang, Zhaohu Xing, Hongqiu Wang, Lei Zhu, Shichuan Zhang, Shaoting Zhang, Wenjun Liao, Guotai Wang
Abstract:
Accurate delineation of Gross Tumor Volume (GTV), Lymph Node Clinical Target Volume (LN CTV), and Organ‑at‑Risk (OAR) from Computed Tomography (CT) scans is essential for precise radiotherapy planning in Nasopharyngeal Carcinoma (NPC). Building upon SegRap2023, which focused on OAR and GTV segmentation using single‑center paired non‑contrast CT (ncCT) and contrast‑enhanced CT (ceCT) scans, the SegRap2025 challenge aims to enhance the generalizability and robustness of segmentation models across imaging centers and modalities. SegRap2025 comprises two tasks: Task01 addresses GTV segmentation using paired CT from the SegRap2023 dataset, with an additional external testing set to evaluate cross‑center generalization, and Task02 focuses on LN CTV segmentation using multi‑center training data and an unseen external testing set, where each case contains paired CT scans or a single modality, emphasizing both cross‑center and cross‑modality robustness. This paper presents the challenge setup and provides a comprehensive analysis of the solutions submitted by ten participating teams. For GTV segmentation task, the top‑performing models achieved average Dice Similarity Coefficient (DSC) of 74.61% and 56.79% on the internal and external testing cohorts, respectively. For LN CTV segmentation task, the highest average DSC values reached 60.24%, 60.50%, and 57.23% on paired CT, ceCT‑only, and ncCT‑only subsets, respectively. SegRap2025 establishes a large‑scale multi‑center, multi‑modality benchmark for evaluating the generalization and robustness in radiotherapy target segmentation, providing valuable insights toward clinically applicable automated radiotherapy planning systems. The benchmark is available at: https://hilab‑git.github.io/SegRap2025_Challenge.
Authors:Haoran Wei, Yaofeng Sun, Yukun Li
Abstract:
We present DeepSeek‑OCR 2 to investigate the feasibility of a novel encoder‑DeepEncoder V2‑capable of dynamically reordering visual tokens upon image semantics. Conventional vision‑language models (VLMs) invariably process visual tokens in a rigid raster‑scan order (top‑left to bottom‑right) with fixed positional encoding when fed into LLMs. However, this contradicts human visual perception, which follows flexible yet semantically coherent scanning patterns driven by inherent logical structures. Particularly for images with complex layouts, human vision exhibits causally‑informed sequential processing. Inspired by this cognitive mechanism, DeepEncoder V2 is designed to endow the encoder with causal reasoning capabilities, enabling it to intelligently reorder visual tokens prior to LLM‑based content interpretation. This work explores a novel paradigm: whether 2D image understanding can be effectively achieved through two‑cascaded 1D causal reasoning structures, thereby offering a new architectural approach with the potential to achieve genuine 2D reasoning. Codes and model weights are publicly accessible at http://github.com/deepseek‑ai/DeepSeek‑OCR‑2.
Authors:Robbyant Team, Zelin Gao, Qiuyu Wang, Yanhong Zeng, Jiapeng Zhu, Ka Leong Cheng, Yixuan Li, Hanlin Wang, Yinghao Xu, Shuailei Ma, Yihang Chen, Jie Liu, Yansong Cheng, Yao Yao, Jiayi Zhu, Yihao Meng, Kecheng Zheng, Qingyan Bai, Jingye Chen, Zehong Shen, Yue Yu, Xing Zhu, Yujun Shen, Hao Ouyang
Abstract:
We present LingBot‑World, an open‑sourced world simulator stemming from video generation. Positioned as a top‑tier world model, LingBot‑World offers the following features. (1) It maintains high fidelity and robust dynamics in a broad spectrum of environments, including realism, scientific contexts, cartoon styles, and beyond. (2) It enables a minute‑level horizon while preserving contextual consistency over time, which is also known as "long‑term memory". (3) It supports real‑time interactivity, achieving a latency of under 1 second when producing 16 frames per second. We provide public access to the code and model in an effort to narrow the divide between open‑source and closed‑source technologies. We believe our release will empower the community with practical applications across areas like content creation, gaming, and robot learning.
Authors:Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj
Abstract:
Zero‑shot anomaly detection aims to detect and localise abnormal regions in the image without access to any in‑domain training images. While recent approaches leverage vision‑language models (VLMs), such as CLIP, to transfer high‑level concept knowledge, methods based on purely vision foundation models (VFMs), like DINOv2, have lagged behind in performance. We argue that this gap stems from two practical issues: (i) limited diversity in existing auxiliary anomaly detection datasets and (ii) overly shallow VFM adaptation strategies. To address both challenges, we propose AnomalyVFM, a general and effective framework that turns any pretrained VFM into a strong zero‑shot anomaly detector. Our approach combines a robust three‑stage synthetic dataset generation scheme with a parameter‑efficient adaptation mechanism, utilising low‑rank feature adapters and a confidence‑weighted pixel loss. Together, these components enable modern VFMs to substantially outperform current state‑of‑the‑art methods. More specifically, with RADIO as a backbone, AnomalyVFM achieves an average image‑level AUROC of 94.1% across 9 diverse datasets, surpassing previous methods by significant 3.3 percentage points. Project Page: https://maticfuc.github.io/anomaly_vfm/
Authors:Qiyan Zhao, Xiaofeng Zhang, Shuochen Chang, Qianyu Chen, Xiaosong Yuan, Xuhang Chen, Luoqi Liu, Jiajun Zhang, Xu-Yao Zhang, Da-Han Wang
Abstract:
Recent diffusion‑based Multimodal Large Language Models (dMLLMs) suffer from high inference latency and therefore rely on caching techniques to accelerate decoding. However, the application of cache mechanisms often introduces undesirable repetitive text generation, a phenomenon we term the Repeat Curse. To better investigate underlying mechanism behind this issue, we analyze repetition generation through the lens of information flow. Our work reveals three key findings: (1) context tokens aggregate semantic information as anchors and guide the final predictions; (2) as information propagates across layers, the entropy of context tokens converges in deeper layers, reflecting the model's growing prediction certainty; (3) Repetition is typically linked to disruptions in the information flow of context tokens and to the inability of their entropy to converge in deeper layers. Based on these insights, we present CoTA, a plug‑and‑play method for mitigating repetition. CoTA enhances the attention of context tokens to preserve intrinsic information flow patterns, while introducing a penalty term to the confidence score during decoding to avoid outputs driven by uncertain context tokens. With extensive experiments, CoTA demonstrates significant effectiveness in alleviating repetition and achieves consistent performance improvements on general tasks. Code is available at https://github.com/ErikZ719/CoTA
Authors:Hang Guo, Zhaoyang Jia, Jiahao Li, Bin Li, Yuanhao Cai, Jiangshan Wang, Yawei Li, Yan Lu
Abstract:
The autoregressive video diffusion model has recently gained considerable research interest due to its causal modeling and iterative denoising. In this work, we identify that the multi‑head self‑attention in these models under‑utilizes historical frames: approximately 25% heads attend almost exclusively to the current frame, and discarding their KV caches incurs only minor performance degradation. Building upon this, we propose Dummy Forcing, a simple yet effective method to control context accessibility across different heads. Specifically, the proposed heterogeneous memory allocation reduces head‑wise context redundancy, accompanied by dynamic head programming to adaptively classify head types. Moreover, we develop a context packing technique to achieve more aggressive cache compression. Without additional training, our Dummy Forcing delivers up to 2.0x speedup over the baseline, supporting video generation at 24.3 FPS with less than 0.5% quality drop. Project page is available at https://csguoh.github.io/project/DummyForcing/.
Authors:Zengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong, Man Zhang, Xiangxiang Chu
Abstract:
Text‑to‑image (T2I) models have achieved remarkable success in generating high‑fidelity images, but they often fail in handling complex spatial relationships, e.g., spatial perception, reasoning, or interaction. These critical aspects are largely overlooked by current benchmarks due to their short or information‑sparse prompt design. In this paper, we introduce SpatialGenEval, a new benchmark designed to systematically evaluate the spatial intelligence of T2I models, covering two key aspects: (1) SpatialGenEval involves 1,230 long, information‑dense prompts across 25 real‑world scenes. Each prompt integrates 10 spatial sub‑domains and corresponding 10 multi‑choice question‑answer pairs, ranging from object position and layout to occlusion and causality. Our extensive evaluation of 21 state‑of‑the‑art models reveals that higher‑order spatial reasoning remains a primary bottleneck. (2) To demonstrate that the utility of our information‑dense design goes beyond simple evaluation, we also construct the SpatialT2I dataset. It contains 15,400 text‑image pairs with rewritten prompts to ensure image consistency while preserving information density. Fine‑tuned results on current foundation models (i.e., Stable Diffusion‑XL, Uniworld‑V1, OmniGen2) yield consistent performance gains (+4.2%, +5.7%, +4.4%) and more realistic effects in spatial relations, highlighting a data‑centric paradigm to achieve spatial intelligence in T2I models.
Authors:Mai Su, Qihan Yu, Zhongtao Wang, Yilong Li, Chengwei Pan, Yisong Chen, Guoping Wang, Fei Zhu
Abstract:
3D Gaussian Splatting (3DGS) enables efficient rendering, yet accurate surface reconstruction remains challenging due to unreliable geometric supervision. Existing approaches predominantly rely on depth‑based reprojection to infer visibility and enforce multi‑view consistency, leading to a fundamental circular dependency: visibility estimation requires accurate depth, while depth supervision itself is conditioned on visibility. In this work, we revisit multi‑view geometric supervision from the perspective of visibility modeling. Instead of inferring visibility from pixel‑wise depth consistency, we explicitly model visibility at the level of Gaussian primitives. We introduce a Gaussian visibility‑aware multi‑view geometric consistency (GVMV) formulation, which aggregates cross‑view visibility of shared Gaussians to construct reliable supervision over co‑visible regions. To further incorporate monocular priors, we propose a progressive quadtree‑calibrated depth alignment (QDC) strategy that performs block‑wise affine calibration under visibility‑aware guidance, effectively mitigating scale ambiguity while preserving local geometric structures. Extensive experiments on DTU and Tanks and Temples demonstrate that our method consistently improves reconstruction accuracy over prior Gaussian‑based approaches. Our code is fully open‑sourced and available at an anonymous repository: https://github.com/GVGScode/GVGS.
Authors:Jiyuan Xu, Wenyu Zhang, Xin Jing, Shuai Chen, Shuai Zhang, Jiahao Nie
Abstract:
Current methods for multivariate time series forecasting can be classified into channel‑dependent and channel‑independent models. Channel‑dependent models learn cross‑channel features but often overfit the channel ordering, which hampers adaptation when channels are added or reordered. Channel‑independent models treat each channel in isolation to increase flexibility, yet this neglects inter‑channel dependencies and limits performance. To address these limitations, we propose CPiRi, a channel permutation invariant (CPI) framework that infers cross‑channel structure from data rather than memorizing a fixed ordering, enabling deployment in settings with structural and distributional co‑drift without retraining. CPiRi couples spatio‑temporal decoupling architecture with permutation‑invariant regularization training strategy: a frozen pretrained temporal encoder extracts high‑quality temporal features, a lightweight spatial module learns content‑driven inter‑channel relations, while a channel shuffling strategy enforces CPI during training. We further ground CPiRi in theory by analyzing permutation equivariance in multivariate time series forecasting. Experiments on multiple benchmarks show state‑of‑the‑art results. CPiRi remains stable when channel orders are shuffled and exhibits strong inductive generalization to unseen channels even when trained on only half of the channels, while maintaining practical efficiency on large‑scale datasets. The source code is released at https://github.com/JasonStraka/CPiRi.
Authors:Shuoyan Wei, Feng Li, Chen Zhou, Runmin Cong, Yao Zhao, Huihui Bai
Abstract:
Diffusion models have demonstrated exceptional success in video super‑resolution (VSR), exhibiting powerful capabilities for generating fine‑grained details. However, their potential for space‑time video super‑resolution (STVSR), which necessitates not only recovering realistic high‑resolution visual content but also improving the frame rate with coherent temporal dynamics, remains largely underexplored. Moreover, existing STVSR methods predominantly address spatiotemporal upsampling under simple degradation assumptions, thus failing in real‑world scenarios with complex unknown degradations. To address these challenges, we propose OSDEnhancer, the first framework that achieves robust STVSR in one‑step diffusion. OSDEnhancer begins with a linear initialization to establish essential spatiotemporal structures and adapt the model for one‑step reconstruction. It then applies a divide‑and‑conquer strategy, introducing the temporal coherence (TC) and texture enrichment (TE) LoRAs that progressively specialize in inter‑frame dynamics modeling and fine‑grained texture recovery, respectively, while collaborating during inference for enhanced overall performance. A bidirectional VAE decoder employs deformable recurrent blocks to leverage the multi‑scale structure of the vanilla VAE, enhancing latent‑to‑pixel reconstruction through joint multi‑scale deformable aggregation and inter‑frame feature propagation. Experimental results demonstrate that the proposed method attains state‑of‑the‑art performance with superior generalization in real‑world scenarios. The code is available at https://github.com/W‑Shuoyan/OSDEnhancer.
Authors:Yanjie Tu, Qingsen Yan, Axi Niu, Jiacong Tang
Abstract:
All‑in‑one image restoration aims to address diverse degradation types using a single unified model. Existing methods typically rely on degradation priors to guide restoration, yet often struggle to reconstruct content in severely degraded regions. Although recent works leverage semantic information to facilitate content generation, integrating it into the shallow layers of diffusion models often disrupts spatial structures (\emphe.g., blurring artifacts). To address this issue, we propose a Triple‑Prior Guided Diffusion (TPGDiff) network for unified image restoration. TPGDiff incorporates degradation priors throughout the diffusion trajectory, while introducing structural priors into shallow layers and semantic priors into deep layers, enabling hierarchical and complementary prior guidance for image reconstruction. Specifically, we leverage multi‑source structural cues as structural priors to capture fine‑grained details and guide shallow layers representations. To complement this design, we further develop a distillation‑driven semantic extractor that yields robust semantic priors, ensuring reliable high‑level guidance at deep layers even under severe degradations. Furthermore, a degradation extractor is employed to learn degradation‑aware priors, enabling stage‑adaptive control of the diffusion process across all timesteps. Extensive experiments on both single‑ and multi‑degradation benchmarks demonstrate that TPGDiff achieves superior performance and generalization across diverse restoration scenarios. Our project page is: https://leoyjtu.github.io/tpgdiff‑project.
Authors:Xiaofeng Zhang, Yuanchao Zhu, Chaochen Gu, Xiaosong Yuan, Qiyan Zhao, Jiawei Cao, Feilong Tang, Sinan Fan, Yaomin Shen, Chen Shen, Hao Tang
Abstract:
Recent studies have examined attention dynamics in large vision‑language models (LVLMs) to detect hallucinations. However, existing approaches remain limited in reliably distinguishing hallucinated from factually grounded outputs, as they rely solely on forward‑pass attention patterns and neglect gradient‑based signals that reveal how token influence propagates through the network. To bridge this gap, we introduce LVLMs‑Saliency, a gradient‑aware diagnostic framework that quantifies the visual grounding strength of each output token by fusing attention weights with their input gradients. Our analysis uncovers a decisive pattern: hallucinations frequently arise when preceding output tokens exhibit low saliency toward the prediction of the next token, signaling a breakdown in contextual memory retention. Leveraging this insight, we propose a dual‑mechanism inference‑time framework to mitigate hallucinations: (1) Saliency‑Guided Rejection Sampling (SGRS), which dynamically filters candidate tokens during autoregressive decoding by rejecting those whose saliency falls below a context‑adaptive threshold, thereby preventing coherence‑breaking tokens from entering the output sequence; and (2) Local Coherence Reinforcement (LocoRE), a lightweight, plug‑and‑play module that strengthens attention from the current token to its most recent predecessors, actively counteracting the contextual forgetting behavior identified by LVLMs‑Saliency. Extensive experiments across multiple LVLMs demonstrate that our method significantly reduces hallucination rates while preserving fluency and task performance, offering a robust and interpretable solution for enhancing model reliability. Code is available at: https://github.com/zhangbaijin/LVLMs‑Saliency
Authors:Wanjun Jia, Kang Li, Fan Yang, Mengfei Duan, Wenrui Chen, Yiming Jiang, Hui Zhang, Kailun Yang, Zhiyong Li, Yaonan Wang
Abstract:
The central challenge in robotic manipulation of deformable objects lies in aligning high‑level semantic instructions with physical interaction points under complex appearance and texture variations. Due to near‑infinite degrees of freedom, complex dynamics, and heterogeneous patterns, existing vision‑based affordance prediction methods often suffer from boundary overflow and fragmented functional regions. To address these issues, we propose TRACER, a Texture‑Robust Affordance Chain‑of‑thought with dEformable‑object Refinement framework, which establishes a cross‑hierarchical mapping from hierarchical semantic reasoning to appearance‑robust and physically consistent functional region refinement. Specifically, a Tree‑structured Affordance Chain‑of‑Thought (TA‑CoT) is formulated to decompose high‑level task intentions into hierarchical sub‑task semantics, providing consistent guidance across various execution stages. To ensure spatial integrity, a Spatial‑Constrained Boundary Refinement (SCBR) mechanism is introduced to suppress prediction spillover, guiding the perceptual response to converge toward authentic interaction manifolds. Furthermore, an Interactive Convergence Refinement Flow (ICRF) is developed to aggregate discrete pixels corrupted by appearance noise, significantly enhancing the spatial continuity and physical plausibility of the identified functional regions. Extensive experiments conducted on the Fine‑AGDDO15 dataset and a real‑world robotic platform demonstrate that TRACER significantly improves affordance grounding precision across diverse textures and patterns inherent to deformable objects. More importantly, it enhances the success rate of long‑horizon tasks, effectively bridging the gap between high‑level semantic reasoning and low‑level physical execution. The source code and dataset will be made publicly available at https://github.com/Dikay1/TRACER.
Authors:Shiwen Zhang, Xiaoyan Yang, Bojia Zi, Haibin Huang, Chi Zhang, Xuelong Li
Abstract:
Content‑preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and style features in their internal representations. In this technical report, we present TeleStyle, a lightweight yet effective model for both image and video stylization. Built upon Qwen‑Image‑Edit, TeleStyle leverages the base model's robust capabilities in content preservation and style customization. To facilitate effective training, we curated a high‑quality dataset of distinct specific styles and further synthesized triplets using thousands of diverse, in‑the‑wild style categories. We introduce a Curriculum Continual Learning framework to train TeleStyle on this hybrid dataset of clean (curated) and noisy (synthetic) triplets. This approach enables the model to generalize to unseen styles without compromising precise content fidelity. Additionally, we introduce a video‑to‑video stylization module to enhance temporal consistency and visual quality. TeleStyle achieves state‑of‑the‑art performance across three core evaluation metrics: style similarity, content consistency, and aesthetic quality. Code and pre‑trained models are available at https://github.com/Tele‑AI/TeleStyle
Authors:Atik Faysal, Mohammad Rostami, Reihaneh Gh. Roshan, Nikhil Muralidhar, Huaxia Wang
Abstract:
We address the challenge of training Vision Transformers (ViTs) when labeled data is scarce but unlabeled data is abundant. We propose Semi‑Supervised Masked Autoencoder (SSMAE), a framework that jointly optimizes masked image reconstruction and classification using both unlabeled and labeled samples with dynamically selected pseudo‑labels. SSMAE introduces a validation‑driven gating mechanism that activates pseudo‑labeling only after the model achieves reliable, high‑confidence predictions that are consistent across both weakly and strongly augmented views of the same image, reducing confirmation bias. On CIFAR‑10 and CIFAR‑100, SSMAE consistently outperforms supervised ViT and fine‑tuned MAE, with the largest gains in low‑label regimes (+9.24% over ViT on CIFAR‑10 with 10% labels). Our results demonstrate that when pseudo‑labels are introduced is as important as how they are generated for data‑efficient transformer training. Codes are available at https://github.com/atik666/ssmae.
Authors:Shubham Patle, Sara Ghaboura, Hania Tariq, Mohammad Usman Khan, Omkar Thawakar, Rao Muhammad Anwer, Salman Khan
Abstract:
Arabic calligraphy represents one of the richest visual traditions of the Arabic language, blending linguistic meaning with artistic form. Although multimodal models have advanced across languages, their ability to process Arabic script, especially in artistic and stylized calligraphic forms, remains largely unexplored. To address this gap, we present DuwatBench, a benchmark of 1,272 curated samples containing about 1,475 unique words across six classical and modern calligraphic styles, each paired with sentence‑level detection annotations. The dataset reflects real‑world challenges in Arabic writing, such as complex stroke patterns, dense ligatures, and stylistic variations that often challenge standard text recognition systems. Using DuwatBench, we evaluated 13 leading Arabic and multilingual multimodal models and showed that while they perform well on clean text, they struggle with calligraphic variation, artistic distortions, and precise visual‑text alignment. By publicly releasing DuwatBench and its annotations, we aim to advance culturally grounded multimodal research, foster fair inclusion of the Arabic language and visual heritage in AI systems, and support continued progress in this area. Our dataset (https://huggingface.co/datasets/MBZUAI/DuwatBench) and evaluation suit (https://github.com/mbzuai‑oryx/DuwatBench) are publicly available.
Authors:Binzhu Xie, Shi Qiu, Sicheng Zhang, Yinqiao Wang, Hao Xu, Muzammal Naseer, Chi-Wing Fu, Pheng-Ann Heng
Abstract:
Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self‑occlusion, and complex hand‑object interactions. Prior methods mitigate these issues by scaling training data or adding auxiliary cues, but they often struggle in unseen contexts. We present EgoHandICL, the first in‑context learning (ICL) framework for 3D hand reconstruction that improves semantic alignment, visual consistency, and robustness under challenging egocentric conditions. EgoHandICL introduces complementary exemplar retrieval guided by vision‑language models (VLMs), an ICL‑tailored tokenizer for multimodal context, and a masked autoencoder (MAE)‑based architecture trained with hand‑guided geometric and perceptual objectives. Experiments on ARCTIC and EgoExo4D show consistent gains over state‑of‑the‑art methods. We also demonstrate real‑world generalization and improve EgoVLM hand‑object interaction reasoning by using reconstructed hands as visual prompts. Code and data: https://github.com/Nicous20/EgoHandICL
Authors:Xinrui Zhang, Yufeng Wang, Shuangkang Fang, Zesheng Wang, Dacheng Qi, Wenrui Ding
Abstract:
Underwater 3D reconstruction and appearance restoration are hindered by the complex optical properties of water, such as wavelength‑dependent attenuation and scattering. Existing Neural Radiance Fields (NeRF)‑based methods struggle with slow rendering speeds and suboptimal color restoration, while 3D Gaussian Splatting (3DGS) inherently lacks the capability to model complex volumetric scattering effects. To address these issues, we introduce WaterClear‑GS, the first pure 3DGS‑based framework that explicitly integrates underwater optical properties of local attenuation and scattering into Gaussian primitives, eliminating the need for an auxiliary medium network. Our method employs a dual‑branch optimization strategy to ensure underwater photometric consistency while naturally recovering water‑free appearances. This strategy is enhanced by depth‑guided geometry regularization and perception‑driven image loss, together with exposure constraints, spatially‑adaptive regularization, and physically guided spectral regularization, which collectively enforce local 3D coherence and maintain natural visual perception. Experiments on standard benchmarks and our newly collected dataset demonstrate that WaterClear‑GS achieves outstanding performance on both novel view synthesis (NVS) and underwater image restoration (UIR) tasks, while maintaining real‑time rendering. The code will be available at https://buaaxrzhang.github.io/WaterClear‑GS/.
Authors:Renrong Shao, Dongyang Li, Dong Xia, Lin Shao, Jiangdong Lu, Fen Zheng, Lulu Zhang
Abstract:
Vision Mamba models have been extensively researched in various fields, which address the limitations of previous models by effectively managing long‑range dependencies with a linear‑time overhead. Several prospective studies have further designed Vision Mamba based on UNet(VM‑UNet) for medical image segmentation. These approaches primarily focus on optimizing architectural designs by creating more complex structures to enhance the model's ability to perceive semantic features. In this paper, we propose a simple yet effective approach to improve the model by Dual Self‑distillation for VM‑UNet (DSVM‑UNet) without any complex architectural designs. To achieve this goal, we develop double self‑distillation methods to align the features at both the global and local levels. Extensive experiments conducted on the ISIC2017, ISIC2018, and Synapse benchmarks demonstrate that our approach achieves state‑of‑the‑art performance while maintaining computational efficiency. Code is available at https://github.com/RoryShao/DSVM‑UNet.git.
Authors:Ziyue Wang, Sheng Jin, Zhongrong Zuo, Jiawei Wu, Han Qiu, Qi She, Hao Zhang, Xudong Jiang
Abstract:
Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models, yet existing video reasoning methods often rely on coarse sequence‑level rewards or single‑factor token selection, neglecting fine‑grained links among visual inputs, temporal dynamics, and linguistic outputs, limiting both accuracy and interpretability. We propose Video‑KTR, a modality‑aware policy shaping framework that performs selective, token‑level RL by combining three attribution signals: (1) visual‑aware tokens identified via counterfactual masking to reveal perceptual dependence; (2) temporal‑aware tokens detected through frame shuffling to expose temporal sensitivity; and (3) high‑entropy tokens signaling predictive uncertainty. By reinforcing only these key tokens, Video‑KTR focuses learning on semantically informative, modality‑sensitive content while filtering out low‑value tokens. Across five challenging benchmarks, Video‑KTR achieves state‑of‑the‑art or highly competitive results, achieving 42.7% on Video‑Holmes (surpassing GPT‑4o) with consistent gains on both reasoning and general video understanding tasks. Ablation studies verify the complementary roles of the attribution signals and the robustness of targeted token‑level updates. Overall, Video‑KTR improves accuracy and interpretability, offering a simple, drop‑in extension to RL for complex video reasoning. Our code and models are available at https://github.com/zywang0104/Video‑KTR.
Authors:Mao-Lin Luo, Zi-Hao Zhou, Yi-Lin Zhang, Yuanyu Wan, Tong Wei, Min-Ling Zhang
Abstract:
Continual learning for pre‑trained vision‑language models requires balancing three competing objectives: retaining pre‑trained knowledge, preserving knowledge from a sequence of learned tasks, and maintaining the plasticity to acquire new knowledge. This paper presents a simple but effective approach called KeepLoRA to effectively balance these objectives. We first analyze the knowledge retention mechanism within the model parameter space and find that general knowledge is mainly encoded in the principal subspace, while task‑specific knowledge is encoded in the residual subspace. Motivated by this finding, KeepLoRA learns new tasks by restricting LoRA parameter updates in the residual subspace to prevent interfering with previously learned capabilities. Specifically, we infuse knowledge for a new task by projecting its gradient onto a subspace orthogonal to both the principal subspace of pre‑trained model and the dominant directions of previous task features. Our theoretical and empirical analyses confirm that KeepLoRA balances the three objectives and achieves state‑of‑the‑art performance. The implementation code is available at https://github.com/MaolinLuo/KeepLoRA.
Authors:Yujin Wang, Yutong Zheng, Wenxian Fan, Tianyi Wang, Hongqing Chu, Li Zhang, Bingzhao Gao, Daxin Tian, Jianqiang Wang, Hong Chen
Abstract:
In this paper, we introduce ScenePilot‑4K, a large‑scale first‑person dataset for safety‑aware vision‑language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot‑4K contains 3,847 hours of video and 27.7M front‑view frames spanning 63 countries/regions and 1,210 cities. It jointly provides scene‑level natural‑language descriptions, risk assessment labels, key‑participant annotations, ego trajectories, and camera parameters through a unified multi‑stage annotation pipeline. Building on this dataset, we establish ScenePilot‑Bench, a standardized benchmark that evaluates vision‑language models along four complementary axes: scene understanding, spatial perception, motion planning, and GPT‑based semantic alignment. The benchmark includes fine‑grained metrics and geographic generalization settings that expose model robustness under cross‑region and cross‑traffic domain shifts. Baseline results on representative open‑source and proprietary vision‑language models show that current models remain competitive in high‑level scene semantics but still exhibit substantial limitations in geometry‑aware perception and planning‑oriented reasoning. Beyond the released dataset itself, the proposed annotation pipeline serves as a reusable and extensible recipe for scalable dataset construction from public Internet driving videos. The codes and supplementary materials are available at: https://github.com/yjwangtj/ScenePilot‑4K, with the dataset available at https://huggingface.co/datasets/larswangtj/ScenePilot‑4K.
Authors:Cuong Le, Pavlo Melnyk, Urs Waldmann, Mårten Wadenbäck, Bastian Wandt
Abstract:
Vision‑based 3D human motion capture from videos remains a challenge in computer vision. Traditional 3D pose estimation approaches often ignore the temporal consistency between frames, causing implausible and jittery motion. The emerging field of kinematics‑based 3D motion capture addresses these issues by estimating the temporal transitioning between poses instead. A major drawback in current kinematics approaches is their reliance on Euler angles. Despite their simplicity, Euler angles suffer from discontinuity that leads to unstable motion reconstructions, especially in online settings where trajectory refinement is unavailable. Contrarily, quaternions have no discontinuity and can produce continuous transitions between poses. In this paper, we propose QuaMo, a novel Quaternion Motions method using quaternion differential equations (QDE) for human kinematics capture. We utilize the state‑space model, an effective system for describing real‑time kinematics estimations, with quaternion state and the QDE describing quaternion velocity. The corresponding angular acceleration is computed from a meta‑PD controller with a novel acceleration enhancement that adaptively regulates the control signals as the human quickly changes to a new pose. Unlike previous work, our QDE is solved under the quaternion unit‑sphere constraint that results in more accurate estimations. Experimental results show that our novel formulation of the QDE with acceleration enhancement accurately estimates 3D human kinematics with no discontinuity and minimal implausibilities. QuaMo outperforms comparable state‑of‑the‑art methods on multiple datasets, namely Human3.6M, Fit3D, SportsPose and AIST. The code is available at https://github.com/cuongle1206/QuaMo
Authors:Ziyu Zhang, Tianle Liu, Diantao Tu, Shuhan Shen
Abstract:
We present a fast 3DGS reconstruction pipeline designed to converge within one minute, developed for the SIGGRAPH Asia 3DGS Fast Reconstruction Challenge. The challenge consists of an initial round using SLAM‑generated camera poses (with noisy trajectories) and a final round using COLMAP poses (highly accurate). To robustly handle these heterogeneous settings, we develop a two‑stage solution. In the first round, we use reverse per‑Gaussian parallel optimization and compact forward splatting based on Taming‑GS and Speedy‑splat, load‑balanced tiling, an anchor‑based Neural‑Gaussian representation enabling rapid convergence with fewer learnable parameters, initialization from monocular depth and partially from feed‑forward 3DGS models, and a global pose refinement module for noisy SLAM trajectories. In the final round, the accurate COLMAP poses change the optimization landscape; we disable pose refinement, revert from Neural‑Gaussians back to standard 3DGS to eliminate MLP inference overhead, introduce multi‑view consistency‑guided Gaussian splitting inspired by Fast‑GS, and introduce a depth estimator to supervise the rendered depth. Together, these techniques enable high‑fidelity reconstruction under a strict one‑minute budget. Our method achieved the top performance with a PSNR of 28.43 and ranked first in the competition.
Authors:Jisheng Chu, Wenrui Li, Rui Zhao, Wangmeng Zuo, Shifeng Chen, Xiaopeng Fan
Abstract:
Generating immersive 3D scenes from texts is a core task in computer vision, crucial for applications in virtual reality and game development. Despite the promise of leveraging 2D diffusion priors, existing methods suffer from spatial blindness and rely on predefined trajectories that fail to exploit the inner relationships among salient objects. Consequently, these approaches are unable to comprehend the semantic layout, preventing them from exploring the scene adaptively to infer occluded content. Moreover, current inpainting models operate in 2D image space, struggling to plausibly fill holes caused by camera motion. To address these limitations, we propose RoamScene3D, a novel framework that bridges the gap between semantic guidance and spatial generation. Our method reasons about the semantic relations among objects and produces consistent and photorealistic scenes. Specifically, we employ a vision‑language model (VLM) to construct a scene graph that encodes object relations, guiding the camera to perceive salient object boundaries and plan an adaptive roaming trajectory. Furthermore, to mitigate the limitations of static 2D priors, we introduce a Motion‑Injected Inpainting model that is fine‑tuned on a synthetic panoramic dataset integrating authentic camera trajectories, making it adaptive to camera motion. Extensive experiments demonstrate that with semantic reasoning and geometric constraints, our method significantly outperforms state‑of‑the‑art approaches in producing consistent and photorealistic scenes. Our code is available at https://github.com/JS‑CHU/RoamScene3D.
Authors:Yao Xiao, Weiyan Chen, Jiahao Chen, Zijie Cao, Weijian Deng, Binbin Yang, Ziyi Dong, Xiangyang Ji, Wei Ke, Pengxu Wei, Liang Lin
Abstract:
Current AI‑Generated Image (AIGI) detection approaches predominantly rely on binary classification to distinguish real from synthetic images, often lacking interpretable or convincing evidence to substantiate their decisions. This limitation stems from existing AIGI detection benchmarks, which, despite featuring a broad collection of synthetic images, remain restricted in their coverage of artifact diversity and lack detailed, localized annotations. To bridge this gap, we introduce a fine‑grained benchmark towards eXplainable AI‑Generated image Detection, named X‑AIGD, which provides pixel‑level, categorized annotations of perceptual artifacts, spanning low‑level distortions, high‑level semantics, and cognitive‑level counterfactuals. These comprehensive annotations facilitate fine‑grained interpretability evaluation and deeper insight into model decision‑making processes. Our extensive investigation using X‑AIGD provides several key insights: (1) Existing AIGI detectors demonstrate negligible reliance on perceptual artifacts, even at the most basic distortion level. (2) While AIGI detectors can be trained to identify specific artifacts, they still substantially base their judgment on uninterpretable features. (3) Explicitly aligning model attention with artifact regions can increase the interpretability and generalization of detectors. The data and code are available at: https://github.com/Coxy7/X‑AIGD.
Authors:Chengxiang Guo, Jian Wang, Junhua Fei, Xiao Li, Chunling Chen, Yun Jin
Abstract:
Multimodal MRI is essential for brain tumor segmentation, yet missing modalities in clinical practice cause existing methods to exhibit >40% performance variance across modality combinations, rendering them clinically unreliable. We propose AMGFormer, achieving significantly improved stability through three synergistic modules: (1) QuadIntegrator Bridge (QIB) enabling spatially adaptive fusion maintaining consistent predictions regardless of available modalities, (2) Multi‑Granular Attention Orchestrator (MGAO) focusing on pathological regions to reduce background sensitivity, and (3) Modality Quality‑Aware Enhancement (MQAE) preventing error propagation from corrupted sequences. On BraTS 2018, our method achieves 89.33% WT, 82.70% TC, 67.23% ET Dice scores with <0.5% variance across 15 modality combinations, solving the stability crisis. Single‑modality ET segmentation shows 40‑81% relative improvements over state‑of‑the‑art methods. The method generalizes to BraTS 2020/2021, achieving up to 92.44% WT, 89.91% TC, 84.57% ET. The model demonstrates potential for clinical deployment with 1.2s inference. Code: https://github.com/guochengxiangives/AMGFormer.
Authors:Chaozheng Wen, Jingwen Tong, Zehong Lin, Chenghong Bian, Jun Zhang
Abstract:
The emerging applications of next‑generation wireless networks demand high‑fidelity environmental intelligence. 3D radio maps bridge physical environments and electromagnetic propagation for spectrum planning and environment‑aware sensing. However, most existing methods treat visual and wireless data as independent modalities and fail to leverage shared electromagnetic propagation principles. To bridge this gap, we propose URF‑GS, a unified radio‑optical radiation field framework based on 3D Gaussian splatting and inverse rendering for 3D radio map construction. By fusing cross‑modal observations, our method recovers scene geometry and material properties to predict radio signals under arbitrary transceiver configurations without retraining. Experiments demonstrate up to a 24.7% improvement in spatial spectrum accuracy and a 10x increase in sample efficiency compared with NeRF‑based methods. We further showcase URF‑GS in Wi‑Fi AP deployment and robot path planning tasks. This unified visual‑wireless representation supports holistic radiation field modeling for future wireless communication systems.
Authors:Sen Nie, Jie Zhang, Zhuo Wang, Shiguang Shan, Xilin Chen
Abstract:
Vision‑language models (VLMs) such as CLIP have demonstrated remarkable zero‑shot generalization, yet remain highly vulnerable to adversarial examples (AEs). While test‑time defenses are promising, existing methods fail to provide sufficient robustness against strong attacks and are often hampered by high inference latency and task‑specific applicability. To address these limitations, we start by investigating the intrinsic properties of AEs, which reveals that AEs exhibit severe feature inconsistency under progressive frequency attenuation. We further attribute this to the model's inherent spectral bias. Leveraging this insight, we propose an efficient test‑time defense named Contrastive Spectral Rectification (CSR). CSR optimizes a rectification perturbation to realign the input with the natural manifold under a spectral‑guided contrastive objective, which is applied input‑adaptively. Extensive experiments across 16 classification benchmarks demonstrate that CSR outperforms the SOTA by an average of 18.1% against strong AutoAttack with modest inference overhead. Furthermore, CSR exhibits broad applicability across diverse visual tasks. Code is available at https://github.com/Summu77/CSR.
Authors:Zhixi Cai, Fucai Ke, Kevin Leo, Sukai Huang, Maria Garcia de la Banda, Peter J. Stuckey, Hamid Rezatofighi
Abstract:
Recent vision‑language models have strong perceptual ability but their implicit reasoning is hard to explain and easily generates hallucinations on complex queries. Compositional methods improve interpretability, but most rely on a single agent or hand‑crafted pipeline and cannot decide when to collaborate across complementary agents or compete among overlapping ones. We introduce MATA (Multi‑Agent hierarchical Trainable Automaton), a multi‑agent system presented as a hierarchical finite‑state automaton for visual reasoning whose top‑level transitions are chosen by a trainable hyper agent. Each agent corresponds to a state in the hyper automaton, and runs a small rule‑based sub‑automaton for reliable micro‑control. All agents read and write a shared memory, yielding transparent execution history. To supervise the hyper agent's transition policy, we build transition‑trajectory trees and transform to memory‑to‑next‑state pairs, forming the MATA‑SFT‑90K dataset for supervised finetuning (SFT). The finetuned LLM as the transition policy understands the query and the capacity of agents, and it can efficiently choose the optimal agent to solve the task. Across multiple visual reasoning benchmarks, MATA achieves the state‑of‑the‑art results compared with monolithic and compositional baselines. The code and dataset are available at https://github.com/ControlNet/MATA.
Authors:Iftekhar Ahmed, Shakib Absar, Aftar Ahmad Sami, Shadman Sakib, Debojyoti Biswas, Seraj Al Mahmud Mostafa
Abstract:
Precise segmentation of retinal arteries and veins carries the diagnosis of systemic cardiovascular conditions. However, standard convolutional architectures often yield topologically disjointed segmentations, characterized by gaps and discontinuities that render reliable graph‑based clinical analysis impossible despite high pixel‑level accuracy. To address this, we introduce a topology‑aware framework engineered to maintain vascular connectivity. Our architecture fuses a Topological Feature Fusion Module (TFFM) that maps local feature representations into a latent graph space, deploying Graph Attention Networks to capture global structural dependencies often missed by fixed receptive fields. Furthermore, we drive the learning process with a hybrid objective function, coupling Tversky loss for class imbalance with soft clDice loss to explicitly penalize topological disconnects. Evaluation on the Fundus‑AVSeg dataset reveals state‑of‑the‑art performance, achieving a combined Dice score of 90.97% and a 95% Hausdorff Distance of 3.50 pixels. Notably, our method decreases vessel fragmentation by approximately 38% relative to baselines, yielding topologically coherent vascular trees viable for automated biomarker quantification. We open‑source our code at https://tffm‑module.github.io/.
Authors:Linshan Wu, Jiaxin Zhuang, Hao Chen
Abstract:
Pan‑cancer screening in large‑scale CT scans remains challenging for existing AI methods, primarily due to the difficulty of localizing diverse types of tiny lesions in large CT volumes. The extreme foreground‑background imbalance significantly hinders models from focusing on diseased regions, while redundant focus on healthy regions not only decreases the efficiency but also increases false positives. Inspired by radiologists' glance and focus diagnostic strategy, we introduce GF‑Screen, a Glance and Focus reinforcement learning framework for pan‑cancer screening. GF‑Screen employs a Glance model to localize the diseased regions and a Focus model to precisely segment the lesions, where segmentation results of the Focus model are leveraged to reward the Glance model via Reinforcement Learning (RL). Specifically, the Glance model crops a group of sub‑volumes from the entire CT volume and learns to select the sub‑volumes with lesions for the Focus model to segment. Given that the selecting operation is non‑differentiable for segmentation training, we propose to employ the segmentation results to reward the Glance model. To optimize the Glance model, we introduce a novel group relative learning paradigm, which employs group relative comparison to prioritize high‑advantage predictions and discard low‑advantage predictions within sub‑volume groups, not only improving efficiency but also reducing false positives. In this way, for the first time, we effectively extend cutting‑edge RL techniques to tackle the specific challenges in pan‑cancer screening. Extensive experiments on 16 internal and 7 external datasets across 9 lesion types demonstrated the effectiveness of GF‑Screen. Notably, GF‑Screen leads the public validation leaderboard of MICCAI FLARE25 pan‑cancer challenge, surpassing the FLARE24 champion solution by a large margin (+25.6% DSC and +28.2% NSD).
Authors:Yanxi Wang, Zhiling Zhang, Wenbo Zhou, Weiming Zhang, Jie Zhang, Qiannan Zhu, Yu Shi, Shuxin Zheng, Jiyan He
Abstract:
As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive information such as identities, accounts, locations, and behavioral traces. While existing benchmarks primarily focus on task completion, grounding, or defenses against third‑party attacks, current visual privacy datasets remain largely restricted to static natural images, limiting their ability to capture the contextual dependence and task relevance of privacy risks in GUI task trajectories. To bridge this gap, we introduce GUIGuard‑Bench, a first‑step benchmark for studying privacy‑preserving GUI agents in trajectory‑based GUI workflows. GUIGuard‑Bench contains 241 real GUI‑agent trajectories with 4,080 screenshots across Android and PC environments. Each screenshot is annotated at the region level with privacy bounding boxes, semantic privacy categories, risk levels, and whether the private information is necessary for completing the task. Built on these annotations, GUIGuard‑Bench supports three complementary evaluations: privacy recognition, offline planning fidelity under protected screenshots, and the utility impact of different protection strategies. Our results show that current models can often detect whether a screenshot contains private information, but they struggle with fine‑grained localization, category recognition, risk assessment, and task‑necessity judgment. We also find that closed‑source models, exemplified by Claude Sonnet 4.6, can maintain largely consistent planner semantics in Android environments after privacy protection is applied. Our results highlight privacy recognition as a critical bottleneck for practical GUI agents. Project: https://futuresis.github.io/GUIGuard‑page/
Authors:Li Kang, Heng Zhou, Xiufeng Song, Rui Li, Bruno N. Y. Chen, Ziye Wang, Ximeng Meng, Stone Tao, Yiran Qin, Xiaohong Liu, Ruimao Zhang, Lei Bai, Yilun Du, Hao Su, Philip Torr, Zhenfei Yin, Ruihao Gong, Yejun Zeng, Fengjun Zhong, Shenghao Jin, Jinyang Guo, Xianglong Liu, Xiaojun Jia, Tianqi Shan, Wenqi Ren, Simeng Qin, Jialing Yang, Xiaoyu Ma, Tianxing Chen, Zixuan Li, Zijian Cai, Yan Qin, Yusen Qin, Qiangyu Chen, Kaixuan Wang, Zhaoming Han, Yao Mu, Ping Luo, Yuanqi Yao, Haoming Song, Jan-Nico Zaech, Fabien Despinoy, Danda Pani Paudel, Luc Van Gool
Abstract:
Recent advancements in multimodal large language models and vision‑languageaction models have significantly driven progress in Embodied AI. As the field transitions toward more complex task scenarios, multi‑agent system frameworks are becoming essential for achieving scalable, efficient, and collaborative solutions. This shift is fueled by three primary factors: increasing agent capabilities, enhancing system efficiency through task delegation, and enabling advanced human‑agent interactions. To address the challenges posed by multi‑agent collaboration, we propose the Multi‑Agent Robotic System (MARS) Challenge, held at the NeurIPS 2025 Workshop on SpaVLE. The competition focuses on two critical areas: planning and control, where participants explore multi‑agent embodied planning using vision‑language models (VLMs) to coordinate tasks and policy execution to perform robotic manipulation in dynamic environments. By evaluating solutions submitted by participants, the challenge provides valuable insights into the design and coordination of embodied multi‑agent systems, contributing to the future development of advanced collaborative AI systems.
Authors:Wei Wu, Fan Lu, Yunnan Wang, Shuai Yang, Shi Liu, Fangjing Wang, Qian Zhu, He Sun, Yong Wang, Shuailei Ma, Yiyu Ren, Kejia Zhang, Hui Yu, Jingmei Zhao, Shuai Zhou, Zhenqi Qiu, Houlong Xiong, Ziyu Wang, Zechen Wang, Ran Cheng, Yong-Lu Li, Yongtao Huang, Xing Zhu, Yujun Shen, Kecheng Zheng
Abstract:
Offering great potential in robotic manipulation, a capable Vision‑Language‑Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensuring cost efficiency (e.g., data and GPU hours required for adaptation). To this end, we develop LingBot‑VLA with around 20,000 hours of real‑world data from 9 popular dual‑arm robot configurations. Through a systematic assessment on 3 robotic platforms, each completing 100 tasks with 130 post‑training episodes per task, our model achieves clear superiority over competitors, showcasing its strong performance and broad generalizability. We have also built an efficient codebase, which delivers a throughput of 261 samples per second with an 8‑GPU training setup, representing a 1.5~2.8× (depending on the relied VLM base model) speedup over existing VLA‑oriented codebases. The above features ensure that our model is well‑suited for real‑world deployment. To advance the field of robot learning, we provide open access to the code, base model, and benchmark data, with a focus on enabling more challenging tasks and promoting sound evaluation standards.
Authors:Tong Shi, Melonie de Almeida, Daniela Ivanova, Nicolas Pugeault, Paul Henderson
Abstract:
Talking Head Generation aims at synthesizing natural‑looking talking videos from speech and a single portrait image. Previous 3D talking head generation methods have relied on domain‑specific heuristics such as warping‑based facial motion representation priors to animate talking motions, yet still produce inaccurate 3D avatar reconstructions, thus undermining the realism of generated animations. We introduce Splat‑Portrait, a Gaussian‑splatting‑based method that addresses the challenges of 3D head reconstruction and lip motion synthesis. Our approach automatically learns to disentangle a single portrait image into a static 3D reconstruction represented as static Gaussian Splatting, and a predicted whole‑image 2D background. It then generates natural lip motion conditioned on input audio, without any motion driven priors. Training is driven purely by 2D reconstruction and score‑distillation losses, without 3D supervision nor landmarks. Experimental results demonstrate that Splat‑Portrait exhibits superior performance on talking head generation and novel view synthesis, achieving better visual quality compared to previous works. Our project code and supplementary documents are public available at https://github.com/stonewalking/Splat‑portrait.
Authors:Zequn Xie
Abstract:
Text‑Based Person Search (TBPS) aims to retrieve pedestrian images from large galleries using natural language descriptions. This task, essential for public safety applications, is hindered by cross‑modal discrepancies and ambiguous user queries. We introduce CONQUER, a two‑stage framework designed to address these challenges by enhancing cross‑modal alignment during training and adaptively refining queries at inference. During training, CONQUER employs multi‑granularity encoding, complementary pair mining, and context‑guided optimal matching based on Optimal Transport to learn robust embeddings. At inference, a plug‑and‑play query enhancement module refines vague or incomplete queries via anchor selection and attribute‑driven enrichment, without requiring retraining of the backbone. Extensive experiments on CUHK‑PEDES, ICFG‑PEDES, and RSTPReid demonstrate that CONQUER consistently outperforms strong baselines in both Rank‑1 accuracy and mAP, yielding notable improvements in cross‑domain and incomplete‑query scenarios. These results highlight CONQUER as a practical and effective solution for real‑world TBPS deployment. Source code is available at https://github.com/zqxie77/CONQUER.
Authors:Sangwon Jang, Taekyung Ki, Jaehyeong Jo, Saining Xie, Jaehong Yoon, Sung Ju Hwang
Abstract:
Modern video generators still struggle with complex physical dynamics, often falling short of physical realism. Existing approaches address this using external verifiers or additional training on augmented data, which is computationally expensive and still limited in capturing fine‑grained motion. In this work, we present self‑refining video sampling, a simple method that uses a pre‑trained video generator trained on large‑scale datasets as its own self‑refiner. By interpreting the generator as a denoising autoencoder, we enable iterative inner‑loop refinement at inference time without any external verifier or additional training. We further introduce an uncertainty‑aware refinement strategy that selectively refines regions based on self‑consistency, which prevents artifacts caused by over‑refinement. Experiments on state‑of‑the‑art video generators demonstrate significant improvements in motion coherence and physics alignment, achieving over 70% human preference compared to the default sampler and guidance‑based sampler.
Authors:Roberto Di Via, Vito Paolo Pastore, Francesca Odone, Siôn Glyn-Jones, Irina Voiculescu
Abstract:
Many clinical screening decisions are based on angle measurements. In particular, FemoroAcetabular Impingement (FAI) screening relies on angles traditionally measured on X‑rays. However, assessing the height and span of the impingement area requires also a 3D view through an MRI scan. The two modalities inform the surgeon on different aspects of the condition. In this work, we conduct a matched‑cohort validation study (89 patients, paired MRI/X‑ray) using standard heatmap regression architectures to assess cross‑modality clinical equivalence. Seen that landmark detection has been proven effective on X‑rays, we show that MRI also achieves equivalent localisation and diagnostic accuracy for cam‑type impingement. Our method demonstrates clinical feasibility for FAI assessment in coronal views of 3D MRI volumes, opening the possibility for volumetric analysis through placing further landmarks. These results support integrating automated FAI assessment into routine MRI workflows. Code is released at https://github.com/Malga‑Vision/Landmarks‑Hip‑Conditions
Authors:Kaixun Jiang, Yuzheng Wang, Junjie Zhou, Pandeng Li, Zhihang Liu, Chen-Wei Xie, Zhaoyu Chen, Yun Zheng, Wenqiang Zhang
Abstract:
We introduce GenAgent, unifying visual understanding and generation through an agentic multimodal model. Unlike unified models that face expensive training costs and understanding‑generation trade‑offs, GenAgent decouples these capabilities through an agentic framework: understanding is handled by the multimodal model itself, while generation is achieved by treating image generation models as invokable tools. Crucially, unlike existing modular systems constrained by static pipelines, this design enables autonomous multi‑turn interactions where the agent generates multimodal chains‑of‑thought encompassing reasoning, tool invocation, judgment, and reflection to iteratively refine outputs. We employ a two‑stage training strategy: first, cold‑start with supervised fine‑tuning on high‑quality tool invocation and reflection data to bootstrap agent behaviors; second, end‑to‑end agentic reinforcement learning combining pointwise rewards (final image quality) and pairwise rewards (reflection accuracy), with trajectory resampling for enhanced multi‑turn exploration. GenAgent significantly boosts base generator(FLUX.1‑dev) performance on GenEval++ (+23.6%) and WISE (+14%). Beyond performance gains, our framework demonstrates three key properties: 1) cross‑tool generalization to generators with varying capabilities, 2) test‑time scaling with consistent improvements across interaction rounds, and 3) task‑adaptive reasoning that automatically adjusts to different tasks. Our code will be available at \hrefhttps://github.com/deep‑kaixun/GenAgentthis url.
Authors:Zhengyang Li, Thomas Graave, Björn Möller, Zehang Wu, Matthias Franz, Tim Fingscheidt
Abstract:
In audiovisual automatic speech recognition (AV‑ASR) systems, information fusion of visual features in a pre‑trained ASR has been proven as a promising method to improve noise robustness. In this work, based on the prominent Whisper ASR, first, we propose a simple and effective visual fusion method ‑‑ use of visual features both in encoder and decoder (dual‑use) ‑‑ to learn the audiovisual interactions in the encoder and to weigh modalities in the decoder. Second, we compare visual fusion methods in Whisper models of various sizes. Our proposed dual‑use method shows consistent noise robustness improvement, e.g., a 35% relative improvement (WER: 4.41% vs. 6.83%) based on Whisper small, and a 57% relative improvement (WER: 4.07% vs. 9.53%) based on Whisper medium, compared to typical reference middle fusion in babble noise with a signal‑to‑noise ratio (SNR) of 0dB. Third, we conduct ablation studies examining the impact of various module designs and fusion options. Fine‑tuned on 1929 hours of audiovisual data, our dual‑use method using Whisper medium achieves 4.08% (MUSAN babble noise) and 4.43% (NoiseX babble noise) average WER across various SNRs, thereby establishing a new state‑of‑the‑art in noisy conditions on the LRS3 AV‑ASR benchmark. Our code is at https://github.com/ifnspaml/Dual‑Use‑AVASR
Authors:Isaac Deutsch, Nicolas Moënne-Loccoz, Gavriel State, Zan Gojcic
Abstract:
Multi‑view 3D reconstruction methods remain highly sensitive to photometric inconsistencies arising from camera optical characteristics and variations in image signal processing (ISP). Existing mitigation strategies such as per‑frame latent variables or affine color corrections lack physical grounding and generalize poorly to novel views. We propose the Physically‑Plausible ISP (PPISP) correction module, which disentangles camera‑intrinsic and capture‑dependent effects through physically based and interpretable transformations. A dedicated PPISP controller, trained on the input views, predicts ISP parameters for novel viewpoints, analogous to auto exposure and auto white balance in real cameras. This design enables realistic and fair evaluation on novel views without access to ground‑truth images. PPISP achieves state‑of‑the‑art performance on standard benchmarks, while providing intuitive control and supporting the integration of metadata when available. The source code is available at: https://github.com/nv‑tlabs/ppisp
Authors:Zhixian Zhao, Wenjie Tian, Lei Xie
Abstract:
Multimodal emotion analysis is shifting from static classification to generative reasoning. Beyond simple label prediction, robust affective reasoning must synthesize fine‑grained signals such as facial micro‑expressions and prosodic which shifts to decode the latent causality within complex social contexts. However, current Multimodal Large Language Models (MLLMs) face significant limitations in fine‑grained perception, primarily due to data scarcity and insufficient cross‑modal fusion. As a result, these models often exhibit unimodal dominance which leads to hallucinations in complex multimodal interactions, particularly when visual and acoustic cues are subtle, ambiguous, or even contradictory (e.g., in sarcastic scenery). To address this, we introduce SABER‑LLM, a framework designed for robust multimodal reasoning. First, we construct SABER, a large‑scale emotion reasoning dataset comprising 600K video clips, annotated with a novel six‑dimensional schema that jointly captures audiovisual cues and causal logic. Second, we propose the structured evidence decomposition paradigm, which enforces a "perceive‑then‑reason" separation between evidence extraction and reasoning to alleviate unimodal dominance. The ability to perceive complex scenes is further reinforced by consistency‑aware direct preference optimization, which explicitly encourages alignment among modalities under ambiguous or conflicting perceptual conditions. Experiments on EMER, EmoBench‑M, and SABER‑Test demonstrate that SABER‑LLM significantly outperforms open‑source baselines and achieves robustness competitive with closed‑source models in decoding complex emotional dynamics. The dataset and model are available at https://github.com/zxzhao0/SABER‑LLM.
Authors:Mengfan He, Liangzheng Sun, Chunyu Li, Ziyang Meng
Abstract:
Deep homography estimation has broad applications in computer vision and robotics. Remarkable progresses have been achieved while the existing methods typically treat it as a direct regression or iterative refinement problem and often struggling to capture complex geometric transformations or generalize across different domains. In this work, we propose HomoFM, a new framework that introduces the flow matching technique from generative modeling into the homography estimation task for the first time. Unlike the existing methods, we formulate homography estimation problem as a velocity field learning problem. By modeling a continuous and point‑wise velocity field that transforms noisy distributions into registered coordinates, the proposed network recovers high‑precision transformations through a conditional flow trajectory. Furthermore, to address the challenge of domain shifts issue, e.g., the cases of multimodal matching or varying illumination scenarios, we integrate a gradient reversal layer (GRL) into the feature extraction backbone. This domain adaptation strategy explicitly constrains the encoder to learn domain‑invariant representations, significantly enhancing the network's robustness. Extensive experiments demonstrate the effectiveness of the proposed method, showing that HomoFM outperforms state‑of‑the‑art methods in both estimation accuracy and robustness on standard benchmarks. Code and data resource are available at https://github.com/hmf21/HomoFM.
Authors:Linhan Cao, Wei Sun, Weixia Zhang, Xiangyang Zhu, Kaiwei Zhang, Jun Jia, Dandan Zhu, Guangtao Zhai, Xiongkuo Min
Abstract:
Visual quality assessment (VQA) is increasingly shifting from scalar score prediction toward interpretable quality understanding ‑‑ a paradigm that demands fine‑grained spatiotemporal perception and auxiliary contextual information. Current approaches rely on supervised fine‑tuning or reinforcement learning on curated instruction datasets, which involve labor‑intensive annotation and are prone to dataset‑specific biases. To address these challenges, we propose QualiRAG, a training‑free Retrieval‑Augmented Generation (RAG) framework that systematically leverages the latent perceptual knowledge of large multimodal models (LMMs) for visual quality perception. Unlike conventional RAG that retrieves from static corpora, QualiRAG dynamically generates auxiliary knowledge by decomposing questions into structured requests and constructing four complementary knowledge sources: visual metadata, subject localization, global quality summaries, and local quality descriptions, followed by relevance‑aware retrieval for evidence‑grounded reasoning. Extensive experiments show that QualiRAG achieves substantial improvements over open‑source general‑purpose LMMs and VQA‑finetuned LMMs on visual quality understanding tasks, and delivers competitive performance on visual quality comparison tasks, demonstrating robust quality assessment capabilities without any task‑specific training. The code will be publicly available at https://github.com/clh124/QualiRAG.
Authors:Yifan Li, Shiying Wang, Jianqiang Huang
Abstract:
Vision‑Language Pre‑training (VLP) models like CLIP have significantly advanced Remote Sensing Image‑Text Retrieval (RSITR). However, existing methods predominantly rely on coarse‑grained global alignment, which often overlooks the dense, multi‑scale semantics inherent in overhead imagery. Moreover, adapting these heavy models via full fine‑tuning incurs prohibitive computational costs and risks catastrophic forgetting. To address these challenges, we propose MPS‑CLIP, a parameter‑efficient framework designed to shift the retrieval paradigm from global matching to keyword‑guided fine‑grained alignment. Specifically, we leverage a Large Language Model (LLM) to extract core semantic keywords, guiding the Segment Anything Model (SamGeo) to generate semantically relevant sub‑perspectives. To efficiently adapt the frozen backbone, we introduce a Gated Global Attention (G^2A) adapter, which captures global context and long‑range dependencies with minimal overhead. Furthermore, a Multi‑Perspective Representation (MPR) module aggregates these local cues into robust multi‑perspective embeddings. The framework is optimized via a hybrid objective combining multi‑perspective contrastive and weighted triplet losses, which dynamically selects maximum‑response perspectives to suppress noise and enforce precise semantic matching. Extensive experiments on the RSICD and RSITMD benchmarks demonstrate that MPS‑CLIP achieves state‑of‑the‑art performance with 35.18% and 48.40% mean Recall (mR), respectively, significantly outperforming full fine‑tuning baselines and recent competitive methods. Code is available at https://github.com/Lcrucial1f/MPS‑CLIP.
Authors:Zehua Liu, Shihao Zou, Jincai Huang, Yanfang Zhang, Chao Tong, Weixin Si
Abstract:
Transarterial chemoembolization (TACE) is a preferred treatment option for hepatocellular carcinoma and other liver malignancies, yet it remains a highly challenging procedure due to complex intra‑operative vascular navigation and anatomical variability. Accurate and robust 2D‑3D vessel registration is essential to guide microcatheter and instruments during TACE, enabling precise localization of vascular structures and optimal therapeutic targeting. To tackle this issue, we develop a coarse‑to‑fine registration strategy. First, we introduce a global alignment module, structure‑aware perspective n‑point (SA‑PnP), to establish correspondence between 2D and 3D vessel structures. Second, we propose TempDiffReg, a temporal diffusion model that performs vessel deformation iteratively by leveraging temporal context to capture complex anatomical variations and local structural changes. We collected data from 23 patients and constructed 626 paired multi‑frame samples for comprehensive evaluation. Experimental results demonstrate that the proposed method consistently outperforms state‑of‑the‑art (SOTA) methods in both accuracy and anatomical plausibility. Specifically, our method achieves a mean squared error (MSE) of 0.63 mm and a mean absolute error (MAE) of 0.51 mm in registration accuracy, representing 66.7% lower MSE and 17.7% lower MAE compared to the most competitive existing approaches. It has the potential to assist less‑experienced clinicians in safely and efficiently performing complex TACE procedures, ultimately enhancing both surgical outcomes and patient care. Code and data are available at: \textcolorbluehttps://github.com/LZH970328/TempDiffReg.git
Authors:Aniket Rege, Arka Sadhu, Yuliang Li, Kejie Li, Ramya Korlakai Vinayak, Yuning Chai, Yong Jae Lee, Hyo Jin Kim
Abstract:
The advent of always‑on personal AI assistants, enabled by all‑day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous, longitudinal stream of egocentric video. Achieving this vision requires advances in long‑horizon video understanding, where systems must interpret and recall visual and audio information spanning days or even weeks. Existing methods, including large language models and retrieval‑augmented generation, are constrained by limited context windows and lack the ability to perform compositional, multi‑hop reasoning over very long video streams. In this work, we address these challenges through EGAgent, an enhanced agentic framework centered on entity scene graphs, which represent people, places, objects, and their relationships over time. Our system equips a planning agent with tools for structured search and reasoning over these graphs, as well as hybrid visual and audio search capabilities, enabling detailed, cross‑modal, and temporally coherent reasoning. Experiments on the EgoLifeQA and Video‑MME (Long) datasets show that our method achieves state‑of‑the‑art performance on EgoLifeQA (57.5%) and competitive performance on Video‑MME (Long) (74.1%) for complex longitudinal video understanding tasks. Code is available at https://github.com/facebookresearch/egagent.
Authors:Asiegbu Miracle Kanu-Asiegbu, Nitin Jotwani, Xiaoxiao Du
Abstract:
Pedestrian detection is a critical task in robot perception. Multispectral modalities (visible light and thermal) can boost pedestrian detection performance by providing complementary visual information. Several gaps remain with multispectral pedestrian detection methods. First, existing approaches primarily focus on spatial fusion and often neglect temporal information. Second, RGB and thermal image pairs in multispectral benchmarks may not always be perfectly aligned. Pedestrians are also challenging to detect due to varying lighting conditions, occlusion, etc. This work proposes Strip‑Fusion, a spatial‑temporal fusion network that is robust to misalignment in input images, as well as varying lighting conditions and heavy occlusions. The Strip‑Fusion pipeline integrates temporally adaptive convolutions to dynamically weigh spatial‑temporal features, enabling our model to better capture pedestrian motion and context over time. A novel Kullback‑Leibler divergence loss was designed to mitigate modality imbalance between visible and thermal inputs, guiding feature alignment toward the more informative modality during training. Furthermore, a novel post‑processing algorithm was developed to reduce false positives. Extensive experimental results show that our method performs competitively for both the KAIST and the CVC‑14 benchmarks. We also observed significant improvements compared to previous state‑of‑the‑art on challenging conditions such as heavy occlusion and misalignment.
Authors:Vi Vu, Thanh-Huy Nguyen, Tien-Thinh Nguyen, Ba-Thinh Lam, Hoang-Thien Nguyen, Tianyang Wang, Xingjian Li, Min Xu
Abstract:
Foundation models like the Segment Anything Model (SAM) show strong generalization, yet adapting them to medical images remains difficult due to domain shift, scarce labels, and the inability of Parameter‑Efficient Fine‑Tuning (PEFT) to exploit unlabeled data. While conventional models like U‑Net excel in semi‑supervised medical learning, their potential to assist a PEFT SAM has been largely overlooked. We introduce SC‑SAM, a specialist‑generalist framework where U‑Net provides point‑based prompts and pseudo‑labels to guide SAM's adaptation, while SAM serves as a powerful generalist supervisor to regularize U‑Net. This reciprocal guidance forms a bidirectional co‑training loop that allows both models to effectively exploit the unlabeled data. Across prostate MRI and polyp segmentation benchmarks, our method achieves state‑of‑the‑art results, outperforming other existing semi‑supervised SAM variants and even medical foundation models like MedSAM, highlighting the value of specialist‑generalist cooperation for label‑efficient medical image segmentation. Our code is available at https://github.com/vnlvi2k3/SC‑SAM.
Authors:Dain Kim, Jiwoo Lee, Jaehoon Yun, Yong Hoe Koo, Qingyu Chen, Hyunjae Kim, Jaewoo Kang
Abstract:
Large Vision‑Language Models (LVLMs) hold significant promise for medical applications, yet their deployment is often constrained by insufficient alignment and reliability. While Direct Preference Optimization (DPO) has emerged as a potent framework for refining model responses, its efficacy in high‑stakes medical contexts remains underexplored, lacking the rigorous empirical groundwork necessary to guide future methodological advances. To bridge this gap, we present the first comprehensive examination of diverse DPO variants within the medical domain, evaluating nine distinct formulations across two medical LVLMs: LLaVA‑Med and HuatuoGPT‑Vision. Our results reveal several critical limitations: current DPO approaches often yield inconsistent gains over supervised fine‑tuning, with their efficacy varying significantly across different tasks and backbones. Furthermore, they frequently fail to resolve fundamental visual misinterpretation errors. Building on these insights, we present a targeted preference construction strategy as a proof‑of‑concept that explicitly addresses visual misinterpretation errors frequently observed in existing DPO models. This design yields a 3.6% improvement over the strongest existing DPO baseline on visual question‑answering tasks. To support future research, we release our complete framework, including all training data, model checkpoints, and our codebase at https://github.com/dmis‑lab/med‑vlm‑dpo.
Authors:Zhongyu Xiao, Zhiwei Hao, Jianyuan Guo, Yong Luo, Jia Liu, Jie Xu, Han Hu
Abstract:
Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirectional attention to achieve superior global coherence compared to autoregressive models. While recent works have accelerated inference via KV cache reuse or heuristic decoding, they overlook the intrinsic inefficiencies within the block‑wise diffusion process. Specifically, they suffer from spatial redundancy by modeling informative‑sparse suffix regions uniformly and temporal inefficiency by applying fixed denoising schedules across all the decoding process. To address this, we propose Streaming‑dLLM, a training‑free framework that streamlines inference across both spatial and temporal dimensions. Spatially, we introduce attenuation guided suffix modeling to approximate the full context by pruning redundant mask tokens. Temporally, we employ a dynamic confidence aware strategy with an early exit mechanism, allowing the model to skip unnecessary iterations for converged tokens. Extensive experiments show that Streaming‑dLLM achieves up to 68.2X speedup while maintaining generation quality, highlighting its effectiveness in diffusion decoding. The code is available at https://github.com/xiaoshideta/Streaming‑dLLM.
Authors:Qingyu Fan, Zhaoxiang Li, Yi Lu, Wang Chen, Qiu Shen, Xiao-xiao Long, Yinghao Cai, Tao Lu, Shuo Wang, Xun Cao
Abstract:
Bimanual manipulation in cluttered scenes requires policies that remain stable under occlusions, viewpoint and scene variations. Existing vision‑language‑action models often fail to generalize because (i) multi‑view features are fused via view‑agnostic token concatenation, yielding weak 3D‑consistent spatial understanding, and (ii) language is injected as global conditioning, resulting in coarse instruction grounding.
In this paper, we introduce PEAfowl, a perception‑enhanced multi‑view VLA policy for bimanual manipulation. For spatial reasoning, PEAfowl predicts per‑token depth distributions, performs differentiable 3D lifting, and aggregates local cross‑view neighbors to form geometrically grounded, cross‑view consistent representations. For instruction grounding, we propose to replace global conditioning with a Perceiver‑style text‑aware readout over frozen CLIP visual features, enabling iterative evidence accumulation. To overcome noisy and incomplete commodity depth without adding inference overhead, we apply training‑only depth distillation from a pretrained depth teacher to supervise the depth‑distribution head, providing perception front‑end with geometry‑aware priors.
On RoboTwin 2.0 under domain‑randomized setting, PEAfowl improves the strongest baseline by 23.0 pp in success rate, and real‑robot experiments further demonstrate reliable sim‑to‑real transfer and consistent improvements from depth distillation.
Project website: https://peafowlvla.github.io/.
Authors:Zhihao He, Tieyuan Chen, Kangyu Wang, Ziran Qin, Yang Shao, Chaofan Gan, Shijie Li, Zuxuan Wu, Weiyao Lin
Abstract:
Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this AR paradigm inevitably faces a dual efficiency bottleneck: strictly unidirectional attention compromises understanding efficiency by hindering global spatiotemporal aggregation, while serial decoding restricts generation efficiency. To address this, we propose VidLaDA, a Video LLM based on Diffusion Language Models (DLMs) that leverages bidirectional attention to unlock comprehensive spatiotemporal modeling and decode tokens in parallel. To further mitigate the computational overhead of diffusion decoding, we introduce MARS‑Cache, an acceleration strategy that prunes redundancy by combining asynchronous visual cache refreshing with frame‑wise chunk attention. Experiments show VidLaDA rivals state‑of‑the‑art AR baselines (e.g., Qwen2.5‑VL and LLaVA‑Video) and outperforms DLM baselines, with MARS‑Cache delivering over 12x speedup without compromising accuracy. Code and checkpoints are open‑sourced at https://github.com/ziHoHe/VidLaDA.
Authors:Yoonwoo Jeong, Cheng Sun, Yu-Chiang Frank Wang, Minsu Cho, Jaesung Choe
Abstract:
Promptable segmentation has emerged as a powerful paradigm in computer vision, enabling users to guide models in parsing complex scenes with prompts such as clicks, boxes, or textual cues. Recent advances, exemplified by the Segment Anything Model (SAM), have extended this paradigm to videos and multi‑view images. However, the lack of 3D awareness often leads to inconsistent results, necessitating costly per‑scene optimization to enforce 3D consistency. In this work, we introduce MV‑SAM, a framework for multi‑view segmentation that achieves 3D consistency using pointmaps ‑‑ 3D points reconstructed from unposed images by recent visual geometry models. Leveraging the pixel‑point one‑to‑one correspondence of pointmaps, MV‑SAM lifts images and prompts into 3D space, eliminating the need for explicit 3D networks or annotated 3D data. Specifically, MV‑SAM extends SAM by lifting image embeddings from its pretrained encoder into 3D point embeddings, which are decoded by a transformer using cross‑attention with 3D prompt embeddings. This design aligns 2D interactions with 3D geometry, enabling the model to implicitly learn consistent masks across views through 3D positional embeddings. Trained on the SA‑1B dataset, our method generalizes well across domains, outperforming SAM2‑Video and achieving comparable performance with per‑scene optimization baselines on NVOS, SPIn‑NeRF, ScanNet++, uCo3D, and DL3DV benchmarks. Code will be released.
Authors:Ziyang Song, Xinyu Gong, Bangya Liu, Zelin Zhao
Abstract:
Existing Subject‑to‑Video Generation (S2V) methods have achieved high‑fidelity and subject‑consistent video generation, yet remain constrained to single‑view subject references. This limitation renders the S2V task reducible to an S2I + I2V pipeline, failing to exploit the full potential of video subject control. In this work, we propose and address the challenging Multi‑View S2V (MV‑S2V) task, which synthesizes videos from multiple reference views to enforce 3D‑level subject consistency. Regarding the scarcity of training data, we first develop a synthetic data curation pipeline to generate highly customized synthetic data, complemented by a small‑scale real‑world captured dataset to boost the training of MV‑S2V. Another key issue lies in the potential confusion between cross‑subject and cross‑view references in conditional generation. To overcome this, we further introduce Temporally Shifted RoPE (TS‑RoPE) to distinguish between different subjects and distinct views of the same subject in reference conditioning. Our framework achieves superior 3D subject consistency w.r.t. multi‑view reference images and high‑quality visual outputs, establishing a new meaningful direction for subject‑driven video generation. Code and data are available at: https://szy‑young.github.io/mv‑s2v
Authors:Saptarshi Ghosh, Linfeng Liu, Tianyu Jiang
Abstract:
Images often communicate more than they literally depict: a set of tools can suggest an occupation and a cultural artifact can suggest a tradition. This kind of indirect visual reference, known as visual metonymy, invites viewers to recover a target concept via associated cues rather than explicit depiction. In this work, we present the first computational investigation of visual metonymy. We introduce a novel pipeline grounded in semiotic theory that leverages large language models and text‑to‑image models to generate metonymic visual representations. Using this framework, we construct ViMET, the first visual metonymy dataset comprising 2,000 multiple‑choice questions to evaluate the cognitive reasoning abilities in multimodal language models. Experimental results on our dataset reveal a significant gap between human performance (86.9%) and state‑of‑the‑art vision‑language models (65.9%), highlighting limitations in machines' ability to interpret indirect visual references. Our dataset is publicly available at: https://github.com/cincynlp/ViMET.
Authors:Taewan Cho, Taeryang Kim, Andrew Jaeyong Choi
Abstract:
Robotic and autonomous systems need dense spatial cues, but many monocular depth models are heavy, task‑specific, or hard to attach to an existing multimodal stack. CLIP offers strong semantic representations, yet most CLIP‑based depth methods still depend on text prompts or backbone updates, which complicate deployment in integrated control pipelines. We present SPACE‑CLIP, a decoder‑only depth framework that reads geometric cues directly from a frozen CLIP vision encoder and bypasses the text encoder at inference time. The model combines FiLM‑conditioned semantic features from deep layers with structural features from shallow layers to recover both global scene layout and local geometric detail. Under the TFI‑FB constraint (text‑free inference and frozen vision backbone), SPACE‑CLIP achieves AbsRel 0.0901 on KITTI and 0.1042 on NYU Depth V2, and the same dual‑pathway decoder transfers to a frozen SigLIP backbone with comparable results. These findings show that a compact decoder can turn a shared foundation‑model backbone into a reusable spatial perception module for embodied AI and autonomous robotic systems. Our model is available at https://github.com/taewan2002/space‑clip
Authors:Sebastian Doerrich, Francesco Di Salvo, Jonas Alle, Christian Ledig
Abstract:
Deep learning models in medical image analysis often struggle with generalizability across domains and demographic groups due to data heterogeneity and scarcity. Traditional augmentation improves robustness, but fails under substantial domain shifts. Recent advances in stylistic augmentation enhance domain generalization by varying image styles but fall short in terms of style diversity or by introducing artifacts into the generated images. To address these limitations, we propose Stylizing ViT, a novel Vision Transformer encoder that utilizes weight‑shared attention blocks for both self‑ and cross‑attention. This design allows the same attention block to maintain anatomical consistency through self‑attention while performing style transfer via cross‑attention. We assess the effectiveness of our method for domain generalization by employing it for data augmentation on three distinct image classification tasks in the context of histopathology and dermatology. Results demonstrate an improved robustness (up to +13% accuracy) over the state of the art while generating perceptually convincing images without artifacts. Additionally, we show that Stylizing ViT is effective beyond training, achieving a 17% performance improvement during inference when used for test‑time augmentation. The source code is available at https://github.com/sdoerrich97/stylizing‑vit .
Authors:Fengting Zhang, Yue He, Qinghao Liu, Yaonan Wang, Xiang Chen, Hang Zhang
Abstract:
Deep learning has revolutionized medical image registration by achieving unprecedented speeds, yet its clinical application is hindered by a limited ability to generalize beyond the training domain, a critical weakness given the typically small scale of medical datasets. In this paper, we introduce FMIR, a foundation model‑based registration framework that overcomes this limitation.Combining a foundation model‑based feature encoder for extracting anatomical structures with a general registration head, and trained with a channel regularization strategy on just a single dataset, FMIR achieves state‑of‑the‑art(SOTA) in‑domain performance while maintaining robust registration on out‑of‑domain images.Our approach demonstrates a viable path toward building generalizable medical imaging foundation models with limited resources. The code is available at https://github.com/Monday0328/FMIR.git.
Authors:Yan Zhou, Zhen Huang, Yingqiu Li, Yue Ouyang, Suncheng Xiang, Zehua Wang
Abstract:
Accurate brain tumor segmentation from multi‑modal magnetic resonance imaging (MRI) is a prerequisite for precise radiotherapy planning and surgical navigation. While recent Transformer‑based models such as Swin UNETR have achieved impressive benchmark performance, their clinical utility is often compromised by two critical issues: sensitivity to missing modalities (common in clinical practice) and a lack of confidence calibration. Merely chasing higher Dice scores on idealized data fails to meet the safety requirements of real‑world medical deployment. In this work, we propose BMDS‑Net, a unified framework that prioritizes clinical robustness and trustworthiness over simple metric maximization. Our contribution is three‑fold. First, we construct a robust deterministic backbone by integrating a Zero‑Init Multimodal Contextual Fusion (MMCF) module and a Residual‑Gated Deep Decoder Supervision (DDS) mechanism, enabling stable feature learning and precise boundary delineation with significantly reduced Hausdorff Distance, even under modality corruption. Second, and most importantly, we introduce a memory‑efficient Bayesian fine‑tuning strategy that transforms the network into a probabilistic predictor, providing voxel‑wise uncertainty maps to highlight potential errors for clinicians. Third, comprehensive experiments on the BraTS 2021 dataset demonstrate that BMDS‑Net not only maintains competitive accuracy but, more importantly, exhibits superior stability in missing‑modality scenarios where baseline models fail. The source code is publicly available at https://github.com/RyanZhou168/BMDS‑Net.
Authors:Chia-Ming Lee, Yu-Fan Lin, Yu-Jou Hsiao, Jin-Hui Jiang, Yu-Lun Liu, Chih-Chung Hsu
Abstract:
Shadow removal under diverse lighting conditions requires disentangling illumination from intrinsic reflectance, a challenge compounded when physical priors are not properly aligned. We propose PhaSR (Physically Aligned Shadow Removal), addressing this through dual‑level prior alignment to enable robust performance from single‑light shadows to multi‑source ambient lighting. First, Physically Aligned Normalization (PAN) performs closed‑form illumination correction via Gray‑world normalization, log‑domain Retinex decomposition, and dynamic range recombination, suppressing chromatic bias. Second, Geometric‑Semantic Rectification Attention (GSRA) extends differential attention to cross‑modal alignment, harmonizing depth‑derived geometry with DINO‑v2 semantic embeddings to resolve modal conflicts under varying illumination. Experiments show competitive performance in shadow removal with lower complexity and generalization to ambient lighting where traditional methods fail under multi‑source illumination. Our source code is available at https://github.com/ming053l/PhaSR.
Authors:Chia-Ming Lee, Yu-Fan Lin, Jin-Hui Jiang, Yu-Jou Hsiao, Chih-Chung Hsu, Yu-Lun Liu
Abstract:
Single Image Reflection Separation (SIRS) disentangles mixed images into transmission and reflection layers. Existing methods suffer from transmission‑reflection confusion under nonlinear mixing, particularly in deep decoder layers, due to implicit fusion mechanisms and inadequate multi‑scale coordination. We propose ReflexSplit, a dual‑stream framework with three key innovations. (1) Cross‑scale Gated Fusion (CrGF) adaptively aggregates semantic priors, texture details, and decoder context across hierarchical depths, stabilizing gradient flow and maintaining feature consistency. (2) Layer Fusion‑Separation Blocks (LFSB) alternate between fusion for shared structure extraction and differential separation for layer‑specific disentanglement. Inspired by Differential Transformer, we extend attention cancellation to dual‑stream separation via cross‑stream subtraction. (3) Curriculum training progressively strengthens differential separation through depth‑dependent initialization and epoch‑wise warmup. Extensive experiments on synthetic and real‑world benchmarks demonstrate state‑of‑the‑art performance with superior perceptual quality and robust generalization. Our code is available at https://github.com/wuw2135/ReflexSplit.
Authors:Fangyijie Wang, Siteng Ma, Guénolé Silvestre, Kathleen M. Curran
Abstract:
Fetal ultrasound (US) data is often limited due to privacy and regulatory restrictions, posing challenges for training deep learning (DL) models. While semi‑supervised learning (SSL) is commonly used for fetal US image analysis, existing SSL methods typically rely on random limited selection, which can lead to suboptimal model performance by overfitting to homogeneous labeled data. To address this, we propose a two‑stage Active Learning (AL) sampler, Entropy‑Guided Agreement‑Diversity (EGAD), for fetal head segmentation. Our method first selects the most uncertain samples using predictive entropy, and then refines the final selection using the agreement‑diversity score combining cosine similarity and mutual information. Additionally, our SSL framework employs a consistency learning strategy with feature downsampling to further enhance segmentation performance. In experiments, SSL‑EGAD achieves an average Dice score of 94.57% and 96.32% on two public datasets for fetal head segmentation, using 5% and 10% labeled data for training, respectively. Our method outperforms current SSL models and showcases consistent robustness across diverse pregnancy stage data. The code is available on \hrefhttps://github.com/13204942/Semi‑supervised‑EGADGitHub.
Authors:Shiu-hong Kao, Chak Ho Huang, Huaiqian Liu, Yu-Wing Tai, Chi-Keung Tang
Abstract:
Existing works of reasoning segmentation often fall short in complex cases, particularly when addressing complicated queries and out‑of‑domain images. Inspired by the chain‑of‑thought reasoning, where harder problems require longer thinking steps/time, this paper aims to explore a system that can think step‑by‑step, look up information if needed, generate results, self‑evaluate its own results, and refine the results, in the same way humans approach harder questions. We introduce CoT‑Seg, a training‑free framework that rethinks reasoning segmentation by combining chain‑of‑thought reasoning with self‑correction. Instead of fine‑tuning, CoT‑Seg leverages the inherent reasoning ability of pre‑trained MLLMs (GPT‑4o) to decompose queries into meta‑instructions, extract fine‑grained semantics from images, and identify target objects even under implicit or complex prompts. Moreover, CoT‑Seg incorporates a self‑correction stage: the model evaluates its own segmentation against the original query and reasoning trace, identifies mismatches, and iteratively refines the mask. This tight integration of reasoning and correction significantly improves reliability and robustness, especially in ambiguous or error‑prone cases. Furthermore, our CoT‑Seg framework allows easy incorporation of retrieval‑augmented reasoning, enabling the system to access external knowledge when the input lacks sufficient information. To showcase CoT‑Seg's ability to handle very challenging cases ,we introduce a new dataset ReasonSeg‑Hard. Our results highlight that combining chain‑of‑thought reasoning, self‑correction, offers a powerful paradigm for vision‑language integration driven segmentation.
Authors:Xuan Ding, Xiu Yan, Chuanlong Xie, Yao Zhu
Abstract:
Watermarking methods have always been effective means of protecting intellectual property, yet they face significant challenges. Although existing deep learning‑based watermarking systems can hide watermarks in images with minimal impact on image quality, they often lack robustness when encountering image corruptions during transmission, which undermines their practical application value. To this end, we propose a high‑quality and robust watermark framework based on the diffusion model. Our method first converts the clean image into inversion noise through a null‑text optimization process, and after optimizing the inversion noise in the latent space, it produces a high‑quality watermarked image through an iterative denoising process of the diffusion model. The iterative denoising process serves as a powerful purification mechanism, ensuring both the visual quality of the watermarked image and enhancing the robustness of the watermark against various corruptions. To prevent the optimizing of inversion noise from distorting the original semantics of the image, we specifically introduced self‑attention constraints and pseudo‑mask strategies. Extensive experimental results demonstrate the superior performance of our method against various image corruptions. In particular, our method outperforms the stable signature method by an average of 10% across 12 different image transformations on COCO datasets. Our codes are available at https://github.com/920927/ONRW.
Authors:Chen Ling, Kai Hu, Hangcheng Liu, Xingshuo Han, Tianwei Zhang, Changhai Ou
Abstract:
Large Vision‑Language Models (LVLMs) are increasingly deployed in real‑world intelligent systems for perception and reasoning in open physical environments. While LVLMs are known to be vulnerable to prompt injection attacks, existing methods either require access to input channels or depend on knowledge of user queries, assumptions that rarely hold in practical deployments. We propose the first Physical Prompt Injection Attack (PPIA), a black‑box, query‑agnostic attack that embeds malicious typographic instructions into physical objects perceivable by the LVLM. PPIA requires no access to the model, its inputs, or internal pipeline, and operates solely through visual observation. It combines offline selection of highly recognizable and semantically effective visual prompts with strategic environment‑aware placement guided by spatiotemporal attention, ensuring that the injected prompts are both perceivable and influential on model behavior. We evaluate PPIA across 10 state‑of‑the‑art LVLMs in both simulated and real‑world settings on tasks including visual question answering, planning, and navigation, PPIA achieves attack success rates up to 98%, with strong robustness under varying physical conditions such as distance, viewpoint, and illumination. Our code is publicly available at https://github.com/2023cghacker/Physical‑Prompt‑Injection‑Attack.
Authors:Chengbo Ding, Fenghe Tang, Shaohua Kevin Zhou
Abstract:
Existing displacement strategies in semi‑supervised segmentation only operate on rectangular regions, ignoring anatomical structures and resulting in boundary distortions and semantic inconsistency. To address these issues, we propose UCAD, an Uncertainty‑Guided Contour‑Aware Displacement framework for semi‑supervised medical image segmentation that preserves contour‑aware semantics while enhancing consistency learning. Our UCAD leverages superpixels to generate anatomically coherent regions aligned with anatomy boundaries, and an uncertainty‑guided selection mechanism to selectively displace challenging regions for better consistency learning. We further propose a dynamic uncertainty‑weighted consistency loss, which adaptively stabilizes training and effectively regularizes the model on unlabeled regions. Extensive experiments demonstrate that UCAD consistently outperforms state‑of‑the‑art semi‑supervised segmentation methods, achieving superior segmentation accuracy under limited annotation. The code is available at:https://github.com/dcb937/UCAD.
Authors:Fabian Vazquez, Jose A. Nuñez, Diego Adame, Alissen Moreno, Augustin Zhan, Huimin Li, Jinghao Yang, Haoteng Tang, Bin Fu, Pengfei Gu
Abstract:
Accurate and robust polyp segmentation is essential for early colorectal cancer detection and for computer‑aided diagnosis. While convolutional neural network‑, Transformer‑, and Mamba‑based U‑Net variants have achieved strong performance, they still struggle to capture geometric and structural cues, especially in low‑contrast or cluttered colonoscopy scenes. To address this challenge, we propose a novel Geometric Prior‑guided Module (GPM) that injects explicit geometric priors into U‑Net‑based architectures for polyp segmentation. Specifically, we fine‑tune the Visual Geometry Grounded Transformer (VGGT) on a simulated ColonDepth dataset to estimate depth maps of polyp images tailored to the endoscopic domain. These depth maps are then processed by GPM to encode geometric priors into the encoder's feature maps, where they are further refined using spatial and channel attention mechanisms that emphasize both local spatial and global channel information. GPM is plug‑and‑play and can be seamlessly integrated into diverse U‑Net variants. Extensive experiments on five public polyp segmentation datasets demonstrate consistent gains over three strong baselines. Code and the generated depth maps are available at: https://github.com/fvazqu/GPM‑PolypSeg
Authors:Debang Li, Zhengcong Fei, Tuanhui Li, Yikun Dou, Zheng Chen, Jiangping Yang, Mingyuan Fan, Jingtao Xu, Jiahua Wang, Baoxuan Gu, Mingshan Chang, Wenjing Cai, Yuqiang Xie, Binjie Mao, Youqiang Zhang, Nuo Pang, Hao Zhang, Yuzhe Jin, Zhiheng Xu, Dixuan Lin, Guibin Chen, Yahui Zhou
Abstract:
Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReels‑V3, a conditional video generation model, built upon a unified multimodal in‑context learning framework with diffusion Transformers. SkyReels‑V3 model supports three core generative paradigms within a single architecture: reference images‑to‑video synthesis, video‑to‑video extension and audio‑guided video generation. (i) reference images‑to‑video model is designed to produce high‑fidelity videos with strong subject identity preservation, temporal coherence, and narrative consistency. To enhance reference adherence and compositional stability, we design a comprehensive data processing pipeline that leverages cross frame pairing, image editing, and semantic rewriting, effectively mitigating copy paste artifacts. During training, an image video hybrid strategy combined with multi‑resolution joint optimization is employed to improve generalization and robustness across diverse scenarios. (ii) video extension model integrates spatio‑temporal consistency modeling with large‑scale video understanding, enabling both seamless single‑shot continuation and intelligent multi‑shot switching with professional cinematographic patterns. (iii) Talking avatar model supports minute‑level audio‑conditioned video generation by training first‑and‑last frame insertion patterns and reconstructing key‑frame inference paradigms. On the basis of ensuring visual quality, synchronization of audio and videos has been optimized.
Extensive evaluations demonstrate that SkyReels‑V3 achieves state‑of‑the‑art or near state‑of‑the‑art performance on key metrics including visual quality, instruction following, and specific aspect metrics, approaching leading closed‑source systems. Github: https://github.com/SkyworkAI/SkyReels‑V3.
Authors:Kun Huang, Fang-Lue Zhang, Neil Dodgson
Abstract:
360° depth estimation is a challenging research problem due to the difficulty of finding a representation that both preserves global continuity and avoids distortion in spherical images. Existing methods attempt to leverage complementary information from multiple projections, but struggle with balancing global and local consistency. Their local patch features have limited global perception, and the combined global representation does not address discrepancies in feature extraction at the boundaries between patches. To address these issues, we propose Cross360, a novel cross‑attention‑based architecture integrating local and global information using less‑distorted tangent patches along with equirectangular features. Our Cross Projection Feature Alignment module employs cross‑attention to align local tangent projection features with the equirectangular projection's 360° field of view, ensuring each tangent projection patch is aware of the global context. Additionally, our Progressive Feature Aggregation with Attention module refines multi‑scaled features progressively, enhancing depth estimation accuracy. Cross360 significantly outperforms existing methods across most benchmark datasets, especially those in which the entire 360° image is available, demonstrating its effectiveness in accurate and globally consistent depth estimation. The code and model are available at https://github.com/huangkun101230/Cross360.
Authors:Bin Lin, Zongjian Li, Yuwei Niu, Kaixiong Gong, Yunyang Ge, Yunlong Lin, Mingzhe Zheng, JianWei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Li Yuan
Abstract:
The field of image generation is currently bifurcated into autoregressive (AR) models operating on discrete tokens and diffusion models utilizing continuous latents. This divide, rooted in the distinction between VQ‑VAEs and VAEs, hinders unified modeling and fair benchmarking. Finite Scalar Quantization (FSQ) offers a theoretical bridge, yet vanilla FSQ suffers from a critical flaw: its equal‑interval quantization can cause activation collapse. This mismatch forces a trade‑off between reconstruction fidelity and information efficiency. In this work, we resolve this dilemma by simply replacing the activation function in original FSQ with a distribution‑matching mapping to enforce a uniform prior. Termed iFSQ, this simple strategy requires just one line of code yet mathematically guarantees both optimal bin utilization and reconstruction precision. Leveraging iFSQ as a controlled benchmark, we uncover two key insights: (1) The optimal equilibrium between discrete and continuous representations lies at approximately 4 bits per dimension. (2) Under identical reconstruction constraints, AR models exhibit rapid initial convergence, whereas diffusion models achieve a superior performance ceiling, suggesting that strict sequential ordering may limit the upper bounds of generation quality. Finally, we extend our analysis by adapting Representation Alignment (REPA) to AR models, yielding LlamaGen‑REPA. Codes is available at https://github.com/Tencent‑Hunyuan/iFSQ
Authors:Qinkai Yu, Chong Zhang, Gaojie Jin, Tianjin Huang, Wei Zhou, Wenhui Li, Xiaobo Jin, Bo Huang, Yitian Zhao, Guang Yang, Gregory Y. H. Lip, Yalin Zheng, Aline Villavicencio, Yanda Meng
Abstract:
Annotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive patient information. As a result, well‑trained medical segmentation models on private datasets constitute valuable intellectual property requiring robust protection mechanisms. Existing model protection techniques primarily focus on classification and generative tasks, while segmentation models‑crucial to medical image analysis‑remain largely underexplored. In this paper, we propose a novel, stealthy, and harmless method, StealthMark, for verifying the ownership of medical segmentation models under black‑box conditions. Our approach subtly modulates model uncertainty without altering the final segmentation outputs, thereby preserving the model's performance. To enable ownership verification, we incorporate model‑agnostic explanation methods, e.g. LIME, to extract feature attributions from the model outputs. Under specific triggering conditions, these explanations reveal a distinct and verifiable watermark. We further design the watermark as a QR code to facilitate robust and recognizable ownership claims. We conducted extensive experiments across four medical imaging datasets and five mainstream segmentation models. The results demonstrate the effectiveness, stealthiness, and harmlessness of our method on the original model's segmentation performance. For example, when applied to the SAM model, StealthMark consistently achieved ASR above 95% across various datasets while maintaining less than a 1% drop in Dice and AUC scores, significantly outperforming backdoor‑based watermarking methods and highlighting its strong potential for practical deployment. Our implementation code is made available at: https://github.com/Qinkaiyu/StealthMark.
Authors:Rui-Yang Ju, Jen-Shiun Chiang
Abstract:
Virtual try‑on systems allow users to interactively try different products within VR scenarios. However, most existing VTON methods operate only on predefined eyewear templates and lack support for fine‑grained, user‑driven customization. While GlassesGAN enables personalized 2D eyewear design, its capability remains limited to 2D image generation. Motivated by the success of 3D Gaussian Blendshapes in head reconstruction, we integrate these two techniques and propose GlassesGB, a framework that supports customizable eyewear generation for 3D head avatars. GlassesGB effectively bridges 2D generative customization with 3D head avatar rendering, addressing the challenge in achieving personalized eyewear design for VR applications. The implementation code is available at https://ruiyangju.github.io/GlassesGB.
Authors:Aahana Basappa, Pranay Goel, Anusri Karra, Anish Karra, Asa Gilmore, Kevin Zhu
Abstract:
We investigated visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel benchmark to systematically compare failure modes across image‑to‑text and text‑to‑image tasks, enabling cross‑modal evaluation of visual understanding. Despite rapid growth in machine learning, vision language models (VLMs) still fail to understand or generate basic visual concepts such as object orientation, quantity, or spatial relationships, which highlighted gaps in elementary visual reasoning. By adapting MMVP benchmark questions into explicit and implicit prompts, we create AMVICC, a novel benchmark for profiling failure modes across various modalities. After testing 11 MLLMs and 3 IGMs in nine categories of visual reasoning, our results show that failure modes are often shared between models and modalities, but certain failures are model‑specific and modality‑specific, and this can potentially be attributed to various factors. IGMs consistently struggled to manipulate specific visual components in response to prompts, especially in explicit prompts, suggesting poor control over fine‑grained visual attributes. Our findings apply most directly to the evaluation of existing state‑of‑the‑art models on structured visual reasoning tasks. This work lays the foundation for future cross‑modal alignment studies, offering a framework to probe whether generation and interpretation failures stem from shared limitations to guide future improvements in unified vision‑language modeling.
Authors:Basile Van Hoorick, Dian Chen, Shun Iwase, Pavel Tokmakov, Muhammad Zubair Irshad, Igor Vasiljevic, Swati Gupta, Fangzhou Cheng, Sergey Zakharov, Vitor Campagnolo Guizilini
Abstract:
Modern generative video models excel at producing convincing, high‑quality outputs, but struggle to maintain multi‑view and spatiotemporal consistency in highly dynamic real‑world environments. In this work, we introduce AnyView, a diffusion‑based video generation framework for \emphdynamic view synthesis with minimal inductive biases or geometric assumptions. We leverage multiple data sources with various levels of supervision, including monocular (2D), multi‑view static (3D) and multi‑view dynamic (4D) datasets, to train a generalist spatiotemporal implicit representation capable of producing zero‑shot novel videos from arbitrary camera locations and trajectories. We evaluate AnyView on standard benchmarks, showing competitive results with the current state of the art, and propose AnyViewBench, a challenging new benchmark tailored towards \emphextreme dynamic view synthesis in diverse real‑world scenarios. In this more dramatic setting, we find that most baselines drastically degrade in performance, as they require significant overlap between viewpoints, while AnyView maintains the ability to produce realistic, plausible, and spatiotemporally consistent videos when prompted from \emphany viewpoint. Results, data, code, and models can be viewed at: https://tri‑ml.github.io/AnyView/
Authors:Zirui Wang, Junyi Zhang, Jiaxin Ge, Long Lian, Letian Fu, Lisa Dunlap, Ken Goldberg, XuDong Wang, Ion Stoica, David M. Chan, Sewon Min, Joseph E. Gonzalez
Abstract:
Modern Vision‑Language Models (VLMs) remain poorly characterized in multi‑step visual interactions, particularly in how they integrate perception, memory, and action over long horizons. We introduce VisGym, a gymnasium of 17 environments for evaluating and training VLMs. The suite spans symbolic puzzles, real‑image understanding, navigation, and manipulation, and provides flexible controls over difficulty, input representation, planning horizon, and feedback. We also provide multi‑step solvers that generate structured demonstrations, enabling supervised finetuning. Our evaluations show that all frontier models struggle in interactive settings, achieving low success rates in both the easy (46.6%) and hard (26.0%) configurations. Our experiments reveal notable limitations: models struggle to effectively leverage long context, performing worse with an unbounded history than with truncated windows. Furthermore, we find that several text‑based symbolic tasks become substantially harder once rendered visually. However, explicit goal observations, textual feedback, and exploratory demonstrations in partially observable or unknown‑dynamics settings for supervised finetuning yield consistent gains, highlighting concrete failure modes and pathways for improving multi‑step visual decision‑making. Code, data, and models can be found at: https://visgym.github.io/.
Authors:Yangfan Xu, Lilian Zhang, Xiaofeng He, Pengdong Wu, Wenqi Wu, Jun Mao
Abstract:
Transformer‑based general visual geometry frameworks have shown promising performance in camera pose estimation and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise in camera pose estimation and 3D reconstruction. However, these models typically rely on ground truth labels for training, posing challenges when adapting to unlabeled and unseen scenes. In this paper, we propose a self‑supervised framework to train VGGT with unlabeled data, thereby enhancing its localization capability in large‑scale environments. To achieve this, we extend conventional pair‑wise relations to sequence‑wise geometric constraints for self‑supervised learning. Specifically, in each sequence, we sample multiple source frames and geometrically project them onto different target frames, which improves temporal feature consistency. We formulate physical photometric consistency and geometric constraints as a joint optimization loss to circumvent the requirement for hard labels. By training the model with this proposed method, not only the local and global cross‑view attention layers but also the camera and depth heads can effectively capture the underlying multi‑view geometry. Experiments demonstrate that the model converges within hundreds of iterations and achieves significant improvements in large‑scale localization. Our code will be released at https://github.com/X‑yangfan/GPA‑VGGT.
Authors:Cuong Le, Pavlo Melnyk, Bastian Wandt, Mårten Wadenbäck
Abstract:
Recovering 3D human poses from a monocular camera view is a highly ill‑posed problem due to the depth ambiguity. Earlier studies on 3D human pose lifting from 2D often contain incorrect‑yet‑overconfident 3D estimations. To mitigate the problem, emerging probabilistic approaches treat the 3D estimations as a distribution, taking into account the uncertainty measurement of the poses. Falling in a similar category, we proposed FMPose, a probabilistic 3D human pose estimation method based on the flow matching generative approach. Conditioned on the 2D cues, the flow matching scheme learns the optimal transport from a simple source distribution to the plausible 3D human pose distribution via continuous normalizing flows. The 2D lifting condition is modeled via graph convolutional networks, leveraging the learnable connections between human body joints as the graph structure for feature aggregation. While trade‑offs between processing time and precision exist, already in the equal‑accuracy comparison, FMPose exhibits significantly faster processing time than the diffusion model, and also offers another faster and more accurate configuration. Experimental results show major improvements of our FMPose over current state‑of‑the‑art methods on two common benchmarks for 3D human pose estimation, namely Human3.6M, MPI‑INF‑3DHP. Additionally, FMPose shows competitive performance on the more challenging 3DPW dataset. The code implementation is available at https://github.com/cuongle1206/FMPose
Authors:Hongda Liu, Yunfan Liu, Min Ren, Lin Sui, Yunlong Wang, Zhenan Sun
Abstract:
In skeleton‑based human activity understanding, existing methods often adopt the contrastive learning paradigm to construct a discriminative feature space. However, many of these approaches fail to exploit the structural inter‑class similarities and overlook the impact of anomalous positive samples. In this study, we introduce ACLNet, an Affinity Contrastive Learning Network that explores the intricate clustering relationships among human activity classes to improve feature discrimination. Specifically, we propose an affinity metric to refine similarity measurements, thereby forming activity superclasses that provide more informative contrastive signals. A dynamic temperature schedule is also introduced to adaptively adjust the penalty strength for various superclasses. In addition, we employ a margin‑based contrastive strategy to improve the separation of hard positive and negative samples within classes. Extensive experiments on NTU RGB+D 60, NTU RGB+D 120, Kinetics‑Skeleton, PKU‑MMD, FineGYM, and CASIA‑B demonstrate the superiority of our method in skeleton‑based action recognition, gait recognition, and person re‑identification. The source code is available at https://github.com/firework8/ACLNet.
Authors:Minsu Gong, Nuri Ryu, Jungseul Ok, Sunghyun Cho
Abstract:
Recent advances in image editing leverage latent diffusion models (LDMs) for versatile, text‑prompt‑driven edits across diverse tasks. Yet, maintaining pixel‑level edge structures‑crucial for tasks such as photorealistic style transfer or image tone adjustment‑remains as a challenge for latent‑diffusion‑based editing. To overcome this limitation, we propose a novel Structure Preservation Loss (SPL) that leverages local linear models to quantify structural differences between input and edited images. Our training‑free approach integrates SPL directly into the diffusion model's generative process to ensure structural fidelity. This core mechanism is complemented by a post‑processing step to mitigate LDM decoding distortions, a masking strategy for precise edit localization, and a color preservation loss to preserve hues in unedited areas. Experiments confirm SPL enhances structural fidelity, delivering state‑of‑the‑art performance in latent‑diffusion‑based image editing. Our code will be publicly released at https://github.com/gongms00/SPL.
Authors:Ming Kang, Fung Fung Ting, Raphaël C. -W. Phan, Zongyuan Ge, Chee-Ming Ting
Abstract:
Nuclei panoptic segmentation supports cancer diagnostics by integrating both semantic and instance segmentation of different cell types to analyze overall tissue structure and individual nuclei in histopathology images. Major challenges include detecting small objects, handling ambiguous boundaries, and addressing class imbalance. To address these issues, we propose PanopMamba, a novel hybrid encoder‑decoder architecture that integrates Mamba and Transformer with additional feature‑enhanced fusion via state space modeling. We design a multiscale Mamba backbone and a State Space Model (SSM)‑based fusion network to enable efficient long‑range perception in pyramid features, thereby extending the pure encoder‑decoder framework while facilitating information sharing across multiscale features of nuclei. The proposed SSM‑based feature‑enhanced fusion integrates pyramid feature networks and dynamic feature enhancement across different spatial scales, enhancing the feature representation of densely overlapping nuclei in both semantic and spatial dimensions. To the best of our knowledge, this is the first Mamba‑based approach for panoptic segmentation. Additionally, we introduce alternative evaluation metrics, including image‑level Panoptic Quality (iPQ), boundary‑weighted PQ (wPQ), and frequency‑weighted PQ (fwPQ), which are specifically designed to address the unique challenges of nuclei segmentation and thereby mitigate the potential bias inherent in vanilla PQ. Experimental evaluations on two multiclass nuclei segmentation benchmark datasets, MoNuSAC2020 and NuInsSeg, demonstrate the superiority of PanopMamba for nuclei panoptic segmentation over state‑of‑the‑art methods. Consequently, the robustness of PanopMamba is validated across various metrics, while the distinctiveness of PQ variants is also demonstrated. Code is available at https://github.com/mkang315/PanopMamba.
Authors:Shuying Li, Yuchen Wang, San Zhang, Chuang Yang
Abstract:
Remote sensing change detection (RSCD) aims to identify the spatio‑temporal changes of land cover, providing critical support for multi‑disciplinary applications (e.g., environmental monitoring, disaster assessment, and climate change studies). Existing methods focus either on extracting features from localized patches, or pursue processing entire images holistically, which leads to the cross temporal feature matching deviation and exhibiting sensitivity to radiometric and geometric noise. Following the above issues, we propose a dual‑module collaboration guided hierarchical adaptive aggregation framework, namely HA2F, which consists of dynamic hierarchical feature calibration module (DHFCM) and noise‑adaptive feature refinement module (NAFRM). The former dynamically fuses adjacent‑level features through perceptual feature selection, suppressing irrelevant discrepancies to address multi‑temporal feature alignment deviations. The NAFRM utilizes the dual feature selection mechanism to highlight the change sensitive regions and generate spatial masks, suppressing the interference of irrelevant regions or shadows. Extensive experiments verify the effectiveness of the proposed HA2F, which achieves state‑of‑the‑art performance on LEVIR‑CD, WHU‑CD, and SYSU‑CD datasets, surpassing existing comparative methods in terms of both precision metrics and computational efficiency. In addition, ablation experiments show that DHFCM and NAFRM are effective. \hrefhttps://huggingface.co/InPeerReview/RemoteSensingChangeDetection‑RSCD.HA2FHA2F Official Code is Available Here!
Authors:Erik Wallin, Fredrik Kahl, Lars Hammarstrand
Abstract:
Hierarchical open‑set classification handles previously unseen classes by assigning them to the most appropriate high‑level category in a class taxonomy. We extend this paradigm to the semi‑supervised setting, enabling the use of large‑scale, uncurated datasets containing a mixture of known and unknown classes to improve the hierarchical open‑set performance. To this end, we propose a teacher‑student framework based on pseudo‑labeling. Two key components are introduced: 1) subtree pseudo‑labels, which provide reliable supervision in the presence of unknown data, and 2) age‑gating, a mechanism that mitigates overconfidence in pseudo‑labels. Experiments show that our framework outperforms self‑supervised pretraining followed by supervised adaptation, and even matches the fully supervised counterpart when using only 20 labeled samples per class on the iNaturalist19 benchmark. Our code is available at https://github.com/walline/semihoc.
Authors:Chen Long, Dian Chen, Ruifei Ding, Zhe Chen, Zhen Dong, Bisheng Yang
Abstract:
Accurate fine‑grained tree species classification is critical for forest inventory and biodiversity monitoring. Existing methods predominantly focus on designing complex architectures to fit local data distributions. However, they often overlook the long‑tailed distributions and high inter‑class similarity inherent in limited data, thereby struggling to distinguish between few‑shot or confusing categories. In the process of knowledge dissemination in the human world, individuals will actively seek expert assistance to transcend the limitations of local thinking. Inspired by this, we introduce an external "Domain Expert" and propose an Expert Knowledge‑Guided Classification Decision Calibration Network (EKDC‑Net) to overcome these challenges. Our framework addresses two core issues: expert knowledge extraction and utilization. Specifically, we first develop a Local Prior Guided Knowledge Extraction Module (LPKEM). By leveraging Class Activation Map (CAM) analysis, LPKEM guides the domain expert to focus exclusively on discriminative features essential for classification. Subsequently, to effectively integrate this knowledge, we design an Uncertainty‑Guided Decision Calibration Module (UDCM). This module dynamically corrects the local model's decisions by considering both overall category uncertainty and instance‑level prediction uncertainty. Furthermore, we present a large‑scale classification dataset covering 102 tree species, named CU‑Tree102 to address the issue of scarce diversity in current benchmarks. Experiments on three benchmark datasets demonstrate that our approach achieves state‑of‑the‑art performance. Crucially, as a lightweight plug‑and‑play module, EKDC‑Net improves backbone accuracy by 6.42% and precision by 11.46% using only 0.08M additional learnable parameters. The dataset, code, and pre‑trained models are available at https://github.com/WHU‑USI3DV/TreeCLS.
Authors:Peixian Liang, Songhao Li, Shunsuke Koga, Yutong Li, Zahra Alipour, Yucheng Tang, Daguang Xu, Zhi Huang
Abstract:
Accurate semantic segmentation for histopathology image is crucial for quantitative tissue analysis and downstream clinical modeling. Recent segmentation foundation models have improved generalization through large‑scale pretraining, yet remain poorly aligned with pathology because they treat segmentation as a static visual prediction task. Here we present VISTA‑PATH, an interactive, class‑aware pathology segmentation foundation model designed to resolve heterogeneous structures, incorporate expert feedback, and produce pixel‑level segmentation that are directly meaningful for clinical interpretation. VISTA‑PATH jointly conditions segmentation on visual context, semantic tissue descriptions, and optional expert‑provided spatial prompts, enabling precise multi‑class segmentation across heterogeneous pathology images. To support this paradigm, we curate VISTA‑PATH Data, a large‑scale pathology segmentation corpus comprising over 1.6 million image‑mask‑text triplets spanning 9 organs and 93 tissue classes. Across extensive held‑out and external benchmarks, VISTA‑PATH consistently outperforms existing segmentation foundation models. Importantly, VISTA‑PATH supports dynamic human‑in‑the‑loop refinement by propagating sparse, patch‑level bounding‑box annotation feedback into whole‑slide segmentation. Finally, we show that the high‑fidelity, class‑aware segmentation produced by VISTA‑PATH is a preferred model for computational pathology. It improve tissue microenvironment analysis through proposed Tumor Interaction Score (TIS), which exhibits strong and significant associations with patient survival. Together, these results establish VISTA‑PATH as a foundation model that elevates pathology image segmentation from a static prediction to an interactive and clinically grounded representation for digital pathology. Source code and demo can be found at https://github.com/zhihuanglab/VISTA‑PATH.
Authors:Jongmin Yu, Hyeontaek Oh, Zhongtian Sun, Angelica I Aviles-Rivero, Moongu Jeon, Jinhong Yang
Abstract:
Existing face‑swapping methods often deliver competitive results in constrained settings but exhibit substantial quality degradation when handling extreme facial poses. To improve facial pose robustness, explicit geometric features are applied, but this approach remains problematic since it introduces additional dependencies and increases computational cost. Diffusion‑based methods have achieved remarkable results; however, they are impractical for real‑time processing. We introduce AlphaFace, which leverages an open‑source vision‑language model and CLIP image and text embeddings to apply novel visual and textual semantic contrastive losses. AlphaFace enables stronger identity representation and more precise attribute preservation, all while maintaining real‑time performance. Comprehensive experiments across FF++, MPIE, and LPFF demonstrate that AlphaFace surpasses state‑of‑the‑art methods in pose‑challenging cases. The project is publicly available on `https://github.com/andrewyu90/Alphaface_Official.git'.
Authors:Shuying Li, Qiang Ma, San Zhang, Chuang Yang
Abstract:
Infrared small target detection (IRSTD) is critical for applications like remote sensing and surveillance, which aims to identify small, low‑contrast targets against complex backgrounds. However, existing methods often struggle with inadequate joint modeling of local‑global features (harming target‑background discrimination) or feature redundancy and semantic dilution (degrading target representation quality). To tackle these issues, we propose DCCS‑Det (Directional Context and Cross‑Scale Aware Detector for Infrared Small Target), a novel detector that incorporates a Dual‑stream Saliency Enhancement (DSE) block and a Latent‑aware Semantic Extraction and Aggregation (LaSEA) module. The DSE block integrates localized perception with direction‑aware context aggregation to help capture long‑range spatial dependencies and local details. On this basis, the LaSEA module mitigates feature degradation via cross‑scale feature extraction and random pooling sampling strategies, enhancing discriminative features and suppressing noise. Extensive experiments show that DCCS‑Det achieves state‑of‑the‑art detection accuracy with competitive efficiency across multiple datasets. Ablation studies further validate the contributions of DSE and LaSEA in improving target perception and feature representation under complex scenarios. \hrefhttps://huggingface.co/InPeerReview/InfraredSmallTargetDetection‑IRSTD.DCCSDCCS‑Det Official Code is Available Here!
Authors:Soumitri Chattopadhyay, Basar Demir, Marc Niethammer
Abstract:
While 3D foundational models have shown promise for promptable segmentation of medical volumes, their robustness to imprecise prompts remains under‑explored. In this work, we aim to address this gap by systematically studying the effect of various controlled perturbations of dense visual prompts, that closely mimic real‑world imprecision. By conducting experiments with two recent foundational models on a multi‑organ abdominal segmentation task, we reveal several facets of promptable medical segmentation, especially pertaining to reliance on visual shape and spatial cues, and the extent of resilience of models towards certain perturbations. Codes are available at: https://github.com/ucsdbiag/Prompt‑Robustness‑MedSegFMs
Authors:Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen, Jong Chul Ye, Duygu Ceylan, Hyeonho Jeong
Abstract:
Video‑to‑video diffusion models achieve impressive single‑turn editing performance, but practical editing workflows are inherently iterative. When edits are applied sequentially, existing models treat each turn independently, often causing previously generated regions to drift or be overwritten. We identify this failure mode as the problem of cross‑turn consistency in multi‑turn video editing. We introduce Memory‑V2V, a memory‑augmented framework that treats prior edits as structured constraints for subsequent generations. Memory‑V2V maintains an external memory of previous outputs, retrieves task‑relevant edits, and integrates them through relevance‑aware tokenization and adaptive compression. These technical ingredients enable scalable conditioning without linear growth in computation. We demonstrate Memory‑V2V on iterative video novel view synthesis and text‑guided long video editing. Memory‑V2V substantially enhances cross‑turn consistency while maintaining visual quality, outperforming strong baselines with modest overhead.
Authors:Wenhang Ge, Guibao Shen, Jiawei Feng, Luozhou Wang, Hao Lu, Xingye Tian, Xin Tao, Ying-Cong Chen
Abstract:
Recent advances in camera‑controlled video diffusion models have significantly improved video‑camera alignment. However, the camera controllability still remains limited. In this work, we build upon Reward Feedback Learning and aim to further improve camera controllability. However, directly borrowing existing ReFL approaches faces several challenges. First, current reward models lack the capacity to assess video‑camera alignment. Second, decoding latent into RGB videos for reward computation introduces substantial computational overhead. Third, 3D geometric information is typically neglected during video decoding. To address these limitations, we introduce an efficient camera‑aware 3D decoder that decodes video latent into 3D representations for reward quantization. Specifically, video latent along with the camera pose are decoded into 3D Gaussians. In this process, the camera pose not only acts as input, but also serves as a projection parameter. Misalignment between the video latent and camera pose will cause geometric distortions in the 3D structure, resulting in blurry renderings. Based on this property, we explicitly optimize pixel‑level consistency between the rendered novel views and ground‑truth ones as reward. To accommodate the stochastic nature, we further introduce a visibility term that selectively supervises only deterministic regions derived via geometric warping. Extensive experiments conducted on RealEstate10K and WorldScore benchmarks demonstrate the effectiveness of our proposed method. Project page: \hrefhttps://a‑bigbao.github.io/CamPilot/CamPilot Page.
Authors:Geo Ahn, Inwoong Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, Jinwoo Choi
Abstract:
Zero‑Shot Compositional Action Recognition (ZS‑CAR) requires recognizing novel verb‑object combinations composed of previously observed primitives. In this work, we tackle a key failure mode: models predict verbs via object‑driven shortcuts (i.e., relying on the labeled object class) rather than temporal evidence. We argue that sparse compositional supervision and verb‑object learning asymmetry can promote object‑driven shortcut learning. Our analysis with proposed diagnostic metrics shows that existing methods overfit to training co‑occurrence patterns and underuse temporal verb cues, resulting in weak generalization to unseen compositions. To address object‑driven shortcuts, we propose Robust COmpositional REpresentations (RCORE) with two components. Co‑occurrence Prior Regularization (CPR) adds explicit supervision for unseen compositions and regularizes the model against frequent co‑occurrence priors by treating them as hard negatives. Temporal Order Regularization for Composition (TORC) enforces temporal‑order sensitivity to learn temporally grounded verb representations. Across Sth‑com and EK100‑com, RCORE reduces shortcut diagnostics and consequently improves compositional generalization.
Authors:Shengbang Tong, Boyang Zheng, Ziteng Wang, Bingda Tang, Nanye Ma, Ellis Brown, Jihan Yang, Rob Fergus, Yann LeCun, Saining Xie
Abstract:
Representation Autoencoders (RAEs) have shown distinct advantages in diffusion modeling on ImageNet by training in high‑dimensional semantic latent spaces. In this work, we investigate whether this framework can scale to large‑scale, freeform text‑to‑image (T2I) generation. We first scale RAE decoders on the frozen representation encoder (SigLIP‑2) beyond ImageNet by training on web, synthetic, and text‑rendering data, finding that while scale improves general fidelity, targeted data composition is essential for specific domains like text. We then rigorously stress‑test the RAE design choices originally proposed for ImageNet. Our analysis reveals that scaling simplifies the framework: while dimension‑dependent noise scheduling remains critical, architectural complexities such as wide diffusion heads and noise‑augmented decoding offer negligible benefits at scale Building on this simplified framework, we conduct a controlled comparison of RAE against the state‑of‑the‑art FLUX VAE across diffusion transformer scales from 0.5B to 9.8B parameters. RAEs consistently outperform VAEs during pretraining across all model scales. Further, during finetuning on high‑quality datasets, VAE‑based models catastrophically overfit after 64 epochs, while RAE models remain stable through 256 epochs and achieve consistently better performance. Across all experiments, RAE‑based diffusion models demonstrate faster convergence and better generation quality, establishing RAEs as a simpler and stronger foundation than VAEs for large‑scale T2I generation. Additionally, because both visual understanding and generation can operate in a shared representation space, the multimodal model can directly reason over generated latents, opening new possibilities for unified models.
Authors:Ziyi Wu, Daniel Watson, Andrea Tagliasacchi, David J. Fleet, Marcus A. Brubaker, Saurabh Saxena
Abstract:
Lifting perspective images and videos to 360° panoramas enables immersive 3D world generation. Existing approaches often rely on explicit geometric alignment between the perspective and the equirectangular projection (ERP) space. Yet, this requires known camera metadata, obscuring the application to in‑the‑wild data where such calibration is typically absent or noisy. We propose 360Anything, a geometry‑free framework built upon pre‑trained diffusion transformers. By treating the perspective input and the panorama target simply as token sequences, 360Anything learns the perspective‑to‑equirectangular mapping in a purely data‑driven way, eliminating the need for camera information. Our approach achieves state‑of‑the‑art performance on both image and video perspective‑to‑360° generation, outperforming prior works that use ground‑truth camera information. We also trace the root cause of the seam artifacts at ERP boundaries to zero‑padding in the VAE encoder, and introduce Circular Latent Encoding to facilitate seamless generation. Finally, we show competitive results in zero‑shot camera FoV and orientation estimation benchmarks, demonstrating 360Anything's deep geometric understanding and broader utility in computer vision tasks. Additional results are available at https://360anything.github.io/.
Authors:Remy Sabathier, David Novotny, Niloy J. Mitra, Tom Monnier
Abstract:
Generating animated 3D objects is at the heart of many applications, yet most advanced works are typically difficult to apply in practice because of their limited setup, their long runtime, or their limited quality. We introduce ActionMesh, a generative model that predicts production‑ready 3D meshes "in action" in a feed‑forward manner. Drawing inspiration from early video models, our key insight is to modify existing 3D diffusion models to include a temporal axis, resulting in a framework we dubbed "temporal 3D diffusion". Specifically, we first adapt the 3D diffusion stage to generate a sequence of synchronized latents representing time‑varying and independent 3D shapes. Second, we design a temporal 3D autoencoder that translates a sequence of independent shapes into the corresponding deformations of a pre‑defined reference shape, allowing us to build an animation. Combining these two components, ActionMesh generates animated 3D meshes from different inputs like a monocular video, a text description, or even a 3D mesh with a text prompt describing its animation. Besides, compared to previous approaches, our method is fast and produces results that are rig‑free and topology consistent, hence enabling rapid iteration and seamless applications like texturing and retargeting. We evaluate our model on standard video‑to‑4D benchmarks (Consistent4D, Objaverse) and report state‑of‑the‑art performances on both geometric accuracy and temporal consistency, demonstrating that our model can deliver animated 3D meshes with unprecedented speed and quality.
Authors:Sylvestre-Alvise Rebuffi, Tuan Tran, Valeriu Lacatusu, Pierre Fernandez, Tomáš Souček, Nikola Jovanović, Tom Sander, Hady Elsahar, Alexandre Mourachko
Abstract:
Existing approaches for watermarking AI‑generated images often rely on post‑hoc methods applied in pixel space, introducing computational overhead and potential visual artifacts. In this work, we explore latent space watermarking and introduce DistSeal, a unified approach for latent watermarking that works across both diffusion and autoregressive models. Our approach works by training post‑hoc watermarking models in the latent space of generative models. We demonstrate that these latent watermarkers can be effectively distilled either into the generative model itself or into the latent decoder, enabling in‑model watermarking. The resulting latent watermarks achieve competitive robustness while offering similar imperceptibility and up to 20x speedup compared to pixel‑space baselines. Our experiments further reveal that distilling latent watermarkers outperforms distilling pixel‑space ones, providing a solution that is both more efficient and more robust.
Authors:Zhiyin Qian, Siwei Zhang, Bharat Lal Bhatnagar, Federica Bogo, Siyu Tang
Abstract:
Human motion reconstruction from monocular videos is a fundamental challenge in computer vision, with broad applications in AR/VR, robotics, and digital content creation, but remains challenging under frequent occlusions in real‑world settings. Existing regression‑based methods are efficient but fragile to missing observations, while optimization‑ and diffusion‑based approaches improve robustness at the cost of slow inference speed and heavy preprocessing steps. To address these limitations, we leverage recent advances in generative masked modeling and present MoRo: Masked Modeling for human motion Recovery under Occlusions. MoRo is an occlusion‑robust, end‑to‑end generative framework that formulates motion reconstruction as a video‑conditioned task, and efficiently recover human motion in a consistent global coordinate system from RGB videos. By masked modeling, MoRo naturally handles occlusions while enabling efficient, end‑to‑end inference. To overcome the scarcity of paired video‑motion data, we design a cross‑modality learning scheme that learns multi‑modal priors from a set of heterogeneous datasets: (i) a trajectory‑aware motion prior trained on MoCap datasets, (ii) an image‑conditioned pose prior trained on image‑pose datasets, capturing diverse per‑frame poses, and (iii) a video‑conditioned masked transformer that fuses motion and pose priors, finetuned on video‑motion datasets to integrate visual cues with motion dynamics for robust inference. Extensive experiments on EgoBody and RICH demonstrate that MoRo substantially outperforms state‑of‑the‑art methods in accuracy and motion realism under occlusions, while performing on‑par in non‑occluded scenarios. MoRo achieves real‑time inference at 70 FPS on a single H200 GPU.
Authors:Junha Lee, Eunha Park, Minsu Cho
Abstract:
Language‑driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand‑object interactions. While vision‑language models have been applied to this problem, existing approaches directly map observations to grasp parameters without intermediate reasoning about physical interactions. We present DextER, Dexterous Grasp Generation with Embodied Reasoning, which introduces contact‑based embodied reasoning for multi‑finger manipulation. Our key insight is that predicting which hand links contact where on the object surface provides an embodiment‑aware intermediate representation, bridging task semantics with physical constraints. DextER autoregressively generates embodied contact tokens specifying which finger links contact where on the object surface, followed by grasp tokens encoding the hand configuration. On DexGYS, DextER achieves 67.14% success rate, outperforming state‑of‑the‑art by 3.83 p.p. with 96.4% improvement in intention alignment. We also demonstrate steerable generation through partial contact specification, providing fine‑grained control over grasp synthesis.
Authors:Yifan Chen, Fei Yin, Hao Chen, Jia Wu, Chao Li
Abstract:
Contrast‑enhanced imaging is central to oncologic diagnosis, but contrast agents can be contraindicated for many of the patients who need them most. Synthesizing contrast scans from non‑contrast inputs is the natural response. Two obstacles stand in the way: no benchmark provides paired contrast data with lesion‑level evaluation, and no single model handles the arbitrary missing patterns seen in practice. We introduce Contrast‑X, a benchmark of paired contrast‑enhanced and non‑contrast imaging spanning 10 organs in CT (1,526 patients) and multi‑phase breast DCE‑MRI (1116 patients). Every case carries radiologist‑verified phase labels and tumor masks. We further propose FlowMI, a single model that handles arbitrary subsets of available modalities through a unified multi‑modal latent space and flow matching. We benchmark a range of missing‑modality configurations, reporting standard image‑quality metrics, radiologist reader studies, and downstream lesion analysis on the synthesized scans. We further evaluate cross‑organ generalization to test whether the model has learned a transferable contrast‑enhancement operation. Dataset, code, and leaderboard will be released. Our code are available at https://github.com/YifanChen02/Contrast‑X.
Authors:Yonghao Xu, Pedram Ghamisi, Qihao Weng
Abstract:
Recent years have witnessed the remarkable success of deep learning in remote sensing image interpretation, driven by the availability of large‑scale benchmark datasets. However, this reliance on massive training data also brings two major challenges: (1) high storage and computational costs, and (2) the risk of data leakage, especially when sensitive categories are involved. To address these challenges, this study introduces the concept of dataset distillation into the field of remote sensing image interpretation for the first time. Specifically, we train a text‑to‑image diffusion model to condense a large‑scale remote sensing dataset into a compact and representative distilled dataset. To improve the discriminative quality of the synthesized samples, we propose a classifier‑driven guidance by injecting a classification consistency loss from a pre‑trained model into the diffusion training process. Besides, considering the rich semantic complexity of remote sensing imagery, we further perform latent space clustering on training samples to select representative and diverse prototypes as visual style guidance, while using a visual language model to provide aggregated text descriptions. Experiments on three high‑resolution remote sensing scene classification benchmarks show that the proposed method can distill realistic and diverse samples for downstream model training. Code and pre‑trained models are available online (https://github.com/YonghaoXu/DPD).
Authors:Liuyun Jiang, Yanchao Zhang, Jinyue Guo, Yizhuo Lu, Ruining Zhou, Hua Han
Abstract:
Neuron segmentation in electron microscopy (EM) aims to reconstruct the complete neuronal connectome; however, current deep learning‑based methods are limited by their reliance on large‑scale training data and extensive, time‑consuming manual annotations. Traditional methods augment the training set through geometric and photometric transformations; however, the generated samples remain highly correlated with the original images and lack structural diversity. To address this limitation, we propose a diffusion‑based data augmentation framework capable of generating diverse and structurally plausible image‑label pairs for neuron segmentation. Specifically, the framework employs a resolution‑aware conditional diffusion model with multi‑scale conditioning and EM resolution priors to enable voxel‑level image synthesis from 3D masks. It further incorporates a biology‑guided mask remodeling module that produces augmented masks with enhanced structural realism. Together, these components effectively enrich the training set and improve segmentation performance. On the AC3 and AC4 datasets under low‑annotation regimes, our method improves the ARAND metric by 32.1% and 30.7%, respectively, when combined with two different post‑processing methods. Our code is available at https://github.com/HeadLiuYun/NeuroDiff.
Authors:Yikui Zhai, Shikuang Liu, Wenlve Zhou, Hongsheng Zhang, Zhiheng Zhou, Xiaolin Tian, C. L. Philip Chen
Abstract:
Few‑shot recognition in synthetic aperture radar (SAR) imagery remains a critical bottleneck for real‑world applications due to extreme data scarcity. A promising strategy involves synthesizing a large dataset with a generative adversarial network (GAN), pre‑training a model via self‑supervised learning (SSL), and then fine‑tuning on the few labeled samples. However, this approach faces a fundamental paradox: conventional GANs themselves require abundant data for stable training, contradicting the premise of few‑shot learning. To resolve this, we propose the consistency‑regularized generative adversarial network (Cr‑GAN), a novel framework designed to synthesize diverse, high‑fidelity samples even when trained under these severe data limitations. Cr‑GAN introduces a dual‑branch discriminator that decouples adversarial training from representation learning. This architecture enables a channel‑wise feature interpolation strategy to create novel latent features, complemented by a dual‑domain cycle consistency mechanism that ensures semantic integrity. Our Cr‑GAN framework is adaptable to various GAN architectures, and its synthesized data effectively boosts multiple SSL algorithms. Extensive experiments on the MSTAR and SRSDD datasets validate our approach, with Cr‑GAN achieving a highly competitive accuracy of 71.21% and 51.64%, respectively, in the 8‑shot setting, significantly outperforming leading baselines, while requiring only ~5 of the parameters of state‑of‑the‑art diffusion models. Code is available at: https://github.com/yikuizhai/Cr‑GAN.
Authors:Zichen Yu, Quanli Liu, Wei Wang, Liyong Zhang, Xiaoguang Zhao
Abstract:
3D occupancy prediction plays a pivotal role in the realm of autonomous driving, as it provides a comprehensive understanding of the driving environment. Most existing methods construct dense scene representations for occupancy prediction, overlooking the inherent sparsity of real‑world driving scenes. Recently, 3D superquadric representation has emerged as a promising sparse alternative to dense scene representations due to the strong geometric expressiveness of superquadrics. However, existing superquadric frameworks still suffer from insufficient temporal modeling, a challenging trade‑off between query sparsity and geometric expressiveness, and inefficient superquadric‑to‑voxel splatting. To address these issues, we propose SuperOcc, a novel framework for superquadric‑based 3D occupancy prediction. SuperOcc incorporates three key designs: (1) a cohesive temporal modeling mechanism to simultaneously exploit view‑centric and object‑centric temporal cues; (2) a multi‑superquadric decoding strategy to enhance geometric expressiveness without sacrificing query sparsity; and (3) an efficient superquadric‑to‑voxel splatting scheme to improve computational efficiency. Extensive experiments on the SurroundOcc and Occ3D benchmarks demonstrate that SuperOcc achieves state‑of‑the‑art performance while maintaining superior efficiency. The code is available at https://github.com/Yzichen/SuperOcc.
Authors:Ning Jiang, Dingheng Zeng, Yanhong Liu, Haiyang Yi, Shijie Yu, Minghe Weng, Haifeng Shen, Ying Li
Abstract:
Most prior deepfake detection methods lack explainable outputs. With the growing interest in multimodal large language models (MLLMs), researchers have started exploring their use in interpretable deepfake detection. However, a major obstacle in applying MLLMs to this task is the scarcity of high‑quality datasets with detailed forgery attribution annotations, as textual annotation is both costly and challenging ‑ particularly for high‑fidelity forged images or videos. Moreover, multiple studies have shown that reinforcement learning (RL) can substantially enhance performance in visual tasks, especially in improving cross‑domain generalization. To facilitate the adoption of mainstream MLLM frameworks in deepfake detection with reduced annotation cost, and to investigate the potential of RL in this context, we propose an automated Chain‑of‑Thought (CoT) data generation framework based on Self‑Blended Images, along with an RL‑enhanced deepfake detection framework. Extensive experiments validate the effectiveness of our CoT data construction pipeline, tailored reward mechanism, and feedback‑driven synthetic data generation approach. Our method achieves performance competitive with state‑of‑the‑art (SOTA) approaches across multiple cross‑dataset benchmarks. Implementation details are available at https://github.com/deon1219/rlsbi.
Authors:Weiwei Wu, Yueyang Li, Yuhu Shi, Weiming Zeng, Lang Qin, Yang Yang, Ke Zhou, Zhiguo Zhang, Wai Ting Siok, Nizhuan Wang
Abstract:
Cross‑subject EEG‑based emotion recognition (EER) remains challenging due to strong inter‑subject variability, which induces substantial distribution shifts in EEG signals, as well as the high complexity of emotion‑related neural representations in both spatial organization and temporal evolution. Existing approaches typically improve spatial modeling, temporal modeling, or generalization strategies in isolation, which limits their ability to align representations across subjects while capturing multi‑scale dynamics and suppressing subject‑specific bias within a unified framework. To address these gaps, we propose a Region‑aware Spatiotemporal Modeling framework with Collaborative Domain Generalization (RSM‑CoDG) for cross‑subject EEG emotion recognition. RSM‑CoDG incorporates neuroscience priors derived from functional brain region partitioning to construct region‑level spatial representations, thereby improving cross‑subject comparability. It also employs multi‑scale temporal modeling to characterize the dynamic evolution of emotion‑evoked neural activity. In addition, the framework employs a collaborative domain generalization strategy, incorporating multidimensional constraints to reduce subject‑specific bias in a fully unseen target subject setting, which enhances the generalization to unknown individuals. Extensive experimental results on SEED series datasets demonstrate that RSM‑CoDG consistently outperforms existing competing methods, providing an effective approach for improving robustness. The source code is available at https://github.com/RyanLi‑X/RSM‑CoDG.
Authors:William Huang, Siyou Pei, Leyi Zou, Eric J. Gonzalez, Ishan Chatterjee, Yang Zhang
Abstract:
The proliferation of XR devices has made egocentric hand pose estimation a vital task, yet this perspective is inherently challenged by frequent finger occlusions. To address this, we propose a novel approach that leverages the rich information in dorsal hand skin deformation, unlocked by recent advances in dense visual featurizers. We introduce a dual‑stream delta encoder that learns pose by contrasting features from a dynamic hand with a baseline relaxed position. Our evaluation demonstrates that, using only cropped dorsal images, our method reduces the Mean Per Joint Angle Error (MPJAE) by 18% in self‑occluded scenarios (fingers >= 50% occluded) compared to state‑of‑the‑art techniques that depend on the whole hand's geometry and large model backbones. Consequently, our method not only enhances the reliability of downstream tasks like index finger pinch and tap estimation in occluded scenarios but also unlocks new interaction paradigms, such as detecting isometric force for a surface "click" without visible movement while minimizing model size.
Authors:Yunshan Qi, Lin Zhu, Nan Bao, Yifan Zhao, Jia Li
Abstract:
Novel view synthesis from low dynamic range (LDR) blurry images, which are common in the wild, struggles to recover high dynamic range (HDR) and sharp 3D representations in extreme lighting conditions. Although existing methods employ event data to address this issue, they ignore the sensor‑physics mismatches between the camera output and physical world radiance, resulting in suboptimal HDR and deblurring results. To cope with this problem, we propose a unified sensor‑physics grounded NeRF framework for sharp HDR novel view synthesis from single‑exposure blurry LDR images and corresponding events. We employ NeRF to directly represent the actual radiance of the 3D scene in the HDR domain and model raw HDR scene rays hitting the sensor pixels as in the physical world. A 2D pixel‑wise RGB CRF model is introduced to align the NeRF rendered pixel values with the sensor‑recorded LDR pixel values of the input images. A novel event CRF model is also designed to bridge the gap between physical scene dynamics and event sensor output. The two models are jointly optimized with the NeRF network, leveraging the spatial and temporal dynamic information in events to enhance the sharp HDR 3D representation learning. Experiments on the collected and public datasets demonstrate that our method achieves state‑of‑the‑art HDR and deblurring novel view synthesis results with single‑exposure blurry LDR images and corresponding events.
Authors:Yinghan Xu, Théo Morales, John Dingliana
Abstract:
Radiance field‑based rendering methods have attracted significant interest from the computer vision and computer graphics communities. They enable high‑fidelity rendering with complex real‑world lighting effects, but at the cost of high rendering time. 3D Gaussian Splatting solves this issue with a rasterisation‑based approach for real‑time rendering, enabling applications such as autonomous driving, robotics, virtual reality, and extended reality. However, current 3DGS implementations are difficult to integrate into traditional mesh‑based rendering pipelines, which is a common use case for interactive applications and artistic exploration. To address this limitation, this software solution uses Nvidia's interprocess communication (IPC) APIs to easily integrate into implementations and allow the results to be viewed in external clients such as Unity, Blender, Unreal Engine, and OpenGL viewers. The code is available at https://github.com/RockyXu66/splatbus.
Authors:Pablo Messina, Andrés Villa, Juan León Alcázar, Karen Sánchez, Carlos Hinojosa, Denis Parra, Álvaro Soto, Bernard Ghanem
Abstract:
Medical vision‑language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present CURE, an error‑aware curriculum learning framework that improves grounding and report quality without any additional data. CURE fine‑tunes a multimodal instructional model on phrase grounding, grounded report generation, and anatomy‑grounded report generation using public datasets. The method dynamically adjusts sampling based on model performance, emphasizing harder samples to improve spatial and textual alignment. CURE improves grounding accuracy by +0.37 IoU, boosts report quality by +0.188 CXRFEScore, and reduces hallucinations by 18.6%. CURE is a data‑efficient framework that enhances both grounding accuracy and report reliability. Code is available at https://github.com/PabloMessina/CURE and model weights at https://huggingface.co/pamessina/medgemma‑4b‑it‑cure
Authors:Francesca Pia Panaccione, Carlo Sgaravatti, Pietro Pinoli
Abstract:
Biomedical research increasingly relies on integrating diverse data modalities, including gene expression profiles, medical images, and clinical metadata. While medical images and clinical metadata are routinely collected in clinical practice, gene expression data presents unique challenges for widespread research use, mainly due to stringent privacy regulations and costly laboratory experiments. To address these limitations, we present GeMM‑GAN, a novel Generative Adversarial Network conditioned on histopathology tissue slides and clinical metadata, designed to synthesize realistic gene expression profiles. GeMM‑GAN combines a Transformer Encoder for image patches with a final Cross Attention mechanism between patches and text tokens, producing a conditioning vector to guide a generative model in generating biologically coherent gene expression profiles. We evaluate our approach on the TCGA dataset and demonstrate that our framework outperforms standard generative models and generates more realistic and functionally meaningful gene expression profiles, improving by more than 11% the accuracy on downstream disease type prediction compared to current state‑of‑the‑art generative models. Code will be available at: https://github.com/francescapia/GeMM‑GAN
Authors:Wei Ai, Yilong Tan, Yuntao Shou, Tao Meng, Haowen Chen, Zhixiong He, Keqin Li
Abstract:
In recent years, the rapid evolution of large vision‑language models (LVLMs) has driven a paradigm shift in multimodal fake news detection (MFND), transforming it from traditional feature‑engineering approaches to unified, end‑to‑end multimodal reasoning frameworks. Early methods primarily relied on shallow fusion techniques to capture correlations between text and images, but they struggled with high‑level semantic understanding and complex cross‑modal interactions. The emergence of LVLMs has fundamentally changed this landscape by enabling joint modeling of vision and language with powerful representation learning, thereby enhancing the ability to detect misinformation that leverages both textual narratives and visual content. Despite these advances, the field lacks a systematic survey that traces this transition and consolidates recent developments. To address this gap, this paper provides a comprehensive review of MFND through the lens of LVLMs. We first present a historical perspective, mapping the evolution from conventional multimodal detection pipelines to foundation model‑driven paradigms. Next, we establish a structured taxonomy covering model architectures, datasets, and performance benchmarks. Furthermore, we analyze the remaining technical challenges, including interpretability, temporal reasoning, and domain generalization. Finally, we outline future research directions to guide the next stage of this paradigm shift. To the best of our knowledge, this is the first comprehensive survey to systematically document and analyze the transformative role of LVLMs in combating multimodal fake news. The summary of existing methods mentioned is in our Github: \hrefhttps://github.com/Tan‑YiLong/Overview‑of‑Fake‑News‑Detectionhttps://github.com/Tan‑YiLong/Overview‑of‑Fake‑News‑Detection.
Authors:Jiwon Kang, Yeji Choi, JoungBin Lee, Wooseok Jang, Jinhyeok Choi, Taekeun Kang, Yongjae Park, Myungin Kim, Seungryong Kim
Abstract:
Face swapping aims to transfer the identity of a source face onto a target face while preserving target‑specific attributes such as pose, expression, lighting, skin tone, and makeup. However, since real ground truth for face swapping is unavailable, achieving both accurate identity transfer and high‑quality attribute preservation remains challenging. Recent diffusion‑based approaches attempt to improve visual fidelity through conditional inpainting on masked target images, but the masked condition removes crucial appearance cues, resulting in plausible yet misaligned attributes. To address this limitation, we propose APPLE (Attribute‑Preserving Pseudo‑Labeling), a fully diffusion‑based teacher‑student framework for attribute‑preserving face swapping. Our approach introduces a teacher design to produce pseudo‑labels aligned with the target attributes through (1) a conditional deblurring formulation that improves the preservation of global attributes such as skin tone and illumination, and (2) an attribute‑aware inversion scheme that further enhances fine‑grained attribute preservation such as makeup. APPLE conditions the student on clean pseudo‑labels rather than degraded masked inputs, enabling more faithful attribute preservation. As a result, APPLE achieves state‑of‑the‑art performance in attribute preservation while maintaining competitive identity transferability.
Authors:Gautom Das, Vincent La, Ethan Lau, Abhinav Shrivastava, Matthew Gwilliam
Abstract:
Large language models (LLMs) deliver impressive results for a variety of tasks, but state‑of‑the‑art systems require fast GPUs with large amounts of memory. To reduce both the memory and latency of these systems, practitioners quantize their learned parameters, typically at half precision. A growing body of research focuses on preserving the model performance with more aggressive bit widths, and some work has been done to apply these strategies to other models, like vision transformers. In our study we investigate how a variety of quantization methods, including state‑of‑the‑art GPTQ and AWQ, can be applied effectively to multimodal pipelines comprised of vision models, language models, and their connectors. We address how performance on captioning, retrieval, and question answering can be affected by bit width, quantization method, and which portion of the pipeline the quantization is used for. Results reveal that ViT and LLM exhibit comparable importance in model performance, despite significant differences in parameter size, and that lower‑bit quantization of the LLM achieves high accuracy at reduced bits per weight (bpw). These findings provide practical insights for efficient deployment of MLLMs and highlight the value of exploration for understanding component sensitivities in multimodal models. Our code is available at https://github.com/gautomdas/mmq.
Authors:Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, Daquan Zhou
Abstract:
Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing high‑quality videos that accurately reflect real‑world robotic interactions remains challenging, and the lack of a standardized benchmark limits fair comparisons and progress. To address this gap, we introduce a comprehensive robotics benchmark, RBench, designed to evaluate robot‑oriented video generation across five task domains and four distinct embodiments. It assesses both task‑level correctness and visual fidelity through reproducible sub‑metrics, including structural consistency, physical plausibility, and action completeness. Evaluation of 25 representative models highlights significant deficiencies in generating physically realistic robot behaviors. Furthermore, the benchmark achieves a Spearman correlation coefficient of 0.96 with human evaluations, validating its effectiveness. While RBench provides the necessary lens to identify these deficiencies, achieving physical realism requires moving beyond evaluation to address the critical shortage of high‑quality training data. Driven by these insights, we introduce a refined four‑stage data pipeline, resulting in RoVid‑X, the largest open‑source robotic dataset for video generation with 4 million annotated video clips, covering thousands of tasks and enriched with comprehensive physical property annotations. Collectively, this synergistic ecosystem of evaluation and data establishes a robust foundation for rigorous assessment and scalable training of video models, accelerating the evolution of embodied AI toward general intelligence.
Authors:Dominik Rößle, Xujun Xie, Adithya Mohan, Venkatesh Thirugnana Sambandham, Daniel Cremers, Torsten Schön
Abstract:
Perception is a cornerstone of autonomous driving, enabling vehicles to understand their surroundings and make safe, reliable decisions. Developing robust perception algorithms requires large‑scale, high‑quality datasets that cover diverse driving conditions and support thorough evaluation. Existing datasets often lack a high‑fidelity digital twin, limiting systematic testing, edge‑case simulation, sensor modification, and sim‑to‑real evaluations. To address this gap, we present DrivIng, a large‑scale multimodal dataset with a complete geo‑referenced digital twin of a ~18 km route spanning urban, suburban, and highway segments. Our dataset provides continuous recordings from six RGB cameras, one LiDAR, and high‑precision ADMA‑based localization, captured across day, dusk, and night. All sequences are annotated at 10 Hz with 3D bounding boxes and track IDs across 12 classes, yielding ~1.2 million annotated instances. Alongside the benefits of a digital twin, DrivIng enables a 1‑to‑1 transfer of real traffic into simulation, preserving agent interactions while enabling realistic and flexible scenario testing. To support reproducible research and robust validation, we benchmark DrivIng with state‑of‑the‑art perception models and publicly release the dataset, digital twin, HD map, and codebase.
Authors:Jianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran Lu, Dingcheng Wang, Letian Xue, Han Liu
Abstract:
Estimating task progress requires reasoning over long‑horizon dynamics rather than recognizing static visual content. While modern Vision‑Language Models (VLMs) excel at describing what is visible, it remains unclear whether they can infer how far a task has progressed from partial observations. To this end, we introduce Progress‑Bench, a benchmark for systematically evaluating progress reasoning in VLMs. Beyond benchmarking, we further explore a human‑inspired two‑stage progress reasoning paradigm through both training‑free prompting and training‑based approach based on curated dataset ProgressLM‑45K. Experiments on 14 VLMs show that most models are not yet ready for task progress estimation, exhibiting sensitivity to demonstration modality and viewpoint changes, as well as poor handling of unanswerable cases. While training‑free prompting that enforces structured progress reasoning yields limited and model‑dependent gains, the training‑based ProgressLM‑3B achieves consistent improvements even at a small model scale, despite being trained on a task set fully disjoint from the evaluation tasks. Further analyses reveal characteristic error patterns and clarify when and why progress reasoning succeeds or fails. Website: https://progresslm.github.io/ProgressLM/
Authors:Miroslav Purkrabek, Constantin Kolomiiets, Jiri Matas
Abstract:
Most 2D human pose estimation benchmarks are nearly saturated, with the exception of crowded scenes. We introduce PMPose, a top‑down 2D pose estimator that incorporates the probabilistic formulation and the mask‑conditioning. PMPose improves crowded pose estimation without sacrificing performance on standard scenes. Building on this, we present BBoxMaskPose v2 (BMPv2) integrating PMPose and an enhanced SAM‑based mask refinement module. BMPv2 surpasses state‑of‑the‑art by 1.5 average precision (AP) points on COCO and 6 AP points on OCHuman, becoming the first method to exceed 50 AP on OCHuman. We demonstrate that BMP's 2D prompting of 3D model improves 3D pose estimation in crowded scenes and that advances in 2D pose quality directly benefit 3D estimation. Results on the new OCHuman‑Pose dataset show that multi‑person performance is more affected by pose prediction accuracy than by detection. The code, models, and data are available on https://MiraPurkrabek.github.io/BBox‑Mask‑Pose/.
Authors:Bostan Khan, Masoud Daneshtalab
Abstract:
Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet existing Federated Neural Architecture Search (FedNAS) methods suffer from unguided supernet training and prohibitively costly post‑training search pipelines that demand over 20 GPU‑hours per deployment target. We introduce DeepFedNAS, a two‑phase framework built on a multi‑objective fitness function that synthesizes information‑theoretic network metrics with architectural heuristics. In the first phase, Federated Pareto Optimal Supernet Training replaces random subnet sampling with a pre‑computed cache of elite, high‑fitness architectures, yielding a superior supernet. In the second phase, a Predictor‑Free Search uses this fitness function as a zero‑cost accuracy proxy, discovering hardware‑optimized subnets in ~20 seconds, a ~61x speedup over the baseline pipeline. Experiments on CIFAR‑10, CIFAR‑100, and CINIC‑10 demonstrate state‑of‑the‑art accuracy (up to +1.21% on CIFAR‑100), a 2.8x reduction in per‑round transmission size, and robust performance under extreme non‑IID conditions (α = 0.1), making DeepFedNAS practical for scalable, communication‑constrained IoT federations. Source code: https://github.com/bostankhan6/DeepFedNAS
Authors:Andrey Moskalenko, Danil Kuznetsov, Irina Dudko, Anastasiia Iasakova, Nikita Boldyrev, Denis Shepelev, Andrei Spiridonov, Andrey Kuznetsov, Vlad Shakhuro
Abstract:
Promptable segmentation models such as SAM have established a powerful paradigm, enabling strong generalization to unseen objects and domains with minimal user input, including points, bounding boxes, and text prompts. Among these, bounding boxes stand out as particularly effective, often outperforming points while significantly reducing annotation costs. However, current training and evaluation protocols typically rely on synthetic prompts generated through simple heuristics, offering limited insight into real‑world robustness. In this paper, we investigate the robustness of promptable segmentation models to natural variations in bounding box prompts. First, we conduct a controlled user study and collect thousands of real bounding box annotations. Our analysis reveals substantial variability in segmentation quality across users for the same model and instance, indicating that SAM‑like models are highly sensitive to natural prompt noise. Then, since exhaustive testing of all possible user inputs is computationally prohibitive, we reformulate robustness evaluation as a white‑box optimization problem over the bounding box prompt space. We introduce BREPS, a method for generating adversarial bounding boxes that minimize or maximize segmentation error while adhering to naturalness constraints. Finally, we benchmark state‑of‑the‑art models across 10 datasets, spanning everyday scenes to medical imaging. Code ‑ https://github.com/emb‑ai/BREPS.
Authors:Shuonan Yang, Yuchen Zhang, Zeyu Fu
Abstract:
Hateful videos pose serious risks by amplifying discrimination, inciting violence, and undermining online safety. Existing training‑based hateful video detection methods are constrained by limited training data and lack of interpretability, while directly prompting large vision‑language models often struggle to deliver reliable hate detection. To address these challenges, this paper introduces MARS, a training‑free Multi‑stage Adversarial ReaSoning framework that enables reliable and interpretable hateful content detection. MARS begins with the objective description of video content, establishing a neutral foundation for subsequent analysis. Building on this, it develops evidence‑based reasoning that supports potential hateful interpretations, while in parallel incorporating counter‑evidence reasoning to capture plausible non‑hateful perspectives. Finally, these perspectives are synthesized into a conclusive and explainable decision. Extensive evaluation on two real‑world datasets shows that MARS achieves up to 10% improvement under certain backbones and settings compared to other training‑free approaches and outperforms state‑of‑the‑art training‑based methods on one dataset. In addition, MARS produces human‑understandable justifications, thereby supporting compliance oversight and enhancing the transparency of content moderation workflows. The code is available at https://github.com/Multimodal‑Intelligence‑Lab‑MIL/MARS.
Authors:Tianyu Li, Zongqian Wu, Songyue Cai, Ping Hu, Xiaofeng Zhu
Abstract:
CLIP‑based foreground‑background (FG‑BG) decomposition methods have demonstrated remarkable effectiveness in improving few‑shot out‑of‑distribution (OOD) detection performance. However, existing approaches still suffer from several limitations. For background regions obtained from decomposition, existing methods adopt a uniform suppression strategy for all patches, overlooking the varying contributions of different patches to the prediction. For foreground regions, existing methods fail to adequately consider that some local patches may exhibit appearance or semantic similarity to other classes, which may mislead the training process. To address these issues, we propose a new plug‑and‑play framework. This framework consists of three core components: (1) a Foreground‑Background Decomposition module, which follows previous FG‑BG methods to separate an image into foreground and background regions; (2) an Adaptive Background Suppression module, which adaptively weights patch classification entropy; and (3) a Confusable Foreground Rectification module, which identifies and rectifies confusable foreground patches. Extensive experimental results demonstrate that the proposed plug‑and‑play framework significantly improves the performance of existing FG‑BG decomposition methods. Code is available at: https://github.com/lounwb/FoBoR.
Authors:Adam Rokah, Daniel Veress, Caleb Caulk, Sourav Sharan
Abstract:
Mixture‑of‑Experts (MoE) architectures enable conditional computation by routing inputs to multiple expert subnetworks and are often motivated as a mechanism for scaling large language models. In this project, we instead study MoE behavior in an image classification setting, focusing on predictive performance, expert utilization, and generalization. We compare dense, SoftMoE, and SparseMoE classifier heads on the CIFAR10 dataset under comparable model capacity. Both MoE variants achieve slightly higher validation accuracy than the dense baseline while maintaining balanced expert utilization through regularization, avoiding expert collapse. To analyze generalization, we compute Hessian‑based sharpness metrics at convergence, including the largest eigenvalue and trace of the loss Hessian, evaluated on both training and test data. We find that SoftMoE exhibits higher sharpness by these metrics, while Dense and SparseMoE lie in a similar curvature regime, despite all models achieving comparable generalization performance. Complementary loss surface perturbation analyses reveal qualitative differences in non‑local behavior under finite parameter perturbations between dense and MoE models, which help contextualize curvature‑based measurements without directly explaining validation accuracy. We further evaluate empirical inference efficiency and show that naively implemented conditional routing does not yield inference speedups on modern hardware at this scale, highlighting the gap between theoretical and realized efficiency in sparse MoE models.
Authors:Yanan Wang, Linjie Ren, Zihao Li, Junyi Wang, Tian Gan
Abstract:
While video‑to‑audio generation has achieved remarkable progress in semantic and temporal alignment, most existing studies focus solely on these aspects, paying limited attention to the spatial perception and immersive quality of the synthesized audio. This limitation stems largely from current models' reliance on mono audio datasets, which lack the binaural spatial information needed to learn visual‑to‑spatial audio mappings. To address this gap, we introduce two key contributions: we construct BinauralVGGSound, the first large‑scale video‑binaural audio dataset designed to support spatially aware video‑to‑audio generation; and we propose a end‑to‑end spatial audio generation framework guided by visual cues, which explicitly models spatial features. Our framework incorporates a visual‑guided audio spatialization module that ensures the generated audio exhibits realistic spatial attributes and layered spatial depth while maintaining semantic and temporal alignment. Experiments show that our approach substantially outperforms state‑of‑the‑art models in spatial fidelity and delivers a more immersive auditory experience, without sacrificing temporal or semantic consistency. The demo page can be accessed at https://github.com/renlinjie868‑web/SpatialV2A.
Authors:Xinyu Peng, Han Li, Yuyang Huang, Ziyang Zheng, Yaoming Wang, Xin Chen, Wenrui Dai, Chenglin Li, Junni Zou, Hongkai Xiong
Abstract:
Existing video frame interpolation (VFI) methods often adopt a frame‑centric approach, processing videos as independent short segments (e.g., triplets), which leads to temporal inconsistencies and motion artifacts. To overcome this, we propose a holistic, video‑centric paradigm named Local Diffusion Forcing for Video Frame Interpolation (LDF‑VFI). Our framework is built upon an auto‑regressive diffusion transformer that models the entire video sequence to ensure long‑range temporal coherence. To mitigate error accumulation inherent in auto‑regressive generation, we introduce a novel skip‑concatenate sampling strategy that effectively maintains temporal stability. Furthermore, LDF‑VFI incorporates sparse, local attention and tiled VAE encoding, a combination that not only enables efficient processing of long sequences but also allows generalization to arbitrary spatial resolutions (e.g., 4K) at inference without retraining. An enhanced conditional VAE decoder, which leverages multi‑scale features from the input video, further improves reconstruction fidelity. Empirically, LDF‑VFI achieves state‑of‑the‑art performance on challenging VFI benchmarks, demonstrating superior per‑frame quality and temporal consistency, especially in scenes with large motion. The source code is available at https://github.com/xypeng9903/LDF‑VFI.
Authors:Donnate Hooft, Stefan M. Fischer, Cosmin Bercea, Jan C. Peeken, Julia A. Schnabel
Abstract:
Patch‑based methods are widely used in 3D medical image segmentation to address memory constraints in processing high‑resolution volumetric data. However, these approaches often neglect the patch's location within the global volume, which can limit segmentation performance when anatomical context is important. In this paper, we investigate the role of location context in patch‑based 3D segmentation and propose a novel attention mechanism, LocBAM, that explicitly processes spatial information. Experiments on BTCV, AMOS22, and KiTS23 demonstrate that incorporating location context stabilizes training and improves segmentation performance, particularly under low patch‑to‑volume coverage where global context is missing. Furthermore, LocBAM consistently outperforms classical coordinate encoding via CoordConv. Code is publicly available at https://github.com/compai‑lab/2026‑ISBI‑hooft
Authors:Jing Lan, Hexiao Ding, Hongzhao Chen, Yufeng Jiang, Nga-Chun Ng, Gwing Kei Yip, Gerald W. Y. Cheng, Yunlin Mao, Jing Cai, Liang-ting Lin, Jung Sun Yoo
Abstract:
AI models for drug discovery and chemical literature mining must interpret molecular images and generate outputs consistent with 3D geometry and stereochemistry. Most molecular language models rely on strings or graphs, while vision‑language models often miss stereochemical details and struggle to map continuous 3D structures into discrete tokens. We propose DeepMoLM: Deep Molecular Language M odeling, a dual‑view framework that grounds high‑resolution molecular images in geometric invariants derived from molecular conformations. DeepMoLM preserves high‑frequency evidence from 1024 × 1024 inputs, encodes conformer neighborhoods as discrete Extended 3‑Dimensional Fingerprints, and fuses visual and geometric streams with cross‑attention, enabling physically grounded generation without atom coordinates. DeepMoLM improves PubChem captioning with a 12.3% relative METEOR gain over the strongest generalist baseline while staying competitive with specialist methods. It produces valid numeric outputs for all property queries and attains MAE 13.64 g/mol on Molecular Weight and 37.89 on Complexity in the specialist setting. On ChEBI‑20 description generation from images, it exceeds generalist baselines and matches state‑of‑the‑art vision‑language models. Code is available at https://github.com/1anj/DeepMoLM.
Authors:Gensmo. ai, Chao Gao, Siqiao Xue, Jiwen Fu, Tingyi Gu, Shanshan Li, Fan Zhou
Abstract:
In this paper, we present LookBench (We use the term "look" to reflect retrieval that mirrors how people shop ‑‑ finding the exact item, a close substitute, or a visually consistent alternative.), a live, holistic and challenging benchmark for fashion image retrieval in real e‑commerce settings. LookBench includes both recent product images sourced from live websites and AI‑generated fashion images, reflecting contemporary trends and use cases. Each test sample is time‑stamped and we intend to update the benchmark periodically, enabling contamination‑aware evaluation aligned with declared training cutoffs. Grounded in our fine‑grained attribute taxonomy, LookBench covers single‑item and outfit‑level retrieval across. Our experiments reveal that LookBench poses a significant challenge on strong baselines, with many models achieving below 60% Recall@1. Our proprietary model achieves the best performance on LookBench, and we release an open‑source counterpart that ranks second, with both models attaining state‑of‑the‑art results on legacy Fashion200K evaluations. LookBench is designed to be updated semi‑annually with new test samples and progressively harder task variants, providing a durable measure of progress. We publicly release our leaderboard, dataset, evaluation code, and trained models.
Authors:Yian Huang, Qing Qin, Aji Mao, Xiangyu Qiu, Liang Xu, Xian Zhang, Zhenming Peng
Abstract:
Infrared small target detection (ISTD) has been a critical technology in defense and civilian applications over the past several decades, such as missile warning, maritime surveillance, and disaster monitoring. Nevertheless, moving infrared small target detection still faces considerable challenges: existing models suffer from insufficient spatio‑temporal semantic correlation and are not lightweight‑friendly, while algorithms with strong scene generalization capability are in great demand for real‑world applications. To address these issues, we propose FeedbackSTS‑Det, a sparse frames‑based spatio‑temporal semantic feedback network. Our approach introduces a closed‑loop spatio‑temporal semantic feedback strategy with paired forward and backward refinement modules that work cooperatively across the encoder and decoder to enhance information exchange between consecutive frames, effectively improving detection accuracy and reducing false alarms. Moreover, we introduce an embedded sparse semantic module (SSM), which operates by strategically grouping frames by interval, propagating semantics within each group, and reassembling the sequence to efficiently capture long‑range temporal dependencies with low computational overhead. Extensive experiments on many widely adopted multi‑frame infrared small target datasets demonstrate the generalization ability and scene adaptability of our proposed network. Code and models are available at: https://github.com/IDIP‑Lab/FeedbackSTS‑Det.
Authors:Mingyang Xie, Numair Khan, Tianfu Wang, Naina Dhingra, Seonghyeon Nam, Haitao Yang, Zhuo Hui, Christopher Metzler, Andrea Vedaldi, Hamed Pirsiavash, Lei Luo
Abstract:
Given a monocular video, the goal of video re‑rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awareness, leading to drift and deformation under viewpoint changes. On the other hand, geometrically‑conditioned models depend on estimated depth and explicit reconstruction, making them susceptible to depth inaccuracies and calibration errors.
We propose to address these challenges by using the implicit geometric knowledge embedded in the latent space of a large 4D reconstruction model to condition the video generation process. These latents capture scene structure in a continuous space without explicit reconstruction. Therefore, they provide a flexible representation that allows the pretrained diffusion prior to regularize errors more effectively. By jointly conditioning on these latents and source camera poses, we demonstrate that our model achieves state‑of‑the‑art results on the video re‑rendering task. Project webpage is https://lavr‑4d‑scene‑rerender.github.io/.
Authors:Oindrila Saha, Vojtech Krs, Radomir Mech, Subhransu Maji, Matheus Gadelha, Kevin Blackburn-Matzen
Abstract:
Recent progress in large language models (LLMs) has shown that reasoning improves when intermediate thoughts are externalized into explicit workspaces, such as chain‑of‑thought traces or tool‑augmented reasoning. Yet, visual language models (VLMs) lack an analogous mechanism for spatial reasoning, limiting their ability to generate images that accurately reflect geometric relations, object identities, and compositional intent. We introduce the concept of a spatial scratchpad ‑‑ a 3D reasoning substrate that bridges linguistic intent and image synthesis. Given a text prompt, our framework parses subjects and background elements, instantiates them as editable 3D meshes, and employs agentic scene planning for placement, orientation, and viewpoint selection. The resulting 3D arrangement is rendered back into the image domain with identity‑preserving cues, enabling the VLM to generate spatially consistent and visually coherent outputs. Unlike prior 2D layout‑based methods, our approach supports intuitive 3D edits that propagate reliably into final images. Empirically, it achieves a 32% improvement in text alignment on GenAI‑Bench, demonstrating the benefit of explicit 3D reasoning for precise, controllable image generation. Our results highlight a new paradigm for vision‑language models that deliberate not only in language, but also in space. Code and visualizations at https://oindrilasaha.github.io/3DScratchpad/
Authors:Po-Kai Chiu, Hung-Hsuan Chen
Abstract:
The requirement for expert annotations limits the effectiveness of deep learning for medical image analysis. Although 3D self‑supervised methods like volume contrast learning (VoCo) are powerful and partially address the labeling scarcity issue, their high computational cost and memory consumption are barriers. We propose 2D‑VoCo, an efficient adaptation of the VoCo framework for slice‑level self‑supervised pre‑training that learns spatial‑semantic features from unlabeled 2D CT slices via contrastive learning. The pre‑trained CNN backbone is then integrated into a CNN‑LSTM architecture to classify multi‑organ injuries. In the RSNA 2023 Abdominal Trauma dataset, 2D‑VoCo pre‑training significantly improves mAP, precision, recall, and RSNA score over training from scratch. Our framework provides a practical method to reduce the dependency on labeled data and enhance model performance in clinical CT analysis. We release the code for reproducibility. https://github.com/tkz05/2D‑VoCo‑CT‑Classifier
Authors:Yajvan Ravan, Aref Malek, Chester Dolph, Nikhil Behari
Abstract:
High‑altitude, multi‑spectral, aerial imagery is scarce and expensive to acquire, yet it is necessary for algorithmic advances and application of machine learning models to high‑impact problems such as wildfire detection. We introduce a human‑annotated dataset from the NASA Autonomous Modular Sensor (AMS) using 12‑channel, medium to high altitude (3 ‑ 50 km) aerial wildfire images similar to those used in current US wildfire missions. Our dataset combines spectral data from 12 different channels, including infrared (IR), short‑wave IR (SWIR), and thermal. We take imagery from 20 wildfire missions and randomly sample small patches to generate over 4000 images with high variability, including occlusions by smoke/clouds, easily‑confused false positives, and nighttime imagery.
We demonstrate results from a deep‑learning model to automate the human‑intensive process of fire perimeter determination. We train two deep neural networks, one for image classification and the other for pixel‑level segmentation. The networks are combined into a unique real‑time segmentation model to efficiently localize active wildfire on an incoming image feed. Our model achieves 96% classification accuracy, 74% Intersection‑over‑Union(IoU), and 84% recall surpassing past methods, including models trained on satellite data and classical color‑rule algorithms. By leveraging a multi‑spectral dataset, our model is able to detect active wildfire at nighttime and behind clouds, while distinguishing between false positives. We find that data from the SWIR, IR, and thermal bands is the most important to distinguish fire perimeters. Our code and dataset can be found here: https://github.com/nasa/Autonomous‑Modular‑Sensor‑Wildfire‑Segmentation/tree/main and https://drive.google.com/drive/folders/1‑u4vs9rqwkwgdeeeoUhftCxrfe_4QPTn?=usp=drive_link
Authors:Yixiong Chen, Zongwei Zhou, Wenxuan Li, Alan Yuille
Abstract:
Large‑scale medical segmentation datasets often combine manual and pseudo‑labels of uneven quality, which can compromise training and evaluation. Low‑quality labels may hamper performance and make the model training less robust. To address this issue, we propose SegAE (Segmentation Assessment Engine), a lightweight vision‑language model (VLM) that automatically predicts label quality across 142 anatomical structures. Trained on over four million image‑label pairs with quality scores, SegAE achieves a high correlation coefficient of 0.902 with ground‑truth Dice similarity and evaluates a 3D mask in 0.06s. SegAE shows several practical benefits: (I) Our analysis reveals widespread low‑quality labeling across public datasets; (II) SegAE improves data efficiency and training performance in active and semi‑supervised learning, reducing dataset annotation cost by one‑third and quality‑checking time by 70% per label. This tool provides a simple and effective solution for quality control in large‑scale medical segmentation datasets. The dataset, model weights, and codes are released at https://github.com/Schuture/SegAE.
Authors:Zhengyong Huang, Xingwen Sun, Xuting Chang, Ning Jiang, Yao Wang, Jianfei Sun, Hongbin Han, Yao Sui
Abstract:
Deformable image registration is a critical technology in medical image analysis, with broad applications in clinical practice such as disease diagnosis, multi‑modal fusion, and surgical navigation. Traditional methods often rely on iterative optimization, which is computationally intensive and lacks generalizability. Recent advances in deep learning have introduced attention‑based mechanisms that improve feature alignment, yet accurately registering regions with high anatomical variability remains challenging. In this study, we proposed a novel unsupervised deformable image registration framework, LGANet++, which employs a novel local‑global attention mechanism integrated with a unique technique for feature interaction and fusion to enhance registration accuracy, robustness, and generalizability. We evaluated our approach using five publicly available datasets, representing three distinct registration scenarios: cross‑patient, cross‑time, and cross‑modal CT‑MR registration. The results demonstrated that our approach consistently outperforms several state‑of‑the‑art registration methods, improving registration accuracy by 1.39% in cross‑patient registration, 0.71% in cross‑time registration, and 6.12% in cross‑modal CT‑MR registration tasks. These results underscore the potential of LGANet++ to support clinical workflows requiring reliable and efficient image registration. The source code is available at https://github.com/huangzyong/LGANet‑Registration.
Authors:Matthew Gwilliam, Xiao Wang, Xuefeng Hu, Zhenheng Yang
Abstract:
Models for image representation learning are typically designed for either recognition or generation. Various forms of contrastive learning help models learn to convert images to embeddings that are useful for classification, detection, and segmentation. On the other hand, models can be trained to reconstruct images with pixel‑wise, perceptual, and adversarial losses in order to learn a latent space that is useful for image generation. We seek to unify these two directions with a first‑of‑its‑kind model that learns representations which are simultaneously useful for recognition and generation. We train our model as a hyper‑network for implicit neural representation, which learns to map images to model weights for fast, accurate reconstruction. We further integrate our INR hyper‑network with knowledge distillation to improve its generalization and performance. Beyond the novel training design, the model also learns an unprecedented compressed embedding space with outstanding performance for various visual tasks. The complete model competes with state‑of‑the‑art results for image representation learning, while also enabling generative capabilities with its high‑quality tiny embeddings. The code is available at https://github.com/tiktok/huvr.
Authors:Sangbeom Lim, Seoung Wug Oh, Jiahui Huang, Heeji Yoon, Seungryong Kim, Joon-Young Lee
Abstract:
Generalizing video matting models to real‑world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask‑to‑Matte Model (VideoMaMa) that converts coarse segmentation masks into pixel accurate alpha mattes, by leveraging pretrained video diffusion models. VideoMaMa demonstrates strong zero‑shot generalization to real‑world footage, even though it is trained solely on synthetic data. Building on this capability, we develop a scalable pseudo‑labeling pipeline for large‑scale video matting and construct the Matting Anything in Video (MA‑V) dataset, which offers high‑quality matting annotations for more than 50K real‑world videos spanning diverse scenes and motions. To validate the effectiveness of this dataset, we fine‑tune the SAM2 model on MA‑V to obtain SAM2‑Matte, which outperforms the same model trained on existing matting datasets in terms of robustness on in‑the‑wild videos. These findings emphasize the importance of large‑scale pseudo‑labeled video matting and showcase how generative priors and accessible segmentation cues can drive scalable progress in video matting research.
Authors:Hongyuan Chen, Xingyu Chen, Youjia Zhang, Zexiang Xu, Anpei Chen
Abstract:
We present Motion 3‑to‑4, a feed‑forward framework for synthesising high‑quality 4D dynamic objects from a single monocular video and an optional 3D reference mesh. While recent advances have significantly improved 2D, video, and 3D content generation, 4D synthesis remains difficult due to limited training data and the inherent ambiguity of recovering geometry and motion from a monocular viewpoint. Motion 3‑to‑4 addresses these challenges by decomposing 4D synthesis into static 3D shape generation and motion reconstruction. Using a canonical reference mesh, our model learns a compact motion latent representation and predicts per‑frame vertex trajectories to recover complete, temporally coherent geometry. A scalable frame‑wise transformer further enables robustness to varying sequence lengths. Evaluations on both standard benchmarks and a new dataset with accurate ground‑truth geometry show that Motion 3‑to‑4 delivers superior fidelity and spatial consistency compared to prior work. Project page is available at https://motion3‑to‑4.github.io/.
Authors:Pengze Zhang, Yanze Wu, Mengtian Li, Xu Bai, Songtao Zhao, Fulong Ye, Chong Mou, Xinghui Li, Zhuowei Chen, Qian He, Mingyuan Gao
Abstract:
Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task‑specific temporal priors, failing to fully exploit the rich spatio‑temporal information inherent in videos, thereby limiting flexibility and generalization in video generation. To address these limitations, we propose OmniTransfer, a unified framework for spatio‑temporal video transfer. It leverages multi‑view information across frames to enhance appearance consistency and exploits temporal cues to enable fine‑grained temporal control. To unify various video transfer tasks, OmniTransfer incorporates three key designs: Task‑aware Positional Bias that adaptively leverages reference video information to improve temporal alignment or appearance consistency; Reference‑decoupled Causal Learning separating reference and target branches to enable precise reference transfer while improving efficiency; and Task‑adaptive Multimodal Alignment using multimodal semantic guidance to dynamically distinguish and tackle different tasks. Extensive experiments show that OmniTransfer outperforms existing methods in appearance (ID and style) and temporal transfer (camera movement and video effects), while matching pose‑guided methods in motion transfer without using pose, establishing a new paradigm for flexible, high‑fidelity video generation.
Authors:Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov
Abstract:
Pixel‑based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle multiple sources of shift and hinder systematic analysis. We introduce KAGE‑Env, a JAX‑native 2D platformer that factorizes the observation process into independently controllable visual axes while keeping the underlying control problem fixed. By construction, varying a visual axis affects performance only through the induced state‑conditional action distribution of a pixel policy, providing a clean abstraction for visual generalization. Building on this environment, we define KAGE‑Bench, a benchmark of six known‑axis suites comprising 34 train‑evaluation configuration pairs that isolate individual visual shifts. Using a standard PPO‑CNN baseline, we observe strong axis‑dependent failures, with background and photometric shifts often collapsing success, while agent‑appearance shifts are comparatively benign. Several shifts preserve forward motion while breaking task completion, showing that return alone can obscure generalization failures. Finally, the fully vectorized JAX implementation enables up to 33M environment steps per second on a single GPU, enabling fast and reproducible sweeps over visual factors. Code: https://avanturist322.github.io/KAGEBench/.
Authors:Rotem Gatenyo, Ohad Fried
Abstract:
We study zero‑shot 3D alignment of two given meshes, using a text prompt describing their spatial relation ‑‑ an essential capability for content creation and scene assembly. Earlier approaches primarily rely on geometric alignment procedures, while recent work leverages pretrained 2D diffusion models to model language‑conditioned object‑object spatial relationships. In contrast, we directly optimize the relative pose at test time, updating translation, rotation, and isotropic scale with CLIP‑driven gradients via a differentiable renderer, without training a new model. Our framework augments language supervision with geometry‑aware objectives: a variant of soft‑Iterative Closest Point (ICP) term to encourage surface attachment and a penetration loss to discourage interpenetration. A phased schedule strengthens contact constraints over time, and camera control concentrates the optimization on the interaction region. To enable evaluation, we curate a benchmark containing diverse categories and relations, and compare against baselines. Our method outperforms all alternatives, yielding semantically faithful and physically plausible alignments.
Authors:Bin Yu, Shijie Lian, Xiaopeng Lin, Yuliang Wei, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Xinming Wang, Bailing Wang, Cong Huang, Kai Chen
Abstract:
The fundamental premise of Vision‑Language‑Action (VLA) models is to harness the extensive general capabilities of pre‑trained Vision‑Language Models (VLMs) for generalized embodied intelligence. However, standard robotic fine‑tuning inevitably disrupts the pre‑trained feature space, leading to "catastrophic forgetting" that compromises the general visual understanding we aim to leverage. To effectively utilize the uncorrupted general capabilities of VLMs for robotic tasks, we propose TwinBrainVLA, which coordinates two isomorphic VLM pathways: a frozen generalist (also called "Left Brain") and a trainable specialist (also called "Right Brain"). Our architecture utilizes a Asymmetric Mixture‑of‑Transformers (AsyMoT) mechanism, enabling the Right Brain to dynamically query and fuse intact semantic knowledge from the Left Brain with proprioceptive states. This fused representation conditions a flow‑matching action expert for precise continuous control. Empirical results on SimplerEnv and RoboCasa benchmarks demonstrate that by explicitly retaining general capabilities, TwinBrainVLA achieves substantial performance gains over baseline models in complex manipulation tasks.
Authors:Renmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang, Victor Shea-Jay Huang, Shumin Zhang, Chengwei Pan, Han Qiu, Minlie Huang
Abstract:
As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi‑image instructions, this advancement may pose new safety risks. We study this problem by introducing MIR‑SafetyBench, the first benchmark focused on multi‑image reasoning safety, which consists of 2,676 instances across a taxonomy of 9 multi‑image relations. Our extensive evaluations on 19 MLLMs reveal a troubling trend: models with more advanced multi‑image reasoning can be more vulnerable on MIR‑SafetyBench. Beyond attack success rates, we find that many responses labeled as safe are superficial, often driven by misunderstanding or evasive, non‑committal replies. We further observe that unsafe generations exhibit lower attention entropy than safe ones on average. This internal signature suggests a possible risk that models may over‑focus on task solving while neglecting safety constraints. Our code and data are available at https://github.com/thu‑coai/MIR‑SafetyBench.
Authors:Xiaolu Liu, Yicong Li, Qiyuan He, Jiayin Zhu, Wei Ji, Angela Yao, Jianke Zhu
Abstract:
Textured 3D morphing seeks to generate smooth and plausible transitions between two 3D assets, preserving both structural coherence and fine‑grained appearance. This ability is crucial not only for advancing 3D generation research but also for practical applications in animation, editing, and digital content creation. Existing approaches either operate directly on geometry, limiting them to shape‑only morphing while neglecting textures, or extend 2D interpolation strategies into 3D, which often causes semantic ambiguity, structural misalignment, and texture blurring. These challenges underscore the necessity to jointly preserve geometric consistency, texture alignment, and robustness throughout the transition process. To address this, we propose Interp3D, a novel training‑free framework for textured 3D morphing. It harnesses generative priors and adopts a progressive alignment principle to ensure both geometric fidelity and texture coherence. Starting from semantically aligned interpolation in condition space, Interp3D enforces structural consistency via SLAT (Structured Latent)‑guided structure interpolation, and finally transfers appearance details through fine‑grained texture fusion. For comprehensive evaluations, we construct a dedicated dataset, Interp3DData, with graded difficulty levels and assess generation results from fidelity, transition smoothness, and plausibility. Both quantitative metrics and human studies demonstrate the significant advantages of our proposed approach over previous methods. Source code is available at https://github.com/xiaolul2/Interp3D.
Authors:Paul Walker, James A. D. Gardner, Andreea Ardelean, William A. P. Smith, Bernhard Egger
Abstract:
Inverse rendering is an ill‑posed problem, but priors like illumination priors, can simplify it. Existing work either disregards the spherical and rotation‑equivariant nature of illumination environments or does not provide a well‑behaved latent space. We propose a rotation‑equivariant variational autoencoder that models natural illumination on the sphere without relying on 2D projections. To preserve the SO(2)‑equivariance of environment maps, we use a novel Vector Neuron Vision Transformer (VN‑ViT) as encoder and a rotation‑equivariant conditional neural field as decoder. In the encoder, we reduce the equivariance from SO(3) to SO(2) using a novel SO(2)‑equivariant fully connected layer, an extension of Vector Neurons. We show that our SO(2)‑equivariant fully connected layer outperforms standard Vector Neurons when used in our SO(2)‑equivariant model. Compared to previous methods, our variational autoencoder enables smoother interpolation in latent space and offers a more well‑behaved latent space.
Authors:Hendrik Möller, Hanna Schoen, Robert Graf, Matan Atad, Nathan Molinier, Anjany Sekuboyina, Bettina K. Budai, Fabian Bamberg, Steffen Ringhof, Christopher Schlett, Tobias Pischon, Thoralf Niendorf, Josua A. Decker, Marc-André Weber, Bjoern Menze, Daniel Rueckert, Jan S. Kirschke
Abstract:
The human spine commonly consists of seven cervical, twelve thoracic, and five lumbar vertebrae. However, enumeration anomalies may result in individuals having eleven or thirteen thoracic vertebrae and four or six lumbar vertebrae. Although the identification of enumeration anomalies has potential clinical implications for chronic back pain and operation planning, the thoracolumbar junction is often poorly assessed and rarely described in clinical reports. Additionally, even though multiple deep‑learning‑based vertebra labeling algorithms exist, there is a lack of methods to automatically label enumeration anomalies. Our work closes that gap by introducing "Vertebra Identification with Anomaly Handling" (VERIDAH), a novel vertebra labeling algorithm based on multiple classification heads combined with a weighted vertebra sequence prediction algorithm. We show that our approach surpasses existing models on T2w TSE sagittal (98.30% vs. 94.24% of subjects with all vertebrae correctly labeled, p < 0.001) and CT imaging (99.18% vs. 77.26% of subjects with all vertebrae correctly labeled, p < 0.001) and works in arbitrary field‑of‑view images. VERIDAH correctly labeled the presence 2 Möller et al. of thoracic enumeration anomalies in 87.80% and 96.30% of T2w and CT images, respectively, and lumbar enumeration anomalies in 94.48% and 97.22% for T2w and CT, respectively. Our code and models are available at: https://github.com/Hendrik‑code/spineps.
Authors:Yongcong Ye, Kai Zhang, Yanghai Zhang, Enhong Chen, Longfei Li, Jun Zhou
Abstract:
Zero‑shot composed image retrieval (ZS‑CIR) is a rapidly growing area with significant practical applications, allowing users to retrieve a target image by providing a reference image and a relative caption describing the desired modifications. Existing ZS‑CIR methods often struggle to capture fine‑grained changes and integrate visual and semantic information effectively. They primarily rely on either transforming the multimodal query into a single text using image‑to‑text models or employing large language models for target image description generation, approaches that often fail to capture complementary visual information and complete semantic context. To address these limitations, we propose a novel Fine‑Grained Zero‑Shot Composed Image Retrieval method with Complementary Visual‑Semantic Integration (CVSI). Specifically, CVSI leverages three key components: (1) Visual Information Extraction, which not only extracts global image features but also uses a pre‑trained mapping network to convert the image into a pseudo token, combining it with the modification text and the objects most likely to be added. (2) Semantic Information Extraction, which involves using a pre‑trained captioning model to generate multiple captions for the reference image, followed by leveraging an LLM to generate the modified captions and the objects most likely to be added. (3) Complementary Information Retrieval, which integrates information extracted from both the query and database images to retrieve the target image, enabling the system to efficiently handle retrieval queries in a variety of situations. Extensive experiments on three public datasets (e.g., CIRR, CIRCO, and FashionIQ) demonstrate that CVSI significantly outperforms existing state‑of‑the‑art methods. Our code is available at https://github.com/yyc6631/CVSI.
Authors:Kaiyu Wu, Pucheng Han, Hualong Zhang, Naigeng Wu, Keze Wang
Abstract:
While Vision Language Models (VLMs) show advancing reasoning capabilities, their application in meteorology is constrained by a domain gap and a reasoning faithfulness gap. Specifically, mainstream Reinforcement Fine‑Tuning (RFT) can induce Self‑Contradictory Reasoning (Self‑Contra), where the model's reasoning contradicts its final answer, which is unacceptable in such a high‑stakes domain. To address these challenges, we construct WeatherQA, a novel multimodal reasoning benchmark in meteorology. We also propose Logically Consistent Reinforcement Fine‑Tuning (LoCo‑RFT), which resolves Self‑Contra by introducing a logical consistency reward. Furthermore, we introduce Weather‑R1, the first reasoning VLM with logical faithfulness in meteorology, to the best of our knowledge. Experiments demonstrate that Weather‑R1 improves performance on WeatherQA by 9.8 percentage points over the baseline, outperforming Supervised Fine‑Tuning and RFT, and even surpassing the original Qwen2.5‑VL‑32B. These results highlight the effectiveness of our LoCo‑RFT and the superiority of Weather‑R1. Our benchmark and code are available at https://github.com/Marcowky/Weather‑R1.
Authors:Wesam Moustafa, Hossam Elsafty, Helen Schneider, Lorenz Sparrenberg, Rafet Sifa
Abstract:
Label noise is a critical problem in medical image segmentation, often arising from the inherent difficulty of manual annotation. Models trained on noisy data are prone to overfitting, which degrades their generalization performance. While a number of methods and strategies have been proposed to mitigate noisy labels in the segmentation domain, this area remains largely under‑explored. The abstention mechanism has proven effective in classification tasks by enhancing the capabilities of Cross Entropy, yet its potential in segmentation remains unverified. In this paper, we address this gap by introducing a universal and modular abstention framework capable of enhancing the noise‑robustness of a diverse range of loss functions. Our framework improves upon prior work with two key components: an informed regularization term to guide abstention behaviour, and a more flexible power‑law‑based auto‑tuning algorithm for the abstention penalty. We demonstrate the framework's versatility by systematically integrating it with three distinct loss functions to create three novel, noise‑robust variants: GAC, SAC, and ADS. Experiments on the CaDIS and DSAD medical datasets show our methods consistently and significantly outperform their non‑abstaining baselines, especially under high noise levels. This work establishes that enabling models to selectively ignore corrupted samples is a powerful and generalizable strategy for building more reliable segmentation models. Our code is publicly available at https://github.com/wemous/abstention‑for‑segmentation.
Authors:Alexandre Justo Miro, Ludvig af Klinteberg, Bogdan Timus, Aron Asefaw, Ajinkya Khoche, Thomas Gustafsson, Sina Sharif Mansouri, Masoud Daneshtalab
Abstract:
Accurate ground truth annotations are critical to supervised learning and evaluating the performance of autonomous vehicle systems. These vehicles are typically equipped with active sensors, such as LiDAR, which scan the environment in predefined patterns. 3D box annotation based on data from such sensors is challenging in dynamic scenarios, where objects are observed at different timestamps, hence different positions. Without proper handling of this phenomenon, systematic errors are prone to being introduced in the box annotations. Our work is the first to discover such annotation errors in widely used, publicly available datasets. Through our novel offline estimation method, we correct the annotations so that they follow physically feasible trajectories and achieve spatial and temporal consistency with the sensor data. For the first time, we define metrics for this problem; and we evaluate our method on the Argoverse 2, MAN TruckScenes, and our proprietary datasets. Our approach increases the quality of box annotations by more than 17% in these datasets. Furthermore, we quantify the annotation errors in them and find that the original annotations are misplaced by up to 2.5 m, with highly dynamic objects being the most affected. Finally, we test the impact of the errors in benchmarking and find that the impact is larger than the improvements that state‑of‑the‑art methods typically achieve with respect to the previous state‑of‑the‑art methods; showing that accurate annotations are essential for correct interpretation of performance. Our code is available at https://github.com/alexandre‑justo‑miro/annotation‑correction‑3D‑boxes.
Authors:Jiangwei Xie, Zhang Wen, Mike Davies, Dongdong Chen
Abstract:
Hyperspectral image (HSI) restoration is a fundamental challenge in computational imaging and computer vision. It involves ill‑posed inverse problems, such as inpainting and super‑resolution. Although deep learning methods have transformed the field through data‑driven learning, their effectiveness hinges on access to meticulously curated ground‑truth datasets. This fundamentally restricts their applicability in real‑world scenarios where such data is unavailable. This paper presents SHARE (Single Hyperspectral Image Restoration with Equivariance), a fully unsupervised framework that unifies geometric equivariance principles with low‑rank spectral modelling to eliminate the need for ground truth. SHARE's core concept is to exploit the intrinsic invariance of hyperspectral structures under differentiable geometric transformations (e.g. rotations and scaling) to derive self‑supervision signals through equivariance consistency constraints. Our novel Dynamic Adaptive Spectral Attention (DASA) module further enhances this paradigm shift by explicitly encoding the global low‑rank property of HSI and adaptively refining local spectral‑spatial correlations through learnable attention mechanisms. Extensive experiments on HSI inpainting and super‑resolution tasks demonstrate the effectiveness of SHARE. Our method outperforms many state‑of‑the‑art unsupervised approaches and achieves performance comparable to that of supervised methods. We hope that our approach will shed new light on HSI restoration and broader scientific imaging scenarios. The code will be released at https://github.com/xuwayyy/SHARE.
Authors:Yuezhe Yang, Hao Wang, Yige Peng, Jinman Kim, Lei Bi
Abstract:
Automated clinical diagnosis remains a core challenge in medical AI, which usually requires models to integrate multi‑modal data and reason across complex, case‑specific contexts. Although recent methods have advanced medical report generation (MRG) and visual question answering (VQA) with medical vision‑language models (VLMs), these methods, however, predominantly operate under a sample‑isolated inference paradigm, as such processing cases independently without access to longitudinal electronic health records (EHRs) or structurally related patient examples. This paradigm limits reasoning to image‑derived information alone, which ignores external complementary medical evidence for potentially more accurate diagnosis. To overcome this limitation, we propose HyperWalker, a Deep Diagnosis framework that reformulates clinical reasoning via dynamic hypergraphs and test‑time training. First, we construct a dynamic hypergraph, termed iBrochure, to model the structural heterogeneity of EHR data and implicit high‑order associations among multimodal clinical information. Within this hypergraph, a reinforcement learning agent, Walker, navigates to and identifies optimal diagnostic paths. To ensure comprehensive coverage of diverse clinical characteristics in test samples, we incorporate a linger mechanism, a multi‑hop orthogonal retrieval strategy that iteratively selects clinically complementary neighborhood cases reflecting distinct clinical attributes. Experiments on MRG with MIMIC and medical VQA on EHRXQA demonstrate that HyperWalker achieves state‑of‑the‑art performance. Code is available at: https://github.com/Bean‑Young/HyperWalker
Authors:Xu Zhang, Danyang Li, Yingjie Xia, Xiaohang Dong, Hualong Yu, Jianye Wang, Qicheng Li
Abstract:
Change Detection (CD) is a fundamental task in remote sensing. It monitors the evolution of land cover over time. Based on this, Open‑Vocabulary Change Detection (OVCD) introduces a new requirement. It aims to reduce the reliance on predefined categories. Existing training‑free OVCD methods mostly use CLIP to identify categories. These methods also need extra models like DINO to extract features. However, combining different models often causes problems in matching features and makes the system unstable. Recently, the Segment Anything Model 3 (SAM 3) is introduced. It integrates segmentation and identification capabilities within one promptable model, which offers new possibilities for the OVCD task. In this paper, we propose OmniOVCD, a standalone framework designed for OVCD. By leveraging the decoupled output heads of SAM 3, we propose a Synergistic Fusion to Instance Decoupling (SFID) strategy. SFID first fuses the semantic, instance, and presence outputs of SAM 3 to construct land‑cover masks, and then decomposes them into individual instance masks for change comparison. This design preserves high accuracy in category recognition and maintains instance‑level consistency across images. As a result, the model can generate accurate change masks. Experiments on four public benchmarks (LEVIR‑CD, WHU‑CD, S2Looking, and SECOND) demonstrate SOTA performance, achieving IoU scores of 67.2, 66.5, 24.5, and 27.1 (class‑average), respectively, surpassing all previous methods. The code is available at https://github.com/Erxucomeon/OmniOVCD.
Authors:Shangzhe Di, Zhonghua Zhai, Weidi Xie
Abstract:
Current visual representation learning remains bifurcated: vision‑language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self‑supervised methods (e.g., MAE, DINO) capture intricate local structures yet struggle with high‑level semantic context. We argue that these paradigms are fundamentally complementary and can be integrated into a principled multi‑task framework, further enhanced by dense spatial supervision. We introduce MTV, a multi‑task visual pretraining framework that jointly optimizes a shared backbone across vision‑language contrastive, self‑supervised, and dense spatial objectives. To mitigate the need for manual annotations, we leverage high‑capacity "expert" models ‑‑ such as Depth Anything V2 and OWLv2 ‑‑ to synthesize dense, structured pseudo‑labels at scale. Beyond the framework, we provide a systematic investigation into the mechanics of multi‑task visual learning, analyzing: (i) the marginal gain of each objective, (ii) task synergies versus interference, and (iii) scaling behavior across varying data and model scales. Our results demonstrate that MTV achieves "best‑of‑both‑worlds" performance, significantly enhancing fine‑grained spatial reasoning without compromising global semantic understanding. Our findings suggest that multi‑task learning, fueled by high‑quality pseudo‑supervision, is a scalable path toward more general visual encoders.
Authors:Michail Spanakis, Iason Oikonomidis, Antonis Argyros
Abstract:
Class‑Agnostic object Counting (CAC) involves counting instances of objects from arbitrary classes within an image. Due to its practical importance, CAC has received increasing attention in recent years. Most existing methods assume a single object class per image, rely on extensive training of large deep learning models and address the problem by incorporating additional information, such as visual exemplars or text prompts. In this paper, we present OCCAM, the first training‑free approach to CAC that operates without the need of any supplementary information. Moreover, our approach addresses the multi‑class variant of the problem, as it is capable of counting the object instances in each and every class among arbitrary object classes within an image. We leverage Segment Anything Model 2 (SAM2), a foundation model, and a custom threshold‑based variant of the First Integer Neighbor Clustering Hierarchy (FINCH) algorithm to achieve competitive performance on widely used benchmark datasets, FSC‑147 and CARPK. We propose a synthetic multi‑class dataset and F1 score as a more suitable evaluation metric. The code for our method and the proposed synthetic dataset will be made publicly available at https://mikespanak.github.io/OCCAM_counter.
Authors:Qian Chen, Jinlan Fu, Changsong Li, See-Kiong Ng, Xipeng Qiu
Abstract:
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni‑modal perception, their ability to forecast future events from audio‑visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding. To bridge this gap, we introduce FutureOmni, the first benchmark designed to evaluate omni‑modal future forecasting from audio‑visual environments. The evaluated models are required to perform cross‑modal causal and temporal reasoning, as well as effectively leverage internal knowledge to predict future events. FutureOmni is constructed via a scalable LLM‑assisted, human‑in‑the‑loop pipeline and contains 919 videos and 1,034 multiple‑choice QA pairs across 8 primary domains. Evaluations on 13 omni‑modal and 7 video‑only models show that current systems struggle with audio‑visual future prediction, particularly in speech‑heavy scenarios, with the best accuracy of 64.8% achieved by Gemini 3 Flash. To mitigate this limitation, we curate a 7K‑sample instruction‑tuning dataset and propose an Omni‑Modal Future Forecasting (OFF) training strategy. Evaluations on FutureOmni and popular audio‑visual and video‑only benchmarks demonstrate that OFF enhances future forecasting and generalization. We publicly release all code (https://github.com/OpenMOSS/FutureOmni) and datasets (https://huggingface.co/datasets/OpenMOSS‑Team/FutureOmni).
Authors:Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi, Bernt Schiele, Jonas Fischer
Abstract:
Language‑aligned vision foundation models perform strongly across diverse downstream tasks. Yet, their learned representations remain opaque, making interpreting their decision‑making difficult. Recent work decompose these representations into human‑interpretable concepts, but provide poor spatial grounding and are limited to image classification tasks. In this work, we propose CFM, a language‑aligned concept foundation model for vision that provides fine‑grained concepts, which are human‑interpretable and spatially grounded in the input image. When paired with a foundation model with strong semantic representations, we get explanations for any of its downstream tasks. Examining local co‑occurrence dependencies of concepts allows us to define concept relationships through which we improve concept naming and obtain richer explanations. On benchmark data, we show that CFM provides performance on classification, segmentation, and captioning that is competitive with opaque foundation models while providing fine‑grained, high quality concept‑based explanations. Code at https://github.com/kawi19/CFM.
Authors:Daniel Kyselica, Jonáš Herec, Oliver Kutis, Rado Pitoňák
Abstract:
Natural disaster monitoring through continuous satellite observation requires processing multi‑temporal data under strict operational constraints. This paper addresses flood detection, a critical application for hazard management, by developing an onboard change detection system that operates within the memory and computational limits of small satellites. We propose History Injection mechanism for Transformer models (HiT), that maintains historical context from previous observations while reducing data storage by over 99% of original image size. Moreover, testing on the STTORM‑CD flood dataset confirms that the HiT mechanism within the Prithvi‑tiny foundation model maintains detection accuracy compared to the bi‑temporal baseline. The proposed HiT‑Prithvi model achieved 43 FPS on Jetson Orin Nano, a representative onboard hardware used in nanosats. This work establishes a practical framework for satellite‑based continuous monitoring of natural disasters, supporting real‑time hazard assessment without dependency on ground‑based processing infrastructure. Architecture as well as model checkpoints is available at https://github.com/zaitra/HiT‑change‑detection .
Authors:Sam Cantrill, David Ahmedt-Aristizabal, Lars Petersson, Hanna Suominen, Mohammad Ali Armin
Abstract:
Facial remote photoplethysmography (rPPG) methods estimate physiological signals by modeling subtle color changes on the 3D facial surface over time. However, existing methods fail to explicitly align their receptive fields with the 3D facial surface‑the spatial support of the rPPG signal. To address this, we propose the Facial Spatiotemporal Graph (STGraph), a novel representation that encodes facial color and structure using 3D facial mesh sequences‑enabling surface‑aligned spatiotemporal processing. We introduce MeshPhys, a lightweight spatiotemporal graph convolutional network that operates on the STGraph to estimate physiological signals. Across four benchmark datasets, MeshPhys achieves state‑of‑the‑art or competitive performance in both intra‑ and cross‑dataset settings. Ablation studies show that constraining the model's receptive field to the facial surface acts as a strong structural prior, and that surface‑aligned, 3D‑aware node features are critical for robustly encoding facial surface color. Together, the STGraph and MeshPhys constitute a novel, principled modeling paradigm for facial rPPG, enabling robust, interpretable, and generalizable estimation. Code is available at https://samcantrill.github.io/facial‑stgraph‑rppg/ .
Authors:Xinhao Liu, Yu Wang, Xiansheng Guo, Gordon Owusu Boateng, Yu Cao, Haonan Si, Xingchen Guo, Nirwan Ansari
Abstract:
High‑fidelity parking‑lot digital twins provide essential priors for path planning, collision checking, and perception validation in Automated Valet Parking (AVP). Yet robot‑oriented reconstruction faces a trilemma: sparse forward‑facing views cause weak parallax and ill‑posed geometry; dynamic occlusions and extreme lighting hinder stable texture fusion; and neural rendering typically needs expensive offline optimization, violating edge‑side streaming constraints. We propose ParkingTwin, a training‑free, lightweight system for online streaming 3D reconstruction. First, OSM‑prior‑driven geometric construction uses OpenStreetMap semantic topology to directly generate a metric‑consistent TSDF, replacing blind geometric search with deterministic mapping and avoiding costly optimization. Second, geometry‑aware dynamic filtering employs a quad‑modal constraint field (normal/height/depth consistency) to reject moving vehicles and transient occlusions in real time. Third, illumination‑robust fusion in CIELAB decouples luminance and chromaticity via adaptive L‑channel weighting and depth‑gradient suppression, reducing seams under abrupt lighting changes. ParkingTwin runs at 30+ FPS on an entry‑level GTX 1660. On a 68,000 m^2 real‑world dataset, it achieves SSIM 0.87 (+16.0%), delivers about 15x end‑to‑end speedup, and reduces GPU memory by 83.3% compared with state‑of‑the‑art 3D Gaussian Splatting (3DGS) that typically requires high‑end GPUs (RTX 4090D). The system outputs explicit triangle meshes compatible with Unity/Unreal digital‑twin pipelines. Project page: https://mihoutao‑liu.github.io/ParkingTwin/
Authors:Carsten T. Lüth, Jeremias Traub, Kim-Celine Kahl, Till J. Bungert, Lukas Klein, Lars Krämer, Paul F. Jäger, Klaus Maier-Hein, Fabian Isensee
Abstract:
Active learning (AL) has the potential to drastically reduce annotation costs in 3D biomedical image segmentation, where expert labeling of volumetric data is both time‑consuming and expensive. Yet, existing AL methods are unable to consistently outperform improved random sampling baselines adapted to 3D data, leaving the field without a reliable solution. We introduce Class‑stratified Scheduled Power Predictive Entropy (ClaSP PE), a simple and effective query strategy that addresses two key limitations of standard uncertainty‑based AL methods: class imbalance and redundancy in early selections. ClaSP PE combines class‑stratified querying to ensure coverage of underrepresented structures and log‑scale power noising with a decaying schedule to enforce query diversity in early‑stage AL and encourage exploitation later. In our evaluation on 24 experimental settings using four 3D biomedical datasets within the comprehensive nnActive benchmark, ClaSP PE is the only method that generally outperforms improved random baselines in terms of both segmentation quality with statistically significant gains, whilst remaining annotation efficient. Furthermore, we explicitly simulate the real‑world application by testing our method on four previously unseen datasets without manual adaptation, where all experiment parameters are set according to predefined guidelines. The results confirm that ClaSP PE robustly generalizes to novel tasks without requiring dataset‑specific tuning. Within the nnActive framework, we present compelling evidence that an AL method can consistently outperform random baselines adapted to 3D segmentation, in terms of both performance and annotation efficiency in a realistic, close‑to‑production scenario. Our open‑source implementation and clear deployment guidelines make it readily applicable in practice. Code is at https://github.com/MIC‑DKFZ/nnActive.
Authors:Qian Feng, JiaHang Tu, Mintong Kang, Hanbin Zhao, Chao Zhang, Hui Qian
Abstract:
Incremental unlearning (IU) is critical for pre‑trained models to comply with sequential data deletion requests, yet existing methods primarily suppress parameters or confuse knowledge without explicit constraints on both feature and gradient level, resulting in superficial forgetting where residual information remains recoverable. This incomplete forgetting risks security breaches and disrupts retention balance, especially in IU scenarios. We propose FG‑OrIU (Feature‑Gradient Orthogonality for Incremental Unlearning), the first framework unifying orthogonal constraints on both features and gradients level to achieve deep forgetting, where the forgetting effect is irreversible. FG‑OrIU decomposes feature spaces via Singular Value Decomposition (SVD), separating forgetting and remaining class features into distinct subspaces. It then enforces dual constraints: feature orthogonal projection on both forgetting and remaining classes, while gradient orthogonal projection prevents the reintroduction of forgotten knowledge and disruption to remaining classes during updates. Additionally, dynamic subspace adaptation merges newly forgetting subspaces and contracts remaining subspaces, ensuring a stable balance between removal and retention across sequential unlearning tasks. Extensive experiments demonstrate the effectiveness of our method.
Authors:Yu Qin, Shimeng Fan, Fan Yang, Zixuan Xue, Zijie Mai, Wenrui Chen, Kailun Yang, Zhiyong Li
Abstract:
Open‑vocabulary 6D object pose estimation empowers robots to manipulate arbitrary unseen objects guided solely by natural language. However, a critical limitation of existing approaches is their reliance on unconstrained global matching strategies. In open‑world scenarios, trying to match anchor features against the entire query image space introduces excessive ambiguity, as target features are easily confused with background distractors. To resolve this, we propose Fine‑grained Correspondence Pose Estimation (FiCoP), a framework that transitions from noise‑prone global matching to spatially‑constrained patch‑level correspondence. Our core innovation lies in leveraging a patch‑to‑patch correlation matrix as a structural prior to narrowing the matching scope, effectively filtering out irrelevant clutter to prevent it from degrading pose estimation. Firstly, we introduce an object‑centric disentanglement preprocessing to isolate the semantic target from environmental noise. Secondly, a Cross‑Perspective Global Perception (CPGP) module is proposed to fuse dual‑view features, establishing structural consensus through explicit context reasoning. Finally, we design a Patch Correlation Predictor (PCP) that generates a precise block‑wise association map, acting as a spatial filter to enforce fine‑grained, noise‑resilient matching. Experiments on the REAL275 and Toyota‑Light datasets demonstrate that FiCoP improves Average Recall by 8.0% and 6.1%, respectively, compared to the state‑of‑the‑art method, highlighting its capability to deliver robust and generalized perception for robotic agents operating in complex, unconstrained open‑world environments. The source code will be made publicly available at https://github.com/zjjqinyu/FiCoP.
Authors:Zhiguang Liu, Yi Shang
Abstract:
The Abstraction and Reasoning Corpus (ARC) provides a compact laboratory for studying abstract reasoning, an ability central to human intelligence. Modern AI systems, including LLMs and ViTs, largely operate as sequence‑of‑behavior prediction machines: they match observable behaviors by modeling token statistics without a persistent, readable mental state. This creates a gap with human‑like behavior: humans can explain an action by decoding internal state, while AI systems can produce fluent post‑hoc rationalizations that are not grounded in such a state. We hypothesize that reasoning is a modality: reasoning should exist as a distinct channel separate from the low‑level workspace on which rules are applied. To test this hypothesis, on solving ARC tasks as a visual reasoning problem, we designed a novel role‑separated transformer block that splits global controller tokens from grid workspace tokens, enabling iterative rule execution. Trained and evaluated within the VARC vision‑centric protocol, our method achieved 62.6% accuracy on ARC‑1, surpassing average human performance (60.2%) and outperforming prior methods significantly. Qualitatively, our models exhibit more coherent rule‑application structure than the dense ViT baseline, consistent with a shift away from plausible probability blobs toward controller‑driven reasoning.
Authors:Feng Ding, Wenhui Yi, Xinan He, Mengyao Xiao, Jianfeng Xu, Jianqiang Du
Abstract:
Generative models now produce imperceptible, fine‑grained manipulated faces, posing significant privacy risks. However, existing AI‑generated face datasets generally lack focus on samples with fine‑grained regional manipulations. Furthermore, no researchers have yet studied the real impact of splice attacks, which occur between real and manipulated samples, on detectors. We refer to these as detector‑evasive samples. Based on this, we introduce the DiffFace‑Edit dataset, which has the following advantages: 1) It contains over two million AI‑generated fake images. 2) It features edits across eight facial regions (e.g., eyes, nose) and includes a richer variety of editing combinations, such as single‑region and multi‑region edits. Additionally, we specifically analyze the impact of detector‑evasive samples on detection models. We conduct a comprehensive analysis of the dataset and propose a cross‑domain evaluation that combines IMDL methods. Dataset will be available at https://github.com/ywh1093/DiffFace‑Edit.
Authors:Yang Yu, Yunze Deng, Yige Zhang, Yanjie Xiao, Youkun Ou, Wenhao Hu, Mingchao Li, Bin Feng, Wenyu Liu, Dandan Zheng, Jingdong Chen
Abstract:
Existing image‑based virtual try‑on (VTON) methods primarily focus on single‑layer or multi‑garment VTON, neglecting multi‑layer VTON (ML‑VTON), which involves dressing multiple layers of garments onto the human body with realistic deformation and layering to generate visually plausible outcomes. The main challenge lies in accurately modeling occlusion relationships between inner and outer garments to reduce interference from redundant inner garment features. To address this, we propose GO‑MLVTON, the first multi‑layer VTON method, introducing the Garment Occlusion Learning module to learn occlusion relationships and the StableDiffusion‑based Garment Morphing & Fitting module to deform and fit garments onto the human body, producing high‑quality multi‑layer try‑on results. Additionally, we present the MLG dataset for this task and propose a new metric named Layered Appearance Coherence Difference (LACD) for evaluation. Extensive experiments demonstrate the state‑of‑the‑art performance of GO‑MLVTON. Project page: https://upyuyang.github.io/go‑mlvton/.
Authors:Lavsen Dahal, Yubraj Bhandari, Geoffrey D. Rubin, Joseph Y. Lo
Abstract:
There is an urgent need for triage and classification of high‑volume medical imaging modalities such as computed tomography (CT), which can improve patient care and mitigate radiologist burnout. Study‑level CT triage requires calibrated predictions with localized evidence; however, off‑the‑shelf Vision Language Models (VLM) struggle with 3D anatomy, protocol shifts, and noisy report supervision. This study used the two largest publicly available chest CT datasets: CT‑RATE and RADCHEST‑CT (held‑out external test set). Our carefully tuned supervised baseline (instantiated as a simple Global Average Pooling head) establishes a new supervised state of the art, surpassing all reported linear‑probe VLMs. Building on this baseline, we present ORACLE‑CT, an encoder‑agnostic, organ‑aware head that pairs Organ‑Masked Attention (mask‑restricted, per‑organ pooling that yields spatial evidence) with Organ‑Scalar Fusion (lightweight fusion of normalized volume and mean‑HU cues). In the chest setting, ORACLE‑CT masked attention model achieves AUROC 0.86 on CT‑RATE; in the abdomen setting, on MERLIN (30 findings), our supervised baseline exceeds a reproduced zero‑shot VLM baseline obtained by running publicly released weights through our pipeline, and adding masked attention plus scalar fusion further improves performance to AUROC 0.85. Together, these results deliver state‑of‑the‑art supervised classification performance across both chest and abdomen CT under a unified evaluation protocol. The source code is available at https://github.com/lavsendahal/oracle‑ct.
Authors:Wei Wang, Quoc-Toan Ly, Chong Yu, Jun Bai
Abstract:
Spatial transcriptomics (ST) enables transcriptome‑wide profiling while preserving the spatial context of tissues, offering unprecedented opportunities to study tissue organization and cell‑cell interactions in situ. Despite recent advances, existing methods often lack effective integration of histological morphology with molecular profiles, relying on shallow fusion strategies or omitting tissue images altogether, which limits their ability to resolve ambiguous spatial domain boundaries. To address this challenge, we propose MultiST, a unified multimodal framework that jointly models spatial topology, gene expression, and tissue morphology through cross‑attention‑based fusion. MultiST employs graph‑based gene encoders with adversarial alignment to learn robust spatial representations, while integrating color‑normalized histological features to capture molecular‑morphological dependencies and refine domain boundaries. We evaluated the proposed method on 13 diverse ST datasets spanning two organs, including human brain cortex and breast cancer tissue. MultiST yields spatial domains with clearer and more coherent boundaries than existing methods, leading to more stable pseudotime trajectories and more biologically interpretable cell‑cell interaction patterns. The MultiST framework and source code are available at https://github.com/LabJunBMI/MultiST.git.
Authors:Wenxin Ma, Chenlong Wang, Ruisheng Yuan, Hao Chen, Nanru Dai, S. Kevin Zhou, Yijun Yang, Alan Yuille, Jieneng Chen
Abstract:
Humans can look at a static scene and instantly predict what happens next ‑‑ will moving this object cause a collision? We call this ability Causal Spatial Reasoning. However, current multimodal large language models (MLLMs) cannot do this, as they remain largely restricted to static spatial perception, struggling to answer "what‑if" questions in a 3D scene. We introduce CausalSpatial, a diagnostic benchmark evaluating whether models can anticipate consequences of object motions across four tasks: Collision, Compatibility, Occlusion, and Trajectory. Results expose a severe gap: humans score 84% while GPT‑5 achieves only 54%. Why do MLLMs fail? Our analysis uncovers a fundamental deficiency: models over‑rely on textual chain‑of‑thought reasoning that drifts from visual evidence, producing fluent but spatially ungrounded hallucinations. To address this, we propose the Causal Object World model (COW), a framework that externalizes the simulation process by generating videos of hypothetical dynamics. With explicit visual cues of causality, COW enables models to ground their reasoning in physical reality rather than linguistic priors. We make the dataset and code publicly available here: https://github.com/CausalSpatial/CausalSpatial
Authors:Pedro M. Gordaliza, Jaume Banus, Benoît Gérin, Maxence Wynen, Nataliia Molchanova, Jonas Richiardi, Meritxell Bach Cuadra
Abstract:
Developing Foundation Models for medical image analysis is essential to overcome the unique challenges of radiological tasks. The first challenges of this kind for 3D brain MRI, SSL3D and FOMO25, were held at MICCAI 2025. Our solution ranked first in tracks of both contests. It relies on a U‑Net CNN architecture combined with strategies leveraging anatomical priors and neuroimaging domain knowledge. Notably, our models trained 1‑2 orders of magnitude faster and were 10 times smaller than competing transformer‑based approaches. Models are available here: https://github.com/jbanusco/BrainFM4Challenges.
Authors:Richard Shaw, Youngkyoon Jang, Athanasios Papaioannou, Arthur Moreau, Helisa Dhamo, Zhensong Zhang, Eduardo Pérez-Pellitero
Abstract:
This work presents Interactive Conversational 3D Virtual Human (ICo3D), a method for generating an interactive, conversational, and photorealistic 3D human avatar. Based on multi‑view captures of a subject, we create an animatable 3D face model and a dynamic 3D body model, both rendered by splatting Gaussian primitives. Once merged together, they represent a lifelike virtual human avatar suitable for real‑time user interactions. We equip our avatar with an LLM for conversational ability. During conversation, the audio speech of the avatar is used as a driving signal to animate the face model, enabling precise synchronization. We describe improvements to our dynamic Gaussian models that enhance photorealism: SWinGS++ for body reconstruction and HeadGaS++ for face reconstruction, and provide as well a solution to merge the separate face and body models without artifacts. We also present a demo of the complete system, showcasing several use cases of real‑time conversation with the 3D avatar. Our approach offers a fully integrated virtual avatar experience, supporting both oral and written form interactions in immersive environments. ICo3D is applicable to a wide range of fields, including gaming, virtual assistance, and personalized education, among others. Project page: https://ico3d.github.io/
Authors:Muhayy Ud Din, Waseem Akram, Ahsan B. Bakht, Irfan Hussain
Abstract:
Maritime port inspection plays a critical role in ensuring safety, regulatory compliance, and operational efficiency in complex maritime environments. However, existing inspection methods often rely on manual operations and conventional computer vision techniques that lack scalability and contextual understanding. This study introduces a novel integrated engineering framework that utilizes the synergy between Large Language Models (LLMs) and Vision Language Models (VLMs) to enable autonomous maritime port inspection using cooperative aerial and surface robotic platforms. The proposed framework replaces traditional state‑machine mission planners with LLM‑driven symbolic planning and improved perception pipelines through VLM‑based semantic inspection, enabling context‑aware and adaptive monitoring. The LLM module translates natural language mission instructions into executable symbolic plans with dependency graphs that encode operational constraints and ensure safe UAV‑USV coordination. Meanwhile, the VLM module performs real‑time semantic inspection and compliance assessment, generating structured reports with contextual reasoning. The framework was validated using the extended MBZIRC Maritime Simulator with realistic port infrastructure and further assessed through real‑world robotic inspection trials. The lightweight on‑board design ensures suitability for resource‑constrained maritime platforms, advancing the development of intelligent, autonomous inspection systems. Project resources (code and videos) can be found here: https://github.com/Muhayyuddin/llm‑vlm‑fusion‑port‑inspection
Authors:Yulun Guo
Abstract:
Crack detection is critical for concrete infrastructure safety, but real‑world cracks often appear in low‑light environments like tunnels and bridge undersides, degrading computer vision segmentation accuracy. Pixel‑level annotation of low‑light crack images is extremely time‑consuming, yet most deep learning methods require large, well‑illuminated datasets. We propose a dual‑branch prototype learning network integrating Retinex theory with few‑shot learning for low‑light crack segmentation. Retinex‑based reflectance components guide illumination‑invariant global representation learning, while metric learning reduces dependence on large annotated datasets. We introduce a cross‑similarity prior mask generation module that computes high‑dimensional similarities between query and support features to capture crack location and structure, and a multi‑scale feature enhancement module that fuses multi‑scale features with the prior mask to alleviate spatial inconsistency. Extensive experiments on multiple benchmarks demonstrate consistent state‑of‑the‑art performance under low‑light conditions. Code: https://github.com/YulunGuo/CrackFSS.
Authors:Zaibin Zhang, Yuhan Wu, Lianjie Jia, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, Lijun Wang, Huchuan Lu
Abstract:
While contemporary Vision‑Language Models (VLMs) excel at 2D visual understanding, they remain constrained by a passive, 2D‑centric paradigm that severely limits genuine 3D spatial reasoning. To bridge this gap, we introduce Think3D, a novel framework that equips VLM agents with interactive, 3D chain‑of‑thought reasoning capabilities. By integrating a suite of 3D manipulation tools, Think3D transforms passive perception into active spatial exploration, closely mirroring human geometric reasoning. We demonstrate that Think3D acts as a highly effective zero‑shot plug‑in for state‑of‑the‑art closed‑source models (e.g., GPT‑4.1, Gemini 2.5 Pro), yielding absolute performance gains of +7.8% on BLINK Multi‑view and MindCube, and +4.7% on VSI‑Bench. Furthermore, to optimize tool‑use in smaller open‑weight models, we propose Think3D‑RL, a reinforcement learning paradigm designed to autonomously learn spatial exploration strategies. When applied to Qwen3‑VL‑4B, Think3D‑RL amplifies the performance gain from a marginal +0.7% to a substantial +10.7%. Notably, this RL formulation induces an exploration policy that qualitatively aligns with the sophisticated behavior of much larger models, entirely circumventing the need for costly operation‑trajectory annotations. Ultimately, Think3D establishes tool‑augmented active exploration as an effective paradigm for unlocking human‑like 3D reasoning in multimodal agents. Code, models, and data are available at https://github.com/zhangzaibin/spagent.
Authors:Jun Wan, Xinyu Xiong, Ning Chen, Zhihui Lai, Jie Zhou, Wenwen Min
Abstract:
Recently, deep learning based facial landmark detection (FLD) methods have achieved considerable success. However, in challenging scenarios such as large pose variations, illumination changes, and facial expression variations, they still struggle to accurately capture the geometric structure of the face, resulting in performance degradation. Moreover, the limited size and diversity of existing FLD datasets hinder robust model training, leading to reduced detection accuracy. To address these challenges, we propose a Frequency‑Guided Task‑Balancing Transformer (FGTBT), which enhances facial structure perception through frequency‑domain modeling and multi‑dataset unified training. Specifically, we propose a novel Fine‑Grained Multi‑Task Balancing loss (FMB‑loss), which moves beyond coarse task‑level balancing by assigning weights to individual landmarks based on their occurrence across datasets. This enables more effective unified training and mitigates the issue of inconsistent gradient magnitudes. Additionally, a Frequency‑Guided Structure‑Aware (FGSA) model is designed to utilize frequency‑guided structure injection and regularization to help learn facial structure constraints. Extensive experimental results on popular benchmark datasets demonstrate that the integration of the proposed FMB‑loss and FGSA model into our FGTBT framework achieves performance comparable to state‑of‑the‑art methods. The code is available at https://github.com/Xi0ngxinyu/FGTBT.
Authors:Peng Li, Zihan Zhuang, Yangfan Gao, Yi Dong, Sixian Li, Changhao Jiang, Shihan Dou, Zhiheng Xi, Enyu Zhou, Jixuan Huang, Hui Li, Jingjing Gong, Xingjun Ma, Tao Gui, Zuxuan Wu, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Xipeng Qiu
Abstract:
Humanoid robots are capable of performing various actions such as greeting, dancing and even backflipping. However, these motions are often hard‑coded or specifically trained, which limits their versatility. In this work, we present FRoM‑W1, an open‑source framework designed to achieve general humanoid whole‑body motion control using natural language. To universally understand natural language and generate corresponding motions, as well as enable various humanoid robots to stably execute these motions in the physical world under gravity, FRoM‑W1 operates in two stages: (a) H‑GPT: utilizing massive human data, a large‑scale language‑driven human whole‑body motion generation model is trained to generate diverse natural behaviors. We further leverage the Chain‑of‑Thought technique to improve the model's generalization in instruction understanding. (b) H‑ACT: After retargeting generated human whole‑body motions into robot‑specific actions, a motion controller that is pretrained and further fine‑tuned through reinforcement learning in physical simulation enables humanoid robots to accurately and stably perform corresponding actions. It is then deployed on real robots via a modular simulation‑to‑reality module. We extensively evaluate FRoM‑W1 on Unitree H1 and G1 robots. Results demonstrate superior performance on the HumanML3D‑X benchmark for human whole‑body motion generation, and our introduced reinforcement learning fine‑tuning consistently improves both motion tracking accuracy and task success rates of these humanoid robots. We open‑source the entire FRoM‑W1 framework and hope it will advance the development of humanoid intelligence.
Authors:Shuling Zhao, Dan Xu
Abstract:
Building 3D animatable head avatars from a single image is an important yet challenging problem. Existing methods generally collapse under large camera pose variations, compromising the realism of 3D avatars. In this work, we propose a new framework to tackle the novel setting of one‑shot 3D full‑head animatable avatar reconstruction in a single feed‑forward pass, enabling real‑time animation and simultaneous 360^\circ rendering views. To facilitate efficient animation control, we model 3D head avatars with Gaussian primitives embedded on the surface of a parametric face model within the UV space. To obtain knowledge of full‑head geometry and textures, we leverage rich 3D full‑head priors within a pretrained 3D generative adversarial network (GAN) for global full‑head feature extraction and multi‑view supervision. To increase the fidelity of the 3D reconstruction of the input image, we take advantage of the symmetric nature of the UV space and human faces to fuse local fine‑grained input image features with the global full‑head textures. Extensive experiments demonstrate the effectiveness of our method, achieving high‑quality 3D full‑head modeling as well as real‑time animation, thereby improving the realism of 3D talking avatars.
Authors:Zequn Xie, Boyun Zhang, Yuxiao Lin, Tao Jin
Abstract:
Video‑text retrieval (VTR) aims to locate relevant videos using natural language queries. Current methods, often based on pre‑trained models like CLIP, are hindered by video's inherent redundancy and their reliance on coarse, final‑layer features, limiting matching accuracy. To address this, we introduce the HVP‑Net (Hierarchical Visual Perception Network), a framework that mines richer video semantics by extracting and refining features from multiple intermediate layers of a vision encoder. Our approach progressively distills salient visual concepts from raw patch‑tokens at different semantic levels, mitigating redundancy while preserving crucial details for alignment. This results in a more robust video representation, leading to new state‑of‑the‑art performance on challenging benchmarks including MSRVTT, DiDeMo, and ActivityNet. Our work validates the effectiveness of exploiting hierarchical features for advancing video‑text retrieval. Our codes are available at https://github.com/boyun‑zhang/HVP‑Net.
Authors:Lu Yue, Yue Fan, Shiwei Lian, Yu Zhao, Jiaxin Yu, Liang Xie, Feitian Zhang
Abstract:
Zero‑shot Vision‑and‑Language Navigation (VLN) agents leveraging Large Language Models (LLMs) excel in generalization but suffer from insufficient spatial perception. Focusing on complex continuous environments, we categorize key perceptual bottlenecks into three spatial challenges: door interaction,multi‑room navigation, and ambiguous instruction execution, where existing methods consistently suffer high failure rates. We present Spatial‑VLN, a perception‑guided exploration framework designed to overcome these challenges. The framework consists of two main modules. The Spatial Perception Enhancement (SPE) module integrates panoramic filtering with specialized door and region experts to produce spatially coherent, cross‑view consistent perceptual representations. Building on this foundation, our Explored Multi‑expert Reasoning (EMR) module uses parallel LLM experts to address waypoint‑level semantics and region‑level spatial transitions. When discrepancies arise between expert predictions, a query‑and‑explore mechanism is activated, prompting the agent to actively probe critical areas and resolve perceptual ambiguities. Experiments on VLN‑CE demonstrate that Spatial VLN achieves state‑of‑the‑art performance using only low‑cost LLMs. Furthermore, to validate real‑world applicability, we introduce a value‑based waypoint sampling strategy that effectively bridges the Sim2Real gap. Extensive real‑world evaluations confirm that our framework delivers superior generalization and robustness in complex environments. Our codes and videos are available at https://yueluhhxx.github.io/Spatial‑VLN‑web/.
Authors:Qingtian Zhu, Xu Cao, Zhixiang Wang, Yinqiang Zheng, Takafumi Taketomi
Abstract:
We propose KaoLRM to re‑target the learned prior of the Large Reconstruction Model (LRM) for parametric 3D face reconstruction from single‑view images. Parametric 3D Morphable Models (3DMMs) have been widely used for facial reconstruction due to their compact and interpretable parameterization, yet existing 3DMM regressors often exhibit poor consistency across varying viewpoints. To address this, we harness the pre‑trained 3D prior of LRM and incorporate FLAME‑based 2D Gaussian Splatting into LRM's rendering pipeline. Specifically, KaoLRM projects LRM's pre‑trained triplane features into the FLAME parameter space to recover geometry, and models appearance via 2D Gaussian primitives that are tightly coupled to the FLAME mesh. The rich prior enables the FLAME regressor to be aware of the 3D structure, leading to accurate and robust reconstructions under self‑occlusions and diverse viewpoints. Experiments on both controlled and in‑the‑wild benchmarks demonstrate that KaoLRM achieves superior reconstruction accuracy and cross‑view consistency, while existing methods remain sensitive to viewpoint variations. The code is released at https://github.com/CyberAgentAILab/KaoLRM.
Authors:Lin Zhao, Yushu Wu, Aleksei Lebedev, Dishani Lahiri, Meng Dong, Arpit Sahni, Michael Vasilkovsky, Hao Chen, Ju Hu, Aliaksandr Siarohin, Sergey Tulyakov, Yanzhi Wang, Anil Kag, Yanyu Li
Abstract:
Diffusion Transformers (DiTs) have recently improved video generation quality. However, their heavy computational cost makes real‑time or on‑device generation infeasible. In this work, we introduce S2DiT, a Streaming Sandwich Diffusion Transformer designed for efficient, high‑fidelity, and streaming video generation on mobile hardware. S2DiT generates more tokens but maintains efficiency with novel efficient attentions: a mixture of LinConv Hybrid Attention (LCHA) and Stride Self‑Attention (SSA). Based on this, we uncover the sandwich design via a budget‑aware dynamic programming search, achieving superior quality and efficiency. We further propose a 2‑in‑1 distillation framework that transfers the capacity of large teacher models (e.g., Wan 2.2‑14B) to the compact few‑step sandwich model. Together, S2DiT achieves quality on par with state‑of‑the‑art server video models, while streaming at over 10 FPS on an iPhone.
Authors:Raphi Kang, Hongqiao Chen, Georgia Gkioxari, Pietro Perona
Abstract:
Spatio‑temporal reasoning is a remarkable capability of Vision Language Models (VLMs), but the underlying mechanisms of such abilities remain largely opaque. We postulate that visual/geometrical and textual representations of spatial structure must be combined at some point in VLM computations. We search for such confluence, and ask whether the identified representation can causally explain aspects of input‑output model behavior through a linear model. We show empirically that VLMs encode object locations by linearly binding spatial IDs to textual activations, then perform reasoning via language tokens. Through rigorous causal interventions we demonstrate that these IDs, which are ubiquitous across the model, can systematically mediate model beliefs at intermediate VLM layers. Additionally, we find that spatial IDs serve as a diagnostic tool for identifying limitations in existing VLMs, and as a valuable learning signal. We extend our analysis to video VLMs and identify an analogous linear temporal ID mechanism. By characterizing our proposed spatiotemporal ID mechanism, we elucidate a previously underexplored internal reasoning process in VLMs, toward improved interpretability and the principled design of more aligned and capable models. We release our code for reproducibility: https://github.com/Raphoo/linear‑mech‑vlms.
Authors:Jan Fabian Schmid, Annika Hagemann
Abstract:
Sparse keypoint matching is crucial for 3D vision tasks, yet current keypoint detectors often produce spatially inaccurate matches. Existing refinement methods mitigate this issue through alignment of matched keypoint locations, but they are typically detector‑specific, requiring retraining for each keypoint detector. We introduce XRefine, a novel, detector‑agnostic approach for sub‑pixel keypoint refinement that operates solely on image patches centered at matched keypoints. Our cross‑attention‑based architecture learns to predict refined keypoint coordinates without relying on internal detector representations, enabling generalization across detectors. Furthermore, XRefine can be extended to handle multi‑view feature tracks. Experiments on MegaDepth, KITTI, and ScanNet demonstrate that the approach consistently improves geometric estimation accuracy, achieving superior performance compared to existing refinement methods while maintaining runtime efficiency. Our code and trained models can be found at https://github.com/boschresearch/xrefine.
Authors:Richard Liu, Itai Lang, Rana Hanocka
Abstract:
Handle‑based mesh deformation is a classic paradigm in computer graphics which enables intuitive edits from sparse controls. Classical techniques are fast and precise, but require users to know ideal handle placement apriori, which can be unintuitive and inconsistent. Handle sets cannot be adjusted easily, as weights are typically optimized through energies defined by the handles. Modern data‑driven methods, on the other hand, provide semantic edits but sacrifice fine‑grained control and speed. We propose a technique that achieves the best of both worlds: deep feature proximity yields smooth, visual‑aware deformation weights with no additional regularization. Importantly, these weights are computed in real‑time for any surface point, unlike prior methods which require expensive optimization. We introduce barycentric feature distillation, an improved feature distillation pipeline which leverages the full visual signal from shape renders to make distillation complexity robust to mesh resolution. This enables high resolution meshes to be processed in minutes versus potentially hours for prior methods. We preserve and extend classical properties through feature space constraints and locality weighting. Our field representation enables automatic visual symmetry detection, which we use to produce symmetry‑preserving deformations. We show a proof‑of‑concept application which can produce deformations for meshes up to 1 million faces in real‑time on a consumer‑grade machine. Project page at https://threedle.github.io/dfd.
Authors:Ruo Qi, Linhui Dai, Yusong Qin, Chaolei Yang, Yanshan Li
Abstract:
In remote sensing images, complex backgrounds, weak object signals, and small object scales make accurate detection particularly challenging, especially under low‑quality imaging conditions. A common strategy is to integrate single‑image super‑resolution (SR) before detection; however, such serial pipelines often suffer from misaligned optimization objectives, feature redundancy, and a lack of effective interaction between SR and detection. To address these issues, we propose a Saliency‑Driven multi‑task Collaborative Network (SDCoNet) that couples SR and detection through implicit feature sharing while preserving task specificity. SDCoNet employs the swin transformer‑based shared encoder, where hierarchical window‑shifted self‑attention supports cross‑task feature collaboration and adaptively balances the trade‑off between texture refinement and semantic representation. In addition, a multi‑scale saliency prediction module produces importance scores to select key tokens, enabling focused attention on weak object regions, suppression of background clutter, and suppression of adverse features introduced by multi‑task coupling. Furthermore, a gradient routing strategy is introduced to mitigate optimization conflicts. It first stabilizes detection semantics and subsequently routes SR gradients along a detection‑oriented direction, enabling the framework to guide the SR branch to generate high‑frequency details that are explicitly beneficial for detection. Experiments on public datasets, including NWPU VHR‑10‑Split, DOTAv1.5‑Split, and HRSSD‑Split, demonstrate that the proposed method, while maintaining competitive computational efficiency, significantly outperforms existing mainstream algorithms in small object detection on low‑quality remote sensing images. Our code is available at https://github.com/qiruo‑ya/SDCoNet.
Authors:Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma, Wanli Ouyang, Antoni B. Chan
Abstract:
Counting and tracking dense crowds in large‑scale scenes is a highly practical yet challenging problem. Existing methods mostly rely on fixed‑camera datasets with limited scene coverage, making them inadequate for crowd analysis in large‑scale scenes. To bridge this gap, we introduce MovingDroneCrowd++, the largest video‑level dataset dedicated to dense crowd counting and tracking with fast‑moving drones, captured under diverse flight altitudes, camera angles, and illumination conditions. Existing methods, however, still fail to achieve satisfactory video individual counting or tracking performance under these challenging aerial conditions. To this end, we propose GD3A (Global Density map Decomposition via group‑wise Descriptor Association), a video individual counting method that first establishes pixel‑level correspondences between pedestrian descriptors across frames via optimal transport with an adaptive dustbin score. Then, group‑wise association is adopted to guide the decomposition of the global density map into shared, inflow, and outflow density maps. We further introduce a pedestrian tracking method, DVTrack (Descriptor Voting Track), which converts descriptor‑level matching into instance‑level association through descriptor voting. Our methods rely on the association results of group‑wise multiple descriptors for each pedestrian rather than a single vector. Since intra‑group matching errors do not affect the final counting and tracking results, our methods are more robust in dense crowds and challenging aerial conditions. Experiments show that our methods achieve substantial gains in both crowd counting and tracking on moving‑drone videos with dense crowds and complex motions, reducing counting error by 47.4% and improving tracking accuracy by 64.6%. Code, dataset, and pretrained models are available at https://github.com/fyw1999/MovingDroneCrowd.
Authors:Mehrdad Noori, Gustavo Adolfo Vargas Hakim, David Osowiechi, Fereshteh Shakeri, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Ismail Ben Ayed, Christian Desrosiers
Abstract:
Medical Vision‑language models (VLMs) have shown remarkable performances in various medical imaging domains such as histo\‑pathology by leveraging pre‑trained, contrastive models that exploit visual and textual information. However, histopathology images may exhibit severe domain shifts, such as staining, contamination, blurring, and noise, which may severely degrade the VLM's downstream performance. In this work, we introduce Histopath‑C, a new benchmark with realistic synthetic corruptions designed to mimic real‑world distribution shifts observed in digital histopathology. Our framework dynamically applies corruptions to any available dataset and evaluates Test‑Time Adaptation (TTA) mechanisms on the fly. We then propose LATTE, a transductive, low‑rank adaptation strategy that exploits multiple text templates, mitigating the sensitivity of histopathology VLMs to diverse text inputs. Our approach outperforms state‑of‑the‑art TTA methods originally designed for natural images across a breadth of histopathology datasets, demonstrating the effectiveness of our proposed design for robust adaptation in histopathology images. Code and data are available at https://github.com/Mehrdad‑Noori/Histopath‑C.
Authors:Shunyu Huang, Yunjiao Zhou, Jianfei Yang
Abstract:
Skeleton‑based action recognition leverages human pose keypoints to categorize human actions, which shows superior generalization and interoperability compared to regular end‑to‑end action recognition. Existing solutions use RGB cameras to annotate skeletal keypoints, but their performance declines in dark environments and raises privacy concerns, limiting their use in smart homes and hospitals. This paper explores non‑invasive wireless sensors, i.e., LiDAR and mmWave, to mitigate these challenges as a feasible alternative. Two problems are addressed: (1) insufficient data on wireless sensor modality to train an accurate skeleton estimation model, and (2) skeletal keypoints derived from wireless sensors are noisier than RGB, causing great difficulties for subsequent action recognition models. Our work, SkeFi, overcomes these gaps through a novel cross‑modal knowledge transfer method acquired from the data‑rich RGB modality. We propose the enhanced Temporal Correlation Adaptive Graph Convolution (TC‑AGC) with frame interactive enhancement to overcome the noise from missing or inconsecutive frames. Additionally, our research underscores the effectiveness of enhancing multiscale temporal modeling through dual temporal convolution. By integrating TC‑AGC with temporal modeling for cross‑modal transfer, our framework can extract accurate poses and actions from noisy wireless sensors. Experiments demonstrate that SkeFi realizes state‑of‑the‑art performances on mmWave and LiDAR. The code is available at https://github.com/Huang0035/Skefi.
Authors:Jiahui Sheng, Yidan Shi, Shu Xiang, Xiaorun Li, Shuhan Chen
Abstract:
Hyperspectral images (HSIs) are a type of image that contains abundant spectral information. As a type of real‑world data, the high‑dimensional spectra in hyperspectral images are actually determined by only a few factors, such as chemical composition and illumination. Thus, spectra in hyperspectral images are highly likely to satisfy the manifold hypothesis. Based on the hyperspectral manifold hypothesis, we propose a novel hyperspectral anomaly detection method (named ScoreAD) that leverages the time‑dependent gradient field of the data distribution (i.e., the score), as learned by a score‑based generative model (SGM). Our method first trains the SGM on the entire set of spectra from the hyperspectral image. At test time, each spectrum is passed through a perturbation kernel, and the resulting perturbed spectrum is fed into the trained SGM to obtain the estimated score. The manifold hypothesis of HSIs posits that background spectra reside on one or more low‑dimensional manifolds. Conversely, anomalous spectra, owing to their unique spectral signatures, are considered outliers that do not conform to the background manifold. Based on this fundamental discrepancy in their manifold distributions, we leverage a generative SGM to achieve hyperspectral anomaly detection. Experiments on the four hyperspectral datasets demonstrate the effectiveness of the proposed method. The code is available at https://github.com/jiahuisheng/ScoreAD.
Authors:Hailing Jin, Huiying Li
Abstract:
Recent advances in semantic correspondence have been largely driven by the use of pre‑trained large‑scale models. However, a limitation of these approaches is their dependence on high‑resolution input images to achieve optimal performance, which results in considerable computational overhead. In this work, we address a fundamental limitation in current methods: the irreversible fusion of adjacent keypoint features caused by deep downsampling operations. This issue is triggered when semantically distinct keypoints fall within the same downsampled receptive field (e.g., 16x16 patches). To address this issue, we present SimpleMatch, a simple yet effective framework for semantic correspondence that delivers strong performance even at low resolutions. We propose a lightweight upsample decoder that progressively recovers spatial detail by upsampling deep features to 1/4 resolution, and a multi‑scale supervised loss that ensures the upsampled features retain discriminative features across different spatial scales. In addition, we introduce sparse matching and window‑based localization to optimize training memory usage and reduce it by 51%. At a resolution of 252x252 (3.3x smaller than current SOTA methods), SimpleMatch achieves superior performance with 84.1% PCK@0.1 on the SPair‑71k benchmark. We believe this framework provides a practical and efficient baseline for future research in semantic correspondence. Code is available at: https://github.com/hailong23‑jin/SimpleMatch.
Authors:Jiahui Sheng, Xiaorun Li, Shuhan Chen
Abstract:
As a key task in hyperspectral image processing, hyperspectral anomaly detection has garnered significant attention and undergone extensive research. Existing methods primarily relt on two prior assumption: low‑rank background and sparse anomaly, along with additional spatial assumptions of the background. However, most methods only utilize the sparsity prior assumption for anomalies and rarely expand on this hypothesis. From observations of hyperspectral images, we find that anomalous pixels exhibit certain spatial distribution characteristics: they often manifest as small, clustered groups in space, which we refer to as cluster sparsity of anomalies. Then, we combined the cluster sparsity prior with the classical GoDec algorithm, incorporating the cluster sparsity prior into the S‑step of GoDec. This resulted in a new hyperspectral anomaly detection method, which we called Turbo‑GoDec. In this approach, we modeled the cluster sparsity prior of anomalies using a Markov random field and computed the marginal probabilities of anomalies through message passing on a factor graph. Locations with high anomalous probabilities were treated as the sparse component in the Turbo‑GoDec. Experiments are conducted on three real hyperspectral image (HSI) datasets which demonstrate the superior performance of the proposed Turbo‑GoDec method in detecting small‑size anomalies comparing with the vanilla GoDec (LSMAD) and state‑of‑the‑art anomaly detection methods. The code is available at https://github.com/jiahuisheng/Turbo‑GoDec.
Authors:Jianhao Jiao, Changkun Liu, Jingwen Yu, Boyi Liu, Qianyi Zhang, Yue Wang, Dimitrios Kanoulas
Abstract:
Scalable and maintainable map representations are fundamental to enabling large‑scale visual navigation and facilitating the deployment of robots in real‑world environments. While collaborative localization across multi‑session mapping enhances efficiency, traditional structure‑based methods struggle with high maintenance costs and fail in feature‑less environments or under significant viewpoint changes typical of crowd‑sourced data. To address this, we propose OPENNAVMAP, a lightweight, structure‑free topometric system leveraging 3D geometric foundation models for on‑demand reconstruction. Our method unifies dynamic programming‑based sequence matching, geometric verification, and confidence‑calibrated optimization to robust, coarse‑to‑fine submap alignment without requiring pre‑built 3D models. Evaluations on the Map‑Free benchmark demonstrate superior accuracy over structure‑from‑motion and regression baselines, achieving an average translation error of 0.62m. Furthermore, the system maintains global consistency across 15km of multi‑session data with an absolute trajectory error below 3m for map merging. Finally, we validate practical utility through 12 successful autonomous image‑goal navigation tasks on simulated and physical robots. Code and datasets will be publicly available in https://rpl‑cs‑ucl.github.io/OpenNavMap_page.
Authors:Chunyang Fu, Ge Li, Wei Gao, Shiqi Wang, Zhu Li, Shan Liu
Abstract:
Recently, deep learning has significantly advanced the performance of point cloud geometry compression. However, the learning‑based lossless attribute compression of point clouds with varying densities is under‑explored. In this paper, we develop a learning‑based framework, namely DALD‑PCAC that leverages Levels of Detail (LoD) to tailor for point cloud lossless attribute compression. We develop a point‑wise attention model using a permutation‑invariant Transformer to tackle the challenges of sparsity and irregularity of point clouds during context modeling. We also propose a Density‑Adaptive Learning Descriptor (DALD) capable of capturing structure and correlations among points across a large range of neighbors. In addition, we develop a prior‑guided block partitioning to reduce the attribute variance within blocks and enhance the performance. Experiments on LiDAR and object point clouds show that DALD‑PCAC achieves the state‑of‑the‑art performance on most data. Our method boosts the compression performance and is robust to the varying densities of point clouds. Moreover, it guarantees a good trade‑off between performance and complexity, exhibiting great potential in real‑world applications. The source code is available at https://github.com/zb12138/DALD_PCAC.
Authors:Chunyang Fu, Tai Qin, Shiqi Wang, Zhu Li
Abstract:
Regional Adaptive Hierarchical Transform (RAHT) is an effective point cloud attribute compression (PCAC) method. However, its application in deep learning lacks research. In this paper, we propose an end‑to‑end RAHT framework for lossy PCAC based on the sparse tensor, called DeepRAHT. The RAHT transform is performed within the learning reconstruction process, without requiring manual RAHT for preprocessing. We also introduce the predictive RAHT to reduce bitrates and design a learning‑based prediction model to enhance performance. Moreover, we devise a bitrate proxy that applies run‑length coding to entropy model, achieving seamless variable‑rate coding and improving robustness. DeepRAHT is a reversible and distortion‑controllable framework, ensuring its lower bound performance and offering significant application potential. The experiments demonstrate that DeepRAHT is a high‑performance, faster, and more robust solution than the baseline methods. Project Page: https://github.com/zb12138/DeepRAHT.
Authors:Anzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin, Mingxi Cheng, Shahin Nazarian, Paul Thompson, Paul Bogdan
Abstract:
The relentless scaling of deep learning models has led to unsustainable computational demands, positioning Mixture‑of‑Experts (MoE) architectures as a promising path towards greater efficiency. However, MoE models are plagued by two fundamental challenges: 1) a load imbalance problem known as the``rich get richer" phenomenon, where a few experts are over‑utilized, and 2) an expert homogeneity problem, where experts learn redundant representations, negating their purpose. Current solutions typically employ an auxiliary load‑balancing loss that, while mitigating imbalance, often exacerbates homogeneity by enforcing uniform routing at the expense of specialization. To resolve this, we introduce the Eigen‑Mixture‑of‑Experts (EMoE), a novel architecture that leverages a routing mechanism based on a learned orthonormal eigenbasis. EMoE projects input tokens onto this shared eigenbasis and routes them based on their alignment with the principal components of the feature space. This principled, geometric partitioning of data intrinsically promotes both balanced expert utilization and the development of diverse, specialized experts, all without the need for a conflicting auxiliary loss function. Our code is publicly available at https://github.com/Belis0811/EMoE.
Authors:Xiaotong Zhou, Zhenhui Yuan, Yi Han, Tianhua Xu, Laurence T. Yang
Abstract:
Accurate trajectory prediction of vehicles at roundabouts is critical for reducing traffic accidents, yet it remains highly challenging due to their circular road geometry, continuous merging and yielding interactions, and absence of traffic signals. Developing accurate prediction algorithms relies on reliable, multimodal, and realistic datasets; however, such datasets for roundabout scenarios are scarce, as real‑world data collection is often limited by incomplete observations and entangled factors that are difficult to isolate. We present CARLA‑Round, a systematically designed simulation dataset for roundabout trajectory prediction. The dataset varies weather conditions (five types) and traffic density levels (spanning Level‑of‑Service A‑E) in a structured manner, resulting in 25 controlled scenarios. Each scenario incorporates realistic mixtures of driving behaviors and provides explicit annotations that are largely absent from existing datasets. Unlike randomly sampled simulation data, this structured design enables precise analysis of how different conditions influence trajectory prediction performance. Validation experiments using standard baselines (LSTM, GCN, GRU+GCN) reveal traffic density dominates prediction difficulty with strong monotonic effects, while weather shows non‑linear impacts. The best model achieves 0.312m ADE on real‑world rounD dataset, demonstrating effective sim‑to‑real transfer. This systematic approach quantifies factor impacts impossible to isolate in confounded real‑world datasets. Our CARLA‑Round dataset is available at https://github.com/Rebecca689/CARLA‑Round.
Authors:Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Saïd Mahmoudi, Benoît Macq, Christophe De Vleeschouwer
Abstract:
Assisting pathologists in the analysis of histopathological images has high clinical value, as it supports cancer detection and staging. In this context, histology foundation models have recently emerged. Among them, Vision‑Language Models (VLMs) provide strong yet imperfect zero‑shot predictions. We propose to refine these predictions by adapting Conditional Random Fields (CRFs) to histopathological applications, requiring no additional model training. We present HistoCRF, a CRF‑based framework, with a novel definition of the pairwise potential that promotes label diversity and leverages expert annotations. We consider three experiments: without annotations, with expert annotations, and with iterative human‑in‑the‑loop annotations that progressively correct misclassified patches. Experiments on five patch‑level classification datasets covering different organs and diseases demonstrate average accuracy gains of 16.0% without annotations and 27.5% with only 100 annotations, compared to zero‑shot predictions. Moreover, integrating a human in the loop reaches a further gain of 32.6% with the same number of annotations. The code will be made available on https://github.com/tgodelaine/HistoCRF.
Authors:Jing Zhang, Bingjie Fan, Jixiang Zhu, Zhe Wang
Abstract:
We propose EmoLat, a novel emotion latent space that enables fine‑grained, text‑driven image sentiment transfer by modeling cross‑modal correlations between textual semantics and visual emotion features. Within EmoLat, an emotion semantic graph is constructed to capture the relational structure among emotions, objects, and visual attributes. To enhance the discriminability and transferability of emotion representations, we employ adversarial regularization, aligning the latent emotion distributions across modalities. Building upon EmoLat, a cross‑modal sentiment transfer framework is proposed to manipulate image sentiment via joint embedding of text and EmoLat features. The network is optimized using a multi‑objective loss incorporating semantic consistency, emotion alignment, and adversarial regularization. To support effective modeling, we construct EmoSpace Set, a large‑scale benchmark dataset comprising images with dense annotations on emotions, object semantics, and visual attributes. Extensive experiments on EmoSpace Set demonstrate that our approach significantly outperforms existing state‑of‑the‑art methods in both quantitative metrics and qualitative transfer fidelity, establishing a new paradigm for controllable image sentiment editing guided by textual input. The EmoSpace Set and all the code are available at http://github.com/JingVIPLab/EmoLat.
Authors:Zijie Lou, Xiangwei Feng, Jiaxin Wang, Jiangtao Yao, Fei Che, Tianbao Liu, Chengjing Wu, Xiaochao Qu, Luoqi Liu, Ting Liu
Abstract:
Existing video object removal methods predominantly rely on diffusion models following a noise‑to‑data paradigm, where generation starts from uninformative Gaussian noise. This approach discards the rich structural and contextual priors present in the original input video. Consequently, such methods often lack sufficient guidance, leading to incomplete object erasure or the synthesis of implausible content that conflicts with the scene's physical logic. In this paper, we reformulate video object removal as a video‑to‑video translation task via a stochastic bridge model. Unlike noise‑initialized methods, our framework establishes a direct stochastic path from the source video (with objects) to the target video (objects removed). This bridge formulation effectively leverages the input video as a strong structural prior, guiding the model to perform precise removal while ensuring that the filled regions are logically consistent with the surrounding environment. To address the trade‑off where strong bridge priors hinder the removal of large objects, we propose a novel adaptive mask modulation strategy. This mechanism dynamically modulates input embeddings based on mask characteristics, balancing background fidelity with generative flexibility. Extensive experiments demonstrate that our approach significantly outperforms existing methods in both visual quality and temporal consistency. The project page is https://bridgeremoval.github.io/.
Authors:Weixin Ye, Wei Wang, Yahui Liu, Yue Song, Bin Ren, Wei Bi, Rita Cucchiara, Nicu Sebe
Abstract:
In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient information, which can be used to reconstruct the input data. To mitigate this issue, we introduce a Masked Jigsaw Puzzle (MJP) framework. MJP starts with random token shuffling to break the token order, and then a learnable unknown (unk) position embedding is used to mask out the PEs of the shuffled tokens. In this manner, the local spatial information which is encoded in the position embeddings is disrupted, and the models are forced to learn feature representations that are less reliant on the local spatial information. Notably, with the careful use of MJP, we can not only improve models' robustness against gradient attacks, but also boost their performance in both vision and text application scenarios, such as classification for images (e.g., ImageNet‑1K) and sentiment analysis for text (e.g., Yelp and Amazon). Experimental results suggest that MJP is a unified framework for different Transformer‑based models in both vision and language tasks. Code is publicly available via https://github.com/ywxsuperstar/transformerattack
Authors:Jian Lang, Rongpei Hong, Ting Zhong, Yong Wang, Fan Zhou
Abstract:
Fake News Video Detection (FNVD) is critical for social stability. Existing methods typically assume consistent news topic distribution between training and test phases, failing to detect fake news videos tied to emerging events and unseen topics. To bridge this gap, we introduce RADAR, the first framework that enables test‑time adaptation to unseen news videos. RADAR pioneers a new retrieval‑guided adaptation paradigm that leverages stable (source‑close) videos from the target domain to guide robust adaptation of semantically related but unstable instances. Specifically, we propose an Entropy Selection‑Based Retrieval mechanism that provides videos with stable (low‑entropy), relevant references for adaptation. We also introduce a Stable Anchor‑Guided Alignment module that explicitly aligns unstable instances' representations to the source domain via distribution‑level matching with their stable references, mitigating severe domain discrepancies. Finally, our novel Target‑Domain Aware Self‑Training paradigm can generate informative pseudo‑labels augmented by stable references, capturing varying and imbalanced category distributions in the target domain and enabling RADAR to adapt to the fast‑changing label distributions. Extensive experiments demonstrate that RADAR achieves superior performance for test‑time FNVD, enabling strong on‑the‑fly adaptation to unseen fake news video topics.
Authors:Zongmin Li, Yachuan Li, Lei Kang, Dimosthenis Karatzas, Wenkang Ma
Abstract:
Multi‑page Document Visual Question Answering (MP‑DocVQA) remains challenging because long documents not only strain computational resources but also reduce the effectiveness of the attention mechanism in large vision‑language models (LVLMs). We tackle these issues with an Adaptive Visual In‑document Retrieval (AVIR) framework. A lightweight retrieval model first scores each page for question relevance. Pages are then clustered according to the score distribution to adaptively select relevant content. The clustered pages are screened again by Top‑K to keep the context compact. However, for short documents, clustering reliability decreases, so we use a relevance probability threshold to select pages. The selected pages alone are fed to a frozen LVLM for answer generation, eliminating the need for model fine‑tuning. The proposed AVIR framework reduces the average page count required for question answering by 70%, while achieving an ANLS of 84.58% on the MP‑DocVQA dataset‑surpassing previous methods with significantly lower computational cost. The effectiveness of the proposed AVIR is also verified on the SlideVQA and DUDE benchmarks. The code is available at https://github.com/Li‑yachuan/AVIR.
Authors:Lexin Ren, Jiamiao Lu, Weichuan Zhang, Benqing Wu, Tuo Wang, Yi Liao, Jiapan Guo, Changming Sun, Liang Guo
Abstract:
Preterm infants (born between 28 and 37 weeks of gestation) face elevated risks of neurodevelopmental delays, making early identification crucial for timely intervention. While deep learning‑based volumetric segmentation of brain MRI scans offers a promising avenue for assessing neonatal neurodevelopment, achieving accurate segmentation of white matter (WM) and gray matter (GM) in preterm infants remains challenging due to their comparable signal intensities (isointense appearance) on MRI during early brain development. To address this, we propose a novel segmentation neural network, named Hierarchical Dense Attention Network. Our architecture incorporates a 3D spatial‑channel attention mechanism combined with an attention‑guided dense upsampling strategy to enhance feature discrimination in low‑contrast volumetric data. Quantitative experiments demonstrate that our method achieves superior segmentation performance compared to state‑of‑the‑art baselines, effectively tackling the challenge of isointense tissue differentiation. Furthermore, application of our algorithm confirms that WM and GM volumes in preterm infants are significantly lower than those in term infants, providing additional imaging evidence of the neurodevelopmental delays associated with preterm birth. The code is available at: https://github.com/ICL‑SUST/HDAN.
Authors:Zhengxian Wu, Chuanrui Zhang, Shenao Jiang, Hangrui Xu, Zirui Liao, Luyuan Zhang, Huaqiu Li, Peng Jiao, Haoqian Wang
Abstract:
Gait recognition is emerging as a promising technology and an innovative field within computer vision, with a wide range of applications in remote human identification. However, existing methods typically rely on complex architectures to directly extract features from images and apply pooling operations to obtain sequence‑level representations. Such designs often lead to overfitting on static noise (e.g., clothing), while failing to effectively capture dynamic motion regions, such as the arms and legs. This bottleneck is particularly challenging in the presence of intra‑class variation, where gait features of the same individual under different environmental conditions are significantly distant in the feature space. To address the above challenges, we present a Languageguided and Motion‑aware gait recognition framework, named LMGait. To the best of our knowledge, LMGait is the first method to introduce natural language descriptions as explicit semantic priors into the gait recognition task. In particular, we utilize designed gait‑related language cues to capture key motion features in gait sequences. To improve cross‑modal alignment, we propose the Motion Awareness Module (MAM), which refines the language features by adaptively adjusting various levels of semantic information to ensure better alignment with the visual representations. Furthermore, we introduce the Motion Temporal Capture Module (MTCM) to enhance the discriminative capability of gait features and improve the model's motion tracking ability. We conducted extensive experiments across multiple datasets, and the results demonstrate the significant advantages of our proposed network. Specifically, our model achieved accuracies of 88.5%, 97.1%, and 97.5% on the CCPG, SUSTech1K, and CASIAB datasets, respectively, achieving state‑of‑the‑art performance. Homepage: https://dingwu1021.github.io/LMGait/
Authors:Xulei Shi, Maoyu Wang, Yuning Peng, Guanbo Wang, Xin Wang, Yifan Liao, Qi Chen, Pengjie Tao
Abstract:
Image retrieval is a critical step for reducing the quadratic cost of image matching in unconstrained Structure‑from‑Motion (SfM). Unlike generic image retrieval, however, the relevant goal of SfM is to identify geometrically matchable image pairs rather than merely semantically similar images. Prevailing methods are largely trained under anchor‑centric tuple guidance, which organizes the training around isolated tuples and under‑utilizes the dense, graded overlap structure naturally established within a SfM scene. In this work, we present SupScene, a scene‑structured training framework that samples connected local subgraphs from SfM overlap graphs and jointly supervises all valid within‑subgraph pairwise relations. To explicitly align the trained descriptor with geometric co‑visibility, we further introduce an overlap‑ordered objective that combines multi‑similarity optimization with a continuous relative‑overlap ranking term. In addition, the proposed framework is instantiated with a lightweight Structural Context Probe Pooling (SCPP) head that aggregates complementary structural responses into a compact global descriptor. Extensive experimental results on multiple benchmarks demonstrate that our method can significantly improve overall retrieval performance and enhance the completeness of downstream SfM reconstructions. Code and models are available at https://github.com/Suxilan/SupScene.
Authors:Yilmaz Korkmaz, Vishal M. Patel
Abstract:
Remote sensing change detection aims to localize and characterize scene changes between two time points and is central to applications such as environmental monitoring and disaster assessment. Meanwhile, visual autoregressive models (VARs) have recently shown impressive image generation capability, but their adoption for pixel‑level discriminative tasks remains limited due to weak controllability, suboptimal dense prediction performance and exposure bias. We introduce RemoteVAR, a new VAR‑based change detection framework that addresses these limitations by conditioning autoregressive prediction on multi‑resolution fused bi‑temporal features via cross‑attention, and by employing an autoregressive training strategy designed specifically for change map prediction. Extensive experiments on standard change detection benchmarks show that RemoteVAR delivers consistent and significant improvements over strong diffusion‑based and transformer‑based baselines, establishing a competitive autoregressive alternative for remote sensing change detection. Code will be available \hrefhttps://github.com/yilmazkorkmaz1/RemoteVAR\underlinehere.
Authors:Kaustubh Shivshankar Shejole, Gaurav Mishra
Abstract:
Interactive graph‑based segmentation methods partition an image into foreground and background regions with the aid of user inputs. However, existing approaches often suffer from high computational costs, sensitivity to user interactions, and degraded performance when the foreground and background share similar color distributions. A key factor influencing segmentation performance is the similarity measure used for assigning edge weights in the graph. To address these challenges, we propose a novel Pixel Segment Similarity Index (PSSI), which leverages the harmonic mean of inter‑channel similarities by incorporating both pixel intensity and spatial smoothness features. The harmonic mean effectively penalizes dissimilarities in any individual channel, enhancing robustness. The computational complexity of PSSI is \mathcalO(B), where B denotes the number of histogram bins. Our segmentation framework begins with low‑level segmentation using MeanShift, which effectively captures color, texture, and segment shape. Based on the resulting pixel segments, we construct a pixel‑segment graph with edge weights determined by PSSI. For partitioning, we employ the Maximum Spanning Tree (MaxST), which captures strongly connected local neighborhoods beneficial for precise segmentation. The integration of the proposed PSSI, MeanShift, and MaxST allows our method to jointly capture color similarity, smoothness, texture, shape, and strong local connectivity. Experimental evaluations on the GrabCut and Images250 datasets demonstrate that our method consistently outperforms current graph‑based interactive segmentation methods such as AMOE, OneCut, and SSNCut in terms of segmentation quality, as measured by Jaccard Index (IoU), F_1 score, execution time and Mean Error (ME). Code is publicly available at: https://github.com/KaustubhShejole/PSSI‑MaxST.
Authors:Xuchen Li, Xuzhao Li, Renjie Pi, Shiyu Hu, Jian Zhao, Jiahui Gao
Abstract:
Despite the remarkable progress of Vision‑Language Models (VLMs) in adopting "Thinking‑with‑Images" capabilities, accurately evaluating the authenticity of their reasoning process remains a critical challenge. Existing benchmarks mainly rely on outcome‑oriented accuracy, lacking the capability to assess whether models can accurately leverage fine‑grained visual cues for multi‑step reasoning. To address these limitations, we propose ViEBench, a process‑verifiable benchmark designed to evaluate faithful visual reasoning. Comprising 200 multi‑scenario high‑resolution images with expert‑annotated visual evidence, ViEBench uniquely categorizes tasks by difficulty into perception and reasoning dimensions, where reasoning tasks require utilizing localized visual details with prior knowledge. To establish comprehensive evaluation criteria, we introduce a dual‑axis matrix that provides fine‑grained metrics through four diagnostic quadrants, enabling transparent diagnosis of model behavior across varying task complexities. Our experiments yield several interesting observations: (1) VLMs can sometimes produce correct final answers despite grounding on irrelevant regions, and (2) they may successfully locate the correct evidence but still fail to utilize it to reach accurate conclusions. Our findings demonstrate that ViEBench can serve as a more explainable and practical benchmark for comprehensively evaluating the effectiveness agentic VLMs. The codes will be released at: https://github.com/Xuchen‑Li/ViEBench.
Authors:Arnav S. Sonavane
Abstract:
We investigate the impact of domain‑specific self‑supervised pre‑training on agricultural disease classification using hierarchical vision transformers. Our key finding is that SimCLR pre‑training on just 3,000 unlabeled agricultural images provides a +4.57% accuracy improvement‑‑exceeding the +3.70% gain from hierarchical architecture design. Critically, we show this SSL benefit is architecture‑agnostic: applying the same pre‑training to Swin‑Base yields +4.08%, to ViT‑Base +4.20%, confirming practitioners should prioritize domain data collection over architectural choices. Using HierarchicalViT (HVT), a Swin‑style hierarchical transformer, we evaluate on three datasets: Cotton Leaf Disease (7 classes, 90.24%), PlantVillage (38 classes, 96.3%), and PlantDoc (27 classes, 87.1%). At matched parameter counts, HVT‑Base (78M) achieves 88.91% vs. Swin‑Base (88M) at 87.23%, a +1.68% improvement. For deployment reliability, we report calibration analysis showing HVT achieves 3.56% ECE (1.52% after temperature scaling). Code: https://github.com/w2sg‑arnav/HierarchicalViT
Authors:Ruiheng Zhang, Jingfeng Yao, Huangxuan Zhao, Hao Yan, Xiao He, Lei Chen, Zhou Wei, Yong Luo, Zengmao Wang, Lefei Zhang, Dacheng Tao, Bo Du
Abstract:
Despite recent progress, medical foundation models still struggle to unify visual understanding and generation, as these tasks have inherently conflicting goals: semantic abstraction versus pixel‑level reconstruction. Existing approaches, typically based on parameter‑shared autoregressive architectures, frequently lead to compromised performance in one or both tasks. To address this, we present UniX, a next‑generation unified medical foundation model for chest X‑ray understanding and generation. UniX decouples the two tasks into an autoregressive branch for understanding and a diffusion branch for high‑fidelity generation. Crucially, a cross‑modal self‑attention mechanism is introduced to dynamically guide the generation process with understanding features. Coupled with a rigorous data cleaning pipeline and a multi‑stage training strategy, this architecture enables synergistic collaboration between tasks while leveraging the strengths of diffusion models for superior generation. On two representative benchmarks, UniX achieves a 46.1% improvement in understanding performance (Micro‑F1) and a 24.2% gain in generation quality (FD‑RadDino), using only a quarter of the parameters of LLM‑CXR. By achieving performance on par with task‑specific models, our work establishes a scalable paradigm for synergistic medical image understanding and generation. Codes and models are available at https://github.com/ZrH42/UniX.
Authors:Yawar Siddiqui, Duncan Frost, Samir Aroudj, Armen Avetisyan, Henry Howard-Jenkins, Daniel DeTone, Pierre Moulon, Qirui Wu, Zhengqin Li, Julian Straub, Richard Newcombe, Jakob Engel
Abstract:
Recent advances in 3D shape generation have achieved impressive results, but most existing methods rely on clean, unoccluded, and well‑segmented inputs. Such conditions are rarely met in real‑world scenarios. We present ShapeR, a novel approach for conditional 3D object shape generation from casually captured sequences. Given an image sequence, we leverage off‑the‑shelf visual‑inertial SLAM, 3D detection algorithms, and vision‑language models to extract, for each object, a set of sparse SLAM points, posed multi‑view images, and machine‑generated captions. A rectified flow transformer trained to effectively condition on these modalities then generates high‑fidelity metric 3D shapes. To ensure robustness to the challenges of casually captured data, we employ a range of techniques including on‑the‑fly compositional augmentations, a curriculum training scheme spanning object‑ and scene‑level datasets, and strategies to handle background clutter. Additionally, we introduce a new evaluation benchmark comprising 178 in‑the‑wild objects across 7 real‑world scenes with geometry annotations. Experiments show that ShapeR significantly outperforms existing approaches in this challenging setting, achieving an improvement of 2.7x in Chamfer distance compared to state of the art.
Authors:Oishee Bintey Hoque, Nibir Chandra Mandal, Kyle Luong, Amanda Wilson, Samarth Swarup, Madhav Marathe, Abhijin Adiga
Abstract:
Large‑scale livestock operations pose significant risks to human health and the environment, while also being vulnerable to threats such as infectious diseases and extreme weather events. As the number of such operations continues to grow, accurate and scalable mapping has become increasingly important. In this work, we present an infrastructure‑first, explainable pipeline for identifying and characterizing Concentrated Animal Feeding Operations (CAFOs) from aerial and satellite imagery. Our method (i) detects candidate infrastructure (e.g., barns, feedlots, manure lagoons, silos) with a domain‑tuned YOLOv8 detector, then derives SAM2 masks from these boxes and filters component‑specific criteria; (ii) extracts structured descriptors (e.g., counts, areas, orientations, and spatial relations) and fuses them with deep visual features using a lightweight spatial cross‑attention classifier; and (iii) outputs both CAFO type predictions and mask‑level attributions that link decisions to visible infrastructure. Through comprehensive evaluation, we show that our approach achieves state‑of‑the‑art performance, with Swin‑B+PRISM‑CAFO surpassing the best performing baseline by up to 15%. Beyond strong predictive performance across diverse U.S. regions, we run systematic gradient‑‑activation analyses that quantify the impact of domain priors and show how specific infrastructure (e.g., barns, lagoons) shapes classification decisions. We release code, infrastructure masks, and descriptors to support transparent, scalable monitoring of livestock infrastructure, enabling risk modeling, change detection, and targeted regulatory action.
Github: https://github.com/Nibir088/PRISM‑CAFO.
Authors:Raphaël Razafindralambo, Rémy Sun, Frédéric Precioso, Damien Garreau, Pierre-Alexandre Mattei
Abstract:
Diffusion models now generate high‑quality, diverse samples, with an increasing focus on more powerful models. Although ensembling is a well‑known way to improve supervised models, its application to unconditional score‑based diffusion models remains largely unexplored. In this work we investigate whether it provides tangible benefits for generative modelling. We find that while ensembling the scores generally improves the score‑matching loss and model likelihood, it fails to consistently enhance perceptual quality metrics such as FID on image datasets. We confirm this observation across a breadth of aggregation rules using Deep Ensembles, Monte Carlo Dropout, on CIFAR‑10 and FFHQ. We attempt to explain this discrepancy by investigating possible explanations, such as the link between score estimation and image quality. We also look into tabular data through random forests, and find that one aggregation strategy outperforms the others. Finally, we provide theoretical insights into the summing of score models, which shed light not only on ensembling but also on several model composition techniques (e.g. guidance).
Authors:Mark Eastwood, Thomas McKee, Zedong Hu, Sabine Tejpar, Fayyaz Minhas
Abstract:
Separating the contributions of individual chromogenic stains in RGB histology whole slide images (WSIs) is essential for stain normalization, quantitative assessment of marker expression, and cell‑level readouts in immunohistochemistry (IHC). Classical Beer‑Lambert (BL) color deconvolution is well‑established for two‑ or three‑stain settings, but becomes under‑determined and unstable for multiplex IHC (mIHC) with K>3 chromogens. We present a simple, data‑driven encoder‑decoder architecture that learns cohort‑specific stain characteristics for mIHC RGB WSIs and yields crisp, well‑separated per‑stain concentration maps. The encoder is a compact U‑Net that predicts K nonnegative concentration channels; the decoder is a differentiable BL forward model with a learnable stain matrix initialized from typical chromogen hues. Training is unsupervised with a perceptual reconstruction objective augmented by loss terms that discourage unnecessary stain mixing. On a colorectal mIHC panel comprising 5 stains (H, CDX2, MUC2, MUC5, CD8) we show excellent RGB reconstruction, and significantly reduced inter‑channel bleed‑through compared with matrix‑based deconvolution. Code and model are available at https://github.com/measty/StainQuant.git.
Authors:Cheng-Zhuang Liu, Si-Bao Chen, Qing-Ling Shu, Chris Ding, Jin Tang, Bin Luo
Abstract:
Recent advances in video anomaly detection (VAD) mainly focus on ground‑based surveillance or unmanned aerial vehicle (UAV) videos with static backgrounds, whereas research on UAV videos with dynamic backgrounds remains limited. Unlike static scenarios, dynamically captured UAV videos exhibit multi‑source motion coupling, where the motion of objects and UAV‑induced global motion are intricately intertwined. Consequently, existing methods may misclassify normal UAV movements as anomalies or fail to capture true anomalies concealed within dynamic backgrounds. Moreover, many approaches do not adequately address the joint modeling of inter‑frame continuity and local spatial correlations across diverse temporal scales. To overcome these limitations, we propose the Frequency‑Assisted Temporal Dilation Mamba (FTDMamba) network for UAV VAD, including two core components: (1) a Frequency Decoupled Spatiotemporal Correlation Module, which disentangles coupled motion patterns and models global spatiotemporal dependencies through frequency analysis; and (2) a Temporal Dilation Mamba Module, which leverages Mamba's sequence modeling capability to jointly learn fine‑grained temporal dynamics and local spatial structures across multiple temporal receptive fields. Additionally, unlike existing UAV VAD datasets which focus on static backgrounds, we construct a large‑scale Moving UAV VAD dataset (MUVAD), comprising 222,736 frames with 240 anomaly events across 12 anomaly types. Extensive experiments demonstrate that FTDMamba achieves state‑of‑the‑art (SOTA) performance on two public static benchmarks and the new MUVAD dataset. The code and MUVAD dataset will be available at: https://github.com/uavano/FTDMamba.
Authors:Ana Davila, Jacinto Colan, Yasuhisa Hasegawa
Abstract:
Deep learning has significantly advanced image analysis across diverse domains but often depends on large, annotated datasets for success. Transfer learning addresses this challenge by utilizing pre‑trained models to tackle new tasks with limited labeled data. However, discrepancies between source and target domains can hinder effective transfer learning. We introduce BioTune, a novel adaptive fine‑tuning technique utilizing evolutionary optimization. BioTune enhances transfer learning by optimally choosing which layers to freeze and adjusting learning rates for unfrozen layers. Through extensive evaluation on nine image classification datasets, spanning natural and specialized domains such as medical imaging, BioTune demonstrates superior accuracy and efficiency over state‑of‑the‑art fine‑tuning methods, including AutoRGN and LoRA, highlighting its adaptability to various data characteristics and distribution changes. Additionally, BioTune consistently achieves top performance across four different CNN architectures, underscoring its flexibility. Ablation studies provide valuable insights into the impact of BioTune's key components on overall performance. The source code is available at https://github.com/davilac/BioTune.
Authors:Pascal Schlachter, Bin Yang
Abstract:
Unsupervised domain adaptation tackles the problem that domain shifts between training and test data impair the performance of neural networks in many real‑world applications. Thereby, in realistic scenarios, the source data may no longer be available during adaptation, and the label space of the target domain may differ from the source label space. This setting, known as source‑free universal domain adaptation (SF‑UniDA), has recently gained attention, but all existing approaches only assume a single domain shift from source to target. In this work, we present the first study on continual SF‑UniDA, where the model must adapt sequentially to a stream of multiple different unlabeled target domains. Building upon our previous methods for online SF‑UniDA, we combine their key ideas by integrating Gaussian mixture model‑based pseudo‑labeling within a mean teacher framework for improved stability over long adaptation sequences. Additionally, we introduce consistency losses for further robustness. The resulting method GMM‑COMET provides a strong first baseline for continual SF‑UniDA and is the only approach in our experiments to consistently improve upon the source‑only model across all evaluated scenarios. Our code is available at https://github.com/pascalschlachter/GMM‑COMET.
Authors:Shaofeng Yin, Jiaxin Ge, Zora Zhiruo Wang, Chenyang Wang, Xiuyu Li, Michael J. Black, Trevor Darrell, Angjoo Kanazawa, Haiwen Feng
Abstract:
Vision‑as‑inverse‑graphics, the concept of reconstructing images into editable programs, remains challenging for Vision‑Language Models (VLMs), which inherently lack fine‑grained spatial grounding in one‑shot settings. To address this, we introduce VIGA (Vision‑as‑Inverse‑Graphics Agent), an interleaved multimodal reasoning framework where symbolic logic and visual perception actively cross‑verify each other. VIGA operates through a tightly coupled code‑render‑inspect loop: synthesizing symbolic programs, projecting them into visual states, and inspecting discrepancies to guide iterative edits. Equipped with high‑level semantic skills and an evolving multimodal memory, VIGA sustains evidence‑based modifications over long horizons. This training‑free, task‑agnostic framework seamlessly supports 2D document generation, 3D reconstruction, multi‑step 3D editing, and 4D physical interaction. Finally, we introduce BlenderBench, a challenging visual‑to‑code benchmark. Empirically, VIGA substantially improves accuracy compared with one‑shot baselines in BlenderGym (35.32%), SlideBench (117.17%) and our proposed BlenderBench (124.70%).
Authors:Shuai Tan, Biao Gong, Ke Ma, Yutong Feng, Qiyuan Zhang, Yan Wang, Yujun Shen, Hengshuang Zhao
Abstract:
Character image animation is gaining significant importance across various domains, driven by the demand for robust and flexible multi‑subject rendering. While existing methods excel in single‑person animation, they struggle to handle arbitrary subject counts, diverse character types, and spatial misalignment between the reference image and the driving poses. We attribute these limitations to an overly rigid spatial binding that forces strict pixel‑wise alignment between the pose and reference, and an inability to consistently rebind motion to intended subjects. To address these challenges, we propose CoDance, a novel Unbind‑Rebind framework that enables the animation of arbitrary subject counts, types, and spatial configurations conditioned on a single, potentially misaligned pose sequence. Specifically, the Unbind module employs a novel pose shift encoder to break the rigid spatial binding between the pose and the reference by introducing stochastic perturbations to both poses and their latent features, thereby compelling the model to learn a location‑agnostic motion representation. To ensure precise control and subject association, we then devise a Rebind module, leveraging semantic guidance from text prompts and spatial guidance from subject masks to direct the learned motion to intended characters. Furthermore, to facilitate comprehensive evaluation, we introduce a new multi‑subject CoDanceBench. Extensive experiments on CoDanceBench and existing datasets show that CoDance achieves SOTA performance, exhibiting remarkable generalization across diverse subjects and spatial layouts. The code and weights will be open‑sourced.
Authors:Takuya Murakawa, Takumi Fukuzawa, Ning Ding, Toru Tamaki
Abstract:
M3DDM provides a computationally efficient framework for video outpainting via latent diffusion modeling. However, it exhibits significant quality degradation ‑‑ manifested as spatial blur and temporal inconsistency ‑‑ under challenging scenarios characterized by limited camera motion or large outpainting regions, where inter‑frame information is limited. We identify the cause as a training‑inference mismatch in the masking strategy: M3DDM's training applies random mask directions and widths across frames, whereas inference requires consistent directional outpainting throughout the video. To address this, we propose M3DDM+, which applies uniform mask direction and width across all frames during training, followed by fine‑tuning of the pretrained M3DDM model. Experiments demonstrate that M3DDM+ substantially improves visual fidelity and temporal coherence in information‑limited scenarios while maintaining computational efficiency. The code is available at https://github.com/tamaki‑lab/M3DDM‑Plus.
Authors:Long Ma, Zihao Xue, Yan Wang, Zhiyuan Yan, Jin Xu, Xiaorui Jiang, Haiyang Yu, Yong Liao, Zhen Bi
Abstract:
Recent advances in generative modeling can create remarkably realistic synthetic videos, making it increasingly difficult for humans to distinguish them from real ones and necessitating reliable detection methods.
However, two key limitations hinder the development of this field.
From the dataset perspective, existing datasets are often limited in scale and constructed using outdated or narrowly scoped generative models, making it difficult to capture the diversity and rapid evolution of modern generative techniques. Moreover, the dataset construction process frequently prioritizes quantity over quality, neglecting essential aspects such as semantic diversity, scenario coverage, and technological representativeness.
From the benchmark perspective, current benchmarks largely remain at the stage of dataset creation, leaving many fundamental issues and in‑depth analysis yet to be systematically explored.
Addressing this gap, we propose AIGVDBench, a benchmark designed to be comprehensive and representative, covering 31 state‑of‑the‑art generation models and over 440,000 videos. By executing more than 1,500 evaluations on 33 existing detectors belonging to four distinct categories. This work presents 8 in‑depth analyses from multiple perspectives and identifies 4 novel findings that offer valuable insights for future research. We hope this work provides a solid foundation for advancing the field of AI‑generated video detection.
Our benchmark is open‑sourced at https://github.com/LongMa‑2025/AIGVDBench.
Authors:Santiago Martínez Novoa, María Catalina Ibáñez, Lina Gómez Mesa, Jeremias Kramer
Abstract:
Multi‑Classification Chest X‑Ray Images are one of the most prevalent forms of radiological examination used for diagnosing thoracic diseases. In this study, we offer a concise overview of several methods employed for tackling this task, including DenseNet121. In addition, we deploy an open‑source web‑based application. In our study, we conduct tests to compare different methods and see how well they work. We also look closely at the weaknesses of the methods we propose and suggest ideas for making them better in the future. Our code is available at: https://github.com/AML4206‑MINE20242/Proyecto_AML
Authors:Chuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger, Gerard Pons-Moll
Abstract:
Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence‑level or action‑level descriptions due to the absence of fine‑grained, part‑level motion annotations. This limits their controllability over individual body parts. In this work, we construct a high‑quality motion dataset with atomic, temporally‑aware part‑level text annotations, leveraging the reasoning capabilities of large language models (LLMs). Unlike prior datasets that either provide synchronized part captions with fixed time segments or rely solely on global sequence labels, our dataset captures asynchronous and semantically distinct part movements at fine temporal resolution. Based on this dataset, we introduce a diffusion‑based part‑aware motion generation framework, namely FrankenMotion, where each body part is guided by its own temporally‑structured textual prompt. This is, to our knowledge, the first work to provide atomic, temporally‑aware part‑level motion annotations and have a model that allows motion generation with both spatial (body part) and temporal (atomic action) control. Experiments demonstrate that FrankenMotion outperforms all previous baseline models adapted and retrained for our setting, and our model can compose motions unseen during training. Our code and dataset will be publicly available upon publication.
Authors:Chongcong Jiang, Tianxingjian Ding, Chuhan Song, Jiachen Tu, Ziyang Yan, Yihua Shao, Zhenyi Wang, Yuzhang Shang, Tianyu Han, Yu Tian
Abstract:
Promptable segmentation foundation models such as SAM3 have demonstrated strong generalization capabilities through interactive and concept‑based prompting. However, their direct applicability to medical image segmentation remains limited by severe domain shifts, the absence of privileged spatial prompts, and the need to reason over complex anatomical and volumetric structures. Here we present Medical SAM3, a foundation model for universal prompt‑driven medical image segmentation, obtained by fully fine‑tuning SAM3 on large‑scale, heterogeneous 2D and 3D medical imaging datasets with paired segmentation masks and text prompts. Through a systematic analysis of vanilla SAM3, we observe that its performance degrades substantially on medical data, with its apparent competitiveness largely relying on strong geometric priors such as ground‑truth‑derived bounding boxes. These findings motivate full model adaptation beyond prompt engineering alone. By fine‑tuning SAM3's model parameters on 33 datasets spanning 10 medical imaging modalities, Medical SAM3 acquires robust domain‑specific representations while preserving prompt‑driven flexibility. Extensive experiments across organs, imaging modalities, and dimensionalities demonstrate consistent and significant performance gains, particularly in challenging scenarios characterized by semantic ambiguity, complex morphology, and long‑range 3D context. Our results establish Medical SAM3 as a universal, text‑guided segmentation foundation model for medical imaging and highlight the importance of holistic model adaptation for achieving robust prompt‑driven segmentation under severe domain shift. Code and model will be made available at https://github.com/AIM‑Research‑Lab/Medical‑SAM3.
Authors:Gerhard Krumpl, Henning Avenhaus, Horst Possegger
Abstract:
Current progress in out‑of‑distribution (OOD) detection is limited by the lack of large, high‑quality datasets with clearly defined OOD categories across varying difficulty levels (near‑ to far‑OOD) that support both fine‑ and coarse‑grained computer vision tasks. To address this limitation, we introduce ICONIC‑444 (Image Classification and OOD Detection with Numerous Intricate Complexities), a specialized large‑scale industrial image dataset containing over 3.1 million RGB images spanning 444 classes tailored for OOD detection research. Captured with a prototype industrial sorting machine, ICONIC‑444 closely mimics real‑world tasks. It complements existing datasets by offering structured, diverse data suited for rigorous OOD evaluation across a spectrum of task complexities. We define four reference tasks within ICONIC‑444 to benchmark and advance OOD detection research and provide baseline results for 22 state‑of‑the‑art post‑hoc OOD detection methods.
Authors:Sen Wang, Bangwei Liu, Zhenkun Gao, Lizhuang Ma, Xuhong Wang, Yuan Xie, Xin Tan
Abstract:
An ideal embodied agent should possess lifelong learning capabilities to handle long‑horizon and complex tasks, enabling continuous operation in general environments. This not only requires the agent to accurately accomplish given tasks but also to leverage long‑term episodic memory to optimize decision‑making. However, existing mainstream one‑shot embodied tasks primarily focus on task completion results, neglecting the crucial process of exploration and memory utilization. To address this, we propose Long‑term Memory Embodied Exploration (LMEE), which aims to unify the agent's exploratory cognition and decision‑making behaviors to promote lifelong learning. We further construct a corresponding dataset and benchmark, LMEE‑Bench, incorporating multi‑goal navigation and memory‑based question answering to comprehensively evaluate both the process and outcome of embodied exploration. To enhance the agent's memory recall and proactive exploration capabilities, we propose MemoryExplorer, a novel method that fine‑tunes a multimodal large language model through reinforcement learning to encourage active memory querying. By incorporating a multi‑task reward function that includes action prediction, frontier selection, and question answering, our model achieves proactive exploration. Extensive experiments against state‑of‑the‑art embodied exploration models demonstrate that our approach achieves significant advantages in long‑horizon embodied tasks. Our dataset and code will be released at https://wangsen99.github.io/papers/lmee/
Authors:Tal Reiss, Daniel Winter, Matan Cohen, Alex Rav-Acha, Yael Pritch, Ariel Shamir, Yedid Hoshen
Abstract:
We introduce Alterbute, a diffusion‑based method for editing an object's intrinsic attributes in an image. We allow changing color, texture, material, and even the shape of an object, while preserving its perceived identity and scene context. Existing approaches either rely on unsupervised priors that often fail to preserve identity or use overly restrictive supervision that prevents meaningful intrinsic variations. Our method relies on: (i) a relaxed training objective that allows the model to change both intrinsic and extrinsic attributes conditioned on an identity reference image, a textual prompt describing the target intrinsic attributes, and a background image and object mask defining the extrinsic context. At inference, we restrict extrinsic changes by reusing the original background and object mask, thereby ensuring that only the desired intrinsic attributes are altered; (ii) Visual Named Entities (VNEs) ‑ fine‑grained visual identity categories (e.g., ''Porsche 911 Carrera'') that group objects sharing identity‑defining features while allowing variation in intrinsic attributes. We use a vision‑language model to automatically extract VNE labels and intrinsic attribute descriptions from a large public image dataset, enabling scalable, identity‑preserving supervision. Alterbute outperforms existing methods on identity‑preserving object intrinsic attribute editing.
Authors:Darshan Singh, Arsha Nagrani, Kawshik Manikantan, Harman Singh, Dinesh Tewari, Tobias Weyand, Cordelia Schmid, Anelia Angelova, Shachi Dave
Abstract:
Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western‑centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce MINERVA‑Cultural, a challenging benchmark for multicultural and multilingual video reasoning. MINERVA‑Cultural comprises high‑quality, entirely human‑generated annotations from diverse, region‑specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, MINERVA‑Cultural provides complex questions, answers, and multi‑step reasoning steps, all crafted in native languages. Making progress on MINERVA‑Cultural requires a deeply situated understanding of visual cultural context. Furthermore, we leverage MINERVA‑Cultural's reasoning traces to construct evidence‑based graphs and propose a novel iterative strategy using these graphs to identify fine‑grained errors in reasoning. Our evaluations reveal that SoTA Video‑LLMs struggle significantly, performing substantially below human‑level accuracy, with errors primarily stemming from the visual perception of cultural elements. MINERVA‑Cultural will be publicly available under https://github.com/google‑deepmind/neptune?tab=readme‑ov‑file\#minerva‑cultural
Authors:Chengfeng Zhao, Jiazhi Shu, Yubo Zhao, Tianyu Huang, Jiahao Lu, Zekai Gu, Chengwei Ren, Zhiyang Dou, Qing Shuai, Yuan Liu
Abstract:
In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior for plausibility and consistency in videos, while pre‑trained video models offer strong generalization capabilities for motions. Based on this, we present CoMoVi, a co‑generative framework that generates 3D human motions and videos synchronously within a single diffusion denoising loop. However, since the 3D human motions and the 2D human‑centric videos have a modality gap between each other, we propose to project the 3D human motion into an effective 2D human motion representation that effectively aligns with the 2D videos. Then, we design a dual‑branch diffusion model to couple human motion and the video generation process with mutual feature interaction and 3D‑2D cross attentions. To train and evaluate our model, we curate CoMoVi‑Dataset, a large‑scale real‑world human video dataset with text and motion annotations, covering diverse and challenging human motions. Extensive experiments demonstrate that our method generates high‑quality 3D human motion with a better generalization ability and that our method can generate high‑quality human‑centric videos without external motion references.
Authors:Wenqing Wang, Da Li, Xiatian Zhu, Josef Kittler
Abstract:
Fine‑tuning vision‑language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adaptation; however, forgetting often remains inevitable during this process. We introduce a novel paradigm, continued fine‑tuning (CFT), which seeks to recover pretrained knowledge after a zero‑shot model has already been adapted. We propose a simple, model‑agnostic CFT strategy (named MERGETUNE) guided by linear mode connectivity (LMC), which can be applied post hoc to existing fine‑tuned models without requiring architectural changes. Given a fine‑tuned model, we continue fine‑tuning its trainable parameters (e.g., soft prompts or linear heads) to search for a continued model which has two low‑loss paths to the zero‑shot (e.g., CLIP) and the fine‑tuned (e.g., CoOp) solutions. By exploiting the geometry of the loss landscape, the continued model implicitly merges the two solutions, restoring pretrained knowledge lost in the fine‑tuned counterpart. A challenge is that the vanilla LMC constraint requires data replay from the pretraining task. We approximate this constraint for the zero‑shot model via a second‑order surrogate, eliminating the need for large‑scale data replay. Experiments show that MERGETUNE improves the harmonic mean of CoOp by +5.6% on base‑novel generalisation without adding parameters. On robust fine‑tuning evaluations, the LMC‑merged model from MERGETUNE surpasses ensemble baselines with lower inference cost, achieving further gains and state‑of‑the‑art results when ensembled with the zero‑shot model. Our code is available at https://github.com/Surrey‑UP‑Lab/MERGETUNE.
Authors:Yu Wang, Yi Wang, Rui Dai, Yujie Wang, Kaikui Liu, Xiangxiang Chu, Yansheng Li
Abstract:
As hubs of human activity, urban surfaces consist of a wealth of semantic entities. Segmenting these various entities from satellite imagery is crucial for a range of downstream applications. Current advanced segmentation models can reliably segment entities defined by physical attributes (e.g., buildings, water bodies) but still struggle with socially defined categories (e.g., schools, parks). In this work, we achieve socio‑semantic segmentation by vision‑language model reasoning. To facilitate this, we introduce the Urban Socio‑Semantic Segmentation dataset named SocioSeg, a new resource comprising satellite imagery, digital maps, and pixel‑level labels of social semantic entities organized in a hierarchical structure. Additionally, we propose a novel vision‑language reasoning framework called SocioReasoner that simulates the human process of identifying and annotating social semantic entities via cross‑modal recognition and multi‑stage reasoning. We employ reinforcement learning to optimize this non‑differentiable process and elicit the reasoning capabilities of the vision‑language model. Experiments demonstrate our approach's gains over state‑of‑the‑art models and strong zero‑shot generalization. The dataset and code are open‑sourced under the Apache License 2.0 at https://github.com/AMAP‑ML/SocioReasoner.
Authors:Ahmad Mustapha, Charbel Toumieh, Mariette Awad
Abstract:
With advancements in deep learning (DL) and computer vision techniques, the field of chart understanding is evolving rapidly. In particular, multimodal large language models (MLLMs) are proving to be efficient and accurate in understanding charts. To accurately measure the performance of MLLMs, the research community has developed multiple datasets to serve as benchmarks. By examining these datasets, we found that they are all limited to a small set of chart types. To bridge this gap, we propose the ChartComplete dataset. The dataset is based on a chart taxonomy borrowed from the visualization community, and it covers thirty different chart types. The dataset is a collection of classified chart images and does not include a learning signal. We present the ChartComplete dataset as is to the community to build upon it.
Authors:Clementine Grethen, Nicolas Menga, Roland Brochard, Geraldine Morin, Simone Gasparini, Jeremy Lebreton, Manuel Sanchez Gestido
Abstract:
We address the problem of estimating realistic, spatially varying reflectance for complex planetary surfaces such as the lunar regolith, which is critical for high‑fidelity rendering and vision‑based navigation. Existing lunar rendering pipelines rely on simplified or spatially uniform BRDF models whose parameters are difficult to estimate and fail to capture local reflectance variations, limiting photometric realism. We propose Lunar‑G2R, a geometry‑to‑reflectance learning framework that predicts spatially varying BRDF parameters directly from a lunar digital elevation model (DEM), without requiring multi‑view imagery, controlled illumination, or dedicated reflectance‑capture hardware at inference time. The method leverages a U‑Net trained with differentiable rendering to minimize photometric discrepancies between real orbital images and physically based renderings under known viewing and illumination geometry. Experiments on a geographically held‑out region of the Tycho crater show that our approach reduces photometric error by 38 % compared to a state‑of‑the‑art baseline, while achieving higher PSNR and SSIM and improved perceptual similarity, capturing fine‑scale reflectance variations absent from spatially uniform models. To our knowledge, this is the first method to infer a spatially varying reflectance model directly from terrain geometry.
Authors:Yiming Zhang, Weibo Qin, Yuntian Liu, Feng Wang
Abstract:
Synthetic aperture radar (SAR) imagery exhibits intrinsic information sparsity due to its unique electromagnetic scattering mechanism. Despite the widespread adoption of deep neural network (DNN)‑based SAR automatic target recognition (SAR‑ATR) systems, they remain vulnerable to adversarial examples and tend to over‑rely on background regions, leading to degraded adversarial robustness. Existing adversarial attacks for SAR‑ATR often require visually perceptible distortions to achieve effective performance, thereby necessitating an attack method that balances effectiveness and stealthiness. In this paper, a novel attack method termed Space‑Reweighted Adversarial Warping (SRAW) is proposed, which generates adversarial examples through optimized spatial deformation with reweighted budgets across foreground and background regions. Extensive experiments demonstrate that SRAW significantly degrades the performance of state‑of‑the‑art SAR‑ATR models and consistently outperforms existing methods in terms of imperceptibility and adversarial transferability. Code is made available at https://github.com/boremycin/SAR‑ATR‑TransAttack.
Authors:Xueyun Tian, Wei Li, Bingbing Xu, Heng Dong, Yuanzhuo Wang, Huawei Shen
Abstract:
Recent Omni‑multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio‑video understanding remains challenging, as existing approaches suffer from disjointed capabilities: they typically exhibit incomplete modality support or lack autonomous proactive monitoring. To address this, we present ROMA, a real‑time omni‑multimodal assistant for unified reactive and proactive interaction. ROMA processes continuous inputs as synchronized multimodal units, aligning dense audio with discrete video frames to handle granularity mismatches. For online decision‑making, we introduce a lightweight speak head that decouples response initiation from generation to ensure precise triggering without task conflict. We train ROMA with a curated streaming dataset and a two‑stage curriculum that progressively optimizes for streaming format adaptation and proactive responsiveness. To standardize the fragmented evaluation landscape, we reorganize diverse benchmarks into a unified suite covering both proactive (alert, narration) and reactive (QA) settings. Extensive experiments across 12 benchmarks demonstrate ROMA achieves state‑of‑the‑art performance on proactive tasks while competitive in reactive settings, validating its robustness in unified real‑time omni‑multimodal understanding.
Authors:Megha Mariam K M, C. V. Jawahar
Abstract:
Imagine sitting in a presentation, trying to follow the speaker while simultaneously scanning the slides for relevant information. While the entire slide is visible, identifying the relevant regions can be challenging. As you focus on one part of the slide, the speaker moves on to a new sentence, leaving you scrambling to catch up visually. This constant back‑and‑forth creates a disconnect between what is being said and the most important visual elements, making it hard to absorb key details, especially in fast‑paced or content‑heavy presentations such as conference talks. This requires an understanding of slides, including text, graphics, and layout. We introduce a method that automatically identifies and highlights the most relevant slide regions based on the speaker's narrative. By analyzing spoken content and matching it with textual or graphical elements in the slides, our approach ensures better synchronization between what listeners hear and what they need to attend to. We explore different ways of solving this problem and assess their success and failure cases. Analyzing multimedia documents is emerging as a key requirement for seamless understanding of content‑rich videos, such as educational videos and conference talks, by reducing cognitive strain and improving comprehension. Code and dataset are available at: https://github.com/meghamariamkm2002/Slide_Highlight
Authors:Sicheng Yang, Yukai Huang, Shitong Sun, Weitong Cai, Jiankang Deng, Jifei Song, Zhensong Zhang
Abstract:
Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD‑EPIC VQA due to ambiguous queries/options, poor long‑range temporal reasoning, and non‑standardized outputs. We propose a framework integrating query/choice pre‑processing, domain‑specific Qwen2.5‑VL fine‑tuning, a novel Temporal Chain‑of‑Thought (T‑CoT) prompting for multi‑step reasoning, and robust post‑processing. This system achieves 41.6% accuracy on HD‑EPIC VQA, highlighting the need for holistic pipeline optimization in demanding video understanding. Our code, fine‑tuned models are available at https://github.com/YoungSeng/Egocentric‑Co‑Pilot.
Authors:Dong-Yu Chen, Yixin Guo, Shuojin Yang, Tai-Jiang Mu, Shi-Min Hu
Abstract:
Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to achieving precise camera control is warping a 3D representation according to the target trajectory. However, such methods fail to fully leverage the 3D priors of video diffusion models (VDMs) and often fall into the Inpainting Trap, resulting in subject inconsistency and degraded generation quality. To address this problem, we propose DepthDirector, a video re‑rendering framework with precise camera controllability. By leveraging the depth video from explicit 3D representation as camera‑control guidance, our method can faithfully reproduce the dynamic scene of an input video under novel camera trajectories. Specifically, we design a View‑Content Dual‑Stream Condition mechanism that injects both the source video and the warped depth sequence rendered under the target viewpoint into the pretrained video generation model. This geometric guidance signal enables VDMs to comprehend camera movements and leverage their 3D understanding capabilities, thereby facilitating precise camera control and consistent content generation. Next, we introduce a lightweight LoRA‑based video diffusion adapter to train our framework, fully preserving the knowledge priors of VDMs. Additionally, we construct a large‑scale multi‑camera synchronized dataset named MultiCam‑WarpData using Unreal Engine 5, containing 8K videos across 1K dynamic scenes. Extensive experiments show that DepthDirector outperforms existing methods in both camera controllability and visual quality. Our code and dataset will be publicly available.
Authors:Kim Youwang, Lee Hyoseok, Subin Park, Gerard Pons-Moll, Tae-Hyun Oh
Abstract:
We introduce ELITE, an Efficient Gaussian head avatar synthesis from a monocular video via Learned Initialization and TEst‑time generative adaptation. Prior works rely either on a 3D data prior or a 2D generative prior to compensate for missing visual cues in monocular videos. However, 3D data prior methods often struggle to generalize in‑the‑wild, while 2D generative prior methods are computationally heavy and prone to identity hallucination. We identify a complementary synergy between these two priors and design an efficient system that achieves high‑fidelity animatable avatar synthesis with strong in‑the‑wild generalization. Specifically, we introduce a feed‑forward Mesh2Gaussian Prior Model (MGPM) that enables fast initialization of a Gaussian avatar. To further bridge the domain gap at test time, we design a test‑time generative adaptation stage, leveraging both real and synthetic images as supervision. Unlike previous full diffusion denoising strategies that are slow and hallucination‑prone, we propose a rendering‑guided single‑step diffusion enhancer that restores missing visual details, grounded on Gaussian avatar renderings. Our experiments demonstrate that ELITE produces visually superior avatars to prior works, even for challenging expressions, while achieving 60x faster synthesis than the 2D generative prior method.
Authors:Sicheng Yang, Zhaohu Xing, Lei Zhu
Abstract:
Consistency learning with feature perturbation is a widely used strategy in semi‑supervised medical image segmentation. However, many existing perturbation methods rely on dropout, and thus require a careful manual tuning of the dropout rate, which is a sensitive hyperparameter and often difficult to optimize and may lead to suboptimal regularization. To overcome this limitation, we propose VQ‑Seg, the first approach to employ vector quantization (VQ) to discretize the feature space and introduce a novel and controllable Quantized Perturbation Module (QPM) that replaces dropout. Our QPM perturbs discrete representations by shuffling the spatial locations of codebook indices, enabling effective and controllable regularization. To mitigate potential information loss caused by quantization, we design a dual‑branch architecture where the post‑quantization feature space is shared by both image reconstruction and segmentation tasks. Moreover, we introduce a Post‑VQ Feature Adapter (PFA) to incorporate guidance from a foundation model (FM), supplementing the high‑level semantic information lost during quantization. Furthermore, we collect a large‑scale Lung Cancer (LC) dataset comprising 828 CT scans annotated for central‑type lung carcinoma. Extensive experiments on the LC dataset and other public benchmarks demonstrate the effectiveness of our method, which outperforms state‑of‑the‑art approaches. Code available at: https://github.com/script‑Yang/VQ‑Seg.
Authors:Chenyue Zhou, Jiayi Tuo, Shitong Qin, Wei Dai, Mingxuan Wang, Ziwei Zhao, Duoyang Li, Shiyang Su, Yanxi Lu, Yanbiao Ma
Abstract:
The automated extraction of structured questions from paper‑based mathematics exams is fundamental to intelligent education, yet remains challenging in real‑world settings due to severe visual noise. Existing benchmarks mainly focus on clean documents or generic layout analysis, overlooking both the structural integrity of mathematical problems and the ability of models to actively reject incomplete inputs. We introduce MathDoc, the first benchmark for document‑level information extraction from authentic high school mathematics exam papers. MathDoc contains 3,609 carefully curated questions with real‑world artifacts and explicitly includes unrecognizable samples to evaluate active refusal behavior. We propose a multi‑dimensional evaluation framework covering stem accuracy, visual similarity, and refusal capability. Experiments on SOTA MLLMs, including Qwen3‑VL and Gemini‑2.5‑Pro, show that although end‑to‑end models achieve strong extraction performance, they consistently fail to refuse illegible inputs, instead producing confident but invalid outputs. These results highlight a critical gap in current MLLMs and establish MathDoc as a benchmark for assessing model reliability under degraded document conditions. Our project repository is available at \hrefhttps://github.com/winnk123/papers/tree/masterGitHub repository
Authors:Han Wang, Yi Yang, Jingyuan Hu, Minfeng Zhu, Wei Chen
Abstract:
Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision‑language models (VLMs). However, state‑of‑the‑art approaches rely heavily on large‑scale human‑annotated datasets, which are costly and time‑consuming to acquire. To overcome this limitation, we introduce V‑Zero, a general post‑training framework that facilitates self‑improvement using exclusively unlabeled images. V‑Zero establishes a co‑evolutionary loop by instantiating two distinct roles: a Questioner and a Solver. The Questioner learns to synthesize high‑quality, challenging questions by leveraging a dual‑track reasoning reward that contrasts intuitive guesses with reasoned results. The Solver is optimized using pseudo‑labels derived from majority voting over its own sampled responses. Both roles are trained iteratively via Group Relative Policy Optimization (GRPO), driving a cycle of mutual enhancement. Remarkably, without a single human annotation, V‑Zero achieves consistent performance gains on Qwen2.5‑VL‑7B‑Instruct, improving visual mathematical reasoning by +1.7 and general vision‑centric by +2.6, demonstrating the potential of self‑improvement in multimodal systems. Code is available at https://github.com/SatonoDia/V‑Zero
Authors:Nick Truong, Pritam P. Karmokar, William J. Beksi
Abstract:
Underwater imaging is fundamentally challenging due to wavelength‑dependent light attenuation, strong scattering from suspended particles, turbidity‑induced blur, and non‑uniform illumination. These effects impair standard cameras and make ground‑truth motion nearly impossible to obtain. On the other hand, event cameras offer microsecond resolution and high dynamic range. Nonetheless, progress on investigating event cameras for underwater environments has been limited due to the lack of datasets that pair realistic underwater optics with accurate optical flow. To address this problem, we introduce the first synthetic underwater benchmark dataset for event‑based optical flow derived from physically‑based ray‑traced RGBD sequences. Using a modern video‑to‑event pipeline applied to rendered underwater videos, we produce realistic event data streams with dense ground‑truth flow, depth, and camera motion. Moreover, we benchmark state‑of‑the‑art learning‑based and model‑based optical flow prediction methods to understand how underwater light transport affects event formation and motion estimation accuracy. Our dataset establishes a new baseline for future development and evaluation of underwater event‑based perception algorithms. The source code and dataset for this project are publicly available at https://robotic‑vision‑lab.github.io/ueof.
Authors:Carlo Sgaravatti, Riccardo Pieroni, Matteo Corno, Sergio M. Savaresi, Luca Magri, Giacomo Boracchi
Abstract:
Accurately localizing 3D objects like pedestrians, cyclists, and other vehicles is essential in Autonomous Driving. To ensure high detection performance, Autonomous Vehicles complement RGB cameras with LiDAR sensors, but effectively combining these data sources for 3D object detection remains challenging. We propose LCF3D, a novel sensor fusion framework that combines a 2D object detector on RGB images with a 3D object detector on LiDAR point clouds. By leveraging multimodal fusion principles, we compensate for inaccuracies in the LiDAR object detection network. Our solution combines two key principles: (i) late fusion, to reduce LiDAR False Positives by matching LiDAR 3D detections with RGB 2D detections and filtering out unmatched LiDAR detections; and (ii) cascade fusion, to recover missed objects from LiDAR by generating new 3D frustum proposals corresponding to unmatched RGB detections. Experiments show that LCF3D is beneficial for domain generalization, as it turns out to be successful in handling different sensor configurations between training and testing domains. LCF3D achieves significant improvements over LiDAR‑based methods, particularly for challenging categories like pedestrians and cyclists in the KITTI dataset, as well as motorcycles and bicycles in nuScenes. Code can be downloaded from: https://github.com/CarloSgaravatti/LCF3D.
Authors:Chi-Pin Huang, Yunze Man, Zhiding Yu, Min-Hung Chen, Jan Kautz, Yu-Chiang Frank Wang, Fu-En Yang
Abstract:
Vision‑Language‑Action (VLA) tasks require reasoning over complex visual scenes and executing adaptive actions in dynamic environments. While recent studies on reasoning VLAs show that explicit chain‑of‑thought (CoT) can improve generalization, they suffer from high inference latency due to lengthy reasoning traces. We propose Fast‑ThinkAct, an efficient reasoning framework that achieves compact yet performant planning through verbalizable latent reasoning. Fast‑ThinkAct learns to reason efficiently with latent CoTs by distilling from a teacher, driven by a preference‑guided objective to align manipulation trajectories that transfers both linguistic and visual planning capabilities for embodied control. This enables reasoning‑enhanced policy learning that effectively connects compact reasoning to action execution. Extensive experiments across diverse embodied manipulation and reasoning benchmarks demonstrate that Fast‑ThinkAct achieves strong performance with up to 89.3% reduced inference latency over state‑of‑the‑art reasoning VLAs, while maintaining effective long‑horizon planning, few‑shot adaptation, and failure recovery.
Authors:Ruiqi Shen, Chang Liu, Henghui Ding
Abstract:
Segment Anything 3 (SAM3) has established a powerful foundation that robustly detects, segments, and tracks specified targets in videos. However, in its original implementation, its group‑level collective memory selection is suboptimal for complex multi‑object scenarios, as it employs a synchronized decision across all concurrent targets conditioned on their average performance, often overlooking individual reliability. To this end, we propose SAM3‑DMS, a training‑free decoupled strategy that utilizes fine‑grained memory selection on individual objects. Experiments demonstrate that our approach achieves robust identity preservation and tracking stability. Notably, our advantage becomes more pronounced with increased target density, establishing a solid foundation for simultaneous multi‑target video segmentation in the wild.
Authors:Xuyang Fang, Sion Hannuna, Edwin Simpson, Neill Campbell
Abstract:
Identifying individual animals in long‑duration videos is essential for behavioral ecology, wildlife monitoring, and livestock management. Traditional methods require extensive manual annotation, while existing self‑supervised approaches are computationally demanding and ill‑suited for long sequences due to memory constraints and temporal error propagation. We introduce a highly efficient, self‑supervised method that reframes animal identification as a global clustering task rather than a sequential tracking problem. Our approach assumes a known, fixed number of individuals within a single video ‑‑ a common scenario in practice ‑‑ and requires only bounding box detections and the total count. By sampling pairs of frames, using a frozen pre‑trained backbone, and employing a self‑bootstrapping mechanism with the Hungarian algorithm for in‑batch pseudo‑label assignment, our method learns discriminative features without identity labels. We adapt a Binary Cross Entropy loss from vision‑language models, enabling state‑of‑the‑art accuracy (>97%) while consuming less than 1 GB of GPU memory per batch ‑‑ an order of magnitude less than standard contrastive methods. Evaluated on challenging real‑world datasets (3D‑POP pigeons and 8‑calves feeding videos), our framework matches or surpasses supervised baselines trained on over 1,000 labeled frames, effectively removing the manual annotation bottleneck. This work enables practical, high‑accuracy animal identification on consumer‑grade hardware, with broad applicability in resource‑constrained research settings. All code written for this paper are \hrefhttps://huggingface.co/datasets/tonyFang04/8‑calveshere.
Authors:Yonglin Tian, Qiyao Zhang, Wei Xu, Yutong Wang, Yihao Wu, Xinyi Li, Xingyuan Dai, Hui Zhang, Zhiyong Cui, Baoqing Guo, Zujun Yu, Yisheng Lv
Abstract:
Accurate and early perception of potential intrusion targets is essential for ensuring the safety of railway transportation systems. However, most existing systems focus narrowly on object classification within fixed visual scopes and apply rule‑based heuristics to determine intrusion status, often overlooking targets that pose latent intrusion risks. Anticipating such risks requires the cognition of spatial context and temporal dynamics for the object of interest (OOI), which presents challenges for conventional visual models. To facilitate deep intrusion perception, we introduce a novel benchmark, CogRail, which integrates curated open‑source datasets with cognitively driven question‑answer annotations to support spatio‑temporal reasoning and prediction. Building upon this benchmark, we conduct a systematic evaluation of state‑of‑the‑art visual‑language models (VLMs) using multimodal prompts to identify their strengths and limitations in this domain. Furthermore, we fine‑tune VLMs for better performance and propose a joint fine‑tuning framework that integrates three core tasks, position perception, movement prediction, and threat analysis, facilitating effective adaptation of general‑purpose foundation models into specialized models tailored for cognitive intrusion perception. Extensive experiments reveal that current large‑scale multimodal models struggle with the complex spatial‑temporal reasoning required by the cognitive intrusion perception task, underscoring the limitations of existing foundation models in this safety‑critical domain. In contrast, our proposed joint fine‑tuning framework significantly enhances model performance by enabling targeted adaptation to domain‑specific reasoning demands, highlighting the advantages of structured multi‑task learning in improving both accuracy and interpretability. Code will be available at https://github.com/Hub‑Tian/CogRail.
Authors:Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng Sun
Abstract:
We propose OpenVoxel, a training‑free algorithm for grouping and captioning sparse voxels for the open‑vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization (SVR) model obtained from multi‑view images of a 3D scene, our OpenVoxel is able to produce meaningful groups that describe different objects in the scene. Also, by leveraging powerful Vision Language Models (VLMs) and Multi‑modal Large Language Models (MLLMs), our OpenVoxel successfully build an informative scene map by captioning each group, enabling further 3D scene understanding tasks such as open‑vocabulary segmentation (OVS) or referring expression segmentation (RES). Unlike previous methods, our method is training‑free and does not introduce embeddings from a CLIP/BERT text encoder. Instead, we directly proceed with text‑to‑text search using MLLMs. Through extensive experiments, our method demonstrates superior performance compared to recent studies, particularly in complex referring expression segmentation (RES) tasks. The code will be open.
Authors:Lennart Eing, Cristina Luna-Jiménez, Silvan Mertes, Elisabeth André
Abstract:
This paper introduces a novel application of Video Joint‑Embedding Predictive Architectures (V‑JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre‑training methods for video understanding that rely on pixel‑level reconstructions, V‑JEPAs learn by predicting embeddings of masked regions from the embeddings of unmasked regions. This enables the trained encoder to not capture irrelevant information about a given video like the color of a region of pixels in the background. Using a pre‑trained V‑JEPA video encoder, we train shallow classifiers using the RAVDESS and CREMA‑D datasets, achieving state‑of‑the‑art performance on RAVDESS and outperforming all other vision‑based methods on CREMA‑D (+1.48 WAR). Furthermore, cross‑dataset evaluations reveal strong generalization capabilities, demonstrating the potential of purely embedding‑based pre‑training approaches to advance FER. We release our code at https://github.com/lennarteingunia/vjepa‑for‑fer.
Authors:Ritabrata Chakraborty, Hrishit Mitra, Shivakumara Palaiahnakote, Umapada Pal
Abstract:
Object detectors often perform well in‑distribution, yet degrade sharply on a different benchmark. We study cross‑dataset object detection (CD‑OD) through a lens of setting specificity. We group benchmarks into setting‑agnostic datasets with diverse everyday scenes and setting‑specific datasets tied to a narrow environment, and evaluate a standard detector family across all train‑‑test pairs. This reveals a clear structure in CD‑OD: transfer within the same setting type is relatively stable, while transfer across setting types drops substantially and is often asymmetric. The most severe breakdowns occur when transferring from specific sources to agnostic targets, and persist after open‑label alignment, indicating that domain shift dominates in the hardest regimes. To disentangle domain shift from label mismatch, we compare closed‑label transfer with an open‑label protocol that maps predicted classes to the nearest target label using CLIP similarity. Open‑label evaluation yields consistent but bounded gains, and many corrected cases correspond to semantic near‑misses supported by the image evidence. Overall, we provide a principled characterization of CD‑OD under setting specificity and practical guidance for evaluating detectors under distribution shift. Code will be released at \href[https://github.com/Ritabrata04/cdod‑icpr.githttps://github.com/Ritabrata04/cdod‑icpr.
Authors:Ahmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord, Alexandre Alahi
Abstract:
Recent video diffusion models generate photorealistic, temporally coherent videos, yet they fall short as reliable world models for autonomous driving, where structured motion and physically consistent interactions are essential. Adapting these generalist video models to driving domains has shown promise but typically requires massive domain‑specific data and costly fine‑tuning. We propose an efficient adaptation framework that converts generalist video diffusion models into controllable driving world models with minimal supervision. The key idea is to decouple motion learning from appearance synthesis. First, the model is adapted to predict structured motion in a simplified form: videos of skeletonized agents and scene elements, focusing learning on physical and social plausibility. Then, the same backbone is reused to synthesize realistic RGB videos conditioned on these motion sequences, effectively "dressing" the motion with texture and lighting. This two‑stage process mirrors a reasoning‑rendering paradigm: first infer dynamics, then render appearance. Our experiments show this decoupled approach is exceptionally efficient: adapting SVD, we match prior SOTA models with less than 6% of their compute. Scaling to LTX, our MAD‑LTX model outperforms all open‑source competitors, and supports a comprehensive suite of text, ego, and object controls. Project page: https://vita‑epfl.github.io/MAD‑World‑Model/
Authors:Rui Zhu, Xin Shen, Shuchen Wu, Chenxi Miao, Xin Yu, Yang Li, Weikang Li, Deguo Xia, Jizhou Huang
Abstract:
Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmarks primarily focus on single‑step perception‑to‑judgment tasks, leaving scenarios requiring complex visual‑spatial logical chains significantly underexplored. To bridge this gap, we introduce Video‑MSR, the first benchmark specifically designed to evaluate Multi‑hop Spatial Reasoning (MSR) in dynamic video scenarios. Video‑MSR systematically probes MSR capabilities through four distinct tasks: Constrained Localization, Chain‑based Reference Retrieval, Route Planning, and Counterfactual Physical Deduction. Our benchmark comprises 3,052 high‑quality video instances with 4,993 question‑answer pairs, constructed via a scalable, visually‑grounded pipeline combining advanced model generation with rigorous human verification. Through a comprehensive evaluation of 20 state‑of‑the‑art MLLMs, we uncover significant limitations, revealing that while models demonstrate proficiency in surface‑level perception, they exhibit distinct performance drops in MSR tasks, frequently suffering from spatial disorientation and hallucination during multi‑step deductions. To mitigate these shortcomings and empower models with stronger MSR capabilities, we further curate MSR‑9K, a specialized instruction‑tuning dataset, and fine‑tune Qwen‑VL, achieving a +7.82% absolute improvement on Video‑MSR. Our results underscore the efficacy of multi‑hop spatial instruction data and establish Video‑MSR as a vital foundation for future research. The code and data will be available at https://github.com/ruiz‑nju/Video‑MSR.
Authors:Yaxi Chen, Zi Ye, Shaheer U. Saeed, Oliver Yu, Simin Ni, Jie Huang, Yipeng Hu
Abstract:
Osteosarcoma (OS) is an aggressive primary bone malignancy. Accurate histopathological assessment of viable versus non‑viable tumor regions after neoadjuvant chemotherapy is critical for prognosis and treatment planning, yet manual evaluation remains labor‑intensive, subjective, and prone to inter‑observer variability. Recent advances in digital pathology have enabled automated necrosis quantification. Evaluating on test data, independently sampled on patient‑level, revealed that the deep learning model performance dropped significantly from the tile‑level generalization ability reported in previous studies. First, this work proposes the use of radiomic features as additional input in model training. We show that, despite that they are derived from the images, such a multimodal input effectively improved the classification performance, in addition to its added benefits in interpretability. Second, this work proposes to optimize two binary classification tasks with hierarchical classes (i.e. tumor‑vs‑non‑tumor and viable‑vs‑non‑viable), as opposed to the alternative ``flat'' three‑class classification task (i.e. non‑tumor, non‑viable tumor, viable tumor), thereby enabling a hierarchical loss. We show that such a hierarchical loss, with trainable weightings between the two tasks, the per‑class performance can be improved significantly. Using the TCIA OS Tumor Assessment dataset, we experimentally demonstrate the benefits from each of the proposed new approaches and their combination, setting a what we consider new state‑of‑the‑art performance on this open dataset for this application. Code and trained models: https://github.com/YaxiiC/RadiomicsOS.git.
Authors:Xinming Fang, Chaoyan Huang, Juncheng Li, Jun Wang, Jun Shi, Guixu Zhang
Abstract:
Magnetic resonance imaging (MRI) plays a vital role in clinical diagnostics, yet it remains hindered by long acquisition times and motion artifacts. Multi‑contrast MRI reconstruction has emerged as a promising direction by leveraging complementary information from fully‑sampled reference scans. However, existing approaches suffer from three major limitations: (1) superficial reference fusion strategies, such as simple concatenation, (2) insufficient utilization of the complementary information provided by the reference contrast, and (3) fixed under‑sampling patterns. We propose an efficient and interpretable frequency error‑guided reconstruction framework to tackle these issues. We first employ a conditional diffusion model to learn a Frequency Error Prior (FEP), which is then incorporated into a unified framework for jointly optimizing both the under‑sampling pattern and the reconstruction network. The proposed reconstruction model employs a model‑driven deep unfolding framework that jointly exploits frequency‑ and image‑domain information. In addition, a spatial alignment module and a reference feature decomposition strategy are incorporated to improve reconstruction quality and bridge model‑based optimization with data‑driven learning for improved physical interpretability. Comprehensive validation across multiple imaging modalities, acceleration rates (4‑30x), and sampling schemes demonstrates consistent superiority over state‑of‑the‑art methods in both quantitative metrics and visual quality. All codes are available at https://github.com/fangxinming/JUF‑MRI.
Authors:Maria Sdraka, Dimitrios Michail, Ioannis Papoutsis
Abstract:
Delineating wildfire affected areas using satellite imagery remains challenging due to irregular and spatially heterogeneous spectral changes across the electromagnetic spectrum. While recent deep learning approaches achieve high accuracy when high‑resolution multispectral data are available, their applicability in operational settings, where a quick delineation of the burn scar shortly after a wildfire incident is required, is limited by the trade‑off between spatial resolution and temporal revisit frequency of current satellite systems. To address this limitation, we propose a novel deep learning model, namely BAM‑MRCD, which employs multi‑resolution, multi‑source satellite imagery (MODIS and Sentinel‑2) for the timely production of detailed burnt area maps with high spatial and temporal resolution. Our model manages to detect even small scale wildfires with high accuracy, surpassing similar change detection models as well as solid baselines. All data and code are available in the GitHub repository: https://github.com/Orion‑AI‑Lab/BAM‑MRCD.
Authors:Jiajun Chen, Jing Xiao, Shaohan Cao, Yuming Zhu, Liang Liao, Jun Pan, Mi Wang
Abstract:
Satellite videos provide continuous observations of surface dynamics but pose significant challenges for multi‑object tracking (MOT), especially under unstabilized conditions where platform jitter and the weak appearance of tiny objects jointly degrade tracking performance. To address this problem, we propose DeTracker, a joint‑detection‑and‑tracking framework tailored for unstabilized satellite videos. DeTracker introduces a task‑driven Global‑Local Motion Decoupling (GLMD) module to address the motion imbalance between dominant platform motion and weak target motion. It suppresses background‑dominated motion via global semantic alignment at the feature level and captures target‑specific motion through local refinement, improving trajectory stability and identity consistency. In addition, a Temporal Dependency Feature Pyramid (TDFP) module is developed to perform cross‑frame temporal feature fusion, enhancing the continuity and discriminability of tiny‑object representations. We further construct a new benchmark dataset, SDM‑Car‑SU, which simulates multi‑directional and multi‑speed platform motions to enable systematic evaluation of tracking robustness under varying motion perturbations. Extensive experiments on both simulated and real unstabilized satellite videos demonstrate that DeTracker significantly outperforms existing methods, achieving 61.1% MOTA on SDM‑Car‑SU and 45.3% MOTA on real satellite video data. The code and dataset will be publicly available at https://github.com/alex‑chenjiajun/DeTracker.
Authors:Bahar Khodabakhshian, Nima Hashemi, Armin Saadat, Zahra Gholami, In-Chang Hwang, Samira Sojoudi, Christina Luong, Purang Abolmaesumi, Teresa Tsang
Abstract:
Purpose: Myocardium segmentation in echocardiography videos is a challenging task due to low contrast, noise, and anatomical variability. Traditional deep learning models either process frames independently, ignoring temporal information, or rely on memory‑based feature propagation, which accumulates error over time. Methods: We propose Point‑Seg, a transformer‑based segmentation framework that integrates point tracking as a temporal cue to ensure stable and consistent segmentation of myocardium across frames. Our method leverages a point‑tracking module trained on a synthetic echocardiography dataset to track key anatomical landmarks across video sequences. These tracked trajectories provide an explicit motion‑aware signal that guides segmentation, reducing drift and eliminating the need for memory‑based feature accumulation. Additionally, we incorporate a temporal smoothing loss to further enhance temporal consistency across frames. Results: We evaluate our approach on both public and private echocardiography datasets. Experimental results demonstrate that Point‑Seg has statistically similar accuracy in terms of Dice to state‑of‑the‑art segmentation models in high quality echo data, while it achieves better segmentation accuracy in lower quality echo with improved temporal stability. Furthermore, Point‑Seg has the key advantage of pixel‑level myocardium motion information as opposed to other segmentation methods. Such information is essential in the computation of other downstream tasks such as myocardial strain measurement and regional wall motion abnormality detection. Conclusion: Point‑Seg demonstrates that point tracking can serve as an effective temporal cue for consistent video segmentation, offering a reliable and generalizable approach for myocardium segmentation in echocardiography videos. The code is available at https://github.com/DeepRCL/PointSeg.
Authors:Yanguang Sun, Chao Wang, Jian Yang, Lei Luo
Abstract:
Accurately localizing and segmenting relevant objects from optical remote sensing images (ORSIs) is critical for advancing remote sensing applications. Existing methods are typically built upon moderate‑scale pre‑trained models and employ diverse optimization strategies to achieve promising performance under full‑parameter fine‑tuning. In fact, deeper and larger‑scale foundation models can provide stronger support for performance improvement. However, due to their massive number of parameters, directly adopting full‑parameter fine‑tuning leads to pronounced training difficulties, such as excessive GPU memory consumption and high computational costs, which result in extremely limited exploration of large‑scale models in existing works. In this paper, we propose a novel dynamic wavelet expert‑guided fine‑tuning paradigm with fewer trainable parameters, dubbed WEFT, which efficiently adapts large‑scale foundation models to ORSIs segmentation tasks by leveraging the guidance of wavelet experts. Specifically, we introduce a task‑specific wavelet expert extractor to model wavelet experts from different perspectives and dynamically regulate their outputs, thereby generating trainable features enriched with task‑specific information for subsequent fine‑tuning. Furthermore, we construct an expert‑guided conditional adapter that first enhances the fine‑grained perception of frozen features for specific tasks by injecting trainable features, and then iteratively updates the information of both types of feature, allowing for efficient fine‑tuning. Extensive experiments show that our WEFT not only outperforms 21 state‑of‑the‑art (SOTA) methods on three ORSIs datasets, but also achieves optimal results in camouflage, natural, and medical scenarios. The source code is available at: https://github.com/CSYSI/WEFT.
Authors:Constantin Kolomiiets, Miroslav Purkrabek, Jiri Matas
Abstract:
Segment Anything (SAM) provides an unprecedented foundation for human segmentation, but may struggle under occlusion, where keypoints may be partially or fully invisible. We adapt SAM 2.1 for pose‑guided segmentation with minimal encoder modifications, retaining its strong generalization. Using a fine‑tuning strategy called PoseMaskRefine, we incorporate pose keypoints with high visibility into the iterative correction process originally employed by SAM, yielding improved robustness and accuracy across multiple datasets. During inference, we simplify prompting by selecting only the three keypoints with the highest visibility. This strategy reduces sensitivity to common errors, such as missing body parts or misclassified clothing, and allows accurate mask prediction from as few as a single keypoint. Our results demonstrate that pose‑guided fine‑tuning of SAM enables effective, occlusion‑aware human segmentation while preserving the generalization capabilities of the original model. The code and pretrained models will be available at https://mirapurkrabek.github.io/BBox‑Mask‑Pose/.
Authors:Anush Lakshman S, Adam Haroon, Beiwen Li
Abstract:
Machine learning approaches for fringe projection profilometry (FPP) are hindered by the lack of large, diverse datasets and standardized benchmarking protocols. This paper introduces the first open‑source, photorealistic synthetic dataset for FPP, generated using NVIDIA Isaac Sim, comprising 15,600 fringe images and 300 depth reconstructions across 50 objects. We apply this dataset to single‑shot FPP, where models predict 3D depth maps directly from individual fringe images without temporal phase shifting. Through systematic ablation studies, we identify optimal learning configurations for long‑range (1.5‑2.1 m) depth prediction. We compare three depth normalization strategies and show that individual normalization, which decouples object shape from absolute scale, yields a 9.1x improvement in object reconstruction accuracy over raw depth. We further show that removing background fringe patterns severely degrades performance across all normalizations, demonstrating that background fringes provide essential spatial phase reference rather than noise. We evaluate six loss functions and identify Hybrid L1 loss as optimal. Using the best configuration, we benchmark four architectures and find UNet achieves the strongest performance, though errors remain far above the sub‑millimeter accuracy of classical FPP. The small performance gap between architectures indicates that the dominant limitation is information deficit rather than model design: single fringe images lack sufficient information for accurate depth recovery without explicit phase cues. This work provides a standardized benchmark and evidence motivating hybrid approaches combining phase‑based FPP with learned refinement. The dataset is available at https://huggingface.co/datasets/aharoon/fpp‑ml‑bench and code at https://github.com/AnushLak/fpp‑ml‑bench.
Authors:Yu Xu, Hongbin Yan, Juan Cao, Yiji Cheng, Tiankai Hang, Runze He, Zijin Yin, Shiyi Zhang, Yuxin Zhang, Jintao Li, Chunyu Wang, Qinglin Lu, Tong-Yee Lee, Fan Tang
Abstract:
Unified image generation and editing models suffer from severe task interference in dense diffusion transformers architectures, where a shared parameter space must compromise between conflicting objectives (e.g., local editing v.s. subject‑driven generation). While the sparse Mixture‑of‑Experts (MoE) paradigm is a promising solution, its gating networks remain task‑agnostic, operating based on local features, unaware of global task intent. This task‑agnostic nature prevents meaningful specialization and fails to resolve the underlying task interference. In this paper, we propose a novel framework to inject semantic intent into MoE routing. We introduce a Hierarchical Task Semantic Annotation scheme to create structured task descriptors (e.g., scope, type, preservation). We then design Predictive Alignment Regularization to align internal routing decisions with the task's high‑level semantics. This regularization evolves the gating network from a task‑agnostic executor to a dispatch center. Our model effectively mitigates task interference, outperforming dense baselines in fidelity and quality, and our analysis shows that experts naturally develop clear and semantically correlated specializations.
Authors:Jiahao Qin, Yiwen Wang
Abstract:
Image registration under domain shift remains a fundamental challenge in computer vision and medical imaging: when source and target images exhibit systematic intensity differences, the brightness constancy assumption underlying conventional registration methods is violated, rendering correspondence estimation ill‑posed. We propose SAR‑Net, a unified framework that addresses this challenge through principled scene‑appearance disentanglement. Our key insight is that observed images can be decomposed into domain‑invariant scene representations and domain‑specific appearance codes, enabling registration via re‑rendering rather than direct intensity matching. We establish theoretical conditions under which this decomposition enables consistent cross‑domain alignment (Proposition 1) and prove that our scene consistency loss provides a sufficient condition for geometric correspondence in the shared latent space (Proposition 2). Empirically, we validate SAR‑Net on the ANHIR (Automatic Non‑rigid Histological Image Registration) challenge benchmark, where multi‑stain histopathology images exhibit coupled domain shift from different staining protocols and geometric distortion from tissue preparation. Our method achieves a median relative Target Registration Error (rTRE) of 0.25%, outperforming the state‑of‑the‑art MEVIS method (0.27% rTRE) by 7.4%, with robustness of 99.1%. Code is available at https://github.com/D‑ST‑Sword/SAR‑NET .
Authors:Qingyu Liu, Zhongjie Ba, Jianmin Guo, Qiu Wang, Zhibo Wang, Jie Shi, Kui Ren
Abstract:
Recently, reconstruction‑based methods have gained attention for AIGC image detection. These methods leverage pre‑trained diffusion models to reconstruct inputs and measure residuals for distinguishing real from fake images. Their key advantage lies in reducing reliance on dataset‑specific artifacts and improving generalization under distribution shifts. However, they are limited by significant inefficiency due to multi‑step inversion and reconstruction, and their reliance on diffusion backbones further limits generalization to other generative paradigms such as GANs.
In this paper, we propose a novel fake image detection framework, called R^2BD, built upon two key designs: (1) G‑LDM, a unified reconstruction model that simulates the generation behaviors of VAEs, GANs, and diffusion models, thereby broadening the detection scope beyond prior diffusion‑only approaches; and (2) a residual bias calculation module that distinguishes real and fake images in a single inference step, which is a significant efficiency improvement over existing methods that typically require 20+ steps.
Extensive experiments on the benchmark from 10 public datasets demonstrate that R^2BD is over 22× faster than existing reconstruction‑based methods while achieving superior detection accuracy. In cross‑dataset evaluations, it outperforms state‑of‑the‑art methods by an average of 13.87%, showing strong efficiency and generalization across diverse generative methods. The code and dataset used for evaluation are available at https://github.com/QingyuLiu/RRBD.
Authors:Yang-Che Sun, Cheng Sun, Chin-Yang Lin, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu
Abstract:
Video object segmentation methods like SAM2 achieve strong performance through memory‑based architectures but struggle under large viewpoint changes due to reliance on appearance features. Traditional 3D instance segmentation methods address viewpoint consistency but require camera poses, depth maps, and expensive preprocessing. We introduce 3AM, a training‑time enhancement that integrates 3D‑aware features from MUSt3R into SAM2. Our lightweight Feature Merger fuses multi‑level MUSt3R features that encode implicit geometric correspondence. Combined with SAM2's appearance features, the model achieves geometry‑consistent recognition grounded in both spatial position and visual similarity. We propose a field‑of‑view aware sampling strategy ensuring frames observe spatially consistent object regions for reliable 3D correspondence learning. Critically, our method requires only RGB input at inference, with no camera poses or preprocessing. On challenging datasets with wide‑baseline motion (ScanNet++, Replica), 3AM substantially outperforms SAM2 and extensions, achieving 90.6% IoU and 71.7% Tracking Recall on ScanNet++'s Selected Subset, improving over state‑of‑the‑art VOS methods by +15.9 and +30.4 points. Project page: https://jayisaking.github.io/3AM‑Page/
Authors:Zhi Qin Tan, Xiatian Zhu, Owen Addison, Yunpeng Li
Abstract:
Diagnosing dental diseases from radiographs is time‑consuming and challenging due to the subtle nature of diagnostic evidence. Existing methods, which rely on object detection models designed for natural images with more distinct target patterns, struggle to detect dental diseases that present with far less visual support. To address this challenge, we propose \bf DentalX, a novel context‑aware dental disease detection approach that leverages oral structure information to mitigate the visual ambiguity inherent in radiographs. Specifically, we introduce a structural context extraction module that learns an auxiliary task: semantic segmentation of dental anatomy. The module extracts meaningful structural context and integrates it into the primary disease detection task to enhance the detection of subtle dental diseases. Extensive experiments on a dedicated benchmark demonstrate that DentalX significantly outperforms prior methods in both tasks. This mutual benefit arises naturally during model optimization, as the correlation between the two tasks is effectively captured. Our code is available at https://github.com/zhiqin1998/DentYOLOX.
Authors:Juntao Jiang, Jiangning Zhang, Yali Bi, Jinsheng Bai, Weixuan Liu, Weiwei Jin, Zhucun Xue, Yong Liu, Xiaobin Hu, Shuicheng Yan
Abstract:
Chain‑of‑Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step‑by‑step intermediate reasoning, and recent advances have extended this paradigm to Multimodal Large Language Models (MLLMs). In the medical domain, where diagnostic decisions depend on nuanced visual cues and sequential reasoning, CoT aligns naturally with clinical thinking processes. However, current benchmarks for medical image understanding generally focus on the final answer while ignoring the reasoning path. Such opaque reasoning processes lack reliable bases for judgment, making it difficult to assist doctors in diagnosis. To address this gap, we introduce a new M3CoTBench benchmark specifically designed to evaluate the correctness, efficiency, impact, and consistency of CoT reasoning in medical image understanding. M3CoTBench features 1) a diverse, multi‑level difficulty dataset covering 24 examination types, 2) 13 varying‑difficulty tasks, 3) a suite of CoT‑specific evaluation metrics (correctness, efficiency, impact, and consistency) tailored to clinical reasoning, and 4) a performance analysis of multiple MLLMs. M3CoTBench systematically evaluates CoT reasoning across diverse medical imaging tasks, revealing current limitations of MLLMs in generating reliable and clinically interpretable reasoning, and aims to foster the development of transparent, trustworthy, and diagnostically accurate AI systems for healthcare. Project page at https://juntaojianggavin.github.io/projects/M3CoTBench/.
Authors:Naren Medarametla, Sreejon Mondal
Abstract:
Localization is a fundamental capability for autonomous robots, enabling them to operate effectively in dynamic environments. In Robocon 2025, accurate and reliable localization is crucial for improving shooting precision, avoiding collisions with other robots, and navigating the competition field efficiently. In this paper, we propose a hybrid localization algorithm that integrates classical techniques with learning based methods that rely solely on visual data from the court's floor to achieve self‑localization on the basketball field.
Authors:Shaoan Wang, Yuanfei Luo, Xingyu Chen, Aocheng Luo, Dongyue Li, Chang Liu, Sheng Chen, Yangang Zhang, Junzhi Yu
Abstract:
VLA models have shown promising potential in embodied navigation by unifying perception and planning while inheriting the strong generalization abilities of large VLMs. However, most existing VLA models rely on reactive mappings directly from observations to actions, lacking the explicit reasoning capabilities and persistent memory required for complex, long‑horizon navigation tasks. To address these challenges, we propose VLingNav, a VLA model for embodied navigation grounded in linguistic‑driven cognition. First, inspired by the dual‑process theory of human cognition, we introduce an adaptive chain‑of‑thought mechanism, which dynamically triggers explicit reasoning only when necessary, enabling the agent to fluidly switch between fast, intuitive execution and slow, deliberate planning. Second, to handle long‑horizon spatial dependencies, we develop a visual‑assisted linguistic memory module that constructs a persistent, cross‑modal semantic memory, enabling the agent to recall past observations to prevent repetitive exploration and infer movement trends for dynamic environments. For the training recipe, we construct Nav‑AdaCoT‑2.9M, the largest embodied navigation dataset with reasoning annotations to date, enriched with adaptive CoT annotations that induce a reasoning paradigm capable of adjusting both when to think and what to think about. Moreover, we incorporate an online expert‑guided reinforcement learning stage, enabling the model to surpass pure imitation learning and to acquire more robust, self‑explored navigation behaviors. Extensive experiments demonstrate that VLingNav achieves state‑of‑the‑art performance across a wide range of embodied navigation benchmarks. Notably, VLingNav transfers to real‑world robotic platforms in a zero‑shot manner, executing various navigation tasks and demonstrating strong cross‑domain and cross‑task generalization.
Authors:Renyang Liu, Kangjie Chen, Han Qiu, Jie Zhang, Kwok-Yan Lam, Tianwei Zhang, See-Kiong Ng
Abstract:
Image generation models (IGMs), while capable of producing impressive and creative content, often memorize a wide range of undesirable concepts from their training data, leading to the reproduction of unsafe content such as NSFW imagery and copyrighted artistic styles. Such behaviors pose persistent safety and compliance risks in real‑world deployments and cannot be reliably mitigated by post‑hoc filtering, owing to the limited robustness of such mechanisms and a lack of fine‑grained semantic control. Recent unlearning methods seek to erase harmful concepts at the model level, which exhibit the limitations of requiring costly retraining, degrading the quality of benign generations, or failing to withstand prompt paraphrasing and adversarial attacks. To address these challenges, we introduce SafeRedir, a lightweight inference‑time framework for robust unlearning via prompt embedding redirection. Without modifying the underlying IGMs, SafeRedir adaptively routes unsafe prompts toward safe semantic regions through token‑level interventions in the embedding space. The framework comprises two core components: a latent‑aware multi‑modal safety classifier for identifying unsafe generation trajectories, and a token‑level delta generator for precise semantic redirection, equipped with auxiliary predictors for token masking and adaptive scaling to localize and regulate the intervention. Empirical results across multiple representative unlearning tasks demonstrate that SafeRedir achieves effective unlearning capability, high semantic and perceptual preservation, robust image quality, and enhanced resistance to adversarial attacks. Furthermore, SafeRedir generalizes effectively across a variety of diffusion backbones and existing unlearned models, validating its plug‑and‑play compatibility and broad applicability. Code and data are available at https://github.com/ryliu68/SafeRedir.
Authors:Xi Chen, Hongxun Yao, Sicheng Zhao, Jiankun Zhu, Jing Jiang, Kui Jiang
Abstract:
Source‑free domain adaptation (SFDA) tackles the critical challenge of adapting source‑pretrained models to unlabeled target domains without access to source data, overcoming data privacy and storage limitations in real‑world applications. However, existing SFDA approaches struggle with the trade‑off between perception field and computational efficiency in domain‑invariant feature learning. Recently, Mamba has offered a promising solution through its selective scan mechanism, which enables long‑range dependency modeling with linear complexity. However, the Visual Mamba (i.e., VMamba) remains limited in capturing channel‑wise frequency characteristics critical for domain alignment and maintaining spatial robustness under significant domain shifts. To address these, we propose a framework called SfMamba to fully explore the stable dependency in source‑free model transfer. SfMamba introduces Channel‑wise Visual State‑Space block that enables channel‑sequence scanning for domain‑invariant feature extraction. In addition, SfMamba involves a Semantic‑Consistent Shuffle strategy that disrupts background patch sequences in 2D selective scan while preserving prediction consistency to mitigate error accumulation. Comprehensive evaluations across multiple benchmarks show that SfMamba achieves consistently stronger performance than existing methods while maintaining favorable parameter efficiency, offering a practical solution for SFDA. Our code is available at https://github.com/chenxi52/SfMamba.
Authors:Zishan Shu, Juntong Wu, Wei Yan, Xudong Liu, Hongyu Zhang, Chang Liu, Youdong Mao, Jie Chen
Abstract:
Vision modeling has advanced rapidly with Transformers, whose attention mechanisms capture visual dependencies but lack a principled account of how semantic information propagates spatially. We revisit this problem from a wave‑based perspective: feature maps are treated as spatial signals whose evolution over an internal propagation time (aligned with network depth) is governed by an underdamped wave equation. In this formulation, spatial frequency‑from low‑frequency global layout to high‑frequency edges and textures‑is modeled explicitly, and its interaction with propagation time is controlled rather than implicitly fixed. We derive a closed‑form, frequency‑time decoupled solution and implement it as the Wave Propagation Operator (WPO), a lightweight module that models global interactions in O(N log N) time‑far lower than attention. Building on WPO, we propose a family of WaveFormer models as drop‑in replacements for standard ViTs and CNNs, achieving competitive accuracy across image classification, object detection, and semantic segmentation, while delivering up to 1.6x higher throughput and 30% fewer FLOPs than attention‑based alternatives. Furthermore, our results demonstrate that wave propagation introduces a complementary modeling bias to heat‑based methods, effectively capturing both global coherence and high‑frequency details essential for rich visual semantics. Codes are available at: https://github.com/ZishanShu/WaveFormer.
Authors:Zhifan Ni, Eckehard Steinbach
Abstract:
Incomplete point clouds captured by 3D sensors often result in the loss of both geometric and semantic information. Most existing point cloud completion methods are built on rotation‑variant frameworks trained with data in canonical poses, limiting their applicability in real‑world scenarios. While data augmentation with random rotations can partially mitigate this issue, it significantly increases the learning burden and still fails to guarantee robust performance under arbitrary poses. To address this challenge, we propose the Rotation‑Equivariant Anchor Transformer (REVNET), a novel framework built upon the Vector Neuron (VN) network for robust point cloud completion under arbitrary rotations. To preserve local details, we represent partial point clouds as sets of equivariant anchors and design a VN Missing Anchor Transformer to predict the positions and features of missing anchors. Furthermore, we extend VN networks with a rotation‑equivariant bias formulation and a ZCA‑based layer normalization to improve feature expressiveness. Leveraging the flexible conversion between equivariant and invariant VN features, our model can generate point coordinates with greater stability. Experimental results show that our method outperforms state‑of‑the‑art approaches on the synthetic MVP dataset in the equivariant setting. On the real‑world KITTI dataset, REVNET delivers competitive results compared to non‑equivariant networks, without requiring input pose alignment. The source code will be released on GitHub under URL: https://github.com/nizhf/REVNET.
Authors:Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen
Abstract:
Hallucinations in video‑capable vision‑language models (Video‑VLMs) remain frequent and high‑confidence, while existing uncertainty metrics often fail to align with correctness. We introduce VideoHEDGE, a modular framework for hallucination detection in video question answering that extends entropy‑based reliability estimation from images to temporally structured inputs. Given a video‑question pair, VideoHEDGE draws a baseline answer and multiple high‑temperature generations from both clean clips and photometrically and spatiotemporally perturbed variants, then clusters the resulting textual outputs into semantic hypotheses using either Natural Language Inference (NLI)‑based or embedding‑based methods. Cluster‑level probability masses yield three reliability scores: Semantic Entropy (SE), RadFlag, and Vision‑Amplified Semantic Entropy (VASE). We evaluate VideoHEDGE on the SoccerChat benchmark using an LLM‑as‑a‑judge to obtain binary hallucination labels. Across three 7B Video‑VLMs (Qwen2‑VL, Qwen2.5‑VL, and a SoccerChat‑finetuned model), VASE consistently achieves the highest ROC‑AUC, especially at larger distortion budgets, while SE and RadFlag often operate near chance. We further show that embedding‑based clustering matches NLI‑based clustering in detection performance at substantially lower computational cost, and that domain fine‑tuning reduces hallucination frequency but yields only modest improvements in calibration. The hedge‑bench PyPI library enables reproducible and extensible benchmarking, with full code and experimental resources available at https://github.com/Simula/HEDGE#videohedge .
Authors:Tolgay Atinc Uzun, Dmitry Ignatov, Radu Timofte
Abstract:
Channel‑configuration search, the optimization of layer specifications such as channel widths in deep neural networks, presents a combinatorial challenge constrained by tensor‑shape compatibility and computational budgets. We investigate whether large language models (LLMs) can support neural architecture search (NAS) by reasoning over architectural code structures in ways that complement traditional search heuristics. We apply an LLM‑driven NAS framework to channel‑configuration search, formulating the task as conditional code generation in which the LLM refines architectural specifications using performance feedback. To address data scarcity, we generate a corpus of valid, shape‑consistent architectures through abstract syntax tree (AST) mutations. Although these mutated networks are not necessarily optimized for performance, they provide structural examples that help the LLM learn executable architectural patterns and relate channel configurations to model performance. Experimental results on CIFAR‑100 show that the closed‑loop LLM improves upon the initial AST‑generated architecture population under the same proxy‑evaluation protocol. Our analysis further shows that the generated architectures reflect domain‑specific design patterns, including non‑standard channel widths and late‑stage expansion, highlighting the potential of language‑driven design for code‑level NAS. The code and prompts are publicly available at https://github.com/ABrain‑One/NN‑GPT, and the generated deep neural networks are published at https://github.com/ABrain‑One/NN‑Dataset under model names with the prefix ast‑dimension‑.
Authors:Aditya Chaudhary, Sneha Barman, Mainak Singha, Ankit Jha, Girish Mishra, Biplab Banerjee
Abstract:
In this paper, we propose a novel multimodal framework, Multimodal Language‑Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision‑language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic‑level understanding. MMLGNet employs modality‑specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi‑directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high‑dimensional remote sensing data and language‑guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN‑based encoders, outperforming several established multimodal visual‑only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.
Authors:Yuan Gao, Di Cao, Xiaohuan Xi, Sheng Nie, Shaobo Xia, Cheng Wang
Abstract:
Semantic segmentation of 3D geospatial point clouds is fundamental to remote sensing applications, yet domain shifts caused by regional and acquisition‑related variations often degrade model performance. Although domain adaptation can mitigate such shifts, existing methods typically require access to source‑domain data, which is often infeasible due to privacy concerns and regulatory policies. To address this, we propose LoGo (Local‑Global Dual‑Consensus), a novel source‑free unsupervised domain adaptation (SFUDA) framework requiring only a pretrained model and unlabeled target data. At the local level, we introduce a class‑balanced prototype estimation module that ensures that robust feature prototypes can be generated even for sample‑scarce tail classes, effectively mitigating the feature collapse caused by long‑tailed distributions. At the global level, we introduce an optimal transport‑based global distribution alignment module that formulates pseudo‑label assignment as a global optimization problem, effectively correcting the over‑dominance of head classes inherent in local greedy assignments, and thereby preventing model predictions from being severely biased towards majority classes. Finally, we propose a dual‑consistency pseudo‑label filtering mechanism that retains only high‑confidence pseudo‑labels where local multi‑augmented ensemble predictions align with global optimal transport assignments for self‑training. Extensive experiments on two challenging benchmarks, encompassing cross‑scene and cross‑sensor settings, demonstrate that LoGo consistently outperforms existing state‑of‑the‑art methods. The source code is available at https://github.com/GYproject/LoGo‑SFUDA.
Authors:Dongting Hu, Aarush Gupta, Magzhan Gabidolla, Arpit Sahni, Huseyin Coskun, Yanyu Li, Yerlan Idelbayev, Ahsan Mahmood, Aleksei Lebedev, Dishani Lahiri, Anujraaj Goyal, Ju Hu, Mingming Gong, Sergey Tulyakov, Anil Kag
Abstract:
Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on‑device deployment due to their high computational and memory costs. In this work, we present an efficient DiT framework tailored for mobile and edge devices that achieves transformer‑level generation quality under strict resource constraints. Our design combines three key components. First, we propose a compact DiT architecture with an adaptive global‑local sparse attention mechanism that balances global context modeling and local detail preservation. Second, we propose an elastic training framework that jointly optimizes sub‑DiTs of varying capacities within a unified supernetwork, allowing a single model to dynamically adjust for efficient inference across different hardware. Finally, we develop Knowledge‑Guided Distribution Matching Distillation, a step‑distillation pipeline that integrates the DMD objective with knowledge transfer from few‑step teacher models, producing high‑fidelity and low‑latency generation (e.g., 4‑step) suitable for real‑time on‑device use. Together, these contributions enable scalable, efficient, and high‑quality diffusion models for deployment on diverse hardware.
Authors:Taminul Islam, Toqi Tahamid Sarker, Mohamed Embaby, Khaled R Ahmed, Amer AbuGhazaleh
Abstract:
Ruminal acidosis is a prevalent metabolic disorder in dairy cattle causing significant economic losses and animal welfare concerns. Current diagnostic methods rely on invasive pH measurement, limiting scalability for continuous monitoring. We present FUME (Fused Unified Multi‑gas Emission Network), the first deep learning approach for rumen acidosis detection from dual‑gas optical imaging under in vitro conditions. Our method leverages complementary carbon dioxide (CO2) and methane (CH4) emission patterns captured by infrared cameras to classify rumen health into Healthy, Transitional, and Acidotic states. FUME employs a lightweight dual‑stream architecture with weight‑shared encoders, modality‑specific self‑attention, and channel attention fusion, jointly optimizing gas plume segmentation and classification of dairy cattle health. We introduce the first dual‑gas OGI dataset comprising 8,967 annotated frames across six pH levels with pixel‑level segmentation masks. Experiments demonstrate that FUME achieves 80.99% mIoU and 98.82% classification accuracy while using only 1.28M parameters and 1.97G MACs‑‑outperforming state‑of‑the‑art methods in segmentation quality with 10x lower computational cost. Ablation studies reveal that CO2 provides the primary discriminative signal and dual‑task learning is essential for optimal performance. Our work establishes the feasibility of gas emission‑based livestock health monitoring, paving the way for practical, in vitro acidosis detection systems. Codes are available at https://github.com/taminulislam/fume.
Authors:Md. Faiyaz Abdullah Sayeedi, Rashedur Rahman, Siam Tahsin Bhuiyan, Sefatul Wasi, Ashraful Islam, Saadia Binte Alam, AKM Mahbubur Rahman
Abstract:
Medical image analysis increasingly relies on large vision‑language models (VLMs), yet most systems remain single‑pass black boxes that offer limited control over reasoning, safety, and spatial grounding. We propose R^4, an agentic framework that decomposes medical imaging workflows into four coordinated agents: a Router that configures task‑ and specialization‑aware prompts from the image, patient history, and metadata; a Retriever that uses exemplar memory and pass@k sampling to jointly generate free‑text reports and bounding boxes; a Reflector that critiques each draft‑box pair for key clinical error modes (negation, laterality, unsupported claims, contradictions, missing findings, and localization errors); and a Repairer that iteratively revises both narrative and spatial outputs under targeted constraints while curating high‑quality exemplars for future cases. Instantiated on chest X‑ray analysis with multiple modern VLM backbones and evaluated on report generation and weakly supervised detection, R^4 consistently boosts LLM‑as‑a‑Judge scores by roughly +1.7‑+2.5 points and mAP50 by +2.5‑+3.5 absolute points over strong single‑VLM baselines, without any gradient‑based fine‑tuning. These results show that agentic routing, reflection, and repair can turn strong but brittle VLMs into more reliable and better grounded tools for clinical image interpretation. Our code can be found at: https://github.com/faiyazabdullah/MultimodalMedAgent
Authors:Yan Zhu, Te Luo, Pei-Yao Fu, Zhen Zhang, Zi-Long Wang, Yi-Fan Qu, Zi-Han Geng, Jia-Qi Xu, Lu Yao, Li-Yun Ma, Wei Su, Wei-Feng Chen, Quan-Lin Li, Shuo Wang, Ping-Hong Zhou
Abstract:
Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state‑of‑the‑art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI‑Bench, a benchmark encompassing 20 fine‑grained lesion categories. Twelve MLLMs were evaluated across a five‑stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro‑F1, mean Intersection‑over‑Union (mIoU), and multi‑dimensional Likert scale. Gemini‑3‑Pro achieved state‑of‑the‑art performance. In diagnostic reasoning, top‑tier models (Macro‑F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p<0.05). Furthermore, qualitative analysis revealed a "fluency‑accuracy paradox": models generated reports with superior linguistic readability compared with humans (p<0.05) but exhibited significantly lower factual correctness (p<0.05) due to "over‑interpretation" and hallucination of visual features. GI‑Bench maintains a dynamic leaderboard that tracks the evolving performance of MLLMs in clinical endoscopy. The current rankings and benchmark results are available at https://roterdl.github.io/GIBench/.
Authors:Anh H. Vo, Tae-Seok Kim, Hulin Jin, Soo-Mi Choi, Yong-Guk Kim
Abstract:
A 3D avatar typically has one of six cardinal facial expressions. To simulate realistic emotional variation, we should be able to render a facial transition between two arbitrary expressions. This study presents a new framework for instruction‑driven facial expression generation that produces a 3D face and, starting from an image of the face, transforms the facial expression from one designated facial expression to another. The Instruction‑driven Facial Expression Decomposer (IFED) module is introduced to facilitate multimodal data learning and capture the correlation between textual descriptions and facial expression features. Subsequently, we propose the Instruction to Facial Expression Transition (I2FET) method, which leverages IFED and a vertex reconstruction loss function to refine the semantic comprehension of latent vectors, thus generating a facial expression sequence according to the given instruction. Lastly, we present the Facial Expression Transition model to generate smooth transitions between facial expressions. Extensive evaluation suggests that the proposed model outperforms state‑of‑the‑art methods on the CK+ and CelebV‑HQ datasets. The results show that our framework can generate facial expression trajectories according to text instruction. Considering that text prompts allow us to make diverse descriptions of human emotional states, the repertoire of facial expressions and the transitions between them can be expanded greatly. We expect our framework to find various practical applications More information about our project can be found at https://vohoanganh.github.io/tg3dfet/
Authors:Feiran Wang, Junyi Wu, Dawen Cai, Yuan Hong, Yan Yan
Abstract:
We present CogniMap3D, a bioinspired framework for dynamic 3D scene understanding and reconstruction that emulates human cognitive processes. Our approach maintains a persistent memory bank of static scenes, enabling efficient spatial knowledge storage and rapid retrieval. CogniMap3D integrates three core capabilities: a multi‑stage motion cue framework for identifying dynamic objects, a cognitive mapping system for storing, recalling, and updating static scenes across multiple visits, and a factor graph optimization strategy for refining camera poses. Given an image stream, our model identifies dynamic regions through motion cues with depth and camera pose priors, then matches static elements against its memory bank. When revisiting familiar locations, CogniMap3D retrieves stored scenes, relocates cameras, and updates memory with new observations. Evaluations on video depth estimation, camera pose reconstruction, and 3D mapping tasks demonstrate its state‑of‑the‑art performance, while effectively supporting continuous scene understanding across extended sequences and multiple visits.
Authors:Guoping Xu, Jayaram K. Udupa, Weiguo Lu, You Zhang
Abstract:
Deep learning‑based automatic medical image segmentation plays a critical role in clinical diagnosis and treatment planning but remains challenging in few‑shot scenarios due to the scarcity of annotated training data. Recently, self‑supervised foundation models such as DINOv3, which were trained on large natural image datasets, have shown strong potential for dense feature extraction that can help with the few‑shot learning challenge. Yet, their direct application to medical images is hindered by domain differences. In this work, we propose DINO‑AugSeg, a novel framework that leverages DINOv3 features to address the few‑shot medical image segmentation challenge. Specifically, we introduce WT‑Aug, a wavelet‑based feature‑level augmentation module that enriches the diversity of DINOv3‑extracted features by perturbing frequency components, and CG‑Fuse, a contextual information‑guided fusion module that exploits cross‑attention to integrate semantic‑rich low‑resolution features with spatially detailed high‑resolution features. Extensive experiments on six public benchmarks spanning five imaging modalities, including MRI, CT, ultrasound, endoscopy, and dermoscopy, demonstrate that DINO‑AugSeg consistently outperforms existing methods under limited‑sample conditions. The results highlight the effectiveness of incorporating wavelet‑domain augmentation and contextual fusion for robust feature representation, suggesting DINO‑AugSeg as a promising direction for advancing few‑shot medical image segmentation. Code and data will be made available on https://github.com/apple1986/DINO‑AugSeg.
Authors:Samet Hicsonmez, Abd El Rahman Shabayek, Djamila Aouada
Abstract:
Zero‑Shot image Anomaly Detection (ZSAD) aims to detect and localise anomalies without access to any normal training samples of the target data. While recent ZSAD approaches leverage additional modalities such as language to generate fine‑grained prompts for localisation, vision‑only methods remain limited to image‑level classification, lacking spatial precision. In this work, we introduce a simple yet effective training‑free vision‑only ZSAD framework that circumvents the need for fine‑grained prompts by leveraging the inversion of a pretrained Denoising Diffusion Implicit Model (DDIM). Specifically, given an input image and a generic text description (e.g., "an image of an [object class]"), we invert the image to obtain latent representations and initiate the denoising process from a fixed intermediate timestep to reconstruct the image. Since the underlying diffusion model is trained solely on normal data, this process yields a normal‑looking reconstruction. The discrepancy between the input image and the reconstructed one highlights potential anomalies. Our method achieves state‑of‑the‑art performance on VISA dataset, demonstrating strong localisation capabilities without auxiliary modalities and facilitating a shift away from prompt dependence for zero‑shot anomaly detection research. Code is available at https://github.com/giddyyupp/DIVAD.
Authors:Haorui Yu, Diji Yang, Hang He, Fengrui Zhang, Qiufeng Yi
Abstract:
We introduce VULCA‑Bench, a multicultural art‑critique benchmark for evaluating Vision‑Language Models' (VLMs) cultural understanding beyond surface‑level visual perception. Existing VLM benchmarks predominantly measure L1‑L2 capabilities (object recognition, scene description, and factual question answering) while under‑evaluate higher‑order cultural interpretation. VULCA‑Bench contains 7,410 matched image‑critique pairs spanning eight cultural traditions, with Chinese‑English bilingual coverage. We operationalise cultural understanding using a five‑layer framework (L1‑L5, from Visual Perception to Philosophical Aesthetics), instantiated as 225 culture‑specific dimensions and supported by expert‑written bilingual critiques. Our pilot results indicate that higher‑layer reasoning (L3‑L5) is consistently more challenging than visual and technical analysis (L1‑L2). The dataset, evaluation scripts, and annotation tools are available under CC BY 4.0 at https://github.com/yha9806/VULCA‑Bench.
Authors:Maxwell Jones, Rameen Abdal, Or Patashnik, Ruslan Salakhutdinov, Sergey Tulyakov, Jun-Yan Zhu, Kuan-Chieh Jackson Wang
Abstract:
We present RefVFX, a new framework that transfers complex temporal effects from a reference video onto a target video or image in a feed‑forward manner. While existing methods excel at prompt‑based or keyframe‑conditioned editing, they struggle with dynamic temporal effects such as dynamic lighting changes or character transformations, which are difficult to describe via text or static conditions. Transferring a video effect is challenging, as the model must integrate the new temporal dynamics with the input video's existing motion and appearance. % To address this, we introduce a large‑scale dataset of triplets, where each triplet consists of a reference effect video, an input image or video, and a corresponding output video depicting the transferred effect. Creating this data is non‑trivial, especially the video‑to‑video effect triplets, which do not exist naturally. To generate these, we propose a scalable automated pipeline that creates high‑quality paired videos designed to preserve the input's motion and structure while transforming it based on some fixed, repeatable effect. We then augment this data with image‑to‑video effects derived from LoRA adapters and code‑based temporal effects generated through programmatic composition. Building on our new dataset, we train our reference‑conditioned model using recent text‑to‑video backbones. Experimental results demonstrate that RefVFX produces visually consistent and temporally coherent edits, generalizes across unseen effect categories, and outperforms prompt‑only baselines in both quantitative metrics and human preference. See our website at https://snap‑research.github.io/RefVFX/
Authors:Kewei Zhang, Ye Huang, Yufan Deng, Jincheng Yu, Junsong Chen, Huan Ling, Enze Xie, Daquan Zhou
Abstract:
While the Transformer architecture dominates many fields, its quadratic self‑attention complexity hinders its use in large‑scale applications. Linear attention offers an efficient alternative, but its direct application often degrades performance, with existing fixes typically re‑introducing computational overhead through extra modules (e.g., depthwise separable convolution) that defeat the original purpose. In this work, we identify a key failure mode in these methods: global context collapse, where the model loses representational diversity. To address this, we propose Multi‑Head Linear Attention (MHLA), which preserves this diversity by computing attention within divided heads along the token dimension. We prove that MHLA maintains linear complexity while recovering much of the expressive power of softmax attention, and verify its effectiveness across multiple domains, achieving a 3.6% improvement on ImageNet classification, a 6.3% gain on NLP, a 12.6% improvement on image generation, and a 41% enhancement on video generation under the same time complexity.
Authors:Anurag Das, Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Bernt Schiele, Georgios Tzimiropoulos, Brais Martinez
Abstract:
Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities, yet their proficiency in understanding and reasoning over multiple images remains largely unexplored. While existing benchmarks have initiated the evaluation of multi‑image models, a comprehensive analysis of their core weaknesses and their causes is still lacking. In this work, we introduce MIMIC (Multi‑Image Model Insights and Challenges), a new benchmark designed to rigorously evaluate the multi‑image capabilities of LVLMs. Using MIMIC, we conduct a series of diagnostic experiments that reveal pervasive issues: LVLMs often fail to aggregate information across images and struggle to track or attend to multiple concepts simultaneously. To address these failures, we propose two novel complementary remedies. On the data side, we present a procedural data‑generation strategy that composes single‑image annotations into rich, targeted multi‑image training examples. On the optimization side, we analyze layer‑wise attention patterns and derive an attention‑masking scheme tailored for multi‑image inputs. Experiments substantially improved cross‑image aggregation, while also enhancing performance on existing multi‑image benchmarks, outperforming prior state of the art across tasks. Data and code will be made available at https://github.com/anurag‑198/MIMIC.
Authors:Sijun Dong, Siming Fu, Kaiyu Li, Xiangyong Cao, Xiaoliang Meng, Bo Du
Abstract:
Remote sensing change detection fundamentally relies on the effective fusion and discrimination of bi‑temporal features. Prevailing paradigms typically utilize Siamese encoders bridged by explicit difference computation modules, such as subtraction or concatenation, to identify changes. In this work, we challenge this complexity with SEED (Siamese Encoder‑Exchange‑Decoder), a streamlined paradigm that replaces explicit differencing with parameter‑free feature exchange. By sharing weights across both Siamese encoders and decoders, SEED effectively operates as a single parameter set model. Theoretically, we formalize feature exchange as an orthogonal permutation operator and prove that, under pixel consistency, this mechanism preserves mutual information and Bayes optimal risk, whereas common arithmetic fusion methods often introduce information loss. Extensive experiments across five benchmarks, including SYSU‑CD, LEVIR‑CD, PX‑CLCD, WaterCD, and CDD, and three backbones, namely SwinT, EfficientNet, and ResNet, demonstrate that SEED matches or surpasses state of the art methods despite its simplicity. Furthermore, we reveal that standard semantic segmentation models can be transformed into competitive change detectors solely by inserting this exchange mechanism, referred to as SEG2CD. The proposed paradigm offers a robust, unified, and interpretable framework for change detection, demonstrating that simple feature exchange is sufficient for high performance information fusion. Code and full training and evaluation protocols will be released at https://github.com/dyzy41/open‑rscd.
Authors:Lingchen Sun, Rongyuan Wu, Zhengqiang Zhang, Ruibin Li, Yujing Sun, Shuaizheng Liu, Lei Zhang
Abstract:
Recent works such as REPA have shown that guiding diffusion models with external semantic features (e.g., DINO) can significantly accelerate the training of diffusion transformers (DiTs). However, the use of pretrained external features as guidance signals introduces additional dependencies. We argue that DiTs actually have the power to guide the training of themselves, and propose SelfTranscendence, an effective method that achieves fast convergence using internal feature supervision only. The desired internal guidance features should meet two requirements: structurally clean to help shallow blocks separate noise from signal, and semantically discriminative to help shallow layers learn effective representations. With this consideration, we first align the DiT features with the clean VAE latent features, a native component of latent diffusion, for a short training phase (e.g., 40 epochs) to improve their structural representations, then apply the classifier‑free guidance to the intermediate features, enhancing their discriminative capability and semantic expressiveness. These enriched internal features, learned entirely within the model, are used as supervision signals to guide a new DiT training from scratch. Compared to existing self‑contained methods, our approach achieves a significant performance boost. It can even surpass REPA, which uses the external DINO features as guidance, in both generation quality and convergence speed for both class‑to‑image and text‑to‑image generation tasks. The source code of our method can be found at https://github.com/csslc/Self‑Transcendence.
Authors:Nicolas Sereyjol-Garros, Ellington Kirby, Victor Besnier, Nermin Samet
Abstract:
LiDAR scene synthesis is an emerging solution to scarcity in 3D data for robotic tasks such as autonomous driving. Recent approaches employ diffusion or flow matching models to generate realistic scenes, but 3D data remains limited compared to RGB datasets with millions of samples. We introduce R3DPA, the first LiDAR scene generation method to unlock image‑pretrained priors for LiDAR point clouds, and leverage self‑supervised 3D representations for state‑of‑the‑art results. Specifically, we (i) align intermediate features of our generative model with self‑supervised 3D features, which substantially improves generation quality; (ii) transfer knowledge from large‑scale image‑pretrained generative models to LiDAR generation, mitigating limited LiDAR datasets; and (iii) enable point cloud control at inference for object inpainting and scene mixing with solely an unconditional model. On the KITTI‑360 benchmark R3DPA achieves state of the art performance. Code and pretrained models are available at https://github.com/valeoai/R3DPA.
Authors:Zijian Wu, Boyao Zhou, Liangxiao Hu, Hongyu Liu, Yuan Sun, Xuan Wang, Xun Cao, Yujun Shen, Hao Zhu
Abstract:
We present UIKA, a feed‑forward animatable Gaussian head model from an arbitrary number of pose‑free inputs, including a single image, multi‑view captures, and smartphone‑captured videos. Unlike the traditional avatar method, which requires a studio‑level multi‑view capture system and reconstructs a human‑specific model through a long‑time optimization process, we rethink the task through the lenses of model representation, network design, and data preparation. First, we introduce a UV‑guided avatar modeling strategy, in which each input image is associated with a pixel‑wise facial correspondence estimation. Such correspondence estimation allows us to reproject each valid pixel color from screen space to UV space, which is independent of camera pose and character expression. Furthermore, we design learnable UV tokens on which the attention mechanism can be applied at both the screen and UV levels. The learned UV tokens can be decoded into canonical Gaussian attributes using aggregated UV information from all input views. To train our large avatar model, we additionally prepare a large‑scale, identity‑rich synthetic training dataset. Our method significantly outperforms existing approaches in both monocular and multi‑view settings.
Authors:Ahmad AlMughrabi, Guillermo Rivo, Carlos Jiménez-Farfán, Umair Haroon, Farid Al-Areqi, Hyunjun Jung, Benjamin Busam, Ricardo Marques, Petia Radeva
Abstract:
Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi‑view data and poor generalization to new viewpoints. We introduce BenchSeg, a novel multi‑view food video segmentation dataset and benchmark. BenchSeg aggregates 55 dish scenes (from Nutrition5k, Vegetables & Fruits, MetaFood3D, and FoodKit) with 25,284 meticulously annotated frames, capturing each dish under free 360° camera motion. We evaluate a diverse set of 20 state‑of‑the‑art segmentation models (e.g., SAM‑based, transformer, CNN, and large multimodal) on the existing FoodSeg103 dataset and evaluate them (alone and combined with video‑memory modules) on BenchSeg. Quantitative and qualitative results demonstrate that while standard image segmenters degrade sharply under novel viewpoints, memory‑augmented methods maintain temporal consistency across frames. Our best model based on a combination of SeTR‑MLA+XMem2 outperforms prior work (e.g., improving over FoodMem by ~2.63% mAP), offering new insights into food segmentation and tracking for dietary analysis. In addition to frame‑wise spatial accuracy, we introduce a dedicated temporal evaluation protocol that explicitly quantifies segmentation stability over time through continuity, flicker rate, and IoU drift metrics. This allows us to reveal failure modes that remain invisible under standard per‑frame evaluations. We release BenchSeg to foster future research. The project page including the dataset annotations and the food segmentation models can be found at https://amughrabi.github.io/benchseg.
Authors:Alvaro Becerra, Ruth Cobos, Roberto Daza
Abstract:
Oral presentation skills are a critical component of higher education, yet comprehensive datasets capturing real‑world student performance across multiple modalities remain scarce. To address this gap, we present SOPHIAS (Student Oral Presentation monitoring for Holistic Insights & Analytics using Sensors), a 12‑hour multimodal dataset containing recordings of 50 oral presentations (10‑15‑minute presentation followed by 5‑15‑minute Q&A) delivered by 65 undergraduate and master's students at the Universidad Autonoma de Madrid. SOPHIAS integrates eight synchronized sensor streams from high‑definition webcams, ambient and webcam audio, eye‑tracking glasses, smartwatch physiological sensors, and clicker, keyboard, and mouse interactions. In addition, the dataset includes slides and rubric‑based evaluations from teachers, peers, and self‑assessments, along with timestamped contextual annotations. The dataset captures presentations conducted in real classroom settings, preserving authentic student behaviors, interactions, and physiological responses. SOPHIAS enables the exploration of relationships between multimodal behavioral and physiological signals and presentation performance, supports the study of peer assessment, and provides a benchmark for developing automated feedback and Multimodal Learning Analytics tools. The dataset is publicly available for research through GitHub and Science Data Bank.
Authors:Bing Yu, Liu Shi, Haitao Wang, Deran Qi, Xiang Cai, Wei Zhong, Qiegen Liu
Abstract:
Accurate three‑dimensional (3D) tooth segmentation from Cone‑Beam Computed Tomography (CBCT) is a prerequisite for digital dental workflows. However, achieving high‑fidelity segmentation remains challenging due to adhesion artifacts in naturally occluded scans, which are caused by low contrast and indistinct inter‑arch boundaries. To address these limitations, we propose the Anatomy Aware Cascade Network (AACNet), a coarse‑to‑fine framework designed to resolve boundary ambiguity while maintaining global structural consistency. Specifically, we introduce two mechanisms: the Ambiguity Gated Boundary Refiner (AGBR) and the Signed Distance Map guided Anatomical Attention (SDMAA). The AGBR employs an entropy based gating mechanism to perform targeted feature rectification in high uncertainty transition zones. Meanwhile, the SDMAA integrates implicit geometric constraints via signed distance map to enforce topological consistency, preventing the loss of spatial details associated with standard pooling. Experimental results on a dataset of 125 CBCT volumes demonstrate that AACNet achieves a Dice Similarity Coefficient of 90.17 % and a 95% Hausdorff Distance of 3.63 mm, significantly outperforming state‑of‑the‑art methods. Furthermore, the model exhibits strong generalization on an external dataset with an HD95 of 2.19 mm, validating its reliability for downstream clinical applications such as surgical planning. Code for AACNet is available at https://github.com/shiliu0114/AACNet.
Authors:Mahdi Chamseddine, Didier Stricker, Jason Rambach
Abstract:
Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates the pre‑trained Segment Anything (SAM) encoder to make use of its extensive training and integrate it into a semantic segmentation model for panoramic images using multiple modalities. We modify the SAM encoder to output multi‑stage features and introduce a novel spatio‑modal fusion module that allows the model to select the relevant modalities and best features from each modality for different areas of the input. Furthermore, our semantic decoder uses spherical attention and dual view fusion to overcome the distortions and edge discontinuity often associated with panoramic images. PanoSAMic achieves state‑of‑the‑art (SotA) results on Stanford2D3DS for RGB, RGB‑D, and RGB‑D‑N modalities and on Matterport3D for RGB and RGB‑D modalities. https://github.com/dfki‑av/PanoSAMic
Authors:Prachet Dev Singh, Shyamsundar Paramasivam, Sneha Barman, Mainak Singha, Ankit Jha, Girish Mishra, Biplab Banerjee
Abstract:
Hyperspectral image (HSI) classification presents unique challenges due to its high spectral dimensionality and limited labeled data. Traditional deep learning models often suffer from overfitting and high computational costs. Self‑distillation (SD), a variant of knowledge distillation where a network learns from its own predictions, has recently emerged as a promising strategy to enhance model performance without requiring external teacher networks. In this work, we explore the application of SD to HSI by treating earlier outputs as soft targets, thereby enforcing consistency between intermediate and final predictions. This process improves intra‑class compactness and inter‑class separability in the learned feature space. Our approach is validated on two benchmark HSI datasets and demonstrates significant improvements in classification accuracy and robustness, highlighting the effectiveness of SD for spectral‑spatial learning. Codes are available at https://github.com/Prachet‑Dev‑Singh/SDHSI.
Authors:Jiao Xu, Xin Chen, Lihe Zhang
Abstract:
In this paper, we present a new dynamic collaborative network for semi‑supervised 3D vessel segmentation, termed DiCo. Conventional mean teacher (MT) methods typically employ a static approach, where the roles of the teacher and student models are fixed. However, due to the complexity of 3D vessel data, the teacher model may not always outperform the student model, leading to cognitive biases that can limit performance. To address this issue, we propose a dynamic collaborative network that allows the two models to dynamically switch their teacher‑student roles. Additionally, we introduce a multi‑view integration module to capture various perspectives of the inputs, mirroring the way doctors conduct medical analysis. We also incorporate adversarial supervision to constrain the shape of the segmented vessels in unlabeled data. In this process, the 3D volume is projected into 2D views to mitigate the impact of label inconsistencies. Experiments demonstrate that our DiCo method sets new state‑of‑the‑art performance on three 3D vessel segmentation benchmarks. The code repository address is https://github.com/xujiaommcome/DiCo
Authors:Mohit Jaiswal, Naman Jain, Shivani Pathak, Mainak Singha, Nikunja Bihari Kar, Ankit Jha, Biplab Banerjee
Abstract:
Few‑shot remote sensing image classification is challenging due to limited labeled samples and high variability in land‑cover types. We propose a reconstruction‑guided few‑shot network (RGFS‑Net) that enhances generalization to unseen classes while preserving consistency for seen categories. Our method incorporates a masked image reconstruction task, where parts of the input are occluded and reconstructed to encourage semantically rich feature learning. This auxiliary task strengthens spatial understanding and improves class discrimination under low‑data settings. We evaluated the efficacy of EuroSAT and PatternNet datasets under 1‑shot and 5‑shot protocols, our approach consistently outperforms existing baselines. The proposed method is simple, effective, and compatible with standard backbones, offering a robust solution for few‑shot remote sensing classification. Codes are available at https://github.com/stark0908/RGFS.
Authors:Zhongming Liu, Bingbing Jiang
Abstract:
Attention mechanisms have become a core component of deep learning models, with Channel Attention and Spatial Attention being the two most representative architectures. Current research on their fusion strategies primarily bifurcates into sequential and parallel paradigms, yet the selection process remains largely empirical, lacking systematic analysis and unified principles. We systematically compare channel‑spatial attention combinations under a unified framework, building an evaluation suite of 18 topologies across four classes: sequential, parallel, multi‑scale, and residual. Across two vision and nine medical datasets, we uncover a "data scale‑method‑performance" coupling law: (1) in few‑shot tasks, the "Channel‑Multi‑scale Spatial" cascaded structure achieves optimal performance; (2) in medium‑scale tasks, parallel learnable fusion architectures demonstrate superior results; (3) in large‑scale tasks, parallel structures with dynamic gating yield the best performance. Additionally, experiments indicate that the "Spatial‑Channel" order is more stable and effective for fine‑grained classification, while residual connections mitigate vanishing gradient problems across varying data scales. We thus propose scenario‑based guidelines for building future attention modules. Code is open‑sourced at https://github.com/DWlzm.
Authors:Weidong Tang, Xinyan Wan, Siyu Li, Xiumei Wang
Abstract:
While inference‑time scaling has significantly enhanced generative quality in large language and diffusion models, its application to vector‑quantized (VQ) visual autoregressive modeling (VAR) remains unexplored. We introduce VAR‑Scaling, the first general framework for inference‑time scaling in VAR, addressing the critical challenge of discrete latent spaces that prohibit continuous path search. We find that VAR scales exhibit two distinct pattern types: general patterns and specific patterns, where later‑stage specific patterns conditionally optimize early‑stage general patterns. To overcome the discrete latent space barrier in VQ models, we map sampling spaces to quasi‑continuous feature spaces via kernel density estimation (KDE), where high‑density samples approximate stable, high‑quality solutions. This transformation enables effective navigation of sampling distributions. We propose a density‑adaptive hybrid sampling strategy: Top‑k sampling focuses on high‑density regions to preserve quality near distribution modes, while Random‑k sampling explores low‑density areas to maintain diversity and prevent premature convergence. Consequently, VAR‑Scaling optimizes sample fidelity at critical scales to enhance output quality. Experiments in class‑conditional and text‑to‑image evaluations demonstrate significant improvements in inference process. The code is available at https://github.com/WD7ang/VAR‑Scaling.
Authors:Taekbeom Lee, Dabin Kim, Youngseok Jang, H. Jin Kim
Abstract:
We present HERE, an active 3D scene reconstruction framework based on neural radiance fields, enabling high‑fidelity implicit mapping. Our approach centers around an active learning strategy for camera trajectory generation, driven by accurate identification of unseen regions, which supports efficient data acquisition and precise scene reconstruction. The key to our approach is epistemic uncertainty quantification based on evidential deep learning, which directly captures data insufficiency and exhibits a strong correlation with reconstruction errors. This allows our framework to more reliably identify unexplored or poorly reconstructed regions compared to existing methods, leading to more informed and targeted exploration. Additionally, we design a hierarchical exploration strategy that leverages learned epistemic uncertainty, where local planning extracts target viewpoints from high‑uncertainty voxels based on visibility for trajectory generation, and global planning uses uncertainty to guide large‑scale coverage for efficient and comprehensive reconstruction. The effectiveness of the proposed method in active 3D reconstruction is demonstrated by achieving higher reconstruction completeness compared to previous approaches on photorealistic simulated scenes across varying scales, while a hardware demonstration further validates its real‑world applicability. Project page: https://taekbum.github.io/here/
Authors:Yuetao Li, Zhizhou Jia, Yu Zhang, Qun Hao, Shaohui Zhang
Abstract:
Autonomous high‑fidelity object reconstruction is fundamental for creating digital assets and bridging the simulation‑to‑reality gap in robotics. We present ObjSplat, an active reconstruction framework that leverages Gaussian surfels as a unified representation to progressively reconstruct unknown objects with both photorealistic appearance and accurate geometry. Addressing the limitations of conventional opacity or depth‑based cues, we introduce a geometry‑aware viewpoint evaluation pipeline that explicitly models back‑face visibility and occlusion‑aware multi‑view covisibility, reliably identifying under‑reconstructed regions even on geometrically complex objects. Furthermore, to overcome the limitations of greedy planning strategies, ObjSplat employs a next‑best‑path (NBP) planner that performs multi‑step lookahead on a dynamically constructed spatial graph. By jointly optimizing information gain and movement cost, this planner generates globally efficient trajectories. Extensive experiments in simulation and on real‑world cultural artifacts demonstrate that ObjSplat produces physically consistent models within minutes, achieving superior reconstruction fidelity and surface completeness while significantly reducing scan time and path length compared to state‑of‑the‑art approaches. Project page: https://li‑yuetao.github.io/ObjSplat‑page/ .
Authors:Jie Zhu, Yiyang Su, Xiaoming Liu
Abstract:
Multi‑modal large language models (MLLMs) exhibit strong general‑purpose capabilities, yet still struggle on Fine‑Grained Visual Classification (FGVC), a core perception task that requires subtle visual discrimination and is crucial for many real‑world applications. A widely adopted strategy for boosting performance on challenging tasks such as math and coding is Chain‑of‑Thought (CoT) reasoning. However, several prior works have reported that CoT can actually harm performance on visual perception tasks. These studies, though, examine the issue from relatively narrow angles and leave open why CoT degrades perception‑heavy performance. We systematically re‑examine the role of CoT in FGVC through the lenses of zero‑shot evaluation and multiple training paradigms. Across these settings, we uncover a central paradox: the degradation induced by CoT is largely driven by the reasoning length, in which longer textual reasoning consistently lowers classification accuracy. We term this phenomenon the ``Cost of Thinking''. Building on this finding, we make two key contributions: (1) MRN, a simple and general plug‑and‑play normalization method for multi‑reward optimization that balances heterogeneous reward signals, and (2) ReFine‑RFT, a framework that combines ensemble rewards with MRN to constrain reasoning length while providing dense accuracy‑oriented feedback. Extensive experiments demonstrate the effectiveness of our findings and the proposed ReFine‑RFT, achieving state‑of‑the‑art performance across FGVC benchmarks. Project page: \hrefhttps://refine‑rft.github.io/ReFine‑RFT.
Authors:Yuhang Su, Mei Wang, Yaoyao Zhong, Guozhang Li, Shixing Li, Yihan Feng, Hua Huang
Abstract:
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual understanding, they often struggle when faced with the unstructured and ambiguous nature of human‑generated sketches. This limitation is particularly pronounced in the underexplored task of visual grading, where models should not only solve a problem but also diagnose errors in hand‑drawn diagrams. Such diagnostic capabilities depend on complex structural, semantic, and metacognitive reasoning. To bridge this gap, we introduce SketchJudge, a novel benchmark tailored for evaluating MLLMs as graders of hand‑drawn STEM diagrams. SketchJudge encompasses 1,015 hand‑drawn student responses across four domains: geometry, physics, charts, and flowcharts, featuring diverse stylistic variations and distinct error types. Evaluations on SketchJudge demonstrate that even advanced MLLMs lag significantly behind humans, validating the benchmark's effectiveness in exposing the fragility of current vision‑language alignment in symbolic and noisy contexts. All data, code, and evaluation scripts are publicly available at https://github.com/yuhangsu82/SketchJudge.
Authors:Zengyuan Zuo, Junjun Jiang, Gang Wu, Xianming Liu
Abstract:
Image dehazing has witnessed significant advancements with the development of deep learning models. However, most existing methods focus solely on single‑modal RGB features, neglecting the inherent correlation between scene depth and haze distribution. Even those that jointly optimize depth estimation and image dehazing often suffer from suboptimal performance due to inadequate utilization of accurate depth information. In this paper, we present UDPNet, a general framework that leverages depth‑based priors from a large‑scale pretrained depth estimation model DepthAnything V2 to boost existing image dehazing models. Specifically, our architecture comprises two key components: the Depth‑Guided Attention Module (DGAM) adaptively modulates features via lightweight depth‑guided channel attention, and the Depth Prior Fusion Module (DPFM) enables hierarchical fusion of multi‑scale depth map features by dual sliding‑window multi‑head cross‑attention mechanism. These modules ensure both computational efficiency and effective integration of depth priors. Moreover, the depth priors empower the network to dynamically adapt to varying haze densities, illumination conditions, and domain gaps across synthetic and real‑world data. Extensive experimental results demonstrate the effectiveness of our UDPNet, outperforming the state‑of‑the‑art methods on popular dehazing datasets, with PSNR improvements of 0.85 dB on SOTS‑indoor, 1.19 dB on Haze4K, and 1.79 dB on NHR. Our proposed solution establishes a new benchmark for depth‑aware dehazing across various scenarios. Pretrained models and codes are released at our project https://github.com/Harbinzzy/UDPNet.
Authors:Changli Wu, Haodong Wang, Jiayi Ji, Yutian Yao, Chunsai Du, Jihua Kang, Yanwei Fu, Liujuan Cao
Abstract:
Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high‑quality point clouds, while real‑world agents such as robots and mobile phones operate with only a few sparse RGB views and strict latency constraints. We introduce Multi‑view 3D Referring Expression Segmentation (MV‑3DRES), where the model must recover scene structure and segment the referred object directly from sparse multi‑view images. Traditional two‑stage pipelines, which first reconstruct a point cloud and then perform segmentation, often yield low‑quality geometry, produce coarse or degraded target regions, and run slowly. We propose the Multimodal Visual Geometry Grounded Transformer (MVGGT), an efficient end‑to‑end framework that integrates language information into sparse‑view geometric reasoning through a dual‑branch design. Training in this setting exposes a critical optimization barrier, termed Foreground Gradient Dilution (FGD), where sparse 3D signals lead to weak supervision. To resolve this, we introduce Per‑view No‑target Suppression Optimization (PVSO), which provides stronger and more balanced gradients across views, enabling stable and efficient learning. To support consistent evaluation, we build MVRefer, a benchmark that defines standardized settings and metrics for MV‑3DRES. Experiments show that MVGGT establishes the first strong baseline and achieves both high accuracy and fast inference, outperforming existing alternatives. The code is available at https://mvggt.github.io/.
Authors:Junyan Lin, Junlong Tong, Hao Wu, Jialiang Zhang, Jinming Liu, Xin Jin, Xiaoyu Shen
Abstract:
Multimodal Large Language Models (MLLMs) have achieved strong performance across many tasks, yet most systems remain limited to offline inference, requiring complete inputs before generating outputs. Recent streaming methods reduce latency by interleaving perception and generation, but still enforce a sequential perception‑generation cycle, limiting real‑time interaction. In this work, we target a fundamental bottleneck that arises when extending MLLMs to real‑time video understanding: the global positional continuity constraint imposed by standard positional encoding schemes. While natural in offline inference, this constraint tightly couples perception and generation, preventing effective input‑output parallelism. To address this limitation, we propose a parallel streaming framework that relaxes positional continuity through three designs: Overlapped, Group‑Decoupled, and Gap‑Isolated. These designs enable simultaneous perception and generation, allowing the model to process incoming inputs while producing responses in real time. Extensive experiments reveal that Group‑Decoupled achieves the best efficiency‑performance balance, maintaining high fluency and accuracy while significantly reducing latency. We further show that the proposed framework yields up to 2x acceleration under balanced perception‑generation workloads, establishing a principled pathway toward speak‑while‑watching real‑time systems. We make all our code publicly available: https://github.com/EIT‑NLP/Speak‑While‑Watching.
Authors:Zhongping Ji
Abstract:
Modern computer vision architectures, from CNNs to Transformers, predominantly rely on the stacking of heuristic modules: spatial mixers (Attention/Conv) followed by channel mixers (FFNs). In this work, we challenge this paradigm by returning to mathematical first principles. We propose the Clifford Algebra Network (CAN), also referred to as CliffordNet, a vision backbone grounded purely in Geometric Algebra. Instead of engineering separate modules for mixing and memory, we derive a unified interaction mechanism based on the Clifford Geometric Product (uv = u \cdot v + u \wedge v). This operation ensures algebraic completeness regarding the Geometric Product by simultaneously capturing feature coherence (via the generalized inner product) and structural variation (via the exterior wedge product).
Implemented via an efficient sparse rolling mechanism with strict linear complexity O(N), our model reveals a surprising emergent property: the geometric interaction is so representationally dense that standard Feed‑Forward Networks (FFNs) become redundant. Empirically, CliffordNet establishes a new Pareto frontier: our Nano variant achieves 77.82% accuracy on CIFAR‑100 with only 1.4M parameters, effectively matching the heavy‑weight ResNet‑18 (11.2M) with 8× fewer parameters, while our Lite variant (2.6M) sets a new SOTA for tiny models at 79.05%. Our results suggest that global understanding can emerge solely from rigorous, algebraically complete local interactions, potentially signaling a shift where geometry is all you need. Code is available at https://github.com/ParaMind2025/CAN.
Authors:Krishna Vinod, Joseph Raj Vishal, Kaustav Chanda, Prithvi Jai Ramesh, Yezhou Yang, Bharatesh Chakravarthi
Abstract:
Tracking skiers in RGB broadcast footage is challenging due to motion blur, static overlays, and clutter that obscure the fast‑moving athlete. Event cameras, with their asynchronous contrast sensing, offer natural robustness to such artifacts, yet a controlled benchmark for winter‑sport tracking has been missing. We introduce event SkiTB (eSkiTB), a synthetic event‑based ski tracking dataset generated from SkiTB using direct video‑to‑event conversion without neural interpolation, enabling an iso‑informational comparison between RGB and event modalities. Benchmarking SDTrack (spiking transformer) against STARK (RGB transformer), we find that event‑based tracking is substantially resilient to broadcast clutter in scenes dominated by static overlays, achieving 0.685 IoU, outperforming RGB by +20.0 points. Across the dataset, SDTrack attains a mean IoU of 0.711, demonstrating that temporal contrast is a reliable cue for tracking ballistic motion in visually congested environments. eSkiTB establishes the first controlled setting for event‑based tracking in winter sports and highlights the promise of event cameras for ski tracking. The dataset and code will be released at https://github.com/eventbasedvision/eSkiTB.
Authors:Yuanting Gao, Shuo Cao, Xiaohui Li, Yuandong Pu, Yihao Liu, Kai Zhang
Abstract:
Image deblurring has advanced rapidly with deep learning, yet most methods exhibit poor generalization beyond their training datasets, with performance dropping significantly in real‑world scenarios. Our analysis shows this limitation stems from two factors: datasets face an inherent trade‑off between realism and coverage of diverse blur patterns, and algorithmic designs remain restrictive, as pixel‑wise losses drive models toward local detail recovery while overlooking structural and semantic consistency, whereas diffusion‑based approaches, though perceptually strong, still fail to generalize when trained on narrow datasets with simplistic strategies. Through systematic investigation, we identify blur pattern diversity as the decisive factor for robust generalization and propose Blur Pattern Pretraining (BPP), which acquires blur priors from simulation datasets and transfers them through joint fine‑tuning on real data. We further introduce Motion and Semantic Guidance (MoSeG) to strengthen blur priors under severe degradation, and integrate it into GLOWDeblur, a Generalizable reaL‑wOrld lightWeight Deblur model that combines convolution‑based pre‑reconstruction & domain alignment module with a lightweight diffusion backbone. Extensive experiments on six widely‑used benchmarks and two real‑world datasets validate our approach, confirming the importance of blur priors for robust generalization and demonstrating that the lightweight design of GLOWDeblur ensures practicality in real‑world applications. The project page is available at https://vegdog007.github.io/GLOWDeblur_Website/.
Authors:Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, Hans Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, Haoyu Lu, Y. charles, Yiping Bao, Yuantao Fan, Guopeng Li, Haiyang Shen, Xuanzhong Chen, Wendong Xu, Shuzheng Si, Zefan Cai, Wenhao Chai, Ziqi Huang, Fangfu Liu, Tianyu Liu, Baobao Chang, Xiaobo Hu, Kaiyuan Chen, Yixin Ren, Yang Liu, Yuan Gong, Kuan Li
Abstract:
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state‑of‑the‑art MLLMs consistently fail on basic visual tasks that humans, even 3‑year‑olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3‑Pro‑Preview scores 49.7, lagging behind 6‑year‑old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge‑heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human‑level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision‑Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat‑AI/BabyVision for reproduction.
Authors:Hao Tang, Ting Huang, Zeyu Zhang
Abstract:
Spatial intelligence refers to the ability to perceive, reason about, and describe objects and their relationships within three‑dimensional environments, forming a foundation for embodied perception and scene understanding. 3D captioning aims to describe 3D scenes in natural language; however, it remains challenging due to the sparsity and irregularity of point clouds and, more critically, the weak grounding and limited out‑of‑distribution (OOD) generalization of existing captioners across drastically different environments, including indoor and outdoor 3D scenes. To address this challenge, we propose 3D CoCa v2, a generalizable 3D captioning framework that unifies contrastive vision‑language learning with 3D caption generation and further improves robustness via test‑time search (TTS) without updating the captioner parameters. 3D CoCa v2 builds on a frozen CLIP‑based semantic prior, a spatially‑aware 3D scene encoder for geometry, and a multimodal decoder jointly optimized with contrastive and captioning objectives, avoiding external detectors or handcrafted proposals. At inference, TTS produces diverse caption candidates and performs reward‑guided selection using a compact scene summary. Experiments show improvements over 3D CoCa of +1.50 CIDEr@0.5IoU on ScanRefer and +1.61 CIDEr@0.5IoU on Nr3D, and +3.8 CIDEr@0.25 in zero‑shot OOD evaluation on TOD3Cap. Code will be released at https://github.com/AIGeeksGroup/3DCoCav2.
Authors:Weihao Hong, Zhiyuan Jiang, Bingyu Shen, Xinlei Guan, Yangyi Feng, Meng Xu, Boyang Li
Abstract:
Vision‑Language Models (VLMs) are increasingly used in safety‑critical applications that require reliable visual grounding. However, these models often hallucinate details that are not present in the image to satisfy user prompts. While recent datasets and benchmarks have been introduced to evaluate systematic hallucinations in VLMs, many hallucination behaviors remain insufficiently characterized. In particular, prior work primarily focuses on object presence or absence, leaving it unclear how prompt phrasing and structural constraints can systematically induce hallucinations. In this paper, we investigate how different forms of prompt pressure influence hallucination behavior. We introduce Ghost‑100, a procedurally generated dataset of synthetic scenes in which key visual details are deliberately removed, enabling controlled analysis of absence‑based hallucinations. Using a structured 5‑Level Prompt Intensity Framework, we vary prompts from neutral queries to toxic demands and rigid formatting constraints. We evaluate three representative open‑weight VLMs: MiniCPM‑V 2.6‑8B, Qwen2‑VL‑7B, and Qwen3‑VL‑8B. Across all three models, hallucination rates do not increase monotonically with prompt intensity. All models exhibit reductions at higher intensity levels at different thresholds, though not all show sustained reduction under maximum coercion. These results suggest that current safety alignment is more effective at detecting semantic hostility than structural coercion, revealing model‑specific limitations in handling compliance pressure. Our dataset is available at: https://github.com/bli1/tone‑matters
Authors:Xianghong Zou, Jianping Li, Yandi Yang, Weitong Wu, Yuan Wang, Qiegen Liu, Zhen Dong
Abstract:
Point Cloud‑based Place Recognition (PCPR) demonstrates considerable potential in applications such as autonomous driving, robot localization and navigation, and map update. In practical applications, point clouds used for place recognition are often acquired from different platforms and LiDARs across varying scene. However, existing PCPR datasets lack diversity in scenes, platforms, and sensors, which limits the effective development of related research. To address this gap, we establish WHU‑PCPR, a cross‑platform heterogeneous point cloud dataset designed for place recognition. The dataset differentiates itself from existing datasets through its distinctive characteristics: 1) cross‑platform heterogeneous point clouds: collected from survey‑grade vehicle‑mounted Mobile Laser Scanning (MLS) systems and low‑cost Portable helmet‑mounted Laser Scanning (PLS) systems, each equipped with distinct mechanical and solid‑state LiDAR sensors. 2) Complex localization scenes: encompassing real‑time and long‑term changes in both urban and campus road scenes. 3) Large‑scale spatial coverage: featuring 82.3 km of trajectory over a 60‑month period and an unrepeated route of approximately 30 km. Based on WHU‑PCPR, we conduct extensive evaluation and in‑depth analysis of several representative PCPR methods, and provide a concise discussion of key challenges and future research directions. The dataset and benchmark code are available at https://github.com/zouxianghong/WHU‑PCPR.
Authors:Yueming Pan, Ruoyu Feng, Jianmin Bao, Chong Luo, Nanning Zheng
Abstract:
Video outpainting extends a video beyond its original boundaries by synthesizing missing border content. Compared with image outpainting, it requires not only per‑frame spatial plausibility but also long‑range temporal coherence, especially when outpainted content becomes visible across time under camera or object motion. We propose GlobalPaint, a diffusion‑based framework for spatiotemporal coherent video outpainting. Our approach adopts a hierarchical pipeline that first outpaints key frames and then completes intermediate frames via an interpolation model conditioned on the completed boundaries, reducing error accumulation in sequential processing. At the model level, we augment a pretrained image inpainting backbone with (i) an Enhanced Spatial‑Temporal module featuring 3D windowed attention for stronger spatiotemporal interaction, and (ii) global feature guidance that distills OpenCLIP features from observed regions across all frames into compact global tokens using a dedicated extractor. Comprehensive evaluations on benchmark datasets demonstrate improved reconstruction quality and more natural motion compared to prior methods. Our demo page is https://yuemingpan.github.io/GlobalPaint/
Authors:Ahmed Abdelkawy, Ahmed Elsayed, Asem Ali, Aly Farag, Thomas Tretter, Michael McIntyre
Abstract:
Understanding student behavior in the classroom is essential to improve both pedagogical quality and student engagement. Existing methods for predicting student engagement typically require substantial annotated data to model the diversity of student behaviors, yet privacy concerns often restrict researchers to their own proprietary datasets. Moreover, the classroom context, represented in peers' actions, is ignored. To address the aforementioned limitation, we propose a novel three‑stage framework for video‑based student engagement measurement. First, we explore the few‑shot adaptation of the vision‑language model for student action recognition, which is fine‑tuned to distinguish among action categories with a few training samples. Second, to handle continuous and unpredictable student actions, we utilize the sliding temporal window technique to divide each student's 2‑minute‑long video into non‑overlapping segments. Each segment is assigned an action category via the fine‑tuned VLM model, generating a sequence of action predictions. Finally, we leverage the large language model to classify this entire sequence of actions, together with the classroom context, as belonging to an engaged or disengaged student. The experimental results demonstrate the effectiveness of the proposed approach in identifying student engagement. The source code will be available at https://github.com/ahmed‑nady/context_aware_student_engagement.
Authors:Saksham Singh Kushwaha, Sayan Nag, Yapeng Tian, Kuldeep Kulkarni
Abstract:
In this paper, we introduce Object‑WIPER, a training‑free framework for removing dynamic objects and their associated visual effects from videos, and inpainting them with semantically consistent and temporally coherent content. Our approach leverages a pre‑trained text‑to‑video diffusion transformer (DiT). Given an input video, a user‑provided object mask, and query tokens describing the target object and its effects, we localize relevant visual tokens via visual‑text cross‑attention and visual self‑attention. This produces an intermediate effect mask that we fuse with the user mask to obtain a final foreground token mask to replace. We first invert the video through the DiT to obtain structured noise, then reinitialize the masked tokens with Gaussian noise while preserving background tokens. During denoising, we copy values for the background tokens saved during inversion to maintain scene fidelity. To address the lack of suitable evaluation, we introduce a new object removal metric that rewards temporal consistency among foreground tokens across consecutive frames, coherence between foreground and background tokens within each frame, and dissimilarity between the input and output foreground tokens. Experiments on DAVIS and a newly curated real‑world associated effect benchmark (WIPER‑Bench) show that Object‑WIPER surpasses both training‑based and training‑free baselines in terms of the metric, achieving clean removal and temporally stable reconstruction without any retraining. Our new benchmark, source code, and pre‑trained models will be publicly available.
Authors:Chen Gong, Kecen Li, Zinan Lin, Tianhao Wang
Abstract:
To improve the quality of Differentially private (DP) synthetic images, most studies have focused on improving the core optimization techniques (e.g., DP‑SGD). Recently, we have witnessed a paradigm shift that takes these techniques off the shelf and studies how to use them together to achieve the best results. One notable work is DP‑FETA, which proposes using `central images' for `warming up' the DP training and then using traditional DP‑SGD.
Inspired by DP‑FETA, we are curious whether there are other such tools we can use together with DP‑SGD. We first observe that using `central images' mainly works for datasets where there are many samples that look similar. To handle scenarios where images could vary significantly, we propose FETA‑Pro, which introduces frequency features as `training shortcuts.' The complexity of frequency features lies between that of spatial features (captured by `central images') and full images, allowing for a finer‑grained curriculum for DP training. To incorporate these two types of shortcuts together, one challenge is to handle the training discrepancy between spatial and frequency features. To address it, we leverage the pipeline generation property of generative models (instead of having one model trained with multiple features/objectives, we can have multiple models working on different features, then feed the generated results from one model into another) and use a more flexible design. Specifically, FETA‑Pro introduces an auxiliary generator to produce images aligned with noisy frequency features. Then, another model is trained with these images, together with spatial features and DP‑SGD. Evaluated across five sensitive image datasets, FETA‑Pro shows an average of 25.7% higher fidelity and 4.1% greater utility than the best‑performing baseline, under a privacy budget ε= 1.
Authors:Stevenson Pather, Niels Martignène, Arnaud Bugnet, Fouad Boutaleb, Fabien D'Hondt, Deise Santana Maia
Abstract:
We introduce EyeTheia, a lightweight and open deep learning pipeline for webcam‑based gaze estimation, designed for browser‑based experimental platforms and real‑world cognitive and clinical research. EyeTheia enables real‑time gaze tracking using only a standard laptop webcam, combining MediaPipe‑based landmark extraction with a convolutional neural network inspired by iTracker and optional user‑specific fine‑tuning. We investigate two complementary strategies: adapting a model pretrained on mobile data and training the same architecture from scratch on a desktop‑oriented dataset. Validation results on MPIIFaceGaze show comparable performance between both approaches prior to calibration, while lightweight user‑specific fine‑tuning consistently reduces gaze prediction error. We further evaluate EyeTheia in a realistic Dot‑Probe task and compare it to the commercial webcam‑based tracker SeeSo SDK. Results indicate strong agreement in left‑right gaze allocation during stimulus presentation, despite higher temporal variability. Overall, EyeTheia provides a transparent and extensible solution for low‑cost gaze tracking, suitable for scalable and reproducible experimental and clinical studies. The code, trained models, and experimental materials are publicly available.
Authors:Kuan Wei Chen, Ting Yi Lin, Wen Ren Yang, Aryan Kesarwani, Riya Singh
Abstract:
We present a cost‑effective two‑step authentication system that integrates face identification and speaker verification using only a camera and microphone available on common devices. The pipeline first performs face recognition to identify a candidate user from a small enrolled group, then performs voice recognition only against the matched identity to reduce computation and improve robustness. For face recognition, a pruned VGG‑16 based classifier is trained on an augmented dataset of 924 images from five subjects, with faces localized by MTCNN; it achieves 95.1% accuracy. For voice recognition, a CNN speaker‑verification model trained on LibriSpeech (train‑other‑360) attains 98.9% accuracy and 3.456% EER on test‑clean. Source code and trained models are available at https://github.com/NCUE‑EE‑AIAL/Two‑step‑Authentication‑Multi‑biometric‑System.
Authors:Shiwen Zhang, Haibin Huang, Chi Zhang, Xuelong Li
Abstract:
Content‑Preserving Style transfer, given content and style references, remains challenging for Diffusion Transformers (DiTs) due to its internal entangled content and style features. In this technical report, we propose the first content‑preserving style transfer model trained on Qwen‑Image‑Edit, which activates Qwen‑Image‑Edit's strong content preservation and style customization capability. We collected and filtered high quality data of limited specific styles and synthesized triplets with thousands categories of style images in‑the‑wild. We introduce the Curriculum Continual Learning framework to train QwenStyle with such mixture of clean and noisy triplets, which enables QwenStyle to generalize to unseen styles without degradation of the precise content preservation capability. Our QwenStyle V1 achieves state‑of‑the‑art performance in three core metrics: style similarity, content consistency, and aesthetic quality.
Authors:Shubham Goel, Farzana S, C V Rishi, Aditya Arun, C V Jawahar
Abstract:
Biryani, one of India's most celebrated dishes, exhibits remarkable regional diversity in its preparation, ingredients, and presentation. With the growing availability of online cooking videos, there is unprecedented potential to study such culinary variations using computational tools systematically. However, existing video understanding methods fail to capture the fine‑grained, multimodal, and culturally grounded differences in procedural cooking videos. This work presents the first large‑scale, curated dataset of biryani preparation videos, comprising 120 high‑quality YouTube recordings across 12 distinct regional styles. We propose a multi‑stage framework leveraging recent advances in vision‑language models (VLMs) to segment videos into fine‑grained procedural units and align them with audio transcripts and canonical recipe text. Building on these aligned representations, we introduce a video comparison pipeline that automatically identifies and explains procedural differences between regional variants. We construct a comprehensive question‑answer (QA) benchmark spanning multiple reasoning levels to evaluate procedural understanding in VLMs. Our approach employs multiple VLMs in complementary roles, incorporates human‑in‑the‑loop verification for high‑precision tasks, and benchmarks several state‑of‑the‑art models under zero‑shot and fine‑tuned settings. The resulting dataset, comparison methodology, and QA benchmark provide a new testbed for evaluating VLMs on structured, multimodal reasoning tasks and open new directions for computational analysis of cultural heritage through cooking videos. We release all data, code, and the project website at https://farzanashaju.github.io/how‑does‑india‑cook‑biryani/.
Authors:Kaiyuan Deng, Bo Hui, Gen Li, Jie Ji, Minghai Qin, Geng Yuan, Xiaolong Ma
Abstract:
The widespread adoption of text‑to‑image (T2I) diffusion models has raised concerns about their potential to generate copyrighted, inappropriate, or sensitive imagery. As a practical solution, machine unlearning aims to erase unwanted concepts without retraining from scratch. While most existing methods are effective for single‑concept unlearning, they often struggle when removing multiple concepts, causing significant challenges in unlearning effectiveness, generation quality, and sensitivity to hyperparameters and datasets. We take a unique perspective on multi‑concept unlearning by leveraging model sparsity and propose the Forget It All (FIA) framework. FIA first introduces Contrastive Concept Saliency to quantify each weight connection's contribution to a target concept. It then identifies Concept Sensitive Neurons by combining temporal and spatial information, ensuring that only neurons consistently responsive to the target concept are selected. Finally, FIA constructs masks from the identified neurons and fuses them into a unified multi‑concept mask, where Concept Agnostic Neurons that broadly support general content generation are preserved while concept‑specific neurons are pruned to remove the targets. FIA is training‑free and requires minimal hyperparameter tuning for new tasks, enabling plug‑and‑play use. Extensive experiments across three distinct unlearning tasks demonstrate that FIA achieves more reliable multi‑concept unlearning, improving forgetting effectiveness while maintaining generation fidelity and quality. Code is available at https://github.com/kaiyuan02415/Forget‑It‑All
Authors:Kaiyuan Deng, Gen Li, Yang Xiao, Bo Hui, Xiaolong Ma
Abstract:
Text‑to‑image diffusion models have achieved remarkable progress, yet their use raises copyright and misuse concerns, prompting research into machine unlearning. However, extending multi‑concept unlearning to large‑scale scenarios remains difficult due to three challenges: (i) conflicting weight updates that hinder unlearning or degrade generation; (ii) imprecise mechanisms that cause collateral damage to similar content; and (iii) reliance on additional data or modules, creating scalability bottlenecks. To address these, we propose Scalable‑Precise Concept Unlearning (ScaPre), a unified framework tailored for large‑scale unlearning. ScaPre introduces a conflict‑aware stable design, integrating spectral trace regularization and geometry alignment to stabilize optimization, suppress conflicts, and preserve global structure. Furthermore, an Informax Decoupler identifies concept‑relevant parameters and adaptively reweights updates, strictly confining unlearning to the target subspace. ScaPre yields an efficient closed‑form solution without requiring auxiliary data or sub‑models. Comprehensive experiments on objects, styles, and explicit content demonstrate that ScaPre effectively removes target concepts while maintaining generation quality. It forgets up to × \mathbf5 more concepts than the best baseline within acceptable quality limits, achieving state‑of‑the‑art precision and efficiency for large‑scale unlearning. Code is available at https://github.com/kaiyuan02415/scapre
Authors:Chimdi Walter Ndubuisi, Toni Kazic
Abstract:
Leaf‑lesion segmentation is topology‑sensitive: small merges, splits, or false holes can be biologically meaningful descriptors of biochemical pathways, yet they are weakly penalized by standard pixel‑wise losses in Euclidean latents. I explore HyperTopo‑Adapters, a lightweight, parameter‑efficient head trained on top of a frozen vision encoder, which embeds features on a product manifold ‑‑ hyperbolic + Euclidean + spherical (H + E + S) ‑‑ to encourage hierarchical separation (H), local linear detail (E), and global closure (S). A topology prior complements Dice/BCE in two forms: (i) persistent‑homology (PH) distance for evaluation and selection, and (ii) a differentiable surrogate that combines a soft Euler‑characteristic match with total variation regularization for stable training. I introduce warm‑ups for both the hyperbolic contrastive term and the topology prior, per‑sample evaluation of structure‑aware metrics (Boundary‑F1, Betti errors, PD distance), and a min‑PD within top‑K Dice rule for checkpoint selection. On a Kaggle leaf‑lesion dataset (N=2,940), early results show consistent gains in boundary and topology metrics (reducing Delta beta_1 hole error by 9%) while Dice/IoU remain competitive. The study is diagnostic by design: I report controlled ablations (curvature learning, latent dimensions, contrastive temperature, surrogate settings), and ongoing tests varying encoder strength (ResNet‑50, DeepLabV3, DINOv2/v3), input resolution, PH weight, and partial unfreezing of late blocks. The contribution is an open, reproducible train/eval suite (available at https://github.com/ChimdiWalter/HyperTopo‑Adapters) that isolates geometric/topological priors and surfaces failure modes to guide stronger, topology‑preserving architectures.
Authors:Yinsong Wang, Xinzhe Luo, Siyi Du, Chen Qin
Abstract:
Deformable multi‑contrast image registration is a challenging yet crucial task due to the complex, non‑linear intensity relationships across different imaging contrasts. Conventional registration methods typically rely on iterative optimization of the deformation field, which is time‑consuming. Although recent learning‑based approaches enable fast and accurate registration during inference, their generalizability remains limited to the specific contrasts observed during training. In this work, we propose an adaptive conditional contrast‑agnostic deformable image registration framework (AC‑CAR) based on a random convolution‑based contrast augmentation scheme. AC‑CAR can generalize to arbitrary imaging contrasts without observing them during training. To encourage contrast‑invariant feature learning, we propose an adaptive conditional feature modulator (ACFM) that adaptively modulates the features and the contrast‑invariant latent regularization to enforce the consistency of the learned feature across different imaging contrasts. Additionally, we enable our framework to provide contrast‑agnostic registration uncertainty by integrating a variance network that leverages the contrast‑agnostic registration encoder to improve the trustworthiness and reliability of AC‑CAR. Experimental results demonstrate that AC‑CAR outperforms baseline methods in registration accuracy and exhibits superior generalization to unseen imaging contrasts. Code is available at https://github.com/Yinsong0510/AC‑CAR.
Authors:Longbin Ji, Xiaoxiong Liu, Junyuan Shang, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang
Abstract:
Recent advances in video generation have been dominated by diffusion and flow‑matching models, which produce high‑quality results but remain computationally intensive and difficult to scale. In this work, we introduce VideoAR, the first large‑scale Visual Autoregressive (VAR) framework for video generation that combines multi‑scale next‑frame prediction with autoregressive modeling. VideoAR disentangles spatial and temporal dependencies by integrating intra‑frame VAR modeling with causal next‑frame prediction, supported by a 3D multi‑scale tokenizer that efficiently encodes spatio‑temporal dynamics. To improve long‑term consistency, we propose Multi‑scale Temporal RoPE, Cross‑Frame Error Correction, and Random Frame Mask, which collectively mitigate error propagation and stabilize temporal coherence. Our multi‑stage pretraining pipeline progressively aligns spatial and temporal learning across increasing resolutions and durations. Empirically, VideoAR achieves new state‑of‑the‑art results among autoregressive models, improving FVD on UCF‑101 from 99.5 to 88.6 while reducing inference steps by over 10x, and reaching a VBench score of 81.74‑competitive with diffusion‑based models an order of magnitude larger. These results demonstrate that VideoAR narrows the performance gap between autoregressive and diffusion paradigms, offering a scalable, efficient, and temporally consistent foundation for future video generation research.
Authors:Chanchan Wang, Yuanfang Wang, Qing Xu, Guanxin Chen
Abstract:
Domain‑generalized retinal vessel segmentation is critical for automated ophthalmic diagnosis, yet faces significant challenges from domain shift induced by non‑uniform illumination and varying contrast, compounded by the difficulty of preserving fine vessel structures. While the Segment Anything Model (SAM) exhibits remarkable zero‑shot capabilities, existing SAM‑based methods rely on simple adapter fine‑tuning while overlooking frequency‑domain information that encodes domain‑invariant features, resulting in degraded generalization under illumination and contrast variations. Furthermore, SAM's direct upsampling inevitably loses fine vessel details. To address these limitations, we propose WaveRNet, a wavelet‑guided frequency learning framework for robust multi‑source domain‑generalized retinal vessel segmentation. Specifically, we devise a Spectral‑guided Domain Modulator (SDM) that integrates wavelet decomposition with learnable domain tokens, enabling the separation of illumination‑robust low‑frequency structures from high‑frequency vessel boundaries while facilitating domain‑specific feature generation. Furthermore, we introduce a Frequency‑Adaptive Domain Fusion (FADF) module that performs intelligent test‑time domain selection through wavelet‑based frequency similarity and soft‑weighted fusion. Finally, we present a Hierarchical Mask‑Prompt Refiner (HMPR) that overcomes SAM's upsampling limitation through coarse‑to‑fine refinement with long‑range dependency modeling. Extensive experiments under the Leave‑One‑Domain‑Out protocol on four public retinal datasets demonstrate that WaveRNet achieves state‑of‑the‑art generalization performance. The source code is available at https://github.com/Chanchan‑Wang/WaveRNet.
Authors:Kaiwen Huang, Yizhe Zhang, Yi Zhou, Tianyang Xu, Tao Zhou
Abstract:
Semi‑supervised medical image segmentation is an effective method for addressing scenarios with limited labeled data. Existing methods mainly rely on frameworks such as mean teacher and dual‑stream consistency learning. These approaches often face issues like error accumulation and model structural complexity, while also neglecting the interaction between labeled and unlabeled data streams. To overcome these challenges, we propose a Bidirectional Channel‑selective Semantic Interaction~(BCSI) framework for semi‑supervised medical image segmentation. First, we propose a Semantic‑Spatial Perturbation~(SSP) mechanism, which disturbs the data using two strong augmentation operations and leverages unsupervised learning with pseudo‑labels from weak augmentations. Additionally, we employ consistency on the predictions from the two strong augmentations to further improve model stability and robustness. Second, to reduce noise during the interaction between labeled and unlabeled data, we propose a Channel‑selective Router~(CR) component, which dynamically selects the most relevant channels for information exchange. This mechanism ensures that only highly relevant features are activated, minimizing unnecessary interference. Finally, the Bidirectional Channel‑wise Interaction~(BCI) strategy is employed to supplement additional semantic information and enhance the representation of important channels. Experimental results on multiple benchmarking 3D medical datasets demonstrate that the proposed method outperforms existing semi‑supervised approaches.
Authors:Yinghan Xu, John Dingliana
Abstract:
We propose a novel framework for decomposing arbitrarily posed humans into animatable multi‑layered 3D human avatars, separating the body and garments. Conventional single‑layer reconstruction methods lock clothing to one identity, while prior multi‑layer approaches struggle with occluded regions. We overcome both limitations by encoding each layer as a set of 2D Gaussians for accurate geometry and photorealistic rendering, and inpainting hidden regions with a pretrained 2D diffusion model via score‑distillation sampling (SDS). Our three‑stage training strategy first reconstructs the coarse canonical garment via single‑layer reconstruction, followed by multi‑layer training to jointly recover the inner‑layer body and outer‑layer garment details. Experiments on two 3D human benchmark datasets (4D‑Dress, Thuman2.0) show that our approach achieves better rendering quality and layer decomposition and recomposition than the previous state‑of‑the‑art, enabling realistic virtual try‑on under novel viewpoints and poses, and advancing practical creation of high‑fidelity 3D human assets for immersive applications. Our code is available at https://github.com/RockyXu66/LayerGS
Authors:ChunTeng Chen, YiChen Hsu, YiWen Liu, WeiFang Sun, TsaiChing Ni, ChunYi Lee, Min Sun, YuanFu Yang
Abstract:
The ability to automatically generate large‑scale, interactive, and physically realistic 3D environments is crucial for advancing robotic learning and embodied intelligence. However, existing generative approaches often fail to capture the functional complexity of real‑world interiors, particularly those containing articulated objects with movable parts essential for manipulation and navigation. This paper presents SceneFoundry, a language‑guided diffusion framework that generates apartment‑scale 3D worlds with functionally articulated furniture and semantically diverse layouts for robotic training. From natural language prompts, an LLM module controls floor layout generation, while diffusion‑based posterior sampling efficiently populates the scene with articulated assets from large‑scale 3D repositories. To ensure physical usability, SceneFoundry employs differentiable guidance functions to regulate object quantity, prevent articulation collisions, and maintain sufficient walkable space for robotic navigation. Extensive experiments demonstrate that our framework generates structurally valid, semantically coherent, and functionally interactive environments across diverse scene types and conditions, enabling scalable embodied AI research. project page: https://anc891203.github.io/SceneFoundry‑Demo/
Authors:Hassaan Farooq, Marvin Brenner, Peter Stütz
Abstract:
Unmanned Aerial Vehicles (UAVs) are increasingly deployed in close proximity to humans for applications such as parcel delivery, traffic monitoring, disaster response and infrastructure inspections. Ensuring safe and reliable operation in these human‑populated environments demands accurate perception of human poses and actions from an aerial viewpoint. This perspective challenges existing methods with low resolution, steep viewing angles and (self‑)occlusion, especially if the application demands realtime feasibile models. We train and deploy FlyPose, a lightweight top‑down human pose estimation pipeline for aerial imagery. Through multi‑dataset training, we achieve an average improvement of 6.8 mAP in person detection across the test‑sets of Manipal‑UAV, VisDrone, HIT‑UAV as well as our custom dataset. For 2D human pose estimation we report an improvement of 16.3 mAP on the challenging UAV‑Human dataset. FlyPose runs with an inference latency of ~20 milliseconds including preprocessing on a Jetson Orin AGX Developer Kit and is deployed onboard a quadrotor UAV during flight experiments. We also publish FlyPose‑104, a small but challenging aerial human pose estimation dataset, that includes manual annotations from difficult aerial perspectives: https://github.com/farooqhassaan/FlyPose.
Authors:Zehan Wang, Ziang Zhang, Jiayang Xu, Jialei Wang, Tianyu Pang, Chao Du, HengShuang Zhao, Zhou Zhao
Abstract:
This work presents Orient Anything V2, an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. Building upon Orient Anything V1, which defines orientation via a single unique front face, V2 extends this capability to handle objects with diverse rotational symmetries and directly estimate relative rotations. These improvements are enabled by four key innovations: 1) Scalable 3D assets synthesized by generative models, ensuring broad category coverage and balanced data distribution; 2) An efficient, model‑in‑the‑loop annotation system that robustly identifies 0 to N valid front faces for each object; 3) A symmetry‑aware, periodic distribution fitting objective that captures all plausible front‑facing orientations, effectively modeling object rotational symmetry; 4) A multi‑frame architecture that directly predicts relative object rotations. Extensive experiments show that Orient Anything V2 achieves state‑of‑the‑art zero‑shot performance on orientation estimation, 6DoF pose estimation, and object symmetry recognition across 11 widely used benchmarks. The model demonstrates strong generalization, significantly broadening the applicability of orientation estimation in diverse downstream tasks.
Authors:Pengcheng Xu, Peng Tang, Donghao Luo, Xiaobin Hu, Weichu Cui, Qingdong He, Zhennan Chen, Jiangning Zhang, Charles Ling, Boyu Wang
Abstract:
Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when referencing details across multiple input images. In this work, we propose a scalable multi‑image editing framework for UMMs that explicitly distinguishes image identities and generalizes to variable input counts. Algorithmically, we introduce two innovations: 1) The learnable latent separators explicitly differentiate each reference image in the latent space, enabling accurate and disentangled conditioning. 2) The sinusoidal index encoding assigns visual tokens from the same image a continuous sinusoidal index embedding, which provides explicit image identity while allowing generalization and extrapolation on a variable number of inputs. To facilitate training and evaluation, we establish a high‑fidelity benchmark using an inverse dataset construction methodology to guarantee artifact‑free, achievable outputs. Experiments show clear improvements in semantic consistency, visual fidelity, and cross‑image integration over prior baselines on diverse multi‑image editing tasks, validating our advantages on consistency and generalization ability.
Authors:Bin-Bin Gao, Chengjie Wang
Abstract:
Universal visual anomaly detection (AD) aims to identify anomaly images and segment anomaly regions towards open and dynamic scenarios, following zero‑ and few‑shot paradigms without any dataset‑specific fine‑tuning. We have witnessed significant progress in widely use of visual‑language foundational models in recent approaches. However, current methods often struggle with complex prompt engineering, elaborate adaptation modules, and challenging training strategies, ultimately limiting their flexibility and generality. To address these issues, this paper rethinks the fundamental mechanism behind visual‑language models for AD and presents an embarrassingly simple, general, and effective framework for Universal vision Anomaly Detection (UniADet). Specifically, we first find language encoder is used to derive decision weights for anomaly classification and segmentation, and then demonstrate that it is unnecessary for universal AD. Second, we propose an embarrassingly simple method to completely decouple classification and segmentation, and decouple cross‑level features, i.e., learning independent weights for different tasks and hierarchical features. UniADet is highly simple (learning only decoupled weights), parameter‑efficient (only 0.002M learnable parameters), general (adapting a variety of foundation models), and effective (surpassing state‑of‑the‑art zero‑/few‑shot by a large margin and even full‑shot AD methods for the first time) on 14 real‑world AD benchmarks covering both industrial and medical domains. We will make the code and model of UniADet available at https://github.com/gaobb/UniADet.
Authors:Yanfeng Li, Yue Sun, Keren Fu, Sio-Kei Im, Xiaoming Liu, Guangtao Zhai, Xiaohong Liu, Tao Tan
Abstract:
Existing multi‑object image generation methods face difficulties in achieving precise alignment between localized image generation regions and their corresponding semantics based on language descriptions, frequently resulting in inconsistent object quantities and attribute aliasing. To mitigate this limitation, mainstream approaches typically rely on external control signals to explicitly constrain the spatial layout, local semantic and visual attributes of images. However, this strong dependency makes the input format rigid, rendering it incompatible with the heterogeneous resource conditions of users and diverse constraint requirements. To address these challenges, we propose MoGen, a user‑friendly multi‑object image generation method. First, we design a Regional Semantic Anchor (RSA) module that precisely anchors phrase units in language descriptions to their corresponding image regions during the generation process, enabling text‑to‑image generation that follows quantity specifications for multiple objects. Building upon this foundation, we further introduce an Adaptive Multi‑modal Guidance (AMG) module, which adaptively parses and integrates various combinations of multi‑source control signals to formulate corresponding structured intent. This intent subsequently guides selective constraints on scene layouts and object attributes, achieving dynamic fine‑grained control. Experimental results demonstrate that MoGen significantly outperforms existing methods in generation quality, quantity consistency, and fine‑grained control, while exhibiting superior accessibility and control flexibility. Code is available at: https://github.com/Tear‑kitty/MoGen/tree/master.
Authors:Qiwei Yang, Pingping Zhang, Yuhao Wang, Zijing Gong
Abstract:
Video‑based Person Re‑IDentification (VPReID) aims to retrieve the same person from videos captured by non‑overlapping cameras. At extreme far distances, VPReID is highly challenging due to severe resolution degradation, drastic viewpoint variation and inevitable appearance noise. To address these issues, we propose a Scale‑Adaptive framework with Shape Priors for VPReID, named SAS‑VPReID. The framework is built upon three complementary modules. First, we deploy a Memory‑Enhanced Visual Backbone (MEVB) to extract discriminative feature representations, which leverages the CLIP vision encoder and multi‑proxy memory. Second, we propose a Multi‑Granularity Temporal Modeling (MGTM) to construct sequences at multiple temporal granularities and adaptively emphasize motion cues across scales. Third, we incorporate Prior‑Regularized Shape Dynamics (PRSD) to capture body structure dynamics. With these modules, our framework can obtain more discriminative feature representations. Experiments on the VReID‑XFD benchmark demonstrate the effectiveness of each module and our final framework ranks the first on the VReID‑XFD challenge leaderboard. The source code is available at https://github.com/YangQiWei3/SAS‑VPReID.
Authors:Fuwen Luo, Zihao Wan, Ziyue Wang, Yaluo Liu, Pau Tong Lin Xu, Xuanjia Qiao, Xiaolong Wang, Peng Li, Yang Liu
Abstract:
Hieroglyphs, as logographic writing systems, encode rich semantic and cultural information within their internal structural composition. Yet, current advanced Large Language Models (LLMs) and Multimodal LLMs (MLLMs) usually remain structurally blind to this information. LLMs process characters as textual tokens, while MLLMs additionally view them as raw pixel grids. Both fall short to model the underlying logic of character strokes. Furthermore, existing structural analysis methods are often script‑specific and labor‑intensive. In this paper, we propose Hieroglyphic Stroke Analyzer (HieroSA), a novel and generalizable framework that enables MLLMs to automatically derive stroke‑level structures from character bitmaps without handcrafted data. It transforms modern logographic and ancient hieroglyphs character images into explicit, interpretable line‑segment representations in a normalized coordinate space, allowing for cross‑lingual generalization. Extensive experiments demonstrate that HieroSA effectively captures character‑internal structures and semantics, bypassing the need for language‑specific priors. Experimental results highlight the potential of our work as a graphematics analysis tool for a deeper understanding of hieroglyphic scripts. View our code at https://github.com/THUNLP‑MT/HieroSA.
Authors:Tingwei Xie, Jinxin He, Yonghong Song
Abstract:
The efficacy of Multimodal Transformers in visually‑rich document understanding (VrDU) is critically constrained by two inherent limitations: the lack of explicit modeling for logical reading order and the interference of visual tokens that dilutes attention on textual semantics.
To address these challenges, this paper presents ROAP, a lightweight and architecture‑agnostic pipeline designed to optimize attention distributions in Layout Transformers without altering their pre‑trained backbones.
The proposed pipeline first employs an Adaptive‑XY‑Gap (AXG‑Tree) to robustly extract hierarchical reading sequences from complex layouts. These sequences are then integrated into the attention mechanism via a Reading‑Order‑Aware Relative Position Bias (RO‑RPB). Furthermore, a Textual‑Token Sub‑block Attention Prior (TT‑Prior) is introduced to adaptively suppress visual noise and enhance fine‑grained text‑text interactions.
Extensive experiments on the FUNSD and CORD benchmarks demonstrate that ROAP consistently improves the performance of representative backbones, including LayoutLMv3 and GeoLayoutLM.
These findings confirm that explicitly modeling reading logic and regulating modality interference are critical for robust document understanding, offering a scalable solution for complex layout analysis. The implementation code will be released at https://github.com/KevinYuLei/ROAP.
Authors:Yuan-Kang Lee, Kuan-Lin Chen, Chia-Che Chang, Yu-Lun Liu
Abstract:
Nighttime color constancy still remains a challenging problem in computational photography due to low‑light noise and complex illumination conditions. We present RL‑AWB, a novel framework combining statistical methods with deep reinforcement learning for nighttime white balance. Our method begins with a statistical algorithm tailored for nighttime scenes, integrating salient gray pixel detection with novel illumination estimation. Building on this foundation, we develop the first deep reinforcement learning approach for color constancy that leverages the statistical algorithm as its core, mimicking professional AWB tuning experts by dynamically optimizing parameters for each image. To facilitate cross‑sensor evaluation, we introduce the first multi‑sensor nighttime dataset. Experiment results show that our method achieves superior generalization capability across low‑light and well‑illuminated images. Project page: https://ntuneillee.github.io/research/rl‑awb/
Authors:Gangwei Xu, Haotong Lin, Hongcheng Luo, Haiyang Sun, Bing Wang, Guang Chen, Sida Peng, Hangjun Ye, Xin Yang
Abstract:
Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying pixels and the loss of fine details. In this paper, we present pixel‑perfect visual geometry models that can predict high‑quality, flying‑pixel‑free point clouds by leveraging generative modeling in the pixel space. We first introduce Pixel‑Perfect Depth (PPD), a monocular depth foundation model built upon pixel‑space diffusion transformers (DiT). To address the high computational complexity associated with pixel‑space diffusion, we propose two key designs: 1) Semantics‑Prompted DiT, which incorporates semantic representations from vision foundation models to prompt the diffusion process, preserving global semantics while enhancing fine‑grained visual details; and 2) Cascade DiT architecture that progressively increases the number of image tokens, improving both efficiency and accuracy. To further extend PPD to video (PPVD), we introduce a new Semantics‑Consistent DiT, which extracts temporally consistent semantics from a multi‑view geometry foundation model. We then perform reference‑guided token propagation within the DiT to maintain temporal coherence with minimal computational and memory overhead. Our models achieve the best performance among all generative monocular and video depth estimation models and produce significantly cleaner point clouds than all other models.
Authors:Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Yu-Gang Jiang
Abstract:
Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) generates an expression for the selected object. Existing datasets and methods commonly support single‑target expressions only, i.e., one expression refers to one object, not considering multi‑target and no‑target expressions. This greatly limits the real applications of REx (RES/REC/REG). This paper introduces three new benchmarks called Generalized Referring Expression Segmentation (GRES), Comprehension (GREC), and Generation (GREG), collectively denoted as GREx, which extend the classic REx to allow expressions to identify an arbitrary number of objects. We construct the first large‑scale GREx dataset gRefCOCO that contains multi‑target, no‑target, and single‑target expressions and their corresponding images with labeled targets. GREx and gRefCOCO are designed to be backward‑compatible with REx, facilitating extensive experiments to study the performance gap of the existing REx methods on GREx tasks. One of the challenges of GRES/GREC is complex relationship modeling, for which we propose a baseline ReLA that adaptively divides the image into regions with sub‑instance clues and explicitly models the region‑region and region‑language dependencies. The proposed ReLA achieves the state‑of‑the‑art results on the both GRES and GREC tasks. The proposed gRefCOCO dataset and method are available at https://henghuiding.github.io/GREx.
Authors:Shuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen, Lemeng Wu, Zechun Liu, Chenchen Zhu, Zhipeng Cai, Chong Zhou, Haozhe Liu, Ernie Chang, Saksham Suri, Hongyu Xu, Qi Qian, Wei Wen, Balakrishnan Varadarajan, Zhuang Liu, Hu Xu, Florian Bordes, Raghuraman Krishnamoorthi, Bernard Ghanem, Vikas Chandra, Yunyang Xiong
Abstract:
Chain‑of‑thought (CoT) reasoning has emerged as a powerful tool for multimodal large language models on video understanding tasks. However, its necessity and advantages over direct answering remain underexplored. In this paper, we first demonstrate that for RL‑trained video models, direct answering often matches or even surpasses CoT performance, despite CoT producing step‑by‑step analyses at a higher computational cost. Motivated by this, we propose VideoAuto‑R1, a video understanding framework that adopts a reason‑when‑necessary strategy. During training, our approach follows a Thinking Once, Answering Twice paradigm: the model first generates an initial answer, then performs reasoning, and finally outputs a reviewed answer. Both answers are supervised via verifiable rewards. During inference, the model uses the confidence score of the initial answer to determine whether to proceed with reasoning. Across video QA and grounding benchmarks, VideoAuto‑R1 achieves state‑of‑the‑art accuracy with significantly improved efficiency, reducing the average response length by ~3.3x, e.g., from 149 to just 44 tokens. Moreover, we observe a low rate of thinking‑mode activation on perception‑oriented tasks, but a higher rate on reasoning‑intensive tasks. This suggests that explicit language‑based reasoning is generally beneficial but not always necessary.
Authors:Haoyu Zhao, Akide Liu, Zeyu Zhang, Weijie Wang, Feng Chen, Ruihan Zhu, Gholamreza Haffari, Bohan Zhuang
Abstract:
Embodied question answering (EQA) in 3D environments often requires collecting context that is distributed across multiple viewpoints and partially occluded. However, most recent vision‑‑language models (VLMs) are constrained to a fixed and finite set of input views, which limits their ability to acquire question‑relevant context at inference time and hinders complex spatial reasoning. We propose Chain‑of‑View (CoV) prompting, a training‑free, test‑time reasoning framework that transforms a VLM into an active viewpoint reasoner through a coarse‑to‑fine exploration process. CoV first employs a View Selection agent to filter redundant frames and identify question‑aligned anchor views. It then performs fine‑grained view adjustment by interleaving iterative reasoning with discrete camera actions, obtaining new observations from the underlying 3D scene representation until sufficient context is gathered or a step budget is reached.
We evaluate CoV on OpenEQA across four mainstream VLMs and obtain an average +11.56% improvement in LLM‑Match, with a maximum gain of +13.62% on Qwen3‑VL‑Flash. CoV further exhibits test‑time scaling: increasing the minimum action budget yields an additional +2.51% average improvement, peaking at +3.73% on Gemini‑2.5‑Flash. On ScanQA and SQA3D, CoV delivers strong performance (e.g., 116 CIDEr / 31.9 EM@1 on ScanQA and 51.1 EM@1 on SQA3D). Overall, these results suggest that question‑aligned view selection coupled with open‑view search is an effective, model‑agnostic strategy for improving spatial reasoning in 3D EQA without additional training. Code is available on https://github.com/ziplab/CoV .
Authors:Elia Peruzzo, Guillaume Sautière, Amirhossein Habibian
Abstract:
Autoregressive (AR) models have achieved remarkable success in image synthesis, yet their sequential nature imposes significant latency constraints. Speculative Decoding offers a promising avenue for acceleration, but existing approaches are limited by token‑level ambiguity and lack of spatial awareness. In this work, we introduce Multi‑Scale Local Speculative Decoding (MuLo‑SD), a novel framework that combines multi‑resolution drafting with spatially informed verification to accelerate AR image generation. Our method leverages a low‑resolution drafter paired with an up‑sampling step to propose candidate image tokens, which are then verified in parallel by a high‑resolution target model. Crucially, we incorporate a local rejection and resampling mechanism, enabling efficient correction of draft errors by focusing on spatial neighborhoods rather than raster‑scan resampling after the first rejection. When integrated with parallel decoding resampling, MuLo‑SD achieves substantial speedups ‑‑ up to \mathbf5× ‑‑ outperforming both speculative decoding and parallel decoding baselines in terms of acceleration, while maintaining comparable semantic alignment and perceptual quality. These results are validated using GenEval, DPG‑Bench, and FID/HPSv2 on the MS‑COCO 5k validation split. Extensive ablations highlight the impact of up‑sampling design, probability pooling, and local rejection and resampling with neighborhood expansion. Our approach sets a new state‑of‑the‑art in speculative decoding for image synthesis, bridging the gap between efficiency and fidelity. Project page is available at https://qualcomm‑ai‑research.github.io/mulo‑sd‑webpage/ .
Authors:Sixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li, Ying Shan, Yanwei Fu
Abstract:
Video world models aim to simulate dynamic, real‑world environments, yet existing methods struggle to provide unified and precise control over camera and multi‑object motion, as videos inherently capture dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a geometry‑driven video world model that generates dynamic, realistic videos from a unified 4D geometric world state. Our approach is centered on a novel 4D Geometric Control representation, which encodes the world state as a static background point cloud and per‑object 3D Gaussian trajectories. This representation captures each object's motion path and probabilistic 3D occupancy over time, providing a flexible, category‑agnostic alternative to rigid bounding boxes and parametric models. We render 4D Geometric Control into 4D control maps for a pretrained video diffusion model, enabling high‑fidelity, view‑consistent video generation that faithfully follows the specified dynamics. To enable training at scale, we develop an automatic data engine and construct VerseControl4D, a real‑world dataset of 35K training samples with automatically derived prompts and rendered 4D control maps. Extensive experiments show that VerseCrafter achieves superior visual quality and more accurate control over camera and multi‑object motion than prior methods.
Authors:Runze He, Yiji Cheng, Tiankai Hang, Zhimin Li, Yu Xu, Zijin Yin, Shiyi Zhang, Wenxun Dai, Penghui Du, Ao Ma, Chunyu Wang, Qinglin Lu, Jizhong Han, Jiao Dai
Abstract:
In‑context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image‑text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengths often fail to transfer effectively to image generation. We introduce Re‑Align, a unified framework that bridges the gap between understanding and generation through structured reasoning‑guided alignment. At its core lies the In‑Context Chain‑of‑Thought (IC‑CoT), a structured reasoning paradigm that decouples semantic guidance and reference association, providing clear textual target and mitigating confusion among reference images. Furthermore, Re‑Align introduces an effective RL training scheme that leverages a surrogate reward to measure the alignment between structured reasoning text and the generated image, thereby improving the model's overall performance on ICGE tasks. Extensive experiments verify that Re‑Align outperforms competitive methods of comparable model scale and resources on both in‑context image generation and editing tasks.
Authors:Zirui Wu, Zeren Jiang, Martin R. Oswald, Jie Song
Abstract:
Feed‑forward view synthesis models predict a novel view in a single pass with minimal 3D inductive bias. Existing works encode cameras as Plücker ray maps, which tie predictions to the arbitrary world coordinate gauge and make them sensitive to small camera transformations, thereby undermining geometric consistency. In this paper, we ask what inputs best condition a model for robust and consistent view synthesis. We propose projective conditioning, which replaces raw camera parameters with a target‑view projective cue that provides a stable 2D input. This reframes the task from a brittle geometric regression problem in ray space to a well‑conditioned target‑view image‑to‑image translation problem. Additionally, we introduce a masked autoencoding pretraining strategy tailored to this cue, enabling the use of large‑scale uncalibrated data for pretraining. Our method shows improved fidelity and stronger cross‑view consistency compared to ray‑conditioned baselines on our view‑consistency benchmark. It also achieves state‑of‑the‑art quality on standard novel view synthesis benchmarks.
Authors:Jens Bayer, Stefan Becker, David Münch, Michael Arens, Jürgen Beyerer
Abstract:
Higher‑order adversarial attacks can directly be considered the result of a cat‑and‑mouse game ‑‑ an elaborate action involving constant pursuit, near captures, and repeated escapes. This idiom describes the enduring circular training of adversarial attack patterns and adversarial training the best. The following work investigates the impact of higher‑order adversarial attacks on object detectors by successively training attack patterns and hardening object detectors with adversarial training. The YOLOv10 object detector is chosen as a representative, and adversarial patches are used in an evasion attack manner. Our results indicate that higher‑order adversarial patches are not only affecting the object detector directly trained on but rather provide a stronger generalization capacity compared to lower‑order adversarial patches. Moreover, the results highlight that solely adversarial training is not sufficient to harden an object detector efficiently against this kind of adversarial attack. Code: https://github.com/JensBayer/HigherOrder
Authors:Juyuan Kang, Hao Zhu, Yan Zhu, Wei Zhang, Jianing Chen, Tianxiang Xiao, Yike Ma, Hao Jiang, Feng Dai
Abstract:
Crop mapping based on satellite images time‑series (SITS) holds substantial economic value in agricultural production settings, in which parcel segmentation is an essential step. Existing approaches have achieved notable advancements in SITS segmentation with predetermined sequence lengths. However, we found that these approaches overlooked the generalization capability of models across scenarios with varying temporal length, leading to markedly poor segmentation results in such cases. To address this issue, we propose TEA, a TEmporal Adaptive SITS semantic segmentation method to enhance the model's resilience under varying sequence lengths. We introduce a teacher model that encapsulates the global sequence knowledge to guide a student model with adaptive temporal input lengths. Specifically, teacher shapes the student's feature space via intermediate embedding, prototypes and soft label perspectives to realize knowledge transfer, while dynamically aggregating student model to mitigate knowledge forgetting. Finally, we introduce full‑sequence reconstruction as an auxiliary task to further enhance the quality of representations across inputs of varying temporal lengths. Through extensive experiments, we demonstrate that our method brings remarkable improvements across inputs of different temporal lengths on common benchmarks. Our code will be publicly available.
Authors:Matan Kleiner, Lior Michaeli, Tomer Michaeli
Abstract:
Diffractive neural networks have recently emerged as a promising framework for all‑optical computing. However, these networks are typically trained for a single task, limiting their potential adoption in systems requiring multiple functionalities. Existing approaches to achieving multi‑task functionality either modify the mechanical configuration of the network per task or use a different illumination wavelength or polarization state for each task. In this work, we propose a new control mechanism, which is based on the illumination's angular spectrum. Specifically, we shape the illumination using an amplitude mask that selectively controls its angular spectrum. We employ different illumination masks for achieving different network functionalities, so that the mask serves as a unique task encoder. Interestingly, we show that effective control can be achieved over a very narrow angular range, within the paraxial regime. We numerically illustrate the proposed approach by training a single diffractive network to perform multiple image‑to‑image translation tasks. In particular, we demonstrate translating handwritten digits into typeset digits of different values, and translating handwritten English letters into typeset numbers and typeset Greek letters, where the type of the output is determined by the illumination's angular components. As we show, the proposed framework can work under different coherence conditions, and can be combined with existing control strategies, such as different wavelengths. Our results establish the illumination angular spectrum as a powerful degree of freedom for controlling diffractive networks, enabling a scalable and versatile framework for multi‑task all‑optical computing.
Authors:Denis Korzhenkov, Adil Karjauv, Animesh Karnewar, Mohsen Ghafoorian, Amirhossein Habibian
Abstract:
Recently proposed pyramidal models decompose the conventional forward and backward diffusion processes into multiple stages operating at varying resolutions. These models handle inputs with higher noise levels at lower resolutions, while less noisy inputs are processed at higher resolutions. This hierarchical approach significantly reduces the computational cost of inference in multi‑step denoising models. However, existing open‑source pyramidal video models have been trained from scratch and tend to underperform compared to state‑of‑the‑art systems in terms of visual plausibility. In this work, we present a pipeline that converts a pretrained diffusion model into a pyramidal one through low‑cost finetuning, achieving this transformation without degradation in quality of output videos. Furthermore, we investigate and compare various strategies for step distillation within pyramidal models, aiming to further enhance the inference efficiency. Our results are available at https://qualcomm‑ai‑research.github.io/PyramidalWan.
Authors:Yen-Jen Chiou, Wei-Tse Cheng, Yuan-Fu Yang
Abstract:
We present ProFuse, an efficient context‑aware framework for open‑vocabulary 3D scene understanding with 3D Gaussian Splatting (3DGS). The pipeline enhances cross‑view consistency and intra‑mask cohesion within a direct registration setup, adding minimal overhead and requiring no render‑supervised fine‑tuning. Instead of relying on a pretrained 3DGS scene, we introduce a dense correspondence‑guided pre‑registration phase that initializes Gaussians with accurate geometry while jointly constructing 3D Context Proposals via cross‑view clustering. Each proposal carries a global feature obtained through weighted aggregation of member embeddings, and this feature is fused onto Gaussians during direct registration to maintain per‑primitive language coherence across views. With associations established in advance, semantic fusion requires no additional optimization beyond standard reconstruction, and the model retains geometric refinement without densification. ProFuse achieves strong open‑vocabulary 3DGS understanding while completing semantic attachment in about five minutes per scene, which is two times faster than SOTA. Additional details are available at our project page https://chiou1203.github.io/ProFuse/.
Authors:Yanbing Zeng, Jia Wang, Hanghang Ma, Junqiang Wu, Jie Zhu, Xiaoming Wei, Jie Hu
Abstract:
Integrating image generation and understanding into a single framework has become a pivotal goal in the multimodal domain. However, how understanding can effectively assist generation has not been fully explored. Unlike previous works that focus on leveraging reasoning abilities and world knowledge from understanding models, this paper introduces a novel perspective: leveraging understanding to enhance the fidelity and detail richness of generated images. To this end, we propose Forge‑and‑Quench, a new unified framework that puts this principle into practice. In the generation process of our framework, an MLLM first reasons over the entire conversational context, including text instructions, to produce an enhanced text instruction. This refined instruction is then mapped to a virtual visual representation, termed the Bridge Feature, via a novel Bridge Adapter. This feature acts as a crucial link, forging insights from the understanding model to quench and refine the generation process. It is subsequently injected into the T2I backbone as a visual guidance signal, alongside the enhanced text instruction that replaces the original input. To validate this paradigm, we conduct comprehensive studies on the design of the Bridge Feature and Bridge Adapter. Our framework demonstrates exceptional extensibility and flexibility, enabling efficient migration across different MLLM and T2I models with significant savings in training overhead, all without compromising the MLLM's inherent multimodal understanding capabilities. Experiments show that Forge‑and‑Quench significantly improves image fidelity and detail across multiple models, while also maintaining instruction‑following accuracy and enhancing world knowledge application. Models and codes are available at https://github.com/YanbingZeng/Forge‑and‑Quench.
Authors:Ali Kurban, Wei Luo, Liangyu Zuo, Zeyu Zhang, Renda Han, Zhaolu Kang, Hao Tang
Abstract:
Cryptocurrency trading increasingly depends on timely integration of heterogeneous web information and market microstructure signals to support short‑horizon decision making under extreme volatility. However, existing trading systems struggle to jointly reason over noisy multi‑source web evidence while maintaining robustness to rapid price shocks at sub‑second timescales. The first challenge lies in synthesizing unstructured web content, social sentiment, and structured OHLCV signals into coherent and interpretable trading decisions without amplifying spurious correlations, while the second challenge concerns risk control, as slow deliberative reasoning pipelines are ill‑suited for handling abrupt market shocks that require immediate defensive responses. To address these challenges, we propose WebCryptoAgent, an agentic trading framework that decomposes web‑informed decision making into modality‑specific agents and consolidates their outputs into a unified evidence document for confidence‑calibrated reasoning. We further introduce a decoupled control architecture that separates strategic hourly reasoning from a real‑time second‑level risk model, enabling fast shock detection and protective intervention independent of the trading loop. Extensive experiments on real‑world cryptocurrency markets demonstrate that WebCryptoAgent improves trading stability, reduces spurious activity, and enhances tail‑risk handling compared to existing baselines. Code will be available at https://github.com/AIGeeksGroup/WebCryptoAgent.
Authors:Yang Zou, Xingyue Zhu, Kaiqi Han, Jun Ma, Xingyuan Li, Zhiying Jiang, Jinyuan Liu
Abstract:
Infrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super‑resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fail to restore turbulence‑induced distortions. Directly cascading turbulence mitigation (TM) algorithms with VSR methods leads to error propagation and accumulation due to the decoupled modeling of degradation between turbulence and resolution. We introduce HATIR, a Heat‑Aware Diffusion for Turbulent InfraRed Video Super‑Resolution, which injects heat‑aware deformation priors into the diffusion sampling path to jointly model the inverse process of turbulent degradation and structural detail loss. Specifically, HATIR constructs a Phasor‑Guided Flow Estimator, rooted in the physical principle that thermally active regions exhibit consistent phasor responses over time, enabling reliable turbulence‑aware flow to guide the reverse diffusion process. To ensure the fidelity of structural recovery under nonuniform distortions, a Turbulence‑Aware Decoder is proposed to selectively suppress unstable temporal cues and enhance edge‑aware feature aggregation via turbulence gating and structure‑aware attention. We built FLIR‑IVSR, the first dataset for turbulent infrared VSR, comprising paired LR‑HR sequences from a FLIR T1050sc camera (1024 X 768) spanning 640 diverse scenes with varying camera and object motion conditions. This encourages future research in infrared VSR. Project page: https://github.com/JZ0606/HATIR
Authors:Wentao Zhang, Mingkun Xu, Qi Zhang, Shangyang Li, Derek F. Wong, Lifei Wang, Yanchao Yang, Lina Lu, Tao Fang
Abstract:
Agricultural disease diagnosis challenges VLMs, as conventional fine‑tuning requires extensive labels, lacks interpretability, and generalizes poorly. While reasoning improves model robustness, existing methods rely on costly expert annotations and rarely address the open‑ended, diverse nature of agricultural queries. To address these limitations, we propose Agri‑R1, a reasoning‑enhanced large model for agriculture. Our framework automates high‑quality reasoning data generation via vision‑language synthesis and LLM‑based filtering, using only 19% of available samples. Training employs Group Relative Policy Optimization (GRPO) with a novel reward function that integrates domain‑specific lexicons and fuzzy matching to assess both correctness and linguistic flexibility in open‑ended responses. Evaluated on CDDMBench, our resulting 3B‑parameter model achieves performance competitive with 7B‑ to 13B‑parameter baselines, showing a +27.9% relative gain in disease recognition accuracy, +33.3% in agricultural knowledge QA, and a +26.10‑point improvement in cross‑domain generalization over standard fine‑tuning. These results suggest that automated reasoning synthesis paired with domain‑aware reward design may provide a broadly applicable paradigm for RL‑based VLM adaptation in data‑scarce specialized domains. Our code and data are publicly available at: https://github.com/CPJ‑Agricultural/Agri‑R1.
Authors:Paul Pu Liang
Abstract:
Our experience of the world is multisensory, spanning a synthesis of language, sight, sound, touch, taste, and smell. Yet, artificial intelligence has primarily advanced in digital modalities like text, vision, and audio. This paper outlines a research vision for multisensory artificial intelligence over the next decade. This new set of technologies can change how humans and AI experience and interact with one another, by connecting AI to the human senses and a rich spectrum of signals from physiological and tactile cues on the body, to physical and social signals in homes, cities, and the environment. We outline how this field must advance through three interrelated themes of sensing, science, and synergy. Firstly, research in sensing should extend how AI captures the world in richer ways beyond the digital medium. Secondly, developing a principled science for quantifying multimodal heterogeneity and interactions, developing unified modeling architectures and representations, and understanding cross‑modal transfer. Finally, we present new technical challenges to learn synergy between modalities and between humans and AI, covering multisensory integration, alignment, reasoning, generation, generalization, and experience. Accompanying this vision paper are a series of projects, resources, and demos of latest advances from the Multisensory Intelligence group at the MIT Media Lab, see https://mit‑mi.github.io/.
Authors:James Brock, Ce Zhang, Nantheera Anantrasirichai
Abstract:
Modern forest monitoring workflows increasingly benefit from the growing availability of high‑resolution satellite imagery and advances in deep learning. Two persistent challenges in this context are accurate pixel‑level change detection and meaningful semantic change captioning for complex forest dynamics. While large language models (LLMs) are being adapted for interactive data exploration, their integration with vision‑language models (VLMs) for remote sensing image change interpretation (RSICI) remains underexplored. To address this gap, we introduce an LLM‑driven agent for integrated forest change analysis that supports natural language querying across multiple RSICI tasks. The proposed system builds upon a multi‑level change interpretation (MCI) vision‑language backbone with LLM‑based orchestration. To facilitate adaptation and evaluation in forest environments, we further introduce the Forest‑Change dataset, which comprises bi‑temporal satellite imagery, pixel‑level change masks, and multi‑granularity semantic change captions generated using a combination of human annotation and rule‑based methods. Experimental results show that the proposed system achieves mIoU and BLEU‑4 scores of 67.10% and 40.17% on the Forest‑Change dataset, and 88.13% and 34.41% on LEVIR‑MCI‑Trees, a tree‑focused subset of LEVIR‑MCI benchmark for joint change detection and captioning. These results highlight the potential of interactive, LLM‑driven RSICI systems to improve accessibility, interpretability, and efficiency of forest change analysis. All data and code are publicly available at https://github.com/JamesBrockUoB/ForestChat.
Authors:Zhexiao Xiong, Xin Ye, Burhan Yaman, Sheng Cheng, Yiren Lu, Jingru Luo, Nathan Jacobs, Liu Ren
Abstract:
World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recent work has explored using vision‑language models (VLMs) for planning, yet existing approaches typically treat perception, prediction, and planning as separate modules. We propose UniDrive‑WM, a unified VLM‑based world model that jointly performs driving‑scene understanding, trajectory planning, and trajectory‑conditioned future image generation within a single architecture. UniDrive‑WM's trajectory planner predicts a future trajectory, which conditions a VLM‑based image generator to produce plausible future frames. These predictions provide additional supervisory signals that enhance scene understanding and iteratively refine trajectory generation. We further compare discrete and continuous output representations for future image prediction, analyzing their influence on downstream driving performance. Experiments on the challenging Bench2Drive benchmark show that UniDrive‑WM produces high‑fidelity future images and improves planning performance by 7.3% in L2 trajectory error and 10.4% in collision rate over the previous best method. These results demonstrate the advantages of tightly integrating VLM‑driven reasoning, planning, and generative world modeling for autonomous driving. The project page is available at https://unidrive‑wm.github.io/UniDrive‑WM.
Authors:Mohsen Ghafoorian, Amirhossein Habibian
Abstract:
Recent advances in video diffusion models have shifted towards transformer‑based architectures, achieving state‑of‑the‑art video generation but at the cost of quadratic attention complexity, which severely limits scalability for longer sequences. We introduce ReHyAt, a Recurrent Hybrid Attention mechanism that combines the fidelity of softmax attention with the efficiency of linear attention, enabling chunk‑wise recurrent reformulation and constant memory usage. Unlike the concurrent linear‑only SANA Video, ReHyAt's hybrid design allows efficient distillation from existing softmax‑based models, reducing the training cost by two orders of magnitude to ~160 GPU hours, while being competitive in the quality. Our light‑weight distillation and finetuning pipeline provides a recipe that can be applied to future state‑of‑the‑art bidirectional softmax‑based models. Experiments on VBench and VBench‑2.0, as well as a human preference study, demonstrate that ReHyAt achieves state‑of‑the‑art video quality while reducing attention cost from quadratic to linear, unlocking practical scalability for long‑duration and on‑device video generation. Project page is available at https://qualcomm‑ai‑research.github.io/rehyat.
Authors:Xueqing Wu, Zihan Xue, Da Yin, Shuyan Zhou, Kai-Wei Chang, Nanyun Peng, Yeming Wen
Abstract:
We present FronTalk, a benchmark for front‑end code generation that pioneers the study of a unique interaction dynamic: conversational code generation with multi‑modal feedback. In front‑end development, visual artifacts such as sketches, mockups and annotated creenshots are essential for conveying design intent, yet their role in multi‑turn code generation remains largely unexplored. To address this gap, we focus on the front‑end development task and curate FronTalk, a collection of 100 multi‑turn dialogues derived from real‑world websites across diverse domains such as news, finance, and art. Each turn features both a textual instruction and an equivalent visual instruction, each representing the same user intent. To comprehensively evaluate model performance, we propose a novel agent‑based evaluation framework leveraging a web agent to simulate users and explore the website, and thus measuring both functional correctness and user experience. Evaluation of 20 models reveals two key challenges that are under‑explored systematically in the literature: (1) a significant forgetting issue where models overwrite previously implemented features, resulting in task failures, and (2) a persistent challenge in interpreting visual feedback, especially for open‑source vision‑language models (VLMs). We propose a strong baseline to tackle the forgetting issue with AceCoder, a method that critiques the implementation of every past instruction using an autonomous web agent. This approach significantly reduces forgetting to nearly zero and improves the performance by up to 9.3% (56.0% to 65.3%). Overall, we aim to provide a solid foundation for future research in front‑end development and the general interaction dynamics of multi‑turn, multi‑modal code generation. Code and data are released at https://github.com/shirley‑wu/frontalk
Authors:Yanzhe Lyu, Chen Geng, Karthik Dharmarajan, Yunzhi Zhang, Hadi Alzayer, Shangzhe Wu, Jiajun Wu
Abstract:
Dynamic objects in our physical 4D (3D + time) world are constantly evolving, deforming, and interacting with other objects, leading to diverse 4D scene dynamics. In this paper, we present a universal generative pipeline, CHORD, for CHOReographing Dynamic objects and scenes and synthesizing this type of phenomena. Traditional rule‑based graphics pipelines to create these dynamics are based on category‑specific heuristics, yet are labor‑intensive and not scalable. Recent learning‑based methods typically demand large‑scale datasets, which may not cover all object categories in interest. Our approach instead inherits the universality from the video generative models by proposing a distillation‑based pipeline to extract the rich Lagrangian motion information hidden in the Eulerian representations of 2D videos. Our method is universal, versatile, and category‑agnostic. We demonstrate its effectiveness by conducting experiments to generate a diverse range of multi‑body 4D dynamics, show its advantage compared to existing methods, and demonstrate its applicability in generating robotics manipulation policies. Project page: https://yanzhelyu.github.io/chord
Authors:Xudong Jiang, Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Marc Pollefeys
Abstract:
Existing visual localization methods are typically either 2D image‑based, which are easy to build and maintain but limited in effective geometric reasoning, or 3D structure‑based, which achieve high accuracy but require a centralized reconstruction and are difficult to update. In this work, we revisit visual localization with a 2D image‑based representation and propose to augment each image with estimated depth maps to capture the geometric structure. Supported by the effective use of dense matchers, this representation is not only easy to build and maintain, but achieves highest accuracy in challenging conditions. With compact compression and a GPU‑accelerated LO‑RANSAC implementation, the whole pipeline is efficient in both storage and computation and allows for a flexible trade‑off between accuracy and highest memory efficiency. Our method achieves a new state‑of‑the‑art accuracy on various standard benchmarks and outperforms existing memory‑efficient methods at comparable map sizes. Code will be available at https://github.com/cvg/Hierarchical‑Localization.
Authors:Yifan Wang, Yanyu Li, Gordon Guocheng Qian, Sergey Tulyakov, Yun Fu, Anil Kag
Abstract:
Video diffusion alignment has been heavily relied on scalar rewards. These rewards are typically derived from learned reward models in human preference datasets, requiring additional training and extensive collection. Moreover, scalar rewards provide coarse, global supervision, offering limited prompt‑generation mismatch credit assignment and making models prone to reward exploitation and unstable optimization. We propose Diffusion‑DRF, a free, rich, and differentiable reward framework for video diffusion fine‑tuning. Diffusion‑DRF employs a frozen, off‑the‑shelf Vision‑Language Model (VLM) as the critic, eliminating the need for reward model training. Instead of relying on a single scalar reward, it decomposes each user prompt into multi‑dimensional questions with freeform dense VQA explanation queries, yielding information‑rich feedback. By direct differentiable optimization over this rich feedback, Diffusion‑DRF achieves stable reward‑based tuning without preference datasets collection. Diffusion‑DRF achieves significant gains both quantitatively and qualitatively, outperforming state‑of‑the‑art Flow‑GRPO by 4.74% in overall performance on unseen VBench‑2.0.
Authors:Jiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma, Yuewen Ma, Yiyi Liao
Abstract:
We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene‑level 3D generation. We repurpose the VGGT reconstruction model to produce geometric latents by training an adapter on its tokens, which are regularized to align with the appearance latents of pre‑trained video diffusion models. By jointly generating these disentangled yet aligned latents, Gen3R produces both RGB videos and corresponding 3D geometry, including camera poses, depth maps, and global point clouds. Experiments demonstrate that our approach achieves state‑of‑the‑art results in single‑ and multi‑image conditioned 3D scene generation. Additionally, our method can enhance the robustness of reconstruction by leveraging generative priors, demonstrating the mutual benefit of tightly coupling reconstruction and generative models.
Authors:Zitong Huang, Kaidong Zhang, Yukang Ding, Chao Gao, Rui Ding, Ying Chen, Wangmeng Zuo
Abstract:
Aligning text‑to‑video diffusion models with human preferences is crucial for generating high‑quality videos. Existing Direct Preference Otimization (DPO) methods rely on multi‑sample ranking and task‑specific critic models, which is inefficient and often yields ambiguous global supervision. To address these limitations, we propose LocalDPO, a novel post‑training framework that constructs localized preference pairs from real videos and optimizes alignment at the spatio‑temporal region level. We design an automated pipeline to efficiently collect preference pair data that generates preference pairs with a single inference per prompt, eliminating the need for external critic models or manual annotation. Specifically, we treat high‑quality real videos as positive samples and generate corresponding negatives by locally corrupting them with random spatio‑temporal masks and restoring only the masked regions using the frozen base model. During training, we introduce a region‑aware DPO loss that restricts preference learning to corrupted areas for rapid convergence. Experiments on Wan2.1 and CogVideoX demonstrate that LocalDPO consistently improves video fidelity, temporal coherence and human preference scores over other post‑training approaches, establishing a more efficient and fine‑grained paradigm for video generator alignment.The code is available at https://github.com/1170300714/Local‑DPO.
Authors:Onur Keleş, A. Murat Tekalp
Abstract:
Neural networks commonly employ the McCulloch‑Pitts neuron model, which is a linear model followed by a point‑wise non‑linear activation. Various researchers have already advanced inherently non‑linear neuron models, such as quadratic neurons, generalized operational neurons, generative neurons, and super neurons, which offer stronger non‑linearity compared to point‑wise activation functions. In this paper, we introduce a novel and better non‑linear neuron model called Padé neurons (Paons), inspired by Padé approximants. Paons offer several advantages, such as diversity of non‑linearity, since each Paon learns a different non‑linear function of its inputs, and layer efficiency, since Paons provide stronger non‑linearity in much fewer layers compared to piecewise linear approximation. Furthermore, Paons include all previously proposed neuron models as special cases, thus any neuron model in any network can be replaced by Paons. We note that there has been a proposal to employ the Padé approximation as a generalized point‑wise activation function, which is fundamentally different from our model. To validate the efficacy of Paons, in our experiments, we replace classic neurons in some well‑known neural image super‑resolution, compression, and classification models based on the ResNet architecture with Paons. Our comprehensive experimental results and analyses demonstrate that neural models built by Paons provide better or equal performance than their classic counterparts with a smaller number of layers. The PyTorch implementation code for Paon is open‑sourced at https://github.com/onur‑keles/Paon.
Authors:Junle Liu, Peirong Zhang, Yuyi Zhang, Pengyu Yan, Hui Zhou, Xinyue Zhou, Fengjun Guo, Lianwen Jin
Abstract:
Commercial‑grade poster design demands the seamless integration of aesthetic appeal with precise, informative content delivery. Current automated poster generation systems face significant limitations, including incomplete design workflows, poor text rendering accuracy, and insufficient flexibility for commercial applications. To address these challenges, we propose PosterVerse, a full‑workflow, commercial‑grade poster generation method that seamlessly automates the entire design process while delivering high‑density and scalable text rendering. PosterVerse replicates professional design through three key stages: (1) blueprint creation using fine‑tuned LLMs to extract key design elements from user requirements, (2) graphical background generation via customized diffusion models to create visually appealing imagery, and (3) unified layout‑text rendering with an MLLM‑powered HTML engine to guarantee high text accuracy and flexible customization. In addition, we introduce PosterDNA, a commercial‑grade, HTML‑based dataset tailored for training and validating poster design models. To the best of our knowledge, PosterDNA is the first Chinese poster generation dataset to introduce HTML typography files, enabling scalable text rendering and fundamentally solving the challenges of rendering small and high‑density text. Experimental results demonstrate that PosterVerse consistently produces commercial‑grade posters with appealing visuals, accurate text alignment, and customizable layouts, making it a promising solution for automating commercial poster design. The code and model are available at https://github.com/wuhaer/PosterVerse.
Authors:Xu Zhang, Cheng Da, Huan Yang, Kun Gai, Ming Lu, Zhan Ma
Abstract:
Existing 1D visual tokenizers for autoregressive (AR) generation largely follow the design principles of language modeling, as they are built directly upon transformers whose priors originate in language, yielding single‑hierarchy latent tokens and treating visual data as flat sequential token streams. However, this language‑like formulation overlooks key properties of vision, particularly the hierarchical and residual network designs that have long been essential for convergence and efficiency in visual models. To bring "vision" back to vision, we propose the Residual Tokenizer (ResTok), a 1D visual tokenizer that builds hierarchical residuals for both image tokens and latent tokens. The hierarchical representations obtained through progressively merging enable cross‑level feature fusion at each layer, substantially enhancing representational capacity. Meanwhile, the semantic residuals between hierarchies prevent information overlap, yielding more concentrated latent distributions that are easier for AR modeling. Cross‑level bindings consequently emerge without any explicit constraints. To accelerate the generation process, we further introduce a hierarchical AR generator that substantially reduces sampling steps by predicting an entire level of latent tokens at once rather than generating them strictly token‑by‑token. Extensive experiments demonstrate that restoring hierarchical residual priors in visual tokenization significantly improves AR image generation, achieving a gFID of 2.34 on ImageNet‑256 with only 9 sampling steps. Code is available at https://github.com/Kwai‑Kolors/ResTok.
Authors:Jan Tagscherer, Sarah de Boer, Lena Philipp, Fennie van der Graaf, Dré Peeters, Joeran Bosma, Lars Leijten, Bogdan Obreja, Ewoud Smit, Alessa Hering
Abstract:
Developing foundation models in medical imaging requires continuous monitoring of downstream performance. Researchers are burdened with tracking numerous experiments, design choices, and their effects on performance, often relying on ad‑hoc, manual workflows that are inherently slow and error‑prone. We introduce EvalBlocks, a modular, plug‑and‑play framework for efficient evaluation of foundation models during development. Built on Snakemake, EvalBlocks supports seamless integration of new datasets, foundation models, aggregation methods, and evaluation strategies. All experiments and results are tracked centrally and are reproducible with a single command, while efficient caching and parallel execution enable scalable use on shared compute infrastructure. Demonstrated on five state‑of‑the‑art foundation models and three medical imaging classification tasks, EvalBlocks streamlines model evaluation, enabling researchers to iterate faster and focus on model innovation rather than evaluation logistics. The framework is released as open source software at https://github.com/DIAGNijmegen/eval‑blocks.
Authors:Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu, Dieter Fox, Kaichun Mo, Li Fei-Fei
Abstract:
Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre‑trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: given one or few RGB‑D images and a sequence of low‑level robot action commands, PointWorld forecasts per‑pixel displacements in 3D that respond to the given actions. By representing actions as 3D point flows instead of embodiment‑specific action spaces (e.g., joint positions), this formulation directly conditions on physical geometries of robots while seamlessly integrating learning across embodiments. To train our 3D world model, we curate a large‑scale dataset spanning real and simulated robotic manipulation in open‑world environments, enabled by recent advances in 3D vision and simulated environments, totaling about 2M trajectories and 500 hours across a single‑arm Franka and a bimanual humanoid. Through rigorous, large‑scale empirical studies of backbones, action representations, learning objectives, partial observability, data mixtures, domain transfers, and scaling, we distill design principles for large‑scale 3D world modeling. With a real‑time (0.1s) inference speed, PointWorld can be efficiently integrated in the model‑predictive control (MPC) framework for manipulation. We demonstrate that a single pre‑trained checkpoint enables a real‑world Franka robot to perform rigid‑body pushing, deformable and articulated object manipulation, and tool use, without requiring any demonstrations or post‑training and all from a single image captured in‑the‑wild. Project website at https://point‑world.github.io/.
Authors:Yunhao Liang, Ruixuan Ying, Bo Li, Hong Li, Kai Yan, Qingwen Li, Min Yang, Okamoto Satoshi, Zhe Cui, Shiwen Ni
Abstract:
DeepSeek‑OCR utilizes an optical 2D mapping approach to achieve high‑ratio vision‑text compression, claiming to decode text tokens exceeding ten times the input visual tokens. While this suggests a promising solution for the LLM long‑context bottleneck, we investigate a critical question: "Visual merit or linguistic crutch ‑ which drives DeepSeek‑OCR's performance?" By employing sentence‑level and word‑level semantic corruption, we isolate the model's intrinsic OCR capabilities from its language priors. Results demonstrate that without linguistic support, DeepSeek‑OCR's performance plummets from approximately 90% to 20%. Comparative benchmarking against 13 baseline models reveals that traditional pipeline OCR methods exhibit significantly higher robustness to such semantic perturbations than end‑to‑end methods. Furthermore, we find that lower visual token counts correlate with increased reliance on priors, exacerbating hallucination risks. Context stress testing also reveals a total model collapse around 10,000 text tokens, suggesting that current optical compression techniques may paradoxically aggravate the long‑context bottleneck. This study empirically defines DeepSeek‑OCR's capability boundaries and offers essential insights for future optimizations of the vision‑text compression paradigm. We release all data, results and scripts used in this study at https://github.com/dududuck00/DeepSeekOCR.
Authors:Siddarth Nilol Kundur Satish, Devesh Jaiswal, Hongyu Chen, Abhishek Bakshi
Abstract:
Current video generation models produce high‑quality aesthetic videos but often struggle to learn representations of real‑world physics dynamics, resulting in artifacts such as unnatural object collisions, inconsistent gravity, and temporal flickering. In this work, we propose PhysVideoGenerator, a proof‑of‑concept framework that explicitly embeds a learnable physics prior into the video generation process. We introduce a lightweight predictor network, PredictorP, which regresses high‑level physical features extracted from a pre‑trained Video Joint Embedding Predictive Architecture (V‑JEPA 2) directly from noisy diffusion latents. These predicted physics tokens are injected into the temporal attention layers of a DiT‑based generator (Latte) via a dedicated cross‑attention mechanism. Our primary contribution is demonstrating the technical feasibility of this joint training paradigm: we show that diffusion latents contain sufficient information to recover V‑JEPA 2 physical representations, and that multi‑task optimization remains stable over training. This report documents the architectural design, technical challenges, and validation of training stability, establishing a foundation for future large‑scale evaluation of physics‑aware generative models.
Authors:Jiangyuan Liu, Yuhao Zhao, Hongxuan Ma, Zhe Liu, Jian Wang, Wei Zou
Abstract:
Point cloud completion aims to recover complete 3D geometry from partial observations caused by limited viewpoints and occlusions. Existing learning‑based works, including 3D Convolutional Neural Network (CNN)‑based, point‑based, and Transformer‑based methods, have achieved strong performance on synthetic benchmarks. However, due to the limitations of modality, scalability, and generative capacity, their generalization to novel objects and real‑world scenarios remains challenging. In this paper, we propose MGPC, a generalizable multimodal point cloud completion framework that integrates point clouds, RGB images, and text within a unified architecture. MGPC introduces an innovative modality dropout strategy, a Transformer‑based fusion module, and a novel progressive generator to improve robustness, scalability, and geometric modeling capability. We further develop an automatic data generation pipeline and construct MGPC‑1M, a large‑scale benchmark with over 1,000 categories and one million training pairs. Extensive experiments on MGPC‑1M and in‑the‑wild data demonstrate that the proposed method consistently outperforms prior baselines and exhibits strong generalization under real‑world conditions.
Authors:Jinsong Zhou, Yihua Du, Xinli Xu, Luozhou Wang, Zijie Zhuang, Yehang Zhang, Shuaibo Li, Xiaojun Hu, Bolan Su, Ying-cong Chen
Abstract:
Maintaining consistent characters, props, and environments across multiple shots is a central challenge in narrative video generation. Existing models can produce high‑quality short clips but often fail to preserve entity identity and appearance when scenes change or when entities reappear after long temporal gaps. We present VideoMemory, an entity‑centric framework that integrates narrative planning with visual generation through a Dynamic Memory Bank. Given a structured script, a multi‑agent system decomposes the narrative into shots, retrieves entity representations from memory, and synthesizes keyframes and videos conditioned on these retrieved states. The Dynamic Memory Bank stores explicit visual and semantic descriptors for characters, props, and backgrounds, and is updated after each shot to reflect story‑driven changes while preserving identity. This retrieval‑update mechanism enables consistent portrayal of entities across distant shots and supports coherent long‑form generation. To evaluate this setting, we construct a 54‑case multi‑shot consistency benchmark covering character‑, prop‑, and background‑persistent scenarios. Extensive experiments show that VideoMemory achieves strong entity‑level coherence and high perceptual quality across diverse narrative sequences.
Authors:Qianyu Guo, Jingrong Wu, Jieji Ren, Weifeng Ge, Wenqiang Zhang
Abstract:
Few‑shot segmentation (FSS) aims to rapidly learn novel class concepts from limited examples to segment specific targets in unseen images, and has been widely applied in areas such as medical diagnosis and industrial inspection. However, existing studies largely overlook the complex environmental factors encountered in real world scenarios‑such as illumination, background, and camera viewpoint‑which can substantially increase the difficulty of test images. As a result, models trained under laboratory conditions often fall short of practical deployment requirements. To bridge this gap, in this paper, an environment‑robust FSS setting is introduced that explicitly incorporates challenging test cases arising from complex environments‑such as motion blur, small objects, and camouflaged targets‑to enhance model's robustness under realistic, dynamic conditions. An environment robust FSS benchmark (ER‑FSS) is established, covering eight datasets across multiple real world scenarios. In addition, an Adaptive Attention Distillation (AAD) method is proposed, which repeatedly contrasts and distills key shared semantics between known (support) and unknown (query) images to derive class‑specific attention for novel categories. This strengthens the model's ability to focus on the correct targets in complex environments, thereby improving environmental robustness. Comparative experiments show that AAD improves mIoU by 3.3% ‑ 8.5% across all datasets and settings, demonstrating superior performance and strong generalization. The source code and dataset are available at: https://github.com/guoqianyu‑alberta/Adaptive‑Attention‑Distillation‑for‑FSS.
Authors:Zhongbin Guo, Zhen Yang, Yushan Li, Xinyue Zhang, Wenyu Gao, Jiacheng Wang, Chengzhi Li, Xiangrui Liu, Ping Jian
Abstract:
Recent advancements in Spatial Intelligence (SI) have predominantly relied on Vision‑Language Models (VLMs), yet a critical question remains: does spatial understanding originate from visual encoders or the fundamental reasoning backbone? Inspired by this question, we introduce SiT‑Bench, a novel benchmark designed to evaluate the SI performance of Large Language Models (LLMs) without pixel‑level input, comprises over 3,800 expert‑annotated items across five primary categories and 17 subtasks, ranging from egocentric navigation and perspective transformation to fine‑grained robotic manipulation. By converting single/multi‑view scenes into high‑fidelity, coordinate‑aware textual descriptions, we challenge LLMs to perform symbolic textual reasoning rather than visual pattern matching. Evaluation results of state‑of‑the‑art (SOTA) LLMs reveals that while models achieve proficiency in localized semantic tasks, a significant "spatial gap" remains in global consistency. Notably, we find that explicit spatial reasoning significantly boosts performance, suggesting that LLMs possess latent world‑modeling potential. Our proposed dataset SiT‑Bench serves as a foundational resource to foster the development of spatially‑grounded LLM backbones for future VLMs and embodied agents. Our code and benchmark will be released at https://github.com/binisalegend/SiT‑Bench .
Authors:Guobin Tu, Di Weng
Abstract:
Sign Language Translation (SLT) is a challenging cross‑modal task requiring joint modeling of manual articulations and non‑manual signals. Existing gloss‑free SLT methods effectively capture gestural dynamics but often underutilize facial expressions, which play crucial grammatical and disambiguating roles. This limitation can cause semantic degradation when distinct concepts share similar manual configurations. To address this issue, we propose FEA‑SLT (Facial‑Expression‑Aware Sign Language Translation), a gloss‑free end‑to‑end framework that uses facial dynamics as semantic anchors for resolving manual ambiguity. FEA‑SLT employs a domain‑transferred facial encoder to extract expression‑sensitive representations and integrates them with manual features through a linguistically constrained Facial‑Expression‑Aware Fusion (FEAF) module. FEAF captures reciprocal dependencies between manual and facial channels via bidirectional modulation, enhancing syntactic fidelity. Experiments on PHOENIX14T and CSL‑Daily show that FEA‑SLT achieves state‑of‑the‑art BLEU performance among gloss‑free methods, while targeted analyses confirm improved translation of facial‑sensitive utterances. Code is available at [https://github.com/TuGuobin/FEA‑SLT](https://github.com/TuGuobin/FEA‑SLT).
Authors:Joshua Salako
Abstract:
Scalability and data sparsity remain critical bottlenecks for collaborative filtering on massive interaction datasets. This work investigates the latent geometry of user preferences using the MovieLens 32M dataset, implementing a high‑performance, parallelized Alternating Least Squares (ALS) framework. Through extensive hyperparameter optimization, we demonstrate that constrained low‑rank models significantly outperform higher dimensional counterparts in generalization, achieving an optimal balance between Root Mean Square Error (RMSE) and ranking precision. We visualize the learned embedding space to reveal the unsupervised emergence of semantic genre clusters, confirming that the model captures deep structural relationships solely from interaction data. Finally, we validate the system's practical utility in a cold‑start scenario, introducing a tunable scoring parameter to manage the trade‑off between popularity bias and personalized affinity effectively. The codebase for this research can be found here: https://github.com/joshsalako/recommender.git
Authors:M. Akın Yılmaz, Ahmet Bilican, Burak Can Biner, A. Murat Tekalp
Abstract:
Image restoration has traditionally required training specialized models on thousands of paired examples per degradation type. We challenge this paradigm by demonstrating that powerful pre‑trained text‑conditioned image editing models can be efficiently adapted for multiple restoration tasks through parameter‑efficient fine‑tuning with remarkably few examples. Our approach fine‑tunes LoRA adapters on FLUX.1 Kontext, a state‑of‑the‑art 12B parameter flow matching model for image‑to‑image translation, using only 16‑128 paired images per task, guided by simple text prompts that specify the restoration operation. Unlike existing methods that train specialized restoration networks from scratch with thousands of samples, we leverage the rich visual priors already encoded in large‑scale pre‑trained editing models, dramatically reducing data requirements while maintaining high perceptual quality. A single unified LoRA adapter, conditioned on task‑specific text prompts, effectively handles multiple degradations including denoising, deraining, and dehazing. Through comprehensive ablation studies, we analyze: (i) the impact of training set size on restoration quality, (ii) trade‑offs between task‑specific versus unified multi‑task adapters, (iii) the role of text encoder fine‑tuning, and (iv) zero‑shot baseline performance. While our method prioritizes perceptual quality over pixel‑perfect reconstruction metrics like PSNR/SSIM, our results demonstrate that pre‑trained image editing models, when properly adapted, offer a compelling and data‑efficient alternative to traditional image restoration approaches, opening new avenues for few‑shot, prompt‑guided image enhancement. The code to reproduce our results are available at: https://github.com/makinyilmaz/Edit2Restore
Authors:Oran Duan, Yinghua Shen, Yingzhu Lv, Luyang Jie, Yaxin Liu, Qiong Wu
Abstract:
Advances in generative models and sequence learning have greatly promoted research in dance motion generation, yet current methods still suffer from coarse semantic control and poor coherence in long sequences. In this work, we present Listen to Rhythm, Choose Movements (LRCM), a multimodal‑guided diffusion framework supporting both diverse input modalities and autoregressive dance motion generation. We explore a feature decoupling paradigm for dance datasets and generalize it to the Motorica Dance dataset, separating motion capture data, audio rhythm, and professionally annotated global and local text descriptions. Our diffusion architecture integrates an audio‑latent Conformer and a text‑latent Cross‑Conformer, and incorporates a Motion Temporal Mamba Module (MTMM) to enable smooth, long‑duration autoregressive synthesis. Experimental results indicate that LRCM delivers strong performance in both functional capability and quantitative metrics, demonstrating notable potential in multimodal input scenarios and extended sequence generation. The project page is available at https://oranduanstudy.github.io/LRCM/.
Authors:Hexiao Lu, Xiaokun Sun, Zeyu Cai, Hao Guo, Ying Tai, Jian Yang, Zhenyu Zhang
Abstract:
We present Muses, the first training‑free method for fantastic 3D creature generation in a feed‑forward paradigm. Previous methods, which rely on part‑aware optimization, manual assembly, or 2D image generation, often produce unrealistic or incoherent 3D assets due to the challenges of intricate part‑level manipulation and limited out‑of‑domain generation. In contrast, Muses leverages the 3D skeleton, a fundamental representation of biological forms, to explicitly and rationally compose diverse elements. This skeletal foundation formalizes 3D content creation as a structure‑aware pipeline of design, composition, and generation. Muses begins by constructing a creatively composed 3D skeleton with coherent layout and scale through graph‑constrained reasoning. This skeleton then guides a voxel‑based assembly process within a structured latent space, integrating regions from different objects. Finally, image‑guided appearance modeling under skeletal conditions is applied to generate a style‑consistent and harmonious texture for the assembled shape. Extensive experiments establish Muses' state‑of‑the‑art performance in terms of visual fidelity and alignment with textual descriptions, and potential on flexible 3D object editing. Project page: https://luhexiao.github.io/Muses.github.io/.
Authors:Anees Ur Rehman Hashmi, Numan Saeed, Christoph Lippert
Abstract:
Multimodal medical large language models have shown substantial progress in chest X‑ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techniques improve overall performance, they often fail to establish a true anatomical correspondence, resulting in incorrect anatomical understanding in the medical domain. To address this gap, we introduce AnatomiX, a multitask multimodal large language model for anatomically grounded chest X‑ray interpretation. Inspired by the radiological workflow, AnatomiX adopts a two stage approach: first, it identifies anatomical structures and extracts their features, and then leverages a large language model to perform diverse downstream tasks such as phrase grounding, report generation, visual question answering, and image understanding. Extensive experiments across multiple benchmarks demonstrate that AnatomiX achieves superior anatomical reasoning and delivers over 25% improvement in performance on anatomy grounding, phrase grounding, grounded diagnosis and grounded captioning tasks compared to existing approaches. Code and pretrained model are available at https://aneesurhashmi.github.io/anatomix
Authors:Matěj Pekár, Vít Musil, Rudolf Nenutil, Petr Holub, Tomáš Brázdil
Abstract:
Precise and scalable instance segmentation of cell nuclei is essential for computational pathology, yet gigapixel Whole‑Slide Images pose major computational challenges. Existing approaches rely on patch‑based processing and costly post‑processing for instance separation, sacrificing context and efficiency. We introduce LSP‑DETR (Local Star Polygon DEtection TRansformer), a fully end‑to‑end framework that uses a lightweight transformer with linear complexity to process substantially larger images without additional computational cost. Nuclei are represented as star‑convex polygons, and a novel radial distance loss function allows the segmentation of overlapping nuclei to emerge naturally, without requiring explicit overlap annotations or handcrafted post‑processing. Evaluations on PanNuke and MoNuSeg show strong generalization across tissues and state‑of‑the‑art efficiency, with LSP‑DETR being over five times faster than the next‑fastest leading method. Code and models are available at https://github.com/RationAI/lsp‑detr.
Authors:Guoqiang Liang, Jianyi Wang, Zhonghua Wu, Shangchen Zhou
Abstract:
Image Quality Assessment (IQA) is a long‑standing problem in computer vision. Previous methods typically focus on predicting numerical scores without explanation or providing low‑level descriptions lacking precise scores. Recent reasoning‑based vision language models (VLMs) have shown strong potential for IQA by jointly generating quality descriptions and scores. However, existing VLM‑based IQA methods often suffer from unreliable reasoning due to their limited capability of integrating visual and textual cues. In this work, we introduce Zoom‑IQA, a VLM‑based IQA model to explicitly emulate key cognitive behaviors: uncertainty awareness, region reasoning, and iterative refinement. Specifically, we present a two‑stage training pipeline: 1) supervised fine‑tuning (SFT) on our Grounded‑Rationale‑IQA (GR‑IQA) dataset to teach the model to ground its assessments in key regions, and 2) reinforcement learning (RL) for dynamic policy exploration, stabilized by our KL‑Coverage regularizer to prevent reasoning and scoring diversity collapse, with a Progressive Re‑sampling Strategy for mitigating annotation bias. Extensive experiments show that Zoom‑IQA achieves improved robustness, explainability, and generalization. The application to downstream tasks, such as image restoration, further demonstrates the effectiveness of Zoom‑IQA.
Authors:Mengtian Li, Jinshu Chen, Songtao Zhao, Wanquan Feng, Pengqi Tu, Qian He
Abstract:
Video stylization, an important downstream task of video generation models, has not yet been thoroughly explored. Its input style conditions typically include text, style image, and stylized first frame. Each condition has a characteristic advantage: text is more flexible, style image provides a more accurate visual anchor, and stylized first frame makes long‑video stylization feasible. However, existing methods are largely confined to a single type of style condition, which limits their scope of application. Additionally, their lack of high‑quality datasets leads to style inconsistency and temporal flicker. To address these limitations, we introduce DreamStyle, a unified framework for video stylization, supporting (1) text‑guided, (2) style‑image‑guided, and (3) first‑frame‑guided video stylization, accompanied by a well‑designed data curation pipeline to acquire high‑quality paired video data. DreamStyle is built on a vanilla Image‑to‑Video (I2V) model and trained using a Low‑Rank Adaptation (LoRA) with token‑specific up matrices that reduces the confusion among different condition tokens. Both qualitative and quantitative evaluations demonstrate that DreamStyle is competent in all three video stylization tasks, and outperforms the competitors in style consistency and video quality.
Authors:Boyu Chang, Qi Wang, Xi Guo, Zhixiong Nan, Yazhou Yao, Tianfei Zhou
Abstract:
Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general‑purpose multimodal reasoning capabilities, they fall short in abductive inference, as compared to human beings. To bridge this gap, we draw inspiration from the interplay between verbal and pictorial abduction in human cognition, and propose to strengthen abduction of MLLMs by mimicking such dual‑mode behavior. Concretely, we introduce AbductiveMLLM comprising of two synergistic components: REASONER and IMAGINER. The REASONER operates in the verbal domain. It first explores a broad space of possible explanations using a blind LLM and then prunes visually incongruent hypotheses based on cross‑modal causal alignment. The remaining hypotheses are introduced into the MLLM as targeted priors, steering its reasoning toward causally coherent explanations. The IMAGINER, on the other hand, further guides MLLMs by emulating human‑like pictorial thinking. It conditions a text‑to‑image diffusion model on both the input video and the REASONER's output embeddings to "imagine" plausible visual scenes that correspond to verbal explanation, thereby enriching MLLMs' contextual grounding. The two components are trained jointly in an end‑to‑end manner. Experiments on standard VAR benchmarks show that AbductiveMLLM achieves state‑of‑the‑art performance, consistently outperforming traditional solutions and advanced MLLMs.
Authors:Xu Zhang, Huan Zhang, Guoli Wang, Qian Zhang, Lefei Zhang
Abstract:
All‑in‑One Image Restoration (AiOIR) has advanced significantly, offering promising solutions for complex real‑world degradations. However, most existing approaches rely heavily on degradation‑specific representations, often resulting in oversmoothing and artifacts. To address this, we propose ClearAIR, a novel AiOIR framework inspired by Human Visual Perception (HVP) and designed with a hierarchical, coarse‑to‑fine restoration strategy. First, leveraging the global priority of early HVP, we employ a Multimodal Large Language Model (MLLM)‑based Image Quality Assessment (IQA) model for overall evaluation. Unlike conventional IQA, our method integrates cross‑modal understanding to more accurately characterize complex, composite degradations. Building upon this overall assessment, we then introduce a region awareness and task recognition pipeline. A semantic cross‑attention, leveraging semantic guidance unit, first produces coarse semantic prompts. Guided by this regional context, a degradation‑aware module implicitly captures region‑specific degradation characteristics, enabling more precise local restoration. Finally, to recover fine details, we propose an internal clue reuse mechanism. It operates in a self‑supervised manner to mine and leverage the intrinsic information of the image itself, substantially enhancing detail restoration. Experimental results show that ClearAIR achieves superior performance across diverse synthetic and real‑world datasets.
Authors:Zeyu Ren, Zeyu Zhang, Wukai Li, Qingxiang Liu, Hao Tang
Abstract:
Monocular depth estimation aims to recover the depth information of 3D scenes from 2D images. Recent work has made significant progress, but its reliance on large‑scale datasets and complex decoders has limited its efficiency and generalization ability. In this paper, we propose a lightweight and data‑centric framework for zero‑shot monocular depth estimation. We first adopt DINOv3 as the visual encoder to obtain high‑quality dense features. Secondly, to address the inherent drawbacks of the complex structure of the DPT, we design the Simple Depth Transformer (SDT), a compact transformer‑based decoder. Compared to the DPT, it uses a single‑path feature fusion and upsampling process to reduce the computational overhead of cross‑scale feature fusion, achieving higher accuracy while reducing the number of parameters by approximately 85%‑89%. Furthermore, we propose a quality‑based filtering strategy to filter out harmful samples, thereby reducing dataset size while improving overall training quality. Extensive experiments on five benchmarks demonstrate that our framework surpasses the DPT in accuracy. This work highlights the importance of balancing model design and data quality for achieving efficient and generalizable zero‑shot depth estimation. Code: https://github.com/AIGeeksGroup/AnyDepth. Website: https://aigeeksgroup.github.io/AnyDepth.
Authors:Hyungtae Lim, Minkyun Seo, Luca Carlone, Jaesik Park
Abstract:
Some deep learning‑based point cloud registration methods struggle with zero‑shot generalization, often requiring dataset‑specific hyperparameter tuning or retraining for new environments. We identify three critical limitations: (a) fixed user‑defined parameters (e.g., voxel size, search radius) that fail to generalize across varying scales, (b) learned keypoint detectors exhibit poor cross‑domain transferability, and (c) absolute coordinates amplify scale mismatches between datasets. To address these three issues, we present BUFFER‑X, a training‑free registration framework that achieves zero‑shot generalization through: (a) geometric bootstrapping for automatic hyperparameter estimation, (b) distribution‑aware farthest point sampling to replace learned detectors, and (c) patch‑level coordinate normalization to ensure scale consistency. Our approach employs hierarchical multi‑scale matching to extract correspondences across local, middle, and global receptive fields, enabling robust registration in diverse environments. For efficiency‑critical applications, we introduce BUFFER‑X‑Lite, which reduces total computation time by 43% (relative to BUFFER‑X) through early exit strategies and fast pose solvers while preserving accuracy. We evaluate on a comprehensive benchmark comprising 12 datasets spanning object‑scale, indoor, and outdoor scenes, including cross‑sensor registration between heterogeneous LiDAR configurations. Results demonstrate that our approach generalizes effectively without manual tuning or prior knowledge of test domains. Code: https://github.com/MIT‑SPARK/BUFFER‑X.
Authors:Zanting Ye, Xiaolong Niu, Xuanbin Wu, Xu Han, Shengyuan Liu, Jing Hao, Zhihao Peng, Hao Sun, Jieqin Lv, Fanghu Wang, Yanchao Huang, Hubing Wu, Yixuan Yuan, Habib Zaidi, Arman Rahmim, Yefeng Zheng, Lijun Lu
Abstract:
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in tasks such as abnormality detection and report generation for anatomical modalities, their capability in functional imaging remains largely unexplored. In this work, we identify and quantify a fundamental functional perception gap: the inability of current vision encoders to decode functional tracer biodistribution independent of morphological priors. Identifying Positron Emission Tomography (PET) as the quintessential modality to investigate this disconnect, we introduce PET‑Bench, the first large‑scale functional imaging benchmark comprising 52,308 hierarchical QA pairs from 9,732 multi‑site, multi‑tracer PET studies. Extensive evaluation of 19 state‑of‑the‑art MLLMs reveals a critical safety hazard termed the Chain‑of‑Thought (CoT) hallucination trap. We observe that standard CoT prompting, widely considered to enhance reasoning, paradoxically decouples linguistic generation from visual evidence in PET, producing clinically fluent but factually ungrounded diagnoses. To resolve this, we propose Atomic Visual Alignment (AVA), a simple fine‑tuning strategy that enforces the mastery of low‑level functional perception prior to high‑level diagnostic reasoning. Our results demonstrate that AVA effectively bridges the perception gap, transforming CoT from a source of hallucination into a robust inference tool and improving diagnostic accuracy by up to 14.83%. Code and data are available at https://github.com/yezanting/PET‑Bench.
Authors:Aniruddha Mahapatra, Long Mai, Cusuh Ham, Feng Liu
Abstract:
Cinemagraphs, which combine static photographs with selective, looping motion, offer unique artistic appeal. Generating them from a single photograph in a controllable manner is particularly challenging. Existing image‑animation techniques are restricted to simple, low‑frequency motions and operate only in narrow domains with repetitive textures like water and smoke. In contrast, large‑scale video diffusion models are not tailored for cinemagraph constraints and lack the specialized data required to generate seamless, controlled loops. We present DreamLoop, a controllable video synthesis framework dedicated to generating cinemagraphs from a single photo without requiring any cinemagraph training data. Our key idea is to adapt a general video diffusion model by training it on two objectives: temporal bridging and motion conditioning. This strategy enables flexible cinemagraph generation. During inference, by using the input image as both the first‑ and last‑ frame condition, we enforce a seamless loop. By conditioning on static tracks, we maintain a static background. Finally, by providing a user‑specified motion path for a target object, our method provides intuitive control over the animation's trajectory and timing. To our knowledge, DreamLoop is the first method to enable cinemagraph generation for general scenes with flexible and intuitive controls. We demonstrate that our method produces high‑quality, complex cinemagraphs that align with user intent, outperforming existing approaches.
Authors:Jyothi Rikhab Chand, Mathews Jacob
Abstract:
Solving inverse problems in imaging requires models that support efficient inference, uncertainty quantification, and principled probabilistic reasoning. Energy‑Based Models (EBMs), with their interpretable energy landscapes and compositional structure, are well‑suited for this task but have historically suffered from high computational costs and training instability. To overcome the historical shortcomings of EBMs, we introduce a fast distillation strategy to transfer the strengths of pre‑trained diffusion models into multi‑scale EBMs. These distilled EBMs enable efficient sampling and preserve the interpretability and compositionality inherent to potential‑based frameworks. Leveraging EBM compositionality, we propose Annealed Langevin Posterior Sampling (ALPS) algorithm for Maximum‑A‑Posteriori (MAP), Minimum Mean Square Error (MMSE), and uncertainty estimates for inverse problems in imaging. Unlike diffusion models that use complex guidance strategies for latent variables, we perform annealing on static posterior distributions that are well‑defined and composable. Experiments on image inpainting and MRI reconstruction demonstrate that our method matches or surpasses diffusion‑based baselines in both accuracy and efficiency, while also supporting MAP recovery. Overall, our framework offers a scalable and principled solution for inverse problems in imaging, with potential for practical deployment in scientific and clinical settings. ALPS code is available at the GitHub repository \hrefhttps://github.com/JyoChand/ALPSALPS.
Authors:Souhail Hadgi, Bingchen Gong, Ramana Sundararaman, Emery Pierson, Lei Li, Peter Wonka, Maks Ovsjanikov
Abstract:
Current foundation models for 3D shapes excel at global tasks (retrieval, classification) but transfer poorly to local part‑level reasoning. Recent approaches leverage vision and language foundation models to directly solve dense tasks through multi‑view renderings and text queries. While promising, these pipelines require expensive inference over multiple renderings, depend heavily on large language‑model (LLM) prompt engineering for captions, and fail to exploit the inherent 3D geometry of shapes. We address this gap by introducing an encoder‑only 3D model that produces language‑aligned patch‑level features directly from point clouds. Our pre‑training approach builds on existing data engines that generate part‑annotated 3D shapes by pairing multi‑view SAM regions with VLM captioning. Using this data, we train a point cloud transformer encoder in two stages: (1) distillation of dense 2D features from visual encoders such as DINOv2 into 3D patches, and (2) alignment of these patch embeddings with part‑level text embeddings through a multi‑positive contrastive objective. Our 3D encoder achieves zero‑shot 3D part segmentation with fast single‑pass inference without any test‑time multi‑view rendering, while significantly outperforming previous rendering‑based and feed‑forward approaches across several 3D part segmentation benchmarks. Project website: https://souhail‑hadgi.github.io/patchalign3dsite/
Authors:Yuan Li, Shin'ya Nishida
Abstract:
Textual reasoning has recently been widely adopted in Blind Image Quality Assessment (BIQA). However, it remains unclear how textual information contributes to quality prediction and to what extent text can represent the score‑related image contents. This work addresses these questions from an information‑flow perspective by comparing existing BIQA models with three paradigms designed to learn the image‑text‑score relationship: Chain‑of‑Thought, Self‑Consistency, and Autoencoder. Our experiments show that the score prediction performance of the existing model significantly drops when only textual information is used for prediction. Whereas the Chain‑of‑Thought paradigm introduces little improvement in BIQA performance, the Self‑Consistency paradigm significantly reduces the gap between image‑ and text‑conditioned predictions, narrowing the PLCC/SRCC difference to 0.02/0.03. The Autoencoder‑like paradigm is less effective in closing the image‑text gap, yet it reveals a direction for further optimization. These findings provide insights into how to improve the textual reasoning for BIQA and high‑level vision tasks.
Authors:Wenting Lu, Didi Zhu, Tao Shen, Donglin Zhu, Ayong Ye, Chao Wu
Abstract:
Multi‑modal reasoning requires the seamless integration of visual and linguistic cues, yet existing Chain‑of‑Thought methods suffer from two critical limitations in cross‑modal scenarios: (1) over‑reliance on single coarse‑grained image regions, and (2) semantic fragmentation between successive reasoning steps. To address these issues, we propose the CoCoT (Collaborative Coross‑modal Thought) framework, built upon two key innovations: a) Dynamic Multi‑Region Grounding to adaptively detect the most relevant image regions based on the question, and b) Relation‑Aware Reasoning to enable multi‑region collaboration by iteratively aligning visual cues to form a coherent and logical chain of thought. Through this approach, we construct the CoCoT‑70K dataset, comprising 74,691 high‑quality samples with multi‑region annotations and structured reasoning chains. Extensive experiments demonstrate that CoCoT significantly enhances complex visual reasoning, achieving an average accuracy improvement of 15.4% on LLaVA‑1.5 and 4.0% on Qwen2‑VL across six challenging benchmarks. The data and code are available at: https://github.com/deer‑echo/CoCoT.
Authors:Kaede Shiohara, Toshihiko Yamasaki, Vladislav Golyanik
Abstract:
Detecting unknown deepfake manipulations remains one of the most challenging problems in face forgery detection. Current state‑of‑the‑art approaches fail to generalize to unseen manipulations, as they primarily rely on supervised training with existing deepfakes or pseudo‑fakes, which leads to overfitting to specific forgery patterns. In contrast, self‑supervised methods offer greater potential for generalization, but existing work struggles to learn discriminative representations only from self‑supervision. In this paper, we propose ExposeAnyone, a fully self‑supervised approach based on a diffusion model that generates expression sequences from audio. The key idea is, once the model is personalized to specific subjects using reference sets, it can compute the identity distances between suspected videos and personalized subjects via diffusion reconstruction errors, enabling person‑of‑interest face forgery detection. Extensive experiments demonstrate that 1) our method outperforms the previous state‑of‑the‑art method by 4.22 percentage points in the average AUC on DF‑TIMIT, DFDCP, KoDF, and IDForge datasets, 2) our model is also capable of detecting Sora2‑generated videos, where the previous approaches perform poorly, and 3) our method is highly robust to corruptions such as blur and compression, highlighting the applicability in real‑world face forgery detection.
Authors:Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, Weicai Ye
Abstract:
We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task‑specific models or independent modules for each modality, VINO uses a shared diffusion backbone that conditions on text, images and videos, enabling a broad range of visual creation and editing tasks under one model. Specifically, VINO couples a vision‑language model (VLM) with a Multimodal Diffusion Transformer (MMDiT), where multimodal inputs are encoded as interleaved conditioning tokens, and then used to guide the diffusion process. This design supports multi‑reference grounding, long‑form instruction following, and coherent identity preservation across static and dynamic content, while avoiding modality‑specific architectural components. To train such a unified system, we introduce a multi‑stage training pipeline that progressively expands a video generation base model into a unified, multi‑task generator capable of both image and video input and output. Across diverse generation and editing benchmarks, VINO demonstrates strong visual quality, faithful instruction following, improved reference and attribute preservation, and more controllable multi‑identity edits. Our results highlight a practical path toward scalable unified visual generation, and the promise of interleaved, in‑context computation as a foundation for general‑purpose visual creation.
Authors:Jing Tan, Zhaoyang Zhang, Yantao Shen, Jiarui Cai, Shuo Yang, Jiajun Wu, Wei Xia, Zhuowen Tu, Stefano Soatto
Abstract:
We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text‑instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for multimodal generation systems. While existing text‑based manipulation methods can adjust appearance or style, they struggle to perform object‑level geometric transformations‑such as translating, rotating, or resizing objects‑due to scarce paired supervision and pixel‑level optimization limits. Talk2Move employs Group Relative Policy Optimization (GRPO) to explore geometric actions through diverse rollouts generated from input images and lightweight textual variations, removing the need for costly paired data. A spatial reward guided model aligns geometric transformations with linguistic description, while off‑policy step evaluation and active step sampling improve learning efficiency by focusing on informative transformation stages. Furthermore, we design object‑centric spatial rewards that evaluate displacement, rotation, and scaling behaviors directly, enabling interpretable and coherent transformations. Experiments on curated benchmarks demonstrate that Talk2Move achieves precise, consistent, and semantically faithful object transformations, outperforming existing text‑guided editing approaches in both spatial accuracy and scene coherence.
Authors:Saurabh Kaushik, Lalit Maurya, Beth Tellman
Abstract:
Geo‑Foundation Models (GFMs), have proven effective in diverse downstream applications, including semantic segmentation, classification, and regression tasks. However, in case of flood mapping using Sen1Flood11 dataset as a downstream task, GFMs struggles to outperform the baseline U‑Net, highlighting model's limitation in capturing critical local nuances. To address this, we present the Prithvi‑Complementary Adaptive Fusion Encoder (CAFE), which integrate Prithvi GFM pretrained encoder with a parallel CNN residual branch enhanced by Convolutional Attention Modules (CAM). Prithvi‑CAFE enables fast and efficient fine‑tuning through adapters in Prithvi and performs multi‑scale, multi‑level fusion with CNN features, capturing critical local details while preserving long‑range dependencies. We achieve state‑of‑the‑art results on two comprehensive flood mapping datasets: Sen1Flood11 and FloodPlanet. On Sen1Flood11 test data, Prithvi‑CAFE (IoU 83.41) outperforms the original Prithvi (IoU 82.50) and other major GFMs (TerraMind 82.90, DOFA 81.54, spectralGPT: 81.02). The improvement is even more pronounced on the hold‑out test site, where Prithvi‑CAFE achieves an IoU of 81.37 compared to the baseline U‑Net (70.57) and original Prithvi (72.42). On FloodPlanet, Prithvi‑CAFE also surpasses the baseline U‑Net and other GFMs, achieving an IoU of 64.70 compared to U‑Net (60.14), Terramind (62.33), DOFA (59.15) and Prithvi 2.0 (61.91). Our proposed simple yet effective Prithvi‑CAFE demonstrates strong potential for improving segmentation tasks where multi‑channel and multi‑modal data provide complementary information and local details are critical. The code is released on \hrefhttps://github.com/Sk‑2103/Prithvi‑CAFEPrithvi‑CAFE Github
Authors:Xiaopeng Guo, Yinzhe Xu, Huajian Huang, Sai-Kit Yeung
Abstract:
Monocular omnidirectional visual odometry (OVO) systems leverage 360‑degree cameras to overcome field‑of‑view limitations of perspective VO systems. However, existing methods, reliant on handcrafted features or photometric objectives, often lack robustness in challenging scenarios, such as aggressive motion and varying illumination. To address this, we present 360DVO, the first deep learning‑based OVO framework. Our approach introduces a distortion‑aware spherical feature extractor (DAS‑Feat) that adaptively learns distortion‑resistant features from 360‑degree images. These sparse feature patches are then used to establish constraints for effective pose estimation within a novel omnidirectional differentiable bundle adjustment (ODBA) module. To facilitate evaluation in realistic settings, we also contribute a new real‑world OVO benchmark. Extensive experiments on this benchmark and public synthetic datasets (TartanAir V2 and 360VO) demonstrate that 360DVO surpasses state‑of‑the‑art baselines (including 360VO and OpenVSLAM), improving robustness by 50% and accuracy by 37.5%. Homepage: https://chris1004336379.github.io/360DVO‑homepage
Authors:Tom Burgert, Leonard Hackel, Paolo Rota, Begüm Demir
Abstract:
Self‑supervised learning (SSL) has become a powerful paradigm for learning from large, unlabeled datasets, particularly in computer vision (CV). However, applying SSL to multispectral remote sensing (RS) images presents unique challenges and opportunities due to the geographical and temporal variability of the data. In this paper, we introduce GeoRank, a novel regularization method for contrastive SSL that improves upon prior techniques by directly optimizing spherical distances to embed geographical relationships into the learned feature space. GeoRank outperforms or matches prior methods that integrate geographical metadata and consistently improves diverse contrastive SSL algorithms (e.g., BYOL, DINO). Beyond this, we present a systematic investigation of key adaptations of contrastive SSL for multispectral RS images, including the effectiveness of data augmentations, the impact of dataset cardinality and image size on performance, and the task dependency of temporal views. Code is available at https://github.com/tomburgert/georank.
Authors:Shuai Yuan, Yantai Yang, Xiaotian Yang, Xupeng Zhang, Zhonghao Zhao, Lingming Zhang, Zhipeng Zhang
Abstract:
The grand vision of enabling persistent, large‑scale 3D visual geometry understanding is shackled by the irreconcilable demands of scalability and long‑term stability. While offline models like VGGT achieve inspiring geometry capability, their batch‑based nature renders them irrelevant for live systems. Streaming architectures, though the intended solution for live operation, have proven inadequate. Existing methods either fail to support truly infinite‑horizon inputs or suffer from catastrophic drift over long sequences. We shatter this long‑standing dilemma with InfiniteVGGT, a causal visual geometry transformer that operationalizes the concept of a rolling memory through a bounded yet adaptive and perpetually expressive KV cache. Capitalizing on this, we devise a training‑free, attention‑agnostic pruning strategy that intelligently discards obsolete information, effectively ``rolling'' the memory forward with each new frame. Fully compatible with FlashAttention, InfiniteVGGT finally alleviates the compromise, enabling infinite‑horizon streaming while outperforming existing streaming methods in long‑term stability. The ultimate test for such a system is its performance over a truly infinite horizon, a capability that has been impossible to rigorously validate due to the lack of extremely long‑term, continuous benchmarks. To address this critical gap, we introduce the Long3D benchmark, which, for the first time, enables a rigorous evaluation of continuous 3D geometry estimation on sequences about 10,000 frames. This provides the definitive evaluation platform for future research in long‑term 3D geometry understanding. Code is available at: https://github.com/AutoLab‑SAI‑SJTU/InfiniteVGGT
Authors:Salim Khazem
Abstract:
Foundation segmentation models such as the Segment Anything Model (SAM) exhibit strong zero‑shot generalization through large‑scale pretraining, but adapting them to domain‑specific semantic segmentation remains challenging, particularly for thin structures (e.g., retinal vessels) and noisy modalities (e.g., SAR imagery). Full fine‑tuning is computationally expensive and risks catastrophic forgetting. We propose TopoLoRA‑SAM, a topology‑aware and parameter‑efficient adaptation framework for binary semantic segmentation. TopoLoRA‑SAM injects Low‑Rank Adaptation (LoRA) into the frozen ViT encoder, augmented with a lightweight spatial convolutional adapter and optional topology‑aware supervision via differentiable clDice. We evaluate our approach on five benchmarks spanning retinal vessel segmentation (DRIVE, STARE, CHASE\_DB1), polyp segmentation (Kvasir‑SEG), and SAR sea/land segmentation (SL‑SSDD), comparing against U‑Net, DeepLabV3+, SegFormer, and Mask2Former. TopoLoRA‑SAM achieves the best retina‑average Dice and the best overall average Dice across datasets, while training only 5.2% of model parameters (~4.9M). On the challenging CHASE\_DB1 dataset, our method substantially improves segmentation accuracy and robustness, demonstrating that topology‑aware parameter‑efficient adaptation can match or exceed fully fine‑tuned specialist models. Code is available at : https://github.com/salimkhazem/Seglab.git
Authors:Renke Wang, Zhenyu Zhang, Ying Tai, Jun Li, Jian Yang
Abstract:
Precise human mesh recovery (HMR) from multi‑view images remains challenging: end‑to‑end methods produce entangled errors hard to localize, while fitting‑based methods rely on sparse keypoints that provide limited surface constraints. We observe that the true bottleneck lies in the quality of intermediate representations, and that dense pixel‑to‑surface correspondences can be effectively generated by repurposing pre‑trained diffusion models with rich visual priors. We propose DiffProxy, a Stable‑Diffusion‑based framework trained on large‑scale synthetic data with pixel‑perfect annotations. A multi‑conditional proxy generator predicts dense correspondences from multi‑view images, providing uniform surface constraints that enable precise fitting. Hand refinement feeds enlarged hand crops alongside full‑body images for fine‑grained detail, while test‑time scaling exploits diffusion stochasticity to estimate per‑pixel uncertainty. Trained only on synthetic data, DiffProxy achieves state‑of‑the‑art results on five diverse real‑world benchmarks. Project page: https://wrk226.github.io/DiffProxy.html
Authors:Shikun Sun, Liao Qu, Huichao Zhang, Yiheng Liu, Yangyang Song, Xian Li, Xu Wang, Yi Jiang, Daniel K. Du, Xinglong Wu, Jia Jia
Abstract:
Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous input structures across their generation steps, which creates severe asynchronous policy conflicts. This issue becomes particularly acute in reinforcement learning (RL) scenarios, leading to unstable training and suboptimal alignment. To resolve this, we propose a novel framework to enhance Group Relative Policy Optimization (GRPO) by explicitly managing these conflicts. Our method integrates three synergistic components: 1) a stabilizing intermediate reward to guide early‑stage generation; 2) a dynamic time‑step reweighting scheme for precise credit assignment; and 3) a novel mask propagation algorithm, derived from principles of Reward Feedback Learning (ReFL), designed to isolate optimization effects both spatially and temporally. Our approach demonstrates significant improvements in sample quality and objective alignment over the vanilla GRPO baseline, enabling robust and effective optimization for VAR models.
Authors:Jingjing Wang, Zhuo Xiao, Xinning Yao, Bo Liu, Lijuan Niu, Xiangzhi Bai, Fugen Zhou
Abstract:
Accurate detection of ultrasound nodules is essential for the early diagnosis and treatment of thyroid and breast cancers. However, this task remains challenging due to irregular nodule shapes, indistinct boundaries, substantial scale variations, and the presence of speckle noise that degrades structural visibility. To address these challenges, we propose a prior‑guided DETR framework specifically designed for ultrasound nodule detection. Instead of relying on purely data‑driven feature learning, the proposed framework progressively incorporates different prior knowledge at multiple stages of the network. First, a Spatially‑adaptive Deformable FFN with Prior Regularization (SDFPR) is embedded into the CNN backbone to inject geometric priors into deformable sampling, stabilizing feature extraction for irregular and blurred nodules. Second, a Multi‑scale Spatial‑Frequency Feature Mixer (MSFFM) is designed to extract multi‑scale structural priors, where spatial‑domain processing emphasizes contour continuity and boundary cues, while frequency‑domain modeling captures global morphology and suppresses speckle noise. Furthermore, a Dense Feature Interaction (DFI) mechanism propagates and exploits these prior‑modulated features across all encoder layers, enabling the decoder to enhance query refinement under consistent geometric and structural guidance. Experiments conducted on two clinically collected thyroid ultrasound datasets (Thyroid I and Thyroid II) and two public benchmarks (TN3K and BUSI) for thyroid and breast nodules demonstrate that the proposed method achieves superior accuracy compared with 18 detection methods, particularly in detecting morphologically complex nodules.The source code is publicly available at https://github.com/wjj1wjj/Ultrasound‑DETR.
Authors:Dachun Kai, Zeyu Xiao, Huyue Zhu, Jiaxiao Wang, Yueyi Zhang, Xiaoyan Sun
Abstract:
This paper addresses low‑light video super‑resolution (LVSR), aiming to restore high‑resolution videos from low‑light, low‑resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high‑frequency information. To overcome these challenges, we present RetinexEVSR, the first event‑driven LVSR framework that leverages high‑contrast event signals and Retinex‑inspired priors to enhance video quality under low‑light scenarios. Unlike previous approaches that directly fuse degraded signals, RetinexEVSR introduces a novel bidirectional cross‑modal fusion strategy to extract and integrate meaningful cues from noisy event data and degraded RGB frames. Specifically, an illumination‑guided event enhancement module is designed to progressively refine event features using illumination maps derived from the Retinex model, thereby suppressing low‑light artifacts while preserving high‑contrast details. Furthermore, we propose an event‑guided reflectance enhancement module that utilizes the enhanced event features to dynamically recover reflectance details via a multi‑scale fusion mechanism. Experimental results show that our RetinexEVSR achieves state‑of‑the‑art performance on three datasets. Notably, on the SDSD benchmark, our method can get up to 2.95 dB gain while reducing runtime by 65% compared to prior event‑based methods. Code: https://github.com/DachunKai/RetinexEVSR.
Authors:Huichao Zhang, Liao Qu, Yiheng Liu, Hang Chen, Yangyang Song, Yongsheng Dong, Shikun Sun, Xian Li, Xu Wang, Yi Jiang, Hu Ye, Bo Chen, Yiming Gao, Peng Liu, Akide Liu, Zhipeng Yang, Qili Deng, Linjie Xing, Jiyang Liu, Zhao Wang, Yang Zhou, Mingcong Liu, Yi Zhang, Qian He, Xiwei Hu, Zhongqi Qi, Jie Shao, Zhiye Fu, Shuai Wang, Fangmin Chen, Xuezhi Chai, Zhihua Wu, Yitong Wang, Zehuan Yuan, Daniel K. Du, Xinglong Wu
Abstract:
We present NextFlow, a unified decoder‑only autoregressive transformer trained on 6 trillion interleaved text‑image discrete tokens. By leveraging a unified vision representation within a unified autoregressive architecture, NextFlow natively activates multimodal understanding and generation capabilities, unlocking abilities of image editing, interleaved content and video generation. Motivated by the distinct nature of modalities ‑ where text is strictly sequential and images are inherently hierarchical ‑ we retain next‑token prediction for text but adopt next‑scale prediction for visual generation. This departs from traditional raster‑scan methods, enabling the generation of 1024x1024 images in just 5 seconds ‑ orders of magnitude faster than comparable AR models. We address the instabilities of multi‑scale generation through a robust training recipe. Furthermore, we introduce a prefix‑tuning strategy for reinforcement learning. Experiments demonstrate that NextFlow achieves state‑of‑the‑art performance among unified models and rivals specialized diffusion baselines in visual quality.
Authors:Jiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang, Shifeng Chen
Abstract:
Amid the surge in generic text‑to‑video generation, the field of personalized human video generation has witnessed notable advancements, primarily concentrated on single‑person scenarios. However, to our knowledge, the domain of two‑person interactions, particularly in the context of martial arts combat, remains uncharted. We identify a significant gap: existing models for single‑person dancing generation prove insufficient for capturing the subtleties and complexities of two engaged fighters, resulting in challenges such as identity confusion, anomalous limbs, and action mismatches. To address this, we introduce a pioneering new task, Personalized Martial Arts Combat Video Generation. Our approach, MagicFight, is specifically crafted to overcome these hurdles. Given this pioneering task, we face a lack of appropriate datasets. Thus, we generate a bespoke dataset using the game physics engine Unity, meticulously crafting a multitude of 3D characters, martial arts moves, and scenes designed to represent the diversity of combat. MagicFight refines and adapts existing models and strategies to generate high‑fidelity two‑person combat videos that maintain individual identities and ensure seamless, coherent action sequences, thereby laying the groundwork for future innovations in the realm of interactive video content creation.
Website: https://MingfuYAN.github.io/MagicFight/
Dataset: https://huggingface.co/datasets/MingfuYAN/KungFu‑Fiesta
Authors:Peizhuo Li, Sebastian Starke, Yuting Ye, Olga Sorkine-Hornung
Abstract:
Ballroom dancing is a structured yet expressive motion category. Its highly diverse movement and complex interactions between leader and follower dancers make the understanding and synthesis challenging. We demonstrate that the three‑point trajectory available from a virtual reality (VR) device can effectively serve as a dancer's motion descriptor, simplifying the modeling and synthesis of interplay between dancers' full‑body motions down to sparse trajectories. Thanks to the low dimensionality, we can employ an efficient MLP network to predict the follower's three‑point trajectory directly from the leader's three‑point input for certain types of ballroom dancing, addressing the challenge of modeling high‑dimensional full‑body interaction. It also prevents our method from overfitting thanks to its compact yet explicit representation. By leveraging the inherent structure of the movements and carefully planning the autoregressive procedure, we show a deterministic neural network is able to translate three‑point trajectories into a virtual embodied avatar, which is typically considered under‑constrained and requires generative models for common motions. In addition, we demonstrate this deterministic approach generalizes beyond small, structured datasets like ballroom dancing, and performs robustly on larger, more diverse datasets such as LaFAN. Our method provides a computationally‑ and data‑efficient solution, opening new possibilities for immersive paired dancing applications. Code and pre‑trained models for this paper are available at https://peizhuoli.github.io/dancing‑points.
Authors:Zhehuan Cao, Fiseha Berhanu Tesema, Ping Fu, Jianfeng Ren, Ahmed Nasr
Abstract:
Glacial segmentation is essential for reconstructing past glacier dynamics and evaluating climate‑driven landscape change. However, weak optical contrast and the limited availability of high‑resolution DEMs hinder automated mapping. This study introduces the first large‑scale optical‑only moraine segmentation dataset, comprising 3,340 manually annotated high‑resolution images from Google Earth covering glaciated regions of Sichuan and Yunnan, China. We develop MCD‑Net, a lightweight baseline that integrates a MobileNetV2 encoder, a Convolutional Block Attention Module (CBAM), and a DeepLabV3+ decoder. Benchmarking against deeper backbones (ResNet152, Xception) shows that MCD‑Net achieves 62.3% mean Intersection over Union (mIoU) and 72.8% Dice coefficient while reducing computational cost by more than 60%. Although ridge delineation remains constrained by sub‑pixel width and spectral ambiguity, the results demonstrate that optical imagery alone can provide reliable moraine‑body segmentation. The dataset and code are publicly available at https://github.com/Lyra‑alpha/MCD‑Net, establishing a reproducible benchmark for moraine‑specific segmentation and offering a deployable baseline for high‑altitude glacial monitoring.
Authors:Matthias Bartolo, Dylan Seychell, Gabriel Hili, Matthew Montebello, Carl James Debono, Saviour Formosa, Konstantinos Makantasis
Abstract:
This paper investigates the integration of the Learning Using Privileged Information (LUPI) paradigm in object detection to exploit fine‑grained, descriptive information available during training but not at inference. We introduce a general, model‑agnostic methodology for injecting privileged information‑such as bounding box masks, saliency maps, and depth cues‑into deep learning‑based object detectors through a teacher‑student architecture. Experiments are conducted across five state‑of‑the‑art object detection models and multiple public benchmarks, including UAV‑based litter detection datasets and Pascal VOC 2012, to assess the impact on accuracy, generalization, and computational efficiency. Our results demonstrate that LUPI‑trained students consistently outperform their baseline counterparts, achieving significant boosts in detection accuracy with no increase in inference complexity or model size. Performance improvements are especially marked for medium and large objects, while ablation studies reveal that intermediate weighting of teacher guidance optimally balances learning from privileged and standard inputs. The findings affirm that the LUPI framework provides an effective and practical strategy for advancing object detection systems in both resource‑constrained and real‑world settings.
Authors:Meng Wang, Wenjing Dai, Jiawan Zhang, Xiaojie Guo
Abstract:
Although recent approaches to face normal estimation have achieved promising results, their effectiveness heavily depends on large‑scale paired data for training. This paper concentrates on relieving this requirement via developing a coarse‑to‑fine normal estimator. Concretely, our method first trains a neat model from a small dataset to produce coarse face normals that perform as guidance (called exemplars) for the following refinement. A self‑attention mechanism is employed to capture long‑range dependencies, thus remedying severe local artifacts left in estimated coarse facial normals. Then, a refinement network is customized for the sake of mapping input face images together with corresponding exemplars to fine‑grained high‑quality facial normals. Such a logical function split can significantly cut the requirement of massive paired data and computational resource. Extensive experiments and ablation studies are conducted to demonstrate the efficacy of our design and reveal its superiority over state‑of‑the‑art methods in terms of both training expense as well as estimation quality. Our code and models are open‑sourced at: https://github.com/AutoHDR/FNR2R.git.
Authors:Jingjing Wang, Qianglin Liu, Zhuo Xiao, Xinning Yao, Bo Liu, Lu Li, Lijuan Niu, Fugen Zhou
Abstract:
Thyroid cancer is the most common endocrine malignancy, and its incidence is rising globally. While ultrasound is the preferred imaging modality for detecting thyroid nodules, its diagnostic accuracy is often limited by challenges such as low image contrast and blurred nodule boundaries. To address these issues, we propose Nodule‑DETR, a novel detection transformer (DETR) architecture designed for robust thyroid nodule detection in ultrasound images. Nodule‑DETR introduces three key innovations: a Multi‑Spectral Frequency‑domain Channel Attention (MSFCA) module that leverages frequency analysis to enhance features of low‑contrast nodules; a Hierarchical Feature Fusion (HFF) module for efficient multi‑scale integration; and Multi‑Scale Deformable Attention (MSDA) to flexibly capture small and irregularly shaped nodules. We conducted extensive experiments on a clinical dataset of real‑world thyroid ultrasound images. The results demonstrate that Nodule‑DETR achieves state‑of‑the‑art performance, outperforming the baseline model by a significant margin of 0.149 in mAP@0.5:0.95. The superior accuracy of Nodule‑DETR highlights its significant potential for clinical application as an effective tool in computer‑aided thyroid diagnosis. The code of work is available at https://github.com/wjj1wjj/Nodule‑DETR.
Authors:Shuhang Chen, Yunqiu Xu, Junjie Xie, Aojun Lu, Tao Feng, Zeying Huang, Ning Zhang, Yi Sun, Yi Yang, Hangjie Yuan
Abstract:
Despite significant progress, multimodal large language models continue to struggle with visual mathematical problem solving. Some recent works recognize that visual perception is a bottleneck in visual mathematical reasoning, but their solutions are limited to improving the extraction and interpretation of visual inputs. Notably, they all ignore the key issue of whether the extracted visual cues are faithfully integrated and properly utilized in subsequent reasoning. Motivated by this, we present CogFlow, a novel cognitive‑inspired three‑stage framework that incorporates a knowledge internalization stage, explicitly simulating the hierarchical flow of human reasoning: perception\Rightarrowinternalization\Rightarrowreasoning. In line with this hierarchical flow, we holistically enhance all its stages. We devise Synergistic Visual Rewards to boost perception capabilities in parametric and semantic spaces, jointly improving visual information extraction from symbols and diagrams. To guarantee faithful integration of extracted visual cues into subsequent reasoning, we introduce a Knowledge Internalization Reward model in the internalization stage, bridging perception and reasoning. Moreover, we design a Visual‑Gated Policy Optimization algorithm to further enforce the reasoning is grounded with the visual knowledge, preventing models seeking shortcuts that appear coherent but are visually ungrounded reasoning chains. Moreover, we contribute a new dataset MathCog for model training, which contains samples with over 120K high‑quality perception‑reasoning aligned annotations. Comprehensive experiments and analysis on commonly used visual mathematical reasoning benchmarks validate the superiority of the proposed CogFlow. Project page: https://shchen233.github.io/cogflow.
Authors:Wenyu Shao, Hongbo Liu, Yunchuan Ma, Ruili Wang
Abstract:
Existing text‑driven infrared and visible image fusion approaches often rely on textual information at the sentence level, which can lead to semantic noise from redundant text and fail to fully exploit the deeper semantic value of textual information. To address these issues, we propose a novel fusion approach named Entity‑Guided Multi‑Task learning for infrared and visible image fusion (EGMT). Our approach includes three key innovative components: (i) A principled method is proposed to extract entity‑level textual information from image captions generated by large vision‑language models, eliminating semantic noise from raw text while preserving critical semantic information; (ii) A parallel multi‑task learning architecture is constructed, which integrates image fusion with a multi‑label classification task. By using entities as pseudo‑labels, the multi‑label classification task provides semantic supervision, enabling the model to achieve a deeper understanding of image content and significantly improving the quality and semantic density of the fused image; (iii) An entity‑guided cross‑modal interactive module is also developed to facilitate the fine‑grained interaction between visual and entity‑level textual features, which enhances feature representation by capturing cross‑modal dependencies at both inter‑visual and visual‑entity levels. To promote the wide application of the entity‑guided image fusion framework, we release the entity‑annotated version of four public datasets (i.e., TNO, RoadScene, M3FD, and MSRS). Extensive experiments demonstrate that EGMT achieves superior performance in preserving salient targets, texture details, and semantic consistency, compared to the state‑of‑the‑art methods. The code and dataset will be publicly available at https://github.com/wyshao‑01/EGMT.
Authors:Joongwon Chae, Lihui Luo, Yang Liu, Runming Wang, Dongmei Yu, Zeming Liang, Xi Yuan, Dayan Zhang, Zhenglin Chen, Peiwu Qin, Ilmoon Chae
Abstract:
Feature‑based anomaly detection is widely adopted in industrial inspection due to the strong representational power of large pre‑trained vision encoders. While most existing methods focus on improving within‑category anomaly scoring, practical deployments increasingly require task‑agnostic operation under continual category expansion, where the category identity is unknown at test time. In this setting, overall performance is often dominated by expert selection, namely routing an input to an appropriate normality model before any head‑specific scoring is applied. However, routing rules that compare head‑specific anomaly scores across independently constructed heads are unreliable in practice, as score distributions can differ substantially across categories in scale and tail behavior.
We propose GCR, a lightweight mixture‑of‑experts framework for stabilizing task‑agnostic continual anomaly detection through geometry‑consistent routing. GCR routes each test image directly in a shared frozen patch‑embedding space by minimizing an accumulated nearest‑prototype distance to category‑specific prototype banks, and then computes anomaly maps only within the routed expert using a standard prototype‑based scoring rule. By separating cross‑head decision making from within‑head anomaly scoring, GCR avoids cross‑head score comparability issues without requiring end‑to‑end representation learning.
Experiments on MVTec AD and VisA show that geometry‑consistent routing substantially improves routing stability and mitigates continual performance collapse, achieving near‑zero forgetting while maintaining competitive detection and localization performance. These results indicate that many failures previously attributed to representation forgetting can instead be explained by decision‑rule instability in cross‑head routing. Code is available at https://github.com/jw‑chae/GCR
Authors:Lakshay Sharma, Alex Marin
Abstract:
Self‑supervised learning (SSL) methods have become a dominant paradigm for creating general purpose models whose capabilities can be transferred to downstream supervised learning tasks. However, most such methods rely on vast amounts of pretraining data. This work introduces Subimage Overlap Prediction, a novel self‑supervised pretraining task to aid semantic segmentation in remote sensing imagery that uses significantly lesser pretraining imagery. Given an image, a sub‑image is extracted and the model is trained to produce a semantic mask of the location of the extracted sub‑image within the original image. We demonstrate that pretraining with this task results in significantly faster convergence, and equal or better performance (measured via mIoU) on downstream segmentation. This gap in convergence and performance widens when labeled training data is reduced. We show this across multiple architecture types, and with multiple downstream datasets. We also show that our method matches or exceeds performance while requiring significantly lesser pretraining data relative to other SSL methods. Code and model weights are provided at \hrefhttps://github.com/sharmalakshay93/subimage‑overlap‑predictiongithub.com/sharmalakshay93/subimage‑overlap‑prediction.
Authors:Hao Lu, Ziniu Qian, Yifu Li, Yang Zhou, Bingzheng Wei, Yan Xu
Abstract:
In this paper, we introduce a clinical diagnosis template‑based pipeline to systematically collect and structure pathological information. In collaboration with pathologists and guided by the the College of American Pathologists (CAP) Cancer Protocols, we design a Clinical Pathology Report Template (CPRT) that ensures comprehensive and standardized extraction of diagnostic elements from pathology reports. We validate the effectiveness of our pipeline on TCGA‑BRCA. First, we extract pathological features from reports using CPRT. These features are then used to build CTIS‑Align, a dataset of 80k slide‑description pairs from 804 WSIs for vision‑language alignment training, and CTIS‑Bench, a rigorously curated VQA benchmark comprising 977 WSIs and 14,879 question‑answer pairs. CTIS‑Bench emphasizes clinically grounded, closed‑ended questions (e.g., tumor grade, receptor status) that reflect real diagnostic workflows, minimize non‑visual reasoning, and require genuine slide understanding. We further propose CTIS‑QA, a Slide‑level Question Answering model, featuring a dual‑stream architecture that mimics pathologists' diagnostic approach. One stream captures global slide‑level context via clustering‑based feature aggregation, while the other focuses on salient local regions through attention‑guided patch perception module. Extensive experiments on WSI‑VQA, CTIS‑Bench, and slide‑level diagnostic tasks show that CTIS‑QA consistently outperforms existing state‑of‑the‑art models across multiple metrics. Code and data are available at https://github.com/HLSvois/CTIS‑QA.
Authors:Yanhao Wu, Haoyang Zhang, Fei He, Rui Wu, Yanhu Shan, Congpei Qiu, Liang Gao, Wei Ke, Tong Zhang
Abstract:
Practical autonomous driving requires models that generalize by reasoning through spatial‑temporal possibilities to exclude unsafe outcomes. While state‑of‑the‑art (SOTA) methods use parallel planning architectures, they fail to explicitly couple speed decisions with agent behavior along the driving path, leading to suboptimal coordination. To address this, we propose a cascaded framework that transforms longitudinal planning from an independent prediction task into a path‑conditioned reasoning process. On the model side, we introduce an anchor‑based regression design that conditions longitudinal prediction on the lateral drive path, and reformulate longitudinal planning as 1D displacement prediction along the path. This reduces geometric uncertainty and sharpens the model's focus on interaction‑driven dynamics. On the data side, we introduce a planning‑oriented data augmentation strategy that simulates rare safety‑critical events by programmatically inserting agents and relabeling longitudinal targets to enforce collision avoidance. Evaluated on the challenging Bench2Drive benchmark, our method achieves SOTA performance with a driving score of 89.07 and a success rate of 73.18%, demonstrating significantly improved coordination and safety. Further evaluation on Fail2Drive confirms strong generalization to rare edge cases where parallel formulations typically fail. Project page:https://yanhaowu.github.io/AlignDrive/.
Authors:Zhengsen Xu, Lanying Wang, Sibo Cheng, Xue Rui, Kyle Gao, Yimin Zhu, Mabel Heffring, Zack Dewis, Saeid Taleghanidoozdoozan, Megan Greenwood, Motasem Alkayid, Quinn Ledingham, Hongjie He, Jonathan Li, Lincoln Linlin Xu
Abstract:
In recent decades, the intensification of wildfire activity in western Canada has resulted in substantial socio‑economic and environmental losses. Accurate wildfire risk prediction is hindered by the intrinsic stochasticity of ignition and spread and by nonlinear interactions among fuel conditions, meteorology, climate variability, topography, and human activities, challenging the reliability and interpretability of purely data‑driven models. We propose a trustworthy data‑driven wildfire risk prediction framework based on long‑sequence, multi‑scale temporal modeling, which integrates heterogeneous drivers while explicitly quantifying predictive uncertainty and enabling process‑level interpretation. Evaluated over western Canada during the record‑breaking 2023 and 2024 fire seasons, the proposed model outperforms existing time‑series approaches, achieving an F1 score of 0.90 and a PR‑AUC of 0.98 with low computational cost. Uncertainty‑aware analysis reveals structured spatial and seasonal patterns in predictive confidence, highlighting increased uncertainty associated with ambiguous predictions and spatiotemporal decision boundaries. SHAP‑based interpretation provides mechanistic understanding of wildfire controls, showing that temperature‑related drivers dominate wildfire risk in both years, while moisture‑related constraints play a stronger role in shaping spatial and land‑cover‑specific contrasts in 2024 compared to the widespread hot and dry conditions of 2023. Data and code are available at https://github.com/SynUW/mmFire.
Authors:Jin Yao, Radowan Mahmud Redoy, Sebastian Elbaum, Matthew B. Dwyer, Zezhou Cheng
Abstract:
Detecting objects in 3D space from monocular input is crucial for applications ranging from robotics to scene understanding. Despite advanced performance in the indoor and autonomous driving domains, existing monocular 3D detection models struggle with in‑the‑wild images due to the lack of 3D in‑the‑wild datasets and the challenges of 3D annotation. We introduce LabelAny3D, an \emphanalysis‑by‑synthesis framework that reconstructs holistic 3D scenes from 2D images to efficiently produce high‑quality 3D bounding box annotations. Built on this pipeline, we present COCO3D, a new benchmark for open‑vocabulary monocular 3D detection, derived from the MS‑COCO dataset and covering a wide range of object categories absent from existing 3D datasets. Experiments show that annotations generated by LabelAny3D improve monocular 3D detection performance across multiple benchmarks, outperforming prior auto‑labeling approaches in quality. These results demonstrate the promise of foundation‑model‑driven annotation for scaling up 3D recognition in realistic, open‑world settings.
Authors:Aymen Mir, Riza Alp Guler, Jian Wang, Gerard Pons-Moll, Bing Zhou
Abstract:
We present a method for consistent lighting and shadows when animated 3D Gaussian Splatting (3DGS) avatars interact with 3DGS scenes or with dynamic objects inserted into otherwise static scenes. Our key contribution is Deep Gaussian Shadow Maps (DGSM), a modern analogue of the classical shadow mapping algorithm tailored to the volumetric 3DGS representation. Building on the classic deep shadow mapping idea, we show that 3DGS admits closed form light accumulation along light rays, enabling volumetric shadow computation without meshing. For each estimated light, we tabulate transmittance over concentric radial shells and store them in octahedral atlases, which modern GPUs can sample in real time per query to attenuate affected scene Gaussians and thus cast and receive shadows consistently. To relight moving avatars, we approximate the local environment illumination with HDRI probes represented in a spherical harmonic (SH) basis and apply a fast per Gaussian radiance transfer, avoiding explicit BRDF estimation or offline optimization. We demonstrate environment consistent lighting for avatars from AvatarX and ActorsHQ, composited into ScanNet++, DL3DV, and SuperSplat scenes, and show interactions with inserted objects. Across single and multi avatar settings, DGSM and SH relighting operate fully in the volumetric 3DGS representation, yielding coherent shadows and relighting while avoiding meshing.
Authors:Zixuan Fu, Lanqing Guo, Chong Wang, Binbin Song, Ding Liu, Bihan Wen
Abstract:
Flexible image tokenizers aim to represent an image using an ordered 1D variable‑length token sequence. This flexible tokenization is typically achieved through nested dropout, where a portion of trailing tokens is randomly truncated during training, and the image is reconstructed using the remaining preceding sequence. However, this tail‑truncation strategy inherently concentrates the image information in the early tokens, limiting the effectiveness of downstream AutoRegressive (AR) image generation as the token length increases. To overcome these limitations, we propose ReToK, a flexible tokenizer with \underlineRedundant \underlineToken Padding and Hierarchical Semantic Regularization, designed to fully exploit all tokens for enhanced latent modeling. Specifically, we introduce Redundant Token Padding to activate tail tokens more frequently, thereby alleviating information over‑concentration in the early tokens. In addition, we apply Hierarchical Semantic Regularization to align the decoding features of earlier tokens with those from a pre‑trained vision foundation model, while progressively reducing the regularization strength toward the tail to allow finer low‑level detail reconstruction. Extensive experiments demonstrate the effectiveness of ReTok: on ImageNet 256×256, our method achieves superior generation performance compared with both flexible and fixed‑length tokenizers. Code will be available at: \hrefhttps://github.com/zfu006/ReTokhttps://github.com/zfu006/ReTok
Authors:Hongbing Li, Linhui Xiao, Zihan Zhao, Qi Shen, Yixiang Huang, Bo Xiao, Zhanyu Ma
Abstract:
Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through one‑tower architectures, they still suffer from two primary limitations: (1) over‑entangled multimodal representations that exacerbate deceptive modality biases, and (2) insufficient semantic reasoning that hinders the comprehension of referential cues. In this paper, we propose BARE, a bias‑aware and reasoning‑enhanced framework for one‑tower visual grounding. BARE introduces a mechanism that preserves modality‑specific features and constructs referential semantics through three novel modules: (i) language salience modulator, (ii) visual bias correction and (iii) referential relationship enhancement, which jointly mitigate multimodal distractions and enhance referential comprehension. Extensive experimental results on five benchmarks demonstrate that BARE not only achieves state‑of‑the‑art performance but also delivers superior computational efficiency compared to existing approaches. The code is publicly accessible at https://github.com/Marloweeee/BARE.
Authors:Ziyue Zhang, Luxi Lin, Xiaolin Hu, Chao Chang, HuaiXi Wang, Yiyi Zhou, Rongrong Ji
Abstract:
Diffusion inversion is a task of recovering the noise of an image in a diffusion model, which is vital for controllable diffusion image editing. At present, diffusion inversion still remains a challenging task due to the lack of viable supervision signals. Thus, most existing methods resort to approximation‑based solutions, which however are often at the cost of performance or efficiency. To remedy these shortcomings, we propose a novel self‑supervised diffusion inversion approach in this paper, termed Deep Inversion (DeepInv). Instead of requiring ground‑truth noise annotations, we introduce a self‑supervised objective as well as a data augmentation strategy to generate high‑quality pseudo noises from real images without manual intervention. Based on these two innovative designs, DeepInv is also equipped with an iterative and multi‑scale training regime to train a parameterized inversion solver, thereby achieving the fast and accurate image‑to‑noise mapping. To the best of our knowledge, this is the first attempt of presenting a trainable solver to predict inversion noise step by step. The extensive experiments show that our DeepInv can achieve much better performance and inference speed than the compared methods, e.g., +40.435% SSIM than EasyInv and +9887.5% speed than ReNoise on COCO dataset. Moreover, our careful designs of trainable solvers can also provide insights to the community. Codes and model parameters will be released in https://github.com/potato‑kitty/DeepInv.
Authors:Zobia Batool, Diala Lteif, Vijaya B. Kolachalama, Huseyin Ozkan, Erchan Aptoula
Abstract:
Despite progress in deep learning for Alzheimer's disease (AD) diagnostics, models trained on structural magnetic resonance imaging (sMRI) often do not perform well when applied to new cohorts due to domain shifts from varying scanners, protocols and patient demographics. AD, the primary driver of dementia, manifests through progressive cognitive and neuroanatomical changes like atrophy and ventricular expansion, making robust, generalizable classification essential for real‑world use. While convolutional neural networks and transformers have advanced feature extraction via attention and fusion techniques, single‑domain generalization (SDG) remains underexplored yet critical, given the fragmented nature of AD datasets. To bridge this gap, we introduce Extended MixStyle (EM), a framework for blending higher‑order feature moments (skewness and kurtosis) to mimic diverse distributional variations. Trained on sMRI data from the National Alzheimer's Coordinating Center (NACC; n=4,647) to differentiate persons with normal cognition (NC) from those with mild cognitive impairment (MCI) or AD and tested on three unseen cohorts (total n=3,126), EM yields enhanced cross‑domain performance, improving macro‑F1 on average by 2.4 percentage points over state‑of‑the‑art SDG benchmarks, underscoring its promise for invariant, reliable AD detection in heterogeneous real‑world settings. The source code will be made available upon acceptance at https://github.com/zobia111/Extended‑Mixstyle.
Authors:Xiao Li, Zilong Liu, Yining Liu, Zhuhong Li, Na Dong, Sitian Qin, Xiaolin Hu
Abstract:
To address the scarcity of high‑quality part annotations in existing datasets, we introduce PartImageNet++ (PIN++), a dataset that provides detailed part annotations for all categories in ImageNet‑1K. With 100 annotated images per category, totaling 100K images, PIN++ represents the most comprehensive dataset covering a diverse range of object categories. Leveraging PIN++, we propose a Multi‑scale Part‑supervised recognition Model (MPM) for robust classification on ImageNet‑1K. We first trained a part segmentation network using PIN++ and used it to generate pseudo part labels for the remaining unannotated images. MPM then integrated a conventional recognition architecture with auxiliary bypass layers, jointly supervised by both pseudo part labels and the original part annotations. Furthermore, we conducted extensive experiments on PIN++, including part segmentation, object segmentation, and few‑shot learning, exploring various ways to leverage part annotations in downstream tasks. Experimental results demonstrated that our approach not only enhanced part‑based models for robust object recognition but also established strong baselines for multiple downstream tasks, highlighting the potential of part annotations in improving model performance. The dataset and the code are available at https://github.com/LixiaoTHU/PartImageNetPP.
Authors:Weiqi Yu, Yiyang Yao, Lin He, Jianming Lv
Abstract:
Neural Radiance Fields (NeRF) achieve remarkable performance in dense multi‑view scenarios, but their reconstruction quality degrades significantly under sparse inputs due to geometric artifacts. Existing methods utilize global depth regularization to mitigate artifacts, leading to the loss of geometric boundary details. To address this problem, we propose EdgeNeRF, an edge‑guided sparse‑view 3D reconstruction algorithm. Our method leverages the prior that abrupt changes in depth and normals generate edges. Specifically, we first extract edges from input images, then apply depth and normal regularization constraints to non‑edge regions, enhancing geometric consistency while preserving high‑frequency details at boundaries. Experiments on LLFF and DTU datasets demonstrate EdgeNeRF's superior performance, particularly in retaining sharp geometric boundaries and suppressing artifacts. Additionally, the proposed edge‑guided depth regularization module can be seamlessly integrated into other methods in a plug‑and‑play manner, significantly improving their performance without substantially increasing training time. Code is available at https://github.com/skyhigh404/edgenerf.
Authors:Xu Guo, Fulong Ye, Xinghui Li, Pengqi Tu, Pengze Zhang, Qichao Sun, Songtao Zhao, Xiangwang Hou, Qian He
Abstract:
Video Face Swapping (VFS) requires seamlessly injecting a source identity into a target video while meticulously preserving the original pose, expression, lighting, background, and dynamic information. Existing methods struggle to maintain identity similarity and attribute preservation while preserving temporal consistency. To address the challenge, we propose a comprehensive framework to seamlessly transfer the superiority of Image Face Swapping (IFS) to the video domain. We first introduce a novel data pipeline SyncID‑Pipe that pre‑trains an Identity‑Anchored Video Synthesizer and combines it with IFS models to construct bidirectional ID quadruplets for explicit supervision. Building upon paired data, we propose the first Diffusion Transformer‑based framework DreamID‑V, employing a core Modality‑Aware Conditioning module to discriminatively inject multi‑model conditions. Meanwhile, we propose a Synthetic‑to‑Real Curriculum mechanism and an Identity‑Coherence Reinforcement Learning strategy to enhance visual realism and identity consistency under challenging scenarios. To address the issue of limited benchmarks, we introduce IDBench‑V, a comprehensive benchmark encompassing diverse scenes. Extensive experiments demonstrate DreamID‑V outperforms state‑of‑the‑art methods and further exhibits exceptional versatility, which can be seamlessly adapted to various swap‑related tasks.
Authors:Yue Zhou, Ran Ding, Xue Yang, Xue Jiang, Xingzhao Liu
Abstract:
Despite notable advancements in remote sensing vision‑language models (VLMs), existing models often struggle with spatial understanding, limiting their effectiveness in real‑world applications. To push the boundaries of VLMs in remote sensing, we specifically address vehicle imagery captured by drones and introduce a spatially‑aware dataset AirSpatial, which comprises over 206K instructions and introduces two novel tasks: Spatial Grounding and Spatial Question Answering. It is also the first remote sensing grounding dataset to provide 3DBB. To effectively leverage existing image understanding of VLMs to spatial domains, we adopt a two‑stage training strategy comprising Image Understanding Pre‑training and Spatial Understanding Fine‑tuning. Utilizing this trained spatially‑aware VLM, we develop an aerial agent, AirSpatialBot, which is capable of fine‑grained vehicle attribute recognition and retrieval. By dynamically integrating task planning, image understanding, spatial understanding, and task execution capabilities, AirSpatialBot adapts to diverse query requirements. Experimental results validate the effectiveness of our approach, revealing the spatial limitations of existing VLMs while providing valuable insights. The model, code, and datasets will be released at https://github.com/VisionXLab/AirSpatialBot
Authors:Habiba Kausar, Saeed Anwar, Omar Jamal Hammad, Abdul Bais
Abstract:
Face super‑resolution aims to recover high‑quality facial images from severely degraded low‑resolution inputs, but remains challenging due to the loss of fine structural details and identity‑specific features. This work introduces SwinIFS, a landmark‑guided super‑resolution framework that integrates structural priors with hierarchical attention mechanisms to achieve identity‑preserving reconstruction at both moderate and extreme upscaling factors. The method incorporates dense Gaussian heatmaps of key facial landmarks into the input representation, enabling the network to focus on semantically important facial regions from the earliest stages of processing. A compact Swin Transformer backbone is employed to capture long‑range contextual information while preserving local geometry, allowing the model to restore subtle facial textures and maintain global structural consistency. Extensive experiments on the CelebA benchmark demonstrate that SwinIFS achieves superior perceptual quality, sharper reconstructions, and improved identity retention; it consistently produces more photorealistic results and exhibits strong performance even under 8x magnification, where most methods fail to recover meaningful structure. SwinIFS also provides an advantageous balance between reconstruction accuracy and computational efficiency, making it suitable for real‑world applications in facial enhancement, surveillance, and digital restoration. Our code, model weights, and results are available at https://github.com/Habiba123‑stack/SwinIFS.
Authors:Xiaobao Wei, Zhangjie Ye, Yuxiang Gu, Zunjie Zhu, Yunfei Guo, Yingying Shen, Shan Zhao, Ming Lu, Haiyang Sun, Bing Wang, Guang Chen, Rongfeng Lu, Hangjun Ye
Abstract:
Parking is a critical task for autonomous driving systems (ADS), with unique challenges in crowded parking slots and GPS‑denied environments. However, existing works focus on 2D parking slot perception, mapping, and localization, 3D reconstruction remains underexplored, which is crucial for capturing complex spatial geometry in parking scenarios. Naively improving the visual quality of reconstructed parking scenes does not directly benefit autonomous parking, as the key entry point for parking is the slots perception module. To address these limitations, we curate the first benchmark named ParkRecon3D, specifically designed for parking scene reconstruction. It includes sensor data from four surround‑view fisheye cameras with calibrated extrinsics and dense parking slot annotations. We then propose ParkGaussian, the first framework that integrates 3D Gaussian Splatting (3DGS) for parking scene reconstruction. To further improve the alignment between reconstruction and downstream parking slot detection, we introduce a slot‑aware reconstruction strategy that leverages existing parking perception methods to enhance the synthesis quality of slot regions. Experiments on ParkRecon3D demonstrate that ParkGaussian achieves state‑of‑the‑art reconstruction quality and better preserves perception consistency for downstream tasks. The code and dataset will be released at: https://github.com/wm‑research/ParkGaussian
Authors:Md. Sadman Haque, Zobaer Ibn Razzaque, Robiul Awoul Robin, Fahim Hafiz, Riasat Azim
Abstract:
Public bus transport systems in developing countries often suffer from a lack of real‑time location updates and for users, making commuting inconvenient and unreliable for passengers. Furthermore, stopping at undesired locations rather than designated bus stops creates safety risks and contributes to roadblocks, often causing traffic congestion. Additionally, issues such as blind spots, along with a lack of following traffic laws, increase the chances of accidents. In this work, we address these challenges by proposing a smart public bus system along with intelligent bus stops that enhance safety, efficiency, and sustainability. Our approach includes a deep learning‑based blind‑spot warning system to help drivers avoid accidents with automated bus‑stop detection to accurately identify bus stops, improving transit efficiency. We also introduce IoT‑based solar‑powered smart bus stops that show real‑time passenger counts, along with an RFID‑based card system to track where passengers board and exit. A smart door system ensures safer and more organised boarding, while real‑time bus tracking keeps passengers informed. To connect all these features, we use an HTTP‑based server for seamless communication between the interconnected network systems. Our proposed system demonstrated approximately 99% efficiency in real‑time blind spot detection while stopping precisely at the bus stops. Furthermore, the server showed real‑time location updates both to the users and at the bus stops, enhancing commuting efficiency. The proposed energy‑efficient bus stop demonstrated 12.71kWh energy saving, promoting sustainable architecture. Full implementation and source code are available at: https://github.com/sadman‑adib/MoveMe‑IoT
Authors:Bac Nguyen, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai, Toshimitsu Uesaka, Stefano Ermon, Yuki Mitsufuji
Abstract:
Slot Attention (SA) with pretrained diffusion models has recently shown promise for object‑centric learning (OCL), but suffers from slot entanglement and weak alignment between object slots and image content. We propose Contrastive Object‑centric Diffusion Alignment (CODA), a simple extension that (i) employs register slots to absorb residual attention and reduce interference between object slots, and (ii) applies a contrastive alignment loss to explicitly encourage slot‑image correspondence. The resulting training objective serves as a tractable surrogate for maximizing mutual information (MI) between slots and inputs, strengthening slot representation quality. On both synthetic (MOVi‑C/E) and real‑world datasets (VOC, COCO), CODA improves object discovery (e.g., +6.1% FG‑ARI on COCO), property prediction, and compositional image generation over strong baselines. Register slots add negligible overhead, keeping CODA efficient and scalable. These results indicate potential applications of CODA as an effective framework for robust OCL in complex, real‑world scenes. Code and pretrained models are available at https://github.com/sony/coda.
Authors:Mengfei Li, Peng Li, Zheng Zhang, Jiahao Lu, Chengfeng Zhao, Wei Xue, Qifeng Liu, Sida Peng, Wenxiao Zhang, Wenhan Luo, Yuan Liu, Yike Guo
Abstract:
We present UniSH, a unified, feed‑forward framework for joint metric‑scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large‑scale, annotated real‑world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim‑to‑real domain gap, leading to poor generalization, low‑fidelity human geometry, and poor alignment on in‑the‑wild videos. To address this, we propose an innovative training paradigm that effectively leverages unlabeled in‑the‑wild data. Our framework bridges strong, disparate priors from scene reconstruction and HMR, and is trained with two core components: (1) a robust distillation strategy to refine human surface details by distilling high‑frequency details from an expert depth model, and (2) a two‑stage supervision scheme, which first learns coarse localization on synthetic data, then fine‑tunes on real data by directly optimizing the geometric correspondence between the SMPL mesh and the human point cloud. This approach enables our feed‑forward model to jointly recover high‑fidelity scene geometry, human point clouds, camera parameters, and coherent, metric‑scale SMPL bodies, all in a single forward pass. Extensive experiments demonstrate that our model achieves state‑of‑the‑art performance on human‑centric scene reconstruction and delivers highly competitive results on global human motion estimation, comparing favorably against both optimization‑based frameworks and HMR‑only methods. Project page: https://murphylmf.github.io/UniSH/
Authors:Zunhai Su, Weihao Ye, Hansen Feng, Keyu Fan, Jing Zhang, Dahai Yu, Zhengwu Liu, Ngai Wong
Abstract:
Learning‑based 3D visual geometry models have benefited substantially from large‑scale transformers. Among these, StreamVGGT leverages frame‑wise causal attention for strong streaming reconstruction, but suffers from unbounded KV cache growth, leading to escalating memory consumption and inference latency as input frames accumulate. We propose XStreamVGGT, a tuning‑free approach that systematically compresses the KV cache through joint pruning and quantization, enabling extremely memory‑efficient streaming inference. Specifically, redundant KVs originating from multi‑view inputs are pruned through efficient token importance identification, enabling a fixed memory budget. Leveraging the unique distribution of KV tensors, we incorporate KV quantization to further reduce memory consumption. Extensive evaluations show that XStreamVGGT achieves mostly negligible performance degradation while substantially reducing memory usage by 4.42× and accelerating inference by 5.48×, enabling scalable and practical streaming 3D applications. The code is available at https://github.com/ywh187/XStreamVGGT/.
Authors:Zhang Chen, Shuai Wan, Yuezhe Zhang, Siyu Ren, Fuzheng Yang, Junhui Hou
Abstract:
The unstructured and irregular nature of points poses a significant challenge for accurate point cloud quality assessment (PCQA), particularly in establishing accurate perceptual feature correspondence. To tackle this, we propose the Multi‑scale Implicit Structural Similarity Measurement (MS‑ISSM). Unlike traditional point‑to‑point matching, MS‑ISSM utilizes radial basis function (RBF) to represent local features continuously, transforming distortion measurement into a comparison of implicit function coefficients. This approach effectively circumvents matching errors inherent in irregular data. Additionally, we propose a ResGrouped‑MLP quality assessment network, which robustly maps multi‑scale feature differences to perceptual scores. The network architecture departs from traditional flat multi‑layer perceptron (MLP) by adopting a grouped encoding strategy integrated with residual blocks and channel‑wise attention mechanisms. This hierarchical design allows the model to preserve the distinct physical semantics of luma, chroma, and geometry while adaptively focusing on the most salient distortion features across High, Medium, and Low scales. Experimental results on multiple benchmarks demonstrate that MS‑ISSM outperforms state‑of‑the‑art metrics in both reliability and generalization. The source code is available at: https://github.com/ZhangChen2022/MS‑ISSM.
Authors:Hao Lu, Xuhui Zhu, Wenjing Zhang, Yanan Li, Xiang Bai
Abstract:
Video Individual Counting (VIC) is a recently introduced task aiming to estimate pedestrian flux from a video. It extends Video Crowd Counting (VCC) beyond the per‑frame pedestrian count. In contrast to VCC that learns to count pedestrians across frames, VIC must identify co‑existent pedestrians between frames, which turns out to be a correspondence problem. Existing VIC approaches, however, can underperform in congested scenes such as metro commuting. To address this, we build WuhanMetroCrowd, one of the first VIC datasets that characterize crowded, dynamic pedestrian flows. It features sparse‑to‑dense density levels, short‑to‑long video clips, slow‑to‑fast flow variations, front‑to‑back appearance changes, and light‑to‑heavy occlusions. To better adapt VIC approaches to crowds, we rethink the nature of VIC and recognize two informative priors: i) the social grouping prior that indicates pedestrians tend to gather in groups and ii) the spatial‑temporal displacement prior that informs an individual cannot teleport physically. The former inspires us to relax the standard one‑to‑one (O2O) matching used by VIC to one‑to‑many (O2M) matching, implemented by an implicit context generator and a O2M matcher; the latter facilitates the design of a displacement prior injector, which strengthens not only O2M matching but also feature extraction and model training. These designs jointly form a novel and strong VIC baseline OMAN++. Extensive experiments show that OMAN++ not only outperforms state‑of‑the‑art VIC baselines on the standard SenseCrowd, CroHD, and MovingDroneCrowd benchmarks, but also indicates a clear advantage in crowded scenes, with a 38.12% error reduction on our WuhanMetroCrowd dataset. Code, data, and pretrained models are available at https://github.com/tiny‑smart/OMAN.
Authors:Tianheng Cheng, Xinggang Wang, Junchao Liao, Wenyu Liu
Abstract:
Semantic segmentation is a fundamental problem in computer vision and it requires high‑resolution feature maps for dense prediction. Current coordinate‑guided low‑resolution feature interpolation methods, e.g., bilinear interpolation, produce coarse high‑resolution features which suffer from feature misalignment and insufficient context information. Moreover, enriching semantics to high‑resolution features requires a high computation burden, so that it is challenging to meet the requirement of lowlatency inference. We propose a novel Guided Attentive Interpolation (GAI) method to adaptively interpolate fine‑grained high‑resolution features with semantic features to tackle these issues. Guided Attentive Interpolation determines both spatial and semantic relations of pixels from features of different resolutions and then leverages these relations to interpolate high‑resolution features with rich semantics. GAI can be integrated with any deep convolutional network for efficient semantic segmentation. In experiments, the GAI‑based semantic segmentation networks, i.e., GAIN, can achieve78.8 mIoU with 22.3 FPS on Cityscapes and 80.6 mIoU with 64.5 on CamVid using an NVIDIA 1080Ti GPU, which are the new state‑of‑the‑art results of low‑latency semantic segmentation. Code and models are available at: https://github.com/hustvl/simpleseg.
Authors:Xingchen Li, Junzhe Zhang, Junqi Shi, Ming Lu, Zhan Ma
Abstract:
While one‑step diffusion models have recently excelled in perceptual image compression, their application to video remains limited. Prior efforts typically rely on pretrained 2D autoencoders that generate per‑frame latent representations independently, thereby neglecting temporal dependencies. We present YODA‑‑Yet Another One‑step Diffusion‑based Video Compressor‑‑which embeds multiscale features from temporal references for both latent generation and latent coding to better exploit spatial‑temporal correlations for more compact representation, and employs a linear Diffusion Transformer (DiT) for efficient one‑step denoising. YODA achieves state‑of‑the‑art perceptual performance, consistently outperforming traditional and deep‑learning baselines on LPIPS, DISTS, FID, and KID. Source code will be publicly available at https://github.com/NJUVISION/YODA.
Authors:Abhinav Attri, Rajeev Ranjan Dwivedi, Samiran Das, Vinod Kumar Kurmi
Abstract:
We present HAQAGen, a unified generative model for resolution‑invariant NIR‑to‑RGB colorization that balances chromatic realism with structural fidelity. The proposed model introduces (i) a combined loss term aligning the global color statistics through differentiable histogram matching, perceptual image quality measure, and feature based similarity to preserve texture information, (ii) local hue‑saturation priors injected via Spatially Adaptive Denormalization (SPADE) to stabilize chromatic reconstruction, and (iii) texture‑aware supervision within a Mamba backbone to preserve fine details. We introduce an adaptive‑resolution inference engine that further enables high‑resolution translation without sacrificing quality. Our proposed NIR‑to‑RGB translation model simultaneously enforces global color statistics and local chromatic consistency, while scaling to native resolutions without compromising texture fidelity or generalization. Extensive evaluations on FANVID, OMSIV, VCIP2020, and RGB2NIR using different evaluation metrics demonstrate consistent improvements over state‑of‑the‑art baseline methods. HAQAGen produces images with sharper textures, natural colors, attaining significant gains as per perceptual metrics. These results position HAQAGen as a scalable and effective solution for NIR‑to‑RGB translation across diverse imaging scenarios. Project Page: https://rajeev‑dw9.github.io/HAQAGen/
Authors:Hyeonjeong Ha, Jinjin Ge, Bo Feng, Kaixin Ma, Gargi Chakraborty
Abstract:
Multimodal large language models (MLLMs) have achieved impressive progress in vision‑language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative understanding requires grounding who is doing what, when, and where, maintaining coherent entity representations across dynamic visual and temporal contexts. We introduce NarrativeTrack, the first benchmark to evaluate narrative understanding in MLLMs through fine‑grained entity‑centric reasoning. Unlike existing benchmarks limited to short clips or coarse scene‑level semantics, we decompose videos into constituent entities and examine their continuity via a Compositional Reasoning Progression (CRP), a structured evaluation framework that progressively increases narrative complexity across three dimensions: entity existence, entity changes, and entity ambiguity. CRP challenges models to advance from temporal persistence to contextual evolution and fine‑grained perceptual reasoning. A fully automated entity‑centric pipeline enables scalable extraction of temporally grounded entity representations, providing the foundation for CRP. Evaluations of state‑of‑the‑art MLLMs reveal that models fail to robustly track entities across visual transitions and temporal dynamics, often hallucinating identity under context shifts. Open‑source general‑purpose MLLMs exhibit strong perceptual grounding but weak temporal coherence, while video‑specific MLLMs capture temporal context yet hallucinate entity's contexts. These findings uncover a fundamental trade‑off between perceptual grounding and temporal reasoning, indicating that narrative understanding emerges only from their integration. NarrativeTrack provides the first systematic framework to diagnose and advance temporally grounded narrative comprehension in MLLMs.
Authors:Jianan Li, Wangcai Zhao, Tingfa Xu
Abstract:
Hyperspectral imaging (HSI) is essential across various disciplines for its capacity to capture rich spectral information. However, efficiently reconstructing hyperspectral images from compressive sensing measurements presents significant challenges. To tackle these, we adopt a divide‑and‑conquer strategy that capitalizes on the unique spectral and spatial characteristics of hyperspectral images. We introduce the Lightweight Separate Spectral Transformer (LSST), an innovative architecture tailored for efficient hyperspectral image reconstruction. This architecture consists of Separate Spectral Transformer Blocks (SSTB) for modeling spectral relationships and Lightweight Spatial Convolution Blocks (LSCB) for spatial processing. The SSTB employs Grouped Spectral Self‑attention and a Spectrum Shuffle operation to effectively manage both local and non‑local spectral relationships. Simultaneously, the LSCB utilizes depth‑wise separable convolutions and strategic ordering to enhance spatial information processing. Furthermore, we implement the Focal Spectrum Loss, a novel loss weighting mechanism that dynamically adjusts during training to improve reconstruction across spectrally complex bands. Extensive testing demonstrates that our LSST achieves superior performance while requiring fewer FLOPs and parameters, underscoring its efficiency and effectiveness. The source code is available at: https://github.com/wcz1124/LSST.
Authors:Tien-Huy Nguyen, Huu-Loc Tran, Thanh Duc Ngo
Abstract:
Vision Language Models (VLMs) have rapidly advanced and show strong promise for text‑based person search (TBPS), a task that requires capturing fine‑grained relationships between images and text to distinguish individuals. Previous methods address these challenges through local alignment, yet they are often prone to shortcut learning and spurious correlations, yielding misalignment. Moreover, injecting prior knowledge can distort intra‑modality structure. Motivated by our finding that encoder attention surfaces spatially precise evidence from the earliest training epochs, and to alleviate these issues, we introduceITSELF, an attention‑guided framework for implicit local alignment. At its core, Guided Representation with Attentive Bank (GRAB) converts the model's own attention into an Attentive Bank of high‑saliency tokens and applies local objectives on this bank, learning fine‑grained correspondences without extra supervision. To make the selection reliable and non‑redundant, we introduce Multi‑Layer Attention for Robust Selection (MARS), which aggregates attention across layers and performs diversity‑aware top‑k selection; and Adaptive Token Scheduler (ATS), which schedules the retention budget from coarse to fine over training, preserving context early while progressively focusing on discriminative details. Extensive experiments on three widely used TBPS benchmarks showstate‑of‑the‑art performance and strong cross‑dataset generalization, confirming the effectiveness and robustness of our approach without additional prior supervision. Our project is publicly available at https://trhuuloc.github.io/itself
Authors:Shiao Wang, Xiao Wang, Haonan Zhao, Jiarui Xu, Bo Jiang, Lin Zhu, Xin Zhao, Yonghong Tian, Jin Tang
Abstract:
Existing RGB‑Event visual object tracking approaches primarily rely on conventional feature‑level fusion, failing to fully exploit the unique advantages of event cameras. In particular, the high dynamic range and motion‑sensitive nature of event cameras are often overlooked, while low‑information regions are processed uniformly, leading to unnecessary computational overhead for the backbone network. To address these issues, we propose a novel tracking framework that performs early fusion in the frequency domain, enabling effective aggregation of high‑frequency information from the event modality. Specifically, RGB and event modalities are transformed from the spatial domain to the frequency domain via the Fast Fourier Transform, with their amplitude and phase components decoupled. High‑frequency event information is selectively fused into RGB modality through amplitude and phase attention, enhancing feature representation while substantially reducing backbone computation. In addition, a motion‑guided spatial sparsification module leverages the motion‑sensitive nature of event cameras to capture the relationship between target motion cues and spatial probability distribution, filtering out low‑information regions and enhancing target‑relevant features. Finally, a sparse set of target‑relevant features is fed into the backbone network for learning, and the tracking head predicts the final target position. Extensive experiments on three widely used RGB‑Event tracking benchmark datasets, including FE108, FELT, and COESOT, demonstrate the high performance and efficiency of our method. The source code of this paper will be released on https://github.com/Event‑AHU/OpenEvTracking
Authors:Zihan Li, Dandan Shan, Yunxiang Li, Paul E. Kinahan, Qingqi Hong
Abstract:
Medical image segmentation faces critical challenges in semi‑supervised learning scenarios due to severe annotation scarcity requiring expert radiological knowledge, significant inter‑annotator variability across different viewpoints and expertise levels, and inadequate multi‑scale feature integration for precise boundary delineation in complex anatomical structures. Existing semi‑supervised methods demonstrate substantial performance degradation compared to fully supervised approaches, particularly in small target segmentation and boundary refinement tasks. To address these fundamental challenges, we propose SASNet (Scale‑aware Adaptive Supervised Network), a dual‑branch architecture that leverages both low‑level and high‑level feature representations through novel scale‑aware adaptive reweight mechanisms. Our approach introduces three key methodological innovations, including the Scale‑aware Adaptive Reweight strategy that dynamically weights pixel‑wise predictions using temporal confidence accumulation, the View Variance Enhancement mechanism employing 3D Fourier domain transformations to simulate annotation variability, and segmentation‑regression consistency learning through signed distance map algorithms for enhanced boundary precision. These innovations collectively address the core limitations of existing semi‑supervised approaches by integrating spatial, temporal, and geometric consistency principles within a unified optimization framework. Comprehensive evaluation across LA, Pancreas‑CT, and BraTS datasets demonstrates that SASNet achieves superior performance with limited labeled data, surpassing state‑of‑the‑art semi‑supervised methods while approaching fully supervised performance levels. The source code for SASNet is available at https://github.com/HUANGLIZI/SASNet.
Authors:Yue Zhou, Jue Chen, Zilun Zhang, Penghui Huang, Ran Ding, Zhentao Zou, PengFei Gao, Yuchen Wei, Ke Li, Xue Yang, Xue Jiang, Hongxin Yang, Jonathan Li
Abstract:
Remote sensing (RS) large vision‑language models (LVLMs) have shown strong promise across visual grounding (VG) tasks. However, existing RS VG datasets predominantly rely on explicit referring expressions‑such as relative position, relative size, and color cues‑thereby constraining performance on implicit VG tasks that require scenario‑specific domain knowledge. This article introduces DVGBench, a high‑quality implicit VG benchmark for drones, covering six major application scenarios: traffic, disaster, security, sport, social activity, and productive activity. Each object provides both explicit and implicit queries. Based on the dataset, we design DroneVG‑R1, an LVLM that integrates the novel Implicit‑to‑Explicit Chain‑of‑Thought (I2E‑CoT) within a reinforcement learning paradigm. This enables the model to take advantage of scene‑specific expertise, converting implicit references into explicit ones and thus reducing grounding difficulty. Finally, an evaluation of mainstream models on both explicit and implicit VG tasks reveals substantial limitations in their reasoning capabilities. These findings provide actionable insights for advancing the reasoning capacity of LVLMs for drone‑based agents. The code and datasets will be released at https://github.com/zytx121/DVGBench
Authors:Julian D. Santamaria, Claudia Isaza, Jhony H. Giraldo
Abstract:
Wildlife monitoring is crucial for studying biodiversity loss and climate change. Camera trap images provide a non‑intrusive method for analyzing animal populations and identifying ecological patterns over time. However, manual analysis is time‑consuming and resource‑intensive. Deep learning, particularly foundation models, has been applied to automate wildlife identification, achieving strong performance when tested on data from the same geographical locations as their training sets. Yet, despite their promise, these models struggle to generalize to new geographical areas, leading to significant performance drops. For example, training an advanced vision‑language model, such as CLIP with an adapter, on an African dataset achieves an accuracy of 84.77%. However, this performance drops significantly to 16.17% when the model is tested on an American dataset. This limitation partly arises because existing models rely predominantly on image‑based representations, making them sensitive to geographical data distribution shifts, such as variation in background, lighting, and environmental conditions. To address this, we introduce WildIng, a Wildlife image Invariant representation model for geographical domain shift. WildIng integrates text descriptions with image features, creating a more robust representation to geographical domain shifts. By leveraging textual descriptions, our approach captures consistent semantic information, such as detailed descriptions of the appearance of the species, improving generalization across different geographical locations. Experiments show that WildIng enhances the accuracy of foundation models such as BioCLIP by 30% under geographical domain shift conditions. We evaluate WildIng on two datasets collected from different regions, namely America and Africa. The code and models are publicly available at https://github.com/Julian075/CATALOG/tree/WildIng.
Authors:Lin Xi, Yingliang Ma, Xiahai Zhuang
Abstract:
We introduce a novel FSVOS model that employs a local matching strategy to restrict the search space to the most relevant neighboring pixels. Rather than relying on inefficient standard im2col‑like implementations (e.g., spatial convolutions, depthwise convolutions and feature‑shifting mechanisms) or hardware‑specific CUDA kernels (e.g., deformable and neighborhood attention), which often suffer from limited portability across non‑CUDA devices, we reorganize the local sampling process through a direction‑based sampling perspective. Specifically, we implement a non‑parametric sampling mechanism that enables dynamically varying sampling regions. This approach provides the flexibility to adapt to diverse spatial structures without the computational costs of parametric layers and the need for model retraining. To further enhance feature coherence across frames, we design a supervised spatio‑temporal contrastive learning scheme that enforces consistency in feature representations. In addition, we introduce a publicly available benchmark dataset for multi‑object segmentation in X‑ray angiography videos (MOSXAV), featuring detailed, manually labeled segmentation ground truth. Extensive experiments on the CADICA, XACV, and MOSXAV datasets show that our proposed FSVOS method outperforms current state‑of‑the‑art video segmentation methods in terms of segmentation accuracy and generalization capability (i.e., seen and unseen categories). This work offers enhanced flexibility and potential for a wide range of clinical applications. Code is available at https://github.com/xilin‑x/XRAVOS
Authors:Megha Mariam K. M, Aditya Arun, Zakaria Laskar, C. V. Jawahar
Abstract:
Generative AI models, particularly Text‑to‑Video (T2V) systems, offer a promising avenue for transforming science education by automating the creation of engaging and intuitive visual explanations. In this work, we take a first step toward evaluating their potential in physics education by introducing a dedicated benchmark for explanatory video generation. The benchmark is designed to assess how well T2V models can convey core physics concepts through visual illustrations. Each physics concept in our benchmark is decomposed into granular teaching points, with each point accompanied by a carefully crafted prompt intended for visual explanation of the teaching point. T2V models are evaluated on their ability to generate accurate videos in response to these prompts. Our aim is to systematically explore the feasibility of using T2V models to generate high‑quality, curriculum‑aligned educational content‑paving the way toward scalable, accessible, and personalized learning experiences powered by AI. Our evaluation reveals that current models produce visually coherent videos with smooth motion and minimal flickering, yet their conceptual accuracy is less reliable. Performance in areas such as mechanics, fluids, and optics is encouraging, but models struggle with electromagnetism and thermodynamics, where abstract interactions are harder to depict. These findings underscore the gap between visual quality and conceptual correctness in educational video generation. We hope this benchmark helps the community close that gap and move toward T2V systems that can deliver accurate, curriculum‑aligned physics content at scale. The benchmark and accompanying codebase are publicly available at https://github.com/meghamariamkm/PhyEduVideo.
Authors:Le-Anh Tran, Chung Nguyen Tran, Nhan Cach Dang, Anh Le Van Quoc, Jordi Carrabina, David Castells-Rufas, Minh Son Nguyen
Abstract:
Semantic segmentation is crucial for medical image analysis, enabling precise disease diagnosis and treatment planning. However, many advanced models employ complex architectures, limiting their use in resource‑constrained clinical settings. This paper proposes MFEnNet, an efficient medical image segmentation framework that incorporates MetaFormer in the encoding phase of the U‑Net backbone. MetaFormer, an architectural abstraction of vision transformers, provides a versatile alternative to convolutional neural networks by transforming tokenized image patches into sequences for global context modeling. To mitigate the substantial computational cost associated with self‑attention, the proposed framework replaces conventional transformer modules with pooling transformer blocks, thereby achieving effective global feature aggregation at reduced complexity. In addition, Swish activation is used to achieve smoother gradients and faster convergence, while spatial pyramid pooling is incorporated at the bottleneck to improve multi‑scale feature extraction. Comprehensive experiments on different medical segmentation benchmarks demonstrate that the proposed MFEnNet approach attains competitive accuracy while significantly lowering computational cost compared to state‑of‑the‑art models. The source code for this work is available at https://github.com/tranleanh/mfennet.
Authors:Subhankar Mishra
Abstract:
3D Gaussian Splatting produces high‑quality scene reconstructions but generates hundreds of thousands of spurious Gaussians (floaters) scattered throughout the environment. These artifacts obscure objects of interest and inflate model sizes, hindering deployment in bandwidth‑constrained applications. We present Clean‑GS, a method for removing background clutter and floaters from 3DGS reconstructions using sparse semantic masks. Our approach combines whitelist‑based spatial filtering with color‑guided validation and outlier removal to achieve 60‑80% model compression while preserving object quality. Unlike existing 3DGS pruning methods that rely on global importance metrics, Clean‑GS uses semantic information from as few as 3 segmentation masks (1% of views) to identify and remove Gaussians not belonging to the target object. Our multi‑stage approach consisting of (1) whitelist filtering via projection to masked regions, (2) depth‑buffered color validation, and (3) neighbor‑based outlier removal isolates monuments and objects from complex outdoor scenes. Experiments on Tanks and Temples show that Clean‑GS reduces file sizes from 125MB to 47MB while maintaining rendering quality, making 3DGS models practical for web deployment and AR/VR applications. Our code is available at https://github.com/smlab‑niser/clean‑gs
Authors:Jiewen Chan, Zhenjun Zhao, Yu-Lun Liu
Abstract:
Reconstructing dynamic 3D scenes from monocular videos requires simultaneously capturing high‑frequency appearance details and temporally continuous motion. Existing methods using single Gaussian primitives are limited by their low‑pass filtering nature, while standard Gabor functions introduce energy instability. Moreover, lack of temporal continuity constraints often leads to motion artifacts during interpolation. We propose AdaGaR, a unified framework addressing both frequency adaptivity and temporal continuity in explicit dynamic scene modeling. We introduce Adaptive Gabor Representation, extending Gaussians through learnable frequency weights and adaptive energy compensation to balance detail capture and stability. For temporal continuity, we employ Cubic Hermite Splines with Temporal Curvature Regularization to ensure smooth motion evolution. An Adaptive Initialization mechanism combining depth estimation, point tracking, and foreground masks establishes stable point cloud distributions in early training. Experiments on Tap‑Vid DAVIS demonstrate state‑of‑the‑art performance (PSNR 35.49, SSIM 0.9433, LPIPS 0.0723) and strong generalization across frame interpolation, depth consistency, video editing, and stereo view synthesis. Project page: https://jiewenchan.github.io/AdaGaR/
Authors:Wei-Tse Cheng, Yen-Jen Chiou, Yuan-Fu Yang
Abstract:
We introduce RGS‑SLAM, a robust Gaussian‑splatting SLAM framework that replaces the residual‑driven densification stage of GS‑SLAM with a training‑free correspondence‑to‑Gaussian initialization. Instead of progressively adding Gaussians as residuals reveal missing geometry, RGS‑SLAM performs a one‑shot triangulation of dense multi‑view correspondences derived from DINOv3 descriptors refined through a confidence‑aware inlier classifier, generating a well‑distributed and structure‑aware Gaussian seed prior to optimization. This initialization stabilizes early mapping and accelerates convergence by roughly 20%, yielding higher rendering fidelity in texture‑rich and cluttered scenes while remaining fully compatible with existing GS‑SLAM pipelines. Evaluated on the TUM RGB‑D and Replica datasets, RGS‑SLAM achieves competitive or superior localization and reconstruction accuracy compared with state‑of‑the‑art Gaussian and point‑based SLAM systems, sustaining real‑time mapping performance at up to 925 FPS. Additional details and resources are available at this URL: https://breeze1124.github.io/rgs‑slam‑project‑page/
Authors:Melonie de Almeida, Daniela Ivanova, Tong Shi, John H. Williamson, Paul Henderson
Abstract:
Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence and 3D consistency in single‑image‑conditioned video generation. However, these methods often lack robust user controllability, such as modifying the camera path, limiting their applicability in real‑world applications. Most existing camera‑controlled image‑to‑video models struggle with accurately modeling camera motion, maintaining temporal consistency, and preserving geometric integrity. Leveraging explicit intermediate 3D representations offers a promising solution by enabling coherent video generation aligned with a given camera trajectory. Although these methods often use 3D point clouds to render scenes and introduce object motion in a later stage, this two‑step process still falls short in achieving full temporal consistency, despite allowing precise control over camera movement. We propose a novel framework that constructs a 3D Gaussian scene representation and samples plausible object motion, given a single image in a single forward pass. This enables fast, camera‑guided video generation without the need for iterative denoising to inject object motion into render frames. Extensive experiments on the KITTI, Waymo, RealEstate10K and DL3DV‑10K datasets demonstrate that our method achieves state‑of‑the‑art video quality and inference efficiency. The project page is available at https://melonienimasha.github.io/Pixel‑to‑4D‑Website.
Authors:Zhaiyu Chen, Yuanyuan Wang, Yilei Shi, Xiao Xiang Zhu
Abstract:
Reliable building height estimation is essential for various urban applications. Spaceborne SAR tomography (TomoSAR) provides weather‑independent, side‑looking observations that capture facade‑level structure, offering a promising alternative to conventional optical methods. However, TomoSAR point clouds often suffer from noise, anisotropic point distributions, and data voids on incoherent surfaces, all of which hinder accurate height reconstruction. To address these challenges, we introduce a learning‑based framework for converting raw TomoSAR points into high‑resolution building height maps. Our dual‑topology network alternates between a point branch that models irregular scatterer features and a grid branch that enforces spatial consistency. By jointly processing these representations, the network denoises the input points and inpaints missing regions to produce continuous height estimates. To our knowledge, this is the first proof of concept for large‑scale urban height mapping directly from TomoSAR point clouds. Extensive experiments on data from Munich and Berlin validate the effectiveness of our approach. Moreover, we demonstrate that our framework can be extended to incorporate optical satellite imagery, further enhancing reconstruction quality. The source code is available at https://github.com/zhu‑xlab/tomosar2height.
Authors:Yiling Wang, Zeyu Zhang, Yiran Wang, Hao Tang
Abstract:
Text‑to‑motion (T2M) generation with diffusion backbones achieves strong realism and alignment. Safety concerns in T2M methods have been raised in recent years; existing methods replace discrete VQ‑VAE codebook entries to steer the model away from unsafe behaviors. However, discrete codebook replacement‑based methods have two critical flaws: firstly, replacing codebook entries which are reused by benign prompts leads to drifts on everyday tasks, degrading the model's benign performance; secondly, discrete token‑based methods introduce quantization and smoothness loss, resulting in artifacts and jerky transitions. Moreover, existing text‑to‑motion datasets naturally contain unsafe intents and corresponding motions, making them unsuitable for safety‑driven machine learning. To address these challenges, we propose SafeMo, a trustworthy motion generative framework integrating Minimal Motion Unlearning (MMU), a two‑stage machine unlearning strategy, enabling safe human motion generation in continuous space, preserving continuous kinematics without codebook loss and delivering strong safety‑utility trade‑offs compared to current baselines. Additionally, we present the first safe text‑to‑motion dataset SafeMoVAE‑29K integrating rewritten safe text prompts and continuous refined motion for trustworthy human motion unlearning. Built upon DiP, SafeMo efficiently generates safe human motions with natural transitions. Experiments demonstrate effective unlearning performance of SafeMo by showing strengthened forgetting on unsafe prompts, reaching 2.5x and 14.4x higher forget‑set FID on HumanML3D and Motion‑X respectively, compared to the previous SOTA human motion unlearning method LCR, with benign performance on safe prompts being better or comparable. Code: https://github.com/AIGeeksGroup/SafeMo. Website: https://aigeeksgroup.github.io/SafeMo.
Authors:Shuang Li, Yibing Wang, Jian Gao, Chulhong Kim, Seongwook Choi, Yu Zhang, Qian Chen, Yao Yao, Changhui Li
Abstract:
High‑quality three‑dimensional (3D) photoacoustic imaging (PAI) is gaining increasing attention in clinical applications. To address the challenges of limited space and high costs, irregular geometric transducer arrays that conform to specific imaging regions are promising for achieving high‑quality 3D PAI with fewer transducers. However, traditional iterative reconstruction algorithms struggle with irregular array configurations, suffering from high computational complexity, substantial memory requirements, and lengthy reconstruction times. In this work, we introduce SlingBAG Pro, an advanced reconstruction algorithm based on the point cloud iteration concept of the Sliding ball adaptive growth (SlingBAG) method, while extending its compatibility to arbitrary array geometries. SlingBAG Pro maintains high reconstruction quality, reduces the number of required transducers, and employs a hierarchical optimization strategy that combines zero‑gradient filtering with progressively increased temporal sampling rates during iteration. This strategy rapidly removes redundant spatial point clouds, accelerates convergence, and significantly shortens overall reconstruction time. Compared to the original SlingBAG algorithm, SlingBAG Pro achieves up to a 2.2‑fold speed improvement in point cloud‑based 3D PA reconstruction under irregular array geometries. The proposed method is validated through both simulation and in vivo mouse experiments, and the source code is publicly available at https://github.com/JaegerCQ/SlingBAG_Pro.
Authors:Guangqian Guo, Pengfei Chen, Yong Guo, Huafeng Chen, Boqiang Zhang, Shan Gao
Abstract:
Segment Anything Model (SAM), known for its remarkable zero‑shot segmentation capabilities, has garnered significant attention in the community. Nevertheless, its performance is challenged when dealing with what we refer to as visually non‑salient scenarios, where there is low contrast between the foreground and background. In these cases, existing methods often cannot capture accurate contours and fail to produce promising segmentation results. In this paper, we propose Visually Non‑Salient SAM (VNS‑SAM), aiming to enhance SAM's perception of visually non‑salient scenarios while preserving its original zero‑shot generalizability. We achieve this by effectively exploiting SAM's low‑level features through two designs: Mask‑Edge Token Interactive decoder and Non‑Salient Feature Mining module. These designs help the SAM decoder gain a deeper understanding of non‑salient characteristics with only marginal parameter increments and computational requirements. The additional parameters of VNS‑SAM can be optimized within 4 hours, demonstrating its feasibility and practicality. In terms of data, we established VNS‑SEG, a unified dataset for various VNS scenarios, with more than 35K images, in contrast to previous single‑task adaptations. It is designed to make the model learn more robust VNS features and comprehensively benchmark the model's segmentation performance and generalizability on VNS scenarios. Extensive experiments across various VNS segmentation tasks demonstrate the superior performance of VNS‑SAM, particularly under zero‑shot settings, highlighting its potential for broad real‑world applications. Codes and datasets are publicly available at https://guangqian‑guo.github.io/VNS‑SAM.
Authors:Wenrui Li, Hongtao Chen, Yao Xiao, Wangmeng Zuo, Jiantao Zhou, Yonghong Tian, Xiaopeng Fan
Abstract:
All‑in‑one image restoration aims to recover clean images from diverse unknown degradations using a single model. But extending this task to videos faces unique challenges. Existing approaches primarily focus on frame‑wise degradation variation, overlooking the temporal continuity that naturally exists in real‑world degradation processes. In practice, degradation types and intensities evolve smoothly over time, and multiple degradations may coexist or transition gradually. In this paper, we introduce the Smoothly Evolving Unknown Degradations (SEUD) scenario, where both the active degradation set and degradation intensity change continuously over time. To support this scenario, we design a flexible synthesis pipeline that generates temporally coherent videos with single, compound, and evolving degradations. To address the challenges in the SEUD scenario, we propose an all‑in‑One Recurrent Conditional and Adaptive prompting Network (ORCANet). First, a Coarse Intensity Estimation Dehazing (CIED) module estimates haze intensity using physical priors and provides coarse dehazed features as initialization. Second, a Flow Prompt Generation (FPG) module extracts degradation features. FPG generates both static prompts that capture segment‑level degradation types and dynamic prompts that adapt to frame‑level intensity variations. Furthermore, a label‑aware supervision mechanism improves the discriminability of static prompt representations under different degradations. Extensive experiments show that ORCANet achieves superior restoration quality, temporal consistency, and robustness over image and video‑based baselines. Code is available at https://github.com/Friskknight/ORCANet‑SEUD.
Authors:Miaowei Wang, Jakub Zadrożny, Oisin Mac Aodha, Amir Vaxman
Abstract:
Accurately simulating existing 3D objects and a wide variety of materials often demands expert knowledge and time‑consuming physical parameter tuning to achieve the desired dynamic behavior. We introduce MotionPhysics, an end‑to‑end differentiable framework that infers plausible physical parameters from a user‑provided natural language prompt for a chosen 3D scene of interest, removing the need for guidance from ground‑truth trajectories or annotated videos. Our approach first utilizes a multimodal large language model to estimate material parameter values, which are constrained to lie within plausible ranges. We further propose a learnable motion distillation loss that extracts robust motion priors from pretrained video diffusion models while minimizing appearance and geometry inductive biases to guide the simulation. We evaluate MotionPhysics across more than thirty scenarios, including real‑world, human‑designed, and AI‑generated 3D objects, spanning a wide range of materials such as elastic solids, metals, foams, sand, and both Newtonian and non‑Newtonian fluids. We demonstrate that MotionPhysics produces visually realistic dynamic simulations guided by natural language, surpassing the state of the art while automatically determining physically plausible parameters. The code and project page are available at: https://wangmiaowei.github.io/MotionPhysics.github.io/.
Authors:Shengjun Zhang, Zhang Zhang, Chensheng Dai, Yueqi Duan
Abstract:
Recent reinforcement learning has enhanced the flow matching models on human preference alignment. While stochastic sampling enables the exploration of denoising directions, existing methods which optimize over multiple denoising steps suffer from sparse and ambiguous reward signals. We observe that the high entropy steps enable more efficient and effective exploration while the low entropy steps result in undistinguished roll‑outs. To this end, we propose E‑GRPO, an entropy aware Group Relative Policy Optimization to increase the entropy of SDE sampling steps. Since the integration of stochastic differential equations suffer from ambiguous reward signals due to stochasticity from multiple steps, we specifically merge consecutive low entropy steps to formulate one high entropy step for SDE sampling, while applying ODE sampling on other steps. Building upon this, we introduce multi‑step group normalized advantage, which computes group‑relative advantages within samples sharing the same consolidated SDE denoising step. Experimental results on different reward settings have demonstrated the effectiveness of our methods.
Authors:Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu
Abstract:
Transformer residual streams evolve by additive accumulation: each layer appends a feature update to a shared hidden state, but has no direct mechanism for replacing content that has become obsolete or conflicting. We introduce Deep Delta Learning (DDL), a residual update rule that preserves the identity path while giving every layer the ability to selectively rewrite residual content. DDL reads the current state along a learned direction, compares it with a learned target value, and writes back a gated correction along the same direction. When the gate is closed, the update reduces to the identity; when the gate is fully open, the selected component is overwritten, yielding a depth‑wise delta‑rule generalization of standard residual addition. We integrate DDL in decoder‑only language models with both scalar and expanded residual states, while keeping attention and MLP sublayers at the original compute width. Controlled pretraining and downstream evaluations show that residual rewrite operations improve language modeling quality relative to pure additive accumulation introduced in ResNet, suggesting that a learned delta‑rule update is an effective mechanism for managing Transformer residual streams.
Authors:Tyler Ward, Abdullah Imran
Abstract:
Functional connectivity (FC) analysis, a valuable tool for computer‑aided brain disorder diagnosis, traditionally relies on atlas‑based parcellation. However, issues relating to selection bias and a lack of regard for subject specificity can arise as a result of such parcellations. Addressing this, we propose ABFR‑KAN, a transformer‑based classification network that incorporates novel advanced brain function representation components with the power of Kolmogorov‑Arnold Networks (KANs) to mitigate structural bias, improve anatomical conformity, and enhance the reliability of FC estimation. Extensive experiments on the ABIDE I dataset, including cross‑site evaluation and ablation studies across varying model backbones and KAN configurations, demonstrate that ABFR‑KAN consistently outperforms state‑of‑the‑art baselines for autism spectrum distorder (ASD) classification. Our code is available at https://github.com/tbwa233/ABFR‑KAN.
Authors:Tao Wu, Qing Xu, Xiangjian He, Oakleigh Weekes, James Brown, Wenting Duan
Abstract:
Roadside litter poses environmental, safety and economic challenges, yet current monitoring relies on labour‑intensive surveys and public reporting, providing limited spatial coverage. Existing vision datasets for litter detection focus on street‑level still images, aerial scenes or aquatic environments, and do not reflect the unique characteristics of dashcam footage, where litter appears extremely small, sparse and embedded in cluttered road‑verge backgrounds. We introduce RoLID‑11K, the first large‑scale dataset for roadside litter detection from dashcams, comprising over 11k annotated images spanning diverse UK driving conditions and exhibiting pronounced long‑tail and small‑object distributions. We benchmark a broad spectrum of modern detectors, from accuracy‑oriented transformer architectures to real‑time YOLO models, and analyse their strengths and limitations on this challenging task. Our results show that while CO‑DETR and related transformers achieve the best localisation accuracy, real‑time models remain constrained by coarse feature hierarchies. RoLID‑11K establishes a challenging benchmark for extreme small‑object detection in dynamic driving scenes and aims to support the development of scalable, low‑cost systems for roadside‑litter monitoring. The dataset is available at https://github.com/xq141839/RoLID‑11K.
Authors:Seungyeon Cho, Tae-kyun Kim
Abstract:
Skeleton‑based human action recognition (HAR) has achieved remarkable progress with graph‑based architectures. However, most existing methods remain body‑centric, focusing on large‑scale motions while neglecting subtle hand articulations that are crucial for fine‑grained recognition. This work presents a probabilistic dual‑stream framework that unifies reliability modeling and multi‑modal integration, generalizing expertized learning under uncertainty across both intra‑skeleton and cross‑modal domains. The framework comprises three key components: (1) a calibration‑free preprocessing pipeline that removes canonical‑space transformations and learns directly from native coordinates; (2) a probabilistic Noisy‑OR fusion that stabilizes reliability‑aware dual‑stream learning without requiring explicit confidence supervision; and (3) an intra‑ to cross‑modal ensemble that couples four skeleton modalities (Joint, Bone, Joint Motion, and Bone Motion) to RGB representations, bridging structural and visual motion cues in a unified cross‑modal formulation. Comprehensive evaluations across multiple benchmarks (NTU RGB+D~60/120, PKU‑MMD, N‑UCLA) and a newly defined hand‑centric benchmark exhibit consistent improvements and robustness under noisy and heterogeneous conditions.
Authors:Yingzhi Tang, Qijian Zhang, Junhui Hou
Abstract:
Achieving consistent and high‑fidelity geometry and appearance reconstruction of 3D digital humans from a single RGB image is inherently a challenging task. Existing studies typically resort to decoupled pipelines for geometry estimation and appearance synthesis, often hindering unified reconstruction and causing inconsistencies. This paper introduces JGA‑LBD, a novel framework that unifies the modeling of geometry and appearance into a joint latent representation and formulates the generation process as bridge diffusion. Observing that directly integrating heterogeneous input conditions (e.g., depth maps, SMPL models) leads to substantial training difficulties, we unify all conditions into the 3D Gaussian representations, which can be further compressed into a unified latent space through a shared sparse variational autoencoder (VAE). Subsequently, the specialized form of bridge diffusion enables to start with a partial observation of the target latent code and solely focuses on inferring the missing components. Finally, a dedicated decoding module extracts the complete 3D human geometric structure and renders novel views from the inferred latent representation. Experiments demonstrate that JGA‑LBD outperforms current state‑of‑the‑art approaches in terms of both geometry fidelity and appearance quality, including challenging in‑the‑wild scenarios. Our code will be made publicly available at https://github.com/haiantyz/JGA‑LBD.
Authors:Bryan Constantine Sadihin, Yihao Meng, Michael Hua Wang, Matteo Jiahao Chen, Hang Su
Abstract:
Most colorization models condition only on a single reference, typically the first frame of the scene. However, this approach ignores other sources of conditional data, such as character sheets, background images, or arbitrary colorized frames. We propose TimeColor, a sketch‑based video colorization model that supports heterogeneous, variable‑count references with the use of explicit per‑reference region assignment. TimeColor encodes references as additional latent frames which are concatenated temporally, permitting them to be processed concurrently in each diffusion step while keeping the model's parameter count fixed. TimeColor also uses spatiotemporal correspondence‑masked attention to enforce subject ‑‑ reference binding in addition to modality‑disjoint RoPE indexing. These mechanisms mitigate shortcutting and cross‑identity palette leakage. Experiments on Sakuga‑42M under both single‑ and multi‑reference protocols show that TimeColor improves color fidelity, identity consistency, and temporal stability over prior baselines. Our project page is available at https://bconstantine.github.io/TimeColor/.
Authors:Aobo Li, Jinjian Wu, Yongxu Liu, Leida Li, Weisheng Dong
Abstract:
Blind Image Quality Assessment (BIQA) has advanced significantly through deep learning, but the scarcity of large‑scale labeled datasets remains a challenge. While synthetic data offers a promising solution, models trained on existing synthetic datasets often show limited generalization ability. In this work, we make a key observation that representations learned from synthetic datasets often exhibit a discrete and clustered pattern that hinders regression performance: features of high‑quality images cluster around reference images, while those of low‑quality images cluster based on distortion types. Our analysis reveals that this issue stems from the distribution of synthetic data rather than model architecture. Consequently, we introduce a novel framework SynDR‑IQA, which reshapes synthetic data distribution to enhance BIQA generalization. Based on theoretical derivations of sample diversity and redundancy's impact on generalization error, SynDR‑IQA employs two strategies: distribution‑aware diverse content upsampling, which enhances visual diversity while preserving content distribution, and density‑aware redundant cluster downsampling, which balances samples by reducing the density of densely clustered areas. Extensive experiments across three cross‑dataset settings (synthetic‑to‑authentic, synthetic‑to‑algorithmic, and synthetic‑to‑synthetic) demonstrate the effectiveness of our method. The code is available at https://github.com/Li‑aobo/SynDR‑IQA.
Authors:Xiaokun Sun, Zeyu Cai, Hao Tang, Ying Tai, Jian Yang, Zhenyu Zhang
Abstract:
3D morphing remains challenging due to the difficulty of generating semantically consistent and temporally smooth deformations, especially across categories. We present MorphAny3D, a training‑free framework that leverages Structured Latent (SLAT) representations for high‑quality 3D morphing. Our key insight is that intelligently blending source and target SLAT features within the attention mechanisms of 3D generators naturally produces plausible morphing sequences. To this end, we introduce Morphing Cross‑Attention (MCA), which fuses source and target information for structural coherence, and Temporal‑Fused Self‑Attention (TFSA), which enhances temporal consistency by incorporating features from preceding frames. An orientation correction strategy further mitigates the pose ambiguity within the morphing steps. Extensive experiments show that our method generates state‑of‑the‑art morphing sequences, even for challenging cross‑category cases. MorphAny3D further supports advanced applications such as decoupled morphing and 3D style transfer, and can be generalized to other SLAT‑based generative models. Project page: https://xiaokunsun.github.io/MorphAny3D.github.io/.
Authors:Brady Zhou, Philipp Krähenbühl
Abstract:
Human drivers rarely travel where no person has gone before. After all, thousands of drivers use busy city roads every day, and only one can claim to be the first. The same holds for autonomous computer vision systems. The vast majority of the deployment area of an autonomous vision system will have been visited before. Yet, most autonomous vehicle vision systems act as if they are encountering each location for the first time. In this work, we present Compressed Map Priors (CMP), a simple but effective framework to learn spatial priors from historic traversals. The map priors use a binarized hashmap that requires only 32\textKB/\textkm^2, a 20× reduction compared to the dense storage. Compressed Map Priors easily integrate into leading 3D perception systems at little to no extra computational costs, and lead to a significant and consistent improvement in 3D object detection on the nuScenes dataset across several architectures.